高级端侧 / 本地 12 分钟阅读

将自己的模型量化为 GGUF

用 llama.cpp 的 quantize 工具将任意 HF 模型转为 GGUF Q4_K_M 本地推理。

已验证技术栈

llama.cpp convert_hf_to_gguf.py · llama-quantize · 目标 Q4_K_M

最近验证 2026-07-22

GGUFllama.cppquantizecustom