<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>quantized.uk — 更新</title>
    <link>https://quantized.uk/zh/</link>
    <description>LLM 量化情报：新模型、数据节奏与站点更新。</description>
    <language>zh-Hans</language>
    <lastBuildDate>Tue, 01 Sep 2026 12:00:00 GMT</lastBuildDate>
    <atom:link href="https://quantized.uk/zh/feed.xml" rel="self" type="application/rss+xml"/>
    
    <item>
      <title>新增：格式对比页。GGUF 与 AWQ、GGUF 与 EXL2 等六组，对比硬件支持、运行时与采用率，并列出同时发布两种格式权重的模型 —— 此时两行描述的是同一个模型，对比不再只是编辑判断。若没有模型同时提供两种格式，页面会如实说明，而不是含糊带过</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-New: format comparison pages. GGUF vs AW</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>新增：格式对比页。GGUF 与 AWQ、GGUF 与 EXL2 等六组，对比硬件支持、运行时与采用率，并列出同时发布两种格式权重的模型 —— 此时两行描述的是同一个模型，对比不再只是编辑判断。若没有模型同时提供两种格式，页面会如实说明，而不是含糊带过</description>
    </item>

    <item>
      <title>计算器改为给结论，而不是罗列。硬件面板此前平铺 43 根柱子且大多为绿色 —— 占了大量篇幅却几乎没有回答问题。现在它先给出你真正想要的事实（&quot;43 张卡中 16 张可从容运行 —— 最小的是 Instinct MI100 32G&quot;），如果你设置了自己的显卡还会单独给出裁决，完整列表则收在一次点击之后。三个工具里 79 项的模型下拉也改为按尺寸分组，不再按录入顺序排列</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-The calculator answers instead of listin</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>计算器改为给结论，而不是罗列。硬件面板此前平铺 43 根柱子且大多为绿色 —— 占了大量篇幅却几乎没有回答问题。现在它先给出你真正想要的事实（&quot;43 张卡中 16 张可从容运行 —— 最小的是 Instinct MI100 32G&quot;），如果你设置了自己的显卡还会单独给出裁决，完整列表则收在一次点击之后。三个工具里 79 项的模型下拉也改为按尺寸分组，不再按录入顺序排列</description>
    </item>

    <item>
      <title>新增：每张显卡一个页面。数据库中全部 43 张卡都已覆盖 —— &quot;RTX 4060 Ti 16G 能跑哪些大模型？&quot;给出 79 个索引模型中的 51 个，各自取 4K 上下文下仍留有余量的最佳量化档位，并按尺寸分组。这些数据一直都在，只是此前仅作为筛选参数存在，从未成为页面。中英合计新增 88 个页面</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-New: a page per GPU. All 43 cards in the</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>新增：每张显卡一个页面。数据库中全部 43 张卡都已覆盖 —— &quot;RTX 4060 Ti 16G 能跑哪些大模型？&quot;给出 79 个索引模型中的 51 个，各自取 4K 上下文下仍留有余量的最佳量化档位，并按尺寸分组。这些数据一直都在，只是此前仅作为筛选参数存在，从未成为页面。中英合计新增 88 个页面</description>
    </item>

    <item>
      <title>四个工具页现在会自我说明。此前每页只有一个标题、没有任何正文 —— 一个孤立的交互控件，读者无从校准。现在它们分别说明显存估算如何得出、为何上下文长度主导结果、CLI 命令里的标识符从哪来、以及为什么有些命令是占位符，并各配一组常见问题，中英双语</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-The four tool pages now explain themselv</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>四个工具页现在会自我说明。此前每页只有一个标题、没有任何正文 —— 一个孤立的交互控件，读者无从校准。现在它们分别说明显存估算如何得出、为何上下文长度主导结果、CLI 命令里的标识符从哪来、以及为什么有些命令是占位符，并各配一组常见问题，中英双语</description>
    </item>

    <item>
      <title>推理速度图恢复可读：18 根柱子里有 13 根标签都是 &quot;RTX 4090&quot;，屏幕上没有任何信息表明每根对应哪个模型。现在每根柱子都标注模型与硬件，并新增图例说明颜色含义（颜色一直代表推理框架，但此前从未说明）。手机端点击目标达到 44px 下限</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-The inference speed chart is readable ag</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>推理速度图恢复可读：18 根柱子里有 13 根标签都是 &quot;RTX 4090&quot;，屏幕上没有任何信息表明每根对应哪个模型。现在每根柱子都标注模型与硬件，并新增图例说明颜色含义（颜色一直代表推理框架，但此前从未说明）。手机端点击目标达到 44px 下限</description>
    </item>

    <item>
      <title>首页的质量数字现在会说明自己描述的是哪一档：Q4_K_M 下保留 97.1%，取全部 79 个模型的中位数，并在旁标注 1.4–5.2% 的区间。旧的&quot;98.4% 平均精度&quot;是对每个模型各自最好的那一档求平均，因此往数据里补一个档位就会让它变动。中文版基准表格也已完整翻译 —— 此前备注列整列是英文</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-The homepage quality figure now names th</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>首页的质量数字现在会说明自己描述的是哪一档：Q4_K_M 下保留 97.1%，取全部 79 个模型的中位数，并在旁标注 1.4–5.2% 的区间。旧的&quot;98.4% 平均精度&quot;是对每个模型各自最好的那一档求平均，因此往数据里补一个档位就会让它变动。中文版基准表格也已完整翻译 —— 此前备注列整列是英文</description>
    </item>

    <item>
      <title>模型索引进入 HTML。此前 /quant-hub/ 完全由浏览器渲染，静态页面里没有任何标题、也没有 79 个模型名 —— 全站最核心的页面对搜索引擎和无 JavaScript 环境等于空白。现在 79 张卡片直接写入 HTML，筛选在其之上叠加</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-The model index is now in the HTML. /qua</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>模型索引进入 HTML。此前 /quant-hub/ 完全由浏览器渲染，静态页面里没有任何标题、也没有 79 个模型名 —— 全站最核心的页面对搜索引擎和无 JavaScript 环境等于空白。现在 79 张卡片直接写入 HTML，筛选在其之上叠加</description>
    </item>

    <item>
      <title>显存计算器不再回答你没问的问题：此前未选模型时会回落到通用 7B，给出 4.98 GB 并对全部 43 张显卡打绿灯。中文读者现在有独立的 /zh/feed.xml 订阅源，每个页面都声明了自己的 feed，/rss.xml 也不再 404</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-VRAM calculator no longer answers a ques</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>显存计算器不再回答你没问的问题：此前未选模型时会回落到通用 7B，给出 4.98 GB 并对全部 43 张显卡打绿灯。中文读者现在有独立的 /zh/feed.xml 订阅源，每个页面都声明了自己的 feed，/rss.xml 也不再 404</description>
    </item>

    <item>
      <title>首页无需等待 JavaScript 即可绘制：此前标题以 opacity:0 发出，要等 bundle 加载后才可见，导致预渲染站点的真实用户 LCP P75 达 3.1 秒。入场动画改为纯 CSS，framer-motion 已移除 —— 首屏 JavaScript 减少 33 kB</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-09-01-Homepage paints without waiting for Java</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>首页无需等待 JavaScript 即可绘制：此前标题以 opacity:0 发出，要等 bundle 加载后才可见，导致预渲染站点的真实用户 LCP P75 达 3.1 秒。入场动画改为纯 CSS，framer-motion 已移除 —— 首屏 JavaScript 减少 33 kB</description>
    </item>

    <item>
      <title>格式相关的展示与索引对齐：首页徽章、&quot;格式追踪&quot;统计与 Hub 筛选均改为从实际有模型的格式推导（4 种）。HQQ 仍保留在格式科普中作为参考，但不再引导到没有结果的浏览路径</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-08-23-Format surfaces now agree with the index</guid>
      <pubDate>Sun, 23 Aug 2026 12:00:00 GMT</pubDate>
      <description>格式相关的展示与索引对齐：首页徽章、&quot;格式追踪&quot;统计与 Hub 筛选均改为从实际有模型的格式推导（4 种）。HQQ 仍保留在格式科普中作为参考，但不再引导到没有结果的浏览路径</description>
    </item>

    <item>
      <title>格式向导：量化档位改为按格式区分 —— 此前选择&quot;最易上手&quot;会给出 EXL2 · Q4_K_M、AWQ · Q4_K_M 这类两种格式里都不存在的档位。框架推荐也改为跟随格式，AMD 读者不会再看到 EXL2 配 ROCm 运行时，却在旁边读到&quot;EXL2 无法在 ROCm 上运行&quot;</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-08-23-Format wizard: quant levels are now per-</guid>
      <pubDate>Sun, 23 Aug 2026 12:00:00 GMT</pubDate>
      <description>格式向导：量化档位改为按格式区分 —— 此前选择&quot;最易上手&quot;会给出 EXL2 · Q4_K_M、AWQ · Q4_K_M 这类两种格式里都不存在的档位。框架推荐也改为跟随格式，AMD 读者不会再看到 EXL2 配 ROCm 运行时，却在旁边读到&quot;EXL2 无法在 ROCm 上运行&quot;</description>
    </item>

    <item>
      <title>首页排版：统计条自 job path 卡片引入后就一直盖在卡片上，遮住全部三行说明；Hero 在手机上被裁切，显存计算器按钮与格式徽章有一部分点不到。导航栏在平板宽度下不再溢出</title>
      <link>https://quantized.uk/zh/#changelog</link>
      <guid isPermaLink="false">changelog-zh-2026-08-23-Homepage layout: the stats bar had been </guid>
      <pubDate>Sun, 23 Aug 2026 12:00:00 GMT</pubDate>
      <description>首页排版：统计条自 job path 卡片引入后就一直盖在卡片上，遮住全部三行说明；Hero 在手机上被裁切，显存计算器按钮与格式徽章有一部分点不到。导航栏在平板宽度下不再溢出</description>
    </item>

    <item>
      <title>模型：Qwen3 4B Instruct</title>
      <link>https://quantized.uk/zh/quant-hub/qwen3-4b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/qwen3-4b/</guid>
      <pubDate>Fri, 26 Jun 2026 12:00:00 GMT</pubDate>
      <description>最小 Qwen3 稠密模型，支持思考模式。Q4 约 3.2GB，适合 8GB 显卡与边缘设备。</description>
    </item>

    <item>
      <title>模型：Qwen3-Coder 30B-A3B Instruct</title>
      <link>https://quantized.uk/zh/quant-hub/qwen3-coder-30b-a3b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/qwen3-coder-30b-a3b/</guid>
      <pubDate>Fri, 26 Jun 2026 12:00:00 GMT</pubDate>
      <description>Agentic 代码 MoE，3.3B 激活参数，原生 256K 上下文。16–24GB 显卡上最强开源代码模型之一。</description>
    </item>

    <item>
      <title>模型：Mistral Large 3 675B Instruct</title>
      <link>https://quantized.uk/zh/quant-hub/mistral-large-3/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/mistral-large-3/</guid>
      <pubDate>Fri, 26 Jun 2026 12:00:00 GMT</pubDate>
      <description>Mistral 3 旗舰 MoE（激活 41B / 总 675B），含视觉编码器。FP8 需 8×H200；GGUF 量化仅供研究集群。</description>
    </item>

    <item>
      <title>模型：GLM-4-9B-Chat</title>
      <link>https://quantized.uk/zh/quant-hub/glm-4-9b-chat/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/glm-4-9b-chat/</guid>
      <pubDate>Fri, 26 Jun 2026 12:00:00 GMT</pubDate>
      <description>智谱 GLM-4 开源 9B，128K 上下文，支持工具调用，中英双语表现突出。</description>
    </item>

    <item>
      <title>模型：Gemma 3 27B IT</title>
      <link>https://quantized.uk/zh/quant-hub/gemma-3-27b-it/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/gemma-3-27b-it/</guid>
      <pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate>
      <description>Gemma 3 大杯指令版，长上下文 + 多模态。Q4 约 16GB，适合 24GB 卡或双卡短上下文。</description>
    </item>

    <item>
      <title>模型：DeepSeek-R1-Distill-Llama-8B</title>
      <link>https://quantized.uk/zh/quant-hub/deepseek-r1-distill-llama-8b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/deepseek-r1-distill-llama-8b/</guid>
      <pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate>
      <description>R1 推理蒸馏到 Llama 3.1 8B。8–12GB 显卡上最佳思维链之一，社区 GGUF 极丰富。</description>
    </item>

    <item>
      <title>模型：Phi-4 14B</title>
      <link>https://quantized.uk/zh/quant-hub/phi-4/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/phi-4/</guid>
      <pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate>
      <description>微软 Phi-4 稠密 14B，同尺寸推理能力突出。Q4 约 9GB，12GB 卡中等上下文可跑。</description>
    </item>

    <item>
      <title>模型：Qwen3 1.7B Instruct</title>
      <link>https://quantized.uk/zh/quant-hub/qwen3-1.7b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/qwen3-1.7b/</guid>
      <pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate>
      <description>迷你 Qwen3，支持思考模式。Q4 约 1.4GB，适合手机旁路、NUC 与常驻本地 Agent。</description>
    </item>

    <item>
      <title>模型：GPT-OSS 20B</title>
      <link>https://quantized.uk/zh/quant-hub/gpt-oss-20b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/gpt-oss-20b/</guid>
      <pubDate>Fri, 07 Aug 2026 12:00:00 GMT</pubDate>
      <description>OpenAI 开放权重 MoE（总 21B / 激活 3.6B），原生以 MXFP4 发布 —— 约 12.8GB，16GB 显卡即可跑且无额外精度损失。激活仅 3.6B，CPU 卸载也还能用。</description>
    </item>

    <item>
      <title>模型：GPT-OSS 120B</title>
      <link>https://quantized.uk/zh/quant-hub/gpt-oss-120b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/gpt-oss-120b/</guid>
      <pubDate>Fri, 07 Aug 2026 12:00:00 GMT</pubDate>
      <description>GPT-OSS 大杯（总 117B / 激活 5.1B）。原生 MXFP4 权重约 61GB —— 单张 80GB 卡或 128GB 统一内存 Mac 可跑。24GB 消费卡部分卸载虽慢但可用。</description>
    </item>

    <item>
      <title>模型：GLM-4.5-Air</title>
      <link>https://quantized.uk/zh/quant-hub/glm-4.5-air/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/glm-4.5-air/</guid>
      <pubDate>Fri, 07 Aug 2026 12:00:00 GMT</pubDate>
      <description>智谱 Agent/推理向 MoE（总 106B / 激活 12B）。Q4 约 64GB —— 96GB+ 统一内存 Mac 或双 48GB 卡最合适。同级别中工具调用能力突出。</description>
    </item>

    <item>
      <title>模型：Devstral Small 1.1 24B</title>
      <link>https://quantized.uk/zh/quant-hub/devstral-small-2507/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/devstral-small-2507/</guid>
      <pubDate>Fri, 07 Aug 2026 12:00:00 GMT</pubDate>
      <description>Mistral 与 All Hands 合作的 Agent 编码模型，基于 Mistral Small 3.1。面向仓库级工具调用而非单文件补全。Q4 约 14GB，16GB 显卡可跑。</description>
    </item>

    <item>
      <title>模型：Qwen3-VL 8B Instruct</title>
      <link>https://quantized.uk/zh/quant-hub/qwen3-vl-8b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/qwen3-vl-8b/</guid>
      <pubDate>Sat, 08 Aug 2026 12:00:00 GMT</pubDate>
      <description>当代视觉语言模型,Q4 约 5.9GB,单张 8–12GB 卡仍可跑。没有 24GB 显卡的人做多模态的现实选择 —— 注意视觉编码器会额外占用显存,下方 KV cache 计算并未包含这部分。</description>
    </item>

    <item>
      <title>模型：Qwen3-VL 30B-A3B Instruct</title>
      <link>https://quantized.uk/zh/quant-hub/qwen3-vl-30b-a3b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/qwen3-vl-30b-a3b/</guid>
      <pubDate>Sat, 08 Aug 2026 12:00:00 GMT</pubDate>
      <description>多模态 MoE,激活参数仅约 3B,因此在苹果统一内存上依然跟手,CPU 卸载的表现也远好于稠密 30B。Q4 约 19GB,24GB 显卡可完整放下。</description>
    </item>

    <item>
      <title>模型：Magistral Small 1.2 24B</title>
      <link>https://quantized.uk/zh/quant-hub/magistral-small-2509/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/magistral-small-2509/</guid>
      <pubDate>Sat, 08 Aug 2026 12:00:00 GMT</pubDate>
      <description>Mistral 基于 Mistral Small 3.2 的推理模型,思维链包在 [THINK] 标签内。Q4 约 14GB,让显式思维链落到 16GB 显卡上 —— 比多数教程默认的 70B 级推理模型低一档。</description>
    </item>

    <item>
      <title>模型：Seed-OSS 36B Instruct</title>
      <link>https://quantized.uk/zh/quant-hub/seed-oss-36b/</link>
      <guid isPermaLink="true">https://quantized.uk/zh/quant-hub/seed-oss-36b/</guid>
      <pubDate>Sat, 08 Aug 2026 12:00:00 GMT</pubDate>
      <description>稠密 36B,原生 512K 上下文。Q4 权重约 22GB,单张 24GB 卡或 2×16GB 拆分即可 —— 但真正的开销是上下文:512K 的 KV cache 单独就约 128GB,所以要先规划上下文再考虑权重。</description>
    </item>
  </channel>
</rss>