Quantize
Everything.
Bridge the gap between research papers and real-world deployment. Run state-of-the-art LLMs on consumer hardware.
Start here
Data updated 2026-08-08
This week’s updates
New models, recency tags, and data cadence2026-08-08
- NewQwen3 4B Instruct4B · 2026-06-26
- NewQwen3-Coder 30B-A3B Instruct30B-A3B · 2026-06-26
- NewMistral Large 3 675B Instruct675B MoE · 2026-06-26
- NewGLM-4-9B-Chat9B · 2026-06-26
- NewGemma 3 27B IT27B · 2026-07-22
- NewDeepSeek-R1-Distill-Llama-8B8B · 2026-07-22
- 2026-08-08New cookbook: running GPT-OSS 20B/120B locally without re-quantizing (23 guides). VRAM calculator now offers MXFP4 — picking Q4_K_M for GPT-OSS overstated weights by ~14%
- 2026-08-07Model index +4 → 75: GPT-OSS 20B/120B (native MXFP4), GLM-4.5-Air 106B-A12B, Devstral Small 1.1 — MoE-heavy batch for 16GB cards and unified-memory Macs
- 2026-07-22Polish: re-rendered og.png (71+ models), all 22 cookbook guides have verified stack banners
Editor's Picks
Curated standout quant releases — click to view model details
- NEWGPT-OSS 20BGGUFMXFP4 · 12.8 GB · native 4-bit·openaiRTX 4070 Ti 16G
- NEWDevstral Small 1.1 24BGGUFQ4_K_M · 14.3 GB · agentic coder·unslothRTX 4080 16G
- NEWGPT-OSS 120BGGUFMXFP4 · 61 GB · 5.1B active·openaiH100 80G / M3 Ultra
- NEWGLM-4.5-AirGGUFQ4_K_M · 64 GB · 12B active·unsloth2× 48GB
- HOTQwen3-Coder 30B-A3B InstructGGUFQ4_K_M · 19 GB · agentic coder·bartowskiRTX 4090
- HOTQwen3 30B-A3B InstructGGUFQ4_K_M · 19 GB · 3B active·bartowskiRTX 4060 Ti 16G
Format Heat Index
Editorial adoption estimate
- 1GGUF89%
- 2AWQ45%
- 3EXL232%
- 4GPTQ28%
- 5HQQ18%
Editorial estimate from HF GGUF share and community discussion volume — not live analytics
Quick Tools
VRAM Calculator
Precise memory requirements for any model × quant × context combination. Red/yellow/green hardware verdict.
CLI Generator
Generate ready-to-run llama.cpp, Ollama, vLLM, ExLlamaV2 commands. One-liner or Docker Compose.
Format Wizard
Answer 3 questions — get a personalised GGUF, AWQ, or EXL2 recommendation.
Model Compare
Side-by-side VRAM, speed, and quality for two models — find the best fit for your GPU.
Explore the Site
Data Changelog
Last updated 2026-08-08
- 2026-08-08New cookbook: running GPT-OSS 20B/120B locally without re-quantizing (23 guides). VRAM calculator now offers MXFP4 — picking Q4_K_M for GPT-OSS overstated weights by ~14%
- 2026-08-07Model index +4 → 75: GPT-OSS 20B/120B (native MXFP4), GLM-4.5-Air 106B-A12B, Devstral Small 1.1 — MoE-heavy batch for 16GB cards and unified-memory Macs
- 2026-07-22Polish: re-rendered og.png (71+ models), all 22 cookbook guides have verified stack banners
- 2026-07-22Cadence pack: +4 models (Gemma 3 27B, R1-Llama-8B, Phi-4, Qwen3 1.7B), superseded tags, measured/estimated labels, Hub “recent”, weekly updates, RSS, cookbook verified stack (71 models)
- 2026-06-26UX for real traffic: job paths, mobile GPU profile, OG/favicon, honest format heat, feedback email
- 2026-06-26Model index +4: Qwen3 4B, Qwen3-Coder 30B-A3B, Mistral Large 3, GLM-4-9B (67 total)
- 2026-06-26QA fixes: HF stats merge on failure, ≤3B filter, CLI/VRAM tool bugs, i18n polish
- 2026-06-26Model index +5: Qwen3 32B, 30B-A3B MoE, 235B-A22B, DeepSeek-V3, DeepSeek-R1 (63 total)
- 2026-06-26Model index +7: Qwen3 8B/14B, Gemma 3 4B/12B, Llama 4 Scout/Maverick, Llama 3.1 405B (58 total)
- 2026-06-25Cookbook TOC scroll highlight, code block copy, Quant Hub Markdown export
- 2026-06-25Cookbook reading progress bar, model HF link copy, Quant Hub shareable filter URLs
- 2026-06-25Breadcrumb nav + JSON-LD, cookbook article TOC, Quant Hub GPU quick-filter chips
- 2026-06-25Related cookbook guides, similar-model cards on detail pages, hero latest-update badge
- 2026-06-25Homepage explore strip, 404 page, multi-model benchmarks (Qwen 7B/32B, DeepSeek-R1 14B), llms.txt for AI crawlers
- 2026-06-24Added About page (/about) — maintainer story, update cadence, contribution guide
- 2026-06-24Quant Hub: default to all 51 models visible; scale stats bar; clearer GPU filter UX
- 2026-06-24SEO: canonical URLs, JSON-LD, per-page metadata, Google/Bing verification env vars
- 2026-06-24Quant Hub: show-all toggle when GPU profile active; Cookbook +7 guides (8GB GPU, WSL2, Docker GPU, Nginx, AMD ROCm)
- 2026-06-24Privacy Policy + Plausible analytics, cookbook standalone pages (/cookbook/[slug]), model index expanded to 51
- 2026-06-24Added Terms & Disclaimer page (/legal) with trademark notice and liability disclaimer
- 2026-06-24Phase 3: HF live stats pipeline, model A vs B compare tool, cookbook expanded to 15 guides
- 2026-06-24Expanded model index to 30+ entries; added format wizard, hardware profile, ExLlamaV2 CLI, SEO (sitemap/OG), data transparency
- 2026-06-24Model detail pages, GPU reverse lookup, shareable VRAM calculator URLs, real homepage stats
- 2025-06-10Initial launch: Quant Hub, VRAM calculator, CLI generator, benchmarks, cookbook
All data is manually curated and verified against community sources. Always cross-check with official Hugging Face model cards before deployment.