Quantize 
Everything.

Bridge the gap between research papers and real-world deployment. Run state-of-the-art LLMs on consumer hardware.

GGUFAWQEXL2GPTQHQQ& more

Start here

75Models Indexed
5Formats Tracked
33GPUs in Database
98.4%Avg Accuracy Retained

Data updated 2026-08-08

This week’s updates

New models, recency tags, and data cadence2026-08-08

  • 2026-08-08New cookbook: running GPT-OSS 20B/120B locally without re-quantizing (23 guides). VRAM calculator now offers MXFP4 — picking Q4_K_M for GPT-OSS overstated weights by ~14%
  • 2026-08-07Model index +4 → 75: GPT-OSS 20B/120B (native MXFP4), GLM-4.5-Air 106B-A12B, Devstral Small 1.1 — MoE-heavy batch for 16GB cards and unified-memory Macs
  • 2026-07-22Polish: re-rendered og.png (71+ models), all 22 cookbook guides have verified stack banners
Full changelog

Explore the Site

Data Changelog

Last updated 2026-08-08

  • 2026-08-08New cookbook: running GPT-OSS 20B/120B locally without re-quantizing (23 guides). VRAM calculator now offers MXFP4 — picking Q4_K_M for GPT-OSS overstated weights by ~14%
  • 2026-08-07Model index +4 → 75: GPT-OSS 20B/120B (native MXFP4), GLM-4.5-Air 106B-A12B, Devstral Small 1.1 — MoE-heavy batch for 16GB cards and unified-memory Macs
  • 2026-07-22Polish: re-rendered og.png (71+ models), all 22 cookbook guides have verified stack banners
  • 2026-07-22Cadence pack: +4 models (Gemma 3 27B, R1-Llama-8B, Phi-4, Qwen3 1.7B), superseded tags, measured/estimated labels, Hub “recent”, weekly updates, RSS, cookbook verified stack (71 models)
  • 2026-06-26UX for real traffic: job paths, mobile GPU profile, OG/favicon, honest format heat, feedback email
  • 2026-06-26Model index +4: Qwen3 4B, Qwen3-Coder 30B-A3B, Mistral Large 3, GLM-4-9B (67 total)
  • 2026-06-26QA fixes: HF stats merge on failure, ≤3B filter, CLI/VRAM tool bugs, i18n polish
  • 2026-06-26Model index +5: Qwen3 32B, 30B-A3B MoE, 235B-A22B, DeepSeek-V3, DeepSeek-R1 (63 total)
  • 2026-06-26Model index +7: Qwen3 8B/14B, Gemma 3 4B/12B, Llama 4 Scout/Maverick, Llama 3.1 405B (58 total)
  • 2026-06-25Cookbook TOC scroll highlight, code block copy, Quant Hub Markdown export
  • 2026-06-25Cookbook reading progress bar, model HF link copy, Quant Hub shareable filter URLs
  • 2026-06-25Breadcrumb nav + JSON-LD, cookbook article TOC, Quant Hub GPU quick-filter chips
  • 2026-06-25Related cookbook guides, similar-model cards on detail pages, hero latest-update badge
  • 2026-06-25Homepage explore strip, 404 page, multi-model benchmarks (Qwen 7B/32B, DeepSeek-R1 14B), llms.txt for AI crawlers
  • 2026-06-24Added About page (/about) — maintainer story, update cadence, contribution guide
  • 2026-06-24Quant Hub: default to all 51 models visible; scale stats bar; clearer GPU filter UX
  • 2026-06-24SEO: canonical URLs, JSON-LD, per-page metadata, Google/Bing verification env vars
  • 2026-06-24Quant Hub: show-all toggle when GPU profile active; Cookbook +7 guides (8GB GPU, WSL2, Docker GPU, Nginx, AMD ROCm)
  • 2026-06-24Privacy Policy + Plausible analytics, cookbook standalone pages (/cookbook/[slug]), model index expanded to 51
  • 2026-06-24Added Terms & Disclaimer page (/legal) with trademark notice and liability disclaimer
  • 2026-06-24Phase 3: HF live stats pipeline, model A vs B compare tool, cookbook expanded to 15 guides
  • 2026-06-24Expanded model index to 30+ entries; added format wizard, hardware profile, ExLlamaV2 CLI, SEO (sitemap/OG), data transparency
  • 2026-06-24Model detail pages, GPU reverse lookup, shareable VRAM calculator URLs, real homepage stats
  • 2025-06-10Initial launch: Quant Hub, VRAM calculator, CLI generator, benchmarks, cookbook

All data is manually curated and verified against community sources. Always cross-check with official Hugging Face model cards before deployment.