Tools
Four calculators and generators, plus the two indexes behind them. Every figure they produce is computed from the model index on this site, not typed in by hand.
- VRAM Calc
Estimate VRAM for any quantized LLM: model weights, KV cache and activation buffer, with a per-GPU verdict across every card in the database.
- CLI Gen
Generate runnable llama.cpp, Ollama, vLLM and ExLlamaV2 commands for any indexed model, including Docker and Compose variants.
- Format Wizard
Answer three questions about your hardware and priorities to get a quantization format, runtime and quant level that actually run together.
- Compare
Put two quantized models side by side: parameters, context, VRAM at your chosen context length, speed and quality loss.
Browse by hardware or format
- GPU pages
One page per card — 61 of them — listing which of the 81 indexed models fit it comfortably at 4K context, at what quant level, with how much headroom.
- Format comparisons
4 head-to-head pages — GGUF vs AWQ, GGUF vs EXL2 and the rest — on hardware support, runtime and the models this index ships in both formats.