OLMo 2 7B Instruct

7B

Allen AI OLMo

Fully open training pipeline from Allen AI. Great for reproducibility research.

2.5K HF downloads12 likesbartowski/OLMo-2-1124-7B-Instruct-GGUF· stats from 9/21/2026
Consumer GPUMac / Apple SiliconCPU / VPS

4K

Max Context

2

Quant Variants

GGUF Q8_0

Best Quality

99.7%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.855.3 GB3.4%150 tok/sEstimated
CalcHF
GGUFQ8_08.58.2 GB0.3%125 tok/sEstimated
CalcHF

Running OLMo 2 7B Instruct locally

At Q4_K_M and 4K of context, OLMo 2 7B Instruct needs about 6.8 GB — 4.2 GB of weights, 2.00 GB of KV cache and a 0.6 GB activation buffer. The smallest card in this index that clears that comfortably is the RTX 5060 Ti 8G at 8 GB, and 62 of the 63 cards here do. These are calculated figures, not measurements: the estimate stops counting a card as comfortable at 88% of its VRAM, which is roughly the room a desktop session needs.

What longer context costs

Going from 4K to 32K adds about 14.00 GB, taking the total to 22.2 GB. Weights do not move with context — only the KV cache does, and it grows linearly, so this is the number to watch when planning for long documents. The model's native window is 4K; holding all of it at Q4_K_M would need about 7 GB.

Which build to download

This index only tracks OLMo 2 7B Instruct in GGUF, across 2 levels (Q4_K_M, Q8_0). Q8_0 carries the lowest published perplexity loss at 0.3%. The fastest level measured here is Q4_K_M at 150 tok/s on an RTX 4090, batch 1. GGUF runs on llama.cpp and Ollama across NVIDIA, AMD and Apple silicon; AWQ and GPTQ target vLLM on CUDA and ROCm; EXL2 is ExLlamaV2 and CUDA only.

Common questions

How much VRAM does OLMo 2 7B Instruct need?
About 6.8 GB at Q4_K_M with 4K of context and batch 1, rising to roughly 22.2 GB at 32K. That figure is weights plus KV cache plus a 10% activation buffer, calculated from the model's architecture rather than measured on a card.
Will OLMo 2 7B Instruct run on a 8GB GPU?
Yes — at Q4_K_M and 4K context it needs about 6.8 GB, which leaves 1.2 GB spare on a RTX 5060 Ti 8G. That is the smallest card in this index that clears it comfortably; 62 of 63 do.
Which quantization of OLMo 2 7B Instruct should I use?
Q8_0 has the lowest published quality loss (0.3%), and Q4_K_M is the level most people run. All 2 levels in the index are Q4_K_M, Q8_0.

Where this model fits

Sized at Q4_K_M with a 4K context window, smallest card first. Comfortable means the estimate uses at most 88% of the memory.

Ships in
GGUF

This page's figures change when the model or the runtime does.Last updated 2026-09-21 RSS → /feed.xml