Skip to content

Quantization Formats — Which to Use When

You pay roughly 40% throughput for GGUF's portability. NF4 is for training (QLoRA), not serving — it was the slowest option measured and doubles disk usage.

Head-to-head RTX 3090 measurements, Llama-2-13B, sorted by perplexity (oobabooga.github.io):

Format Perplexity (wikitext) VRAM (GB) Size (GB) Prompt proc. (3200 tok) tok/s
EXL2 4.900b 4.30752 9.31 7.86 1.76 s 52.1
EXL2 4.650b 4.32136 9.03 7.48 1.74 s 56.5
AWQ 4bit-32g 4.32522 10.57 7.62 3.60 s 39.5
GGUF Q4_K_M 4.33326 8.99 7.50 3.73 s 30.8
GPTQ 4bit-32g 4.33805 8.70 7.63 1.86 s 42.4
Q4_K_S 4.34246 8.55 7.07 3.68 s 35.3
AWQ 4bit-128g 4.34761 9.62 6.92 3.59 s 40.6
GPTQ 4bit-128g 4.35793 7.94 6.92 1.85 s 51.9
bnb load_in_4bit 4.36427 8.19 24.83 3.01 s 23.1
EXL2 4.000b 4.37648 7.88 6.50 1.71 s 56.8

Reading of that table, which is the most useful comparative data available:

  • GGUF Q4_K_M — the default for local/desktop use. Best portability (llama.cpp, Ollama, LM Studio, CPU/AMD/Apple/Vulkan all supported), good quality, smallest VRAM of the popular 4-bit options at ~9 GB for a 13B. Costs ~30 tok/s vs ~52 for EXL2 at the same quality — i.e. you pay roughly 40% throughput for portability. Note Q4_K_M is a real 4.8-bit average, not 4.0. Saves ~72% VRAM vs FP16 (willitrunai.com).
  • GGUF Q8_0 / Q5_K_M — use Q5/Q6 when the model is small enough that VRAM is free, or for reasoning/coding tasks where 4-bit degradation is visible. Q8_0 is effectively lossless for most eval suites but costs 2× the memory of Q4. Practical ladder: Q4_K_M default → Q5_K_M if VRAM allows → Q8_0 only when you're chasing the last point of quality. Comparison of Q4_K_M vs Q5_K_M vs Q8_0 with decision criteria at bmdpat.com; K/I-quant format choice at kaitchup.substack.com; an independent unified evaluation of llama.cpp quants on Llama-3.1-8B at arXiv:2601.14277; GGUF quant behaviour on a low-power ARM board (RK3588) at turingpi.com.
  • AWQ (activation-aware weight quant, GPU-only, NVIDIA-first) — best when you're serving via vLLM/HF Transformers and want a sweet spot of quality and size; 32g groups beat 128g on both perplexity and size but need supported kernels.
  • GPTQ — the legacy option; still fine, still widely available, but AWQ generally matches or beats it at the same size. Note it was the fastest prompt-processor in the 2023 test at 1.85 s, ahead of llama.cpp's 3.73 s at the time.
  • EXL2 / EXL3 (ExLlama, vLLM/TabbyAPI/llama.cpp-TE) — the speed king at ~52–57 tok/s for a 13B on a 3090, roughly double GGUF's 30.8, with equal-or-better perplexity. Choose EXL3 for maximum throughput on NVIDIA cards where you control the stack. EXL3 4.00 bpw used 5.79 GiB at 0.02479 KLD vs GGUF Q4_K_M's 6.62 GiB at 0.03985 — better quality and smaller (atomic.chat), though one benchmark found llama.cpp ~12% faster than ExLlamaV3 (LinkedIn/K. Stoykov), so measure on your own model. EXL2's weak point is long context — it only recently caught up with flash attention and context shifting.
  • FP8 (weights and KV) — the pragmatic "large model on a small card" option. FP8 weights ≈ 1 byte/param, FP8 KV cache halves KV memory with a small quality cost. Best supported on Hopper/Ada/Blackwell where FP8 tensor cores exist.
  • bitsandbytes NF4 (load_in_4bit) — use it for fine-tuning (QLoRA) and quick experimentation, not as your serving format. It was the slowest option measured (23.1 tok/s) and, despite being "4-bit", the load_in_4bit checkpoint was 24.8 GB on disk — double the true 4-bit files — because HF stores the FP16 original alongside. QLoRA itself is well-validated: 4-bit NormalFloat + double quantization matches 16-bit full fine-tuning quality (arXiv:2305.14314, HF blog).

Rule of thumb: Q4_K_M if you want it to run everywhere, EXL3 if you want it to run fast, AWQ/GPTQ if you're inside a vLLM production stack, FP8 if you need a big model in limited VRAM, NF4 only for training.

Sources

  • https://codingprotocols.com/blog/local-llm-vram-requirements-quantization — The weights + KV + overhead formula, bytes/param table, worked VRAM examples, offloading behaviour.
  • https://flaviocopes.com/llm-vram-requirements — Single-line VRAM formula with the 1.05 metadata factor and a 70B worked example.
  • https://localaimaster.com/blog/kv-cache-paged-attention-guide — KV cache math, per-token cost, GQA/MLA/CLA reduction table, PagedAttention, prefix caching, quantized KV.
  • https://presenc.ai/research/local-llm-tokens-per-second-benchmarks-2026 — Measured 2026 tok/s for RTX 5090, DGX Spark, M5 Max/Ultra by model, plus prefill vs decode split.
  • https://modelfit.io/gpu — GPU speed ladder (Qwen3 8B Q4_K @16K), VRAM, street prices, multi-card VRAM pooling.
  • https://oobabooga.github.io/blog/posts/gptq-awq-exl2-llamacpp — Head-to-head GPTQ/AWQ/EXL2/Q4_K_M/NF4 perplexity, VRAM, prefill and tok/s on an RTX 3090.
  • https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF Q4_K_M on size, KLD divergence and speed.
  • https://willitrunai.com/blog/quantization-guide-gguf-explained — GGUF quantization guide: Q4_K_M saves ~72% VRAM, Q4 vs Q5 vs Q8 trade-offs.
  • https://kaitchup.substack.com/p/choosing-a-gguf-model-k-quants-i — Choosing between K-quants, I-quants and legacy GGUF formats.
  • https://arxiv.org/html/2601.14277v1 — Unified evaluation of llama.cpp quantization on Llama-3.1-8B-Instruct.
  • https://turingpi.com/llm-inference-benchmarks-rk3588-gguf-quantization — GGUF quantization compared on an ARM RK3588 board (CPU/ARM inference reality).
  • https://bmdpat.com/blog/gguf-quantization-q4-q5-q8-explained-2026 — Q4_K_M vs Q5_K_M vs Q8_0 selection criteria.
  • https://huggingface.co/blog/4bit-transformers-bitsandbytes — bitsandbytes 4-bit / NF4 mechanics and use cases.
  • https://ar5iv.labs.arxiv.org/html/2305.14314 — QLoRA paper: 4-bit NF4 + double quantization matching 16-bit fine-tuning.
  • https://github.com/ggml-org/llama.cpp/discussions/23470 — Asymmetric q8/q4 KV cache quantization to control VRAM at long context.
  • https://github.com/ggml-org/llama.cpp/discussions/21961 — Paged KV cache and scheduler design for llama.cpp.
  • https://docs.vllm.ai/en/stable/serving/parallelism_scaling — Official vLLM tensor/pipeline/data/expert parallelism configuration.
  • https://jarvislabs.ai/blog/scaling-llm-inference-dp-pp-tp — Practical walkthrough of DP vs PP vs TP trade-offs.
  • https://www.glukhov.org/llm-hosting/comparisons/amd-rocm-vs-vulkan-llm-hosting — ROCm vs Vulkan decision matrix per engine; ROCm 10.0.0 and RDNA 4 support.
  • https://www.phoronix.com/forums/forum/linux-graphics-x-org-drivers/open-source-amd-linux/1592861-amd-rocm-7-1-vs-radv-vulkan-for-llama-cpp-with-the-radeon-ai-pro-r9700 — ROCm 7.1 vs RADV Vulkan llama.cpp benchmarks on Radeon AI Pro R9700.
  • https://www.compute-market.com/blog/best-cpu-for-local-llm-2026 — Memory bandwidth matters more than core count for CPU inference.
  • https://dev.to/maximsaplin/ddr5-speed-and-llm-inference-3cdn — Measured effect of DDR5 speed on LLM inference throughput.
  • https://www.corsair.com/us/en/explorer/diy-builder/how-tos/memory-for-local-llms-how-much-ram-do-you-need-and-when-speed-matters — How much RAM local LLMs need and when speed matters.
  • https://www.promptquorum.com/local-llms/best-cpu-only-llm — CPU-only 2026 reality check: Phi-4-mini at ~12 tok/s.
  • https://www.tomshardware.com/desktops/exploring-apple-silicons-local-ai-performance-with-the-mac-studio-and-m4-max-m4-max-beats-gb10-and-strix-halo-in-decode-throughput-but-memory-bandwidth-isnt-everything — M4 Max decode throughput vs GB10 and Strix Halo; bandwidth isn't everything.
  • https://llmcheck.net/benchmarks — Apple Silicon LLM benchmarks across M1–M6 (figures to 258 tok/s).
  • https://www.newegg.com/insider/best-gpus-for-ai-and-local-llms-in-2026-vram-is-king — 2026 GPU buying guide; VRAM as the binding constraint.
  • https://www.compute-market.com/blog/best-local-llm-rtx-50-series-2026 — RTX 50-series VRAM, bandwidth, MSRP and best-paired model.
  • https://modelfit.io/gpu/rtx-3090 — Used RTX 3090 24 GB value case: runs 32B at Q4, ~$900.
  • https://intuitionlabs.ai/articles/local-llm-deployment-24gb-gpu-optimization — 24 GB GPU deployment and July 2026 used-market pricing.
  • https://www.hostrunway.com/blog/best-gpu-for-running-local-llms-and-private-ai-in-2026-complete-buyers-guide-ollama-lm-studio-llama-cpp — Buyer's guide by tier; RTX 5060 Ti 16 GB pricing.
  • https://fluence.ai/blog/best-gpu-for-llm — 2026 local LLM GPU shortlist and recommendation.
  • https://rocm.blogs.amd.com/software-tools-optimization/vllm-moe-guide/README.html — vLLM MoE playbook: TP/DP/PP/expert parallelism in practice.
  • https://docs.vllm.ai/en/v0.8.0/serving/distributed_serving.html — vLLM distributed/multi-GPU inference docs.
  • https://www.spheron.network/blog/gpu-memory-requirements-llm — ~2 GB/B at FP16, ~0.6 GB/B at INT4 VRAM rules of thumb.

Date: 2026-10-09