Prefill vs Decode — Token Speed Realities
prefill is compute-bound (big batched matmul); decode is bandwidth-bound (one token = stream every weight once). The gap is enormous: RTX 5090 does ~14,000 tok/s prefill vs 290-300 tok/s decode.
Two numbers matter and they bottleneck on different resources (presenc.ai):
- Prompt processing (prefill) is compute-bound — it's a big batched matmul over the whole prompt.
- Token generation (decode) is memory-bandwidth-bound — one token requires streaming every weight once, so
tok/s ≈ memory bandwidth ÷ model size.
The gap is enormous. On Llama 2 7B the llama.cpp scoreboard shows the RTX 5090 processing prompts at ~14,000–15,000 tok/s while generating at 290–300 tok/s; the M5 Max does 3,220 prefill vs 119.9 decode. DGX Spark processes Llama 3.1 8B prompts at 7,614 tok/s under Ollama but generates Llama 3.1 70B at only 4.4 tok/s. For agent and RAG workloads with long context, prefill is the bottleneck and CUDA hardware keeps an advantage (presenc.ai).
Published generation speed, 2026 (presenc.ai) — compare within a row, not across, since backends and quants differ:
| Model / test | RTX 5090 (32 GB) | DGX Spark (128 GB) | Mac Studio M5 Max | M5 Ultra |
|---|---|---|---|---|
| Llama 2 7B Q4_0, llama.cpp | 290–300 | — | 119.9 | — |
| Llama 3.1 8B Q4_K_M, Ollama | 150.0 | 38.0 | — | — |
| gpt-oss 20B, Ollama | 205 | 49.7 | — | — |
| Gemma 3 27B Q4, Ollama | 47.3 | 10.8 | — | — |
| 32B class, Ollama | 57.2 (QwQ 32B) | 9.4 (Qwen3 32B) | — | — |
| Llama 3.1 70B, Ollama | does not fit (Q4 ≈ 43 GB > 32 GB) | 4.4 | — | 40–52 (unverified) |
| gpt-oss 120B, llama.cpp | does not fit | 60.6 (40.6 @ 32K) | — | — |
Useful thresholds: >30 tok/s feels interactive for chat, 8–30 is batch/patient-user territory, <8 is impractical for anything a person waits on (presenc.ai). Also note DGX Spark's split personality — its 273 GB/s bandwidth makes it slow on large dense models (4.4 tok/s on 70B) but fast on large sparse MoE models (60.6 tok/s on gpt-oss 120B), because MoE only reads a fraction of the weights per token. Mixture-of-experts broke the old intuition that parameter count predicts speed (codingprotocols.com).
Reasoning models deserve a separate mental model: a 32B-class model emitting 2,000 thinking tokens plus a 200-token answer takes ~38 s on an RTX 5090 (57.2 tok/s) but ~234 s on a DGX Spark (9.4 tok/s).
An independent GPU speed ladder, Qwen3 8B Q4_K @ 16K context (modelfit.io):
| GPU | VRAM | Est. tok/s | GPU | VRAM | Est. tok/s |
|---|---|---|---|---|---|
| RTX 5090 | 32 GB | 145 | RTX 5070 Ti | 16 GB | 87 |
| RTX PRO 6000 | 96 GB | 145 | RTX 3090 | 24 GB | 87 |
| RTX 4090 | 24 GB | 104 | RTX 4080 SUPER | 16 GB | 79 |
| RTX 5080 | 16 GB | 94 | RTX 4070 Ti SUPER | 16 GB | 72 |
| RTX 5070 | 12 GB | 59 | RTX 5060 Ti | 16 GB | 51 |
Sources
- https://codingprotocols.com/blog/local-llm-vram-requirements-quantization — The weights + KV + overhead formula, bytes/param table, worked VRAM examples, offloading behaviour.
- https://flaviocopes.com/llm-vram-requirements — Single-line VRAM formula with the 1.05 metadata factor and a 70B worked example.
- https://localaimaster.com/blog/kv-cache-paged-attention-guide — KV cache math, per-token cost, GQA/MLA/CLA reduction table, PagedAttention, prefix caching, quantized KV.
- https://presenc.ai/research/local-llm-tokens-per-second-benchmarks-2026 — Measured 2026 tok/s for RTX 5090, DGX Spark, M5 Max/Ultra by model, plus prefill vs decode split.
- https://modelfit.io/gpu — GPU speed ladder (Qwen3 8B Q4_K @16K), VRAM, street prices, multi-card VRAM pooling.
- https://oobabooga.github.io/blog/posts/gptq-awq-exl2-llamacpp — Head-to-head GPTQ/AWQ/EXL2/Q4_K_M/NF4 perplexity, VRAM, prefill and tok/s on an RTX 3090.
- https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF Q4_K_M on size, KLD divergence and speed.
- https://willitrunai.com/blog/quantization-guide-gguf-explained — GGUF quantization guide: Q4_K_M saves ~72% VRAM, Q4 vs Q5 vs Q8 trade-offs.
- https://kaitchup.substack.com/p/choosing-a-gguf-model-k-quants-i — Choosing between K-quants, I-quants and legacy GGUF formats.
- https://arxiv.org/html/2601.14277v1 — Unified evaluation of llama.cpp quantization on Llama-3.1-8B-Instruct.
- https://turingpi.com/llm-inference-benchmarks-rk3588-gguf-quantization — GGUF quantization compared on an ARM RK3588 board (CPU/ARM inference reality).
- https://bmdpat.com/blog/gguf-quantization-q4-q5-q8-explained-2026 — Q4_K_M vs Q5_K_M vs Q8_0 selection criteria.
- https://huggingface.co/blog/4bit-transformers-bitsandbytes — bitsandbytes 4-bit / NF4 mechanics and use cases.
- https://ar5iv.labs.arxiv.org/html/2305.14314 — QLoRA paper: 4-bit NF4 + double quantization matching 16-bit fine-tuning.
- https://github.com/ggml-org/llama.cpp/discussions/23470 — Asymmetric q8/q4 KV cache quantization to control VRAM at long context.
- https://github.com/ggml-org/llama.cpp/discussions/21961 — Paged KV cache and scheduler design for llama.cpp.
- https://docs.vllm.ai/en/stable/serving/parallelism_scaling — Official vLLM tensor/pipeline/data/expert parallelism configuration.
- https://jarvislabs.ai/blog/scaling-llm-inference-dp-pp-tp — Practical walkthrough of DP vs PP vs TP trade-offs.
- https://www.glukhov.org/llm-hosting/comparisons/amd-rocm-vs-vulkan-llm-hosting — ROCm vs Vulkan decision matrix per engine; ROCm 10.0.0 and RDNA 4 support.
- https://www.phoronix.com/forums/forum/linux-graphics-x-org-drivers/open-source-amd-linux/1592861-amd-rocm-7-1-vs-radv-vulkan-for-llama-cpp-with-the-radeon-ai-pro-r9700 — ROCm 7.1 vs RADV Vulkan llama.cpp benchmarks on Radeon AI Pro R9700.
- https://www.compute-market.com/blog/best-cpu-for-local-llm-2026 — Memory bandwidth matters more than core count for CPU inference.
- https://dev.to/maximsaplin/ddr5-speed-and-llm-inference-3cdn — Measured effect of DDR5 speed on LLM inference throughput.
- https://www.corsair.com/us/en/explorer/diy-builder/how-tos/memory-for-local-llms-how-much-ram-do-you-need-and-when-speed-matters — How much RAM local LLMs need and when speed matters.
- https://www.promptquorum.com/local-llms/best-cpu-only-llm — CPU-only 2026 reality check: Phi-4-mini at ~12 tok/s.
- https://www.tomshardware.com/desktops/exploring-apple-silicons-local-ai-performance-with-the-mac-studio-and-m4-max-m4-max-beats-gb10-and-strix-halo-in-decode-throughput-but-memory-bandwidth-isnt-everything — M4 Max decode throughput vs GB10 and Strix Halo; bandwidth isn't everything.
- https://llmcheck.net/benchmarks — Apple Silicon LLM benchmarks across M1–M6 (figures to 258 tok/s).
- https://www.newegg.com/insider/best-gpus-for-ai-and-local-llms-in-2026-vram-is-king — 2026 GPU buying guide; VRAM as the binding constraint.
- https://www.compute-market.com/blog/best-local-llm-rtx-50-series-2026 — RTX 50-series VRAM, bandwidth, MSRP and best-paired model.
- https://modelfit.io/gpu/rtx-3090 — Used RTX 3090 24 GB value case: runs 32B at Q4, ~$900.
- https://intuitionlabs.ai/articles/local-llm-deployment-24gb-gpu-optimization — 24 GB GPU deployment and July 2026 used-market pricing.
- https://www.hostrunway.com/blog/best-gpu-for-running-local-llms-and-private-ai-in-2026-complete-buyers-guide-ollama-lm-studio-llama-cpp — Buyer's guide by tier; RTX 5060 Ti 16 GB pricing.
- https://fluence.ai/blog/best-gpu-for-llm — 2026 local LLM GPU shortlist and recommendation.
- https://rocm.blogs.amd.com/software-tools-optimization/vllm-moe-guide/README.html — vLLM MoE playbook: TP/DP/PP/expert parallelism in practice.
- https://docs.vllm.ai/en/v0.8.0/serving/distributed_serving.html — vLLM distributed/multi-GPU inference docs.
- https://www.spheron.network/blog/gpu-memory-requirements-llm — ~2 GB/B at FP16, ~0.6 GB/B at INT4 VRAM rules of thumb.
Date: 2026-10-09