Skip to content

Prefill vs Decode — Token Speed Realities

prefill is compute-bound (big batched matmul); decode is bandwidth-bound (one token = stream every weight once). The gap is enormous: RTX 5090 does ~14,000 tok/s prefill vs 290-300 tok/s decode.

Two numbers matter and they bottleneck on different resources (presenc.ai):

  • Prompt processing (prefill) is compute-bound — it's a big batched matmul over the whole prompt.
  • Token generation (decode) is memory-bandwidth-bound — one token requires streaming every weight once, so tok/s ≈ memory bandwidth ÷ model size.

The gap is enormous. On Llama 2 7B the llama.cpp scoreboard shows the RTX 5090 processing prompts at ~14,000–15,000 tok/s while generating at 290–300 tok/s; the M5 Max does 3,220 prefill vs 119.9 decode. DGX Spark processes Llama 3.1 8B prompts at 7,614 tok/s under Ollama but generates Llama 3.1 70B at only 4.4 tok/s. For agent and RAG workloads with long context, prefill is the bottleneck and CUDA hardware keeps an advantage (presenc.ai).

Published generation speed, 2026 (presenc.ai) — compare within a row, not across, since backends and quants differ:

Model / test RTX 5090 (32 GB) DGX Spark (128 GB) Mac Studio M5 Max M5 Ultra
Llama 2 7B Q4_0, llama.cpp 290–300 — 119.9 —
Llama 3.1 8B Q4_K_M, Ollama 150.0 38.0 — —
gpt-oss 20B, Ollama 205 49.7 — —
Gemma 3 27B Q4, Ollama 47.3 10.8 — —
32B class, Ollama 57.2 (QwQ 32B) 9.4 (Qwen3 32B) — —
Llama 3.1 70B, Ollama does not fit (Q4 ≈ 43 GB > 32 GB) 4.4 — 40–52 (unverified)
gpt-oss 120B, llama.cpp does not fit 60.6 (40.6 @ 32K) — —

Useful thresholds: >30 tok/s feels interactive for chat, 8–30 is batch/patient-user territory, <8 is impractical for anything a person waits on (presenc.ai). Also note DGX Spark's split personality — its 273 GB/s bandwidth makes it slow on large dense models (4.4 tok/s on 70B) but fast on large sparse MoE models (60.6 tok/s on gpt-oss 120B), because MoE only reads a fraction of the weights per token. Mixture-of-experts broke the old intuition that parameter count predicts speed (codingprotocols.com).

Reasoning models deserve a separate mental model: a 32B-class model emitting 2,000 thinking tokens plus a 200-token answer takes ~38 s on an RTX 5090 (57.2 tok/s) but ~234 s on a DGX Spark (9.4 tok/s).

An independent GPU speed ladder, Qwen3 8B Q4_K @ 16K context (modelfit.io):

GPU VRAM Est. tok/s GPU VRAM Est. tok/s
RTX 5090 32 GB 145 RTX 5070 Ti 16 GB 87
RTX PRO 6000 96 GB 145 RTX 3090 24 GB 87
RTX 4090 24 GB 104 RTX 4080 SUPER 16 GB 79
RTX 5080 16 GB 94 RTX 4070 Ti SUPER 16 GB 72
RTX 5070 12 GB 59 RTX 5060 Ti 16 GB 51

Sources

  • https://codingprotocols.com/blog/local-llm-vram-requirements-quantization — The weights + KV + overhead formula, bytes/param table, worked VRAM examples, offloading behaviour.
  • https://flaviocopes.com/llm-vram-requirements — Single-line VRAM formula with the 1.05 metadata factor and a 70B worked example.
  • https://localaimaster.com/blog/kv-cache-paged-attention-guide — KV cache math, per-token cost, GQA/MLA/CLA reduction table, PagedAttention, prefix caching, quantized KV.
  • https://presenc.ai/research/local-llm-tokens-per-second-benchmarks-2026 — Measured 2026 tok/s for RTX 5090, DGX Spark, M5 Max/Ultra by model, plus prefill vs decode split.
  • https://modelfit.io/gpu — GPU speed ladder (Qwen3 8B Q4_K @16K), VRAM, street prices, multi-card VRAM pooling.
  • https://oobabooga.github.io/blog/posts/gptq-awq-exl2-llamacpp — Head-to-head GPTQ/AWQ/EXL2/Q4_K_M/NF4 perplexity, VRAM, prefill and tok/s on an RTX 3090.
  • https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF Q4_K_M on size, KLD divergence and speed.
  • https://willitrunai.com/blog/quantization-guide-gguf-explained — GGUF quantization guide: Q4_K_M saves ~72% VRAM, Q4 vs Q5 vs Q8 trade-offs.
  • https://kaitchup.substack.com/p/choosing-a-gguf-model-k-quants-i — Choosing between K-quants, I-quants and legacy GGUF formats.
  • https://arxiv.org/html/2601.14277v1 — Unified evaluation of llama.cpp quantization on Llama-3.1-8B-Instruct.
  • https://turingpi.com/llm-inference-benchmarks-rk3588-gguf-quantization — GGUF quantization compared on an ARM RK3588 board (CPU/ARM inference reality).
  • https://bmdpat.com/blog/gguf-quantization-q4-q5-q8-explained-2026 — Q4_K_M vs Q5_K_M vs Q8_0 selection criteria.
  • https://huggingface.co/blog/4bit-transformers-bitsandbytes — bitsandbytes 4-bit / NF4 mechanics and use cases.
  • https://ar5iv.labs.arxiv.org/html/2305.14314 — QLoRA paper: 4-bit NF4 + double quantization matching 16-bit fine-tuning.
  • https://github.com/ggml-org/llama.cpp/discussions/23470 — Asymmetric q8/q4 KV cache quantization to control VRAM at long context.
  • https://github.com/ggml-org/llama.cpp/discussions/21961 — Paged KV cache and scheduler design for llama.cpp.
  • https://docs.vllm.ai/en/stable/serving/parallelism_scaling — Official vLLM tensor/pipeline/data/expert parallelism configuration.
  • https://jarvislabs.ai/blog/scaling-llm-inference-dp-pp-tp — Practical walkthrough of DP vs PP vs TP trade-offs.
  • https://www.glukhov.org/llm-hosting/comparisons/amd-rocm-vs-vulkan-llm-hosting — ROCm vs Vulkan decision matrix per engine; ROCm 10.0.0 and RDNA 4 support.
  • https://www.phoronix.com/forums/forum/linux-graphics-x-org-drivers/open-source-amd-linux/1592861-amd-rocm-7-1-vs-radv-vulkan-for-llama-cpp-with-the-radeon-ai-pro-r9700 — ROCm 7.1 vs RADV Vulkan llama.cpp benchmarks on Radeon AI Pro R9700.
  • https://www.compute-market.com/blog/best-cpu-for-local-llm-2026 — Memory bandwidth matters more than core count for CPU inference.
  • https://dev.to/maximsaplin/ddr5-speed-and-llm-inference-3cdn — Measured effect of DDR5 speed on LLM inference throughput.
  • https://www.corsair.com/us/en/explorer/diy-builder/how-tos/memory-for-local-llms-how-much-ram-do-you-need-and-when-speed-matters — How much RAM local LLMs need and when speed matters.
  • https://www.promptquorum.com/local-llms/best-cpu-only-llm — CPU-only 2026 reality check: Phi-4-mini at ~12 tok/s.
  • https://www.tomshardware.com/desktops/exploring-apple-silicons-local-ai-performance-with-the-mac-studio-and-m4-max-m4-max-beats-gb10-and-strix-halo-in-decode-throughput-but-memory-bandwidth-isnt-everything — M4 Max decode throughput vs GB10 and Strix Halo; bandwidth isn't everything.
  • https://llmcheck.net/benchmarks — Apple Silicon LLM benchmarks across M1–M6 (figures to 258 tok/s).
  • https://www.newegg.com/insider/best-gpus-for-ai-and-local-llms-in-2026-vram-is-king — 2026 GPU buying guide; VRAM as the binding constraint.
  • https://www.compute-market.com/blog/best-local-llm-rtx-50-series-2026 — RTX 50-series VRAM, bandwidth, MSRP and best-paired model.
  • https://modelfit.io/gpu/rtx-3090 — Used RTX 3090 24 GB value case: runs 32B at Q4, ~$900.
  • https://intuitionlabs.ai/articles/local-llm-deployment-24gb-gpu-optimization — 24 GB GPU deployment and July 2026 used-market pricing.
  • https://www.hostrunway.com/blog/best-gpu-for-running-local-llms-and-private-ai-in-2026-complete-buyers-guide-ollama-lm-studio-llama-cpp — Buyer's guide by tier; RTX 5060 Ti 16 GB pricing.
  • https://fluence.ai/blog/best-gpu-for-llm — 2026 local LLM GPU shortlist and recommendation.
  • https://rocm.blogs.amd.com/software-tools-optimization/vllm-moe-guide/README.html — vLLM MoE playbook: TP/DP/PP/expert parallelism in practice.
  • https://docs.vllm.ai/en/v0.8.0/serving/distributed_serving.html — vLLM distributed/multi-GPU inference docs.
  • https://www.spheron.network/blog/gpu-memory-requirements-llm — ~2 GB/B at FP16, ~0.6 GB/B at INT4 VRAM rules of thumb.

Date: 2026-10-09