Hardware Buying Guide — GPUs for Local LLMs - Local AI Agent Wiki __md_scope=new URL("../../..",location),__md_hash=e=>[...e].reduce(((e,_)=>(e<<5)-e+_.charCodeAt(0)),0),__md_get=(e,_=localStorage,t=__md_scope)=>JSON.parse(_.getItem(t.pathname+"."+e)),__md_set=(e,_,t=localStorage,a=__md_scope)=>{try{t.setItem(a.pathname+"."+e,JSON.stringify(_))}catch(e){}}
Skip to content

Hardware Buying Guide — GPUs for Local LLMs

VRAM is the binding constraint, not FLOPS. Buy VRAM first, model second. Avoid 8 GB cards as a primary inference device.

VRAM is the binding constraint, not FLOPS — a model either fits or it doesn't, and no amount of compute speed fixes an OOM (newegg.com). Bandwidth is the second constraint and sets your tok/s.

Tier Target Best picks (Oct 2026) Runs
Entry ~$400–500 RTX 5060 Ti 16 GB GDDR7, ~$429–499 (hostrunway); RTX 3060 12 GB used ~$599 list 7B–14B @ Q4 comfortably
Best value ~$900–1,200 Used RTX 3090 24 GB (~$900 used; runs 32B at Q4) (modelfit.io); RTX 5070 Ti 16 GB ~$979–1,152 (newegg) 27–32B @ Q4 on the 24 GB card
High-end ~$1,600–4,700 RTX 4090 24 GB (~104 tok/s est.) / RTX 5080 16 GB (~94) / RTX 5090 32 GB GDDR7 @ 1,792 GB/s, $1,999–2,199 MSRP but street ~$4,700 (modelfit.io, compute-market) 32–70B at Q2/Q3, 70B on 2×
Apple M4/M5 Max or Ultra 128 GB unified on M5 Max, up to 512 GB on M5 Ultra; 460–1,200 GB/s (presenc.ai) 70B dense comfortably; best for big-model single-user
Big-memory GPU 48 GB+ RTX PRO 6000 96 GB ~$12,912; used A6000/4090-class alternatives 100B+ dense, 120B MoE

Concrete recommendations:

  1. If you want one card and maximum value today: a used RTX 3090 (24 GB). ~$900 used, runs 32B at Q4_K_M, ~87 tok/s on an 8B. Its 936 GB/s bandwidth is respectable and the 24 GB is what matters. 70B needs Q2 (unusable) or dual cards (modelfit.io).
  2. If you're buying new and want headroom: RTX 5090 (32 GB). The best single-card small-model speed (290–300 tok/s on Llama 2 7B) and the only consumer card that holds a 32B dense model with room for context. But it cannot hold a 70B at Q4 (~43 GB) — don't buy it for that (presenc.ai).
  3. Best bang for a new 16 GB card: RTX 5070 Ti — same 16 GB as the 5080, ~87 tok/s, ~$979 (newegg). Several guides call a single 5070 Ti or 5080 the 2026 value sweet spot for everything up to 32B at Q4 without paying the 5090 premium (promptquorum, fluence.ai).
  4. If your models are large and dense rather than small and fast, buy a Mac Studio (M5 Max/Ultra) instead of any single GPU. The unified-memory pool is the only consumer-legal way to run 70B dense at interactive speed.
  5. CPU box: DDR5, high clocks, fewer cores. 64–128 GB of fast DDR5 beats more cores; aim for RAM ≥ 1.5× the largest unquantised model you'd ever run.
  6. Avoid 8 GB cards as a primary inference device — they cap you at ~8–9B and leave no room for KV cache at useful context.

Also note used-market pricing in 2026 tracked well below new MSRP (24 GB class listings averaged ~$1,254 in July 2026 vs the RTX 4090's much higher cost per GB) (intuitionlabs.ai).

Sources

  • https://codingprotocols.com/blog/local-llm-vram-requirements-quantization — The weights + KV + overhead formula, bytes/param table, worked VRAM examples, offloading behaviour.
  • https://flaviocopes.com/llm-vram-requirements — Single-line VRAM formula with the 1.05 metadata factor and a 70B worked example.
  • https://localaimaster.com/blog/kv-cache-paged-attention-guide — KV cache math, per-token cost, GQA/MLA/CLA reduction table, PagedAttention, prefix caching, quantized KV.
  • https://presenc.ai/research/local-llm-tokens-per-second-benchmarks-2026 — Measured 2026 tok/s for RTX 5090, DGX Spark, M5 Max/Ultra by model, plus prefill vs decode split.
  • https://modelfit.io/gpu — GPU speed ladder (Qwen3 8B Q4_K @16K), VRAM, street prices, multi-card VRAM pooling.
  • https://oobabooga.github.io/blog/posts/gptq-awq-exl2-llamacpp — Head-to-head GPTQ/AWQ/EXL2/Q4_K_M/NF4 perplexity, VRAM, prefill and tok/s on an RTX 3090.
  • https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF Q4_K_M on size, KLD divergence and speed.
  • https://willitrunai.com/blog/quantization-guide-gguf-explained — GGUF quantization guide: Q4_K_M saves ~72% VRAM, Q4 vs Q5 vs Q8 trade-offs.
  • https://kaitchup.substack.com/p/choosing-a-gguf-model-k-quants-i — Choosing between K-quants, I-quants and legacy GGUF formats.
  • https://arxiv.org/html/2601.14277v1 — Unified evaluation of llama.cpp quantization on Llama-3.1-8B-Instruct.
  • https://turingpi.com/llm-inference-benchmarks-rk3588-gguf-quantization — GGUF quantization compared on an ARM RK3588 board (CPU/ARM inference reality).
  • https://bmdpat.com/blog/gguf-quantization-q4-q5-q8-explained-2026 — Q4_K_M vs Q5_K_M vs Q8_0 selection criteria.
  • https://huggingface.co/blog/4bit-transformers-bitsandbytes — bitsandbytes 4-bit / NF4 mechanics and use cases.
  • https://ar5iv.labs.arxiv.org/html/2305.14314 — QLoRA paper: 4-bit NF4 + double quantization matching 16-bit fine-tuning.
  • https://github.com/ggml-org/llama.cpp/discussions/23470 — Asymmetric q8/q4 KV cache quantization to control VRAM at long context.
  • https://github.com/ggml-org/llama.cpp/discussions/21961 — Paged KV cache and scheduler design for llama.cpp.
  • https://docs.vllm.ai/en/stable/serving/parallelism_scaling — Official vLLM tensor/pipeline/data/expert parallelism configuration.
  • https://jarvislabs.ai/blog/scaling-llm-inference-dp-pp-tp — Practical walkthrough of DP vs PP vs TP trade-offs.
  • https://www.glukhov.org/llm-hosting/comparisons/amd-rocm-vs-vulkan-llm-hosting — ROCm vs Vulkan decision matrix per engine; ROCm 10.0.0 and RDNA 4 support.
  • https://www.phoronix.com/forums/forum/linux-graphics-x-org-drivers/open-source-amd-linux/1592861-amd-rocm-7-1-vs-radv-vulkan-for-llama-cpp-with-the-radeon-ai-pro-r9700 — ROCm 7.1 vs RADV Vulkan llama.cpp benchmarks on Radeon AI Pro R9700.
  • https://www.compute-market.com/blog/best-cpu-for-local-llm-2026 — Memory bandwidth matters more than core count for CPU inference.
  • https://dev.to/maximsaplin/ddr5-speed-and-llm-inference-3cdn — Measured effect of DDR5 speed on LLM inference throughput.
  • https://www.corsair.com/us/en/explorer/diy-builder/how-tos/memory-for-local-llms-how-much-ram-do-you-need-and-when-speed-matters — How much RAM local LLMs need and when speed matters.
  • https://www.promptquorum.com/local-llms/best-cpu-only-llm — CPU-only 2026 reality check: Phi-4-mini at ~12 tok/s.
  • https://www.tomshardware.com/desktops/exploring-apple-silicons-local-ai-performance-with-the-mac-studio-and-m4-max-m4-max-beats-gb10-and-strix-halo-in-decode-throughput-but-memory-bandwidth-isnt-everything — M4 Max decode throughput vs GB10 and Strix Halo; bandwidth isn't everything.
  • https://llmcheck.net/benchmarks — Apple Silicon LLM benchmarks across M1–M6 (figures to 258 tok/s).
  • https://www.newegg.com/insider/best-gpus-for-ai-and-local-llms-in-2026-vram-is-king — 2026 GPU buying guide; VRAM as the binding constraint.
  • https://www.compute-market.com/blog/best-local-llm-rtx-50-series-2026 — RTX 50-series VRAM, bandwidth, MSRP and best-paired model.
  • https://modelfit.io/gpu/rtx-3090 — Used RTX 3090 24 GB value case: runs 32B at Q4, ~$900.
  • https://intuitionlabs.ai/articles/local-llm-deployment-24gb-gpu-optimization — 24 GB GPU deployment and July 2026 used-market pricing.
  • https://www.hostrunway.com/blog/best-gpu-for-running-local-llms-and-private-ai-in-2026-complete-buyers-guide-ollama-lm-studio-llama-cpp — Buyer's guide by tier; RTX 5060 Ti 16 GB pricing.
  • https://fluence.ai/blog/best-gpu-for-llm — 2026 local LLM GPU shortlist and recommendation.
  • https://rocm.blogs.amd.com/software-tools-optimization/vllm-moe-guide/README.html — vLLM MoE playbook: TP/DP/PP/expert parallelism in practice.
  • https://docs.vllm.ai/en/v0.8.0/serving/distributed_serving.html — vLLM distributed/multi-GPU inference docs.
  • https://www.spheron.network/blog/gpu-memory-requirements-llm — ~2 GB/B at FP16, ~0.6 GB/B at INT4 VRAM rules of thumb.

Date: 2026-10-09

var target=document.getElementById(location.hash.slice(1));target&&target.name&&(target.checked=target.name.startsWith("__tabbed_"))
{"annotate": null, "base": "../../..", "features": [], "search": "../../../assets/javascripts/workers/search.2c215733.min.js", "tags": null, "translations": {"clipboard.copied": "Copied to clipboard", "clipboard.copy": "Copy to clipboard", "search.result.more.one": "1 more on this page", "search.result.more.other": "# more on this page", "search.result.none": "No matching documents", "search.result.one": "1 matching document", "search.result.other": "# matching documents", "search.result.placeholder": "Type to start searching", "search.result.term.missing": "Missing", "select.version": "Select version"}, "version": null}