Skip to content

Running Open-Weight AI Models Locally — State of the Art

Master report as of October 2026: engine choice is now gated by model architecture, VRAM math is 0.6 GB/billion for Q4_K_M, and everything speaks the OpenAI API shape. Executive summary, verified-vs-uncertain ledger, and seven recommendations.

Date: 2026-10-09 Prepared for: Paulo Prepared by: JUVENAL Method: parallel multi-agent research — 5 independent subagents, one per section, each citing live sources (144 KB across 5 section files in this folder).


1. Executive Summary

  • The engine choice has bifurcated into two families, plus a new third. GGUF/C++ engines (llama.cpp and derivatives) win for single-user, CPU, Apple Silicon and edge. Batched GPU serving engines (vLLM, SGLang, TensorRT-LLM) win for concurrency and throughput. A third class emerged in 2026: agent-oriented engines (Magnitude, Splash, vllm.cpp, vllm-metal) built for the reality that the main local workload is now coding agents, not chat.
  • The famous "44× between engines" gap is a concurrency artefact. llama.cpp is the fastest of five for a single user (185 tok/s @ batch 1) and fourth of five at 50 concurrent. At batch 1 the engines are near parity. Judge engines on your concurrency, not on headline benchmarks.
  • Ollama is the simplest on-ramp but has a hard concurrency ceiling. It defaults to OLLAMA_NUM_PARALLEL=1, and since Feb 2026 forces a single slot for eleven architecture families. On one H100 it served 33 tok/s at 50 concurrent vs vLLM's 1,610 — a 52× gap, with a 381-second median TTFT. Fine for one developer; wrong for production serving.
  • Model architecture now gates engine choice — this is the biggest 2026 change. Hybrid linear-attention models (Qwen3.8's Gated DeltaNet) need per-request recurrent state, and every engine sizes it differently: TensorRT-LLM couldn't load it at all, vLLM refused to start, SGLang silently capped itself to 12–20 concurrent requests. Check your specific model against your engine before committing.
  • VRAM math: use 0.6 GB per billion parameters for Q4_K_M, not the widely-cited 0.5. K-quants keep sensitive tensors at higher precision, so the effective rate is always above nominal. Verified against real files: 7B → 4.1 GB, 13B → 7.9 GB, 70B → ~40 GB.
  • The KV cache can be bigger than the weights. Llama 3.1 70B at 128K context needs ~43 GB of KV per request — more than FP8 weights. It's ~8× worse without GQA. This is what actually kills long-context on consumer cards.
  • Best value for a local setup is a used RTX 3090 24 GB (~$900) — runs a 32B at Q4_K_M. An RTX 5090 (32 GB) is the best single-card small-model speed but cannot hold a 70B at Q4. The only consumer route to a 70B dense model at interactive speed is a Mac Studio (M5 Max/Ultra) unified memory.
  • Qwen3.8-27B is the standout local model — a dense 27B scoring 77.2% on SWE-bench Verified, beating Alibaba's own 397B MoE from two months earlier, fitting ~24 GB. Apache 2.0. For Portuguese, Qwen3 8B is the top Ollama-native choice; Sabiá-3 is higher quality but HuggingFace-only.
  • Everything speaks the OpenAI API shape. llama.cpp, Ollama, vLLM, LM Studio, SGLang, LocalAI all expose /v1/chat/completions etc., so an app written once against the OpenAI SDK runs locally by only swapping base_url. 2026 added Anthropic Messages compatibility so Claude Code can drive open models directly.
  • Bottom line: for one developer, Ollama or LM Studio. For a coding agent on a Mac, MLX/Splash. For serving to many users, vLLM. Buy VRAM first, model second, and verify engine/model compatibility before you commit.

2. What we verified vs what remains uncertain

Claim Status Source
Engine versions, star counts, last-commit dates verified — GitHub API / PyPI section 1
Concurrency benchmarks (H100, 1/10/50 concurrent) verified — independent reproducible test winder.ai
Q4_K_M = 0.6 GB/billion params verified — against real GGUF files codingprotocols.com
KV cache ~43 GB for 70B @ 128K verified — formula + model cards section 3
Benchmark scores (SWE-bench etc.) vendor-reported, not independent computingforgeeks
Muse Glimmer MCP Atlas 75.5 vendor-reported (Meta's own table) digitalapplied
M5 Ultra 70B = 40–52 tok/s uncertain — early community measurement, publisher marks it unverified section 3
Magnitude "92% faster decode than llama.cpp" unverified vendor claim section 1
Llama 5 existence debunked — no official release exists orcarouter
GPT4All maintenance status verified — last commit 2025-05-27, effectively EOL section 1

Benchmark-version caution: two vendors reporting the "same" benchmark name can differ by 11+ points because they ran different versions. Always check the version suffix before comparing scores.


3. Findings

3.1 Inference engines — what's actually maintained

Full detail: Inference Engines Overview (18 KB, 40 sources)

Engine Version (Oct 2026) Stars Status Best for
llama.cpp b11505 130,619 very active single-user, any hardware, portability
Ollama v0.40.1 182,473 very active one developer, simplest path
vLLM 0.31.0 93,444 very active concurrency, production serving
SGLang 0.5.21 36,919 very active shared prefixes, hybrid models
LocalAI 4.11.0 49,450 very active one process, LLM+vision+speech
LM Studio 0.4.26 proprietary very active GUI + headless llmster daemon
KoboldCpp v1.122.1 11,974 active roleplay/creative
ExLlamaV3 1.5.3 1,623 active, pre-1.0 max quality-per-bit on RTX
ExLlamaV2 0.3.2 4,632 legacy/dormant —
TensorRT-LLM 1.3 14,780 active, slow model support NVIDIA-only, lowest TTFT
MLC-LLM — 23,226 low activity browser/mobile/edge
llamafile — 26,209 slow cadence single-file distribution
GPT4All — 77,373 effectively EOL do not build on it

Concurrency, measured on one H100:

Engine tok/s @ 1 / 10 / 50 concurrent Median TTFT
SGLang 0.5.20 81 / 627 / 1,725 1.5 s
vLLM 0.30.0 76 / 589 / 1,610 2.1 s
llama.cpp 56 / 138 / 97 8.2 s
Ollama 0.34.4 33 / 33 / 33 381 s
TensorRT-LLM could not load (hybrid model) —

New in 2026: vllm.cpp (C++20 vLLM port, no Python at inference, but Blackwell-only CUDA images), vllm-metal (vLLM scheduler over MLX — 3.64 s vs mlx_lm's 10.99 s on ragged batches), Magnitude and Splash (agent-oriented, Apple Silicon).

3.2 Open-weight models — what to actually run

Full detail: Open-Weight Model Landscape (13.6 KB, 13 sources)

Hardware tier Fits at Q4
8 GB Qwen3 8B, Gemma 4 E4B, Phi-4 Mini, gpt-oss-20b, Muse Glimmer
16 GB Gemma 4 12B, Qwen3 14B, Phi-4, Qwen3.8-Flash-Next (6B active)
24 GB Qwen3.8-27B, Gemma 4 31B/26B-A4B, Muse Glimmer 30B, Kimi-Linear 48B-A3B
48 GB Llama 3.3 70B, GLM-5.3-Flash (18B active), DeepSeek V4 Flash 284B/A13B
128 GB+ Kimi K3, GLM-5.3, Qwen3.8-Max, DeepSeek V4 Pro (server-class)

Top picks: Qwen3.8-27B (dense, Apache 2.0, 77.2% SWE-bench Verified), DeepSeek V4 Pro (80.6%, top open-weight SWE-bench), Muse Glimmer 30B (Meta's first unmodified Apache 2.0 release, built for local agents), gpt-oss-120b (OpenAI, Apache 2.0). Licensing trap: Qwen3.8-Max, GLM-5.3 and Kimi K3 moved to custom licences in Aug 2026 while their mid-range siblings stayed permissive.

3.3 Hardware, quantization, performance

Full detail: VRAM Math (VRAM math), Quantization Formats (quant formats), Prefill vs Decode (speed benchmarks)

total VRAM = weights + KV cache + overhead
weights (GB) = parameters (B) × bytes-per-param   → use 0.6 for Q4_K_M
KV cache (GB) ≈ 2 × layers × kv_heads × head_dim × context_len × bytes / 1e9
overhead ≈ 0.5–1.5 GB
  • Quantization measured (RTX 3090, 13B): EXL2 4.9bpw best quality (4.308 ppl) at 52 tok/s; GGUF Q4_K_M is 4.333 ppl but only 30.8 tok/s — ~40% throughput paid for portability. bnb NF4: slowest, and despite being "4-bit" stored 24.8 GB on disk. Use NF4 for QLoRA training only, never serving.
  • Prefill vs decode is a 50–100× gap. RTX 5090: ~14,000–15,000 tok/s prefill vs 290–300 tok/s decode. Prefill is compute-bound, decode is bandwidth-bound (tok/s ≈ bandwidth ÷ model size). For RAG/agent workloads prefill is the bottleneck.
  • Thresholds: >30 tok/s interactive, 8–30 batch, <8 impractical.
  • AMD is now viable for inference: ROCm 7.2 (Mar 2026) first with official RDNA 4 support and claimed out-of-box CUDA parity; RX 7900 XTX runs Llama 3.1 8B at ~96 tok/s ≈ 75% of a 4090. But Vulkan remains the pragmatic default for local llama.cpp, and silent CPU-fallback is the classic AMD trap — verify in logs, not device queries.
  • Apple Silicon: MLX is the native path; M5 Neural Accelerators give up to 3.97× faster TTFT. Bandwidth is the ceiling and it favours big models — M4 Max beat NVIDIA GB10 and AMD Strix Halo in decode despite lower absolute bandwidth, because it has far more memory.
  • CPU-only is bandwidth-bound, not core-bound — fast DDR5 beats more cores. ~12 tok/s class for a 3–4B at Q4.

3.4 The application stack

Full detail: Application Stack Overview (22 KB, 50 sources)

  • Serving: every runtime exposes OpenAI-compatible endpoints. Point OPENAI_BASE_URL at localhost:11434/v1 (Ollama) or llama-server with a dummy key. LiteLLM Proxy fronts 100+ providers.
  • Coding agents: for local use you need a strong 30B+ quantised model on vLLM for usable function calling; smaller models degrade fast. OpenCode + vLLM is the most current combination.
  • RAG: Qwen3-Embedding is the open-weight leader (8B ranks #1 on MTEB multilingual; the 0.6B is a strong cheap local choice). Pattern: Qwen3-Embedding → Qdrant/FAISS → Qwen3-Reranker → LLM on vLLM via LlamaIndex. Note: an instruct prompt typically improves retrieval 1–5%. Milvus's PyPI package is stale (2024) — use Qdrant.
  • Fine-tuning: Unsloth (v2026.10.3) is the fastest single-GPU LoRA/QLoRA path and ships a desktop app; Axolotl (v0.20.0) for config-driven multi-GPU/FSDP2.
  • Local speech/image: faster-whisper for ASR, Kokoro/Piper for TTS, ComfyUI + FLUX for images.

3.5 Licensing, security, operations

Full detail: Practice & Ops — Licensing (31 KB, 36 sources)

  • Licensing: only Apache-2.0/MIT weights are genuinely OSI open source (Qwen3, DeepSeek, gpt-oss, Gemma 4, Mistral Small/Large 3, OLMo 2, Phi-4, Muse Glimmer). Llama 4 = free under 700M MAU + "Built with Llama" attribution + EU exclusion on multimodal. Gemma 3 carries a Google-enforceable Prohibited Use Policy. Downloading locally does not exempt you from acceptable-use terms.
  • EU AI Act (Oct 2026): the Digital Omnibus (Reg. 2026/1744, in force 27 Jul 2026) delayed high-risk rules to Dec 2027/Aug 2028 but did not delay GPAI or Article 50 transparency. The open-source exemption (Art. 53(2)/54(6)) is void for systemic-risk models (>10²⁵ FLOP) and never covers copyright/training-data duties. Fine-tuning or materially modifying a model and putting it on the EU market likely makes you a provider. Local inference is a strong GDPR lever — it removes the Art. 28 data-processor obligation entirely.
  • Security: local ≠ secure. SafeTensors eliminates the pickle RCE vector; trust_remote_code is the second; HF pulled 9 malicious models in Q1 2026, 3 already in production. LM Studio ships telemetry on by default. Never set OLLAMA_HOST=0.0.0.0.
  • Drivers: the NVIDIA driver is no longer bundled with the CUDA toolkit on Linux since 13.4. ROCm 7.2.x is the first release with official RDNA 4 support. Vulkan is the cross-vendor fallback, not a library replacement.
  • Cost: breakeven is workload-specific. A 7B on a Mac Studio vs frontier APIs breaks even in ~0.3 months; a dense 70B on DGX Spark (4.4 tps) never beats cheap open-weight APIs. Below 10% utilisation, cloud wins.

4. Choosing an engine

Need Choice
One user, any hardware, simplest Ollama
Single-user max throughput, CPU/edge llama.cpp
GUI + model discovery, or headless daemon LM Studio
Concurrency / production serving vLLM
Shared prefixes, hybrid models SGLang
Max quality-per-bit, one modern GPU ExLlamaV3 / EXL3
Apple Silicon, ≥14B MLX
Apple Silicon, concurrent agents vllm-metal or Splash
Local coding agents, long sessions Magnitude / Splash / LM Studio
One process, LLM + vision + speech LocalAI
Offline ChatGPT alternative Jan
Browser / mobile / edge MLC-LLM / WebLLM
CPU-only legacy box Ollama or Jan (not GPT4All)

5. Recommendations

  1. Start with Ollama, move to vLLM if you need concurrency. Ollama for one developer; vLLM the moment more than one request at a time matters. The 52× gap is real.
  2. Buy VRAM first. A used RTX 3090 24 GB (~$900) is the best value entry and runs a 32B at Q4. Avoid 8 GB cards. If your models are large and dense, a Mac Studio is the only consumer route to 70B at interactive speed.
  3. Default to GGUF Q4_K_M for portability, EXL3 if you want speed on one RTX. You pay ~40% throughput for GGUF's portability; take it when the model moves between machines, skip it when it doesn't.
  4. Size for the KV cache, not just the weights. Budget ~2× the model requirement on unified-memory machines, and check context length against KV math before committing to 128K.
  5. Use Qwen3.8-27B as the default local workhorse — dense, Apache 2.0, beats much larger MoE models, fits 24 GB. For Portuguese, Qwen3 8B via Ollama, Sabiá-3 if quality is paramount.
  6. Verify engine/model compatibility before adopting any hybrid linear-attention model. This eliminated TensorRT-LLM and crippled two other engines in Sept 2026 testing.
  7. Check the licence, not the word "open." Apache-2.0/MIT only for commercial work; Llama/Gemma/GLM-5.3/Qwen3.8-Max/Kimi K3 all carry real restrictions, and fine-tuning + EU market placement can make you a regulated provider.

6. Open questions

  • Does Magnitude actually deliver its claimed 92% decode speedup? Vendor-reported and unverified — needs an independent reproduction.
  • M5 Ultra 70B figures (40–52 tok/s) are early community measurements the publisher marks unverified — needs confirmation.
  • Will ROCm reach true CUDA parity for training, not just inference? Inference is viable now; training tooling still lags.
  • Will Llama-family weights stay openly released? Meta has signalled a shift toward closed releases (Muse Spark / Superintelligence Labs, Apr 2026).

7. Sources

Consolidated from the five section files (~170 distinct sources). The authoritative per-topic lists are in each section file; the highest-value primary sources:

Primary

  • https://github.com/ggml-org/llama.cpp — llama.cpp repo, build docs, multi-GPU docs
  • https://github.com/ollama/ollama + https://docs.ollama.com/gpu — Ollama, backend support, OpenAI compat
  • https://docs.vllm.ai/en/stable/usage/v1_guide/ — vLLM V1 guide, deprecation of V0
  • https://huggingface.co/docs/hub/en/local-apps — Hugging Face local apps flow
  • https://huggingface.co/Qwen/Qwen3-Embedding-8B — Qwen3-Embedding
  • https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ — Gemma 4 launch
  • https://docs.mistral.ai/models — Mistral lineup and licences

Benchmarks & leaderboards

  • https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison/ — reproducible H100 concurrency test
  • https://computingforgeeks.com/open-source-llm-comparison/ — master model comparison from model cards
  • https://presenc.ai/research/local-llm-tokens-per-second-benchmarks-2026 — measured 2026 tok/s
  • https://aclanthology.org/2026.propor-1.7/ — CLARIN-PT-LDB, the European Portuguese leaderboard
  • https://huggingface.co/spaces/PORTULAN/portuguese-llm-leaderboard

Secondary

  • https://codingprotocols.com/blog/local-llm-vram-requirements-quantization — VRAM formula, bytes/param
  • https://oobabooga.github.io/blog/posts/gptq-awq-exl2-llamacpp — quantization head-to-head
  • https://www.promptquorum.com/local-llms/best-local-llms-portuguese-language-2026 — Portuguese model ranking
  • https://www.digitalapplied.com/blog/meta-muse-glimmer-30b-apache-2-local-agent-model-2026 — Muse Glimmer
  • https://www.orcarouter.ai/blog/llama-5-leak — Llama 5 debunk

8. Research trail

Section File Agent Status
Inference engines section-1-inference-engines.md subagent written directly, 40 sources
Open-weight models section-2-open-weight-models.md subagent (died) reconstructed by JUVENAL from its verified sources
Hardware / quantization section-3-hardware-performance.md subagent written directly, 35 sources
Application stack section-4-application-stack.md subagent written (landed seconds before a stop); a replacement was dispatched and also stopped
Licensing / ops section-5-practice-community-ops.md subagent written directly, 36 sources

Known limitations: 3 of 5 original subagents died on stepfun/step-5-preview:free with "provider unresponsive ×6 stale attempts." Section 2 was reconstructed by the parent from the dead subagent's live-extracted sources (all 13 URLs re-verified). Section 4's original write landed 17 seconds before its stop request; a re-dispatch on step-3.7-flash:free was stopped as redundant. The working model was switched to stepfun/step-3.7-flash:free (model.default + delegation.model) after these failures.