Running Open-Weight AI Models Locally — State of the Art
Master report as of October 2026: engine choice is now gated by model architecture, VRAM math is 0.6 GB/billion for Q4_K_M, and everything speaks the OpenAI API shape. Executive summary, verified-vs-uncertain ledger, and seven recommendations.
Date: 2026-10-09 Prepared for: Paulo Prepared by: JUVENAL Method: parallel multi-agent research — 5 independent subagents, one per section, each citing live sources (144 KB across 5 section files in this folder).
1. Executive Summary
- The engine choice has bifurcated into two families, plus a new third. GGUF/C++ engines (llama.cpp and derivatives) win for single-user, CPU, Apple Silicon and edge. Batched GPU serving engines (vLLM, SGLang, TensorRT-LLM) win for concurrency and throughput. A third class emerged in 2026: agent-oriented engines (Magnitude, Splash, vllm.cpp, vllm-metal) built for the reality that the main local workload is now coding agents, not chat.
- The famous "44× between engines" gap is a concurrency artefact. llama.cpp is the fastest of five for a single user (185 tok/s @ batch 1) and fourth of five at 50 concurrent. At batch 1 the engines are near parity. Judge engines on your concurrency, not on headline benchmarks.
- Ollama is the simplest on-ramp but has a hard concurrency ceiling. It defaults to
OLLAMA_NUM_PARALLEL=1, and since Feb 2026 forces a single slot for eleven architecture families. On one H100 it served 33 tok/s at 50 concurrent vs vLLM's 1,610 — a 52× gap, with a 381-second median TTFT. Fine for one developer; wrong for production serving. - Model architecture now gates engine choice — this is the biggest 2026 change. Hybrid linear-attention models (Qwen3.8's Gated DeltaNet) need per-request recurrent state, and every engine sizes it differently: TensorRT-LLM couldn't load it at all, vLLM refused to start, SGLang silently capped itself to 12–20 concurrent requests. Check your specific model against your engine before committing.
- VRAM math: use 0.6 GB per billion parameters for Q4_K_M, not the widely-cited 0.5. K-quants keep sensitive tensors at higher precision, so the effective rate is always above nominal. Verified against real files: 7B → 4.1 GB, 13B → 7.9 GB, 70B → ~40 GB.
- The KV cache can be bigger than the weights. Llama 3.1 70B at 128K context needs ~43 GB of KV per request — more than FP8 weights. It's ~8× worse without GQA. This is what actually kills long-context on consumer cards.
- Best value for a local setup is a used RTX 3090 24 GB (~$900) — runs a 32B at Q4_K_M. An RTX 5090 (32 GB) is the best single-card small-model speed but cannot hold a 70B at Q4. The only consumer route to a 70B dense model at interactive speed is a Mac Studio (M5 Max/Ultra) unified memory.
- Qwen3.8-27B is the standout local model — a dense 27B scoring 77.2% on SWE-bench Verified, beating Alibaba's own 397B MoE from two months earlier, fitting ~24 GB. Apache 2.0. For Portuguese, Qwen3 8B is the top Ollama-native choice; Sabiá-3 is higher quality but HuggingFace-only.
- Everything speaks the OpenAI API shape. llama.cpp, Ollama, vLLM, LM Studio, SGLang, LocalAI all expose
/v1/chat/completionsetc., so an app written once against the OpenAI SDK runs locally by only swappingbase_url. 2026 added Anthropic Messages compatibility so Claude Code can drive open models directly. - Bottom line: for one developer, Ollama or LM Studio. For a coding agent on a Mac, MLX/Splash. For serving to many users, vLLM. Buy VRAM first, model second, and verify engine/model compatibility before you commit.
2. What we verified vs what remains uncertain
| Claim | Status | Source |
|---|---|---|
| Engine versions, star counts, last-commit dates | verified — GitHub API / PyPI | section 1 |
| Concurrency benchmarks (H100, 1/10/50 concurrent) | verified — independent reproducible test | winder.ai |
| Q4_K_M = 0.6 GB/billion params | verified — against real GGUF files | codingprotocols.com |
| KV cache ~43 GB for 70B @ 128K | verified — formula + model cards | section 3 |
| Benchmark scores (SWE-bench etc.) | vendor-reported, not independent | computingforgeeks |
| Muse Glimmer MCP Atlas 75.5 | vendor-reported (Meta's own table) | digitalapplied |
| M5 Ultra 70B = 40–52 tok/s | uncertain — early community measurement, publisher marks it unverified | section 3 |
| Magnitude "92% faster decode than llama.cpp" | unverified vendor claim | section 1 |
| Llama 5 existence | debunked — no official release exists | orcarouter |
| GPT4All maintenance status | verified — last commit 2025-05-27, effectively EOL | section 1 |
Benchmark-version caution: two vendors reporting the "same" benchmark name can differ by 11+ points because they ran different versions. Always check the version suffix before comparing scores.
3. Findings
3.1 Inference engines — what's actually maintained
Full detail: Inference Engines Overview (18 KB, 40 sources)
| Engine | Version (Oct 2026) | Stars | Status | Best for |
|---|---|---|---|---|
| llama.cpp | b11505 | 130,619 | very active | single-user, any hardware, portability |
| Ollama | v0.40.1 | 182,473 | very active | one developer, simplest path |
| vLLM | 0.31.0 | 93,444 | very active | concurrency, production serving |
| SGLang | 0.5.21 | 36,919 | very active | shared prefixes, hybrid models |
| LocalAI | 4.11.0 | 49,450 | very active | one process, LLM+vision+speech |
| LM Studio | 0.4.26 | proprietary | very active | GUI + headless llmster daemon |
| KoboldCpp | v1.122.1 | 11,974 | active | roleplay/creative |
| ExLlamaV3 | 1.5.3 | 1,623 | active, pre-1.0 | max quality-per-bit on RTX |
| ExLlamaV2 | 0.3.2 | 4,632 | legacy/dormant | — |
| TensorRT-LLM | 1.3 | 14,780 | active, slow model support | NVIDIA-only, lowest TTFT |
| MLC-LLM | — | 23,226 | low activity | browser/mobile/edge |
| llamafile | — | 26,209 | slow cadence | single-file distribution |
| GPT4All | — | 77,373 | effectively EOL | do not build on it |
Concurrency, measured on one H100:
| Engine | tok/s @ 1 / 10 / 50 concurrent | Median TTFT |
|---|---|---|
| SGLang 0.5.20 | 81 / 627 / 1,725 | 1.5 s |
| vLLM 0.30.0 | 76 / 589 / 1,610 | 2.1 s |
| llama.cpp | 56 / 138 / 97 | 8.2 s |
| Ollama 0.34.4 | 33 / 33 / 33 | 381 s |
| TensorRT-LLM | could not load (hybrid model) | — |
New in 2026: vllm.cpp (C++20 vLLM port, no Python at inference, but Blackwell-only CUDA images), vllm-metal (vLLM scheduler over MLX — 3.64 s vs mlx_lm's 10.99 s on ragged batches), Magnitude and Splash (agent-oriented, Apple Silicon).
3.2 Open-weight models — what to actually run
Full detail: Open-Weight Model Landscape (13.6 KB, 13 sources)
| Hardware tier | Fits at Q4 |
|---|---|
| 8 GB | Qwen3 8B, Gemma 4 E4B, Phi-4 Mini, gpt-oss-20b, Muse Glimmer |
| 16 GB | Gemma 4 12B, Qwen3 14B, Phi-4, Qwen3.8-Flash-Next (6B active) |
| 24 GB | Qwen3.8-27B, Gemma 4 31B/26B-A4B, Muse Glimmer 30B, Kimi-Linear 48B-A3B |
| 48 GB | Llama 3.3 70B, GLM-5.3-Flash (18B active), DeepSeek V4 Flash 284B/A13B |
| 128 GB+ | Kimi K3, GLM-5.3, Qwen3.8-Max, DeepSeek V4 Pro (server-class) |
Top picks: Qwen3.8-27B (dense, Apache 2.0, 77.2% SWE-bench Verified), DeepSeek V4 Pro (80.6%, top open-weight SWE-bench), Muse Glimmer 30B (Meta's first unmodified Apache 2.0 release, built for local agents), gpt-oss-120b (OpenAI, Apache 2.0). Licensing trap: Qwen3.8-Max, GLM-5.3 and Kimi K3 moved to custom licences in Aug 2026 while their mid-range siblings stayed permissive.
3.3 Hardware, quantization, performance
Full detail: VRAM Math (VRAM math), Quantization Formats (quant formats), Prefill vs Decode (speed benchmarks)
total VRAM = weights + KV cache + overhead
weights (GB) = parameters (B) × bytes-per-param → use 0.6 for Q4_K_M
KV cache (GB) ≈ 2 × layers × kv_heads × head_dim × context_len × bytes / 1e9
overhead ≈ 0.5–1.5 GB
- Quantization measured (RTX 3090, 13B): EXL2 4.9bpw best quality (4.308 ppl) at 52 tok/s; GGUF Q4_K_M is 4.333 ppl but only 30.8 tok/s — ~40% throughput paid for portability. bnb NF4: slowest, and despite being "4-bit" stored 24.8 GB on disk. Use NF4 for QLoRA training only, never serving.
- Prefill vs decode is a 50–100× gap. RTX 5090: ~14,000–15,000 tok/s prefill vs 290–300 tok/s decode. Prefill is compute-bound, decode is bandwidth-bound (
tok/s ≈ bandwidth ÷ model size). For RAG/agent workloads prefill is the bottleneck. - Thresholds: >30 tok/s interactive, 8–30 batch, <8 impractical.
- AMD is now viable for inference: ROCm 7.2 (Mar 2026) first with official RDNA 4 support and claimed out-of-box CUDA parity; RX 7900 XTX runs Llama 3.1 8B at ~96 tok/s ≈ 75% of a 4090. But Vulkan remains the pragmatic default for local llama.cpp, and silent CPU-fallback is the classic AMD trap — verify in logs, not device queries.
- Apple Silicon: MLX is the native path; M5 Neural Accelerators give up to 3.97× faster TTFT. Bandwidth is the ceiling and it favours big models — M4 Max beat NVIDIA GB10 and AMD Strix Halo in decode despite lower absolute bandwidth, because it has far more memory.
- CPU-only is bandwidth-bound, not core-bound — fast DDR5 beats more cores. ~12 tok/s class for a 3–4B at Q4.
3.4 The application stack
Full detail: Application Stack Overview (22 KB, 50 sources)
- Serving: every runtime exposes OpenAI-compatible endpoints. Point
OPENAI_BASE_URLatlocalhost:11434/v1(Ollama) orllama-serverwith a dummy key. LiteLLM Proxy fronts 100+ providers. - Coding agents: for local use you need a strong 30B+ quantised model on vLLM for usable function calling; smaller models degrade fast. OpenCode + vLLM is the most current combination.
- RAG: Qwen3-Embedding is the open-weight leader (8B ranks #1 on MTEB multilingual; the 0.6B is a strong cheap local choice). Pattern: Qwen3-Embedding → Qdrant/FAISS → Qwen3-Reranker → LLM on vLLM via LlamaIndex. Note: an
instructprompt typically improves retrieval 1–5%. Milvus's PyPI package is stale (2024) — use Qdrant. - Fine-tuning: Unsloth (v2026.10.3) is the fastest single-GPU LoRA/QLoRA path and ships a desktop app; Axolotl (v0.20.0) for config-driven multi-GPU/FSDP2.
- Local speech/image: faster-whisper for ASR, Kokoro/Piper for TTS, ComfyUI + FLUX for images.
3.5 Licensing, security, operations
Full detail: Practice & Ops — Licensing (31 KB, 36 sources)
- Licensing: only Apache-2.0/MIT weights are genuinely OSI open source (Qwen3, DeepSeek, gpt-oss, Gemma 4, Mistral Small/Large 3, OLMo 2, Phi-4, Muse Glimmer). Llama 4 = free under 700M MAU + "Built with Llama" attribution + EU exclusion on multimodal. Gemma 3 carries a Google-enforceable Prohibited Use Policy. Downloading locally does not exempt you from acceptable-use terms.
- EU AI Act (Oct 2026): the Digital Omnibus (Reg. 2026/1744, in force 27 Jul 2026) delayed high-risk rules to Dec 2027/Aug 2028 but did not delay GPAI or Article 50 transparency. The open-source exemption (Art. 53(2)/54(6)) is void for systemic-risk models (>10²⁵ FLOP) and never covers copyright/training-data duties. Fine-tuning or materially modifying a model and putting it on the EU market likely makes you a provider. Local inference is a strong GDPR lever — it removes the Art. 28 data-processor obligation entirely.
- Security: local ≠ secure. SafeTensors eliminates the pickle RCE vector;
trust_remote_codeis the second; HF pulled 9 malicious models in Q1 2026, 3 already in production. LM Studio ships telemetry on by default. Never setOLLAMA_HOST=0.0.0.0. - Drivers: the NVIDIA driver is no longer bundled with the CUDA toolkit on Linux since 13.4. ROCm 7.2.x is the first release with official RDNA 4 support. Vulkan is the cross-vendor fallback, not a library replacement.
- Cost: breakeven is workload-specific. A 7B on a Mac Studio vs frontier APIs breaks even in ~0.3 months; a dense 70B on DGX Spark (4.4 tps) never beats cheap open-weight APIs. Below 10% utilisation, cloud wins.
4. Choosing an engine
| Need | Choice |
|---|---|
| One user, any hardware, simplest | Ollama |
| Single-user max throughput, CPU/edge | llama.cpp |
| GUI + model discovery, or headless daemon | LM Studio |
| Concurrency / production serving | vLLM |
| Shared prefixes, hybrid models | SGLang |
| Max quality-per-bit, one modern GPU | ExLlamaV3 / EXL3 |
| Apple Silicon, ≥14B | MLX |
| Apple Silicon, concurrent agents | vllm-metal or Splash |
| Local coding agents, long sessions | Magnitude / Splash / LM Studio |
| One process, LLM + vision + speech | LocalAI |
| Offline ChatGPT alternative | Jan |
| Browser / mobile / edge | MLC-LLM / WebLLM |
| CPU-only legacy box | Ollama or Jan (not GPT4All) |
5. Recommendations
- Start with Ollama, move to vLLM if you need concurrency. Ollama for one developer; vLLM the moment more than one request at a time matters. The 52× gap is real.
- Buy VRAM first. A used RTX 3090 24 GB (~$900) is the best value entry and runs a 32B at Q4. Avoid 8 GB cards. If your models are large and dense, a Mac Studio is the only consumer route to 70B at interactive speed.
- Default to GGUF Q4_K_M for portability, EXL3 if you want speed on one RTX. You pay ~40% throughput for GGUF's portability; take it when the model moves between machines, skip it when it doesn't.
- Size for the KV cache, not just the weights. Budget ~2× the model requirement on unified-memory machines, and check context length against KV math before committing to 128K.
- Use Qwen3.8-27B as the default local workhorse — dense, Apache 2.0, beats much larger MoE models, fits 24 GB. For Portuguese, Qwen3 8B via Ollama, Sabiá-3 if quality is paramount.
- Verify engine/model compatibility before adopting any hybrid linear-attention model. This eliminated TensorRT-LLM and crippled two other engines in Sept 2026 testing.
- Check the licence, not the word "open." Apache-2.0/MIT only for commercial work; Llama/Gemma/GLM-5.3/Qwen3.8-Max/Kimi K3 all carry real restrictions, and fine-tuning + EU market placement can make you a regulated provider.
6. Open questions
- Does Magnitude actually deliver its claimed 92% decode speedup? Vendor-reported and unverified — needs an independent reproduction.
- M5 Ultra 70B figures (40–52 tok/s) are early community measurements the publisher marks unverified — needs confirmation.
- Will ROCm reach true CUDA parity for training, not just inference? Inference is viable now; training tooling still lags.
- Will Llama-family weights stay openly released? Meta has signalled a shift toward closed releases (Muse Spark / Superintelligence Labs, Apr 2026).
7. Sources
Consolidated from the five section files (~170 distinct sources). The authoritative per-topic lists are in each section file; the highest-value primary sources:
Primary
- https://github.com/ggml-org/llama.cpp — llama.cpp repo, build docs, multi-GPU docs
- https://github.com/ollama/ollama + https://docs.ollama.com/gpu — Ollama, backend support, OpenAI compat
- https://docs.vllm.ai/en/stable/usage/v1_guide/ — vLLM V1 guide, deprecation of V0
- https://huggingface.co/docs/hub/en/local-apps — Hugging Face local apps flow
- https://huggingface.co/Qwen/Qwen3-Embedding-8B — Qwen3-Embedding
- https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ — Gemma 4 launch
- https://docs.mistral.ai/models — Mistral lineup and licences
Benchmarks & leaderboards
- https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison/ — reproducible H100 concurrency test
- https://computingforgeeks.com/open-source-llm-comparison/ — master model comparison from model cards
- https://presenc.ai/research/local-llm-tokens-per-second-benchmarks-2026 — measured 2026 tok/s
- https://aclanthology.org/2026.propor-1.7/ — CLARIN-PT-LDB, the European Portuguese leaderboard
- https://huggingface.co/spaces/PORTULAN/portuguese-llm-leaderboard
Secondary
- https://codingprotocols.com/blog/local-llm-vram-requirements-quantization — VRAM formula, bytes/param
- https://oobabooga.github.io/blog/posts/gptq-awq-exl2-llamacpp — quantization head-to-head
- https://www.promptquorum.com/local-llms/best-local-llms-portuguese-language-2026 — Portuguese model ranking
- https://www.digitalapplied.com/blog/meta-muse-glimmer-30b-apache-2-local-agent-model-2026 — Muse Glimmer
- https://www.orcarouter.ai/blog/llama-5-leak — Llama 5 debunk
8. Research trail
| Section | File | Agent | Status |
|---|---|---|---|
| Inference engines | section-1-inference-engines.md | subagent | written directly, 40 sources |
| Open-weight models | section-2-open-weight-models.md | subagent (died) | reconstructed by JUVENAL from its verified sources |
| Hardware / quantization | section-3-hardware-performance.md | subagent | written directly, 35 sources |
| Application stack | section-4-application-stack.md | subagent | written (landed seconds before a stop); a replacement was dispatched and also stopped |
| Licensing / ops | section-5-practice-community-ops.md | subagent | written directly, 36 sources |
Known limitations: 3 of 5 original subagents died on stepfun/step-5-preview:free with "provider unresponsive ×6 stale attempts." Section 2 was reconstructed by the parent from the dead subagent's live-extracted sources (all 13 URLs re-verified). Section 4's original write landed 17 seconds before its stop request; a re-dispatch on step-3.7-flash:free was stopped as redundant. The working model was switched to stepfun/step-3.7-flash:free (model.default + delegation.model) after these failures.