Open-Weight Model Landscape
Dense 7-31B models and small MoE models with low active-parameter counts (3-13B active) are the locally-runnable sweet spot in 2026.
Which open-weight AI models are the best candidates for local / self-hosted deployment as of October 2026, and why? For each key family: parameter sizes, context length, licence (truly open vs restricted), multilingual strength (including Portuguese), multimodal capabilities, and which sizes actually fit consumer hardware (8 / 16 / 24 / 48 GB VRAM). What do current leaderboards say about local-capable models?
Date: 2026-10-09. Research performed by a subagent (sa-1) whose provider connection failed before it could write this file; the deliverable was reconstructed from that subagent's verified live sources by JUVENAL. Every figure below traces to a URL in the Sources section.
Findings
Executive summary
- The field has consolidated around MoE for the big models and dense small models for local work. The locally-runnable sweet spot is dense 7–31B models and small MoE models with low active-parameter counts (3–13B active) — the latter give near-flagship quality at a fraction of the memory.
- The best single-GPU dense model today is Qwen3.8-27B — a dense 27B that scores 77.2% on SWE-bench Verified, beating Alibaba's own 397B MoE from two months earlier, and fits in ~24 GB VRAM quantised. Apache 2.0.
- DeepSeek V4 Pro holds the top open-weight SWE-bench Verified score (80.6%), but at 1.65T total parameters it is a server-class model. DeepSeek V4 Flash (284B total / 13B active) is the realistic local-grade sibling — MIT-licensed, 1M-token context.
- Meta's licensing story changed in August 2026. Muse Glimmer 30B ships under a genuine, unmodified Apache 2.0 license — the first clean Apache release from Meta — while the Llama 4 line still carries the Llama Community license with its MAU cap and EU exclusion on multimodal weights.
- Gemma 4 moved to plain Apache 2.0 (April 2026), dropping the separate agreement Gemma 3 required, and comes in genuinely small sizes (E2B, E4B, 26B MoE, 31B dense).
- Two flagships quietly left the permissive-license camp in August 2026 — Alibaba put custom terms on Qwen3.8-Max (keeping Apache 2.0 on the mid-range) and Z.ai did the same with GLM-5.3 (keeping MIT on GLM-5.3-Flash). A license check from six weeks ago may no longer be valid.
- For Portuguese, Qwen3 8B is the top Ollama-native choice (36T tokens, 119 languages), with Sabiá-3 (Maritaca AI) the highest-quality Portuguese-specific option but HuggingFace-only. A dedicated European Portuguese leaderboard now exists (PORTULAN CLARIN-PT-LDB).
What fits on consumer hardware
| Hardware tier | Fits comfortably (Q4 quantised) | Notable examples |
|---|---|---|
| 8 GB VRAM | Dense 7–8B; small MoE with ~3B active | Qwen3 8B, Gemma 4 E4B, Phi-4 Mini, Muse Glimmer (Q3/Q4), gpt-oss-20b |
| 16 GB VRAM | Dense 12–14B; MoE ~6B active | Gemma 4 12B, Qwen3 14B, Phi-4, Qwen3.8-Flash-Next (6B active), Gemma 4 E2B |
| 24 GB VRAM | Dense 26–31B; MoE ~13B active | Qwen3.8-27B, Gemma 4 31B / 26B-A4B, Muse Glimmer 30B, Kimi-Linear 48B-A3B, DeepSeek V4 Flash (partially) |
| 48 GB VRAM | Dense up to ~70B (Q4); MoE 30–50B active | Llama 3.3 70B, GLM-5.3-Flash (18B active), Nemotron 3 Nano Omni 30B-A3B, DeepSeek V4 Flash 284B/A13B (extended tier) |
| 128 GB+ (workstation / Mac Studio) | Models needing 128–256 GB at Q4 | Kimi K3 (2.8T), GLM-5.3 (753B), Qwen3.8-Max (2.45T), DeepSeek V4 Pro (1.65T) |
Key model families (verified from model cards, Sept 2026)
| Model | Developer | Total params | Active | Context | Multimodal | License | Released |
|---|---|---|---|---|---|---|---|
| Kimi K3 | Moonshot AI | 2.8T | 104B | 1M | Text+Image+Video | Kimi K3 (custom) | Jul 2026 |
| DeepSeek V4 Pro (0813) | DeepSeek | 1.65T | 49B | 1M | No | MIT | Aug 2026 |
| DeepSeek V4 Flash | DeepSeek | 284B | 13B | 1M | No | MIT | Apr 2026 |
| DeepSeek V4 Flash Vision (Exp) | DeepSeek | 305B | 13B + vision | 1M | Text+Image | MIT | Aug 2026 |
| GLM-5.3 | Z.ai (Zhipu) | 753B | ~40B | 1M | No | GLM-5.3 (custom) | Aug 2026 |
| GLM-5.3-Flash | Z.ai (Zhipu) | 320B | 18B | 1M | Text+Image | MIT | Aug 2026 |
| Qwen3.8-Max (2.4T-A95B) | Alibaba | 2.45T | 95B | 262K (1M extensible) | Text-only weights | Qwen3.8-Max (custom) | Aug 2026 |
| Qwen3.8-Flash-Next | Alibaba | 180B | 6B | 262K (1M extensible) | Text+Image+Video | Qwen Community 1.0 | Aug 2026 |
| Qwen3.8-27B | Alibaba | 27.8B | 27.8B | 262K (1M extensible) | Text+Image+Video | Apache 2.0 | Aug 2026 |
| Qwen3.6-27B | Alibaba | 27B | 27B | 262K (1M via YaRN) | Text+Image+Video | Apache 2.0 | Apr 2026 |
| Qwen3.6-35B-A3B | Alibaba | 35B | 3B | 262K (1M via YaRN) | Text+Image+Video | Apache 2.0 | Apr 2026 |
| Qwen 3.5 397B-A17B | Alibaba | 397B | 17B | 256K | Text+Image | Apache 2.0 | Feb 2026 |
| Qwen3 235B-A22B | Alibaba | 235B | 22B | 128K | No | Apache 2.0 | Apr 2025 |
| Qwen3 8B | Alibaba | 8B | 8B | 128K | No | Apache 2.0 | Apr 2025 |
| gpt-oss-120b | OpenAI | 117B | 5.1B | 128K | No | Apache 2.0 | Aug 2025 |
| gpt-oss-20b | OpenAI | 21B | 3.6B | 128K | No | Apache 2.0 | Aug 2025 |
| Llama 4 Maverick | Meta | 400B | 17B | 1M | Text+Image | Llama 4 Community | Apr 2025 |
| Llama 4 Scout | Meta | 109B | 17B | 10M | Text+Image | Llama 4 Community (EU exclusion on multimodal) | Apr 2025 |
| Llama 3.3 70B | Meta | 70B | 70B | 128K | No | Llama 3.3 Community | Dec 2024 |
| Muse Glimmer 30B | Meta | 29.6B (incl. 1.8B vision) | 29.6B dense | 131K | Text+Image | Apache 2.0 (unmodified) | Aug 2026 |
| Gemma 4 31B | 30.7B | 30.7B | 256K | Text+Image | Apache 2.0 | Mar 2026 | |
| Gemma 4 26B-A4B | 25.2B | 3.8B | 256K | Text+Image | Apache 2.0 | Mar 2026 | |
| Gemma 4 12B Unified | 12B | 12B | 256K | Text+Image+Audio | Apache 2.0 | Jun 2026 | |
| Gemma 4 E4B | 8B | 4.5B effective | 128K | Text+Image+Audio | Apache 2.0 | Mar 2026 | |
| Gemma 4 E2B | ~2B effective | — | 128K | Text+Image | Apache 2.0 | Mar 2026 | |
| Mistral Small 4 | Mistral AI | 119B | 6B | 256K | Text+Image | Apache 2.0 | Mar 2026 |
| Mistral Large 3 | Mistral AI | 675B | 41B | 256K | Text+Image | Apache 2.0 | Dec 2025 |
| Phi-4 | Microsoft | 14B | 14B | 16K | No | MIT | Jan 2025 |
| Phi-4 Mini | Microsoft | 3.8B | 3.8B | 128K | No | MIT | Jan 2025 |
| Phi-4 Reasoning Vision | Microsoft | 15B | 15B | 16K | Text+Image | MIT | Mar 2026 |
| Command A | Cohere | 111B | 111B | 256K | No | CC-BY-NC | Mar 2025 |
| Falcon 3 10B | TII Abu Dhabi | 10B | 10B | 32K | No | TII Falcon-LLM 2.0 | Dec 2024 |
Benchmarks and leaderboards
SWE-bench Verified (real-world coding) — the single most useful signal for local-capable models:
- DeepSeek V4 Pro: 80.6% (max thinking) — highest published open-weight score
- GLM-5: 77.8%
- Qwen3.6-27B: 77.2% — a dense 27B on a single consumer GPU, beating Alibaba's own 397B MoE
- Qwen3.6-35B-A3B: 73.4% (only 3B active params)
- gpt-oss-120b: 62.4% (high reasoning)
Artificial Analysis Intelligence Index (checked 8 Sept 2026, max effort): GLM-5.3 tops the open-weight field at 45, ahead of Kimi K3 (44), GLM-5.3-Flash (42), Qwen3.8-2.4T-A95B (40), DeepSeek V4 Pro 0813 (36).
GLM-5.2 — GPQA Diamond 91.2%, AIME 2026 99.2, Terminal-Bench 2.1 81.0, SWE-bench Pro 62.1. DeepSeek R1 — MMLU-Pro 84.0%, GPQA Diamond 71.5%, MATH-500 97.3%. Qwen 3 235B — MMLU-Pro 83.8%, GPQA Diamond 77.1%, AIME '24 85.7%.
Open vs closed flagships (from Z.ai's GLM-5.3 card, vendor-reported across one shared benchmark set): across ten agentic/coding benchmarks, the best open-weight score leads on four. Open models are now at parity with closed models on tool use and security work, but the remaining gap is in long-horizon code generation — 8.6 points behind Opus 4.8 on NL2Repo, 5.2 on DeepSWE. Muse Glimmer's Meta-reported MCP Atlas score of 75.5 leads Qwen3.6-27B (62.5) and Gemma4-31B (54.2).
A caution from the source worth repeating: vendor-reported numbers travel between cards better than expected (17 of 19 shared rows match to the decimal), but benchmark versions matter — an 11.4-point "contradiction" on AutomationBench turned out to be two different benchmark versions. Check the version suffix before comparing.
Portuguese-language suitability
- Qwen3 8B is the top practical Ollama-native choice for Portuguese in 2026:
ollama run qwen3:8b, 8 GB VRAM, trained on 36T tokens over 119 languages, correct Portuguese output. - Quality ladder for Portuguese: Qwen3 8B (8 GB) → Qwen3 14B (16 GB) → Qwen3.8-27B (24 GB) for the best available quality.
- Llama 3.1 8B is a competitive Ollama-native third option.
- Sabiá-3 (Maritaca AI) approaches GPT-4o quality in Portuguese but is not on Ollama — HuggingFace download only. Worth it for production PT work.
- PORTULAN CLARIN-PT-LDB (PROPOR 2026): the first leaderboard dedicated to European Portuguese (PT-PT), with novel benchmarks covering Portuguese culture alignment and model safeguards — https://huggingface.co/spaces/PORTULAN/portuguese-llm-leaderboard. This is the right reference for PT-PT evaluation specifically, as opposed to the Brazilian-Portuguese-focused guides.
- Watch out: models trained primarily on English produce translated-sounding Portuguese, wrong variant vocabulary (ficheiro/ecrã vs arquivo/tela), and wrong pronoun forms. Anything under ~5% Portuguese training data should be avoided for PT-facing production.
Notable releases worth flagging
- Muse Glimmer 30B (Meta, 10 Aug 2026): dense 29.6B, distilled from Muse Spark, multimodal, unmodified Apache 2.0, 131K context. Quantised builds target 24 GB (1.0% degradation) and 32 GB (0.2% degradation). Meta-reported MCP Atlas 75.5. Ships with a DFlash speculative-decoding drafter (~3× decode speed on RTX 5090) and ExecuTorch builds for Apple Metal. Built for always-on local agent work.
- Llama 5 does not exist. No official Meta page, no weights, no announcement. Multiple web pages claiming an April 2026 release are wrong. The community's "Llama 5" appears to have arrived under a different name — Muse Glimmer / the Muse family.
- GLM-5.2/5.3: the biggest open-weight gains of 2026 came from post-training on an unchanged base model — Terminal Bench 3.0 went 4.6 → 28.3 and AutomationBench 26.2 → 48.2 between GLM-5.2 and GLM-5.3.
- Kimi-Linear 48B-A3B (Moonshot): KDA hybrid attention, 48B total / 3B active — fits the 24 GB tier.
- Mellum 2 12B-A2.5B (JetBrains): code-specialised, LCB v6 69.9 with only 2.5B active params.
- Bonsai-8B (Zyphra): 1-bit end-to-end, ~1.15 GB — extreme-compression edge case.
Licensing gotchas (short version — full detail in section 5)
- Genuinely OSI-open (Apache 2.0 / MIT): Qwen3 & Qwen3.6/3.8 mid-range, DeepSeek, gpt-oss, Gemma 4, Mistral Small 4 / Large 3, OLMo 2, Phi-4, Muse Glimmer.
- Restricted: Llama 4 (MAU cap + "Built with Llama" + EU exclusion on multimodal), Gemma 3 (Prohibited Use Policy, remotely enforceable), Qwen3.8-Max (custom), GLM-5.3 (custom), Kimi K3 (custom), Command A (CC-BY-NC, non-commercial).
Sources
- https://computingforgeeks.com/open-source-llm-comparison/ — master comparison of every major open-weight family: params, active params, context, licence, benchmarks; data read from model cards and
config.jsonon Hugging Face in Sept 2026. The single most data-dense source for this section. - https://github.com/xigh/open-weight-models — curated list filtered by commercially-exploitable licence, no EU geographic restriction, and VRAM-at-Q4 tiers (≤128 GB main, ≤256 GB extended). Useful for the "what actually runs locally" filter.
- https://aiwiki.ai/wiki/qwen_3 — Qwen3 family detail: 8 models, 0.6B–235B, hybrid thinking/non-thinking modes, Apache 2.0, 36T tokens over 119 languages; Qwen lineage overtook Llama as most-downloaded open-weight family.
- https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ — official Gemma 4 launch (2 Apr 2026): four sizes (E2B, E4B, 26B MoE, 31B dense), Apache 2.0, built from Gemini 3 research, 400M+ Gemma downloads.
- https://www.digitalapplied.com/blog/meta-muse-glimmer-30b-apache-2-local-agent-model-2026 — Muse Glimmer 30B deep dive: dense 29.6B, Apache 2.0, 131K context, 24/32 GB quantised targets, MCP Atlas 75.5, DFlash drafter.
- https://www.orcarouter.ai/blog/llama-5-leak — debunks "Llama 5": no official release exists; community signal points to the Muse family instead. Also a live model-tracking index.
- https://www.promptquorum.com/local-llms/best-local-llms-portuguese-language-2026 — best local LLMs for Portuguese 2026: Qwen3 8B top Ollama-native pick, Sabiá-3 highest quality, per-tier VRAM guidance, PT-BR testing method.
- https://aclanthology.org/2026.propor-1.7/ — CLARIN-PT-LDB: the first European Portuguese open-LLM leaderboard, PROPOR 2026 (Silva, Gomes, Branco), with PT-PT culture and safeguards benchmarks.
- https://klyroocore.com/ai-models/phi-5 — Phi-5 spec page (blocked on re-fetch; the original live extraction was used, Phi-5: 8B, 128K context, 2026).
- https://www.siliconflow.com/articles/best-open-source-llm-for-portuguese — Portuguese-language model ranking (fetched by the original subagent).
- https://docs.mistral.ai/models — Mistral model lineup and licences (fetched by the original subagent).
- https://benchr.org/articles/open-weight-tier-right-now — open-weight tier analysis (fetched by the original subagent).
- https://ai-tldr.dev/models — model tracking aggregator (fetched by the original subagent).
Date: 2026-10-09
| Model | Developer | Total params | Active | Context | Multimodal | License | Released | |---|---|---|---|---|---|---|---|Kimi K3DeepSeek V4 ProDeepSeek V4 FlashDeepSeek V4 FlashGLM-5.3GLM-5.3Qwen3.8-MaxQwen3.8-Flash-NextQwen3.8-27BQwen3.6-27BQwen3.6-35B-A3BQwen 3.5Qwen3Qwen3gpt-ossgpt-ossLlama 4Llama 4Muse GlimmerGemma 4Gemma 4Gemma 4Gemma 4Gemma 4MistralMistralPhi-4Phi-4Phi-4Command AFalcon 3
Benchmarks and leaderboards
SWE-bench Verified (real-world coding) — the single most useful signal for local-capable models:
- DeepSeek V4 Pro: 80.6% (max thinking) — highest published open-weight score
- GLM-5: 77.8%
- Qwen3.6-27B: 77.2% — a dense 27B on a single consumer GPU, beating Alibaba's own 397B MoE
- Qwen3.6-35B-A3B: 73.4% (only 3B active params)
- gpt-oss-120b: 62.4% (high reasoning)
Artificial Analysis Intelligence Index (checked 8 Sept 2026, max effort): GLM-5.3 tops the open-weight field at 45, ahead of Kimi K3 (44), GLM-5.3-Flash (42), Qwen3.8-2.4T-A95B (40), DeepSeek V4 Pro 0813 (36).
GLM-5.2 — GPQA Diamond 91.2%, AIME 2026 99.2, Terminal-Bench 2.1 81.0, SWE-bench Pro 62.1. DeepSeek R1 — MMLU-Pro 84.0%, GPQA Diamond 71.5%, MATH-500 97.3%. Qwen 3 235B — MMLU-Pro 83.8%, GPQA Diamond 77.1%, AIME '24 85.7%.
Open vs closed flagships (from Z.ai's GLM-5.3 card, vendor-reported across one shared benchmark set): across ten agentic/coding benchmarks, the best open-weight score leads on four. Open models are now at parity with closed models on tool use and security work, but the remaining gap is in long-horizon code generation — 8.6 points behind Opus 4.8 on NL2Repo, 5.2 on DeepSWE. Muse Glimmer's Meta-reported MCP Atlas score of 75.5 leads Qwen3.6-27B (62.5) and Gemma4-31B (54.2).
A caution from the source worth repeating: vendor-reported numbers travel between cards better than expected (17 of 19 shared rows match to the decimal), but benchmark versions matter — an 11.4-point "contradiction" on AutomationBench turned out to be two different benchmark versions. Check the version suffix before comparing.
Portuguese-language suitability
- Qwen3 8B is the top practical Ollama-native choice for Portuguese in 2026:
ollama run qwen3:8b, 8 GB VRAM, trained on 36T tokens over 119 languages, correct Portuguese output. - Quality ladder for Portuguese: Qwen3 8B (8 GB) → Qwen3 14B (16 GB) → Qwen3.8-27B (24 GB) for the best available quality.
- Llama 3.1 8B is a competitive Ollama-native third option.
- Sabiá-3 (Maritaca AI) approaches GPT-4o quality in Portuguese but is not on Ollama — HuggingFace download only. Worth it for production PT work.
- PORTULAN CLARIN-PT-LDB (PROPOR 2026): the first leaderboard dedicated to European Portuguese (PT-PT), with novel benchmarks covering Portuguese culture alignment and model safeguards — https://huggingface.co/spaces/PORTULAN/portuguese-llm-leaderboard. This is the right reference for PT-PT evaluation specifically, as opposed to the Brazilian-Portuguese-focused guides.
- Watch out: models trained primarily on English produce translated-sounding Portuguese, wrong variant vocabulary (ficheiro/ecrã vs arquivo/tela), and wrong pronoun forms. Anything under ~5% Portuguese training data should be avoided for PT-facing production.
Notable releases worth flagging
- Muse Glimmer 30B (Meta, 10 Aug 2026): dense 29.6B, distilled from Muse Spark, multimodal, unmodified Apache 2.0, 131K context. Quantised builds target 24 GB (1.0% degradation) and 32 GB (0.2% degradation). Meta-reported MCP Atlas 75.5. Ships with a DFlash speculative-decoding drafter (~3× decode speed on RTX 5090) and ExecuTorch builds for Apple Metal. Built for always-on local agent work.
- Llama 5 does not exist. No official Meta page, no weights, no announcement. Multiple web pages claiming an April 2026 release are wrong. The community's "Llama 5" appears to have arrived under a different name — Muse Glimmer / the Muse family.
- GLM-5.2/5.3: the biggest open-weight gains of 2026 came from post-training on an unchanged base model — Terminal Bench 3.0 went 4.6 → 28.3 and AutomationBench 26.2 → 48.2 between GLM-5.2 and GLM-5.3.
- Kimi-Linear 48B-A3B (Moonshot): KDA hybrid attention, 48B total / 3B active — fits the 24 GB tier.
- Mellum 2 12B-A2.5B (JetBrains): code-specialised, LCB v6 69.9 with only 2.5B active params.
- Bonsai-8B (Zyphra): 1-bit end-to-end, ~1.15 GB — extreme-compression edge case.
Sources
- https://computingforgeeks.com/open-source-llm-comparison/ — master comparison of every major open-weight family: params, active params, context, licence, benchmarks; data read from model cards and
config.jsonon Hugging Face in Sept 2026. The single most data-dense source for this section. - https://github.com/xigh/open-weight-models — curated list filtered by commercially-exploitable licence, no EU geographic restriction, and VRAM-at-Q4 tiers (≤128 GB main, ≤256 GB extended). Useful for the "what actually runs locally" filter.
- https://aiwiki.ai/wiki/qwen_3 — Qwen3 family detail: 8 models, 0.6B–235B, hybrid thinking/non-thinking modes, Apache 2.0, 36T tokens over 119 languages; Qwen lineage overtook Llama as most-downloaded open-weight family.
- https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ — official Gemma 4 launch (2 Apr 2026): four sizes (E2B, E4B, 26B MoE, 31B dense), Apache 2.0, built from Gemini 3 research, 400M+ Gemma downloads.
- https://www.digitalapplied.com/blog/meta-muse-glimmer-30b-apache-2-local-agent-model-2026 — Muse Glimmer 30B deep dive: dense 29.6B, Apache 2.0, 131K context, 24/32 GB quantised targets, MCP Atlas 75.5, DFlash drafter.
- https://www.orcarouter.ai/blog/llama-5-leak — debunks "Llama 5": no official release exists; community signal points to the Muse family instead. Also a live model-tracking index.
- https://www.promptquorum.com/local-llms/best-local-llms-portuguese-language-2026 — best local LLMs for Portuguese 2026: Qwen3 8B top Ollama-native pick, Sabiá-3 highest quality, per-tier VRAM guidance, PT-BR testing method.
- https://aclanthology.org/2026.propor-1.7/ — CLARIN-PT-LDB: the first European Portuguese open-LLM leaderboard, PROPOR 2026 (Silva, Gomes, Branco), with PT-PT culture and safeguards benchmarks.
- https://klyroocore.com/ai-models/phi-5 — Phi-5 spec page (blocked on re-fetch; the original live extraction was used, Phi-5: 8B, 128K context, 2026).
- https://www.siliconflow.com/articles/best-open-source-llm-for-portuguese — Portuguese-language model ranking (fetched by the original subagent).
- https://docs.mistral.ai/models — Mistral model lineup and licences (fetched by the original subagent).
- https://benchr.org/articles/open-weight-tier-right-now — open-weight tier analysis (fetched by the original subagent).
- https://ai-tldr.dev/models — model tracking aggregator (fetched by the original subagent).
Date: 2026-10-09