Open-Weight Model Landscape - Local AI Agent Wiki __md_scope=new URL("../../..",location),__md_hash=e=>[...e].reduce(((e,_)=>(e<<5)-e+_.charCodeAt(0)),0),__md_get=(e,_=localStorage,t=__md_scope)=>JSON.parse(_.getItem(t.pathname+"."+e)),__md_set=(e,_,t=localStorage,a=__md_scope)=>{try{t.setItem(a.pathname+"."+e,JSON.stringify(_))}catch(e){}}
Skip to content

Open-Weight Model Landscape

Dense 7-31B models and small MoE models with low active-parameter counts (3-13B active) are the locally-runnable sweet spot in 2026.

Which open-weight AI models are the best candidates for local / self-hosted deployment as of October 2026, and why? For each key family: parameter sizes, context length, licence (truly open vs restricted), multilingual strength (including Portuguese), multimodal capabilities, and which sizes actually fit consumer hardware (8 / 16 / 24 / 48 GB VRAM). What do current leaderboards say about local-capable models?

Date: 2026-10-09. Research performed by a subagent (sa-1) whose provider connection failed before it could write this file; the deliverable was reconstructed from that subagent's verified live sources by JUVENAL. Every figure below traces to a URL in the Sources section.

Findings

Executive summary

  1. The field has consolidated around MoE for the big models and dense small models for local work. The locally-runnable sweet spot is dense 7–31B models and small MoE models with low active-parameter counts (3–13B active) — the latter give near-flagship quality at a fraction of the memory.
  2. The best single-GPU dense model today is Qwen3.8-27B — a dense 27B that scores 77.2% on SWE-bench Verified, beating Alibaba's own 397B MoE from two months earlier, and fits in ~24 GB VRAM quantised. Apache 2.0.
  3. DeepSeek V4 Pro holds the top open-weight SWE-bench Verified score (80.6%), but at 1.65T total parameters it is a server-class model. DeepSeek V4 Flash (284B total / 13B active) is the realistic local-grade sibling — MIT-licensed, 1M-token context.
  4. Meta's licensing story changed in August 2026. Muse Glimmer 30B ships under a genuine, unmodified Apache 2.0 license — the first clean Apache release from Meta — while the Llama 4 line still carries the Llama Community license with its MAU cap and EU exclusion on multimodal weights.
  5. Gemma 4 moved to plain Apache 2.0 (April 2026), dropping the separate agreement Gemma 3 required, and comes in genuinely small sizes (E2B, E4B, 26B MoE, 31B dense).
  6. Two flagships quietly left the permissive-license camp in August 2026 — Alibaba put custom terms on Qwen3.8-Max (keeping Apache 2.0 on the mid-range) and Z.ai did the same with GLM-5.3 (keeping MIT on GLM-5.3-Flash). A license check from six weeks ago may no longer be valid.
  7. For Portuguese, Qwen3 8B is the top Ollama-native choice (36T tokens, 119 languages), with Sabiá-3 (Maritaca AI) the highest-quality Portuguese-specific option but HuggingFace-only. A dedicated European Portuguese leaderboard now exists (PORTULAN CLARIN-PT-LDB).

What fits on consumer hardware

Hardware tier Fits comfortably (Q4 quantised) Notable examples
8 GB VRAM Dense 7–8B; small MoE with ~3B active Qwen3 8B, Gemma 4 E4B, Phi-4 Mini, Muse Glimmer (Q3/Q4), gpt-oss-20b
16 GB VRAM Dense 12–14B; MoE ~6B active Gemma 4 12B, Qwen3 14B, Phi-4, Qwen3.8-Flash-Next (6B active), Gemma 4 E2B
24 GB VRAM Dense 26–31B; MoE ~13B active Qwen3.8-27B, Gemma 4 31B / 26B-A4B, Muse Glimmer 30B, Kimi-Linear 48B-A3B, DeepSeek V4 Flash (partially)
48 GB VRAM Dense up to ~70B (Q4); MoE 30–50B active Llama 3.3 70B, GLM-5.3-Flash (18B active), Nemotron 3 Nano Omni 30B-A3B, DeepSeek V4 Flash 284B/A13B (extended tier)
128 GB+ (workstation / Mac Studio) Models needing 128–256 GB at Q4 Kimi K3 (2.8T), GLM-5.3 (753B), Qwen3.8-Max (2.45T), DeepSeek V4 Pro (1.65T)

Key model families (verified from model cards, Sept 2026)

Model Developer Total params Active Context Multimodal License Released
Kimi K3 Moonshot AI 2.8T 104B 1M Text+Image+Video Kimi K3 (custom) Jul 2026
DeepSeek V4 Pro (0813) DeepSeek 1.65T 49B 1M No MIT Aug 2026
DeepSeek V4 Flash DeepSeek 284B 13B 1M No MIT Apr 2026
DeepSeek V4 Flash Vision (Exp) DeepSeek 305B 13B + vision 1M Text+Image MIT Aug 2026
GLM-5.3 Z.ai (Zhipu) 753B ~40B 1M No GLM-5.3 (custom) Aug 2026
GLM-5.3-Flash Z.ai (Zhipu) 320B 18B 1M Text+Image MIT Aug 2026
Qwen3.8-Max (2.4T-A95B) Alibaba 2.45T 95B 262K (1M extensible) Text-only weights Qwen3.8-Max (custom) Aug 2026
Qwen3.8-Flash-Next Alibaba 180B 6B 262K (1M extensible) Text+Image+Video Qwen Community 1.0 Aug 2026
Qwen3.8-27B Alibaba 27.8B 27.8B 262K (1M extensible) Text+Image+Video Apache 2.0 Aug 2026
Qwen3.6-27B Alibaba 27B 27B 262K (1M via YaRN) Text+Image+Video Apache 2.0 Apr 2026
Qwen3.6-35B-A3B Alibaba 35B 3B 262K (1M via YaRN) Text+Image+Video Apache 2.0 Apr 2026
Qwen 3.5 397B-A17B Alibaba 397B 17B 256K Text+Image Apache 2.0 Feb 2026
Qwen3 235B-A22B Alibaba 235B 22B 128K No Apache 2.0 Apr 2025
Qwen3 8B Alibaba 8B 8B 128K No Apache 2.0 Apr 2025
gpt-oss-120b OpenAI 117B 5.1B 128K No Apache 2.0 Aug 2025
gpt-oss-20b OpenAI 21B 3.6B 128K No Apache 2.0 Aug 2025
Llama 4 Maverick Meta 400B 17B 1M Text+Image Llama 4 Community Apr 2025
Llama 4 Scout Meta 109B 17B 10M Text+Image Llama 4 Community (EU exclusion on multimodal) Apr 2025
Llama 3.3 70B Meta 70B 70B 128K No Llama 3.3 Community Dec 2024
Muse Glimmer 30B Meta 29.6B (incl. 1.8B vision) 29.6B dense 131K Text+Image Apache 2.0 (unmodified) Aug 2026
Gemma 4 31B Google 30.7B 30.7B 256K Text+Image Apache 2.0 Mar 2026
Gemma 4 26B-A4B Google 25.2B 3.8B 256K Text+Image Apache 2.0 Mar 2026
Gemma 4 12B Unified Google 12B 12B 256K Text+Image+Audio Apache 2.0 Jun 2026
Gemma 4 E4B Google 8B 4.5B effective 128K Text+Image+Audio Apache 2.0 Mar 2026
Gemma 4 E2B Google ~2B effective — 128K Text+Image Apache 2.0 Mar 2026
Mistral Small 4 Mistral AI 119B 6B 256K Text+Image Apache 2.0 Mar 2026
Mistral Large 3 Mistral AI 675B 41B 256K Text+Image Apache 2.0 Dec 2025
Phi-4 Microsoft 14B 14B 16K No MIT Jan 2025
Phi-4 Mini Microsoft 3.8B 3.8B 128K No MIT Jan 2025
Phi-4 Reasoning Vision Microsoft 15B 15B 16K Text+Image MIT Mar 2026
Command A Cohere 111B 111B 256K No CC-BY-NC Mar 2025
Falcon 3 10B TII Abu Dhabi 10B 10B 32K No TII Falcon-LLM 2.0 Dec 2024

Benchmarks and leaderboards

SWE-bench Verified (real-world coding) — the single most useful signal for local-capable models:

  • DeepSeek V4 Pro: 80.6% (max thinking) — highest published open-weight score
  • GLM-5: 77.8%
  • Qwen3.6-27B: 77.2% — a dense 27B on a single consumer GPU, beating Alibaba's own 397B MoE
  • Qwen3.6-35B-A3B: 73.4% (only 3B active params)
  • gpt-oss-120b: 62.4% (high reasoning)

Artificial Analysis Intelligence Index (checked 8 Sept 2026, max effort): GLM-5.3 tops the open-weight field at 45, ahead of Kimi K3 (44), GLM-5.3-Flash (42), Qwen3.8-2.4T-A95B (40), DeepSeek V4 Pro 0813 (36).

GLM-5.2 — GPQA Diamond 91.2%, AIME 2026 99.2, Terminal-Bench 2.1 81.0, SWE-bench Pro 62.1. DeepSeek R1 — MMLU-Pro 84.0%, GPQA Diamond 71.5%, MATH-500 97.3%. Qwen 3 235B — MMLU-Pro 83.8%, GPQA Diamond 77.1%, AIME '24 85.7%.

Open vs closed flagships (from Z.ai's GLM-5.3 card, vendor-reported across one shared benchmark set): across ten agentic/coding benchmarks, the best open-weight score leads on four. Open models are now at parity with closed models on tool use and security work, but the remaining gap is in long-horizon code generation — 8.6 points behind Opus 4.8 on NL2Repo, 5.2 on DeepSWE. Muse Glimmer's Meta-reported MCP Atlas score of 75.5 leads Qwen3.6-27B (62.5) and Gemma4-31B (54.2).

A caution from the source worth repeating: vendor-reported numbers travel between cards better than expected (17 of 19 shared rows match to the decimal), but benchmark versions matter — an 11.4-point "contradiction" on AutomationBench turned out to be two different benchmark versions. Check the version suffix before comparing.

Portuguese-language suitability

  • Qwen3 8B is the top practical Ollama-native choice for Portuguese in 2026: ollama run qwen3:8b, 8 GB VRAM, trained on 36T tokens over 119 languages, correct Portuguese output.
  • Quality ladder for Portuguese: Qwen3 8B (8 GB) → Qwen3 14B (16 GB) → Qwen3.8-27B (24 GB) for the best available quality.
  • Llama 3.1 8B is a competitive Ollama-native third option.
  • Sabiá-3 (Maritaca AI) approaches GPT-4o quality in Portuguese but is not on Ollama — HuggingFace download only. Worth it for production PT work.
  • PORTULAN CLARIN-PT-LDB (PROPOR 2026): the first leaderboard dedicated to European Portuguese (PT-PT), with novel benchmarks covering Portuguese culture alignment and model safeguards — https://huggingface.co/spaces/PORTULAN/portuguese-llm-leaderboard. This is the right reference for PT-PT evaluation specifically, as opposed to the Brazilian-Portuguese-focused guides.
  • Watch out: models trained primarily on English produce translated-sounding Portuguese, wrong variant vocabulary (ficheiro/ecrã vs arquivo/tela), and wrong pronoun forms. Anything under ~5% Portuguese training data should be avoided for PT-facing production.

Notable releases worth flagging

  • Muse Glimmer 30B (Meta, 10 Aug 2026): dense 29.6B, distilled from Muse Spark, multimodal, unmodified Apache 2.0, 131K context. Quantised builds target 24 GB (1.0% degradation) and 32 GB (0.2% degradation). Meta-reported MCP Atlas 75.5. Ships with a DFlash speculative-decoding drafter (~3× decode speed on RTX 5090) and ExecuTorch builds for Apple Metal. Built for always-on local agent work.
  • Llama 5 does not exist. No official Meta page, no weights, no announcement. Multiple web pages claiming an April 2026 release are wrong. The community's "Llama 5" appears to have arrived under a different name — Muse Glimmer / the Muse family.
  • GLM-5.2/5.3: the biggest open-weight gains of 2026 came from post-training on an unchanged base model — Terminal Bench 3.0 went 4.6 → 28.3 and AutomationBench 26.2 → 48.2 between GLM-5.2 and GLM-5.3.
  • Kimi-Linear 48B-A3B (Moonshot): KDA hybrid attention, 48B total / 3B active — fits the 24 GB tier.
  • Mellum 2 12B-A2.5B (JetBrains): code-specialised, LCB v6 69.9 with only 2.5B active params.
  • Bonsai-8B (Zyphra): 1-bit end-to-end, ~1.15 GB — extreme-compression edge case.

Licensing gotchas (short version — full detail in section 5)

  • Genuinely OSI-open (Apache 2.0 / MIT): Qwen3 & Qwen3.6/3.8 mid-range, DeepSeek, gpt-oss, Gemma 4, Mistral Small 4 / Large 3, OLMo 2, Phi-4, Muse Glimmer.
  • Restricted: Llama 4 (MAU cap + "Built with Llama" + EU exclusion on multimodal), Gemma 3 (Prohibited Use Policy, remotely enforceable), Qwen3.8-Max (custom), GLM-5.3 (custom), Kimi K3 (custom), Command A (CC-BY-NC, non-commercial).

Sources

  1. https://computingforgeeks.com/open-source-llm-comparison/ — master comparison of every major open-weight family: params, active params, context, licence, benchmarks; data read from model cards and config.json on Hugging Face in Sept 2026. The single most data-dense source for this section.
  2. https://github.com/xigh/open-weight-models — curated list filtered by commercially-exploitable licence, no EU geographic restriction, and VRAM-at-Q4 tiers (≤128 GB main, ≤256 GB extended). Useful for the "what actually runs locally" filter.
  3. https://aiwiki.ai/wiki/qwen_3 — Qwen3 family detail: 8 models, 0.6B–235B, hybrid thinking/non-thinking modes, Apache 2.0, 36T tokens over 119 languages; Qwen lineage overtook Llama as most-downloaded open-weight family.
  4. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ — official Gemma 4 launch (2 Apr 2026): four sizes (E2B, E4B, 26B MoE, 31B dense), Apache 2.0, built from Gemini 3 research, 400M+ Gemma downloads.
  5. https://www.digitalapplied.com/blog/meta-muse-glimmer-30b-apache-2-local-agent-model-2026 — Muse Glimmer 30B deep dive: dense 29.6B, Apache 2.0, 131K context, 24/32 GB quantised targets, MCP Atlas 75.5, DFlash drafter.
  6. https://www.orcarouter.ai/blog/llama-5-leak — debunks "Llama 5": no official release exists; community signal points to the Muse family instead. Also a live model-tracking index.
  7. https://www.promptquorum.com/local-llms/best-local-llms-portuguese-language-2026 — best local LLMs for Portuguese 2026: Qwen3 8B top Ollama-native pick, Sabiá-3 highest quality, per-tier VRAM guidance, PT-BR testing method.
  8. https://aclanthology.org/2026.propor-1.7/ — CLARIN-PT-LDB: the first European Portuguese open-LLM leaderboard, PROPOR 2026 (Silva, Gomes, Branco), with PT-PT culture and safeguards benchmarks.
  9. https://klyroocore.com/ai-models/phi-5 — Phi-5 spec page (blocked on re-fetch; the original live extraction was used, Phi-5: 8B, 128K context, 2026).
  10. https://www.siliconflow.com/articles/best-open-source-llm-for-portuguese — Portuguese-language model ranking (fetched by the original subagent).
  11. https://docs.mistral.ai/models — Mistral model lineup and licences (fetched by the original subagent).
  12. https://benchr.org/articles/open-weight-tier-right-now — open-weight tier analysis (fetched by the original subagent).
  13. https://ai-tldr.dev/models — model tracking aggregator (fetched by the original subagent).

Date: 2026-10-09

| Model | Developer | Total params | Active | Context | Multimodal | License | Released | |---|---|---|---|---|---|---|---|Kimi K3DeepSeek V4 ProDeepSeek V4 FlashDeepSeek V4 FlashGLM-5.3GLM-5.3Qwen3.8-MaxQwen3.8-Flash-NextQwen3.8-27BQwen3.6-27BQwen3.6-35B-A3BQwen 3.5Qwen3Qwen3gpt-ossgpt-ossLlama 4Llama 4Muse GlimmerGemma 4Gemma 4Gemma 4Gemma 4Gemma 4MistralMistralPhi-4Phi-4Phi-4Command AFalcon 3

Benchmarks and leaderboards

SWE-bench Verified (real-world coding) — the single most useful signal for local-capable models:

  • DeepSeek V4 Pro: 80.6% (max thinking) — highest published open-weight score
  • GLM-5: 77.8%
  • Qwen3.6-27B: 77.2% — a dense 27B on a single consumer GPU, beating Alibaba's own 397B MoE
  • Qwen3.6-35B-A3B: 73.4% (only 3B active params)
  • gpt-oss-120b: 62.4% (high reasoning)

Artificial Analysis Intelligence Index (checked 8 Sept 2026, max effort): GLM-5.3 tops the open-weight field at 45, ahead of Kimi K3 (44), GLM-5.3-Flash (42), Qwen3.8-2.4T-A95B (40), DeepSeek V4 Pro 0813 (36).

GLM-5.2 — GPQA Diamond 91.2%, AIME 2026 99.2, Terminal-Bench 2.1 81.0, SWE-bench Pro 62.1. DeepSeek R1 — MMLU-Pro 84.0%, GPQA Diamond 71.5%, MATH-500 97.3%. Qwen 3 235B — MMLU-Pro 83.8%, GPQA Diamond 77.1%, AIME '24 85.7%.

Open vs closed flagships (from Z.ai's GLM-5.3 card, vendor-reported across one shared benchmark set): across ten agentic/coding benchmarks, the best open-weight score leads on four. Open models are now at parity with closed models on tool use and security work, but the remaining gap is in long-horizon code generation — 8.6 points behind Opus 4.8 on NL2Repo, 5.2 on DeepSWE. Muse Glimmer's Meta-reported MCP Atlas score of 75.5 leads Qwen3.6-27B (62.5) and Gemma4-31B (54.2).

A caution from the source worth repeating: vendor-reported numbers travel between cards better than expected (17 of 19 shared rows match to the decimal), but benchmark versions matter — an 11.4-point "contradiction" on AutomationBench turned out to be two different benchmark versions. Check the version suffix before comparing.

Portuguese-language suitability

  • Qwen3 8B is the top practical Ollama-native choice for Portuguese in 2026: ollama run qwen3:8b, 8 GB VRAM, trained on 36T tokens over 119 languages, correct Portuguese output.
  • Quality ladder for Portuguese: Qwen3 8B (8 GB) → Qwen3 14B (16 GB) → Qwen3.8-27B (24 GB) for the best available quality.
  • Llama 3.1 8B is a competitive Ollama-native third option.
  • Sabiá-3 (Maritaca AI) approaches GPT-4o quality in Portuguese but is not on Ollama — HuggingFace download only. Worth it for production PT work.
  • PORTULAN CLARIN-PT-LDB (PROPOR 2026): the first leaderboard dedicated to European Portuguese (PT-PT), with novel benchmarks covering Portuguese culture alignment and model safeguards — https://huggingface.co/spaces/PORTULAN/portuguese-llm-leaderboard. This is the right reference for PT-PT evaluation specifically, as opposed to the Brazilian-Portuguese-focused guides.
  • Watch out: models trained primarily on English produce translated-sounding Portuguese, wrong variant vocabulary (ficheiro/ecrã vs arquivo/tela), and wrong pronoun forms. Anything under ~5% Portuguese training data should be avoided for PT-facing production.

Notable releases worth flagging

  • Muse Glimmer 30B (Meta, 10 Aug 2026): dense 29.6B, distilled from Muse Spark, multimodal, unmodified Apache 2.0, 131K context. Quantised builds target 24 GB (1.0% degradation) and 32 GB (0.2% degradation). Meta-reported MCP Atlas 75.5. Ships with a DFlash speculative-decoding drafter (~3× decode speed on RTX 5090) and ExecuTorch builds for Apple Metal. Built for always-on local agent work.
  • Llama 5 does not exist. No official Meta page, no weights, no announcement. Multiple web pages claiming an April 2026 release are wrong. The community's "Llama 5" appears to have arrived under a different name — Muse Glimmer / the Muse family.
  • GLM-5.2/5.3: the biggest open-weight gains of 2026 came from post-training on an unchanged base model — Terminal Bench 3.0 went 4.6 → 28.3 and AutomationBench 26.2 → 48.2 between GLM-5.2 and GLM-5.3.
  • Kimi-Linear 48B-A3B (Moonshot): KDA hybrid attention, 48B total / 3B active — fits the 24 GB tier.
  • Mellum 2 12B-A2.5B (JetBrains): code-specialised, LCB v6 69.9 with only 2.5B active params.
  • Bonsai-8B (Zyphra): 1-bit end-to-end, ~1.15 GB — extreme-compression edge case.

Sources

  1. https://computingforgeeks.com/open-source-llm-comparison/ — master comparison of every major open-weight family: params, active params, context, licence, benchmarks; data read from model cards and config.json on Hugging Face in Sept 2026. The single most data-dense source for this section.
  2. https://github.com/xigh/open-weight-models — curated list filtered by commercially-exploitable licence, no EU geographic restriction, and VRAM-at-Q4 tiers (≤128 GB main, ≤256 GB extended). Useful for the "what actually runs locally" filter.
  3. https://aiwiki.ai/wiki/qwen_3 — Qwen3 family detail: 8 models, 0.6B–235B, hybrid thinking/non-thinking modes, Apache 2.0, 36T tokens over 119 languages; Qwen lineage overtook Llama as most-downloaded open-weight family.
  4. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ — official Gemma 4 launch (2 Apr 2026): four sizes (E2B, E4B, 26B MoE, 31B dense), Apache 2.0, built from Gemini 3 research, 400M+ Gemma downloads.
  5. https://www.digitalapplied.com/blog/meta-muse-glimmer-30b-apache-2-local-agent-model-2026 — Muse Glimmer 30B deep dive: dense 29.6B, Apache 2.0, 131K context, 24/32 GB quantised targets, MCP Atlas 75.5, DFlash drafter.
  6. https://www.orcarouter.ai/blog/llama-5-leak — debunks "Llama 5": no official release exists; community signal points to the Muse family instead. Also a live model-tracking index.
  7. https://www.promptquorum.com/local-llms/best-local-llms-portuguese-language-2026 — best local LLMs for Portuguese 2026: Qwen3 8B top Ollama-native pick, Sabiá-3 highest quality, per-tier VRAM guidance, PT-BR testing method.
  8. https://aclanthology.org/2026.propor-1.7/ — CLARIN-PT-LDB: the first European Portuguese open-LLM leaderboard, PROPOR 2026 (Silva, Gomes, Branco), with PT-PT culture and safeguards benchmarks.
  9. https://klyroocore.com/ai-models/phi-5 — Phi-5 spec page (blocked on re-fetch; the original live extraction was used, Phi-5: 8B, 128K context, 2026).
  10. https://www.siliconflow.com/articles/best-open-source-llm-for-portuguese — Portuguese-language model ranking (fetched by the original subagent).
  11. https://docs.mistral.ai/models — Mistral model lineup and licences (fetched by the original subagent).
  12. https://benchr.org/articles/open-weight-tier-right-now — open-weight tier analysis (fetched by the original subagent).
  13. https://ai-tldr.dev/models — model tracking aggregator (fetched by the original subagent).

Date: 2026-10-09

var target=document.getElementById(location.hash.slice(1));target&&target.name&&(target.checked=target.name.startsWith("__tabbed_"))
{"annotate": null, "base": "../../..", "features": [], "search": "../../../assets/javascripts/workers/search.2c215733.min.js", "tags": null, "translations": {"clipboard.copied": "Copied to clipboard", "clipboard.copy": "Copy to clipboard", "search.result.more.one": "1 more on this page", "search.result.more.other": "# more on this page", "search.result.none": "No matching documents", "search.result.one": "1 matching document", "search.result.other": "# matching documents", "search.result.placeholder": "Type to start searching", "search.result.term.missing": "Missing", "select.version": "Select version"}, "version": null}