Skip to content

What Fits Your VRAM — Model Sizing by GPU Tier

Q4_K_M quantised. Use 0.6 GB per billion parameters, not 0.5.

Hardware tier Fits comfortably (Q4 quantised) Notable examples
8 GB VRAM Dense 7–8B; small MoE with ~3B active Qwen3 8B, Gemma 4 E4B, Phi-4 Mini, Muse Glimmer (Q3/Q4), gpt-oss-20b
16 GB VRAM Dense 12–14B; MoE ~6B active Gemma 4 12B, Qwen3 14B, Phi-4, Qwen3.8-Flash-Next (6B active), Gemma 4 E2B
24 GB VRAM Dense 26–31B; MoE ~13B active Qwen3.8-27B, Gemma 4 31B / 26B-A4B, Muse Glimmer 30B, Kimi-Linear 48B-A3B, DeepSeek V4 Flash (partially)
48 GB VRAM Dense up to ~70B (Q4); MoE 30–50B active Llama 3.3 70B, GLM-5.3-Flash (18B active), Nemotron 3 Nano Omni 30B-A3B, DeepSeek V4 Flash 284B/A13B (extended tier)
128 GB+ (workstation / Mac Studio) Models needing 128–256 GB at Q4 Kimi K3 (2.8T), GLM-5.3 (753B), Qwen3.8-Max (2.45T), DeepSeek V4 Pro (1.65T)

Sources

  1. https://computingforgeeks.com/open-source-llm-comparison/ — master comparison of every major open-weight family: params, active params, context, licence, benchmarks; data read from model cards and config.json on Hugging Face in Sept 2026. The single most data-dense source for this section.
  2. https://github.com/xigh/open-weight-models — curated list filtered by commercially-exploitable licence, no EU geographic restriction, and VRAM-at-Q4 tiers (≤128 GB main, ≤256 GB extended). Useful for the "what actually runs locally" filter.
  3. https://aiwiki.ai/wiki/qwen_3 — Qwen3 family detail: 8 models, 0.6B–235B, hybrid thinking/non-thinking modes, Apache 2.0, 36T tokens over 119 languages; Qwen lineage overtook Llama as most-downloaded open-weight family.
  4. https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ — official Gemma 4 launch (2 Apr 2026): four sizes (E2B, E4B, 26B MoE, 31B dense), Apache 2.0, built from Gemini 3 research, 400M+ Gemma downloads.
  5. https://www.digitalapplied.com/blog/meta-muse-glimmer-30b-apache-2-local-agent-model-2026 — Muse Glimmer 30B deep dive: dense 29.6B, Apache 2.0, 131K context, 24/32 GB quantised targets, MCP Atlas 75.5, DFlash drafter.
  6. https://www.orcarouter.ai/blog/llama-5-leak — debunks "Llama 5": no official release exists; community signal points to the Muse family instead. Also a live model-tracking index.
  7. https://www.promptquorum.com/local-llms/best-local-llms-portuguese-language-2026 — best local LLMs for Portuguese 2026: Qwen3 8B top Ollama-native pick, Sabiá-3 highest quality, per-tier VRAM guidance, PT-BR testing method.
  8. https://aclanthology.org/2026.propor-1.7/ — CLARIN-PT-LDB: the first European Portuguese open-LLM leaderboard, PROPOR 2026 (Silva, Gomes, Branco), with PT-PT culture and safeguards benchmarks.
  9. https://klyroocore.com/ai-models/phi-5 — Phi-5 spec page (blocked on re-fetch; the original live extraction was used, Phi-5: 8B, 128K context, 2026).
  10. https://www.siliconflow.com/articles/best-open-source-llm-for-portuguese — Portuguese-language model ranking (fetched by the original subagent).
  11. https://docs.mistral.ai/models — Mistral model lineup and licences (fetched by the original subagent).
  12. https://benchr.org/articles/open-weight-tier-right-now — open-weight tier analysis (fetched by the original subagent).
  13. https://ai-tldr.dev/models — model tracking aggregator (fetched by the original subagent).

Date: 2026-10-09