ExLlamaV3 — Max Quality-Per-Bit on NVIDIA/AMD
Self-contained guide to ExLlamaV3 — the local inference engine.
- ExLlamaV2 is legacy: last commit 2026-03-04, last PyPI release 0.3.2 (2025-07-14). Multiple sources describe it as "a legacy, archived project" superseded by V3 (mortalapps). Don't start new work on EXL2.
- ExLlamaV3: active (last commit 2026-10-08, PyPI 1.5.3 on 2026-09-27, 1,996 commits). New EXL3 format, a streamlined variant of QTIP (Cornell RelaxML) using trellis-coded quantisation rather than scalar block rounding, with 2–8 bit KV cache quantisation. EXL3 supports 2.5–6.0+ bpw with a choice of LM-head precision (e.g. 2.5bpw ≈ 9.80 GiB, 4.0bpw ≈ 15.25 GiB for a 70B-class model).
- Backends: CUDA (NVIDIA consumer GPUs) and, newly in 2026, ROCm wheels (commits landed Oct 6–7, 2026 fixing hipBLASLt fp32 bugs and disabling CUDA graphs by default on ROCm). AMD support is young.
- Honest maturity caveat: the repo itself still says "ExLlamaV3 is still in development… the framework is not yet fully optimized. Performance is lacking, especially on Ampere."
- Quality/speed: an independent comparison found EXL3 delivered better quality per bit than GGUF and much faster prompt processing on an RTX 5090, though GGUF loaded and decoded faster in that test (atomic.chat).
- Best for: maximum quality-per-bit on a single modern NVIDIA (and increasingly AMD) consumer GPU, single-user, especially large dense models.
Quickstart
Assembled from the verified facts on this page and its Sources list. Versions are the ones recorded in research on 2026-10-09.
Install
# Pre-1.0 project; install from source with active ROCm wheels (Oct 2026)
git clone https://github.com/turboderp-org/exllamav3
cd exllamav3 && pip install -e .
Run
# Best quality-per-bit on a single modern NVIDIA/AMD card
python examples/generator.py \
-m models/Qwen3-8B-EXL3-4.0bpw \
-p 'Explain tensor parallelism.'
Verify
Expect a token rate near the measured 52 tok/s for EXL2 4.9bpw on an RTX 3090.
Expected output: EXL3/QTIP quantisation with 2-8 bit cache quantisation. On a 3090 13B test EXL2 4.9bpw gave the best quality (4.308 ppl) of any format.
Pitfalls: The repo still says 'in development'. Version 1.5.3, 1,623 stars — small user base means slower issue response than the majors.
Sources
- https://github.com/ggml-org/llama.cpp — llama.cpp repo; 130,619 stars, 11,523 commits, active 2026-10-09.
- https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md — authoritative list of llama.cpp backends (CUDA, HIP, Metal, Vulkan, SYCL, MUSA, CANN, ZenDNN, OpenCL, OpenVINO, Hexagon, KleidiAI, WebGPU, BLAS vendors).
- https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md — llama.cpp multi-GPU split modes (
none/layer/row), flags and recipes. - https://github.com/ollama/ollama — Ollama repo; 182,473 stars, releases v0.40.1 (2026-10-07).
- https://ollama.com/blog — Ollama blog; MLX engine on Apple Silicon (Mar/Jun 2026), GGUF via llama.cpp (0.30, Jun 2026), $88M raise, 8.9M developers, Anthropic API + Claude Desktop support.
- https://docs.ollama.com/gpu — Ollama hardware support: CUDA compute 5.0+, ROCm v7 on Linux and Windows, Metal, Vulkan defaults and
GGML_VK_VISIBLE_DEVICES. - https://lmstudio.ai/blog/0.4.0 — LM Studio 0.4.0:
llmsterheadless daemon, parallel requests with continuous batching, stateful/v1/chat, unified KV cache. - https://lmstudio.ai/changelog/lmstudio — LM Studio changelog through 0.4.26 (2026-10-08): Splash engine, DFlash/DSpark/MTP drafters, CUDA 13 on Windows ARM, llama.cpp 2.50.0/b11337.
- https://docs.vllm.ai/en/stable/usage/v1_guide/ — vLLM V1 guide: V0 fully deprecated; hardware status (NVIDIA/AMD/Intel/TPU/CPU all functional); plugins.
- https://docs.vllm.ai/en/latest/features/quantization/ — vLLM quantization formats and per-hardware compatibility matrix (AWQ, GPTQ, Marlin, FP8, GGUF, etc.).
- https://vllm.ai/blog/2026-09-22-vllm-metal-v0-28-0 — Announcing vllm-metal: paged varlen Metal kernel, MTP, GGUF/hybrid support, ragged-batch benchmark vs mlx_lm/oMLX/llama.cpp.
- https://www.sglang.io/ — SGLang site: 0.5.21 highlights (Rust prefix-cache core, PD role switching, radix tree), supported hardware NVIDIA/AMD/CPU/TPU/Ascend/XPU.
- https://github.com/sgl-project/sglang — SGLang repo; 36,919 stars, active 2026-10-09, DeepSeek-V4 Rust processor parity work.
- https://github.com/turboderp-org/exllamav3 — ExLlamaV3 repo: EXL3/QTIP quantisation, 2–8 bit cache quant, active ROCm wheels (Oct 2026), "still in development" caveat.
- https://github.com/turboderp-org/exllamav2 — ExLlamaV2 repo: last commit 2026-03-04 — confirms legacy/dormant status.
- https://nvidia.github.io/TensorRT-LLM/latest/release-notes.html — TensorRT-LLM 1.3 release notes: TRITON MoE deprecation; 1.2 removed the TensorRT backend entirely, added DGX Spark beta.
- https://github.com/mlc-ai/mlc-llm — MLC-LLM repo; 23,226 stars, last commit 2026-10-06, TVM refactor work.
- https://github.com/mozilla-ai/llamafile — llamafile repo under Mozilla AI; active 2026-10-08, agent.cpp subtree and "Agentfile" work.
- https://github.com/nomic-ai/gpt4all — GPT4All repo: last commit 2025-05-27 — confirms end-of-life status.
- https://github.com/mudler/localai — LocalAI repo; 49,450 stars, 8,447 commits, 4.11.0 (Oct 2026), 73 backends.
- https://localai.io/blog/what-landed-in-localai-4-8/ — LocalAI 4.8: model variants auto-selection, vllm.cpp (alpha) engine, 3.48× lighter web UI.
- https://localai.io/docs/features/vllm-cpp/index.html — vllm.cpp backend docs: Blackwell-only CUDA images (sm_120a/sm_121a), CUDA 13 required, Vulkan/CPU fallback, KV sizing (144 KiB/token).
- https://github.com/localai-org/vllm.cpp — vllm.cpp repo: C++20 port of vLLM with continuous batching, paged KV, RadixAttention; 6,363 commits, active 2026-10-09.
- https://www.jan.ai/docs/desktop — Jan docs: 0.8.4, Jan Server OpenAI-compatible API, CLI, llama.cpp + experimental MLX engines.
- https://freedom.tech/posts/2026-09-26-koboldcpp-1-122-1/ — KoboldCpp 1.122.1: integrated agent with 9 tools, MCP, AGENTS.md, ubatch handling.
- https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison/ — Reproducible H100 benchmark of vLLM 0.30 / SGLang 0.5.20 / TensorRT-LLM 1.2.1 / llama.cpp b11179 / Ollama 0.34.4 on Llama 3.1 8B and Qwen3.8-27B.
- https://www.soothill.io/blog/2026/08/10/sglang-vllm-llamacpp-evox3/ — 14-point AMD Strix Halo (
gfx1151) comparison of SGLang 0.5.17 / vLLM 0.26.0 / llama.cpp b10333; ROCm 7.14 support gaps. - https://ai-tldr.dev/tools/magnitude/ — Magnitude: Rust agent-oriented engine, kernel autotuning, OpenAI+Anthropic API on port 10100, v0.2.0 (2026-09-30), vendor benchmark claims.
- https://www.developersdigest.tech/blog/magnitude-self-tuning-local-inference-engine-2026 — Magnitude launch coverage: Apache-2.0, YC-backed, claims up to 2× faster decode than llama.cpp.
- https://ai-tldr.dev/tools/inco-splash/ — Splash: Apple Silicon engine with precompiled model-specific Metal kernels, DFlash 2 speculative decoding, M3+/macOS 26.4+.
- https://presenc.ai/research/mlx-vs-llama-cpp-throughput-benchmarks-2026 — MLX vs llama.cpp Apple Silicon benchmarks: 10–20% decode advantage above 14B, 4.06× TTFT on M5, crossover past ~40k context.
- https://www.glukhov.org/llm-hosting/comparisons/hosting-llms-ollama-localai-jan-lmstudio-vllm-comparison/ — 14-tool comparison across API maturity, tool calling, GPU support, formats, production readiness.
- https://docs.litellm.ai/docs/proxy_server — LiteLLM Proxy: OpenAI-compatible gateway over 100+ providers including local engines.
- https://particula.tech/blog/ollama-num-ctx-silent-prompt-truncation — Ollama single-slot architectures (mllama, qwen3vl, nemotron_h) forced to one slot since 2026-02-02.
- https://helix.ml/blog/the-ceiling-was-a-state-cache — SGLang mamba state cache silently capping concurrency; 3,833 → 7,883 tok/s after fix.
- https://markaicode.com/benchmarks/gpt4all-production-benchmark-latency/ — GPT4All end-of-life confirmation and CPU-only status.
- https://mortalapps.com/blog/gguf-vs-exl2-vs-mlx-quantization/ — ExLlamaV2 legacy/archived status and EXL3 succession in 2026.
- https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF quality/speed comparison on RTX 5090.
- https://localai.io/blog/what-landed-in-localai-4-11/ — LocalAI 4.11: diarisation, speaker profiles, model failover chains, Operate → This machine.
- https://developer.apple.com/videos/play/wwdc2026/232/ — WWDC26: MLX-LM and the OpenAI-compatible MLX-LM Server for local agentic AI on Mac.
Date: 2026-10-09