vLLM โ The Throughput Standard
Self-contained guide to vLLM โ the local inference engine.
- What it is: Python GPU serving engine built on PagedAttention, 93k stars, 0.31.0 on PyPI (2026-10-05).
- V1 is now the only engine: the docs state plainly "We have fully deprecated V0" (RFC #18571). V1 re-architected the scheduler, KV cache manager, worker, sampler and API server, with chunked prefill and CUDA graphs on by default, near-zero CPU overhead and "zero configs" (v1_guide).
- Hardware (all ๐ข in the V1 guide): NVIDIA, AMD, Intel GPU, TPU, CPU. More via plugins: vllm-ascend, vllm-spyre, vllm-gaudi, vllm-openvino. A dedicated vllm-metal plugin now covers Apple Silicon (ยง5).
- Quantisation (quantization docs): AutoAWQ, BitsAndBytes, GPTQModel, Intel Neural Compressor, LLM Compressor (FP8 W8A8, INT4 W4A16, INT8 W4A8, INT8 W8A8), NVIDIA Model Optimizer, AMD Quark, TorchAO, online quantisation, quantised KV cache, and GGUF. A per-quantisation linear-backend selector (
--linear-backend cutlass,linear_backend_per_quant: {nvfp4_w4a16: humming}) was added in 2026. Note the hardware matrix: AWQ/GPTQ/Marlin are NVIDIA-only; FP8 requires Ada+ or AMD; GGUF is not supported on Intel GPU or x86/Arm CPU in vLLM. - API: full OpenAI-compatible server with streaming and tool-call parsing; documented integrations for Claude Code and Codex.
- Best for: concurrency, throughput, the widest ecosystem (llm-d and Red Hat AI Inference Server are built on it). Multi-GPU tensor/pipeline/expert parallelism is first-class.
Quickstart
Assembled from the verified facts on this page and its Sources list. Versions are the ones recorded in research on 2026-10-09.
Install
# Requires a CUDA-capable GPU; vLLM 0.31.0 (2026-10-05)
pip install 'vllm==0.31.0'
Run
# OpenAI-compatible server, default port 8000
vllm serve Qwen/Qwen3-8B \
--max-model-len 8192 \
--port 8000
Verify
curl http://localhost:8000/v1/models
curl http://localhost:8000/metrics # Prometheus endpoint
Expected output: Startup logs the GPU memory headroom, then /v1/models answers. Use this when you need real concurrency โ it is the throughput standard.
Pitfalls: The V0 engine path is fully deprecated in the V1 guide; old V0 tuning flags and blog recipes no longer apply.
Sources
- https://github.com/ggml-org/llama.cpp โ llama.cpp repo; 130,619 stars, 11,523 commits, active 2026-10-09.
- https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md โ authoritative list of llama.cpp backends (CUDA, HIP, Metal, Vulkan, SYCL, MUSA, CANN, ZenDNN, OpenCL, OpenVINO, Hexagon, KleidiAI, WebGPU, BLAS vendors).
- https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md โ llama.cpp multi-GPU split modes (
none/layer/row), flags and recipes. - https://github.com/ollama/ollama โ Ollama repo; 182,473 stars, releases v0.40.1 (2026-10-07).
- https://ollama.com/blog โ Ollama blog; MLX engine on Apple Silicon (Mar/Jun 2026), GGUF via llama.cpp (0.30, Jun 2026), $88M raise, 8.9M developers, Anthropic API + Claude Desktop support.
- https://docs.ollama.com/gpu โ Ollama hardware support: CUDA compute 5.0+, ROCm v7 on Linux and Windows, Metal, Vulkan defaults and
GGML_VK_VISIBLE_DEVICES. - https://lmstudio.ai/blog/0.4.0 โ LM Studio 0.4.0:
llmsterheadless daemon, parallel requests with continuous batching, stateful/v1/chat, unified KV cache. - https://lmstudio.ai/changelog/lmstudio โ LM Studio changelog through 0.4.26 (2026-10-08): Splash engine, DFlash/DSpark/MTP drafters, CUDA 13 on Windows ARM, llama.cpp 2.50.0/b11337.
- https://docs.vllm.ai/en/stable/usage/v1_guide/ โ vLLM V1 guide: V0 fully deprecated; hardware status (NVIDIA/AMD/Intel/TPU/CPU all functional); plugins.
- https://docs.vllm.ai/en/latest/features/quantization/ โ vLLM quantization formats and per-hardware compatibility matrix (AWQ, GPTQ, Marlin, FP8, GGUF, etc.).
- https://vllm.ai/blog/2026-09-22-vllm-metal-v0-28-0 โ Announcing vllm-metal: paged varlen Metal kernel, MTP, GGUF/hybrid support, ragged-batch benchmark vs mlx_lm/oMLX/llama.cpp.
- https://www.sglang.io/ โ SGLang site: 0.5.21 highlights (Rust prefix-cache core, PD role switching, radix tree), supported hardware NVIDIA/AMD/CPU/TPU/Ascend/XPU.
- https://github.com/sgl-project/sglang โ SGLang repo; 36,919 stars, active 2026-10-09, DeepSeek-V4 Rust processor parity work.
- https://github.com/turboderp-org/exllamav3 โ ExLlamaV3 repo: EXL3/QTIP quantisation, 2โ8 bit cache quant, active ROCm wheels (Oct 2026), "still in development" caveat.
- https://github.com/turboderp-org/exllamav2 โ ExLlamaV2 repo: last commit 2026-03-04 โ confirms legacy/dormant status.
- https://nvidia.github.io/TensorRT-LLM/latest/release-notes.html โ TensorRT-LLM 1.3 release notes: TRITON MoE deprecation; 1.2 removed the TensorRT backend entirely, added DGX Spark beta.
- https://github.com/mlc-ai/mlc-llm โ MLC-LLM repo; 23,226 stars, last commit 2026-10-06, TVM refactor work.
- https://github.com/mozilla-ai/llamafile โ llamafile repo under Mozilla AI; active 2026-10-08, agent.cpp subtree and "Agentfile" work.
- https://github.com/nomic-ai/gpt4all โ GPT4All repo: last commit 2025-05-27 โ confirms end-of-life status.
- https://github.com/mudler/localai โ LocalAI repo; 49,450 stars, 8,447 commits, 4.11.0 (Oct 2026), 73 backends.
- https://localai.io/blog/what-landed-in-localai-4-8/ โ LocalAI 4.8: model variants auto-selection, vllm.cpp (alpha) engine, 3.48ร lighter web UI.
- https://localai.io/docs/features/vllm-cpp/index.html โ vllm.cpp backend docs: Blackwell-only CUDA images (sm_120a/sm_121a), CUDA 13 required, Vulkan/CPU fallback, KV sizing (144 KiB/token).
- https://github.com/localai-org/vllm.cpp โ vllm.cpp repo: C++20 port of vLLM with continuous batching, paged KV, RadixAttention; 6,363 commits, active 2026-10-09.
- https://www.jan.ai/docs/desktop โ Jan docs: 0.8.4, Jan Server OpenAI-compatible API, CLI, llama.cpp + experimental MLX engines.
- https://freedom.tech/posts/2026-09-26-koboldcpp-1-122-1/ โ KoboldCpp 1.122.1: integrated agent with 9 tools, MCP, AGENTS.md, ubatch handling.
- https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison/ โ Reproducible H100 benchmark of vLLM 0.30 / SGLang 0.5.20 / TensorRT-LLM 1.2.1 / llama.cpp b11179 / Ollama 0.34.4 on Llama 3.1 8B and Qwen3.8-27B.
- https://www.soothill.io/blog/2026/08/10/sglang-vllm-llamacpp-evox3/ โ 14-point AMD Strix Halo (
gfx1151) comparison of SGLang 0.5.17 / vLLM 0.26.0 / llama.cpp b10333; ROCm 7.14 support gaps. - https://ai-tldr.dev/tools/magnitude/ โ Magnitude: Rust agent-oriented engine, kernel autotuning, OpenAI+Anthropic API on port 10100, v0.2.0 (2026-09-30), vendor benchmark claims.
- https://www.developersdigest.tech/blog/magnitude-self-tuning-local-inference-engine-2026 โ Magnitude launch coverage: Apache-2.0, YC-backed, claims up to 2ร faster decode than llama.cpp.
- https://ai-tldr.dev/tools/inco-splash/ โ Splash: Apple Silicon engine with precompiled model-specific Metal kernels, DFlash 2 speculative decoding, M3+/macOS 26.4+.
- https://presenc.ai/research/mlx-vs-llama-cpp-throughput-benchmarks-2026 โ MLX vs llama.cpp Apple Silicon benchmarks: 10โ20% decode advantage above 14B, 4.06ร TTFT on M5, crossover past ~40k context.
- https://www.glukhov.org/llm-hosting/comparisons/hosting-llms-ollama-localai-jan-lmstudio-vllm-comparison/ โ 14-tool comparison across API maturity, tool calling, GPU support, formats, production readiness.
- https://docs.litellm.ai/docs/proxy_server โ LiteLLM Proxy: OpenAI-compatible gateway over 100+ providers including local engines.
- https://particula.tech/blog/ollama-num-ctx-silent-prompt-truncation โ Ollama single-slot architectures (mllama, qwen3vl, nemotron_h) forced to one slot since 2026-02-02.
- https://helix.ml/blog/the-ceiling-was-a-state-cache โ SGLang mamba state cache silently capping concurrency; 3,833 โ 7,883 tok/s after fix.
- https://markaicode.com/benchmarks/gpt4all-production-benchmark-latency/ โ GPT4All end-of-life confirmation and CPU-only status.
- https://mortalapps.com/blog/gguf-vs-exl2-vs-mlx-quantization/ โ ExLlamaV2 legacy/archived status and EXL3 succession in 2026.
- https://atomic.chat/blog/guides/exl3-vs-gguf โ EXL3 vs GGUF quality/speed comparison on RTX 5090.
- https://localai.io/blog/what-landed-in-localai-4-11/ โ LocalAI 4.11: diarisation, speaker profiles, model failover chains, Operate โ This machine.
- https://developer.apple.com/videos/play/wwdc2026/232/ โ WWDC26: MLX-LM and the OpenAI-compatible MLX-LM Server for local agentic AI on Mac.
Date: 2026-10-09