Skip to content

llama.cpp — The Substrate of Local Inference

Self-contained guide to llama.cpp — the local inference engine.

  • What it is: C/C++ inference engine defining the GGUF file format and the quantisation ecosystem. Now hosted under ggml-org (moved from ggerganov), 11,523 commits, commits landing hours before this report (repo, build docs).
  • Backends (broadest of any engine): CUDA, HIP/ROCm, Metal, Vulkan, SYCL (Intel), MUSA, CANN (Ascend), ZenDNN, OpenCL, OpenVINO, Hexagon (Snapdragon NPU), Arm KleidiAI, WebGPU, plus CPU with BLAS vendors (OpenBLAS, BLIS, AMD AOCL, Intel oneMKL, Apple Accelerate by default on macOS).
  • Quantisation: GGUF, from Q2_K up to Q8_0 and F16, plus importance-matrix mixing (imatrix) and --allow-requantize / --leave-output-tensor to keep embedding/output tensors at higher precision.
  • API: llama-server ships an OpenAI-compatible HTTP server (/v1/chat/completions, /v1/completions, /v1/embeddings).
  • Multi-GPU: yes, via --split-mode (none, layer = pipeline parallelism, default; row = tensor parallelism) and --main-gpu, documented in docs/multi-gpu.md.
  • Best for: single-user on any hardware; the default choice when portability or a low-bit GGUF is what matters.
  • Notable 2026 development: continuous batching was added upstream and is now used by LM Studio. A caveat found in benchmark testing: on hybrid (linear-attention) models, llama.cpp re-processes whole prompts unless resuming from a saved checkpoint (PR #16382), which degrades throughput at concurrency.

Quickstart

Assembled from the verified facts on this page and its Sources list. Versions are the ones recorded in research on 2026-10-09.

Install

# Build from source (any backend; CUDA shown)
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j

Run

# Serve any GGUF over an OpenAI-compatible API (default port 8080)
./build/bin/llama-server \
  -m models/qwen3-8b-q4_k_m.gguf \
  -c 8192 \
  --port 8080

Verify

curl http://localhost:8080/v1/models

Expected output: A JSON list of the loaded model. llama-server also serves a browser chat UI on the same port.

Pitfalls: Vulkan is the pragmatic default backend when you have no CUDA/ROCm stack, but a silent CPU fallback is the classic trap — check the startup log for the backend actually selected, not a device query.

Sources

  1. https://github.com/ggml-org/llama.cpp — llama.cpp repo; 130,619 stars, 11,523 commits, active 2026-10-09.
  2. https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md — authoritative list of llama.cpp backends (CUDA, HIP, Metal, Vulkan, SYCL, MUSA, CANN, ZenDNN, OpenCL, OpenVINO, Hexagon, KleidiAI, WebGPU, BLAS vendors).
  3. https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md — llama.cpp multi-GPU split modes (none/layer/row), flags and recipes.
  4. https://github.com/ollama/ollama — Ollama repo; 182,473 stars, releases v0.40.1 (2026-10-07).
  5. https://ollama.com/blog — Ollama blog; MLX engine on Apple Silicon (Mar/Jun 2026), GGUF via llama.cpp (0.30, Jun 2026), $88M raise, 8.9M developers, Anthropic API + Claude Desktop support.
  6. https://docs.ollama.com/gpu — Ollama hardware support: CUDA compute 5.0+, ROCm v7 on Linux and Windows, Metal, Vulkan defaults and GGML_VK_VISIBLE_DEVICES.
  7. https://lmstudio.ai/blog/0.4.0 — LM Studio 0.4.0: llmster headless daemon, parallel requests with continuous batching, stateful /v1/chat, unified KV cache.
  8. https://lmstudio.ai/changelog/lmstudio — LM Studio changelog through 0.4.26 (2026-10-08): Splash engine, DFlash/DSpark/MTP drafters, CUDA 13 on Windows ARM, llama.cpp 2.50.0/b11337.
  9. https://docs.vllm.ai/en/stable/usage/v1_guide/ — vLLM V1 guide: V0 fully deprecated; hardware status (NVIDIA/AMD/Intel/TPU/CPU all functional); plugins.
  10. https://docs.vllm.ai/en/latest/features/quantization/ — vLLM quantization formats and per-hardware compatibility matrix (AWQ, GPTQ, Marlin, FP8, GGUF, etc.).
  11. https://vllm.ai/blog/2026-09-22-vllm-metal-v0-28-0 — Announcing vllm-metal: paged varlen Metal kernel, MTP, GGUF/hybrid support, ragged-batch benchmark vs mlx_lm/oMLX/llama.cpp.
  12. https://www.sglang.io/ — SGLang site: 0.5.21 highlights (Rust prefix-cache core, PD role switching, radix tree), supported hardware NVIDIA/AMD/CPU/TPU/Ascend/XPU.
  13. https://github.com/sgl-project/sglang — SGLang repo; 36,919 stars, active 2026-10-09, DeepSeek-V4 Rust processor parity work.
  14. https://github.com/turboderp-org/exllamav3 — ExLlamaV3 repo: EXL3/QTIP quantisation, 2–8 bit cache quant, active ROCm wheels (Oct 2026), "still in development" caveat.
  15. https://github.com/turboderp-org/exllamav2 — ExLlamaV2 repo: last commit 2026-03-04 — confirms legacy/dormant status.
  16. https://nvidia.github.io/TensorRT-LLM/latest/release-notes.html — TensorRT-LLM 1.3 release notes: TRITON MoE deprecation; 1.2 removed the TensorRT backend entirely, added DGX Spark beta.
  17. https://github.com/mlc-ai/mlc-llm — MLC-LLM repo; 23,226 stars, last commit 2026-10-06, TVM refactor work.
  18. https://github.com/mozilla-ai/llamafile — llamafile repo under Mozilla AI; active 2026-10-08, agent.cpp subtree and "Agentfile" work.
  19. https://github.com/nomic-ai/gpt4all — GPT4All repo: last commit 2025-05-27 — confirms end-of-life status.
  20. https://github.com/mudler/localai — LocalAI repo; 49,450 stars, 8,447 commits, 4.11.0 (Oct 2026), 73 backends.
  21. https://localai.io/blog/what-landed-in-localai-4-8/ — LocalAI 4.8: model variants auto-selection, vllm.cpp (alpha) engine, 3.48× lighter web UI.
  22. https://localai.io/docs/features/vllm-cpp/index.html — vllm.cpp backend docs: Blackwell-only CUDA images (sm_120a/sm_121a), CUDA 13 required, Vulkan/CPU fallback, KV sizing (144 KiB/token).
  23. https://github.com/localai-org/vllm.cpp — vllm.cpp repo: C++20 port of vLLM with continuous batching, paged KV, RadixAttention; 6,363 commits, active 2026-10-09.
  24. https://www.jan.ai/docs/desktop — Jan docs: 0.8.4, Jan Server OpenAI-compatible API, CLI, llama.cpp + experimental MLX engines.
  25. https://freedom.tech/posts/2026-09-26-koboldcpp-1-122-1/ — KoboldCpp 1.122.1: integrated agent with 9 tools, MCP, AGENTS.md, ubatch handling.
  26. https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison/ — Reproducible H100 benchmark of vLLM 0.30 / SGLang 0.5.20 / TensorRT-LLM 1.2.1 / llama.cpp b11179 / Ollama 0.34.4 on Llama 3.1 8B and Qwen3.8-27B.
  27. https://www.soothill.io/blog/2026/08/10/sglang-vllm-llamacpp-evox3/ — 14-point AMD Strix Halo (gfx1151) comparison of SGLang 0.5.17 / vLLM 0.26.0 / llama.cpp b10333; ROCm 7.14 support gaps.
  28. https://ai-tldr.dev/tools/magnitude/ — Magnitude: Rust agent-oriented engine, kernel autotuning, OpenAI+Anthropic API on port 10100, v0.2.0 (2026-09-30), vendor benchmark claims.
  29. https://www.developersdigest.tech/blog/magnitude-self-tuning-local-inference-engine-2026 — Magnitude launch coverage: Apache-2.0, YC-backed, claims up to 2× faster decode than llama.cpp.
  30. https://ai-tldr.dev/tools/inco-splash/ — Splash: Apple Silicon engine with precompiled model-specific Metal kernels, DFlash 2 speculative decoding, M3+/macOS 26.4+.
  31. https://presenc.ai/research/mlx-vs-llama-cpp-throughput-benchmarks-2026 — MLX vs llama.cpp Apple Silicon benchmarks: 10–20% decode advantage above 14B, 4.06× TTFT on M5, crossover past ~40k context.
  32. https://www.glukhov.org/llm-hosting/comparisons/hosting-llms-ollama-localai-jan-lmstudio-vllm-comparison/ — 14-tool comparison across API maturity, tool calling, GPU support, formats, production readiness.
  33. https://docs.litellm.ai/docs/proxy_server — LiteLLM Proxy: OpenAI-compatible gateway over 100+ providers including local engines.
  34. https://particula.tech/blog/ollama-num-ctx-silent-prompt-truncation — Ollama single-slot architectures (mllama, qwen3vl, nemotron_h) forced to one slot since 2026-02-02.
  35. https://helix.ml/blog/the-ceiling-was-a-state-cache — SGLang mamba state cache silently capping concurrency; 3,833 → 7,883 tok/s after fix.
  36. https://markaicode.com/benchmarks/gpt4all-production-benchmark-latency/ — GPT4All end-of-life confirmation and CPU-only status.
  37. https://mortalapps.com/blog/gguf-vs-exl2-vs-mlx-quantization/ — ExLlamaV2 legacy/archived status and EXL3 succession in 2026.
  38. https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF quality/speed comparison on RTX 5090.
  39. https://localai.io/blog/what-landed-in-localai-4-11/ — LocalAI 4.11: diarisation, speaker profiles, model failover chains, Operate → This machine.
  40. https://developer.apple.com/videos/play/wwdc2026/232/ — WWDC26: MLX-LM and the OpenAI-compatible MLX-LM Server for local agentic AI on Mac.

Date: 2026-10-09