Emergent Agent-Oriented Engines (2026) - Local AI Agent Wiki __md_scope=new URL("../../..",location),__md_hash=e=>[...e].reduce(((e,_)=>(e<<5)-e+_.charCodeAt(0)),0),__md_get=(e,_=localStorage,t=__md_scope)=>JSON.parse(_.getItem(t.pathname+"."+e)),__md_set=(e,_,t=localStorage,a=__md_scope)=>{try{t.setItem(a.pathname+"."+e,JSON.stringify(_))}catch(e){}}
Skip to content

Emergent Agent-Oriented Engines (2026)

New in 2026: engines purpose-built for coding agents, not chat. Includes vendor-claimed benchmarks that are not independently verified.

vllm.cpp (localai-org/vllm.cpp) — a C++20 port of vLLM: the same continuous-batching scheduler, paged KV cache and automatic prefix caching, with no Python at inference time. Apache-2.0, maintained by the LocalAI team, 6,363 commits, active within an hour of this report. It loads HuggingFace safetensors directories or .gguf files, and handles chat templates, tool-call parsing and reasoning splits in-engine. LocalAI ships it as the native vllm-cpp backend with images for CPU, CUDA 13, Vulkan, Metal and Jetson L4T (docs). - Important limitation: the CUDA images are built for Blackwell architectures only (sm_120a, sm_121a) and require CUDA 13 — there is no CUDA 12 variant. On Ampere/Ada/Hopper/Orin the backend installs then fails at first request with "no kernel image is available"; use the Vulkan or CPU image there. The engine itself builds ten architectures. - Gallery entries include Qwen3.6-27B NVFP4 (with MTP and DFlash speculative variants), Qwen3.6-35B-A3B NVFP4, Qwen3-Coder-30B-A3B, and small Qwen3 models that run on CPU/Metal/Vulkan. A useful sizing note from the docs: the 4B entry spends 144 KiB of KV per token, so 1024 blocks ≈ 4.5 GB on top of the weights. - This is the most interesting structural development of 2026: vLLM's serving architecture without its Python dependency.

Magnitude (magnitudedev/magnitude, YC S25) — an open-source Rust inference engine built for agents rather than chat. Its premise is precisely the 2026 gap: vLLM/SGLang are built for batched datacenter serving, llama.cpp/Ollama trade peak speed for compatibility, and neither targets "long, concurrent agent sessions on one personal machine." Features: on-device kernel compilation and autotuning for Metal/CUDA/AMD/CPU; memory reserved only for weights up front with the heap growing per agent session and freed when agents stop; hybrid paged attention so concurrent sessions share prefix caches; DFlash/DSpark/DFlash2 speculative decoding; OpenAI- and Anthropic-compatible API on port 10100; one-click connections for Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline. Makers' benchmark vs llama.cpp on Qwen 3.6 35B A3B (4-bit, 64k): 92% faster decode on an M4 Pro, 19% faster on a DGX Spark, and 27% less memory per agent (ai-tldr, developersdigest). Latest v0.2.0, 2026-09-30. Vendor-claimed numbers — not independently verified.

Splash (incoai/splash) — a model-specialised C++ engine for Apple Silicon only. Rather than a general runtime, it pairs a small set of models (Qwen3.8-27B, Qwen3.6-35B-A3B, Ternary Bonsai 2) with a trained DFlash 2 draft model and precompiled, model-specific Metal kernels — no Xcode or local tuning needed. Serves on 127.0.0.1:8000 behind OpenAI Chat Completions/Responses/Completions and Anthropic Messages, with tool calls, JSON Schema, images and inline PDFs. Requires M3+ and macOS 26.4+; 4-bit models need ≥36 GB unified memory. Apache-2.0 (GGUF kernels include MIT-licensed llama.cpp material). Notably, LM Studio 0.4.25 added Splash as a selectable engine, and vllm-metal's own benchmark table lists Splash alongside Lily, Uzu, oMLX, mlx_lm and llama.cpp as an Apple Silicon serving engine (ai-tldr, vllm blog).

vllm-metal (vllm-project/vllm-metal) — vLLM's official Apple Silicon plugin using MLX as the compute backend. v0.28.0 (Sept 22, 2026) aligned versioning with upstream vLLM and added batched multi-token prediction, GGUF and hybrid-model support, and faster M5 prefill; v0.29.0 installs via Homebrew. It brings vLLM's V1 scheduler, paged KV block management, chunked prefill and OpenAI-compatible frontend to Macs, replacing mlx_lm's padded attention with a paged varlen Metal kernel and packing queries into [total_q, H, D] with cu_seqlens. Measured batch wall-time advantage over mlx_lm/oMLX/llama.cpp on Qwen3.6-35B-A3B 4-bit with 8 concurrent requests: 3.87 s vs 4.52 / 7.32 / 5.29 on uniform prompts, and 3.64 s vs 10.99 / 7.22 / 5.33 on ragged prompts — i.e. mlx_lm degraded +143% with ragged prompt lengths while vllm-metal improved 6% (vLLM blog).

MLX / MLX-LM — Apple's own array framework has effectively become the default Mac inference backend. LM Studio, Ollama (MLX engine since March 2026), Jan (experimental), Splash and vllm-metal all build on it. mlx_lm.server is an OpenAI-compatible HTTP server, and WWDC26 sessions 232/233 covered running the full agentic loop locally on a Mac and distributed inference via JACCL (Apple).

Sources

  1. https://github.com/ggml-org/llama.cpp — llama.cpp repo; 130,619 stars, 11,523 commits, active 2026-10-09.
  2. https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md — authoritative list of llama.cpp backends (CUDA, HIP, Metal, Vulkan, SYCL, MUSA, CANN, ZenDNN, OpenCL, OpenVINO, Hexagon, KleidiAI, WebGPU, BLAS vendors).
  3. https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md — llama.cpp multi-GPU split modes (none/layer/row), flags and recipes.
  4. https://github.com/ollama/ollama — Ollama repo; 182,473 stars, releases v0.40.1 (2026-10-07).
  5. https://ollama.com/blog — Ollama blog; MLX engine on Apple Silicon (Mar/Jun 2026), GGUF via llama.cpp (0.30, Jun 2026), $88M raise, 8.9M developers, Anthropic API + Claude Desktop support.
  6. https://docs.ollama.com/gpu — Ollama hardware support: CUDA compute 5.0+, ROCm v7 on Linux and Windows, Metal, Vulkan defaults and GGML_VK_VISIBLE_DEVICES.
  7. https://lmstudio.ai/blog/0.4.0 — LM Studio 0.4.0: llmster headless daemon, parallel requests with continuous batching, stateful /v1/chat, unified KV cache.
  8. https://lmstudio.ai/changelog/lmstudio — LM Studio changelog through 0.4.26 (2026-10-08): Splash engine, DFlash/DSpark/MTP drafters, CUDA 13 on Windows ARM, llama.cpp 2.50.0/b11337.
  9. https://docs.vllm.ai/en/stable/usage/v1_guide/ — vLLM V1 guide: V0 fully deprecated; hardware status (NVIDIA/AMD/Intel/TPU/CPU all functional); plugins.
  10. https://docs.vllm.ai/en/latest/features/quantization/ — vLLM quantization formats and per-hardware compatibility matrix (AWQ, GPTQ, Marlin, FP8, GGUF, etc.).
  11. https://vllm.ai/blog/2026-09-22-vllm-metal-v0-28-0 — Announcing vllm-metal: paged varlen Metal kernel, MTP, GGUF/hybrid support, ragged-batch benchmark vs mlx_lm/oMLX/llama.cpp.
  12. https://www.sglang.io/ — SGLang site: 0.5.21 highlights (Rust prefix-cache core, PD role switching, radix tree), supported hardware NVIDIA/AMD/CPU/TPU/Ascend/XPU.
  13. https://github.com/sgl-project/sglang — SGLang repo; 36,919 stars, active 2026-10-09, DeepSeek-V4 Rust processor parity work.
  14. https://github.com/turboderp-org/exllamav3 — ExLlamaV3 repo: EXL3/QTIP quantisation, 2–8 bit cache quant, active ROCm wheels (Oct 2026), "still in development" caveat.
  15. https://github.com/turboderp-org/exllamav2 — ExLlamaV2 repo: last commit 2026-03-04 — confirms legacy/dormant status.
  16. https://nvidia.github.io/TensorRT-LLM/latest/release-notes.html — TensorRT-LLM 1.3 release notes: TRITON MoE deprecation; 1.2 removed the TensorRT backend entirely, added DGX Spark beta.
  17. https://github.com/mlc-ai/mlc-llm — MLC-LLM repo; 23,226 stars, last commit 2026-10-06, TVM refactor work.
  18. https://github.com/mozilla-ai/llamafile — llamafile repo under Mozilla AI; active 2026-10-08, agent.cpp subtree and "Agentfile" work.
  19. https://github.com/nomic-ai/gpt4all — GPT4All repo: last commit 2025-05-27 — confirms end-of-life status.
  20. https://github.com/mudler/localai — LocalAI repo; 49,450 stars, 8,447 commits, 4.11.0 (Oct 2026), 73 backends.
  21. https://localai.io/blog/what-landed-in-localai-4-8/ — LocalAI 4.8: model variants auto-selection, vllm.cpp (alpha) engine, 3.48× lighter web UI.
  22. https://localai.io/docs/features/vllm-cpp/index.html — vllm.cpp backend docs: Blackwell-only CUDA images (sm_120a/sm_121a), CUDA 13 required, Vulkan/CPU fallback, KV sizing (144 KiB/token).
  23. https://github.com/localai-org/vllm.cpp — vllm.cpp repo: C++20 port of vLLM with continuous batching, paged KV, RadixAttention; 6,363 commits, active 2026-10-09.
  24. https://www.jan.ai/docs/desktop — Jan docs: 0.8.4, Jan Server OpenAI-compatible API, CLI, llama.cpp + experimental MLX engines.
  25. https://freedom.tech/posts/2026-09-26-koboldcpp-1-122-1/ — KoboldCpp 1.122.1: integrated agent with 9 tools, MCP, AGENTS.md, ubatch handling.
  26. https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison/ — Reproducible H100 benchmark of vLLM 0.30 / SGLang 0.5.20 / TensorRT-LLM 1.2.1 / llama.cpp b11179 / Ollama 0.34.4 on Llama 3.1 8B and Qwen3.8-27B.
  27. https://www.soothill.io/blog/2026/08/10/sglang-vllm-llamacpp-evox3/ — 14-point AMD Strix Halo (gfx1151) comparison of SGLang 0.5.17 / vLLM 0.26.0 / llama.cpp b10333; ROCm 7.14 support gaps.
  28. https://ai-tldr.dev/tools/magnitude/ — Magnitude: Rust agent-oriented engine, kernel autotuning, OpenAI+Anthropic API on port 10100, v0.2.0 (2026-09-30), vendor benchmark claims.
  29. https://www.developersdigest.tech/blog/magnitude-self-tuning-local-inference-engine-2026 — Magnitude launch coverage: Apache-2.0, YC-backed, claims up to 2× faster decode than llama.cpp.
  30. https://ai-tldr.dev/tools/inco-splash/ — Splash: Apple Silicon engine with precompiled model-specific Metal kernels, DFlash 2 speculative decoding, M3+/macOS 26.4+.
  31. https://presenc.ai/research/mlx-vs-llama-cpp-throughput-benchmarks-2026 — MLX vs llama.cpp Apple Silicon benchmarks: 10–20% decode advantage above 14B, 4.06× TTFT on M5, crossover past ~40k context.
  32. https://www.glukhov.org/llm-hosting/comparisons/hosting-llms-ollama-localai-jan-lmstudio-vllm-comparison/ — 14-tool comparison across API maturity, tool calling, GPU support, formats, production readiness.
  33. https://docs.litellm.ai/docs/proxy_server — LiteLLM Proxy: OpenAI-compatible gateway over 100+ providers including local engines.
  34. https://particula.tech/blog/ollama-num-ctx-silent-prompt-truncation — Ollama single-slot architectures (mllama, qwen3vl, nemotron_h) forced to one slot since 2026-02-02.
  35. https://helix.ml/blog/the-ceiling-was-a-state-cache — SGLang mamba state cache silently capping concurrency; 3,833 → 7,883 tok/s after fix.
  36. https://markaicode.com/benchmarks/gpt4all-production-benchmark-latency/ — GPT4All end-of-life confirmation and CPU-only status.
  37. https://mortalapps.com/blog/gguf-vs-exl2-vs-mlx-quantization/ — ExLlamaV2 legacy/archived status and EXL3 succession in 2026.
  38. https://atomic.chat/blog/guides/exl3-vs-gguf — EXL3 vs GGUF quality/speed comparison on RTX 5090.
  39. https://localai.io/blog/what-landed-in-localai-4-11/ — LocalAI 4.11: diarisation, speaker profiles, model failover chains, Operate → This machine.
  40. https://developer.apple.com/videos/play/wwdc2026/232/ — WWDC26: MLX-LM and the OpenAI-compatible MLX-LM Server for local agentic AI on Mac.

Date: 2026-10-09

var target=document.getElementById(location.hash.slice(1));target&&target.name&&(target.checked=target.name.startsWith("__tabbed_"))
{"annotate": null, "base": "../../..", "features": [], "search": "../../../assets/javascripts/workers/search.2c215733.min.js", "tags": null, "translations": {"clipboard.copied": "Copied to clipboard", "clipboard.copy": "Copy to clipboard", "search.result.more.one": "1 more on this page", "search.result.more.other": "# more on this page", "search.result.none": "No matching documents", "search.result.one": "1 matching document", "search.result.other": "# matching documents", "search.result.placeholder": "Type to start searching", "search.result.term.missing": "Missing", "select.version": "Select version"}, "version": null}