Local AI Application Stack — Overview - Local AI Agent Wiki __md_scope=new URL("../../..",location),__md_hash=e=>[...e].reduce(((e,_)=>(e<<5)-e+_.charCodeAt(0)),0),__md_get=(e,_=localStorage,t=__md_scope)=>JSON.parse(_.getItem(t.pathname+"."+e)),__md_set=(e,_,t=localStorage,a=__md_scope)=>{try{t.setItem(a.pathname+"."+e,JSON.stringify(_))}catch(e){}}
Skip to content

Local AI Application Stack — Overview

Everything speaks OpenAI-compatible APIs. An app written against the OpenAI SDK runs locally by swapping base_url. The stack: vLLM → LlamaIndex/CrewAI → MCP → Qdrant → Qwen3-Embedding.

The single most important architectural fact about the 2026 local stack is that everything speaks the OpenAI API shape. Nearly every local runtime — llama.cpp, Ollama, vLLM, LM Studio, SGLang, LocalAI — exposes /v1/chat/completions, /v1/completions, /v1/models, and /v1/embeddions, so an application written once against the OpenAI SDK runs unchanged against a local endpoint by only swapping base_url (vllm.ai, ollama.readthedocs.io/en/openai). This is the "drop-in replacement" property that makes the whole local stack composable.

Runtime OpenAI endpoint Best for Version / last activity (verified)
llama.cpp llama-server (/v1/chat/completions etc.) CPU/Metal/mixed-backend, GGUF, single user, edge v0.6.0 released Oct 5 2026; repo committed ~1h before Oct 9 2026; 130.6k★ (releases, repo)
Ollama /v1/... at localhost:11434/v1 fast local dev, model management, desktop PyPI ollama v0.6.3 (2026-09-29); compat layer still marked experimental (docs)
vLLM full OpenAI-compatible server multi-user, high throughput, batched/production serving PyPI vllm v0.31.0 (2026-10-05) (vllm.ai)
SGLang OpenAI-compatible + native /generate latency-critical, speculative decoding referenced in Qwen docs (HF Qwen3-VL-Embedding)
LM Studio / LocalAI OpenAI-compatible desktop GUI / drop-in API (daily.dev) actively developed

How apps consume it. The standard pattern is to point the OpenAI SDK (or any OpenAI-compatible client) at the local base URL with a dummy API key:

from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama")
client.chat.completions.create(model="qwen3:8b", messages=[...])

Source: Ollama OpenAI compatibility docs.

Known caveats (still true in 2026). Ollama's compat layer explicitly does not support logprobs, tool_choice, logit_bias, or image URLs (base64 only), and tool-call streaming was still listed as incomplete (Ollama docs). vLLM is the safer choice when you need strict OpenAI parity plus concurrency. A LiteLLM-style proxy layer (PyPI litellm v1.104.2, 2026-10-08) is commonly placed in front to normalise multiple local backends and give apps one stable endpoint.

Maintenance scorecard (as of 2026-10-09, from live version/commit data)

Tool Verified signal Verdict
llama.cpp v0.6.0 (Oct 5 2026), commits Oct 9 2026 ✅ core, extremely active
vLLM v0.31.0 (Oct 5 2026) ✅ core, active
Ollama v0.6.3 (Sep 29 2026); compat still experimental ✅ active
OpenCode commits Oct 8 2026, 212k★ ✅ very active
CrewAI v1.15.26 (Oct 8 2026) ✅ very active
LangChain / LlamaIndex v1.4.4 / v0.14.25 (Oct 2026) ✅ active
MCP 2026-07-28 stateless spec + SDKs ✅ active, now foundational
Unsloth v2026.10.3 (Oct 8 2026) ✅ very active
Axolotl v0.20.0 (Sep 30 2026) ✅ active
mlx-lm / mlx (Apple) v0.32.x (Oct 1 2026) ✅ active
ComfyUI commits Oct 9 2026 ✅ very active
Piper TTS v1.8.0 (Sep 4 2026) ✅ active
lm-eval-harness v0.4.13 (Aug 31 2026) ✅ active (steady)
whisper.cpp commits Oct 9 2026 ✅ active
faster-whisper v1.2.1 (Oct 31 2025) ⚠️ stable, slow cadence
Aider v0.86.2 (Feb 12 2026); last commit May 2026 ⚠️ mature, slowed
AutoGen last PyPI upload Sep 30 2025 ⚠️ slowing — verify
Milvus (PyPI) v2.3.9 (Nov 2024) ⚠️ stale on PyPI
Kokoro (base PyPI) v0.9.4 (Apr 2025) ⚠️ use kokoro-onnx instead

Bottom line. The 2026 local stack is coherent because OpenAI-compatible serving is universal: pick vLLM for serving, LlamaIndex/CrewAI (backed by MCP) for agents, Qwen3-Embedding + reranker + Qdrant for RAG, Unsloth for 24GB fine-tuning, lm-eval-harness for eval, whisper.cpp + Kokoro/Piper for speech, and ComfyUI + FLUX for images — all wired to one local base URL. The projects to treat with caution are AutoGen, Aider (slowing), faster-whisper (stable but slow-moving), and the Milvus/Kokoro PyPI packages (stale; use alternatives).

Sources

  • https://vllm.ai — vLLM: high-throughput OpenAI-compatible local serving engine (docs entry point).
  • https://ollama.readthedocs.io/en/openai — Ollama's experimental OpenAI-compatible /v1/... API and its feature/field gaps.
  • https://github.com/ggml-org/llama.cpp — llama.cpp repo: v0.6.0 (Oct 2026), 130.6k★, commits within hours of Oct 9 2026.
  • https://github.com/ggml-org/llama.cpp/releases — llama.cpp release history confirming v0.6.0 on Oct 5 2026.
  • https://pypi.org/project/vllm/ — vLLM v0.31.0, last upload 2026-10-05 (maintenance verification).
  • https://pypi.org/project/ollama/ — Ollama v0.6.3, last upload 2026-09-29.
  • https://pypi.org/project/litellm/ — LiteLLM v1.104.2, 2026-10-08: proxy layer for normalising local endpoints.
  • https://modelcontextprotocol.io/specification/2026-07-28 — MCP 2026-07-28 stateless spec: JSON-RPC, tools/resources/prompts, security model.
  • https://blog.cloudflare.com/mcp-v2 — Cloudflare on the 2026-07-28 MCP spec and updated SDKs (stateless MCP).
  • https://developers.googleblog.com/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates — Google on MCP stateless updates as foundational agent infra.
  • https://www.descope.com/learn/post/mcp — MCP security: local-server privilege and startup-command risks.
  • https://pypi.org/project/langchain/ — LangChain v1.4.4, 2026-10-08.
  • https://pypi.org/project/llama-index/ — LlamaIndex v0.14.25, 2026-09-21.
  • https://pypi.org/project/llama-index-vector-stores-qdrant/ — LlamaIndex↔Qdrant integration v0.10.4, 2026-10-08.
  • https://pypi.org/project/crewai/ — CrewAI v1.15.26, 2026-10-08 (multi-agent crews, actively maintained).
  • https://pypi.org/project/pyautogen/ — AutoGen v0.10.0, last upload 2025-07-15 (slowing cadence).
  • https://pypi.org/project/dspy/ — DSPy v3.4.0, 2026-09-25.
  • https://pypi.org/project/haystack-ai/ — Haystack v3.3.0, 2026-10-01.
  • https://github.com/sst/opencode — OpenCode (now anomalyco/opencode): open-source coding agent, commits Oct 8 2026, 212k★.
  • https://github.com/Aider-AI/aider — Aider: v0.86.2 (Feb 2026), last commit May 2026 — mature but slowed.
  • https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk — Anthropic docs: using Claude Code via an OpenAI-compatible SDK/gateway.
  • https://github.com/RichardAtCT/claude-code-openai-wrapper — community wrapper to point Claude Code at OpenAI-compatible endpoints.
  • https://huggingface.co/Qwen/Qwen3-Embedding-8B — Qwen3-Embedding series (0.6B/4B/8B + rerankers); #1 MTEB multilingual 70.58; instruct adds 1–5%.
  • https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B — Qwen3-VL-Embedding/Reranker for multimodal RAG; MMEB-V2 leader.
  • https://pypi.org/project/qdrant-client/ — Qdrant client v1.19.1, 2026-09-16.
  • https://pypi.org/project/chromadb/ — Chroma v1.5.9, 2026-05-05.
  • https://pypi.org/project/faiss-cpu/ — FAISS v1.15.1, 2026-09-16.
  • https://pypi.org/project/weaviate-client/ — Weaviate client v4.23.1, 2026-09-07.
  • https://pypi.org/project/milvus/ — Milvus v2.3.9, last upload 2024-11-12 (PyPI stale).
  • https://pypi.org/project/sentence-transformers/ — sentence-transformers v6.1.0, 2026-09-18 (embedding runtime).
  • https://pypi.org/project/unsloth/ — Unsloth v2026.10.3, 2026-10-08 (fast single-GPU LoRA/QLoRA).
  • https://github.com/axolotl-ai-cloud/axolotl — Axolotl repo: v0.20.0 (Sep 2026), multi-GPU/FSDP2, QAT, MoE LoRA.
  • https://pypi.org/project/axolotl/ — Axolotl v0.20.0, 2026-09-30.
  • https://pypi.org/project/mlx-lm/ — MLX-LM v0.32.0, 2026-10-01 (Apple Silicon fine-tuning).
  • https://pypi.org/project/mlx/ — MLX v0.32.3, 2026-09-29.
  • https://explore.n1n.ai/blog/fine-tune-llm-lora-qlora-guide-2026-2026-04-17 — 2026 VRAM table: QLoRA ~5–8GB for 3–8B, up to 70B at 4-bit.
  • https://gigagpu.com/rtx-4090-24gb-for-fine-tuning — RTX 4090 24GB: LoRA to ~14B FP16, QLoRA to ~70B NF4.
  • https://www.stack-junkie.com/blog/lora-fine-tuning-consumer-gpu — QLoRA on consumer GPU with Unsloth, export to GGUF for llama.cpp/Ollama.
  • https://github.com/EleutherAI/lm-evaluation-harness — lm-evaluation-harness repo (standard open-model benchmarking).
  • https://pypi.org/project/lm-eval/ — lm-eval v0.4.13, 2026-08-31 (maintenance verification).
  • https://github.com/ggml-org/whisper.cpp — whisper.cpp: active local ASR (commits Oct 2026).
  • https://github.com/SYSTRAN/faster-whisper — faster-whisper (CTranslate2) v1.2.1, last upload 2025-10-31.
  • https://pypi.org/project/kokoro-onnx/ — kokoro-onnx v0.6.1, 2026-08-19 (active Kokoro TTS runtime).
  • https://pypi.org/project/piper-tts/ — Piper TTS v1.8.0, 2026-09-04.
  • https://pypi.org/project/mlx-audio/ — mlx-audio v0.5.8, 2026-10-05 (local speech on Apple Silicon).
  • https://github.com/comfyanonymous/ComfyUI — ComfyUI (Comfy-Org/ComfyUI): modular diffusion GUI+API, commits Oct 9 2026.
  • https://bfl.ai/blog/flux-1-kontext — FLUX.1 Kontext announcement (local/ComfyUI image generation).
  • https://bfl.ai/blog/flux-2 — FLUX.2 announcement (Black Forest Labs open weights).
  • https://daily.dev/blog/running-llms-locally-ollama-llama-cpp-self-hosted-ai-developers — 2026 overview: Ollama/llama.cpp/LM Studio/LocalAI all expose OpenAI-compatible endpoints.
  • https://www.spheron.network/blog/openai-compatible-api-self-hosted — vLLM's OpenAI-compatible server endpoints (/v1/chat/completions etc.).
  • https://www.langchain.com/resources/langchain-vs-llamaindex — LangChain vs LlamaIndex capability comparison for retrieval and agents.
var target=document.getElementById(location.hash.slice(1));target&&target.name&&(target.checked=target.name.startsWith("__tabbed_"))
{"annotate": null, "base": "../../..", "features": [], "search": "../../../assets/javascripts/workers/search.2c215733.min.js", "tags": null, "translations": {"clipboard.copied": "Copied to clipboard", "clipboard.copy": "Copy to clipboard", "search.result.more.one": "1 more on this page", "search.result.more.other": "# more on this page", "search.result.none": "No matching documents", "search.result.one": "1 matching document", "search.result.other": "# matching documents", "search.result.placeholder": "Type to start searching", "search.result.term.missing": "Missing", "select.version": "Select version"}, "version": null}