Local AI Application Stack — Overview
Everything speaks OpenAI-compatible APIs. An app written against the OpenAI SDK runs locally by swapping base_url. The stack: vLLM → LlamaIndex/CrewAI → MCP → Qdrant → Qwen3-Embedding.
The single most important architectural fact about the 2026 local stack is that everything speaks the OpenAI API shape. Nearly every local runtime — llama.cpp, Ollama, vLLM, LM Studio, SGLang, LocalAI — exposes /v1/chat/completions, /v1/completions, /v1/models, and /v1/embeddions, so an application written once against the OpenAI SDK runs unchanged against a local endpoint by only swapping base_url (vllm.ai, ollama.readthedocs.io/en/openai). This is the "drop-in replacement" property that makes the whole local stack composable.
| Runtime | OpenAI endpoint | Best for | Version / last activity (verified) |
|---|---|---|---|
| llama.cpp | llama-server (/v1/chat/completions etc.) |
CPU/Metal/mixed-backend, GGUF, single user, edge | v0.6.0 released Oct 5 2026; repo committed ~1h before Oct 9 2026; 130.6k★ (releases, repo) |
| Ollama | /v1/... at localhost:11434/v1 |
fast local dev, model management, desktop | PyPI ollama v0.6.3 (2026-09-29); compat layer still marked experimental (docs) |
| vLLM | full OpenAI-compatible server | multi-user, high throughput, batched/production serving | PyPI vllm v0.31.0 (2026-10-05) (vllm.ai) |
| SGLang | OpenAI-compatible + native /generate |
latency-critical, speculative decoding | referenced in Qwen docs (HF Qwen3-VL-Embedding) |
| LM Studio / LocalAI | OpenAI-compatible | desktop GUI / drop-in API (daily.dev) | actively developed |
How apps consume it. The standard pattern is to point the OpenAI SDK (or any OpenAI-compatible client) at the local base URL with a dummy API key:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama")
client.chat.completions.create(model="qwen3:8b", messages=[...])
Source: Ollama OpenAI compatibility docs.
Known caveats (still true in 2026). Ollama's compat layer explicitly does not support logprobs, tool_choice, logit_bias, or image URLs (base64 only), and tool-call streaming was still listed as incomplete (Ollama docs). vLLM is the safer choice when you need strict OpenAI parity plus concurrency. A LiteLLM-style proxy layer (PyPI litellm v1.104.2, 2026-10-08) is commonly placed in front to normalise multiple local backends and give apps one stable endpoint.
Maintenance scorecard (as of 2026-10-09, from live version/commit data)
| Tool | Verified signal | Verdict |
|---|---|---|
| llama.cpp | v0.6.0 (Oct 5 2026), commits Oct 9 2026 | ✅ core, extremely active |
| vLLM | v0.31.0 (Oct 5 2026) | ✅ core, active |
| Ollama | v0.6.3 (Sep 29 2026); compat still experimental | ✅ active |
| OpenCode | commits Oct 8 2026, 212k★ | ✅ very active |
| CrewAI | v1.15.26 (Oct 8 2026) | ✅ very active |
| LangChain / LlamaIndex | v1.4.4 / v0.14.25 (Oct 2026) | ✅ active |
| MCP | 2026-07-28 stateless spec + SDKs | ✅ active, now foundational |
| Unsloth | v2026.10.3 (Oct 8 2026) | ✅ very active |
| Axolotl | v0.20.0 (Sep 30 2026) | ✅ active |
| mlx-lm / mlx (Apple) | v0.32.x (Oct 1 2026) | ✅ active |
| ComfyUI | commits Oct 9 2026 | ✅ very active |
| Piper TTS | v1.8.0 (Sep 4 2026) | ✅ active |
| lm-eval-harness | v0.4.13 (Aug 31 2026) | ✅ active (steady) |
| whisper.cpp | commits Oct 9 2026 | ✅ active |
| faster-whisper | v1.2.1 (Oct 31 2025) | ⚠️ stable, slow cadence |
| Aider | v0.86.2 (Feb 12 2026); last commit May 2026 | ⚠️ mature, slowed |
| AutoGen | last PyPI upload Sep 30 2025 | ⚠️ slowing — verify |
| Milvus (PyPI) | v2.3.9 (Nov 2024) | ⚠️ stale on PyPI |
| Kokoro (base PyPI) | v0.9.4 (Apr 2025) | ⚠️ use kokoro-onnx instead |
Bottom line. The 2026 local stack is coherent because OpenAI-compatible serving is universal: pick vLLM for serving, LlamaIndex/CrewAI (backed by MCP) for agents, Qwen3-Embedding + reranker + Qdrant for RAG, Unsloth for 24GB fine-tuning, lm-eval-harness for eval, whisper.cpp + Kokoro/Piper for speech, and ComfyUI + FLUX for images — all wired to one local base URL. The projects to treat with caution are AutoGen, Aider (slowing), faster-whisper (stable but slow-moving), and the Milvus/Kokoro PyPI packages (stale; use alternatives).
Sources
- https://vllm.ai — vLLM: high-throughput OpenAI-compatible local serving engine (docs entry point).
- https://ollama.readthedocs.io/en/openai — Ollama's experimental OpenAI-compatible
/v1/...API and its feature/field gaps. - https://github.com/ggml-org/llama.cpp — llama.cpp repo: v0.6.0 (Oct 2026), 130.6k★, commits within hours of Oct 9 2026.
- https://github.com/ggml-org/llama.cpp/releases — llama.cpp release history confirming v0.6.0 on Oct 5 2026.
- https://pypi.org/project/vllm/ — vLLM v0.31.0, last upload 2026-10-05 (maintenance verification).
- https://pypi.org/project/ollama/ — Ollama v0.6.3, last upload 2026-09-29.
- https://pypi.org/project/litellm/ — LiteLLM v1.104.2, 2026-10-08: proxy layer for normalising local endpoints.
- https://modelcontextprotocol.io/specification/2026-07-28 — MCP 2026-07-28 stateless spec: JSON-RPC, tools/resources/prompts, security model.
- https://blog.cloudflare.com/mcp-v2 — Cloudflare on the 2026-07-28 MCP spec and updated SDKs (stateless MCP).
- https://developers.googleblog.com/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates — Google on MCP stateless updates as foundational agent infra.
- https://www.descope.com/learn/post/mcp — MCP security: local-server privilege and startup-command risks.
- https://pypi.org/project/langchain/ — LangChain v1.4.4, 2026-10-08.
- https://pypi.org/project/llama-index/ — LlamaIndex v0.14.25, 2026-09-21.
- https://pypi.org/project/llama-index-vector-stores-qdrant/ — LlamaIndex↔Qdrant integration v0.10.4, 2026-10-08.
- https://pypi.org/project/crewai/ — CrewAI v1.15.26, 2026-10-08 (multi-agent crews, actively maintained).
- https://pypi.org/project/pyautogen/ — AutoGen v0.10.0, last upload 2025-07-15 (slowing cadence).
- https://pypi.org/project/dspy/ — DSPy v3.4.0, 2026-09-25.
- https://pypi.org/project/haystack-ai/ — Haystack v3.3.0, 2026-10-01.
- https://github.com/sst/opencode — OpenCode (now anomalyco/opencode): open-source coding agent, commits Oct 8 2026, 212k★.
- https://github.com/Aider-AI/aider — Aider: v0.86.2 (Feb 2026), last commit May 2026 — mature but slowed.
- https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk — Anthropic docs: using Claude Code via an OpenAI-compatible SDK/gateway.
- https://github.com/RichardAtCT/claude-code-openai-wrapper — community wrapper to point Claude Code at OpenAI-compatible endpoints.
- https://huggingface.co/Qwen/Qwen3-Embedding-8B — Qwen3-Embedding series (0.6B/4B/8B + rerankers); #1 MTEB multilingual 70.58; instruct adds 1–5%.
- https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B — Qwen3-VL-Embedding/Reranker for multimodal RAG; MMEB-V2 leader.
- https://pypi.org/project/qdrant-client/ — Qdrant client v1.19.1, 2026-09-16.
- https://pypi.org/project/chromadb/ — Chroma v1.5.9, 2026-05-05.
- https://pypi.org/project/faiss-cpu/ — FAISS v1.15.1, 2026-09-16.
- https://pypi.org/project/weaviate-client/ — Weaviate client v4.23.1, 2026-09-07.
- https://pypi.org/project/milvus/ — Milvus v2.3.9, last upload 2024-11-12 (PyPI stale).
- https://pypi.org/project/sentence-transformers/ — sentence-transformers v6.1.0, 2026-09-18 (embedding runtime).
- https://pypi.org/project/unsloth/ — Unsloth v2026.10.3, 2026-10-08 (fast single-GPU LoRA/QLoRA).
- https://github.com/axolotl-ai-cloud/axolotl — Axolotl repo: v0.20.0 (Sep 2026), multi-GPU/FSDP2, QAT, MoE LoRA.
- https://pypi.org/project/axolotl/ — Axolotl v0.20.0, 2026-09-30.
- https://pypi.org/project/mlx-lm/ — MLX-LM v0.32.0, 2026-10-01 (Apple Silicon fine-tuning).
- https://pypi.org/project/mlx/ — MLX v0.32.3, 2026-09-29.
- https://explore.n1n.ai/blog/fine-tune-llm-lora-qlora-guide-2026-2026-04-17 — 2026 VRAM table: QLoRA ~5–8GB for 3–8B, up to 70B at 4-bit.
- https://gigagpu.com/rtx-4090-24gb-for-fine-tuning — RTX 4090 24GB: LoRA to ~14B FP16, QLoRA to ~70B NF4.
- https://www.stack-junkie.com/blog/lora-fine-tuning-consumer-gpu — QLoRA on consumer GPU with Unsloth, export to GGUF for llama.cpp/Ollama.
- https://github.com/EleutherAI/lm-evaluation-harness — lm-evaluation-harness repo (standard open-model benchmarking).
- https://pypi.org/project/lm-eval/ — lm-eval v0.4.13, 2026-08-31 (maintenance verification).
- https://github.com/ggml-org/whisper.cpp — whisper.cpp: active local ASR (commits Oct 2026).
- https://github.com/SYSTRAN/faster-whisper — faster-whisper (CTranslate2) v1.2.1, last upload 2025-10-31.
- https://pypi.org/project/kokoro-onnx/ — kokoro-onnx v0.6.1, 2026-08-19 (active Kokoro TTS runtime).
- https://pypi.org/project/piper-tts/ — Piper TTS v1.8.0, 2026-09-04.
- https://pypi.org/project/mlx-audio/ — mlx-audio v0.5.8, 2026-10-05 (local speech on Apple Silicon).
- https://github.com/comfyanonymous/ComfyUI — ComfyUI (Comfy-Org/ComfyUI): modular diffusion GUI+API, commits Oct 9 2026.
- https://bfl.ai/blog/flux-1-kontext — FLUX.1 Kontext announcement (local/ComfyUI image generation).
- https://bfl.ai/blog/flux-2 — FLUX.2 announcement (Black Forest Labs open weights).
- https://daily.dev/blog/running-llms-locally-ollama-llama-cpp-self-hosted-ai-developers — 2026 overview: Ollama/llama.cpp/LM Studio/LocalAI all expose OpenAI-compatible endpoints.
- https://www.spheron.network/blog/openai-compatible-api-self-hosted — vLLM's OpenAI-compatible server endpoints (
/v1/chat/completionsetc.). - https://www.langchain.com/resources/langchain-vs-llamaindex — LangChain vs LlamaIndex capability comparison for retrieval and agents.