Skip to content

Local Speech Stack — ASR and TTS

whisper.cpp is the reference CPU/local Whisper (same repo cadence as llama.cpp). Kokoro (via kokoro-onnx) for quality neural TTS; Piper for lightweight embedded.

Component Tool Status (verified)
ASR whisper.cpp (ggml-org) Very active; same repo cadence as llama.cpp (commits on Oct 9 2026), 130k★ class repo (repo) — the reference CPU/local Whisper
ASR faster-whisper (SYSTRAN, CTranslate2) ⚠️ PyPI v1.2.1 last uploaded 2025-10-31 — stable/mature, slower release cadence (repo)
TTS Kokoro (82M) kokoro v0.9.4 (2025-04-05) on PyPI but the ecosystem moves via kokoro-onnx v0.6.1 (2026-08-19) and mlx-audio v0.5.8 (2026-10-05) — actively used in 2026 (kokoro-onnx)
TTS Piper piper-tts v1.8.0 (2026-09-04) — actively maintained, great for embedded/Home Assistant (voices)
Speech-to-speech / local voice mlx-audio, mlx-vlm v0.7.6 (2026-10-05) Apple Silicon local voice pipeline

Recommended: whisper.cpp for portable/CPU/on-device ASR (the ggml ecosystem keeps it current), faster-whisper if you want easy GPU batched ASR and don't mind the slower cadence; Kokoro (via kokoro-onnx) for quality neural TTS, Piper for lightweight/embedded TTS.

Sources

  • https://vllm.ai — vLLM: high-throughput OpenAI-compatible local serving engine (docs entry point).
  • https://ollama.readthedocs.io/en/openai — Ollama's experimental OpenAI-compatible /v1/... API and its feature/field gaps.
  • https://github.com/ggml-org/llama.cpp — llama.cpp repo: v0.6.0 (Oct 2026), 130.6k★, commits within hours of Oct 9 2026.
  • https://github.com/ggml-org/llama.cpp/releases — llama.cpp release history confirming v0.6.0 on Oct 5 2026.
  • https://pypi.org/project/vllm/ — vLLM v0.31.0, last upload 2026-10-05 (maintenance verification).
  • https://pypi.org/project/ollama/ — Ollama v0.6.3, last upload 2026-09-29.
  • https://pypi.org/project/litellm/ — LiteLLM v1.104.2, 2026-10-08: proxy layer for normalising local endpoints.
  • https://modelcontextprotocol.io/specification/2026-07-28 — MCP 2026-07-28 stateless spec: JSON-RPC, tools/resources/prompts, security model.
  • https://blog.cloudflare.com/mcp-v2 — Cloudflare on the 2026-07-28 MCP spec and updated SDKs (stateless MCP).
  • https://developers.googleblog.com/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates — Google on MCP stateless updates as foundational agent infra.
  • https://www.descope.com/learn/post/mcp — MCP security: local-server privilege and startup-command risks.
  • https://pypi.org/project/langchain/ — LangChain v1.4.4, 2026-10-08.
  • https://pypi.org/project/llama-index/ — LlamaIndex v0.14.25, 2026-09-21.
  • https://pypi.org/project/llama-index-vector-stores-qdrant/ — LlamaIndex↔Qdrant integration v0.10.4, 2026-10-08.
  • https://pypi.org/project/crewai/ — CrewAI v1.15.26, 2026-10-08 (multi-agent crews, actively maintained).
  • https://pypi.org/project/pyautogen/ — AutoGen v0.10.0, last upload 2025-07-15 (slowing cadence).
  • https://pypi.org/project/dspy/ — DSPy v3.4.0, 2026-09-25.
  • https://pypi.org/project/haystack-ai/ — Haystack v3.3.0, 2026-10-01.
  • https://github.com/sst/opencode — OpenCode (now anomalyco/opencode): open-source coding agent, commits Oct 8 2026, 212k★.
  • https://github.com/Aider-AI/aider — Aider: v0.86.2 (Feb 2026), last commit May 2026 — mature but slowed.
  • https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk — Anthropic docs: using Claude Code via an OpenAI-compatible SDK/gateway.
  • https://github.com/RichardAtCT/claude-code-openai-wrapper — community wrapper to point Claude Code at OpenAI-compatible endpoints.
  • https://huggingface.co/Qwen/Qwen3-Embedding-8B — Qwen3-Embedding series (0.6B/4B/8B + rerankers); #1 MTEB multilingual 70.58; instruct adds 1–5%.
  • https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B — Qwen3-VL-Embedding/Reranker for multimodal RAG; MMEB-V2 leader.
  • https://pypi.org/project/qdrant-client/ — Qdrant client v1.19.1, 2026-09-16.
  • https://pypi.org/project/chromadb/ — Chroma v1.5.9, 2026-05-05.
  • https://pypi.org/project/faiss-cpu/ — FAISS v1.15.1, 2026-09-16.
  • https://pypi.org/project/weaviate-client/ — Weaviate client v4.23.1, 2026-09-07.
  • https://pypi.org/project/milvus/ — Milvus v2.3.9, last upload 2024-11-12 (PyPI stale).
  • https://pypi.org/project/sentence-transformers/ — sentence-transformers v6.1.0, 2026-09-18 (embedding runtime).
  • https://pypi.org/project/unsloth/ — Unsloth v2026.10.3, 2026-10-08 (fast single-GPU LoRA/QLoRA).
  • https://github.com/axolotl-ai-cloud/axolotl — Axolotl repo: v0.20.0 (Sep 2026), multi-GPU/FSDP2, QAT, MoE LoRA.
  • https://pypi.org/project/axolotl/ — Axolotl v0.20.0, 2026-09-30.
  • https://pypi.org/project/mlx-lm/ — MLX-LM v0.32.0, 2026-10-01 (Apple Silicon fine-tuning).
  • https://pypi.org/project/mlx/ — MLX v0.32.3, 2026-09-29.
  • https://explore.n1n.ai/blog/fine-tune-llm-lora-qlora-guide-2026-2026-04-17 — 2026 VRAM table: QLoRA ~5–8GB for 3–8B, up to 70B at 4-bit.
  • https://gigagpu.com/rtx-4090-24gb-for-fine-tuning — RTX 4090 24GB: LoRA to ~14B FP16, QLoRA to ~70B NF4.
  • https://www.stack-junkie.com/blog/lora-fine-tuning-consumer-gpu — QLoRA on consumer GPU with Unsloth, export to GGUF for llama.cpp/Ollama.
  • https://github.com/EleutherAI/lm-evaluation-harness — lm-evaluation-harness repo (standard open-model benchmarking).
  • https://pypi.org/project/lm-eval/ — lm-eval v0.4.13, 2026-08-31 (maintenance verification).
  • https://github.com/ggml-org/whisper.cpp — whisper.cpp: active local ASR (commits Oct 2026).
  • https://github.com/SYSTRAN/faster-whisper — faster-whisper (CTranslate2) v1.2.1, last upload 2025-10-31.
  • https://pypi.org/project/kokoro-onnx/ — kokoro-onnx v0.6.1, 2026-08-19 (active Kokoro TTS runtime).
  • https://pypi.org/project/piper-tts/ — Piper TTS v1.8.0, 2026-09-04.
  • https://pypi.org/project/mlx-audio/ — mlx-audio v0.5.8, 2026-10-05 (local speech on Apple Silicon).
  • https://github.com/comfyanonymous/ComfyUI — ComfyUI (Comfy-Org/ComfyUI): modular diffusion GUI+API, commits Oct 9 2026.
  • https://bfl.ai/blog/flux-1-kontext — FLUX.1 Kontext announcement (local/ComfyUI image generation).
  • https://bfl.ai/blog/flux-2 — FLUX.2 announcement (Black Forest Labs open weights).
  • https://daily.dev/blog/running-llms-locally-ollama-llama-cpp-self-hosted-ai-developers — 2026 overview: Ollama/llama.cpp/LM Studio/LocalAI all expose OpenAI-compatible endpoints.
  • https://www.spheron.network/blog/openai-compatible-api-self-hosted — vLLM's OpenAI-compatible server endpoints (/v1/chat/completions etc.).
  • https://www.langchain.com/resources/langchain-vs-llamaindex — LangChain vs LlamaIndex capability comparison for retrieval and agents.