Master Index — Every Page in the Local AI Agent Wiki
51 pages across 6 categories. This is the source-side companion to the published /llms.txt endpoint: it is generated from each page's front-matter by scripts/build-index.py, so it cannot drift from the pages themselves.
Pages by category
Inference Engines
| Page | URL | Status | Sources | Summary |
|---|---|---|---|---|
| Emergent Agent-Oriented Engines (2026) | https://wiki.local-ai.dev/inference-engines/emergent-engines/ | unverified | 49 | vllm.cpp, Magnitude, Splash, vllm-metal, MLX — the new agent-oriented engines built for concurrent coding-agent sessions on personal machines. |
| Choosing an Inference Engine — Decision Matrix | https://wiki.local-ai.dev/inference-engines/engine-choice/ | verified | 49 | When you need one user, use Ollama; concurrency, use vLLM; Apple Silicon agents, use Spllash or vllm-metal. A decision table derived from real H100 benchmarks. |
| ExLlamaV3 — Max Quality-Per-Bit on NVIDIA/AMD | https://wiki.local-ai.dev/inference-engines/engines-exllamav3/ | ready | 49 | What ExLlamaV3 is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| GPT4All — Effectively End-of-Life | https://wiki.local-ai.dev/inference-engines/engines-gpt4all/ | ready | 49 | What GPT4All is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| Jan — Open-Source ChatGPT Alternative | https://wiki.local-ai.dev/inference-engines/engines-jan/ | ready | 49 | What Jan is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| KoboldCpp — Single-Binary Runner with Built-In Agent | https://wiki.local-ai.dev/inference-engines/engines-koboldcpp/ | ready | 49 | What KoboldCpp is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| llama.cpp — The Substrate of Local Inference | https://wiki.local-ai.dev/inference-engines/engines-llama-cpp/ | ready | 49 | What llama.cpp is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| llamafile — Single-File Model Distribution | https://wiki.local-ai.dev/inference-engines/engines-llamafile/ | ready | 49 | What llamafile is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| LM Studio — Polished GUI + Headless Daemon | https://wiki.local-ai.dev/inference-engines/engines-lm-studio/ | ready | 49 | What LM Studio is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| LocalAI — The Run-Anything OpenAI-Compatible Engine | https://wiki.local-ai.dev/inference-engines/engines-localai/ | ready | 49 | What LocalAI is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| MLC-LLM — Compilation-First Cross-Platform Option | https://wiki.local-ai.dev/inference-engines/engines-mlc-llm/ | ready | 49 | What MLC-LLM is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| Ollama — The Default Developer Entry Point | https://wiki.local-ai.dev/inference-engines/engines-ollama/ | ready | 49 | What Ollama is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| SGLang — Prefix-Caching and Hybrid Model Specialist | https://wiki.local-ai.dev/inference-engines/engines-sglang/ | ready | 49 | What SGLang is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| TensorRT-LLM — NVIDIA's Engine, With Model Support Limits | https://wiki.local-ai.dev/inference-engines/engines-tensorrt-llm/ | ready | 49 | What TensorRT-LLM is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| vLLM — The Throughput Standard | https://wiki.local-ai.dev/inference-engines/engines-vllm/ | ready | 49 | What vLLM is, its backends, quantisation, API, multi-GPU, and what it's best for. |
| Engine Trade-Offs: Latency vs Throughput vs Memory | https://wiki.local-ai.dev/inference-engines/latency-throughput-memory/ | verified | 49 | The headline 44x engine gaps are a concurrency artefact. At batch 1, engines are near parity. Concurrency, not single-user speed, decides the engine. |
| Inference Engines — Overview | https://wiki.local-ai.dev/inference-engines/overview/ | verified | 40 | The 2026 local inference stack has bifurcated into GGUF/C++ single-user engines, batched GPU serving engines, and agent-oriented engines; engine choice is now gated by model architecture. |
| Serving Local Models Over OpenAI-Compatible Endpoints | https://wiki.local-ai.dev/inference-engines/serving-openai-compatible/ | ready | 49 | The 2026 standard: every local engine speaks the OpenAI API shape. How to wire apps, caveats per runtime, and Anthropic Messages API compatibility. |
Models
| Page | URL | Status | Sources | Summary |
|---|---|---|---|---|
| Open-Weight Model Landscape | https://wiki.local-ai.dev/models/landscape/ | verified | 14 | The field has consolidated around MoE for big models and dense small models for local work. Qwen3.8-27B is the standout single-GPU dense model. |
| Model Licenses and Restrictions | https://wiki.local-ai.dev/models/licenses-and-restrictions/ | partial | 14 | Apache-2.0/MIT = genuinely open; Llama/Gemma/Qwen-Max/GLM-5.3/Kimi K3 = open-weight but restricted. License checks expire — verify before committing. |
| What Fits Your VRAM — Model Sizing by GPU Tier | https://wiki.local-ai.dev/models/what-fits-your-vram/ | verified | 14 | 8 GB → 7-8B dense; 16 GB → 12-14B; 24 GB → Qwen3.8-27B; 48 GB → 70B; 128 GB+ → server-class MoE. Q4 quantised. |
Hardware
| Page | URL | Status | Sources | Summary |
|---|---|---|---|---|
| AMD ROCm for Local Inference — State of 2026 | https://wiki.local-ai.dev/hardware/amd-rocms/ | verified | 44 | ROCm 7.2.x is the first release with official RDNA 4 support and claims CUDA parity for inference. Vulkan remains the pragmatic default for llama.cpp. |
| Running LLMs on Apple Silicon (MLX, Metal) | https://wiki.local-ai.dev/hardware/apple-silicon/ | verified | 49 | MLX is the native Apple Silicon path. M5 Neural Accelerators give up to 3.97x TTFT. Unified-memory sizing: budget ~2x the model requirement. M4 Max beats GB10/Strix Halo on dense models because it has more memory, not more bandwidth. |
| Hardware Buying Guide — GPUs for Local LLMs | https://wiki.local-ai.dev/hardware/buying-guide/ | verified | 44 | VRAM is king. Used RTX 3090 24GB (~$900) is best value for 32T. RTX 5090 (32GB) for speed but can't hold 70B at Q4. Mac Studio for dense 70B single-user. |
| CPU & ARM Inference — Offloading, DDR5, Realistic Speeds | https://wiki.local-ai.dev/hardware/cpu-offloading/ | verified | 44 | CPU is bandwidth-bound, not core-bound. Fast DDR5 beats more cores. Offloading splits the model across GPU+CPU but crosses PCIe per token — dramatically slower. |
| KV Cache Management — Context Length Trade-Offs | https://wiki.local-ai.dev/hardware/kv-cache-and-offloading/ | verified | 44 | The KV cache grows linearly with context length and can rival model weights. GQA/MLA/CLA reduce it. At 128K context, Llama 3.1 70B needs ~43 GB of KV per request. |
| Multi-GPU Scaling — Tensor vs Pipeline vs Data Parallelism | https://wiki.local-ai.dev/hardware/multi-gpu/ | verified | 44 | TP splits layers across GPUs (best latency, needs fast NVLink). PP splits by layer (low comms, bubbles). DP replicates the full model (throughput only). Consumer: 2x RTX 5090 runs 70B at ~27 tok/s. |
| Prefill vs Decode — Token Speed Realities | https://wiki.local-ai.dev/hardware/prefill-decode/ | verified | 44 | Prompt processing (prefill) is compute-bound; token generation (decode) is memory-bandwidth-bound. RTX 5090: ~14,000 tok/s prefill vs 290-300 tok/s decode. For RAG/agent workloads prefill is the bottleneck. |
| Quantization Formats — Which to Use When | https://wiki.local-ai.dev/hardware/quantization-formats/ | verified | 44 | Q4_K_M for portability (pay ~40% throughput), EXL3 for speed on NVIDIA, AWQ/GPTQ for vLLM stacks, FP8 for big models in limited VRAM, NF4 for training only. |
| VRAM Math: Sizing Your GPU Before You Buy | https://wiki.local-ai.dev/hardware/vram-math/ | verified | 44 | total VRAM = weights + KV cache + overhead. Use 0.6 GB/B for Q4_K_M, not 0.5. KV cache can exceed weight size at long context. |
Application Stack
| Page | URL | Status | Sources | Summary |
|---|---|---|---|---|
| Agent Frameworks — LangChain, LlamaIndex, CrewAI, AutoGen | https://wiki.local-ai.dev/application-stack/agent-frameworks/ | partial | 56 | MCP is the standard 2026 tool bus (stateless spec since July). CrewAI most active for multi-agent. AutoGen slowing — verify before adopting. Local tool calls are a code-execution surface. |
| Coding-Agent CLIs Against Local Endpoints | https://wiki.local-ai.dev/application-stack/coding-agents/ | verified | 56 | OpenCode is most active (212k stars). For local coding agents, a strong 30B+ model on vLLM is the minimum for usable function calling. Claude Code works via gateway config. |
| Evaluating Local Models — lm-evaluation-harness | https://wiki.local-ai.dev/application-stack/evaluation/ | verified | 56 | lm-evaluation-harness is the standard for reproducible open-model benchmarking. Evaluates local models directly via transformers or vLLM backend for fully-offline loops. |
| Local Fine-Tuning — LoRA/QLoRA on Consumer GPUs | https://wiki.local-ai.dev/application-stack/fine-tuning/ | verified | 56 | Unsloth is fastest for single-GPU LoRA/QLoRA (7-14B on 24GB comfortably, up to 70B via NF4). Axolotl for multi-GPU/production. MLX-LM on Mac. Export to GGUF afterwards. |
| Local Image Generation — ComfyUI and FLUX | https://wiki.local-ai.dev/application-stack/image-generation/ | verified | 56 | ComfyUI is the de-facto local diffusion stack — node-graph GUI and API. FLUX-dev/schnell class models need 12GB+ consumer GPU. Don't confuse the comfyui PyPI package (v0.0.1, 2024) with the real product. |
| Fully-Local Multi-Agent Orchestration | https://wiki.local-ai.dev/application-stack/multi-agent/ | verified | 56 | Topology: Qwen3 on vLLM -> LiteLLM proxy -> CrewAI crews -> MCP servers -> Qdrant + Qwen3-Embedding. Speech and image layers bolt on via whisper.cpp and ComfyUI on the same box. |
| Local AI Application Stack — Overview | https://wiki.local-ai.dev/application-stack/overview/ | verified | 56 | The 2026 local stack is coherent because OpenAI-compatible serving is universal: vLLM for serving, LlamaIndex/CrewAI for agents, Qwen3-Embedding+Qdrant for RAG, Unsloth for fine-tuning, whisper.cpp+Kokoro for speech, ComfyUI for images. |
| RAG Pipeline — Embeddings, Vector DBs, Rerankers | https://wiki.local-ai.dev/application-stack/rag-pipeline/ | verified | 56 | Recommended local pattern: Qwen3-Embedding -> Qdrant/FAISS -> Qwen3-Reranker -> LLM on vLLM, orchestrated by LlamaIndex. Qwen3-Embedding-8B ranks #1 on MTEB multilingual. Milvus PyPI package is stale — use Qdrant. |
| Local Speech Stack — ASR and TTS | https://wiki.local-ai.dev/application-stack/speech-stack/ | verified | 56 | whisper.cpp for portable/CPU ASR; faster-whisper if you want GPU batched ASR. Kokoro (via kokoro-onnx) for quality neural TTS; Piper for lightweight embedded. |
Practice & Ops
| Page | URL | Status | Sources | Summary |
|---|---|---|---|---|
| Cost — Local Inference vs Cloud APIs (2026 Numbers) | https://wiki.local-ai.dev/practice-ops/cost-vs-cloud/ | verified | 38 | Breakeven depends entirely on what you compare. Local wins for frontier-class models at moderate-high utilisation. Cloud wins for sporadic use (<10% utilisation, 2-4 year payback). Local removes GDPR Article 28 obligations. |
| GPU Driver Setup — CUDA vs ROCm vs Vulkan | https://wiki.local-ai.dev/practice-ops/driver-setup/ | verified | 38 | Breaking change as of CUDA 13.4 Linux/13.1 Windows: the NVIDIA driver is no longer bundled with the CUDA Toolkit. ROCm 7.2.x is first with official RDNA 4 support. Vulkan is the cross-vendor llama.cpp fallback. |
| EU AI Act — Compliance Calendar and Self-Hosting Impact | https://wiki.local-ai.dev/practice-ops/eu-ai-act/ | verified | 38 | Digital Omnibus (Aug 2026) pushed high-risk rules to Dec 2027/Aug 2028. Article 50 transparency applies Aug 2026. The open-source exemption does NOT cover systemic-risk models (>10^25 FLOP). Fine-tuning + EU market placement can make you a provider. |
| Where to Follow the Local AI Field | https://wiki.local-ai.dev/practice-ops/following-the-field/ | verified | 38 | r/LocalLLaMA, Hugging Face Daily Papers, Simon Willison's blog, Latent Space podcast, The Batch newsletter, AA-AgentPerf-Local benchmark, llama.cpp release feed. |
| Learning Resources & Community — Best-Practice Guides | https://wiki.local-ai.dev/practice-ops/learning-resources/ | verified | 38 | Practical starting stack: Ollama for zero-to-running, llama.cpp for efficiency, vLLM for production. r/LocalLLaMA, Simon Willison's blog, Hugging Face courses, Latent Space podcast. |
| Licensing — Open vs Open-Source for Local Models | https://wiki.local-ai.dev/practice-ops/licensing-open-vs-open-source/ | partial | 38 | Open-source vs open-weight: Apache-2.0/MIT are genuinely OSI-open. Llama 4 has a 700M MAU cap. Gemma 3 carries Google-enforceable restrictions. Downloading locally does not exempt you from usage terms. |
| Running on macOS (Apple Silicon) — MLX, M5 Neural Engines | https://wiki.local-ai.dev/practice-ops/mac-deployment/ | verified | 38 | MLX is Apple's native framework. M5 Neural Accelerators (macOS 26.2+) give up to 3.97x TTFT via Metal 4 TensorOps. Requires macOS 26.2+ for M5 NPU access. |
| Model Management, Storage & Versioning | https://wiki.local-ai.dev/practice-ops/model-management/ | verified | 38 | Quantized 7-8B GGUF is ~4-5 GB. Maintain a manifest with SHA-256 digests per model version. Ollama and LM Studio each maintain separate stores — plan for duplication. |
| Monitoring & Observability — OTel, Prometheus, vLLM Metrics | https://wiki.local-ai.dev/practice-ops/monitoring/ | verified | 38 | OpenTelemetry GenAI semantic conventions are the 2026 standard. vLLM exposes Prometheus metrics at /metrics. Aspire Dashboard via Docker for zero-infrastructure local backend. gen_ai content attributes off by default (sensitive data). |
| Security — Local vs Cloud, Supply Chain, SafeTensors | https://wiki.local-ai.dev/practice-ops/security-local-vs-cloud/ | verified | 38 | Local is private by default, not automatically secure. Three leak vectors: tool telemetry, malicious model files, network exposure. SafeTensors eliminates the pickle RCE vector. trust_remote_code=True is the second vector. |
overview
| Page | URL | Status | Sources | Summary |
|---|---|---|---|---|
| Running Open-Weight AI Models Locally — State of the Art | https://wiki.local-ai.dev/_overview/master-report/ | ready | 18 | Master report: executive summary, verified-vs-uncertain, engine comparison, recommendations for local open-weight LLM deployment as of October 2026. |
| Reading Guide for the Local AI Agent Wiki | https://wiki.local-ai.dev/_overview/reading-guide/ | ready | 0 | How to navigate this agent-facing wiki and use its endpoints (llms.txt, raw .md, JSON index) to reproduce local-AI workflows. |
Status legend
- verified — instructions confirmed against live sources on the reference machine.
- partial — from primary docs; not every path tested end-to-end.
- unverified — carried from research; workflow not yet run, or a vendor claim.