Skip to content

Master Index — Every Page in the Local AI Agent Wiki

51 pages across 6 categories. This is the source-side companion to the published /llms.txt endpoint: it is generated from each page's front-matter by scripts/build-index.py, so it cannot drift from the pages themselves.

Pages by category

Inference Engines

Page URL Status Sources Summary
Emergent Agent-Oriented Engines (2026) https://wiki.local-ai.dev/inference-engines/emergent-engines/ unverified 49 vllm.cpp, Magnitude, Splash, vllm-metal, MLX — the new agent-oriented engines built for concurrent coding-agent sessions on personal machines.
Choosing an Inference Engine — Decision Matrix https://wiki.local-ai.dev/inference-engines/engine-choice/ verified 49 When you need one user, use Ollama; concurrency, use vLLM; Apple Silicon agents, use Spllash or vllm-metal. A decision table derived from real H100 benchmarks.
ExLlamaV3 — Max Quality-Per-Bit on NVIDIA/AMD https://wiki.local-ai.dev/inference-engines/engines-exllamav3/ ready 49 What ExLlamaV3 is, its backends, quantisation, API, multi-GPU, and what it's best for.
GPT4All — Effectively End-of-Life https://wiki.local-ai.dev/inference-engines/engines-gpt4all/ ready 49 What GPT4All is, its backends, quantisation, API, multi-GPU, and what it's best for.
Jan — Open-Source ChatGPT Alternative https://wiki.local-ai.dev/inference-engines/engines-jan/ ready 49 What Jan is, its backends, quantisation, API, multi-GPU, and what it's best for.
KoboldCpp — Single-Binary Runner with Built-In Agent https://wiki.local-ai.dev/inference-engines/engines-koboldcpp/ ready 49 What KoboldCpp is, its backends, quantisation, API, multi-GPU, and what it's best for.
llama.cpp — The Substrate of Local Inference https://wiki.local-ai.dev/inference-engines/engines-llama-cpp/ ready 49 What llama.cpp is, its backends, quantisation, API, multi-GPU, and what it's best for.
llamafile — Single-File Model Distribution https://wiki.local-ai.dev/inference-engines/engines-llamafile/ ready 49 What llamafile is, its backends, quantisation, API, multi-GPU, and what it's best for.
LM Studio — Polished GUI + Headless Daemon https://wiki.local-ai.dev/inference-engines/engines-lm-studio/ ready 49 What LM Studio is, its backends, quantisation, API, multi-GPU, and what it's best for.
LocalAI — The Run-Anything OpenAI-Compatible Engine https://wiki.local-ai.dev/inference-engines/engines-localai/ ready 49 What LocalAI is, its backends, quantisation, API, multi-GPU, and what it's best for.
MLC-LLM — Compilation-First Cross-Platform Option https://wiki.local-ai.dev/inference-engines/engines-mlc-llm/ ready 49 What MLC-LLM is, its backends, quantisation, API, multi-GPU, and what it's best for.
Ollama — The Default Developer Entry Point https://wiki.local-ai.dev/inference-engines/engines-ollama/ ready 49 What Ollama is, its backends, quantisation, API, multi-GPU, and what it's best for.
SGLang — Prefix-Caching and Hybrid Model Specialist https://wiki.local-ai.dev/inference-engines/engines-sglang/ ready 49 What SGLang is, its backends, quantisation, API, multi-GPU, and what it's best for.
TensorRT-LLM — NVIDIA's Engine, With Model Support Limits https://wiki.local-ai.dev/inference-engines/engines-tensorrt-llm/ ready 49 What TensorRT-LLM is, its backends, quantisation, API, multi-GPU, and what it's best for.
vLLM — The Throughput Standard https://wiki.local-ai.dev/inference-engines/engines-vllm/ ready 49 What vLLM is, its backends, quantisation, API, multi-GPU, and what it's best for.
Engine Trade-Offs: Latency vs Throughput vs Memory https://wiki.local-ai.dev/inference-engines/latency-throughput-memory/ verified 49 The headline 44x engine gaps are a concurrency artefact. At batch 1, engines are near parity. Concurrency, not single-user speed, decides the engine.
Inference Engines — Overview https://wiki.local-ai.dev/inference-engines/overview/ verified 40 The 2026 local inference stack has bifurcated into GGUF/C++ single-user engines, batched GPU serving engines, and agent-oriented engines; engine choice is now gated by model architecture.
Serving Local Models Over OpenAI-Compatible Endpoints https://wiki.local-ai.dev/inference-engines/serving-openai-compatible/ ready 49 The 2026 standard: every local engine speaks the OpenAI API shape. How to wire apps, caveats per runtime, and Anthropic Messages API compatibility.

Models

Page URL Status Sources Summary
Open-Weight Model Landscape https://wiki.local-ai.dev/models/landscape/ verified 14 The field has consolidated around MoE for big models and dense small models for local work. Qwen3.8-27B is the standout single-GPU dense model.
Model Licenses and Restrictions https://wiki.local-ai.dev/models/licenses-and-restrictions/ partial 14 Apache-2.0/MIT = genuinely open; Llama/Gemma/Qwen-Max/GLM-5.3/Kimi K3 = open-weight but restricted. License checks expire — verify before committing.
What Fits Your VRAM — Model Sizing by GPU Tier https://wiki.local-ai.dev/models/what-fits-your-vram/ verified 14 8 GB → 7-8B dense; 16 GB → 12-14B; 24 GB → Qwen3.8-27B; 48 GB → 70B; 128 GB+ → server-class MoE. Q4 quantised.

Hardware

Page URL Status Sources Summary
AMD ROCm for Local Inference — State of 2026 https://wiki.local-ai.dev/hardware/amd-rocms/ verified 44 ROCm 7.2.x is the first release with official RDNA 4 support and claims CUDA parity for inference. Vulkan remains the pragmatic default for llama.cpp.
Running LLMs on Apple Silicon (MLX, Metal) https://wiki.local-ai.dev/hardware/apple-silicon/ verified 49 MLX is the native Apple Silicon path. M5 Neural Accelerators give up to 3.97x TTFT. Unified-memory sizing: budget ~2x the model requirement. M4 Max beats GB10/Strix Halo on dense models because it has more memory, not more bandwidth.
Hardware Buying Guide — GPUs for Local LLMs https://wiki.local-ai.dev/hardware/buying-guide/ verified 44 VRAM is king. Used RTX 3090 24GB (~$900) is best value for 32T. RTX 5090 (32GB) for speed but can't hold 70B at Q4. Mac Studio for dense 70B single-user.
CPU & ARM Inference — Offloading, DDR5, Realistic Speeds https://wiki.local-ai.dev/hardware/cpu-offloading/ verified 44 CPU is bandwidth-bound, not core-bound. Fast DDR5 beats more cores. Offloading splits the model across GPU+CPU but crosses PCIe per token — dramatically slower.
KV Cache Management — Context Length Trade-Offs https://wiki.local-ai.dev/hardware/kv-cache-and-offloading/ verified 44 The KV cache grows linearly with context length and can rival model weights. GQA/MLA/CLA reduce it. At 128K context, Llama 3.1 70B needs ~43 GB of KV per request.
Multi-GPU Scaling — Tensor vs Pipeline vs Data Parallelism https://wiki.local-ai.dev/hardware/multi-gpu/ verified 44 TP splits layers across GPUs (best latency, needs fast NVLink). PP splits by layer (low comms, bubbles). DP replicates the full model (throughput only). Consumer: 2x RTX 5090 runs 70B at ~27 tok/s.
Prefill vs Decode — Token Speed Realities https://wiki.local-ai.dev/hardware/prefill-decode/ verified 44 Prompt processing (prefill) is compute-bound; token generation (decode) is memory-bandwidth-bound. RTX 5090: ~14,000 tok/s prefill vs 290-300 tok/s decode. For RAG/agent workloads prefill is the bottleneck.
Quantization Formats — Which to Use When https://wiki.local-ai.dev/hardware/quantization-formats/ verified 44 Q4_K_M for portability (pay ~40% throughput), EXL3 for speed on NVIDIA, AWQ/GPTQ for vLLM stacks, FP8 for big models in limited VRAM, NF4 for training only.
VRAM Math: Sizing Your GPU Before You Buy https://wiki.local-ai.dev/hardware/vram-math/ verified 44 total VRAM = weights + KV cache + overhead. Use 0.6 GB/B for Q4_K_M, not 0.5. KV cache can exceed weight size at long context.

Application Stack

Page URL Status Sources Summary
Agent Frameworks — LangChain, LlamaIndex, CrewAI, AutoGen https://wiki.local-ai.dev/application-stack/agent-frameworks/ partial 56 MCP is the standard 2026 tool bus (stateless spec since July). CrewAI most active for multi-agent. AutoGen slowing — verify before adopting. Local tool calls are a code-execution surface.
Coding-Agent CLIs Against Local Endpoints https://wiki.local-ai.dev/application-stack/coding-agents/ verified 56 OpenCode is most active (212k stars). For local coding agents, a strong 30B+ model on vLLM is the minimum for usable function calling. Claude Code works via gateway config.
Evaluating Local Models — lm-evaluation-harness https://wiki.local-ai.dev/application-stack/evaluation/ verified 56 lm-evaluation-harness is the standard for reproducible open-model benchmarking. Evaluates local models directly via transformers or vLLM backend for fully-offline loops.
Local Fine-Tuning — LoRA/QLoRA on Consumer GPUs https://wiki.local-ai.dev/application-stack/fine-tuning/ verified 56 Unsloth is fastest for single-GPU LoRA/QLoRA (7-14B on 24GB comfortably, up to 70B via NF4). Axolotl for multi-GPU/production. MLX-LM on Mac. Export to GGUF afterwards.
Local Image Generation — ComfyUI and FLUX https://wiki.local-ai.dev/application-stack/image-generation/ verified 56 ComfyUI is the de-facto local diffusion stack — node-graph GUI and API. FLUX-dev/schnell class models need 12GB+ consumer GPU. Don't confuse the comfyui PyPI package (v0.0.1, 2024) with the real product.
Fully-Local Multi-Agent Orchestration https://wiki.local-ai.dev/application-stack/multi-agent/ verified 56 Topology: Qwen3 on vLLM -> LiteLLM proxy -> CrewAI crews -> MCP servers -> Qdrant + Qwen3-Embedding. Speech and image layers bolt on via whisper.cpp and ComfyUI on the same box.
Local AI Application Stack — Overview https://wiki.local-ai.dev/application-stack/overview/ verified 56 The 2026 local stack is coherent because OpenAI-compatible serving is universal: vLLM for serving, LlamaIndex/CrewAI for agents, Qwen3-Embedding+Qdrant for RAG, Unsloth for fine-tuning, whisper.cpp+Kokoro for speech, ComfyUI for images.
RAG Pipeline — Embeddings, Vector DBs, Rerankers https://wiki.local-ai.dev/application-stack/rag-pipeline/ verified 56 Recommended local pattern: Qwen3-Embedding -> Qdrant/FAISS -> Qwen3-Reranker -> LLM on vLLM, orchestrated by LlamaIndex. Qwen3-Embedding-8B ranks #1 on MTEB multilingual. Milvus PyPI package is stale — use Qdrant.
Local Speech Stack — ASR and TTS https://wiki.local-ai.dev/application-stack/speech-stack/ verified 56 whisper.cpp for portable/CPU ASR; faster-whisper if you want GPU batched ASR. Kokoro (via kokoro-onnx) for quality neural TTS; Piper for lightweight embedded.

Practice & Ops

Page URL Status Sources Summary
Cost — Local Inference vs Cloud APIs (2026 Numbers) https://wiki.local-ai.dev/practice-ops/cost-vs-cloud/ verified 38 Breakeven depends entirely on what you compare. Local wins for frontier-class models at moderate-high utilisation. Cloud wins for sporadic use (<10% utilisation, 2-4 year payback). Local removes GDPR Article 28 obligations.
GPU Driver Setup — CUDA vs ROCm vs Vulkan https://wiki.local-ai.dev/practice-ops/driver-setup/ verified 38 Breaking change as of CUDA 13.4 Linux/13.1 Windows: the NVIDIA driver is no longer bundled with the CUDA Toolkit. ROCm 7.2.x is first with official RDNA 4 support. Vulkan is the cross-vendor llama.cpp fallback.
EU AI Act — Compliance Calendar and Self-Hosting Impact https://wiki.local-ai.dev/practice-ops/eu-ai-act/ verified 38 Digital Omnibus (Aug 2026) pushed high-risk rules to Dec 2027/Aug 2028. Article 50 transparency applies Aug 2026. The open-source exemption does NOT cover systemic-risk models (>10^25 FLOP). Fine-tuning + EU market placement can make you a provider.
Where to Follow the Local AI Field https://wiki.local-ai.dev/practice-ops/following-the-field/ verified 38 r/LocalLLaMA, Hugging Face Daily Papers, Simon Willison's blog, Latent Space podcast, The Batch newsletter, AA-AgentPerf-Local benchmark, llama.cpp release feed.
Learning Resources & Community — Best-Practice Guides https://wiki.local-ai.dev/practice-ops/learning-resources/ verified 38 Practical starting stack: Ollama for zero-to-running, llama.cpp for efficiency, vLLM for production. r/LocalLLaMA, Simon Willison's blog, Hugging Face courses, Latent Space podcast.
Licensing — Open vs Open-Source for Local Models https://wiki.local-ai.dev/practice-ops/licensing-open-vs-open-source/ partial 38 Open-source vs open-weight: Apache-2.0/MIT are genuinely OSI-open. Llama 4 has a 700M MAU cap. Gemma 3 carries Google-enforceable restrictions. Downloading locally does not exempt you from usage terms.
Running on macOS (Apple Silicon) — MLX, M5 Neural Engines https://wiki.local-ai.dev/practice-ops/mac-deployment/ verified 38 MLX is Apple's native framework. M5 Neural Accelerators (macOS 26.2+) give up to 3.97x TTFT via Metal 4 TensorOps. Requires macOS 26.2+ for M5 NPU access.
Model Management, Storage & Versioning https://wiki.local-ai.dev/practice-ops/model-management/ verified 38 Quantized 7-8B GGUF is ~4-5 GB. Maintain a manifest with SHA-256 digests per model version. Ollama and LM Studio each maintain separate stores — plan for duplication.
Monitoring & Observability — OTel, Prometheus, vLLM Metrics https://wiki.local-ai.dev/practice-ops/monitoring/ verified 38 OpenTelemetry GenAI semantic conventions are the 2026 standard. vLLM exposes Prometheus metrics at /metrics. Aspire Dashboard via Docker for zero-infrastructure local backend. gen_ai content attributes off by default (sensitive data).
Security — Local vs Cloud, Supply Chain, SafeTensors https://wiki.local-ai.dev/practice-ops/security-local-vs-cloud/ verified 38 Local is private by default, not automatically secure. Three leak vectors: tool telemetry, malicious model files, network exposure. SafeTensors eliminates the pickle RCE vector. trust_remote_code=True is the second vector.

overview

Page URL Status Sources Summary
Running Open-Weight AI Models Locally — State of the Art https://wiki.local-ai.dev/_overview/master-report/ ready 18 Master report: executive summary, verified-vs-uncertain, engine comparison, recommendations for local open-weight LLM deployment as of October 2026.
Reading Guide for the Local AI Agent Wiki https://wiki.local-ai.dev/_overview/reading-guide/ ready 0 How to navigate this agent-facing wiki and use its endpoints (llms.txt, raw .md, JSON index) to reproduce local-AI workflows.

Status legend

  • verified — instructions confirmed against live sources on the reference machine.
  • partial — from primary docs; not every path tested end-to-end.
  • unverified — carried from research; workflow not yet run, or a vendor claim.