# Local AI Agent Wiki > A knowledge wiki about running open-weight LLMs locally, built from verified research. Machine-readable for AI agents. Point your agent at `/llms-full.txt` for the full markdown, `/index.json` for the structured index, or follow any `.md` link below to read a page's source text directly. Every page is self-contained with install/run/verify commands and cited sources. ## Inference Engines - [Emergent Agent-Oriented Engines (2026)](https://wiki.imparlabs.com/raw/inference-engines/emergent-engines.md): vllm.cpp, Magnitude, Splash, vllm-metal, MLX — the new agent-oriented engines built for concurrent coding-agent sessions on personal machines. - [Choosing an Inference Engine — Decision Matrix](https://wiki.imparlabs.com/raw/inference-engines/engine-choice.md): When you need one user, use Ollama; concurrency, use vLLM; Apple Silicon agents, use Spllash or vllm-metal. A decision table derived from real H100 benchmarks. - [ExLlamaV3 — Max Quality-Per-Bit on NVIDIA/AMD](https://wiki.imparlabs.com/raw/inference-engines/engines-exllamav3.md): What ExLlamaV3 is, its backends, quantisation, API, multi-GPU, and what it's best for. - [GPT4All — Effectively End-of-Life](https://wiki.imparlabs.com/raw/inference-engines/engines-gpt4all.md): What GPT4All is, its backends, quantisation, API, multi-GPU, and what it's best for. - [Jan — Open-Source ChatGPT Alternative](https://wiki.imparlabs.com/raw/inference-engines/engines-jan.md): What Jan is, its backends, quantisation, API, multi-GPU, and what it's best for. - [KoboldCpp — Single-Binary Runner with Built-In Agent](https://wiki.imparlabs.com/raw/inference-engines/engines-koboldcpp.md): What KoboldCpp is, its backends, quantisation, API, multi-GPU, and what it's best for. - [llama.cpp — The Substrate of Local Inference](https://wiki.imparlabs.com/raw/inference-engines/engines-llama-cpp.md): What llama.cpp is, its backends, quantisation, API, multi-GPU, and what it's best for. - [llamafile — Single-File Model Distribution](https://wiki.imparlabs.com/raw/inference-engines/engines-llamafile.md): What llamafile is, its backends, quantisation, API, multi-GPU, and what it's best for. - [LM Studio — Polished GUI + Headless Daemon](https://wiki.imparlabs.com/raw/inference-engines/engines-lm-studio.md): What LM Studio is, its backends, quantisation, API, multi-GPU, and what it's best for. - [LocalAI — The Run-Anything OpenAI-Compatible Engine](https://wiki.imparlabs.com/raw/inference-engines/engines-localai.md): What LocalAI is, its backends, quantisation, API, multi-GPU, and what it's best for. - [MLC-LLM — Compilation-First Cross-Platform Option](https://wiki.imparlabs.com/raw/inference-engines/engines-mlc-llm.md): What MLC-LLM is, its backends, quantisation, API, multi-GPU, and what it's best for. - [Ollama — The Default Developer Entry Point](https://wiki.imparlabs.com/raw/inference-engines/engines-ollama.md): What Ollama is, its backends, quantisation, API, multi-GPU, and what it's best for. - [SGLang — Prefix-Caching and Hybrid Model Specialist](https://wiki.imparlabs.com/raw/inference-engines/engines-sglang.md): What SGLang is, its backends, quantisation, API, multi-GPU, and what it's best for. - [TensorRT-LLM — NVIDIA's Engine, With Model Support Limits](https://wiki.imparlabs.com/raw/inference-engines/engines-tensorrt-llm.md): What TensorRT-LLM is, its backends, quantisation, API, multi-GPU, and what it's best for. - [vLLM — The Throughput Standard](https://wiki.imparlabs.com/raw/inference-engines/engines-vllm.md): What vLLM is, its backends, quantisation, API, multi-GPU, and what it's best for. - [Engine Trade-Offs: Latency vs Throughput vs Memory](https://wiki.imparlabs.com/raw/inference-engines/latency-throughput-memory.md): The headline 44x engine gaps are a concurrency artefact. At batch 1, engines are near parity. Concurrency, not single-user speed, decides the engine. - [Inference Engines — Overview](https://wiki.imparlabs.com/raw/inference-engines/overview.md): The 2026 local inference stack has bifurcated into GGUF/C++ single-user engines, batched GPU serving engines, and agent-oriented engines; engine choice is now gated by model architecture. - [Serving Local Models Over OpenAI-Compatible Endpoints](https://wiki.imparlabs.com/raw/inference-engines/serving-openai-compatible.md): The 2026 standard: every local engine speaks the OpenAI API shape. How to wire apps, caveats per runtime, and Anthropic Messages API compatibility. ## Models - [Open-Weight Model Landscape](https://wiki.imparlabs.com/raw/models/landscape.md): The field has consolidated around MoE for big models and dense small models for local work. Qwen3.8-27B is the standout single-GPU dense model. - [Model Licenses and Restrictions](https://wiki.imparlabs.com/raw/models/licenses-and-restrictions.md): Apache-2.0/MIT = genuinely open; Llama/Gemma/Qwen-Max/GLM-5.3/Kimi K3 = open-weight but restricted. License checks expire — verify before committing. - [What Fits Your VRAM — Model Sizing by GPU Tier](https://wiki.imparlabs.com/raw/models/what-fits-your-vram.md): 8 GB → 7-8B dense; 16 GB → 12-14B; 24 GB → Qwen3.8-27B; 48 GB → 70B; 128 GB+ → server-class MoE. Q4 quantised. ## Hardware - [AMD ROCm for Local Inference — State of 2026](https://wiki.imparlabs.com/raw/hardware/amd-rocms.md): ROCm 7.2.x is the first release with official RDNA 4 support and claims CUDA parity for inference. Vulkan remains the pragmatic default for llama.cpp. - [Running LLMs on Apple Silicon (MLX, Metal)](https://wiki.imparlabs.com/raw/hardware/apple-silicon.md): MLX is the native Apple Silicon path. M5 Neural Accelerators give up to 3.97x TTFT. Unified-memory sizing: budget ~2x the model requirement. M4 Max beats GB10/Strix Halo on dense models because it has more memory, not more bandwidth. - [Hardware Buying Guide — GPUs for Local LLMs](https://wiki.imparlabs.com/raw/hardware/buying-guide.md): VRAM is king. Used RTX 3090 24GB (~$900) is best value for 32T. RTX 5090 (32GB) for speed but can't hold 70B at Q4. Mac Studio for dense 70B single-user. - [CPU & ARM Inference — Offloading, DDR5, Realistic Speeds](https://wiki.imparlabs.com/raw/hardware/cpu-offloading.md): CPU is bandwidth-bound, not core-bound. Fast DDR5 beats more cores. Offloading splits the model across GPU+CPU but crosses PCIe per token — dramatically slower. - [KV Cache Management — Context Length Trade-Offs](https://wiki.imparlabs.com/raw/hardware/kv-cache-and-offloading.md): The KV cache grows linearly with context length and can rival model weights. GQA/MLA/CLA reduce it. At 128K context, Llama 3.1 70B needs ~43 GB of KV per request. - [Multi-GPU Scaling — Tensor vs Pipeline vs Data Parallelism](https://wiki.imparlabs.com/raw/hardware/multi-gpu.md): TP splits layers across GPUs (best latency, needs fast NVLink). PP splits by layer (low comms, bubbles). DP replicates the full model (throughput only). Consumer: 2x RTX 5090 runs 70B at ~27 tok/s. - [Prefill vs Decode — Token Speed Realities](https://wiki.imparlabs.com/raw/hardware/prefill-decode.md): Prompt processing (prefill) is compute-bound; token generation (decode) is memory-bandwidth-bound. RTX 5090: ~14,000 tok/s prefill vs 290-300 tok/s decode. For RAG/agent workloads prefill is the bottleneck. - [Quantization Formats — Which to Use When](https://wiki.imparlabs.com/raw/hardware/quantization-formats.md): Q4_K_M for portability (pay ~40% throughput), EXL3 for speed on NVIDIA, AWQ/GPTQ for vLLM stacks, FP8 for big models in limited VRAM, NF4 for training only. - [VRAM Math: Sizing Your GPU Before You Buy](https://wiki.imparlabs.com/raw/hardware/vram-math.md): total VRAM = weights + KV cache + overhead. Use 0.6 GB/B for Q4_K_M, not 0.5. KV cache can exceed weight size at long context. ## Application Stack - [Agent Frameworks — LangChain, LlamaIndex, CrewAI, AutoGen](https://wiki.imparlabs.com/raw/application-stack/agent-frameworks.md): MCP is the standard 2026 tool bus (stateless spec since July). CrewAI most active for multi-agent. AutoGen slowing — verify before adopting. Local tool calls are a code-execution surface. - [Coding-Agent CLIs Against Local Endpoints](https://wiki.imparlabs.com/raw/application-stack/coding-agents.md): OpenCode is most active (212k stars). For local coding agents, a strong 30B+ model on vLLM is the minimum for usable function calling. Claude Code works via gateway config. - [Evaluating Local Models — lm-evaluation-harness](https://wiki.imparlabs.com/raw/application-stack/evaluation.md): lm-evaluation-harness is the standard for reproducible open-model benchmarking. Evaluates local models directly via transformers or vLLM backend for fully-offline loops. - [Local Fine-Tuning — LoRA/QLoRA on Consumer GPUs](https://wiki.imparlabs.com/raw/application-stack/fine-tuning.md): Unsloth is fastest for single-GPU LoRA/QLoRA (7-14B on 24GB comfortably, up to 70B via NF4). Axolotl for multi-GPU/production. MLX-LM on Mac. Export to GGUF afterwards. - [Local Image Generation — ComfyUI and FLUX](https://wiki.imparlabs.com/raw/application-stack/image-generation.md): ComfyUI is the de-facto local diffusion stack — node-graph GUI and API. FLUX-dev/schnell class models need 12GB+ consumer GPU. Don't confuse the comfyui PyPI package (v0.0.1, 2024) with the real product. - [Fully-Local Multi-Agent Orchestration](https://wiki.imparlabs.com/raw/application-stack/multi-agent.md): Topology: Qwen3 on vLLM -> LiteLLM proxy -> CrewAI crews -> MCP servers -> Qdrant + Qwen3-Embedding. Speech and image layers bolt on via whisper.cpp and ComfyUI on the same box. - [Local AI Application Stack — Overview](https://wiki.imparlabs.com/raw/application-stack/overview.md): The 2026 local stack is coherent because OpenAI-compatible serving is universal: vLLM for serving, LlamaIndex/CrewAI for agents, Qwen3-Embedding+Qdrant for RAG, Unsloth for fine-tuning, whisper.cpp+Kokoro for speech, ComfyUI for images. - [RAG Pipeline — Embeddings, Vector DBs, Rerankers](https://wiki.imparlabs.com/raw/application-stack/rag-pipeline.md): Recommended local pattern: Qwen3-Embedding -> Qdrant/FAISS -> Qwen3-Reranker -> LLM on vLLM, orchestrated by LlamaIndex. Qwen3-Embedding-8B ranks #1 on MTEB multilingual. Milvus PyPI package is stale — use Qdrant. - [Local Speech Stack — ASR and TTS](https://wiki.imparlabs.com/raw/application-stack/speech-stack.md): whisper.cpp for portable/CPU ASR; faster-whisper if you want GPU batched ASR. Kokoro (via kokoro-onnx) for quality neural TTS; Piper for lightweight embedded. ## Practice & Ops - [Cost — Local Inference vs Cloud APIs (2026 Numbers)](https://wiki.imparlabs.com/raw/practice-ops/cost-vs-cloud.md): Breakeven depends entirely on what you compare. Local wins for frontier-class models at moderate-high utilisation. Cloud wins for sporadic use (<10% utilisation, 2-4 year payback). Local removes GDPR Article 28 obligations. - [GPU Driver Setup — CUDA vs ROCm vs Vulkan](https://wiki.imparlabs.com/raw/practice-ops/driver-setup.md): Breaking change as of CUDA 13.4 Linux/13.1 Windows: the NVIDIA driver is no longer bundled with the CUDA Toolkit. ROCm 7.2.x is first with official RDNA 4 support. Vulkan is the cross-vendor llama.cpp fallback. - [EU AI Act — Compliance Calendar and Self-Hosting Impact](https://wiki.imparlabs.com/raw/practice-ops/eu-ai-act.md): Digital Omnibus (Aug 2026) pushed high-risk rules to Dec 2027/Aug 2028. Article 50 transparency applies Aug 2026. The open-source exemption does NOT cover systemic-risk models (>10^25 FLOP). Fine-tuning + EU market placement can make you a provider. - [Where to Follow the Local AI Field](https://wiki.imparlabs.com/raw/practice-ops/following-the-field.md): r/LocalLLaMA, Hugging Face Daily Papers, Simon Willison's blog, Latent Space podcast, The Batch newsletter, AA-AgentPerf-Local benchmark, llama.cpp release feed. - [Learning Resources & Community — Best-Practice Guides](https://wiki.imparlabs.com/raw/practice-ops/learning-resources.md): Practical starting stack: Ollama for zero-to-running, llama.cpp for efficiency, vLLM for production. r/LocalLLaMA, Simon Willison's blog, Hugging Face courses, Latent Space podcast. - [Licensing — Open vs Open-Source for Local Models](https://wiki.imparlabs.com/raw/practice-ops/licensing-open-vs-open-source.md): Open-source vs open-weight: Apache-2.0/MIT are genuinely OSI-open. Llama 4 has a 700M MAU cap. Gemma 3 carries Google-enforceable restrictions. Downloading locally does not exempt you from usage terms. - [Running on macOS (Apple Silicon) — MLX, M5 Neural Engines](https://wiki.imparlabs.com/raw/practice-ops/mac-deployment.md): MLX is Apple's native framework. M5 Neural Accelerators (macOS 26.2+) give up to 3.97x TTFT via Metal 4 TensorOps. Requires macOS 26.2+ for M5 NPU access. - [Model Management, Storage & Versioning](https://wiki.imparlabs.com/raw/practice-ops/model-management.md): Quantized 7-8B GGUF is ~4-5 GB. Maintain a manifest with SHA-256 digests per model version. Ollama and LM Studio each maintain separate stores — plan for duplication. - [Monitoring & Observability — OTel, Prometheus, vLLM Metrics](https://wiki.imparlabs.com/raw/practice-ops/monitoring.md): OpenTelemetry GenAI semantic conventions are the 2026 standard. vLLM exposes Prometheus metrics at /metrics. Aspire Dashboard via Docker for zero-infrastructure local backend. gen_ai content attributes off by default (sensitive data). - [Security — Local vs Cloud, Supply Chain, SafeTensors](https://wiki.imparlabs.com/raw/practice-ops/security-local-vs-cloud.md): Local is private by default, not automatically secure. Three leak vectors: tool telemetry, malicious model files, network exposure. SafeTensors eliminates the pickle RCE vector. trust_remote_code=True is the second vector. ## Overview - [Running Open-Weight AI Models Locally — State of the Art](https://wiki.imparlabs.com/raw/_overview/master-report.md): Master report: executive summary, verified-vs-uncertain, engine comparison, recommendations for local open-weight LLM deployment as of October 2026. - [Master Index — Every Page in the Local AI Agent Wiki](https://wiki.imparlabs.com/raw/_overview/page-index.md): Machine-readable index of every wiki page with its URL, category, verification status and one-line summary — the llms.txt companion in source form. - [Reading Guide for the Local AI Agent Wiki](https://wiki.imparlabs.com/raw/_overview/reading-guide.md): How to navigate this agent-facing wiki and use its endpoints (llms.txt, raw .md, JSON index) to reproduce local-AI workflows.