Local AI Agent Wiki
A knowledge wiki that AI agents can read, analyze, and reproduce local AI workflows on their own machines. Structured markdown, cited sources, and machine-readable endpoints — no JS required.
What this is
A research library for running open-weight LLMs locally
This wiki exists because local AI has too much scattered, conflicting, and unverified information. Every page is built from live sources — vendor docs, GitHub repos, academic papers, and benchmarks — with uncertainty flags preserved inline (verified, unverified, vendor-reported). It is readable by humans and structured for AI agents.
The content model is deliberately sparse: a question, findings, sources. Six categories cover inference engines, open-weight models, hardware, the application stack, practice/ops, and an overview. Each page carries front-matter (title, category, tags, summary, source count, status) so agents can ingest it programmatically.
Six categories
Browse by topic
For agents
Machine-readable endpoints
Agents do not browse HTML. Point your agent at these endpoints to discover the full wiki:
GET /llms.txt
Project index with all 52 pages, one-line descriptions (llms.txt v2 spec).
GET /llms-full.txt
Concatenated markdown of every page — ingest the entire wiki at once.
GET /content/{category}/{slug}/
Any page rendered as HTML. Append .md for raw source.
GET /index.json
JSON index: every page with url, raw-markdown URL, title, summary, tags, status.
Features
Built for AI agents, readable by humans
Every page indexed in llms.txt and llms-full.txt, following the aug-2026 spec with
rel=alternate link headers. Agents discover the entire wiki from one entry point.
Every page is served byte-identical to its source .md at a predictable URL
(/content/{category}/{slug}.md), so agents ingest verified source text — no HTML parsing.
A JSON index (/index.json) exposes every page's url, raw-markdown url, title, summary,
category, tags, last-updated, source count, and verification status — the programmatic entry point.
sitemap.xml + robots.txt (allow all crawlers). The site builds as static HTML
with zero runtime JS for content — readable with JavaScript disabled, fast to crawl.
Docs preview
Top pages — verified research, cited sources
Start with any of these curated entries. Each carries inline uncertainty flags (verified, vendor-reported, unverified) so agents can weight claims appropriately.
inference engines
GGUF/C++ single-user engines, batched GPU serving, and agent-oriented engines. Engine choice is now gated by model architecture.
hardware
VRAM is the binding constraint. RTX 3090 24GB for 32T, RTX 5090 for speed, M5 Ultra for Apple Silicon.
models
Qwen3.8-27B is the standout dense line. MoE for big, dense for local. License flags preserved per model.
application stack
Everything speaks OpenAI-compatible APIs. vLLM → LlamaIndex/CrewAI → MCP → Qdrant → Qwen3-Embedding.
security
Local is private by default, not automatically secure. SafeTensors eliminates the pickle RCE vector. Verify SHA-256 before loading.
Point your agent at this wiki
Give your agent the llms.txt URL and it can discover, read, and reproduce
every local AI workflow documented here — no human browsing required.