Skip to content

Content Map — PHASE 2a Decomposition

Source: Decomposed from 5 research section files + 1 master report in ~/Documents/hermes/research/, all dated 2026-10-09.

Method: Each source file used the consistent ## Question / ## Findings / ## Sources structure. Findings were split into the page granularity specified in the PHASE 0 architecture report (§5.3), preserving all citations, URLs, and unverified/uncertain flags. No research substance was rewritten — only restructured.

Source-to-page mapping

local-ai-state-of-the-art-2026-10-09.md (master report, 18 KB)

Master report section Wiki page(s) it informed
Executive Summary content/_overview/master-report.md (kept whole)
Verified vs Uncertain table content/_overview/master-report.md
Finding 3.1–3.5 (engines, models, hardware, app stack, licensing) Cross-referenced from each topic page
§4 Choosing an engine content/inference-engines/engine-choice.md
§5 Recommendations content/_overview/reading-guide.md
§7 Sources Preserved on every page that draws from the master

section-1-inference-engines.md (38 KB, 40 sources)

Source subsection Wiki page Status
§1 Executive summary + comparison table content/inference-engines/overview.md verified
§2 Engine-by-engine detail content/inference-engines/engines-{llama-cpp,ollama,lm-studio,vllm,sglang,exllamav3,exllamav2-legacy,tensorrt-llm,mlc-llm,koboldcpp,jan,llamafile,gpt4all,localai}.md verified
§3 Serving over OpenAI-compatible endpoints content/inference-engines/serving-openai-compatible.md verified
§4 Newer engines (vllm.cpp, Magnitude, Splash, vllm-metal, MLX) content/inference-engines/emergent-engines.md unverified (vendor claims)
§5 Apple Silicon picture content/hardware/apple-silicon.md verified
§6 Trade-offs (latency/throughput/memory) content/inference-engines/latency-throughput-memory.md verified
§7 Choosing an engine content/inference-engines/engine-choice.md verified

section-2-open-weight-models.md (13.6 KB, 13 sources)

Source subsection Wiki page Status
Executive summary content/models/landscape.md verified
What fits on consumer hardware content/models/what-fits-your-vram.md verified
Key model families table content/models/landscape.md verified
Benchmarks and leaderboards content/models/landscape.md partial (vendor-reported)
Portuguese-language suitability content/models/landscape.md verified
Notable releases content/models/landscape.md verified
Licensing gotchas content/models/licenses-and-restrictions.md partial

section-3-hardware-performance.md (30 KB, 35 sources)

Source subsection Wiki page Status
§1 The VRAM formula content/hardware/vram-math.md verified
§2 KV cache and context-length content/hardware/kv-cache-and-offloading.md verified
§3 Quantization formats content/hardware/quantization-formats.md verified
§4 CPU/RAM offloading + §7 CPU-only content/hardware/cpu-offloading.md verified
§5 Prefill vs decode content/hardware/prefill-decode.md verified
§6 Apple Silicon (from s1 §5 cross-ref) content/hardware/apple-silicon.md verified
§8 Multi-GPU content/hardware/multi-gpu.md verified
§9 AMD ROCm content/hardware/amd-rocms.md verified
§10 Hardware buying guidance content/hardware/buying-guide.md verified

section-4-application-stack.md (22 KB, 50 sources)

Source subsection Wiki page Status
§1 Foundation (OpenAI-compat serving) content/application-stack/overview.md verified
§2 Agent frameworks content/application-stack/agent-frameworks.md partial (AutoGen/slowed)
§3 Coding-agent CLIs content/application-stack/coding-agents.md verified
§4 RAG pipeline content/application-stack/rag-pipeline.md verified
§5 Local fine-tuning content/application-stack/fine-tuning.md verified
§6 Evaluation content/application-stack/evaluation.md verified
§7 Local speech stack content/application-stack/speech-stack.md verified
§8 Local image generation content/application-stack/image-generation.md verified
§9 Multi-agent orchestration content/application-stack/multi-agent.md verified
Maintenance scorecard content/application-stack/overview.md verified

section-5-practice-community-ops.md (31 KB, 36 sources)

Source subsection Wiki page Status
§1 Best-practice guides & community content/practice-ops/learning-resources.md verified
§2 Licensing (open vs open-source) content/practice-ops/licensing-open-vs-open-source.md partial
§3 EU AI Act content/practice-ops/eu-ai-act.md verified
§4 Privacy & security (local vs cloud) content/practice-ops/security-local-vs-cloud.md verified
§5 GPU driver setup content/practice-ops/driver-setup.md verified
§6 Running on Macs content/practice-ops/mac-deployment.md verified
§7 Model management, storage content/practice-ops/model-management.md verified
Monitoring & observability content/practice-ops/monitoring.md verified
§8 Cost: local vs cloud content/practice-ops/cost-vs-cloud.md verified
§9 Where to follow the field content/practice-ops/following-the-field.md verified

Uncertainty preservation

Every page that carries unverified or vendor-reported claims is tagged status: unverified or status: partial in its front-matter. The specific flags from the master report are:

  • Vendor-reported, not independent: Magnitude's "92% faster decode than llama.cpp" (section 1), all benchmark scores from vendor model cards (section 2), Astro's benchmark removals (PHASE 0 notes).
  • Uncertain / early community measurement: M5 Ultra 70B figures (40-52 tok/s) — section 3.
  • Debunked: Llama 5 existence — section 2.

All see_also cross-references point to real page slugs. The link checker (see PHASE 2b scripts/smoke_test.py) validates every internal link on every publish.

Three scripts enforce this, all run from the project root:

Script Checks
scripts/check-wiki.py Front-matter (7 required fields), h1-first + no heading skips, code fences closed and language-tagged, inline .md links resolve inside content/
scripts/check-see-also.py Every see_also slug in front-matter resolves to a real page (95+ cross-links)
scripts/check-nav.py Every content/ page appears in mkdocs.yml nav (reports only — PHASE 1b owns mkdocs.yml)
scripts/build-index.py Regenerates content/_overview/page-index.md from front-matter so the index can never drift

Current status (2026-10-09): check-wiki.py and check-see-also.py both PASS. check-nav.py reports one gap — content/_overview/page-index.md, added by PHASE 2a, is not yet in mkdocs.yml nav. PHASE 1b (t_3ee18994) must add it along with the two plugin-config fixes below.

Master index (requirement 4)

content/_overview/page-index.md is the source-side companion to the published /llms.txt endpoint. It lists all 51 content pages with public URL, category, verification status, source count and one-line summary — generated from front-matter, never hand-maintained. Note it is named page-index.md, not index.md, because index.md inside a category directory collides with MkDocs' directory-index convention and silently fails to render.

Handoffs to PHASE 1b (t_3ee18994) — blocking a --strict build

Two issues in the committed mkdocs.yml prevent a strict build. PHASE 2a did not edit the file (FILE OWNERSHIP rule). PHASE 1b must apply both:

  1. Plugin key is wrong. mkdocs-llmstxt 0.5.1 registers its MkDocs entry point as llmstxt, but mkdocs.yml declares - mkdocs-llmstxt. The build aborts with Config value 'plugins': The "mkdocs-llmstxt" plugin is not installed. Change the key to - llmstxt.
  2. The plugin's required sections option is missing. Once the key is fixed, the build aborts again with Plugin 'llmstxt' option 'sections': Required configuration not provided. A working sections.main list of all 51 pages is in the verification config used for this card: /home/paulo/.hermes/profiles/builder/cache/scratch/wiki-build/mkdocs-verify.yml.
  3. Add content/_overview/page-index.md to the Overview nav section.

Once those land, mkdocs build --strict should pass with only one residual warning, which is not PHASE 2a's to fix: docs/00-architecture-research.md (PHASE 0 deliverable) line 230 links to #8-security--ops, but the target heading is ## 6. Security & ops for the VPS (line 480) — a stale anchor from an earlier revision. PHASE 0's owner should repoint it.

Self-contained workflows (requirement 3)

The seed research corpus contains almost no runnable commands — measured across the five section files: 0 fenced code blocks and 0 pip install lines in section 1, and only inline code spans elsewhere. A wiki whose pages must each carry "install command, config, run command, expected output, pitfalls" could not be produced by restructuring alone.

So PHASE 2a generated the missing operational layer with scripts/add-quickstart.py, which lifts the already-verified facts from each engine page and its numbered Sources list (versions, default ports, CLI verbs, environment variables) and assembles them into an Install / Run / Verify / Expected output / Pitfalls block. 13 engine pages got one; all 13 render.

The facts used are the ones the research verified — e.g. Ollama's OLLAMA_NUM_PARALLEL=1 default and eleven forced single-slot architectures, Ollama serving on :11434, llama.cpp's llama-server on :8080, vLLM 0.31.0 on :8000, SGLang 0.5.21 on :30000, TensorRT-LLM's model-support failure on hybrid linear-attention models, GPT4All's end-of-life. No workflow step was invented: the blocks contain no claim that is not cited on the page itself. The generator is idempotent (re-run replaces, never duplicates) and guards its block with <!-- QUICKSTART:BEGIN/END --> markers.

add-quickstart.py is the natural home for future additions: when a workflow is actually executed and confirmed on the reference machine, add it there and the page status can move from ready to verified.

Residual risk: these quickstarts are assembled from documentation-grade sources, not from a live execution in this environment. That is precisely why the engine pages carry status: ready rather than verified. They should be treated as high-confidence but not machine-confirmed until PHASE 5 runs them.