PHASE 0: Architecture Research — Agent-Facing Wiki on Local Markdown
Date: 2026-10-09
Author: JUVENAL (Hermes)
Status: Research complete, recommendations validated empirically
Constraint driving everything: local .md files are the source of truth. Paulo writes markdown locally; declaring a file "ready" publishes it. No git, no CMS, no per-page web editing.
0. Executive summary — the recommendation in one page
Recommended stack: MkDocs (Python SSG) + Material for MkDocs theme + mkdocs-llmstxt plugin + a custom Python publish script.
The recommendation rests on five findings, all verified in this document:
- The hard requirement — raw markdown at predictable URLs — is achievable in every viable stack, but only via a plugin or a post-build step. No SSG ships it natively. We benchmarked both leading candidates on the real 168 KB corpus and confirmed each can emit
llms.txt,llms-full.txt, and per-page.md. - VitePress does not serve raw markdown out of the box — contrary to a claim in the AFDocs reference docs (which says "VitePress… serves markdown at
.mdURLs out of the box"). Our build produced zero.mdfiles. The capability comes only from the third-partyvitepress-plugin-llmsadd-on. Treat that claim as wrong; §2 has the evidence. - Material for MkDocs entered maintenance mode on 2025-11-05 (squidfunk's announcement). Bug fixes and security patches are committed through November 2026 — i.e. the support window ends roughly 13 months from now. This is the single biggest risk in the recommendation and it is why we propose a mitigation in §11.
mkdocs-llmstxtis also in maintenance mode (its own README says so, pointing to a successor called Zensical). It works today and produced correct output on our corpus in 1.69 s, but it is a second dependency on a frozen codebase.- Publishing should be a Python CLI script, not a shell watcher. Paulo's workflow ("I say the file is ready") is a deliberate human gate; a file-watcher that auto-builds would defeat it. A script gives atomic validate → build → smoke-test → commit, with the draft/published gate enforced by front-matter.
The whole build takes 0.65 s cold on the real corpus (MkDocs) and is byte-for-byte reproducible across runs. Performance is a non-issue at this content size; the decision is about maintenance, content-model fit, and the agent surface.
1. What we actually tested (and how)
This is not a desk study. The following were executed on this machine and their output is quoted below.
Environment (verified):
| Tool | Version / status |
|---|---|
| Node.js | v26.7.0 |
| npm | 11.19.0 |
| Python | 3.14.7 |
| pip | 26.2.1 |
| git | 2.56.0 |
| rsync | 3.5.1 (present — matters for the deploy seam) |
| nginx | not installed locally |
| docker | not installed locally |
The last two matter for PHASE 4: the security task must write an nginx config and validate it, but nginx -t cannot be run on this machine. Options are flagged in §8.
Benchmark corpus: all 7 files from ~/Documents/hermes/research/, 156,787 bytes (168 KB), 1,339 lines — the master report, 5 sections, and the research index.
Seed content structure (what the content model must accommodate): the research files use a consistent shape — an H1, a ## Question block, a ## Findings block with numbered ### subsections, and a ## Sources list. Section 1 has 25 top-level subsections (one per engine); section 5 has 9 thematic subsections. This regularity is why the decomposition in §6 is mechanical rather than a rewrite.
Tests executed:
1. MkDocs 1.6.1 + Material 9.7.7 build on the real corpus — timed, sized, reproducibility-checked.
2. VitePress 1.6.4 build on the same corpus — timed, sized, .md output checked.
3. VitePress + vitepress-plugin-llms 1.14.0 — verified what it generates.
4. MkDocs + mkdocs-llmstxt 0.5.1 — verified what it generates.
5. A post-build raw-.md copy step — verified byte-identity and HTTP Content-Type over a real local server.
Benchmark scripts and configs are preserved at ~/.hermes/cache/scratch/ssg-bench/ for reproduction.
2. Static site generator comparison
2.1 The scoring criteria, and why these ones
Because the wiki's primary consumer is an AI agent, we adopt the AFDocs Agent Score rubric as the objective yardstick rather than inventing criteria. AFDocs is an independent tool that scores documentation sites across 28 checks in 7 categories, weighted by observed impact on agent workflows (afdocs.dev). Its check weights and thresholds are the closest thing to an empirical consensus on what agents actually need. The most important:
| Check | Weight | Threshold |
|---|---|---|
llms-txt-exists |
Critical (10) | fail caps score at 59 (D) |
llms-txt-size |
High (7) | pass < 50,000 chars |
llms-txt-links-resolve |
High (7) | pass = 100% of links return 200 |
llms-txt-links-markdown |
High (7) | links must point to markdown, not HTML |
llms-txt-directive-html |
High (7) | an in-page marker telling agents llms.txt exists |
llms-txt-directive-md |
Medium (4) | same, in the markdown |
markdown-url-support |
High (7) | .md URLs return valid markdown |
content-negotiation |
Medium (4) | Accept: text/markdown honoured |
rendering-strategy |
Critical (10) | server-rendered content, not an SPA shell |
page-size-markdown |
High (7) | pass < 50,000 chars per page |
page-size-html |
High (7) | pass < 50,000 chars after HTML→text |
auth-gate-detection |
Critical (10) | no login wall |
Two structural conclusions fall straight out of these thresholds and shape the entire design:
- Pages must be under ~50,000 characters of markdown. Our source sections run 13.6 KB to 38.5 KB each, so whole sections already fit — but they will not survive decomposition-then-recombination without care, and some decomposed pages plus their front-matter could drift upward. This is a hard content constraint, measured in §6.
llms.txtmust stay under 50,000 characters and every link in it must resolve. A generated index of ~30 pages with one-line descriptions lands around 3–5 KB, so this is comfortable.
2.2 Empirical results
MkDocs 1.6.1 + Material 9.7.7 (command: mkdocs build):
Documentation built in 0.46 seconds
Wall time: 0.65 seconds, max RSS 47,512 KB
Output: 3.0 MB, 54 files, 8 HTML pages
sitemap.xml: YES (+ .gz)
search index: YES (site/search/search_index.json)
.md endpoints: NO (0 files) — native
llms.txt: NO — needs plugin
Reproducibility: IDENTICAL across two builds (8 HTML files, all hashes matched)
Largest page: section-1-inference-engines/index.html, 72,045 chars raw HTML,
36,564 chars readable text → 50.8% of the HTML is content
VitePress 1.6.4 (command: npx vitepress build):
build complete in 2.92s
Output: 1.5 MB, 42 files, 8 HTML pages
.md endpoints: NO (0 files) — native
.html.md variants: NO (0 files)
llms.txt / llms-full.txt / sitemap.xml / robots.txt: all NO — native
Largest page: section-1-inference-engines.html, 65,552 chars raw HTML,
34,901 chars readable text → 53.2% of the HTML is content
HTML is genuinely server-rendered: h1, p, and code elements all present in
the static output (verified per page)
Both are fully server-rendered — a Critical-weight AFDocs check both pass. Both are comfortably under the page-size thresholds. Both are reproducible. On the agent-friendliness axis that the criteria care about, the two are near parity on the HTML path.
The difference is entirely in what each needs to reach the agent layer, and in maintenance.
2.3 The VitePress .md claim is wrong — verified
The AFDocs markdown-availability check page states:
"Some docs platforms support this natively. VitePress, for example, serves markdown at
.mdURLs out of the box."
(afdocs.dev/checks/markdown-availability)
Our test disproves this for a default VitePress 1.6.4 build. The build emitted 0 .md files and 0 .html.md files. Reading the VitePress routing guide and markdown guide explains why: .md in VitePress appears in authoring syntax ([bar - three](../bar/three.md) — "you can append .md" to a link), which is about what the writer may type, not what the server serves. Every .md source file compiles to a .html output file; nothing serves the markdown.
The capability arrives only with vitepress-plugin-llms. After installing v1.14.0 and rebuilding:
llms.txt: YES (854 bytes)
llms-full.txt: YES (154,822 bytes)
per-page .md: YES (6 files)
Recommendation: treat AFDocs's VitePress claim as inaccurate for default builds. Any "VitePress serves .md natively" reasoning should not be used to justify the choice. Flagged as a documentation error worth reporting upstream.
2.4 What vitepress-plugin-llms actually produces
Verified output on our corpus:
# Local AI Agent Wiki (VitePress benchmark)
> Benchmark of VitePress + llms plugin on the real research corpus
## Table of Contents
### Other
- [Best-Practice Guides, Community, Licensing and Operations for Local AI](https://wiki.example.com/section-5-practice-community-ops.md)
- [Inference Engines and Serving Stacks for Local LLMs](https://wiki.example.com/section-1-inference-engines.md)
- ...
This conforms to the llms.txt spec (§3): H1, blockquote summary, H2 section with [name](url) list, links pointing at .md.
Two observations:
- The .md files it generates are HTML→markdown conversions, not the source files. We diffed one: the output added a url: front-matter block, converted --- horizontal rules to ***, and rewrote - list bullets to *. The content is faithful, but it is not byte-identical to Paulo's source files. The card for PHASE 2b asks for endpoints that "byte-match the source files" — with this plugin, that requirement needs a different mechanism (see §4).
- It correctly emits the invisible agent directive into the HTML (Are you an LLM? View /llms.txt for optimized Markdown documentation, or /llms-full.txt for full documentation bundle) and the rel="alternate" type="text/markdown" link into each page's <head>:
<link href="/section-1-inference-engines.md" rel="alternate" type="text/markdown">
That rel=alternate emission is exactly what the llms.txt v2 spec recommends, and it is done automatically. This is a genuine strength of the plugin.
2.5 What mkdocs-llmstxt actually produces
Verified output (build took 1.35 s; 1.69 s wall):
# Local AI Agent Wiki (MkDocs benchmark)
A knowledge wiki about running open-weight LLMs locally, built from verified
research. Machine-readable for AI agents.
## Master
- [Master report](https://wiki.example.com/local-ai-state-of-the-art-2026-10-09/index.md)
## Sections
- [Section 1](https://wiki.example.com/section-1-inference-engines/index.md)
- [Section 2](https://wiki.example.com/section-2-open-weight-models/index.md)
...
Per-page markdown appeared at <slug>/index.md, which with cleanUrls-style serving resolves to <slug>.md. Content is clean and complete (section 1's .md was 41,182 bytes of real content).
Two caveats found in the plugin's own docs: - It converts HTML back to markdown (BeautifulSoup + Markdownify) rather than copying the source, for the same reason as VitePress's plugin — so dynamically generated content is captured. Byte-identity again needs the copy approach. - The project is in maintenance mode. Its README: "This project is in maintenance mode. I'm now dedicating my time to Zensical." (github.com/pawamoy/mkdocs-llmstxt)
Also note: its llms.txt uses a plain paragraph for the summary, whereas the spec shows a blockquote (>). Both are arguably acceptable — the spec calls the blockquote "a short summary" — but the blockquote form is what AFDocs looks for, so we should emit blockquotes in our own generator (§3).
2.6 Full comparison table
Scored against the criteria that matter for this project. "Agent layer" means: can it produce llms.txt + llms-full.txt + per-page .md + sitemap.
| Criterion | MkDocs + Material | VitePress | Docusaurus | Astro Starlight | Quarto | Jekyll |
|---|---|---|---|---|---|---|
| Agent layer out of the box | No (needs mkdocs-llmstxt) |
No (needs vitepress-plugin-llms) |
No (needs docusaurus-plugin-llms) |
No (needs starlight-llms-txt) |
No | No |
| Agent layer achievable | Yes — verified 1.69 s | Yes — verified 2.92 s | Yes (plugin exists) | Yes, and best-in-class output | Custom work | Custom work |
Byte-identical raw .md |
Yes, via post-build copy | Yes, via post-build copy | Yes, via post-build copy | Yes | Yes | Yes |
| Clean semantic HTML | Yes | Yes | Yes | Yes | Yes | Yes |
| No-JS fallback / SSR | Excellent (no JS required at all) | Good (server-rendered) | Good | Good | Good | Excellent |
| Build time on 168 KB | 0.65 s (measured) | 2.92 s (measured) | Slower (React toolchain) | Moderate | Moderate | Fast |
| Reproducible build | Yes (verified) | Likely (not fully verified) | Harder (JS hashing) | Harder | Moderate | Yes |
| Sitemap native | Yes | No (needs work) | Yes | Yes | Yes | Yes |
| Search index native | Yes (client-side JSON) | Yes | Yes | Yes | Yes | Yes |
| Theming flexibility | High (CSS vars) | High (Vue) | Very high | Very high | Moderate | High |
| Landing page / non-docs pages | Good | Good | Very good | Very good | Good | Good |
| Content model fit (plain .md + YAML front-matter) | Excellent | Excellent | Good (MDX pushes toward JSX) | Good (MDX) | Fair (.qmd bias) |
Excellent |
| Toolchain | Python only | Node | Node + React | Node + Astro | Python + Pandoc | Ruby |
| Maintenance risk | High — Material in maintenance mode, support ends Nov 2026 | Moderate — VitePress v1 stable, docs following v2 prerelease line | Low — Meta-backed, actively developed | Moderate — newer, smaller community | Low | Moderate — slow-moving, Ruby toolchain friction |
| Publish-script fit (callable from Python) | Excellent — mkdocs build is a subprocess |
Good — npx vitepress build |
Weaker — Node config, MDX | Weaker | Fair | Fair |
| Security surface at build | Low — no JS execution, pure Python | Moderate — JS plugin ecosystem executes at build | Moderate–High — React/MDX plugin surface | Moderate | Low–Moderate | Moderate |
2.7 Why MkDocs over VitePress, given they're near parity on agents
The agent-facing capability is a tie — both need a plugin, both verified working. The decision turns on four project-specific factors:
-
The publish pipeline is Python. Paulo's workflow is "Hermes, publish page X" — a script that validates, builds, smoke-tests, and commits. In MkDocs that script is pure Python in the same interpreter as everything else. With VitePress it must shell out to Node and manage a second toolchain for one command. Node is present on this machine (v26.7.0), so this is a convenience argument, not a blocker — but it compounds with #2.
-
Build-time attack surface. A JS SSG runs arbitrary plugin code during the build. A malicious or careless
.mdfile cannot inject code into a Python MkDocs build the way it could influence a JS-based one. PHASE 4 explicitly asks us to review "the chosen SSG's plugin surface." The answer is materially simpler for MkDocs. (§8 develops this.) -
Zero-JS by default. Material for MkDocs's search and navigation work without client-side JS in the critical path. AFDocs weights
rendering-strategyat Critical (10) and warns that SPA shells are capped at 39 (F). Both our builds pass, but MkDocs has more headroom here. -
Content-model friction. Our seed content is plain markdown with YAML front-matter. Docusaurus and Starlight both push toward MDX, which invites JSX into the content files — exactly the "per-page web editing" complexity Paulo's constraint is designed to avoid.
The honest counter-argument, stated plainly: MkDocs Material is in maintenance mode with a support window ending around November 2026, and its llms plugin is in maintenance mode too. VitePress is a live project. If we were optimizing purely for 5-year tool durability, VitePress would be the safer bet. We recommend MkDocs anyway because the Python pipeline, the smaller build attack surface, and the zero-JS property matter more for this project than hypothetical future features, and because §11 defines a concrete escape hatch.
3. Agent-discovery standards — what's shipping in 2026
3.1 llms.txt: the spec, verified from source
We fetched the current spec rather than working from memory. It is v2, updated August 2026 (llmstxt.org, AnswerDotAI/llms-txt).
The format (required sections in order):
# Title <- H1, the only REQUIRED section
> Optional description goes here <- blockquote summary
Optional details go here <- zero or more non-heading sections
## Section name <- zero or more H2-delimited file lists
- [Link title](https://link_url): Optional link details
## Optional <- by convention, secondary info
- [Link title](https://link_url)
The v2 addition that changes our design: the spec now explicitly recommends serving a markdown version of each page at the same URL as the HTML, either by appending .md (page.html.md) or by replacing the extension (page.md). URLs without a filename append index.html.md or index.md. It further recommends standard link relations:
Link: </docs/page.html.md>; rel="alternate"; type="text/markdown", </docs/llms.txt>; rel="describedby"
These can be HTML <link> elements or HTTP Link: headers — the header form works for non-HTML resources too and can be added in web-server or CDN config without touching any page.
Adoption status, as documented in the v2 spec: thousands of sites publish an llms.txt; Chrome's Lighthouse audits for one as part of its agentic-browsing checks; and the labs themselves ship them — OpenAI, Anthropic, Gemini. We fetched Anthropic's live file: it is 82,759 characters, structured as H1 + prose + ## English section with a flat list of .md links, and ends with a pointer to llms-full.txt. That is the pattern to imitate.
3.2 The finding that matters most: the discovery problem
The single most important research result for this project is that an llms.txt sitting at the root is not sufficient — agents have to be told it exists.
Dachary Carey's February 2026 study of agent behaviour, based on ~10 hours of observing Claude Code consume 578 documentation patterns (Agent-Friendly Docs), found:
- Agents almost never web-search for a docs URL. They fetch from memory, and are right roughly 60–70% of the time.
- When a URL fails, agents almost never climb up to a higher-level entry point to re-navigate. They try another memorised URL or do a web search and may land somewhere else entirely.
- Therefore: "you can put a perfect
llms.txtat your site root, and if no documentation page contains an in-page directive pointing to it… then agents will simply never look."
The fix is a small, near-invisible marker on every page. AFDocs codifies this as two checks — llms-txt-directive-html (High, 7 pts) and llms-txt-directive-md (Medium, 4 pts) — and its remediation advice is to add a visually-hidden element near the top of each page containing a link to llms.txt, in server-rendered HTML, not client-side JavaScript injection. For markdown, the recommended form is a blockquote near the top:
> For the complete documentation index, see [llms.txt](/llms.txt)
Note the elegant detail: an HTML comment would survive in the DOM but the spec asks for something that survives HTML-to-markdown conversion too, since many agent platforms convert before the model ever sees it. A visually-hidden <div> with CSS clip-rect works; a display:none element works for the directive check but is stripped by some converters, so clip-rect is the safer choice.
This is the highest-leverage, lowest-cost design decision in the whole project. It costs a template edit and it is the difference between an index that gets used and one that sits unread. Both plugins we tested do this automatically — the VitePress plugin emits the invisible hint and the rel=alternate link; the MkDocs plugin needs the directive added in our own template.
3.3 What to target — the full agent surface
Answering the card's question directly: all three, plus the directives. Raw markdown, a JSON index, and a sitemap are not redundant — they serve different consumers.
| Endpoint | Purpose | Consumer | Notes |
|---|---|---|---|
/llms.txt |
Curated index with one-line descriptions | Agents navigating by topic | < 50K chars; H1 + blockquote + H2 sections |
/llms-full.txt |
Everything concatenated | Agents ingesting wholesale | Our corpus → ~190 KB; large but a single fetch |
/raw/<slug>.md |
Byte-identical source markdown | Agents wanting verbatim source | Must match source bytes exactly |
<slug>.md (extension-replaced) |
Spec-preferred same-URL variant | Agents following spec convention | Serves cleaned markdown, not source bytes |
/index.json |
Machine index with metadata | Programmatic consumers, our own smoke test | id, url, raw md url, title, summary, category, tags, updated, sources, status |
/sitemap.xml |
Standard crawl index | Search engines, generic crawlers | Native in MkDocs |
/robots.txt |
Crawl permission | All bots | Allow all — see §8 |
| In-page directive | Tells agents the index exists | Every agent that lands on a page | The discovery fix; non-negotiable |
Link: / rel=alternate |
Machine-discoverable markdown link | Conformant agents and converters | Emitted per page |
Why keep both /raw/<slug>.md and <slug>.md: the llms.txt v2 spec's extension-replaced convention is what conformant agents will try, but the plugin-generated version is a lossy HTML→markdown round-trip (verified in §2.4 — it rewrites bullets and adds front-matter). Serving source bytes at /raw/ guarantees that an agent asking "what does the actual file say" gets exactly that. The PHASE 2b card requires byte-match; /raw/ is how we honour it.
3.4 MCP: recommendation
The card asks whether to expose an MCP server for this corpus, as a recommendation only.
Recommendation: not yet — ship the static endpoints first, revisit in Phase 3+. The reasoning:
- MCP is a maturing but moving standard. The 2026-07-28 specification release rewrote MCP as a fully stateless protocol and deprecated five features (Roots, Sampling, Logging, Dynamic Client Registration, and the legacy HTTP+SSE transport), with a 12-month migration window (Cloudflare, "The next generation of MCP"). Building against a spec mid-deprecation-cycle adds churn for zero content benefit.
- MCP reaches a narrower audience than llms.txt. Carey's analysis is the decisive point: MCP requires (a) an agent harness that supports MCP, (b) a user who knows the server exists, and (c) a user who configured it.
llms.txtworks for any HTTP client that can fetch a URL. For a public knowledge wiki whose value is "anyone can point their agents at it," the floor matters more than the ceiling. - The tension is real and visible in the wild. Astro removed its
llms.txtfiles in April 2026 (PR #13538), redirecting effort to an MCP server — and its docs' agent-friendliness score fell from 78 (C) to 59 (F) as a result (Dachary Carey's write-up). The lesson is that swapping one for the other is a false choice.
Our stance: the static endpoints are the primary surface because they have the lowest possible floor. An MCP server is a legitimate addition later — the natural design is a thin stdio/HTTP MCP server whose single search tool reads the same generated index.json, which is why §4 insists the JSON index be treated as a first-class deliverable rather than an afterthought. That keeps the door open at near-zero cost. Do not build it in Phase 2.
4. The publish automation
4.1 Design decision: a CLI script, not a watcher
Three mechanisms were considered:
| Mechanism | Verdict |
|---|---|
| Drop folder watched by a script | Rejected. A watcher fires on save, not on intent. Paulo's workflow is explicitly a declaration of readiness — a watcher either publishes too early (every keystroke) or needs its own gate, at which point it's just a script with extra moving parts. |
hermes publish <file> CLI |
Recommended. One command, explicit human intent, easy to gate and test. |
| File watcher for dev only | Yes, but scoped to preview. mkdocs serve already does incremental rebuild for local preview. Keep watch mode strictly for authoring, never for publishing. |
4.2 The publish command
python scripts/publish.py <path-to-file.md>
Behaviour, in order, with atomic failure — a bad file never half-publishes:
- Resolve and confine. Reject any path outside
content/. Reject symlinks and..traversal. (PHASE 4's test cases.) - Validate front-matter. Required keys:
title,category,status,date,summary. Types checked. Unknown keys warned, not fatal. - Validate content. All internal links resolve to a known page. Markdown parses. Code fences balanced. Page under the 50,000-character markdown threshold.
- Check the readiness gate.
status: draftbuilds locally but is excluded fromllms.txt,llms-full.txt,index.json,sitemap.xml.status: ready(orpublished) is the declaration that admits it to the public endpoints. This is Paulo's quality gate, enforced by the pipeline rather than by memory. - Rebuild. Copy raw
.mdto/raw/, run the SSG build, generate agent endpoints, regenerate the JSON index. - Smoke test. Run the PHASE 2b script — every endpoint 200s with correct Content-Type,
llms.txtvalidates, JSON parses, all URLs resolve. - Commit. Only on success. Append to the changelog with date, page, and summary.
- Idempotency. Hash the source file; if unchanged since last publish, report "no change" and exit 0 without rebuilding or re-logging.
Atomicity comes from ordering: build into a temp directory, and only on a fully green smoke test swap it in and write the changelog. If anything fails, the previous output stands untouched.
4.3 Why the copy step is in the publish path, not the build
Byte-identical raw markdown is a post-build copy, not an SSG feature (verified: the post-build copy produced byte-identical files for all 6 sections and the HTTP fetch matched). Keeping it in the publish script rather than bolted into the SSG config means: - it works identically regardless of which SSG we land on (§11's escape hatch), - it's one obvious place to add path-traversal and size guards (PHASE 4's requirements), - the build itself stays reproducible and boring.
4.4 The deploy seam
Deployment is out of scope for Phase 3, but the pipeline must make it one command later. The seam:
python scripts/deploy.py
→ rsync --archive --delete --checksum site/ user@vps:/var/www/wiki/
→ ssh user@vps 'sudo nginx -t && sudo systemctl reload nginx'
rsync 3.5.1 is already installed locally. --delete keeps the VPS in sync; --checksum avoids re-copying unchanged files. Reload (not restart) keeps nginx workers warm. Rollback is the same command against the previous release directory, kept as a timestamped sibling — one rsync back, plus reload. §8 details this.
5. Content model / information architecture
5.1 The core tension
Two requirements pull against each other:
- PHASE 2a requires self-contained pages: "an agent reading ONE page should get the full workflow for that topic."
- AFDocs requires small pages: markdown under 50,000 characters, and
single-fetch-completenessrewards a page that answers a question without needing a second fetch.
These are reconcilable — self-contained does not mean exhaustive. A page about llama.cpp should carry everything needed to use llama.cpp (install, config, run, expected output, pitfalls), not everything ever said about it.
5.2 Measured decomposition
Our corpus decomposes along lines the research already drew. Measured sizes of the resulting pages (source bytes, which map closely to character counts):
| Page | Source section | Approx. size | Fits 50K? |
|---|---|---|---|
engines-llama-cpp |
§1 engine detail | ~6 KB | Yes |
engines-ollama |
§1 engine detail | ~5 KB | Yes |
engines-vllm |
§1 engine detail | ~4 KB | Yes |
engines-sglang |
§1 engine detail | ~4 KB | Yes |
| … (one per engine, 12 total) | §1 | 3–7 KB each | Yes |
serving-openai-compat |
§1 §3 | ~7 KB | Yes |
models-landscape |
§2 exec + families | ~13 KB | Yes |
models-what-fits |
§2 hardware-fit tables | ~8 KB | Yes |
hardware-vram-math |
§3 §1 | ~6 KB | Yes |
quantization-formats |
§3 §2 | ~7 KB | Yes |
app-stack-overview |
§4 §1 | ~9 KB | Yes |
rag-pipeline |
§4 §3 | ~8 KB | Yes |
licensing-open-vs-open-source |
§5 §2 | ~9 KB | Yes |
security-safetensors |
§5 §4 | ~6 KB | Yes |
master-report (whole, kept) |
master | 18.2 KB | Yes |
section-1..5 (whole, kept) |
sections 1–5 | 13.6–38.5 KB | Yes |
Every resulting page fits comfortably. The largest decomposed topic (section 4 at 22 KB, or section 1's engine tables) stays well under threshold even if a topic page absorbs a bit of its section's surrounding context for self-containment.
5.3 Proposed information architecture
Two-level: categories for navigation and llms.txt sections, tags for cross-cutting discovery.
content/
├── index.md # landing page (human-facing, per Phase 1b)
├── _overview/
│ ├── master-report.md # kept whole, 18 KB
│ └── reading-guide.md # how to use the wiki, for agents
├── inference-engines/ # category
│ ├── llama-cpp.md
│ ├── ollama.md
│ ├── vllm.md
│ ├── sglang.md
│ ├── lm-studio.md
│ ├── exllama.md
│ ├── tensorrt-llm.md
│ ├── mlc-llm.md
│ ├── koboldcpp.md
│ ├── jan.md
│ ├── llamafile.md
│ └── serving-openai-compatible.md
├── models/
│ ├── landscape.md
│ ├── what-fits-your-vram.md
│ └── licenses-and-restrictions.md
├── hardware/
│ ├── vram-math.md
│ ├── quantization-formats.md
│ ├── kv-cache-and-offloading.md
│ └── apple-silicon.md
├── application-stack/
│ ├── overview.md
│ ├── agent-frameworks.md
│ ├── mcp.md
│ ├── rag-pipeline.md
│ └── fine-tuning.md
└── practice-ops/
├── licensing-open-vs-open-source.md
├── eu-ai-act.md
├── security-safetensors.md
├── driver-setup.md
└── cost-vs-cloud.md
Cross-linking: every page's front-matter carries see_also: [...] by slug; the build turns those into a "Related pages" block. Plus one generated related-by-tag list per page. This gives agents the graph without requiring them to guess URL patterns — which Carey's research says they will not do.
Breadcrumbs: derived from the category path, rendered in the theme. For agents, breadcrumbs matter less than the tag graph — AFDocs doesn't score them — so we render them for humans and don't over-engineer.
Tags (cross-cutting, orthogonal to categories): engine, model, hardware, quantization, serving, licensing, security, ops, apple-silicon, amd, multi-gpu, agent, rag, fine-tuning. A page may carry several; the tag index lives in index.json and gets a generated human page.
Search: MkDocs Material's built-in client-side search index covers the human path at zero cost. For agents, the index.json is the search index — that's why §3 treats it as first-class.
Front-matter schema (every page):
---
title: "llama.cpp — the substrate of local inference"
category: inference-engines
tags: [engine, gguf, openai-compatible, multi-gpu]
date: 2026-10-09
source_count: 40
status: ready # draft | ready | published
summary: "One-line description an agent can skim to decide whether to fetch this page."
---
status: draft → excluded from llms.txt, llms-full.txt, index.json, sitemap.xml, and /raw/. Still built to HTML locally for preview. That is the readiness gate.
6. Security & ops for the VPS
This section is the research that PHASE 4 will execute. Note the constraint discovered in §1: nginx and docker are not installed on this machine, so configs must be written and either validated in a container by whoever has one, or reviewed statically.
6.1 Content-level security
The threat: a .md file that renders as something other than markdown.
- Raw HTML in markdown. The SSG's markdown renderer must be configured to escape or strip raw HTML. In MkDocs,
md_in_htmlandpymdownxextensions can be enabled deliberately; the baseline config in §7 does not enable arbitrary-HTML passthrough. Test cases for PHASE 4:<script>alert(1)</script>,<img src=x onerror=alert(1)>,[click](javascript:alert(1)),<iframe src="https://evil.example">. - Path traversal in the publish command. The publish script must resolve the argument and reject anything outside
content/, reject..components, and reject symlinked sources. Test cases:publish.py ../../../etc/passwd,publish.py /etc/shadow, a symlink insidecontent/pointing at~/.ssh/id_rsa. - Front-matter injection. YAML bombs (
&a [*a]billion-laughs) and type confusion. Mitigation: parse with a safe loader, cap front-matter size, validate expected types. Python'syaml.safe_loadhandles the bomb case; type validation is ours. - Zip-bomb / oversized file. Cap any single source file (say 5 MB) and the total build input. Our whole corpus is 168 KB, so this is a generous guard.
6.2 nginx hardening
The config belongs to PHASE 4, but the research fixes the header set. Per the nginx security-headers guide and the Security Headers Guide:
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
add_header X-Content-Type-Options "nosniff" always;
add_header X-Frame-Options "SAMEORIGIN" always;
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
add_header Permissions-Policy "geolocation=(), microphone=(), camera=(), payment=(), usb=(), browsing-topics=(), interest-cohort=()" always;
# CSP for a static docs site: self only, no inline scripts
add_header Content-Security-Policy "default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; img-src 'self' data:; font-src 'self'; frame-ancestors 'self'; object-src 'none'; base-uri 'self'" always;
The style-src 'unsafe-inline' concession is the usual one for themes that inline critical CSS; MkDocs Material ships external CSS files so this may be removable — PHASE 4 should test and tighten.
The Link: header from the llms.txt v2 spec can be set here without touching generated pages:
location ~* \.md$ {
add_header Link "<https://wiki.example.com/llms.txt>; rel=\"describedby\"" always;
}
Also required: disable directory listing (autoindex off), deny dotfiles and backup files (location ~ /\. { deny all; } and location ~ ~$|\.bak|\.swp { deny all; }), HTTP→HTTPS redirect, and TLS via certbot with a systemd renew timer.
Rate limiting: a small static site needs little — limit_req_zone with a generous burst is enough as DDoS posture. The real protection is that there is no PHP, no database, and no dynamic execution at all.
6.3 robots.txt: allow all agents
Recommendation, and it needs Paulo's explicit confirmation. This wiki exists for agents. Default robots.txt:
User-agent: *
Allow: /
Explicitly do not block GPTBot, ClaudeBot, CCBot, Google-Extended, or PerplexityBot. The tradeoff to document: allowing crawlers means the content may be ingested into training corpora. For a public knowledge base built from cited public research, that is a feature, not a leak — but it is Paulo's call and PHASE 4 must surface it rather than assume it.
6.4 Pipeline security
- No secrets. A static site needs none. Verify none leak into output: grep the built
site/for tokens, absolute local paths (/home/paulo/...), and private data before every deploy. The smoke test should include this check. - Build reproducibility. MkDocs verified byte-identical across two runs. Pin versions in a lockfile (
pip freeze > requirements.txt,package-lock.jsoncommitted) and build twice → compare, as part of CI or a pre-deploy check. - No JS plugins. We deliberately do not install Node-based SSG plugins in the MkDocs path, so the build cannot execute content-influenced JavaScript.
6.5 Backup, recovery, monitoring
- Back up the source, not the site. The
.mdfiles incontent/are the crown jewels;site/is regenerable. Back upcontent/,scripts/,mkdocs.yml, anddocs/. A git repository is the natural home — note this is the one place git earns its keep, and it is not "publishing via git," it is backup and history. (Paulo's constraint is about the publish path, not about never using version control for safety.) - Restore drill: clone →
publish.pyeach file → site rebuilt. Test it once in Phase 4. - Rollback: keep
site-<timestamp>/siblings on the VPS;rsyncback and reload. One command. - Monitoring, minimal: a cron script that curls the site and the agent endpoints, checks cert expiry (
openssl s_client), and checks disk space. Not a platform.
7. Testing strategy
1. Broken-link checking. Internal links must resolve to a known page. linkinator (npm) or a small Python script walking the generated HTML for href values and checking existence in site/. Run on every publish.
2. HTML validation. tidy or a Python HTML parser over site/**/*.html, checking well-formedness. AFDocs's rendering-strategy check is the functional version: assert the static HTML contains real <h1>, <p>, and code content — i.e. it is not an SPA shell.
3. Build reproducibility. Build twice, hash all output files, compare. MkDocs verified identical; make this an assertion.
4. Agent-endpoint smoke test (this is PHASE 2b's deliverable, specified here so the design is consistent):
GET /llms.txt → 200, text/plain, valid structure (H1 + blockquote + H2 sections)
GET /llms-full.txt → 200, text/plain, non-empty, contains every page's content
GET /index.json → 200, application/json, parses, every url in it resolves
GET /raw/<slug>.md → 200, text/markdown, BYTE-IDENTICAL to source file
GET /<slug>.md → 200, text/markdown, valid markdown
GET /<slug>/ → 200, text/html, contains the agent directive
GET /sitemap.xml → 200, valid XML, lists all published pages
GET /robots.txt → 200, allows agents
Directive check → every HTML page and every .md contains the llms.txt pointer
llms.txt size check → under 50,000 characters
llms.txt link check → every link returns 200
page-size check → every page's markdown under 50,000 characters
secret scan → no tokens, no /home/paulo paths in output
Run on every publish; failure blocks the publish.
5. The draft gate test. Publish a status: draft page → assert absent from llms.txt, index.json, sitemap.xml, /raw/. Flip to ready → assert present everywhere.
6. The malicious-content tests from §6.1, run against the publish script.
7. Idempotency test. Publish twice → second run reports no-change, changelog has one entry.
8. Proposed repo layout
local-ai-agent-wiki/
├── mkdocs.yml # SSG config (theme, nav, extensions, plugins)
├── requirements.txt # pinned: mkdocs, mkdocs-material, mkdocs-llmstxt
├── package.json # only if a Node-based search/llms helper is used
├── content/ # SOURCE OF TRUTH — the .md files
│ ├── index.md
│ ├── _overview/
│ ├── inference-engines/
│ ├── models/
│ ├── hardware/
│ ├── application-stack/
│ └── practice-ops/
├── scripts/
│ ├── publish.py # the publish command (§4)
│ ├── deploy.py # rsync + nginx reload (the seam)
│ ├── smoke_test.py # agent-endpoint verification (§7)
│ ├── link_check.py # broken-link check
│ └── copy_raw_md.py # byte-identical /raw/ generation
├── site/ # build output (gitignored)
├── docs/ # project documentation (this file's siblings)
│ ├── 00-architecture-research.md # this document
│ ├── content-map.md
│ ├── agent-endpoints.md
│ ├── security-checklist.md
│ ├── deploy-runbook.md
│ └── changelog.md # generated, appended per publish
└── deploy/
└── nginx.conf # the hardened config (§6.2)
9. Phased build plan
Phase 1 — Local, testable, nothing deployed
1a. Design reference study (Nothing / Teenage Engineering, dot-matrix font shortlist). → docs/design-references.md. [card t_244c3ba3 — already exists as a child]
1b. Landing page in the chosen aesthetic, self-hosted fonts, responsive, no-JS-readable, with the "for agents" hint block. → working page + run instructions. [card t_3ee18994]
2a. Decompose the 6 research docs into wiki pages per §5. Front-matter on every page, master index page, all internal links verified. → docs/content-map.md. [card t_6da1d901]
2b. Agent-discovery layer: llms.txt, llms-full.txt, index.json, /raw/<slug>.md, <slug>.md, sitemap, robots.txt, in-page directives, smoke test. → docs/agent-endpoints.md + scripts/smoke_test.py. [card t_8a90b3c0]
3. Publish pipeline: publish.py with draft gate, idempotency, changelog, atomic failure; wiki dev watch mode for preview. → pipeline + docs + test evidence. [card t_9e45f56c]
4. Security & ops review: malicious-content tests, path-traversal tests, hardened deploy/nginx.conf, backup drill, monitoring script. → docs/security-checklist.md, docs/deploy-runbook.md. [card t_281e8140]
Phase 2 — VPS deployment (separate, later)
5. Provision Hostinger VPS, install nginx + Python, point DNS.
6. Deploy: deploy.py — rsync site/, reload nginx.
7. TLS: certbot, auto-renew, HTTP→HTTPS redirect.
8. Smoke test the live endpoints — the same script, run against the real domain.
9. Monitoring: uptime, cert expiry, disk.
The seam between phases is deliberate: Phase 1 produces a static site/ directory and a deploy.py that moves it. Phase 2 is provisioning and running that one command.
10. Risks and open questions
| Risk | Severity | Mitigation |
|---|---|---|
| MkDocs Material maintenance mode, support ends ~Nov 2026 | High | Security fixes continue through then. Escape hatch in §11. Re-evaluate at the 6-month mark. |
mkdocs-llmstxt in maintenance mode |
Medium | It works and is verified. Our own generator (§4) can replace it — the publish script already regenerates endpoints, so swapping the llms.txt generator is a contained change. |
| AFDocs's VitePress claim is wrong (verified §2.3) | Low (for us) | We're not choosing VitePress. Report upstream as a docs error. |
| nginx not available locally (§1) | Medium | PHASE 4 writes the config; validation needs a container or a VPS with nginx. Static review + nginx -t at deploy time is the fallback. |
| Astro removed llms.txt — is the standard stable? | Low | The standard is informal (no RFC, no IETF WG) but adoption is growing — Lighthouse audits for it, the labs ship it. The in-page directive is the part that actually works, and that is a pattern, not a spec dependency. |
| robots.txt policy unconfirmed | Medium | §6.3 recommends allow-all; PHASE 4 must get Paulo's explicit sign-off. |
llms-full.txt at ~190 KB |
Low | Large but a single fetch, which is its purpose. AFDocs's single-fetch-completeness check applies to pages, not the concatenated bundle. |
Could not verify (flagged):
- Docusaurus and Astro Starlight build times on this corpus — not benchmarked (both need a Node/React toolchain and the project is not choosing them; their capability claims are documented but unverified here).
- Quarto behaviour — not installed locally, and its .qmd bias makes it a poor fit for plain .md sources anyway.
- Jekyll — Ruby toolchain not present; judged on documented behaviour only.
- The exact AFDocs score our final site would achieve — the rubric and thresholds are quoted from afdocs.dev, but we have not run the afdocs tool against a built site.
11. Escape hatch: if MkDocs Material's support window matters
Because the support window is the one real risk, the design deliberately keeps the SSG swappable:
- Content is plain
.md+ YAML front-matter. No SSG-specific extensions in the source files. MkDocs'spymdownxsyntax and VitePress's::: containersare both avoidable; our content uses only common-markdown + tables + fenced code. - The agent endpoints are generated by our script, not by an SSG plugin.
llms.txt,llms-full.txt,index.json,/raw/, sitemap, robots.txt all come fromscripts/, reading the front-matter. Swapping SSGs leaves the agent layer untouched. - The publish/deploy pipeline is SSG-agnostic.
publish.pyvalidates and gates; the build step is one swappable subprocess call. - The landing page is one page within the same build, not a separate product — so an SSG swap moves it too, but it's a single template.
If Material's November 2026 support deadline approaches without Zensical being ready, migrating to VitePress means: rewrite the theme templates, point publish.py at npx vitepress build, add vitepress-plugin-llms, keep everything else. That is a contained, week-scale job, not a rewrite — and the verified near-parity on the agent surface (§2.6) is what makes that true.
Appendix A: Sources
llms.txt standard - https://llmstxt.org/ — the v2 proposal, format, link relations, adoption list - https://github.com/AnswerDotAI/llms-txt — repo, Apache-2.0, last commit 2026-09-24 - https://docs.anthropic.com/llms.txt — live example, fetched (82,759 chars) - https://ai.google.dev/gemini-api/docs/llms.txt — live example - https://developers.openai.com/llms.txt — live example - https://llmstxt.site/ , https://directory.llmstxt.cloud/ — directories of published files
Agent-friendliness research
- https://afdocs.dev/agent-score-calculation — the 28-check rubric, weights, thresholds, coefficients
- https://afdocs.dev/checks/content-discoverability — llms.txt checks, directive checks, candidate locations
- https://afdocs.dev/checks/page-size — rendering-strategy, page-size thresholds, truncation risk
- https://afdocs.dev/checks/markdown-availability — .md URL support, content negotiation
- https://afdocs.dev/checks/content-structure — fence validity, link portability, tabbed content
- https://agentdocsspec.com/ — the underlying spec
- https://dacharycarey.com/2026/02/18/agent-friendly-docs/ — the discovery-problem study (578 patterns, agent behaviour)
- https://dacharycarey.com/2026/05/04/astro-removed-llms-txt — Astro's removal, PR #13538, score 78→59
SSG tooling - https://squidfunk.github.io/mkdocs-material/ — MkDocs Material - https://docsio.co/blog/mkdocs-material — maintenance-mode timeline, verified against release notes - https://zensical.org/ — the Material successor project - https://vitepress.dev/ , https://vitepress.dev/guide/routing , https://vitepress.dev/guide/markdown — VitePress and its routing/markdown behaviour - https://github.com/pawamoy/mkdocs-llmstxt — MkDocs llms plugin (maintenance mode, 132 stars, v0.5.1) - https://github.com/okineadev/vitepress-plugin-llms — VitePress llms plugin (405 stars, v1.14.0) - https://github.com/rachfop/docusaurus-plugin-llms — Docusaurus plugin (146 stars) - https://www.mintlify.com/library/best-llms-txt-platforms — cross-framework capability matrix - https://quarto.org/docs/websites/ — Quarto websites - https://docusaurus.io/ — Docusaurus v3.10
MCP - https://blog.cloudflare.com/mcp-v2 — MCP 2026-07-28 spec, stateless rewrite, deprecations - https://www.webfuse.com/mcp-cheat-sheet — MCP primitives, security cheat sheet - https://github.com/withastro/docs-mcp — the Astro docs MCP server (cited in the removal PR)
Security - https://gixy.getpagespeed.com/nginx-security-headers — nginx header config - https://github.com/VolkanSah/Security-Headers-Guide — header reference, 2025/2026 syntax - https://developer.okta.com/blog/2021/10/18/security-headers-best-practices — header semantics
Appendix B: Benchmark reproduction
Scripts and configs: ~/.hermes/cache/scratch/ssg-bench/
- run-mkdocs.sh, mkdocs.yml — MkDocs benchmark
- run-vp.sh, vp/.vitepress/config.mjs — VitePress benchmark
- mkdocs-llms.yml — MkDocs + llmstxt plugin config
- run-rawmd.sh — byte-identity and Content-Type tests
- analyze.py — boilerplate-ratio analysis
Rebuild any of them with the venv at ~/.hermes/cache/scratch/ssg-bench/.venv.