Observed, documented, inferred: a design for codebase documentation you can audit
Most of what a team needs to know about a system is locked in its source code. Developers can read it, slowly. Managers, business stakeholders, support staff and coding agents cannot, so they rely on documents that are missing, out of date or written for someone else. The AI tools that now turn a repository into a wiki help developers and leave everyone else where they were. This is the design I would argue for instead: read the code once into a fact base where every fact carries its evidence and one of three labels, then render that fact base for each audience, locally, on the model the user chooses, without ever changing the source.
In brief
- One fact layer, many views. Extract once; every document, diagram and agent answer reads from the same facts, never from raw source.
- Every claim is verifiable. It links to evidence and is labelled observed, documented or inferred. A claim with no evidence is not written.
- Deterministic first, model last. Indexing spends no model tokens. A model is used only to summarize, group, name and narrate, and its output is always labelled inferred.
- Documents stay current. When a file changes, only what depends on it is regenerated, and what went stale is reported.
- Local and private. The only thing that leaves the machine is text sent to the model the user configured. No telemetry, and the source tree is never modified.
Three problems the wikis leave open
The tools that document a codebase today turn a repository into a developer wiki, and they do it well. They leave three problems open:
- Non-developer readers get nothing, or a one-off summary that is never updated.
- Readers cannot tell which statements were read from the code and which are a model’s interpretation.
- Agents are handed either a wiki written for people or a raw code graph, not structured facts.
Each reader needs something different from the same system:
| Reader | What they need | Output |
|---|---|---|
| Developers | How the system is built and where to change it | Developer docs, diagrams |
| Managers | System summary, health, risks, what changed | Manager brief |
| Business team | What the system does, in plain language | Capability catalog, glossary |
| Operations and support | How to run it and what its errors and logs mean | Operations runbook |
| Coding agents | Structured, cited facts on demand | Knowledge base over MCP |
Audience-flavoured prose is one prompt away; what a prompt cannot give these readers is a reason to trust it. That is the problem the rest of this design is about.
Six principles
- Deterministic first, model last. Everything that can be computed is computed. A model is used only to summarize, group, name and narrate.
- One fact layer, many views. Documents, diagrams and agent answers never read raw source. They read facts.
- Evidence travels with every fact. Provenance and a label are part of the fact, not decoration added at the end.
- Read-only. The source tree is never modified.
- Local first. The only thing that can leave the machine is text sent to the configured model.
- Language-neutral contracts. The index schema, the tool surface and the configuration can be implemented in any language.
One fact layer, many views
The pipeline reads a local folder, with git history and a GitHub remote as optional extras, and ends in three kinds of view.
| Stage | What it does | Uses a model |
|---|---|---|
| Discovery | Finds files, applies ignore rules and the sensitivity policy, detects project structure | No |
| Deterministic extraction | Builds the code graph, package graph, metrics, history signals and inventories | No |
| Index | Stores nodes, edges, the full-text index, summaries and history in one file per source | No |
| Annotation layer | Writes summaries of symbols, files, modules and the repository, lazily and cached | Yes |
| Fact layer | Exposes everything above as facts with provenance and labels | No |
| Views | Render facts as documents, diagrams and knowledge-base answers | Yes, for narrative only |
| Delivery | Writes Markdown, builds the static site, serves MCP and the CLI | No |
A plain folder with no .git directory is the base case, not an edge case. It is indexed fully apart from history; change detection uses a manifest of file content hashes; provenance is file, line range and content hash; and anything history-based says “no history available” rather than estimating. A git repository adds commits, authors and change records on top, and the same folder indexed with and without its .git directory produces the same code graph, with only the history tables differing.
The annotation layer is where the model does its first job. Summaries of symbols, files, modules and the repository live in the index, keyed by content hash, model and prompt version. They are generated lazily, reused until the underlying content changes, and never written into the source tree, so a second request for the same summary costs no tokens and the tree is byte-identical after a run.
The evidence model
Every fact carries exactly one label.
| Label | Meaning | Examples |
|---|---|---|
| Observed | Derived deterministically from code, configuration or history | A call edge, a complexity score, a churn count |
| Documented | Stated in human-written material | A doc comment, a README statement, an ADR, the context file |
| Inferred | Produced by a model | A module summary, a capability description |
Three rules make the labels mean something.
- The weakest label wins. A fact built from several inputs takes the weakest label among them. Inferred is weaker than documented, which is weaker than observed, so a fact built from an observed edge and an inferred summary is inferred.
- No laundering. Model output never becomes documented or observed, even on a later run. Text the tool generated never re-enters the index under a stronger label, so a second run over the tool’s own output does not raise anything.
- One way in for human knowledge. Human-written material in the source, and the context file, are the only documented inputs. A docstring the tool proposes becomes documented only after a person merges it into the source.
The labels are shown in every output: documents, diagrams and MCP results. No output presents an inferred statement without its label, and when evidence is missing the output says so instead of filling the gap.
What a fact carries
A sketch of the record, not a schema.
| Field | Purpose |
|---|---|
| ID | Stable identifier, usable in citations and MCP results |
| Kind | For example: symbol, dependency, metric, summary, capability, inventory entry |
| Subject | The node or nodes the fact is about |
| Statement or value | The content of the fact |
| Label | Observed, documented or inferred |
| Provenance | File, line range, and commit or content hash, for each piece of evidence |
| Derived from | IDs of the facts this one was built from |
| Produced by | Extractor name and version, or model, role and prompt version |
| Freshness | Current, stale or orphaned |
“Derived from” is what makes incremental updates and stale-fact detection possible. When evidence changes, everything downstream of it can be found: the affected edges, summaries, facts and document sections are invalidated, and only those are regenerated, so a one-file change rewrites only the pages that cite facts from that file. Each fact also has a freshness state, current, stale (its evidence changed) or orphaned (its evidence is gone), and an overall percentage per source, which a CI command can turn into an exit code against a threshold.
Where a model is used, and where it is not
| Used | Not used |
|---|---|
| Summaries in the annotation layer | Discovery and parsing |
| Grouping and naming in diagrams | Graph edges, metrics, history signals |
| Narrative text in documents | Diagram structure |
| Capability descriptions | Inventories |
| Answers to questions | Provenance, labels, freshness |
The line matters most for diagrams and numbers. A diagram’s structure, its nodes and edges, is computed from the graph; a model may group and name, but may not add or remove an edge, and an edge that static analysis could not resolve is drawn in a distinct style and labelled inferred. Every number shown to a reader is produced by an index query that can be re-run.
Model access is bring-your-own: a configuration file assigns a model to each role (extract, triage, synthesize, verify, embed), any OpenAI-compatible endpoint can take any role, so a fully local setup is possible, and the tool ships no keys and no bundled model. Model use is lazy, cached by content hash, metered, logged per run and capped by a configurable budget; every command that would call a model says so before it does; and a dry run reports the expected number of calls and tokens without making any. An incremental update costs in proportion to the change, not to the size of the codebase.
The trust boundary
- The dotted lines are the only connections that cross the machine boundary.
- With a local folder and a local model, nothing crosses it.
- The source tree is untrusted input: its content can never direct an agent’s actions.
The rules behind the picture: no telemetry of any kind, and no network destination other than the configured model endpoints and the git host; files matching the sensitivity policy never reach a model, and secrets detected in any file are never written to documents, diagrams or the knowledge base; analysis agents have no shell, no network tools and no write access outside the output folder; MCP runs over stdio by default, and over HTTP only bound to localhost with a token on every request; access tokens for git hosts and model providers are read from the environment or a local secret store and never written to the index or the logs. Output that names individuals, ownership and expertise maps, is off by default and shows aggregated or pseudonymized values until someone opts in.
What each reader gets
Developers get the system overview, the architecture, one page per module found by structure inference, the key flows from the entry points, setup and run, dependencies, and a “where to change what” index, built bottom-up from the annotation layer, with parent pages written from child pages. Every detected module has a page and every page has citations.
Managers get a one-page system summary, a component inventory with size and importance, health signals (hotspots, ownership concentration, bus factor, complexity, test presence), the top risks, and a dependency inventory with versions. The numbers are observed, each produced by an index query that can be re-run; the narrative is inferred and labelled as such; and without git history the history-based sections are replaced by a notice, never an estimate.
The business team gets a plain-language list of what the system does, each capability linked to its entry points and components. Descriptions are inferred unless backed by documented sources or the context file, and no business value, revenue or priority claim appears unless the context file supplies it.
Operations and support get how to build, run and deploy, read from build files, container files and CI workflows; a configuration reference; feature flags; an error catalog; a log-statement catalog with template, level and location; external dependencies; and health endpoints. Every entry is an observed inventory item; only the explanatory text is inferred. A log message copied from the code can be found in the catalog with its location.
Coding agents get the fact layer over MCP, as facts rather than pages: search by text, filtered by label, module and audience; a component, a flow or a document section by stable ID; impact queries (“what is affected if this symbol changes”, “where is this behaviour implemented”); and in every result the label, the provenance and, once computed, the freshness. An agent should be able to answer “what does module X do and where is it used” using only these tools, citing fact IDs. The same query set serves an ask-from-the-CLI command, where a question with no supporting facts returns “no evidence found” and not a guess, and generates AGENTS.md, CLAUDE.md and llms.txt into the output folder, never into the source tree.
Diagrams for all of them: module-level and package-level dependency graphs where every edge exists in the index; architecture diagrams at the C4 context, container and component levels, with deployable units detected from the build system and no model-invented nodes; and sequence diagrams per entry point, traced through the call graph to a depth cap, with unresolved calls visibly marked.
Human knowledge has one way in
The code cannot say why it exists. An optional context file in the source root supplies what it cannot: business goals, component owners, glossary corrections and audience notes. Its content is labelled documented; missing or empty is valid, and business-facing output then says what it could not determine. A glossary correction in the context file changes the term used in generated documents.
The other direction is closed. The tool never writes into source code: no inline comments, no auto-fixes. An opt-in command can write proposed doc comments as a patch file for review, which the tool never applies, and merged comments count as documented on the next run. Generated documents are versioned, human-editable Markdown files, one folder per audience, each recording the index hash, the model and the generation date, and regeneration must not silently overwrite a human edit: the edit survives, or the operator is warned before it is replaced. The index is a rebuildable cache that can be exported, copied and queried elsewhere without re-indexing, and the same content and configuration produce the same index hash.
What done would look like
The design is only worth something if it can be checked. These are the tests I would hold it to, on a few demo repositories plus one plain folder with no git history.
- A full index completes with zero model spend.
- Every statement in the developer docs and the manager brief links to evidence and carries a label, verified by a checker that finds no sentence-level claim without a citation.
- Every number in the manager brief can be reproduced by a query against the index.
- No node, edge or fact exists without provenance, verified by a query that returns zero rows.
- A fact built from mixed inputs takes the weakest label, and a second run over the tool’s own output does not raise any label.
- After a one-file change, only the pages that cite facts from that file are regenerated, and exactly the facts anchored to it are marked stale.
- A coding agent answers questions about an unfamiliar repository using only the MCP server.
- The source tree is byte-identical before and after a run.
- On the folder without git, the run completes and states plainly which history-based sections are unavailable, and no output contains invented history.
- A repository seeded with fake credentials produces no model request and no output containing them.
- A repository containing text written to manipulate an agent produces no action outside reading and writing output.
- A run with local models and a local folder makes no outbound connection.
This is a design, not a release. Nothing above has been measured on a real repository yet, and the point of writing it down first is that each claim in it is a test someone can fail.