Roteiro

The pilot book for your codebase.

The Portuguese roteiros were the guarded pilot books of the Age of Discovery — accumulated route knowledge that made navigation repeatable. Roteiro does the same for software: structure, intent, and context in one queryable, provenance-tagged knowledge graph, for humans and AI agents alike.

What Roteiro is

Roteiro reads a git repository and assembles a single SQLite knowledge graph of everything worth knowing about the code: the symbols and how they call each other, the documents and decisions that govern them, and the fuzzy relationships between them. Every node and edge is provenance-tagged — it records how it was produced — so you always know whether a fact is a hard truth extracted from the source, a human decision, or a machine's suggestion.

ProvenanceSourceNature
derivedtree-sitter AST extractionDeterministic — symbols, calls, imports
authoredADRs, blueprints, // @rto: annotationsCurated intent, drift-checked in CI
inferredDocs, PDFs, images, embeddingsFuzzy suggestions with confidence scores

One store. One query surface. Three renderers — a docs website, an OKF bundle, and an optional MCP server — are all build outputs of the same graph, so what humans review is exactly what agents query.

Offline by default. The default build makes no network call of its own and needs no model to work. It is a small binary that needs no C++/cmake toolchain. It includes the whole of preparing to go offline — roteiro model pull and roteiro security prefetch — so that needs no rebuild. Nothing is ever fetched without an explicit [y/N] consent, and security run refuses without --allow-unsandboxed every time. Everything heavier — running local models, PDF/image ingestion, the model server, the sandboxed analyzer backend — stays opt-in behind a feature flag.
One optional feature sends your content elsewhere, and it is off. --features remote compiles the remote model tier (ADR-0019) — the one capability in Roteiro that can send graph-derived context to a hosted model. It is not in the default build and no release will put it there; it is in --all-features. Compiling it does not enable it: a run needs a grant from your own ~/.roteiro/config.toml and from the invocation (--allow-remote, or a prompt that shows you the exact bytes) — neither alone suffices, and a committed roteiro.toml may deny it for a whole repository but can never grant it, because a merged line must not authorise egress on a teammate's machine. Three commands can send, and each has to be told to: roteiro remote call is the one that exists to, roteiro spec draft --allow-remote drafts with the hosted model instead of a local one, and roteiro serve --allow-remote makes it the model the Ask panel uses — only remote call ever prompts, because a prompt on a command whose default is local turns a habituated “yes” into consent you never quite gave. serve --allow-remote grants for the life of the server process: every Ask it answers sends context to the hosted model for as long as it runs, including requests from anyone else who can reach the port — bounded by loopback-by-default binding, by your user config having granted independently, and by the ledger; it is never persisted and dies with the process. A refused --allow-remote stops the run and names the layer that refused, rather than answering from a local model. roteiro remote dry-run prints exactly what would be sent; roteiro remote log reads the append-only record of what did. Source code is not sent — function bodies are not in the graph — but symbol names, paths and captured prose are, they identify a codebase, and there is no redaction chokepoint on a prompt. And it never degrades quietly: an unreachable endpoint is an error naming that endpoint, not a local model's answer handed over in its place. So “nothing leaves the machine” is true of Roteiro as shipped and as configured, and is not a property of the software as a whole.

One graph, not a stack of tools

Teams end up gluing together a code-graph tool, a docs/spec system, a knowledge-graph builder, and a spec-driven-development kit — each with its own store, its own truth, and no shared notion of where a fact came from. Roteiro does all of it in one provenance-tagged store, and imports the tools you already use so you don't start from scratch.

Instead of…Roteiro gives you…
codegraph — an AST code-symbol graphThe derived layer: tree-sitter symbols, calls, imports and containment across 15+ languages — but git-native and provenance-tagged. roteiro import codegraph uses a codegraph snapshot as a validation oracle.
lat.md — markdown specs with lat checkThe authored layer: ADRs/blueprints linked into code symbols, with roteiro check failing CI on drift. roteiro import lat maps existing lat.md docs to authored nodes.
Graphify — a doc/media knowledge graphThe inferred layer: doc/PDF/image nodes with embedding-based similarity edges. roteiro import graphify brings its doc & media nodes in as inferred facts.
SpecKit — spec-driven developmentroteiro spec: graph-grounded ADR/blueprint authoring — assemble context from the graph, scaffold a house-style, check-clean skeleton, and optionally draft it with a local model.

Because they share one store, an authored ADR can link to a derived symbol and be verified against it, while an inferred edge stays clearly labelled as a suggestion. That cross-layer, provenance-aware query is exactly what a pile of separate tools cannot give you.

How it works

The graph is built by an incremental, content-addressed pipeline. Re-running sync only re-extracts blobs whose content changed, so it stays fast on a large repository.

1 · Extract

tree-sitter parses each source blob into symbols, calls and imports. Prose, PDFs and images can be ingested too. Deterministic and cached per git blob.

2 · Author

ADRs, blueprints and inline // @rto: annotations add human intent. roteiro check verifies these links against the code and fails CI on drift.

3 · Infer

Optional embeddings suggest inferred similarity edges and likely duplicates, each with a confidence score you can threshold.

4 · Render / Serve

Emit a docs site or an OKF bundle, or serve the graph to AI agents over MCP — every output is regenerated from the one store.

Because provenance is a first-class attribute, a query can say not just “what calls this function” but “which of these relationships are verified facts versus a model's guess.” That is the difference between a graph you can build automation on and one you can only browse.

Git-native & shareable

Roteiro's cache is content-addressed the same way git is. Extraction is a pure function of (path, git blob id, bytes), so every fact is keyed by the blob and tree hashes git already computes — a Merkle model layered directly on the object graph. Three consequences fall out of that:

Incremental by hash

A re-sync diffs the last-synced tree oid against HEAD and re-extracts only blobs whose content hash changed — never mtimes, so it's correct across checkouts, merges and rebases.

Shared across a team

The same commit yields the same graph for everyone. CI assembles the graph on merge and publishes a content-addressed artifact; a teammate's hooks fetch it instead of rebuilding — and fall back to a local rebuild offline.

Branch- & worktree-safe

Because keys are content hashes, branches and worktrees share extraction work and never serve each other stale facts. Uncommitted edits overlay the committed graph as a preview.

The upshot: the merged graph is the single source of truth, reproducible byte-for-byte from git alone, with no external database to keep in sync. It works with your repository, not beside it.

How does a team stay in sync? The graph store lives in git's per-worktree directory (.git/roteiro/ in a normal clone; a linked worktree resolves its own git dir) and is never committed — it's a local, reproducible cache, like target/. Nothing team-critical lives only there: the derived and authored layers are a pure function of the committed code, ADRs and @rto: annotations, which git already shares on push/pull, so everyone at the same commit re-derives them identically (the hooks do it on pull). The inferred layer needs models, so it's shared the other way — via the CI-published artifact, which a teammate fetches to get identical inferred edges without running inference. Either way the graph travels with git, not beside it — the SQLite file isn't committed (it would bloat history and conflict on merge).
Dogfood: browse Roteiro's own graph. Every merge to main publishes the build-outputs of Roteiro's own graph to the rolling graph-latest release: Both are build-outputs — regenerate either locally at any commit with roteiro render okf / roteiro export.

Install & build

cargo install roteiro --locked gets the lean default build — pure Rust, no C++ toolchain, and no network call of its own. The feature tiers add inference, serving, MCP, PDF/OCR/vision and the sandboxed analyzer backend. Install & build →

Use it in your project

roteiro init scaffolds the store, installs the git hooks that keep the graph fresh and gate drift, and writes an AGENTS.md snippet so your AI agents are graph-aware. Use it in your project →

The five ways to run it

Offline (the default), online with local models, the network HTTP server, MCP for AI agents, and the explorer in your browser. The five ways to run it →

Ask questions of your code

Structured queries that need no model and run fully offline, and natural-language questions answered by a local model calling the graph's own tools. Ask questions of your code →

Plan a change

roteiro spec turns a topic into a house-style, drift-checked ADR or blueprint, grounded in the graph. Four steps, and only one of them needs a model. Plan a change →

Recommended local models

roteiro model list recommends a pick per section tuned to this machine. All local models are GGUF, run through the shared llama.cpp engine. Recommended local models →

Configuration

A config file is optional. Roteiro merges a per-user ~/.roteiro/config.toml with a per-project roteiro.toml (project wins) — except [remote] enabled, which the committed file may switch off but never on. Configuration →

Cross-repo: a hub and its spokes

One hub application repo and many spoke deployment repos that each pin a version of it and override its configuration — joined at query time, never merged into one store. Cross-repo: a hub and its spokes →

Languages

Roteiro extracts a full symbol graph for Rust today — functions, structs, enums, traits, modules, impl blocks, use imports, and per-function call lists resolved into call edges across files. Broader first-class language extraction (Python, JavaScript/TypeScript, Go, Java, C/C++, C#, Ruby, PHP and more) is rolling out via tree-sitter's tags queries, which surface definitions and references through one shared code path.

Independent of source language, any file is ingested as a graph node with size and content metadata — and prose (Markdown/text), PDFs and images can contribute embeddable content — so a repository is never opaque to the graph, only richer where a language extractor exists.