The pilot book for your codebase.
The Portuguese roteiros were the guarded pilot books of the Age of Discovery — accumulated route knowledge that made navigation repeatable. Roteiro does the same for software: structure, intent, and context in one queryable, provenance-tagged knowledge graph, for humans and AI agents alike.
Roteiro reads a git repository and assembles a single SQLite knowledge graph of everything worth knowing about the code: the symbols and how they call each other, the documents and decisions that govern them, and the fuzzy relationships between them. Every node and edge is provenance-tagged — it records how it was produced — so you always know whether a fact is a hard truth extracted from the source, a human decision, or a machine's suggestion.
| Provenance | Source | Nature |
|---|---|---|
| derived | tree-sitter AST extraction | Deterministic — symbols, calls, imports |
ADRs, blueprints, // @rto: annotations | Curated intent, drift-checked in CI | |
| inferred | Docs, PDFs, images, embeddings | Fuzzy suggestions with confidence scores |
One store. One query surface. Three renderers — a docs website, an OKF bundle, and an optional MCP server — are all build outputs of the same graph, so what humans review is exactly what agents query.
roteiro model pull and roteiro security prefetch — so that needs
no rebuild. Nothing is ever fetched without an explicit [y/N] consent, and
security run refuses without --allow-unsandboxed every time.
Everything heavier — running local models, PDF/image ingestion, the model server, the
sandboxed analyzer backend — stays opt-in behind a feature flag.--features remote compiles the
remote model tier
(ADR-0019)
— the one capability in Roteiro that can send graph-derived context to a hosted model. It
is not in the default build and no release will put it there; it is in
--all-features. Compiling it does not enable it: a run needs a grant from your
own ~/.roteiro/config.toml and from the invocation
(--allow-remote, or a prompt that shows you the exact bytes) — neither alone
suffices, and a committed roteiro.toml may deny it for a whole
repository but can never grant it, because a merged line must not authorise egress on a
teammate's machine. Three commands can send, and each has to be told to: roteiro remote call is the
one that exists to, roteiro spec draft --allow-remote drafts with the hosted
model instead of a local one, and roteiro serve --allow-remote makes it the
model the Ask panel uses — only remote call ever prompts, because a
prompt on a command whose default is local turns a habituated “yes” into consent you never
quite gave. serve --allow-remote grants for the life of the server
process: every Ask it answers sends context to the hosted model for as long as it
runs, including requests from anyone else who can reach the port — bounded by
loopback-by-default binding, by your user config having granted independently, and by the
ledger; it is never persisted and dies with the process. A refused
--allow-remote stops the run and names the layer that refused, rather than
answering from a local model. roteiro remote dry-run prints exactly what would
be sent; roteiro remote log reads the append-only record of what did. Source
code is not sent — function bodies are not in the graph — but symbol names, paths and captured prose
are, they identify a codebase, and there is no redaction chokepoint on a prompt. And it
never degrades quietly: an unreachable endpoint is an error naming that endpoint, not a
local model's answer handed over in its place. So “nothing leaves the machine” is
true of Roteiro as shipped and as configured, and is not a property of the software as a
whole.Teams end up gluing together a code-graph tool, a docs/spec system, a knowledge-graph builder, and a spec-driven-development kit — each with its own store, its own truth, and no shared notion of where a fact came from. Roteiro does all of it in one provenance-tagged store, and imports the tools you already use so you don't start from scratch.
| Instead of… | Roteiro gives you… |
|---|---|
| codegraph — an AST code-symbol graph | The derived layer: tree-sitter symbols, calls, imports and containment across 15+ languages — but git-native and provenance-tagged. roteiro import codegraph uses a codegraph snapshot as a validation oracle. |
lat.md — markdown specs with lat check | The layer: ADRs/blueprints linked into code symbols, with roteiro check failing CI on drift. roteiro import lat maps existing lat.md docs to authored nodes. |
| Graphify — a doc/media knowledge graph | The inferred layer: doc/PDF/image nodes with embedding-based similarity edges. roteiro import graphify brings its doc & media nodes in as inferred facts. |
| SpecKit — spec-driven development | roteiro spec: graph-grounded ADR/blueprint authoring — assemble context from the graph, scaffold a house-style, check-clean skeleton, and optionally draft it with a local model. |
Because they share one store, an authored ADR can link to a
derived symbol and be verified against it, while an inferred
edge stays clearly labelled as a suggestion. That cross-layer, provenance-aware
query is exactly what a pile of separate tools cannot give you.
The graph is built by an incremental, content-addressed pipeline. Re-running
sync only re-extracts blobs whose content changed, so it stays fast on
a large repository.
tree-sitter parses each source blob into symbols, calls and imports. Prose, PDFs and images can be ingested too. Deterministic and cached per git blob.
ADRs, blueprints and inline // @rto: annotations add human intent. roteiro check verifies these links against the code and fails CI on drift.
Optional embeddings suggest inferred similarity edges and likely duplicates, each with a confidence score you can threshold.
Emit a docs site or an OKF bundle, or serve the graph to AI agents over MCP — every output is regenerated from the one store.
Because provenance is a first-class attribute, a query can say not just “what calls this function” but “which of these relationships are verified facts versus a model's guess.” That is the difference between a graph you can build automation on and one you can only browse.
Roteiro's cache is content-addressed the same way git is. Extraction
is a pure function of (path, git blob id, bytes), so every fact is
keyed by the blob and tree hashes git already computes — a Merkle model layered
directly on the object graph. Three consequences fall out of that:
A re-sync diffs the last-synced tree oid against HEAD and re-extracts only blobs whose content hash changed — never mtimes, so it's correct across checkouts, merges and rebases.
The same commit yields the same graph for everyone. CI assembles the graph on merge and publishes a content-addressed artifact; a teammate's hooks fetch it instead of rebuilding — and fall back to a local rebuild offline.
Because keys are content hashes, branches and worktrees share extraction work and never serve each other stale facts. Uncommitted edits overlay the committed graph as a preview.
The upshot: the merged graph is the single source of truth, reproducible byte-for-byte from git alone, with no external database to keep in sync. It works with your repository, not beside it.
.git/roteiro/ in a normal clone; a
linked worktree resolves its own git dir) and is never committed — it's a
local, reproducible cache, like target/. Nothing team-critical
lives only there: the derived and
layers are a pure function of the
committed code, ADRs and @rto: annotations, which git
already shares on push/pull, so everyone at the same commit re-derives them
identically (the hooks do it on pull). The inferred
layer needs models, so it's shared the other way — via the CI-published artifact,
which a teammate fetches to get identical inferred edges without running inference.
Either way the graph travels with git, not beside it — the SQLite file
isn't committed (it would bloat history and conflict on merge).main publishes the build-outputs of Roteiro's own graph to the rolling
graph-latest
release:
index.md).curl -fsSL <url> | roteiro load -.roteiro render okf / roteiro export.cargo install roteiro --locked gets the lean default build — pure Rust,
no C++ toolchain, and no network call of its own. The feature tiers add inference,
serving, MCP, PDF/OCR/vision and the sandboxed analyzer backend. Install & build →
roteiro init scaffolds the store, installs the git hooks that keep the
graph fresh and gate drift, and writes an AGENTS.md snippet so your AI
agents are graph-aware. Use it in your project →
Offline (the default), online with local models, the network HTTP server, MCP for AI agents, and the explorer in your browser. The five ways to run it →
Structured queries that need no model and run fully offline, and natural-language questions answered by a local model calling the graph's own tools. Ask questions of your code →
roteiro spec turns a topic into a house-style, drift-checked ADR or
blueprint, grounded in the graph. Four steps, and only one of them needs a model. Plan a change →
roteiro model list recommends a pick per section tuned to this machine.
All local models are GGUF, run through the shared llama.cpp engine. Recommended local models →
A config file is optional. Roteiro merges a per-user ~/.roteiro/config.toml
with a per-project roteiro.toml (project wins) — except
[remote] enabled, which the committed file may switch off but
never on. Configuration →
One hub application repo and many spoke deployment repos that each pin a version of it and override its configuration — joined at query time, never merged into one store. Cross-repo: a hub and its spokes →
Roteiro extracts a full symbol graph for Rust today — functions,
structs, enums, traits, modules, impl blocks, use imports, and
per-function call lists resolved into call edges across files. Broader
first-class language extraction (Python, JavaScript/TypeScript, Go, Java, C/C++,
C#, Ruby, PHP and more) is rolling out via tree-sitter's tags queries,
which surface definitions and references through one shared code path.
Independent of source language, any file is ingested as a graph node with size and content metadata — and prose (Markdown/text), PDFs and images can contribute embeddable content — so a repository is never opaque to the graph, only richer where a language extractor exists.
GitHub crates.io Project Blueprint Use it from an agent platform Architecture Decision Records