Roteiro — Build Plan

[!IMPORTANT] Archived on 2026-09-02. Kept for links and history; no longer current.

This plan took Roteiro from the v0.0.1 scaffold to a dogfooded v1.0 and was delivered. Its successor is BUILD_PLAN_V2, which is itself archived — work is now tracked in issues.

Nothing here is maintained. Read it as a record of what was planned and decided at the time, not as a description of the project today: version numbers, stage lists, coverage figures and CI details were accurate when written and have moved on. The status: deprecated above is OKF §5.4's value, whose gloss is exactly this case — "kept for links and history; no longer current."

One rule from this document is still in force and has been rehomed so it does not retire with the plan: "stay on the current MSRV until a dependency forces a move; any MSRV bump is an ADR-worthy decision." That now lives in ADR-0001, which is where ADR-0001 v1.4 already cites it from.

Status: Archived · Owner: The Roteiro Project Team · Last-modified: 2026-08-09 · Archived: 2026-09-02 Governing decision: ADR-0001

This plan takes Roteiro from the initial v0.0.1 scaffold to a dogfooded v1.0. It is organised as sequenced stages, each ending in a shippable release cut by release-plz. Every stage names its deliverables, the concrete Rust surface it adds, new dependencies (with licence notes for the cargo deny gate), the CLI it wires up, and an explicit Definition of Done (DoD).

Current position (2026-08-09): Stages 1–13 and 15 are delivered — the graph core, extraction/sync/cache, query surface, renderers; the offline inference core


1. Principles (invariants that constrain every stage)

These come from ADR-0001 and must hold at every release, not just at v1.0:

  1. Provenance is first-class. Every edge is derived | authored | inferred, and inferred edges carry a confidence score. No code path may produce an unlabelled edge.
  2. One query surface. Humans (docs/vault) and agents (CLI --json, MCP) read the same store. Renderers are pure build-outputs of the graph.
  3. Offline by default. No network needed to build or query. Grammars are compiled in; the only optional network call is fetching a CI-published artifact, with local rebuild as the fallback.
  4. Git-native & deterministic. Extraction is a pure function of a git blob; the same blob always yields the same facts. The cache is content-addressed by git object id so branches/worktrees share work.
  5. Precise-where-known, fuzzy-where-suggested. Deterministic derivation and human authoring win; inference is clearly marked as suggestion.
  6. Dogfooded. Roteiro runs on its own repo in CI; roteiro check gates the build from the first stage it exists.
  7. Quality gates from day one. fmt, clippy (-D warnings, pedantic), audit, deny, and roteiro check. unsafe_code = "forbid". Coverage is measured in CI but not gated; the 85% per-file floor is an aspiration, not a check that runs (issue #319, §11).

2. Current state (v0.0.1 — done)

CrateWhat exists todayWhat's missing
rto-graphStore (SQLite, nodes/edges schema, open/open_in_memory/node_count), Provenance enumInsert/query API, node identity, migrations, extraction, cache
rto-specAdrStatus, AdrMeta, status FromStrFull ADR/blueprint parser, wiki-link/annotation edges, check, dedup
rto-renderTarget enum (docs/obsidian)Any actual rendering (site is a shell build.sh + md2html.awk stopgap)
roteiroCLI with init/sync/check/import/render/spec/serve stubs that bail!All behaviour

Infra already in place: workspace (edition 2024, MSRV 1.94, rusqlite =0.39 pin), CI (checks + msrv), release-plz, Cloudflare Pages site, all actions SHA-pinned. This plan assumes that baseline.


3. Crate responsibilities & dependency graph

        roteiro (CLI, arg parsing, wiring, hooks, init)
        /        |          \
  rto-spec   rto-render    (rto-graph re-exported)
        \        |          /
              rto-graph  (store, model, cache, extraction, query)

Rule: libraries return typed errors and data; the roteiro binary owns anyhow, stdout/stderr, and exit codes.


4. Data model & schema plan

The v0.0.1 schema is a starting point. Target model:

Extraction outputs a fact set per blob (nodes+edges scoped to one file), which is what gets content-addressed and cached. Assembly merges fact sets for all blobs in a tree, then resolves cross-file references (e.g. a calls edge's dst symbol) into node ids.


5. Staged roadmap

Each stage is independently shippable and leaves main green + dogfoodable.

Stage 1 — Graph core & query primitives → v0.1.0 ✅ delivered

Goal: a real, typed, transactional store the rest of the system builds on.

Stage 2 — Content-addressed cache & roteiro sync → v0.2.0 ✅ delivered (core)

Goal: git-native incremental graph updates.

Stage 3 — Derived extraction (tree-sitter) → v0.3.0 ✅ delivered (Rust + dirty overlay)

Goal: the derived provenance class for real code.

Stage 4 — Authored layer & roteiro check (rto-spec) → v0.4.0 ✅ delivered

Goal: house ADR/blueprint intent linked into code; drift gating.

Stage 5 — Query surface, --json, init & git hooks → v0.5.0 ✅ delivered

Goal: the agent-facing interface and zero-touch freshness.

Stage 6 — Renderers replace the shell stopgap (rto-render) → v0.6.0 ✅ delivered

Goal: docs site + Obsidian vault as true graph build-outputs.

Stage 7 — MCP server (feature-gated serve) → v0.7.0 ✅ delivered

Goal: agent access over MCP as a thin wrapper on the query API.

Stage 8 — Inference layer (inferred) → v0.8.0 ✅ delivered (offline default + local models)

Goal: fuzzy doc/PDF/image → suggestions with confidence.

Stage 9 — Importers → v0.9.0 ✅ Graphify delivered; lat.md + codegraph completed in Stage 11

Goal: migration path off the three incumbents.

Stage 10 — CI-canonical artifacts → v0.10.x 🚧 artifact format delivered

Goal: the merged graph is the source of truth; ship stable. (The v1.0 hardening that once lived here is now tracked explicitly as Stage 14.)


5b. Remaining-work stages (the honest backlog to v1.0)

Several stages above shipped their core and deferred the rest. Rather than let that deferred work hide inside per-stage footnotes, it is promoted here to first-class, sequenced stages. Order reflects the agreed priority: complete Stage 9 → complete Stage 8 → the spec/blueprint authoring pillar → Stage 10 overflow (v1.0 hardening).

Stage 11 — Importers: lat.md + codegraph → v0.11.x ✅ delivered

Goal: finish the migration path off the remaining two incumbents (completes the Stage 9 deferral).

Stage 12 — Inference ingestion: content, PDF, image + semantic dedup → v0.12.x ✅ delivered (text + PDF + image OCR/vision ingestion, semantic dedup, dependency-aware context cache; completes Stage 8)

Goal: make inferred edges meaningful by embedding real content, not just node names, and extend ingestion to docs/PDFs/images.

Stage 13 — Spec/Blueprint authoring pillar → v0.13.x ✅ delivered (ADR-0004; Tier 0 spec context/spec scaffold + Tier 1 local-model spec draft)

Goal: the intent interview + house-style ADR/blueprint + graph-grounded, correct build/deploy plan generation — the front door ADR-0001 always envisioned (roteiro spec), sharpened by GitHub spec-kit's phases (constitution → specify → clarify → plan → tasks). Grounded in Roteiro's graph so generated plans reference real symbols/ADRs/deps and are check-gated.

Stage 16 — Commit-time correctness gate → v0.x ✅ delivered (execution order: after Stage 13, immediately before Stage 14)

Goal: guarantee the knowledge base is not just fresh (what sync gives) but correct (no authored-vs-code drift) at the point of a commit — and checkable mid-work during a large change — instead of relying on a manual check or only Roteiro's own CI. Today the managed hooks run sync --committed only (freshness on checkout/merge); nothing runs check, and check validates the committed HEAD tree, so as a pre-commit hook it would inspect the parent commit, not the staged change. This stage closes that gap. Resolves the follow-ups noted under Stages 2/4 ("making check working-tree-aware pairs with sync_worktree"; "a working-tree query mode").

Stage 14 — v1.0 hardening → v1.0.0 ✅ delivered and released — v1.0.0 on crates.io; v1.1.0 followed

Goal: the merged graph is the canonical source; ship stable (completes the Stage 10 deferral). Every hardening item below is delivered in code, and the v1.0.0 release has since been cut and published to crates.io (crates are now on v1.1.0).

Stage 15 — Intent-debt tracking (TODOs, stubs, deferred work) → delivered (independent; low-risk)

Goal: deterministically detect, log, and track intent debt — the markers in code and docs that signal missed intent or intent left for the future — so end users and AI can find what's incomplete instead of it hiding in comments and footnotes.

Stage 17 — Agent instructions & context-aware review (tool-agnostic) → ✅ delivered (CLI-first roteiro review + tool-agnostic AGENTS.md/checklist; MCP-for-review feasibility resolved — feasible as an optional enhancement, see §5d)

Goal: make every calling agent — Copilot code review, Claude Code, Cursor, cloud agents, future contributors — Roteiro-aware, via tool-agnostic files rather than one vendor's format, so the same standards drive review and authoring everywhere. Sequenced after Stage 14 so the standards it encodes are final (v1.0-frozen), not a moving target.


5c. New stages (decided post-Stage-12, via ADRs)

Decisions taken after the original roadmap, each with its own ADR. Sequenced around the Stage 14 freeze: config is foundational (before 14), serving and acceleration are features (config first, since serving is configured through it).

Stage 18 — Configuration file (ADR-0007) → ✅ core delivered (more sections as their consumers land)

Goal: a persistent, optional roteiro.toml so per-project preferences are set once, not retyped as flags — reproducible and shareable when committed.

Stage 19 — Local model serving (ADR-0006) → ✅ delivered (models + chat + streaming + graph tools + embeddings + vision)

Goal: reuse the models a user already pulled by exposing them over an opt-in, loopback OpenAI-compatible endpoint — offline, no second download — and make the served model code-aware by handing it Roteiro's graph tools.

Stage 20 — Inference-core direction + coding/reasoning models → ✅ delivered (generation migrated; coding/reasoning models; ADR-0003 amended)

Goal: decide the accelerated inference path across all uses (not just serving), and broaden the opt-in model catalogue.


5d. Ideas & deferred follow-ups (post-v1.0 backlog)

Small, non-blocking refinements surfaced while delivering the stages above. Tracked here so they don't hide in per-stage footnotes. None gate v1.0.

Extraction & graph

Commit gate & review

Authoring & importers

Serving & models

Config & hooks

Detection quality

Agent reviews


6. Cross-cutting concerns

Testing strategy

Coverage: measured in CI, not gated (issue #319). For a long time this line read "cargo-llvm-cov in CI, 85% per-file floor" while .github/workflows/ci.yml contained no coverage tooling of any kind — a false statement about the pipeline, which made every stage DoD below that cites "85% coverage" unverifiable. Three separate agents reported it independently, which is the tell: the docs were actively misleading people trying to comply.

What is true now: a non-blocking coverage job runs cargo llvm-cov --workspace --all-features and publishes the per-file table and the workspace total to the run summary. No threshold is enforced. The 85% per-file floor remains the aspiration recorded in ADR-0001, not a gate — and a per-file floor is stricter than most projects run, so switching one on blind would fail thin wiring files nobody wants to pad with ceremonial tests. Measure first; turning the floor on is a separate, deliberate change with the real numbers in hand.

The enforcing gates are, exactly and only: cargo fmt --check, cargo clippy --workspace --all-targets --all-features -D warnings, cargo test --workspace --all-features, roteiro check (dogfood), cargo audit, and cargo deny --all-features check.

Measured baseline (cargo llvm-cov --workspace --all-features, tests excluded from the denominator, at the commit that added the job):

MetricValue
Workspace total, lines87.51%
Workspace total, regions86.96%
Workspace total, functions86.55%
Files measured64
Files below 85% lines7

The seven: roteiro/src/main.rs (60.08%), rto-llama/src/speculative.rs (34.77%), rto-graph/src/media/producers.rs (49.26%), rto-exec/src/subprocess.rs (56.55%), rto-llama/src/engine.rs (68.18%), rto-llama/src/llama.rs (74.20%), rto-exec/src/assets.rs (81.53%).

These numbers are why the floor is not switched on in the same change that started measuring. The workspace already clears 85% — a workspace-level gate would pass today. A per-file gate would fail seven files, and the reason each one is low is a fact about what the code does, not a gap someone forgot: main.rs is CLI wiring driven through integration tests that llvm-cov attributes to the binary unevenly, and speculative.rs, llama.rs, engine.rs, producers.rs and subprocess.rs need a loaded model, a GPU, or a sandboxed subprocess to exercise their real paths. Padding them with ceremonial tests would move the number without moving the risk. Deciding the threshold — workspace-level now, per-file later, or per-file with a documented exemption list — is the follow-up this measurement exists to inform.

Dependency & licence policy: every new crate (esp. tree-sitter grammars and inference/PDF crates) must pass cargo deny (licence MIT/Apache-compatible, no duplicate/banned crates) and cargo audit. Prefer pure-Rust, offline-capable crates (gix over libgit2). Keep MCP/inference deps behind features so the default build stays lean and the MSRV surface stays small.

MSRV discipline: stay on 1.96 until a dependency forces a move; the rusqlite =0.39 pin is one instance of pinning for it (documented in Cargo.toml, which also records the second, non-MSRV constraint now holding that pin in place — boxlite's rusqlite ^0.39 against a links = "sqlite3" crate). Any MSRV bump is an ADR-worthy decision.

Error handling: libraries expose thiserror enums; the binary uses anyhow + process exit codes. check/sync return structured results so the CLI can render both human and --json forms.

Performance targets (validate at Stage 14): cold full extract of a mid-size repo in seconds; incremental sync proportional to the diff; cache-hit sync effectively instant; --json queries sub-100ms on the dogfood graph.


7. Milestones → releases

NominalStageHeadline capabilityActual tag / status
v0.1.01Typed transactional graph store + migrations✅ v0.0.2
v0.2.02Content-addressed cache; roteiro sync (incremental)✅ v0.0.3
v0.3.03Derived tree-sitter extraction (Rust) + dirty overlay✅ v0.0.4
v0.4.04Authored ADR/blueprint layer; roteiro check gates CI✅ v0.0.5
v0.5.05Query API + --json; init + git hooks✅ v0.0.6
v0.6.06Real docs-site + Obsidian renderers (retire shell stopgap)✅ v0.0.7
v0.7.07MCP serve (rmcp; stdio + HTTP, ADR-0002)✅ v0.0.8
—7+roteiro path + MCP path tool (follow-up)✅ v0.0.9
v0.8.08Inference layer (inferred + confidence)✅ offline core + candle local-models (roteiro infer/model); ingestion → Stage 12
v0.9.09Importers (lat.md / Graphify / codegraph) + reports✅ Graphify shipped; lat.md + codegraph completed in Stage 11
v0.10.x10CI-canonical artifacts🚧 artifact export/load shipped (v0.0.10); CI publish/fetch etc. → Stage 14
v0.11.x11Importers: lat.md + codegraph (completes 9)✅ durable+validated imports, lat.md importer, codegraph oracle (compare_codegraph)
v0.12.x12Inference ingestion: content/PDF/image + semantic dedup (completes 8)✅ prose + PDF + image OCR/vision (ADR-0005) ingestion, semantic dedup (roteiro duplicates), dependency-aware context cache (roteiro context)
v0.13.x13Spec/Blueprint authoring pillar (ADR-0004; tiered, graph-grounded)✅ ADR-0004; Tier 0 (spec context/scaffold) + Tier 1 (spec draft) — now Qwen3 via a GGUF-arch-dispatching candle loader
v0.x16Commit-time correctness gate: worktree-aware check + pre-commit/post-commit hooks✅ delivered (runs just before Stage 14; touches sync+check+init)
v1.0.014v1.0 hardening (completes 10): CI artifacts, TS/JS+Python, deploy, --json freeze✅ shipped v1.0.0 (crates.io; v1.1.0 followed)
v0.x15Intent-debt tracking: TODO/stub/deferred markers as derived facts + roteiro debt✅ marker nodes + debt query/CLI/MCP; check summary line
post-1.017Tool-agnostic agent instructions (AGENTS.md) + context-aware review skill; MCP-for-review (investigate)⛔ after Stage 14 (standards must be v1.0-final)
v0.x18Configuration file (ADR-0007): layered roteiro.toml, TOML-only✅ core — roteiro config, [models]/[infer]/[duplicates]/[ingest], CLI>project>user>default
v0.x19Local model serving (ADR-0006): llama.cpp-backed, code-aware OpenAI /v1✅ opt-in serve — /v1/models+chat+streaming+embeddings+vision (mtmd), auto-registered graph tools
v0.x20Inference-core direction (unify on llama.cpp) + coding/reasoning models✅ candle removed — one rto-llama core for gen/embed/vision + serving (ADR-0003 v1.2); coding/reasoning role entries pull+run
——Shipped alongside Stage 12: curated low/mid/high model matrix (ADR-0003) + streaming, checksum-verified model downloads✅ roteiro model list by section→tier; download_verified (constant memory)

8. Risk register

RiskImpactMitigation
Tree-sitter grammar licence incompatibilityBlocks a languageVerify licence before vendoring; cargo deny gate; drop/replace grammar if needed
Offline embedding model size/licence (Stage 8)Binary bloat or licence conflictResolved by ADR-0003: tiny static int8 default compiled in; GGUF local models opt-in behind inference-local-models; consent-gated fetch
Unmaintained YAML crate flagged by auditcheck gate can't shipResolved: hand-parsed frontmatter in rto-spec (no serde_yaml)
Non-deterministic extraction → cache churnCache/CI-diff noiseSort all emitted facts; snapshot + idempotency tests from Stage 3
MSRV drift from a new depCI msrv job breaksPin (as with rusqlite); gate new deps on 1.96; ADR any bump
Scope creep in check/dedupSlips v0.4Handled: structural check shipped (v0.0.5); semantic dedup → Stage 12
Image OCR/vision has no good pure-Rust path (Stage 12)Blocks image ingestion or forces a C++ FFI / heavy model depDe-risk (MSRV + deny) and ADR before committing; keep behind its own feature; ship text + PDF ingestion first (both low-risk) so image is isolated
Generative local model for authoring is bigger than the embedding default (Stage 13)"Light mode" still needs a pulled modelTier-0 (offline, no model) guarantees planning always works; light-tier reuses the Stage 8 candle/GGUF registry; foundation/agent tier for quality

9. Open questions (decide before the relevant stage)

  1. Node identity scheme — decided: opaque natural keys with the convention sym:<lang>:<path>#<name> / adr:<id>#<section> / file:<path>. The store treats keys as opaque strings, so the scheme can evolve. (Stage 1)
  2. On-disk FactSet codec — decided: JSON (debuggable; revisit a compact binary only if size/speed demands it). (Stage 2)
  3. First language breadth — decided: Rust-only for the first extractor; TS/JS and Python are tracked Stage 3 follow-ups. (Stage 3)
  4. YAML/Markdown parsing — decided: hand-parser for house-style frontmatter/sections (no serde_yaml); pulldown-cmark is used only for rendering, not parsing intent. (Stage 4)
  5. Templating — decided: hand-rolled page chrome for the docs site (no askama/minijinja — no templating dependency). (Stage 6)
  6. MCP SDK — decided: rmcp (the official SDK), for stdio + networked HTTP serving — see ADR-0002. (Stage 7)
  7. Embedding model — decided (ADR-0003): a tiny static int8 (model2vec-style) embedding compiled in as the offline default (single-digit MB budget), plus GGUF pluggable local models via an in-binary registry with platform-aware (Metal/Apple vs standard) variant selection and consent-gated fetch; the candle backend sits behind a second inference-local-models feature. (Stage 8)
  8. Image OCR/vision backend — pure-Rust OCR (weak) vs. tesseract C++ FFI (breaks the pure-Rust stance) vs. a candle vision model (heavy, weights). Decide + ADR before building. (Stage 12)
  9. Authoring model driver — confirmed Roteiro-grounded with tiers (offline scaffolding → small local GGUF instruct model → foundation/agent); the small-model choice + size budget is the open sub-decision. (Stage 13)

This is a living document. Each stage should land with any decisions above resolved in its PR description, and — once roteiro check exists — this plan and ADR-0001 are themselves checked by the tool.