Roteiro — Build Plan V2

[!IMPORTANT] Archived on 2026-09-02. Kept for links and history; no longer current.

V2's stages were delivered and the work moved to tracked issues, which is what supersedes this document. Its predecessor is BUILD_PLAN.

Nothing here is maintained. In particular the Baseline (start of V2) table is a point-in-time snapshot — it still records "Released v1.15.0, all seven crates" against a project now past 5.x with nine — and is deliberately left as written. The status: deprecated above is OKF §5.4's value, whose gloss is exactly this case: "kept for links and history; no longer current."

Status: Archived — delivered, then superseded by tracked issues · Owner: The Roteiro Project Team · Last-modified: 2026-08-19 Governing decisions: ADR-0001, ADR-0012, ADR-0013, ADR-0014, ADR-0018

This plan is finished, and it is a record rather than a roadmap

Every stage below is delivered. Work that was still outstanding when this was closed became tracked issues instead, and new work should be an issue from the start rather than a stage appended here.

Why it stops rather than continues. In a single day this document drifted from itself five times — stage headings marked delivered while the release table still showed the stage pending, and a summary section describing stages as blocked after they had shipped. Each time it was found by a person reading it, never by a gate. Its own standing rule says why that matters:

A plan that lags the tree is worse than no plan, because it is trusted.

The cause is structural rather than careless: the same facts live in three places here — the stage headings, the milestones table, and the dependency summary — and nothing holds them level. That is the shape this project has now removed six times in code ([debt] ignore across three surfaces, limit == 0 across five endpoints and then a sixth, [models] reporting versus resolving, the skill artifacts versus their template, the website versus the gate), and every one of those fixes worked by removing the possibility of divergence rather than the instances of it. An issue cannot drift the same way: it is open or it is closed, and it lives beside the work instead of describing it from a distance.

What is deliberately kept. The measurements. Cost against estimate for the lenses (the corrected 195–500 LOC figure was itself out by 3–8×), the negative result on the review arm, the scale benchmark to 72,509 nodes, and the constraint each stage settled for the next. Those cost real work to establish and would have to be re-established if this were deleted.

This plan succeeds BUILD_PLAN.md, which took Roteiro from the v0.0.1 scaffold through Stage 20 to the released v1.9.0. V2 covers the next arc: Roteiro learns things that are not in the source tree — and stores them without compromising the promise that the graph is a pure function of the source tree.

Stage numbering continues from the v1 plan (which ended at Stage 20). As there, stage numbers are labels, not execution order, and the v1.x/v2.0 headings are nominal targets — release-plz cuts the real tags from conventional commits, so a stage nominally marked v1.16.0 may actually ship in a v1.10.x. Each delivered stage records the version it really shipped in.

Keep this document current. A stage is not finished when its code merges — it is finished when its entry here says what shipped, in which release, and what that settles for later stages. The stage entry is where the next person looks first; a plan that lags the tree is worse than no plan, because it is trusted. Update it in the same PR as the work wherever possible.


1. Thesis of V2

V1 built one thing well: a provenance-tagged graph, deterministically derived from git blobs, that humans and agents query through one surface. V2 adds three kinds of knowledge that do not fit that model and would corrupt it if forced in:

  1. Analyzer findings — asserted by an external tool at a point in time, against rules and advisory databases that change independently of the source.
  2. Agent memory — accumulated across sessions, episodic, unreproducible, and often the record of something that failed. 2b. Generated media content — ASR transcripts and VLM descriptions, invented fluently when the source contains nothing to read (ADR-0015, Stage 28).
  3. Deeper analysis lenses — genuinely derived facts, which stay in the graph, but whose true cost was previously understated by an order of magnitude.

The organising rule for (1) and (2) is one sentence, and it is what makes V2 coherent rather than a list of features:

Knowledge that is not a derived/authored/inferred graph fact gets its own artifact store, and never borrows the graph's trust.

imports already works this way — it exists precisely because sync's rebuild would destroy it, and is re-applied afterwards. V2 generalises that precedent instead of inventing something new.


2. Principles

All seven principles of BUILD_PLAN.md §1 remain binding. V2 adds three invariants that constrain every stage below:

  1. The graph stays a pure function of source. Nothing in V2 writes to nodes/edges unless it is deterministically derived from (path, blob id, bytes). export_factset must remain byte-identical for a given tree.
  2. Artifact stores never borrow graph trust. No V2 record acquires the authored relevance boost, and none is exported in the GraphArtifact.
  3. Offline-capable, not "offline". Optional capabilities may require pre-provisioned assets; they must be digest-pinned, explicitly prefetched, and must fail with a named, actionable error rather than fetching implicitly or silently degrading.

3. Baseline (start of V2)

Verified against main at the time of writing:

FactValueConsequence for V2
Releasedv1.15.0 on crates.io, all seven cratesV2 work is post-1.0 — semver is now real.
MSRVrust-version = "1.94"New deps must respect it.
Lintsunsafe_code = "forbid", clippy pedantic -D warningsNative/FFI deps must be isolated behind a feature.
Coveragemeasured in CI, not gated — cargo llvm-cov runs non-blocking; the 85% per-file floor is an aspiration (ADR-0001), never an enforced check (issue #319)Every stage below still carries test cost, but a DoD may not cite "85% coverage" as if something verified it.
CIUbuntu-only; --all-features and the default set (the default-features job, added by #364 after the default set was found not to compile — issue #360)/dev/kvm may be absent; Apple Silicon untested. Turning features on cannot find defects caused by code being cfg'd out.
Schemamigrations 1–13 applied (1–7 at V2's start)V2 appends only; see §5.
EXTRACT_VERSION11 (crates/rto-graph/src/extract.rs) — the Stage 28 bump landed in #316Bumping it forces full re-extraction for every user. No test pins the value.
Provenance`DerivedAuthored
Eviction idiomin-memory byte-budget LRU (rto-llama ModelCache); nothing persisted is boundedStage 25 ports the existing policy to disk rather than inventing one — done, and tested against lru_evict_count's own numbers.

4. Crate & feature map

CrateChangeNotes
rto-execnewAnalyzerRunner trait + three backends (ADR-0014). Feature execution, subfeatures exec-boxlite, exec-subprocess.
rto-graphextendedArtifact-store tables + accessors; new query fns for lenses. Graph model untouched.
roteiro (CLI)extendedsecurity {prefetch,status,run,ingest}, memory {add,list,recall,forget}, new lens subcommands.
rto-serveextendedNew lenses surfaced to served-chat tools; memory recall exposed only behind explicit opt-in.
rto-renderextendedFindings and lens renderers.

Default install gains no new dependency. Everything in Stages 22/24 is feature-gated and off by default.


5. Schema plan — migration discipline

migrations.rs mandates append-only SQL: never edit a shipped migration. V2 adds three tables across three migrations, deliberately not merged:

MigrationTableLifetimeEvictable
8 ✅analysis runs + findings (ADR-0012)replaceable layer per (analyzer, worktree)replaced wholesale, not aged out
11 ✅agent_memory (ADR-0013 episodic)durable, survives rebuildnever
13 ✅agent_cache + agent_cache_clock (ADR-0013 transient)boundedyes, by capacity

Numbers are assigned in landing order, not reserved in advance. Stage 21 landed first and took 8; Stage 28 took 9 and 10; episodic memory took 11; 12 is sync_worktree, from the guardrails branch; and the cache tier took 13. Splitting memory across two migrations is intentional: different lifetimes and guarantees, so the eviction tier can later be altered without touching durable memory.

Check the other worktrees, not just the refs. Stage 25 was dispatched with "12 is free" after git grep over every local and remote branch found nothing claiming it — which was true of the refs and false of the tree, because the guardrails work was sitting in an unpushed worktree. That is the third numbering collision of the day and all three had the same cause, so the check that actually works is:

git worktree list | awk '{print $1}' | xargs -I{} grep -ho 'version: [0-9]*' {}/crates/rto-graph/src/migrations.rs | sort -u

Two constants both declaring the same version merge cleanly in git and break at runtime on whichever store applies them second, so this is not a conflict the tooling will catch for you. migration_versions_are_unique_and_ascending now fails the build if a merge ever produces one.

Stages 22 and 24 need no migration — RunnerKind shipped in migration 8 already naming all three backends, with the schema CHECK accepting them. Stage 22 confirmed this: two analyzers landed with no schema change at all, because FindingKey takes each analyzer's own ordered identity components. Stage 22b confirmed it again: osv-scanner landed with a five-component identity and no schema change, and a test asserts exactly that.

EXTRACT_VERSION does not change in Stages 21–25. None of that work is extraction output (Stage 21 shipped without touching it, as required). It does change in Stage 26, once — see the note there.


6. Staged roadmap

Dependency shape — four tracks, only one hard chain:

Track A (findings):  21 ✅ ──► 22 ✅ ──► 22b ✅ ──► 24
Track B (memory):    23 ✅ ──────────────► 25
Track C (lenses):    26         (independent of A and B throughout)
Track D (media):     28 ✅ ──► 29
                                          └──► 27 (v2.0 hardening)

Stages 21, 23 and 26 were the parallel-startable set; 21, 22, 22b, 23 and 28 have now landed, leaving 24, 25, 26 and 29 open. Nothing in Track C touches the artifact stores; nothing in Track B blocks Track A. 22b was sequenced after 22 and did not block 24: it added one adapter behind a seam 24 does not touch, and needed no migration.


Stage 21 — Analyzer contract & ingest (ADR-0012, ADR-0014) → v1.10.0 · effort S ✅ delivered

Goal: land the whole value of the findings design with no analyzer and no sandbox — the seam, the schema, and a working ingest path. This is the stage that makes CI ingestion and local execution the same code path.

Delivered in v1.10.0 (PR #293). What shipped, and what it changes for later stages:

Stage 22 — First analyzers: semgrep + cargo-audit → v1.11.0 · effort M + M ✅ delivered

Goal: two real analyzers behind the Stage 21 contract, via the subprocess runner, honestly labelled.

Delivered in v1.11.0 (PR #322). What shipped, and what it changes for later stages:

Coverage is a matrix, not an analyzer list (ADR-0018). The requirement is findings for Rust, Python, SQL, Java and Node. The two analyzers named in Stage 22 deliver that on the SAST axis only; Stage 22b closed the dependency axis:

LanguageSASTDependency vulnerabilities
Rustsemgrep (GA)cargo-audit (RustSec) + osv-scanner — both kept, cross-referenced (22b)
Pythonsemgrep (GA)osv-scanner (OSV PyPI) ✅ 22b
Javasemgrep (GA)osv-scanner (OSV Maven) ✅ 22b
Node (JS/TS)semgrep (GA)osv-scanner (OSV npm) ✅ 22b
SQLsemgrep generic — token matching, no parsern/a — no dependency ecosystem

The adapter is the seam, so ingest and execution agree by construction. One conversion (rto_exec::normalize_native) turns an analyzer's native output into the normalized report Stage 21 already validates, and both paths call it: a subprocess run hands it the analyzer's stdout, roteiro security ingest --analyzer <name> hands it a report file CI produced. Equality of Finding values is a property of the code, not something a test establishes afterwards. A new analyzer is a file in adapter/ and a row in ADAPTERS — no migration, because FindingKey takes each analyzer's own ordered identity components.

Three tool behaviours found by running the tools, not by reading their docs. Each is a trap the next analyzer author would otherwise re-discover:

  1. Semgrep path-prefixes rule ids. With --config <file> it renames every rule to <config.path.components>.<id>. The rule id is the first component of a FindingKey, so this puts the local asset-cache directory into every stored key — user-identifying data in a persisted record, and keys that differ between two machines running the identical scan. --no-rewrite-rule-ids is mandatory, and a test asserts no key contains a local path.
  2. Semgrep redacts the matched source. In the open-source CLI extra.lines and extra.fingerprint are the literal string "requires login" unless the caller is authenticated to Semgrep's hosted platform. ADR-0012's identity recipe ends in a snippet hash, so hashing that field makes the component a constant and changes every stored key the day someone logs in. Snippets are read from the worktree instead (rto_exec::snippet), which is a function of the source rather than of the analyzer's auth state.
  3. cargo audit reports no advisory-database identity when you pin one. Given --db <path> it returns last-commit: null and last-updated: null — at a shallow clone and at its own managed checkout alike (0.22.2). Only the unpinned, self-resolving configuration populates them, so the reproducible offline configuration was the one losing its staleness evidence. Provisioning records the publication date from the database's HEAD commit time and the adapter falls back to it; the tool's own account still wins where it has one.

Provisioning: writing and running are separate. prefetch is the only thing that writes to the asset cache; a run never provisions, so "did this machine have the pinned rules?" always has an answer. As shipped, prefetch verifies and pins but fetches nothing — the rule set is vendored into the binary and the RustSec advisory database is a git checkout with no digest-stable URL, so it is refused with the exact clone command rather than obtained by shelling out to git (which would be the host-tool fallback ADR-0014 forbids). ADR-0014 v1.1 records that clarification. An asset whose bytes change after provisioning is refused, not warned about: a run would otherwise stamp a digest that does not describe what it read.

Rules are ours, vendored and pinned. semgrep --config p/default resolves against a network service, which makes an "offline" analyzer network-dependent and its results irreproducible. Every shipped rule was written for this repository under its own licence; no Semgrep Registry rule is vendored, because those carry the Semgrep Rules License v1.0, which is not on deny.toml's allow-list — and cargo deny governs crates, so it would never have caught a rule file. The position is stated in the rule header, the adapter docs and ADR-0018 rather than assumed.

Accepted gaps, recorded so they are decisions rather than oversights:

No new dependencies — Cargo.lock is untouched; exec-subprocess is std::process over the crates already present, and is off by default.

Stage 22b — osv-scanner: the dependency axis for Python, Java and Node → effort M ✅ delivered

Split out of Stage 22 because it is a different axis, not more of the same one, and because the SAST half is independently useful and independently reviewable.

Delivered. What shipped, and what it changes for later stages:

Stage 23 — Agent memory, episodic tier (ADR-0013) → v1.11.0 · effort M ✅ delivered

Goal: stop losing what sessions learn. Write path only — no retrieval ranking, no graph integration.

Delivered in #317. Migration 11 (agent_memory), the rto_graph::memory store, and roteiro memory add|list|forget. Every DoD item above has a test; memory cannot invalidate the fact cache, and is asserted so as a property of memory writes (tests/sync.rs::memory_writes_do_not_invalidate_the_fact_cache), because memory is not extraction output.

Four deviations from ADR-0013's proposed SQL, each deliberate:

  1. kind is a closed CHECK … IN over the ADR's own five names (lesson|attempt|decision|pattern|outcome), not the free TEXT proposed. Free text makes lesson/Lesson/lessons three kinds, none findable by a filter, and a vocabulary that cannot be filtered cannot later be ranked — which Stage 25 needs. Follows the analysis_runs.runner/isolation and media_content.kind precedent. Cost: a sixth kind is an append-only migration, not a string.
  2. superseded_at is TEXT, not INTEGER. An integer here would hold the generation of supersession — which is superseded_by, since the successor's id is the generation — so it would duplicate the column beside it. As TEXT it is a human timestamp on created_at's terms: written, displayed, never read.
  3. Four extra CHECKs making half-states unrepresentable, per migration 10's precedent: superseded_by/superseded_at stand or fall together (a moment with no successor is supersession inferred, the one thing the ADR rules out); nothing supersedes itself; anchor evidence requires an anchor key; empty scope/body refused.
  4. AUTOINCREMENT kept, and it is load-bearing — not decoration. A plain INTEGER PRIMARY KEY is the rowid, and SQLite reuses the largest deleted one, so forgetting the newest record would hand its number to the next write: ORDER BY id DESC stops being newest-first and a surviving superseded_by silently re-points at an unrelated record.

Scope is settled, so Stage 25 inherits it rather than re-litigating it (ADR-0013 v1.1 §Scope). The owner's rule: a lesson learned on a feature branch is valid on main only if the relevant association is merged to main in the same format — if not, then no. That needs no new machinery, because the anchor is the scope test: a record applies to a tree when its anchor resolves there with the same blob, or when it has no anchor at all (a general lesson, repo-wide). Drifted, vanished or unverifiable ⇒ does not apply here, kept and marked. "Same format" means the blob matches, strictly — a reformat breaks it, failing toward marked rather than toward silently applying a lesson to code that moved. Consequently scope is a coarse per-repo/project namespace and never a branch label; no branch bookkeeping exists anywhere in the schema. Recall in Stage 25 should rank on this predicate (AnchorState::applies), not invent a second one.

Out of scope, still: the bounded cache tier, recall ranking, decay, and any search integration — all Stage 25. Memory currently reaches search through no channel at all, which is asserted rather than assumed.

Stage 24 — boxlite sandboxed backend (ADR-0014) → v1.13.0 · effort L ✅ delivered

Goal: the reproducible, offline-capable local run — one command, pinned inputs, digest-level evidence.

Delivered. BoxliteRunner behind exec-boxlite, boxlite pinned =0.9.7, the DoD executed rather than argued: real semgrep 1.173.0 over one tree, once as a host child process and once in a digest-pinned microVM, 4 identical findings — PARITY OK … subprocess isolation=none image=none, boxlite isolation=microvm image=sha256:67319956…. 22 fault injections, 21 red on the right message; the one that was not is recorded below, because it did not go red and saying so is the point.

What the stage actually turned out to be about. The plan said publication was verified "so this is a dependency addition, not a packaging problem". Publication was real; the inference was not. boxlite from crates.io ships no hypervisor: its three -sys crates each detect a published package (.cargo_vcs_info.json) and disable themselves, and libkrun-sys excludes the sources they would build — enabling its krun feature compiles and then fails to link with 26 undefined symbols. What executes is a prebuilt runtime archive that boxlite's own build script fetches with a bare curl -fsSL, include_bytes!s into the rlib, and extracts and execs at run time. That fetch has no expected digest (searched four ways, NOT FOUND) and an env-overridable URL, so two builds of the same pinned version could embed different bytes undetectably.

So the stage's real work was governing that fetch, not adding a dependency:

  1. AssetSource::PinnedArchive — the first asset source with a compile-time digest, in runtime_pins.rs and shared with build.rs by include! so the two cannot drift. It closes the gap Fetcher's contract has to leave open for Download assets: a fetcher that reports success over a truncated body cannot defeat a pin checked in-crate before install. Tested with a deliberately lying fetcher.
  2. build.rs refuses to build unless BOXLITE_RUNTIME_URL names a local file matching the pin. Unset, remote, or wrong bytes are hard failures with a runnable recipe. boxlite's curl then never reaches the network, because what it is asked to fetch is already on disk.
  3. The licence acceptance is recorded in deny.toml and discharged by crates/rto-exec/NOTICE-boxlite-runtime.md, which is include_str!d and printed by prefetch before installing. The archive embeds GPL-2.0 (mke2fs, debugfs, libkrunfw) and LGPL-2.0-or-later (bwrap) binaries; they are exec'd as separate processes, so this is aggregation and Roteiro stays MIT OR Apache-2.0, but distributing a binary built this way carries GPL-2.0 §3 source-offer and LGPL relinking duties.

The gate that could not see any of this is now closed. crates/rto-exec/tests/build_script_fetch_audit.rs reads the build script of every package in the --all-features graph and fails on anything that looks like it fetches without a recorded pin. Measured: 613 packages, 89 of which have a build script (96 script files), 2 flagged (both governed), 0 false positives, in 0.3s — the audit prints those numbers on every run, so "it flagged nothing" stays checkable rather than trusted. Review asked whether the matcher was narrower than its docs claimed; it was, and the docs were the wrong half: widening to a bare http(s):// was measured at 29 flagged with 27 false positives (serde, quote, anyhow, winapi … all citing a docs URL in a comment), so the claim was narrowed to what the code does and a test now pins the two together. Its module docs state what it does not cover (helper crates, include!d build modules, obfuscation, run-time fetches, already-vendored bytes); read them before trusting it. This hole was general, not boxlite's: cargo deny --all-features check reported licenses ok while 25 MB of GPL binaries were being embedded, and was not wrong to — it governs crates.

Deviations and costs, stated rather than buried:

What this settles for later stages. Analyzer execution now has one provisioning contract with a real digest pin, so a future backend or analyzer image inherits verification rather than re-inventing it; and the build-script audit means the next dependency that fetches something is a failing test rather than a discovery. Apple Silicon microVM execution remains untested in CI — the parity proof above ran on Apple Silicon locally, and CI runners have no /dev/kvm, so the sandbox tests skip there with a visible message and the ingest and subprocess paths carry the functional coverage. That gap is unchanged and still accepted.

Stage 25 — Memory recall: cache tier, decay, supersession → shipped in v1.12.0 · effort L ✅ delivered

Goal: make memory useful — recall that ranks by evidence, plus the bounded cache that stops sessions re-deriving what they already know.

Delivered. Migration 13 (agent_cache + agent_cache_clock), retrieval-time ranking in rto_graph::memory, the byte-budget sweep, and the search memory channel. All four DoD items have a test, and each was fault-injected: the guarded behaviour was broken, the guard watched go red, and the source reverted byte-identically (15 injections, all red — two of them only after the tests they exposed as weak were strengthened).

What shipped, and where it deviated:

  1. Decay is none | linear[:span] | exponential[:half-life], and none is the default. ADR-0013 offered the three modes without saying which one leads; the reproducible answer is the default here, on the same terms as SearchOptions defaulting generated content off. Age is counted in generations — one per record written — so ranking never touches a clock.
  2. base_confidence defaults to 0.5 when a writer states none. Not in the ADR. 1.0 would let every record that claimed nothing outrank one that honestly claimed 0.9, pricing honesty; 0.0 would make the common case (the CLI states no confidence unless asked) unrecallable. The midpoint is the only value that makes stating one worth the trouble in both directions.
  3. anchor_penalty ranks drifted below vanished (valid 1.0, unanchored 0.9, unverifiable 0.5, vanished 0.35, drifted 0.25). Drift is the one state that can actively mislead about code still sitting under the same key; a vanished anchor can mislead nobody, and ranking it lowest would punish exactly the records the ADR says are worth keeping most. Two properties are asserted rather than assumed: nothing is ever zero, and every state that AnchorState::applies outranks every state that does not.
  4. The eviction counters needed a logical clock, so migration 13 adds a single-row agent_cache_clock beside the table (on sync_state's precedent). ADR-0013 §3 rules out wall-clock and the ADR's proposed agent_cache names generation/last_used without saying where the values come from. ticks advances per access — ModelCache's list position, made durable, with no ties for the sweep to break arbitrarily — and generation advances once per sweep, which is what makes "written in the current generation" a window rather than a single row, and what makes the pin lapse instead of becoming permanent.
  5. evict_count is a pure function tested against lru_evict_count's own numbers. Porting an existing policy is worth nothing if the port quietly behaves differently. The always-keep-the-MRU rule moved from a len - evict > 1 guard into the caller's pinned set, because this tier pins other rows too and one rule beats two. A sweep can therefore finish still over budget when everything left is pinned; CacheSweep::over_budget says so rather than leaving a bound that silently failed to bind.
  6. Memory reaches search through a third channel, --include-memory, off by default — beside the graph and generated channels, never merged with either. Its scorer shares no branch with the node scorer, so "memory never takes the authored +40" is structural. Stage 23's total absence assertion is restated rather than relaxed, and is now checked more sharply than absence could check it: a memory hit's score must not exceed the ceiling its own lexical terms can produce, which is what a leaked +40 would breach.
  7. CLI: roteiro memory recall [query] [--decay …] [--applicable-only] …, which prints every term of the ranking and not just the product, and roteiro memory cache [--sweep] [--budget-mb]. The sweep also runs at the maintenance seam in roteiro context --refresh, never on a read path. context --refresh --json keeps its long-standing shape: wrapping it would break callers to pay for maintenance they did not ask about, so the sweep is reported on stderr there and has its own --json under memory cache.
  8. Budget: --budget-mb, else ROTEIRO_CACHE_BUDGET_MB, else the 256 MB of §9.1. An unreadable value is an error, not a fallback — running the default under a name that says otherwise is how an operator ends up believing in a bound that was never applied. The config-file layer was deliberately not touched.

Two absolute assertions on a shared constant disarmed, at store.rs:1709 and :1880 — assert_eq!(store.schema_version()?, 11), which migration 13 breaks. Both now read migrations::latest_version(), the idiom a_later_migration_is_additive_on_a_populated_store already uses. The literal was defended in a comment as making someone confirm a new migration is meant to apply on open, but apply runs every migration newer than the recorded version, so there was never a per-migration choice there to confirm — the literal asserted the value of a shared constant and nothing else. Fault injection confirmed the rewritten assertions still catch the thing they are for: an apply that stops one migration short of latest_version() fails both.

EXTRACT_VERSION is unchanged, and tests/sync.rs::memory_writes_do_not_invalidate_the_fact_cache still passes — also fault-injected, by making a memory write clear sync_env.

Not in this stage, deliberately: the cache tier ships with its policy, its seam and its API, but no producer — node_context is still the context cache. Moving a live cache onto the bounded tier is a data migration, not a policy change, and bundling it would have put a schema move and an eviction policy in one reviewable unit.

Stage 26 — Analysis lenses (A1) → v1.15.0 · effort S–M per lens (independent track) ✅ delivered

Goal: deepen the graph itself — the on-brand work — with honest costs.

Cost correction, which this stage exists to respect: a fully surfaced lens is ~195–500 LOC across 6–8 files, not the ~20-line mirror previously assumed. That figure describes only the internal query fn. There are seven surfacing stages, not four: extraction (scan_markers + augment), the query fn, the query result types, MCP (GraphServer) and served-chat (GraphToolRegistry) as separate registries, Obsidian render, and CLI-side aggregation — plus tests and docs.

Shortlist, in order:

All three shipped, as three separate PRs — a reviewable diff beats a complete one nobody can check.

  1. Q3 — directed coupling ✅ (the standout). Calls edges already retain direction, and today's hotspot view throws that away by incrementing both ends. Highest value per line in the set.
  2. Q1 — debt density ✅. Builds directly on delivered intent-debt tracking.
  3. S1 — config-secret inventory ✅ (renamed, deliberately). Values are redacted before persistence, so this lens can report "secret-named config keys present and safely redacted" with paths and key names. It cannot detect hardcoded credentials in source, judge validity, or distinguish a real secret from a placeholder. The old title promised a scanner that this architecture cannot build.

Explicitly deferred out of this stage, with reasons:

EXTRACT_VERSION bump: required once, if and only if a lens adds derived extraction metadata (Q2 and Q10 do; Q1/Q3/S1 as scoped do not). Bumping invalidates every cached blob for every user and forces full re-extraction — so batch all extraction-touching lenses behind a single bump rather than paying it twice.

Also in scope (documentation debt): normalise the security taxonomy — the prose defines GDS/NNX/EXT/LLM while rows S1–S6 use undefined GPB/CVE/SAST labels — and mark the "SmolVLM is too small to emit <tool_call>" claim as a hypothesis, as it currently rests on no code or benchmark evidence.

Q3 — directed coupling ✅ delivered

rto_graph::coupling reports, per node, fan-in (distinct callers) and fan-out (distinct callees) over Calls edges, plus Martin's instability fan_out / (fan_in + fan_out). Ranked by total | fan_in | fan_out, ties broken by key, so identical input gives byte-identical output.

Surfaced on six of the seven stages; the seventh is a deliberate no-op.

StageQ3
Extraction (scan_markers + augment)Untouched by design — Q3 adds no derived metadata, so EXTRACT_VERSION stays 11. Confirmed live: sync after the change reported 237 of 239 blobs cached.
Query fnrto_graph::coupling
Query result typesCouplingReport / CouplingItem / CouplingOrder
MCP (GraphServer)coupling tool (rto-render/src/mcp.rs)
Served-chat (GraphToolRegistry)coupling tool — a separate registry, with its own test
Obsidian render_Home → "Most depended-on (call fan-in)"
CLI-side aggregationroteiro coupling [--order] [--limit] [--json], GET /v1/graph/{project}/coupling

Not surfaced: the explorer web app has no coupling panel. The /coupling endpoint serves the data; wiring assets/app.js is follow-up, tracked here rather than left to be discovered.

/hotspots is deliberately unchanged. Undirected degree over every edge kind is a different, still-useful question; the explorer depends on its shape. Its doc comment now says so and points at /coupling, rather than the discarded direction being an unremarked accident.

Two counting rules that change the numbers, both reported rather than silent:

Confidence signal, and why there is no CI gate. fan_in is exactly as precise as the edges beneath it, and the residual same-language case is not fixable here: a lone Rust join helper still absorbs every .join(…) in the workspace, reading as 232 callers. That is a limit of call resolution — fixing it is extraction work, and extraction work means bumping EXTRACT_VERSION. So the lens offers no CI gate, says so, and carries the caveat on every surface including both tool descriptions, so a model reporting a high fan_in passes it on. No suppression mechanism was added: coupling is a measurement, not a finding, so there is nothing to suppress — and a second exclusion vocabulary beside [debt] ignore would not have helped anyway (this repo's [debt] ignore does not cover the vendored cytoscape.min.js that dominates the ranking).

Scale benchmark, this repo (5501 nodes, 12036 edges, 6553 calls edges; 2887 coupled nodes), release build, warm, whole-process wall clock, best of 5:

CommandTime
roteiro coupling --limit 200.05 s
roteiro coupling --limit 0 (all 2887)0.07 s
roteiro debt (existing baseline)0.04 s
roteiro search store (scans all nodes)0.05 s

Ranking runs on the counts alone and only the nodes that survive the cap are read back, so a top-N question costs one edge scan plus N node lookups — not the whole-graph node scan /hotspots performs.

Documentation debt (both items), fixed in #288, which is where they live — not in any repo file:

The A1 cost line in #288 ("~20-line mirror of debt") is struck through and replaced with the ~195–500 LOC / 6–8 file figure this stage is built on.

Q1 — debt density ✅ delivered

rto_graph::debt_density ranks files by retained intent-debt markers per 1,000 lines. A raw marker count ranks by file size — the biggest file has the most lines to put a marker on — so a 40-marker file of 4,000 lines and a 40-marker file of 200 lines are indistinguishable under debt and twenty-fold apart under this. Built on debt's own output rather than re-walking markers, so the two lenses cannot disagree about which markers exist, and [debt] ignore applies unchanged.

The denominator was the design decision, and it did not need extraction metadata. lines is the file node's meta.lines — the count of \n bytes in the blob, recorded at extraction since Stage 2 and therefore already in every EXTRACT_VERSION 11 blob. Three alternatives were rejected: SLOC does not exist in the graph and computing it is net-new derived metadata (which would have moved Q1 into the Q2/Q10 batch below); per-symbol is what Node.span's byte offsets cannot give, since a span is not a line range; and the highest marker line is a lower bound on length, not the length. What lines counts is stated on every surface: every line, blanks and comments included — file length, not lines of code.

Two rules keep the arithmetic honest. A min_lines floor (default 50) keeps the denominator's tail out of the ranking — one marker in a 3-line file is 333 per kloc, true and useless — while leaving those files in files_with_markers and total_markers, reported as short_files. And ranking cross-multiplies the exact ratio in u128 rather than the rounded per_kloc, so densities differing in the fourth decimal do not tie and silently reorder on path.

Surfaced on six of the seven stages; the seventh is a deliberate no-op.

StageQ1
Extraction (scan_markers + augment)Untouched by design — the denominator was already extracted, so EXTRACT_VERSION stays 11. Confirmed live: a rebuild from an empty store reported 257 of 258 blobs cached, 0 extracted.
Query fnrto_graph::debt_density
Query result typesDebtDensityReport / DensityItem / DensityOrder / DEFAULT_MIN_LINES
MCP (GraphServer)debt_density tool (rto-render/src/mcp.rs)
Served-chat (GraphToolRegistry)debt_density tool — a separate registry, with its own test
Obsidian render_Home → "Densest files (markers per 1,000 lines)", under the existing intent-debt section
CLI-side aggregationroteiro debt-density [--kind] [--order] [--limit] [--min-lines] [--json], GET /v1/graph/{project}/debt/density

Not surfaced: the explorer web app has no density panel — the same position as Q3's coupling panel. roteiro debt is deliberately unchanged: a marker inventory is a different question, and check depends on its shape.

No CI gate, and the caveat travels on every surface (both tool descriptions, the _Home table, the CLI summary, the endpoint doc) — Q3's precedent, not one buried mention. Density inherits the marker scan's prose false positives (for now, placeholder, tbd fire on ordinary writing, so a design document ranks as dense debt) and adds one of its own: the denominator is file length, so verbose, generated or widely-indented files are systematically flattered and dense languages penalised. A gate would fail builds on prose and on formatting. The suppression story is the existing one — [debt] ignore globs and the roteiro:ignore / roteiro:ignore-file directives, applied before anything is counted.

One follow-up, found in review and fixed on the branch (fix(render): _Home scopes intent debt by [debt] ignore, both tables): the Obsidian _Home overview computed debt with an empty ignore list, so the vault reported a different debt for the same repository than debt, debt-density, check and the graph API — the disagreement ADR-0007 v1.1 was amended to end. Both _Home calls were fixed, not only the one Q1 added: leaving debt's wrong would have had the category totals and the density table on one page disagreeing about which files exist, which is worse than being consistently wrong. This was the third surface with that defect (#321 records the first two) and it was missed for the same reason each time — the earlier fix went to the surfaces that had been reported rather than to this stage's own list of seven, which names the Obsidian render. The enumeration that closes it: debt and debt_density are the only functions in the workspace taking an ignore: &[String], and every call site of either now passes the target project's list wherever one can be reached.

S1 — config-secret inventory ✅ delivered

rto_graph::config_secrets reports secret-named config keys — paths, key names, and whether each value was redacted before persistence — in three states, because collapsing any two would misreport them:

StateMeaning
redactedThe value was read from a config file and replaced with the placeholder before anything was stored. The expected state.
declaredThe key carries no value at all — a @rto:config struct field, declared in Rust with no literal to redact. Neither a redaction nor a leak.
presentA value that is not the placeholder. Extraction cannot produce this; Store::apply_import_layer can, so a non-zero count is a finding about this store, pointing at the importing tool rather than the source repository.

redacted_not_secret_named counts values redacted for where they live (a Kubernetes Secret's data) rather than what they are called, so the redaction figures reconcile against the graph instead of leaving an unexplained surplus. config_keys::REDACTED is now a single constant shared by the two redaction sites and this one reader, so the lens cannot drift from the redactor by a spelling.

[debt] ignore deliberately does not apply here, which is worth stating given the follow-up above. That list is defined as a debt-marker exclusion (ADR-0007), and config_secrets takes no ignore parameter at all — so there is no call site that could be passing the wrong thing. Giving it one would mean inventing a second exclusion vocabulary for a different question, which is the mistake Q3 declined to make for coupling.

The rename is the whole point, and it is enforced rather than merely documented. The lens cannot detect a hardcoded credential in source (that produces no config_key node), cannot judge validity (it never sees a value — no ConfigSecretItem has a value field), and cannot tell a real secret from a placeholder. An empty report means no secret-named config key — a statement about naming, not a clean bill of health: a credential under dsn or endpoint never appears. That limitation is carried in both tool descriptions in the imperative ("state the limits when you report it"; "if asked to scan for secrets, say plainly that this tool cannot do it"), in the CLI summary unconditionally including on an empty report, in the _Home section, and in the endpoint doc — and each of those is fault-injected. tests/config_secrets_cli.rs puts the same token in a .env (where it becomes a redacted config key) and in a Rust function body (where it becomes nothing this lens reads), so the boundary is a test.

Surfaced on six of the seven stages; the seventh is a deliberate no-op.

StageS1
Extraction (scan_markers + augment)Untouched by design — the redaction already happens there, so EXTRACT_VERSION stays 11. Confirmed live: 259 of 260 blobs cached, 0 extracted.
Query fnrto_graph::config_secrets
Query result typesConfigSecretReport / ConfigSecretItem / RedactionState
MCP (GraphServer)config_secrets tool (rto-render/src/mcp.rs)
Served-chat (GraphToolRegistry)config_secrets tool — a separate registry, with its own test
Obsidian render_Home → "Config keys named like secrets" (counts and files, not key names — a vault note is browsed out of context)
CLI-side aggregationroteiro config-secrets [--limit] [--json], GET /v1/graph/{project}/config-secrets

Not surfaced: the explorer web app has no config-secret panel. No ordering knob, deliberately: this is an inventory ordered by (path, name, key), and an order would imply some keys are more secret than others. No CI gate and no non-zero exit, even on an unredacted finding — that finding is about Roteiro's own import layer, so failing a user's build over it would be the wrong response.

Scale benchmark, this repo (5,860 nodes, 13,048 edges; 372 config keys, 48 markers over 20 files, 32,444 ranked lines), release build, warm, whole-process wall clock, best of 5:

CommandTime
roteiro debt-density --limit 200.04 s
roteiro debt-density --limit 0 (all 20 ranked)0.04 s
roteiro debt-density --limit 0 --min-lines 0 (floor off)0.04 s
roteiro config-secrets (default 50)0.04 s
roteiro config-secrets --limit 0 (all 372 keys scanned)0.04 s
roteiro debt (existing baseline)0.04 s
roteiro coupling --limit 0 (Q3 reference)0.08 s
roteiro search store (scans all nodes)0.05 s

Both new lenses are indistinguishable from the debt baseline. Q1 reads back only the files that actually carry a marker — one node lookup each, not a whole-graph file-node scan, which matters because file nodes carry captured meta.content.

On this repository S1 finds 2 secret-named keys among 372 config keys, both redacted, none unredacted — which is also a fair illustration of its reach.

The cost estimate, corrected a second time. Actuals per lens, and the figure each was measured against:

LensFilesInsertions
Q3 (#346)81,121
Q1 (#372)81,509
S1101,523
Stage total113,016

The RFC's original "~20-line mirror of debt" was out by ~75×. This stage's own corrected 195–500 LOC / 6–8 files is out by ~3–8× on lines and right on files: a fully surfaced lens is ~1,100–1,500 insertions across 8–10 files, and that is the figure a future lens should be planned against. The line count is dominated by tests and by doc comments defending the design decisions — Q1's denominator, S1's three states and its limits — not by the query, which is under 120 lines in both cases. S1 costs the same as Q1 despite a simpler query because the rename obliged the limitation to be restated, and fault-injected, on every surface.

41 fault injections across the two lenses (20 for Q1, 21 for S1), one per new behaviour on every surface, each caught by a named test with every file byte-identical after revert. Two were retained after they exposed real test gaps rather than being retargeted: Q1's ranking on the rounded ratio, and S1's CLI warning path, which extraction cannot reach and so needed a unit test on config_secrets_summary rather than a repository fixture.

EXTRACT_VERSION is still 11 and no lens in this stage touched extraction output. The batched bump (§8b: Q2, Q10, cross-language edge resolution) remains unpaid and unblocked.

Stage 28 — Generated media content moves out of derived (ADR-0015) → v1.10.x ✅ delivered (independent track)

Goal: stop generative model output masquerading as deterministic extraction — without losing the ability to search it. Resolves #300.

Delivered across two PRs, both merged:

Closes #300. Measured on a real repo: media build --audio refuses a silent clip in 0.013 s with no model load, versus 12.9 s and ~2 KB of confabulated prose under --force.

Known limit, recorded rather than papered over: MP3 and FLAC are not gated. They are entropy-coded, so measuring amplitude means decoding — behind the very load the gate avoids. The gate abstains, and abstention is a pass, so those formats still reach the model. Whether to close that gap with structural parsing (no dependency) is tracked separately.

Stage 29 — Audio metadata as derived facts (ADR-0016) → v1.11.0 · effort M ✅ delivered (independent track)

Goal: the complement of Stage 28. That stage took generated content out of derived because it is invented; this one puts extracted content in because it is present in the bytes — codec, sample rate, bit depth, channels, duration, frame count and tags, from a format read with no decoding and no model (measured 1–100 µs on this repo's own fixtures).

Stage 30 — MTP speculative decoding (issue #320) → v1.11.0, opt-in only · effort M ✅ delivered (independent track)

Goal: spend the draft head a Qwen3.5+ GGUF already carries. Generation on Metal is memory-bandwidth bound, so a decode that confirms three proposed tokens costs about what a decode producing one costs. The drafter is the model's own nextn block — no second model to install or keep in step — and on the vendored llama.cpp (b10200, pre-upstream #26296) its tensors are loaded whether or not anything uses them, so the memory is already paid for.

Stage 31 — Model lifecycle: resumable pulls, removal, high tier (ADR-0003) → v1.11.0 · effort M ✅ delivered (independent track)

Goal: make a multi-gigabyte model store survivable. Nothing here touches the graph, the schema or EXTRACT_VERSION — it is the store and its CLI only, which is why it rides an independent track.

Stage 32 — Guardrails: four ways a wrong answer looked like a right one → v1.12.0 · effort M ✅ delivered (independent track)

Goal: close four defects that share one shape — each let Roteiro, or its own documentation, state something false confidently. Every one of them was found during this project's own development, and every one passed CI. The fix in each case is to make the failure loud, not merely to make the output correct.

Known gap, not fixed here (separate work). The other direction is still silent: a store at migration 13 opened by a binary that knows only 1..12 opens without complaint. Reads stay sound — migrations are additive in effect, so the columns an older binary reads still exist — but sync would re-extract under an older EXTRACT_VERSION and rewrite the graph with worse content: a silent downgrade, not a crash. A hard error in apply was considered and rejected as the wrong granularity, since it would also block the reads that are provably safe. The fix belongs on the write paths (sync/reconcile/rebuild), refusing to rewrite a graph whose store is newer than the binary, and it needs a StoreError variant — a semver-visible addition on a 1.x crate. Filed rather than folded in.

Deliberately NOT done: per-worktree databases. findings, media_content, agent_memory and imports all live inside graph.db, and ADR-0013 v1.1 depends on that store being shared — its scope rule (a memory applies wherever its anchor resolves, with no branch bookkeeping) was demonstrated with one row in one store giving opposite verdicts on two branches. Splitting the database per worktree would silently reintroduce the branch-scoping that ADR rejected, and would need the ADR amended, not extended. Note the distinction the codebase already draws and this preserves: ObjectCache is content-addressed by blob id, so sharing it across worktrees is correct and valuable; the assembled graph is not, so sharing it would mean last-writer-wins.

Stage 33 — Local model resolution → v1.16.0 · effort S–M ✅ delivered (independent track)

Goal: one place decides which model serves a task, and can say why.

The user-facing gap: [models] has keys for embedding and generative only. Vision, audio and OCR are hard-coded string constants (voxtral-mini-3b, smolvlm-500m-gguf, ocrs-text), so a project cannot pin its ASR model today. Seven surfaces each pick a model by their own rule — spec draft, infer --model, serve load, serve/Ask answer, media generation, OCR during sync — and none knows the others exist.

Shape: one function in rto-graph — the crate that structurally cannot reach the network — taking (task_kind, modality, config, host_platform) and returning the model plus the rule that chose it, folding in the scattered call sites. Deterministic rules over categorical signals, not a classifier: every reliably observable signal here is low-cardinality (installed, modality, build feature, task kind), and a table over categoricals is the correct model.

The seed already exists and is the right one: chat_capable_model_ids filters models that cannot do the job, and exists because routing a BERT encoder through /v1/chat/completions aborts llama.cpp with a GGML_ASSERT. Generalise that, rather than starting from "which model is best".

No network, no new dependency, no ADR — it implements what ADR-0003 and ADR-0007 already document, so it takes amendments with version-history rows, not a new decision.

DoD: vision, audio and OCR models are configurable and pinned per project; roteiro config answers why did it use that model? for every surface; resolution is deterministic and unit-tested without loading a model.

What shipped

rto_graph::model_choice — resolve_with(task, pins) -> Result<ModelChoice, ModelChoiceError>, plus a process-wide pin slot published once at startup beside the existing [paths] model_store one. ModelChoice carries the model, the rule (pinned / built-in default), and whether the weights are on disk; the error type names the offending key.

[models] grew to five keys — embedding, generative, vision, audio, ocr — one per model kind, not per command, so generative governs both spec draft and Ask.

Signature, corrected. The plan said (task_kind, modality, config, host_platform). modality turned out to be the same axis as task_kind — a transcribe task is the audio modality — so a separate parameter would have admitted the meaningless pair (Ocr, Audio). Host platform is read inside, from the registry's existing Platform::host(), rather than passed: making it an argument would have let a caller ask about a machine that is not the one the model must load on. The signature is (task, pins).

Nine call sites, not seven. The plan's seven are all real and all folded in: spec draft (generative), infer (embedding — the config half only; a --model flag still wins and is validated by the embedder), serve load (served_models' kind filter), serve/Ask (chat_capable_model_ids, plus a startup check so a bad pin fails before the listener opens rather than per request), media generation audio, media generation vision, and OCR during sync. Two the plan did not name turned up while enumerating:

Two behaviours changed for a set config, deliberately. spec draft used to filter out a [models] generative that was not a generative model and fall through to the default — a silent fallback, and exactly the failure this stage exists to remove; it now refuses, naming the key. And a pinned model that is not installed is now a hard error there, matching what roteiro infer has always done with a configured embedding model. Unset behaviour is unchanged on every surface.

The one exception to failing loudly is roteiro config itself, which reports a bad key rather than refusing — it is the command an operator runs because a pin is misbehaving, so it must not be the command the pin breaks.

Cost, measured: +1,667 / −106 lines across 12 files (1,635 of the insertions are Rust across 9 files; the rest is the two ADR amendments, this entry, and the website's config sample). Against an estimate of S–M. The bulk is the resolver itself (769 lines, of which roughly half is the module's own documentation and its 12 unit tests) and the new 355-line CLI test. As with Stage 26, the surfaced-everywhere work is what costs: 366 lines of main.rs are the call sites, the roteiro config resolution table, and its --json twin.

Gates: fmt clean; clippy --all-targets and clippy --all-targets --all-features clean at -D warnings; cargo test --workspace --no-fail-fast 853 passed / 0 failed; --all-features --no-fail-fast 1,073 passed / 0 failed. EXTRACT_VERSION unchanged at 11, no migration, no new dependency, no network. Every new test was fault-injected — 12 unit tests and 6 CLI tests, each shown to fail under a mutation of the behaviour it claims to check, with the tree byte-identical afterwards.


Stage 34 — Remote model tier (ADR-0019) → v1.17.0 · effort L (independent track) ✅ delivered in three parts: 1 the guard, 2a the transport, 2b the surface wiring

Unblocked. ADR-0019 is Accepted (2026-08-17), so this stage has a settled contract to build against. It remains the largest posture change in the project: the first capability that sends repository content off the machine.

Cut in three, guard first, then the thing it guards, then the surfaces that use it. Part 1 — the consent gate, the payload allow-list, the dry-run and the egress record — landed in a build that compiled no backend and therefore could not send anything. Part 2a is the ureq transport, the TTY form of the invocation grant, the response reader, and the README/website promise amendments that the transport makes necessary. Part 2b — delivered here — is the model_choice amendment and the spec draft / Ask wiring: it was cut out of 2a because it lands in a shared resolver that seven surfaces already read, and merging that with the first code in the project that can open a socket would put two unrelated risks in one review. See What shipped below.

Goal: an optional, explicitly-consented remote model backend for work local models cannot do.

Why an ADR is a prerequisite rather than paperwork. This is the first capability that sends repository content off the machine, and three written promises currently forbid it:

The framing that decides the design: mis-routing among local models wastes tokens; mis-routing outward sends source off the machine for a reason nobody can inspect. The local→remote edge is not a routing decision — it is a gate the user opened. So no learned router, at any model quality.

And the disclosure gap must be stated in the ADR, not deferred: extraction redacts secret-named config keys before persistence, but that is name-matching over ten needles and there is no redaction chokepoint on a prompt. Prompts carry symbol names and prose; DATABASE_URL=postgres://user:pw@host matches none of those needles.

Also unresolved by design: ADR-0015's Producer identity folds a model_digest. A hosted model has no digest — a vendor model string is a mutable pointer, and the weights behind it can change while the name does not. If remote output is ever stored, it needs ProducerTrust::{PinnedDigest, VendorAsserted} so a record states on its face that its identity is a claim.

Sequencing note: Stage 27 re-audits every "offline" claim. Landing this before it converts a documentation task into a re-litigation of the product's identity.

What shipped — part 1 of 3: the guard, before the capability

The stage is cut at the seam ADR-0019's own structure suggests. Part 1 is everything that decides, shows and records; part 2a is the thing that sends. The reason is not PR size, though the size is real: an egress path whose guard lands in the same change as its transport is a guard nobody reviewed on its own, and ADR-0019 §4 names that failure explicitly — "deferring this is how an egress path ships before its guard". Landing the guard first inverts that, and leaves a build that cannot send anything at all to review it in.

rto-remote — a new crate holding the policy and no HTTP client. call_with takes the transport as a caller-supplied closure, exactly as rto-exec takes its Fetcher, so the code that decides whether bytes may leave is not the code that can make them leave. The guarantee is checkable from a Cargo.toml rather than promised in prose, and every test exercises the whole path with no network — a test cannot accidentally become the first thing that sends data. rto-graph gains nothing: its gix is still pinned default-features = false, and rto-remote depends on it.

Five modules, one per clause:

Config. [remote] grows three keys, of which exactly one inverts. enabled goes through ConfigGrant; endpoint and model are ordinary keys and layer ordinarily, so a project may choose where its gateway is without being able to turn the tier on. roteiro config prints the section as layers rather than one merged value, because a reader applying the general precedence here would be wrong about the one key where being wrong means believing egress is off when it is on. Without the feature the section still prints, saying the build has no tier — an omitted section reads as "no such setting".

CLI. roteiro remote status | dry-run | log, behind an off-by-default remote feature that adds no third-party dependency. status reports the gate layer by layer then the decision; dry-run prints the exact bytes and sends nothing; log reads the ledger and says "nothing has left this machine" rather than leaving that to be inferred from silence.

What shipped — part 2a of 3: the thing the guard was for

Part 1's build could not send. This one can, and the change of posture is stated everywhere it is now false to say otherwise.

The transport, and where it is not. crates/roteiro/src/remote_transport.rs — one ureq call, in the binary, reachable from one command. rto-remote still holds no HTTP client, and call_with is unchanged: the transport is handed to it as a closure, so the crate that decides whether bytes may leave still cannot make them leave, and that stays checkable from crates/rto-remote/Cargo.toml. Three things it does not do, each for a written reason: no reachability probe (ADR-0019 §2 — a probe is egress, and a DNS lookup leaks the query to a resolver), no retry (a retried call is a second disclosure, and that decision belongs to whoever consented), and no redirect (max_redirects(0) — the ledger records the endpoint that was consented to, and a 302 would make that record a lie about where the bytes went). It is not a new dependency: ureq was already in the tree for model pull and security prefetch, exactly as ADR-0019 anticipated.

The credential is an environment variable and cannot be a config key. ROTEIRO_REMOTE_API_KEY, because roteiro.toml is committed by design — the same fact that inverted the precedence for enabled — and a key that could be set there is a key that gets committed. It cannot reach the ledger either, and not by discipline: the ledger records Payload::body, and headers are not part of it. remote status reports whether one is set, never what it is.

response — the receive side, in the crate with no socket. rto_remote::response::parse is the mirror of Payload::body: pure, over a &str, so a truncated body, a malformed body and a body the endpoint filled with its own error are all string literals in a unit test. It refuses a generation that stopped short (finish_reason: length, content_filter, anything outside a short list of completion reasons) rather than returning it, because a completion that stopped early reads as finished — handing it over is the silent downgrade ADR-0019 §6 most needs to prevent, wearing the endpoint's own name. It also reports the one identity check this machine can make: a vendor model string is a mutable pointer, so the weights cannot be verified, but a name that answered under a different name than the one requested is a discrepancy, and it is printed rather than dropped.

The prompt, and the two commands that must never show it. remote call is the only surface that asks. status and dry-run are the commands you run to find out what would happen, and a command that asks permission in order to tell you is useless — so the rule is asserted rather than assumed, with a test that runs both in the exact gate state a prompt could resolve. And exactly one gate state may be resolved by asking: InvocationUnset — the human opted in, this run has not. A prompt may never stand in for the user layer, or the two grants ADR-0019 §3 requires separately collapse into one keystroke; it may never override a project denial; and a non-interactive stdin is refused, not assumed to agree, because a pipe cannot consent.

Two review findings on #386, both about a message that was untrue rather than merely unhelpful — the class this project keeps finding, because such a message passes every gate.

Declining a prompt claimed you had passed a flag. The invocation's two forms were collapsed into one Option<bool>, so answering no produced InvocationDenied, whose text names --no-remote — a flag the person never typed and would not find in their shell history. On the consent path especially, a message that misreports how consent was withheld undermines the thing it reports on. Fixed with Invocation::{Unset, Flag, Prompt} and a seventh Reason::PromptDeclined; decide keeps its Option<bool> flag form and became a thin wrapper over decide_with, so there is still one implementation of which layer outranks which. A distinct variant rather than a payload on the existing one, for a reason that decides it: Reason::as_str is the stable token in remote status --json, so a variant adds a token where a payload would change the shape of one readers already parse. The test asserts the rendered text, not the variant — a Reason that classified right while still printing the flag would be the same bug.

And that variant was a semver break, caught in review. rto-remote had been published at 1.19.0 hours earlier, and Reason was not #[non_exhaustive], so the seventh variant would stop a downstream exhaustive match from compiling. Three options were weighed. Design around it, as StoreError did (#342/#348), does not transfer: there the fact had a home outside the enum, where here the fact is which reason to report, so moving it out would leave --json emitting invocation_denied for a prompt and reinstate the defect one layer down. Accept it and mark the commit ! would cut 2.0.0, the version this plan reserves for Stage 27 — the mistake #341 already made by accident. So: #[non_exhaustive] went on in the same change, along with the rest of the crate's open enums, while the crate had nine downloads and no consumer that could exist. That is the workspace convention rather than a departure from it — ExecError, SubprocessError, AssetError, AssetSource and NetworkPolicy all carry it, and AssetSource's docs record that the attribute "is what made adding it a non-breaking change". StoreError is the one that missed the convention and paid for it. Trigger and ProducerTrust stay exhaustive on purpose — closed sets, where a new member is a redesign a downstream match should be made to notice — and say so at their definitions. The version number will not record the break; Reason's doc comment, the crate README and the commit do.

An absent finish_reason was read as "it finished". parse only refused when the field was present and outside the allow-list, so silence passed — a strictly weaker reading than the length case it already refuses, since length at least says something. This is #367's rule at the other end of the same wire: a length that cannot be established is not a length that checks out. Now ResponseError::Indeterminate, and the message says why (completeness could not be established) rather than that a field was missing. Checked before changing it that nothing this tier addresses legitimately omits the field: rto-serve's own ChatChoice::finish_reason is a non-optional &'static str, and the one shape that legitimately carries null is a streaming delta, which Payload::body's pinned "stream": false means this tier never asks for — so a null here is an endpoint streaming at a request that said not to, which is the least complete a body can be.

Testing the untestable bit. Part 1 could say the binary compiled no backend. That sentence is now false, so the guarantee is re-established on different ground rather than dropped: the granted-path CLI tests point [remote] endpoint at http://127.0.0.1:1/…. Loopback is not a network — the kernel cannot route those bytes to another host — there is no DNS lookup (which would itself be the egress §2 describes), and nothing listens, so the call is refused in microseconds. Everything about a successful response is tested over string literals in rto-remote. The failure paths all have tests: no network, an endpoint refusing by status, a malformed body, a truncated generation, and — the one that matters most — a response arriving after the ledger write, asserted from inside a failing transport closure, which only the write-first ordering can satisfy. Reading the ledger after the call returns would be satisfied by a single line written at the end, and the calls worth knowing about are exactly the ones that never returned.

The promises, amended on the commit that makes them false. Part 1 left "nothing leaves the machine" alone, correctly: with no backend compiled it was still literally true of every build it produced. This build can send, so the README and the website now say the scoped thing — that the sentence is true of Roteiro as shipped and as configured, not of the software as a whole — name the tier in their feature tables, and flag that --all-features includes it. ADR-0006 was already scoped to serving in v1.3. The website additionally carries the terminology guard ADR-0019 asked for: "Online mode" keeps its existing meaning (a one-time, consented download, after which inference is local), and a note beside it says in as many words that enabling inference-local-models did not enable this.

What shipped — part 2b of 3: the surfaces, over a shared resolver

Parts 1 and 2a built a tier you had to name to use. This part puts it behind two surfaces where a person could reach it without typing remote — and most of the work is making that impossible to do unaware.

ModelSource::Remote { trust }, and the three things it forced. The variant carries the trust grade rather than a model name, because that is the whole of what it exists to say: a hosted model has no registry entry and no digest, a vendor model string is a mutable pointer, and a resolution must state on its face that its identity is a claim (ADR-0019 §5). Three consequences, none of them predicted by the plan and all of them structural rather than stylistic:

The blast radius, counted rather than estimated. model_choice::resolve spelled literally has 3 non-test call sites; the resolver family — resolve, resolve_model, resolve_model_with, resolve_models — has 10. Eight needed no edit, because 2b adds resolve_with_remote as a new entry point rather than changing resolve_with's signature; the other two are installed comparisons. A RemoteTier argument that defaults to Unavailable is what makes that safe: a caller who forgets the tier gets the local answer by construction rather than by remembering to.

Precedence, decided in the open. The pin resolves first, and its errors still fail, even on a run whose local answer is about to be discarded — a --allow-remote that made a broken [models] generative stop being reported would hide bugs as a side effect of granting egress. Then the tier wins, because a per-run flag someone typed is more specific than a standing project default — and it wins out loud: ModelChoice::why names [models] <key> as not applying, so a displaced pin is reported rather than silently skipped. Only Draft and Chat are eligible; the other four tasks are refused structurally, because Transcribe, Describe and Ocr run inside extraction where there is no invocation to grant anything, so consent there could only come from a config value — the user layer again, which ADR-0019 §3 says never suffices alone.

spec draft: the flag is the only way, and a refusal stops the run. --allow-remote / --no-remote, and no TTY prompt — deliberately unlike roteiro remote call. That command prompts because sending is what it does; here the default is local, and a prompt on a default path turns a habituated "y" into consent-by-default, which is the thing a two-layer gate exists to prevent. The harder half is the refusal: a --allow-remote the gate turns down fails, naming the layer and its remedy, rather than drafting locally. Handing back a local model's prose there is a different answer with no signal that anything changed — ADR-0019 §6's named failure, arriving through the consent gate instead of through a socket. It is the same silent downgrade wearing a different hat, and it is the failure this part most had to get right.

The remote draft is not the local prompt on a wire, and could not have been. rto_spec::draft_prompt interpolates grounded symbol and ADR names into a string. Sending that string would put graph content on the wire without it ever passing ContextItem::from_node — the allow-list would still exist and simply have nothing to do, which is the "whatever the local path happened to build" assembly §4 forbids by name. So the remote path rebuilds: the instruction carries the task and no graph content, and the nodes travel as allow-listed context items. The cost is stated rather than hidden — a remote draft is not the same prompt as a local one — because keeping the allow-list load-bearing is worth more than prompt parity on the one path that leaves the machine. One call per unfilled section, so each disclosure is its own ledger line rather than several hidden inside one.

Ask: the served Engine, wrapped in the binary. The explorer's panel POSTs /v1/chat/completions, which rto_serve::server handles over an Arc<dyn Engine> — so Ask is not a Roteiro function and the branch had two possible homes. crates/roteiro/src/remote_engine.rs is the one that changes nothing: rto-serve gains no code, no feature and no dependency, and the socket stays in remote_transport. The hosted model becomes one more served id and leads the Ask pool, so models[0] — what the UI sends — is the tier a server was started with --allow-remote to use; every local model stays served and addressable, so a default moved rather than anything being removed. The cost is named rather than discovered: that id is in GET /v1/models, so any client on the port may address it — bounded by loopback-by-default, by the user layer having granted independently, and by the ledger.

A chat transcript cannot be wrapped in an allow-list, so it is reduced. With graph tools on, chat_with_tools injects tool results — raw graph query output — into the conversation before the model sees it. Proxying that array would route graph content past the guard entirely. So the remote path takes the user turns and only those, drops assistant, tool and system turns, grounds the question itself through rto_spec::context, and does not run the tool loop remotely at all. A remote Ask is therefore not the same conversation as a local one. That is a trade, made in the open, in the same direction as the draft path's.

serve --allow-remote is a process-scoped grant, per ADR-0019 v1.2. Decided once, in the command dispatch, before anything is built or bound — so a refused grant stops the server from starting rather than letting it come up and answer locally. The Decision is then a field on the engine: never recomputed, never persisted, never read back from anywhere, dropped with the process. There is no code that could infer a grant from a previous session because there is nothing to infer one from.

remote now implies models, and it is a resolver dependency. The tier does not choose itself: ModelSource::Remote is a variant of the shared resolver, which is gated on rto-graph/models, and both surfaces ask that resolver which backend serves them before either can send. Without it a --features remote build would hold a consent gate and a transport and nothing that decides when to use them. It pulls no llama.cpp and enables nothing — and it is what makes spec draft reachable under remote alone, which is the point of a tier meant for work local models cannot do.

Two defects found in review, and one gap the merge opened. Recorded because each was a class rather than an instance, and the class is the part worth carrying forward.

The promises, amended again on the commit that changes them. 2a's README and website text described a tier reached through roteiro remote …. Two more commands can send now, so both say so, name which one prompts and which two do not, and state the serve exposure in the terms ADR-0019 v1.2 uses — including that the person starting the server consents on behalf of every later request to it.

Gates: fmt clean; clippy --all-targets and --all-targets --all-features clean at -D warnings; cargo test --workspace --no-fail-fast 981 passed / 0 failed and --all-features --no-fail-fast 1,240 passed / 0 failed; cargo deny clean. EXTRACT_VERSION unchanged at 12, no migration, no new dependency, and no network — in either the code or the tests, where the only address any granted call may reach is http://127.0.0.1:1/… (loopback, a literal so there is no DNS query, and nothing listening). Every new test was fault-injected: 6 resolver, 1 re-export, 7 engine, 1 Ask-pool and 10 CLI, each shown to fail under a mutation of the behaviour it claims to check, with the tree byte-identical afterwards. The collision guard's test asserts the rendered message names both the key and the local id, because "that name is taken" is unactionable without saying by what; and the semver guard was shown to fail under a mutation that empties its re-export list, which is the regression the type's move actually caused. The Ask grounding budget is asserted at compile time rather than in a test, so violating it fails the build that edits the literal.


Stage 35 — roteiro review LLM mode → v1.18.0 · effort M–L (independent track) ✅ delivered — the instrument, and a negative result on the reviewer

Depends on Stage 33 (a reviewer must resolve a model without a fourth bespoke rule). Independent of Stage 34 — it can run wholly local.

Split into two PRs at the measurement seam. 35a — the scoring harness and the suppression filter — is delivered: roteiro review --score, the corpus as a typed shipped asset, per-class recall, and rto_graph::compile_claim. It lands before any reviewer because it is what makes a reviewer's value a number instead of an impression, and because a "do not build" verdict needs the same harness a "build it" verdict does.

35b is cut in two at the measurement seam, and PR 1 is delivered: the reviewer, its two surfaces (review --llm and the review --replay harness), ModelTask::Review, and the compile-claim suppression wired through compile_claim's four axes. PR 2 is the graph-context arm and the comparison that is the actual experiment. PR 1 ships no per-class recall figure — see below for why that is a result rather than an omission.

The original 35b framing follows, and its second bullet has since been overturned by measurement. Two measurements from 35a constrain it, and both were taken on this repository rather than assumed:

35b also needs a resolver addition, not a workaround: ModelTask::Review in rto_graph::model_choice, sharing the generative key with Draft/Chat. Adding a seventh task is the Stage 33-sanctioned move; a bespoke selection rule in the reviewer would be the fourth one Stage 33 exists to have removed.

Goal: give the adjudicated review corpus a consumer, and put the graph to work on the one thing a diff-only reviewer structurally cannot see.

The asset that already exists: crates/rto-graph/tests/fixtures/review/ holds 26 adjudicated review comments with verdicts, defect classes, and the reviewed_sha each was left on. That makes a reviewer measurable rather than guessed at — which most projects cannot do. It currently has no consumer, and that is how a fixture rots.

What the graph adds, stated honestly: not access — Copilot has been agentic since March 2026 and reads repository context via tool calls, verified against this corpus (it cited a file outside a PR's changed set, correctly). What roteiro review has is pre-assembled, provenance-tagged context: governing ADRs, authored drift, blast radius, intent debt. A weaker claim than a moat, and the one to test.

Two constraints inherited from the investigation:

DoD: scored against the corpus with per-defect-class recall — not an average, which hides the only thing an implementer needs — at the reviewed_sha of each comment, never the PR head (merged heads contain the fix commits, so scoring against them measures recall on already-fixed code and silently reports zero).

Delivered in 35a, so 35b inherits rather than rebuilds it: review --score, per-class recall with the denominators printed beside the rates, an outright refusal to score a commit the corpus does not know (the PR-head guard), and compile_claim's coverage model. One thing 35a had to fix on the way: the corpus README's own reconstruction recipe — merge-base <base> <reviewed_sha> — produced an empty diff for 13 of the 15 review commits, because a merged PR branch is an ancestor of main and the merge base is then the review commit itself. A reviewer handed an empty diff also scores zero, silently, from the opposite direction. The recipe is corrected and now has an executable form (every_row_reconstructs_a_non_empty_reviewed_diff), which also checks that each reconstructed diff touches the file its comment is anchored to.

What PR 1 measured, and the one that matters most

The budget is not the constraint — the earlier conclusion understated the room. 35a established that whole-diff review does not fit and named ~79k per file as the figure to design to. Measured across all 190 changed paths of the corpus (of which 184 have a reviewable diff; six are binary audio fixtures): mean 2,704 tokens raw and 3,275 as sent, median 1,476/1,758, p90 5,621/6,711. Exactly one of 190 exceeds even the ~30k single-call budget, and it is a generated JSON fixture. The largest reviewable source file-diff in the corpus is 14,034 raw and 17,202 as sent — under two-thirds of the single-call budget. "As sent" carries the line-number column the reviewer adds, a measured 1.21× (9 characters per line, so ~1.2× on source and far worse on very short lines).

That is what makes PR 2 worth running at all: the median file leaves ~28k of the budget unused, so the graph arm has somewhere to put whatever it has to add. Had the budget been tight, the stage's central claim would have been untestable here whatever the graph contained.

The instrument had a vacuous measurement in it, and PR 1 walked into it. A reasoning GGUF opens <think> and deliberates before answering. Run with a 1,200-token generation cap, qwen3.8-27b spent the entire budget inside that block on 4 files of 4 of a held-out commit, and the harness reported "0 finding(s) over 4 file(s)". Nothing was wrong-looking about that output. Scored, it would have produced zero recall across every class and read as a clean, honest negative result about local reviewers — when in fact no review had happened at all.

This is the same silent zero as scoring against a PR head or against an empty diff, arriving from a third direction, and it is worse than both because the other two are refused by construction. It is now detected (reviewer::Parsed::reasoning_truncated), reported on both surfaces as not reviewed, and pinned by a test; the cap is 4,096 and the context window is passed explicitly rather than inheriting LlamaEngine's 4,096 default, which had already killed an earlier run outright. A truncated file is never counted as clean and a run containing one is never scored.

The general lesson, for anyone measuring a reasoning model on anything: a generation cap does not degrade a reasoning model's answer, it removes the answer while leaving the response well-formed. Silence and truncation must be distinguishable at the instrument, not inferred afterwards from a suspicious number.

Not yet scored, deliberately. PR 1 ships no per-class recall. The pass that would produce one had not been run under the fixed harness at the time of writing, and a recall figure from a run where truncation was possible is not a measurement. The remaining observation is honest but not a score: with qwen3-coder-30b-a3b (non-reasoning, so unaffected by the above) the reviewer emitted 60 findings across 4 held-out files — ~15 per file, with visible duplicates — and two prompt revisions written specifically to impose precision discipline changed that number not at all. Against 22 adjudicated rows spread over 184 files, that rate is the number to beat, and it is the one PR 2 has to move.

"Do not build this" remains an acceptable outcome. The corpus keeps its value either way: it is how any future reviewer, hosted or local, gets measured. Nothing in PR 1 is evidence for building PR 2 except the headroom finding; the finding rate is evidence against, and both go into the decision.

What PR 2 measured — the graph arm, and the experiment it was built for

The pre-committed floor was not cleared, and the finding that matters is that it could not have been. The interesting half of this result is not a disappointing number; it is that three measurements taken before the graph arm ever ran bound what any run of it could possibly show, on this corpus, with any model. That bound is stated first because it is the durable part.

The sequencing, and why the baseline had to come first

PR 1 shipped no per-class recall deliberately — a figure from a run where truncation was possible is not a measurement — so there was no baseline, and "the graph arm found N contract-drift rows" would have meant nothing beside it. PR 2's first deliverable was therefore the diff-only arm under the corrected harness: GraphContext::none, the truncation guard armed, scored per class at each row's reviewed_sha. The graph arm followed as the same binary, the same model, the same corpus and the same reconstruction, differing in exactly one variable — which is now recorded in the run document (RunArm) rather than in a filename, because a comparison a reader cannot audit from the artifacts is not one.

Three measurements that bound the experiment, taken before it ran

1. The recall ceiling is 21 of 22, not 22 of 22. Scoring matches a finding to a row within LINE_WINDOW (±10 lines), so a row whose anchored line is not in the reconstructed diff cannot be found by any per-file diff reviewer, whatever the context. Measured over every row: 20 of the 22 real rows have their anchored line inside the shown -U3 diff, 21 fall within the scoring window, and one does not. It is crates/roteiro/src/config.rs:634 — the ignore_reset inherited via .or() — whose line sits in the gap between hunks covering 564–575 and 984–1180. It is a contract-drift row, so that class's real ceiling is 4 of 5, not 5 of 5. The reviewer that found it originally was agentic and read the whole file; a per-file diff reviewer structurally cannot, and no amount of graph context moves a line number into a hunk.

2. The arm's dose is real, not nominal. Over the whole corpus the assembler emits 922 items on 85 of the ~184 reviewable files — 189 authored (governing ADR and blueprint sections) and 733 derived (doc comments from elsewhere in the file) — for ~152k estimated tokens, about 1.8k per file that carries any, with 2,305 further items dropped by the cap. On a median file that roughly doubles the prompt. So the treatment was administered; a null result here is not a null dose.

3. And the one that decides it: on 3 of the 4 reachable contract-drift rows, the two arms send a byte-identical prompt.

contract-drift rowreachable?context the graph arm supplied
rto-graph/src/engine_slot.rs:16yesnone — file added in that commit
rto-graph/tests/fixtures/audio/README.md:11yesnone — markdown fixture, no symbols
docs/adr/0005-image-ocr-vision-ingestion.md:16yesnone — an ADR under review
rto-graph/src/query.rs:716yes19 items (0 authored, 19 derived)
roteiro/src/config.rs:634no21 items (6 authored, 15 derived)

Each zero has a specific and defensible cause, and none of them is a bug in the assembler:

So at most one contract-drift row could ever change — and the baseline below makes that bound tighter still. The single reachable row that does receive context, query.rs:716, is the one row the diff-only arm already matched. The graph arm therefore cannot gain a contract-drift row from any starting point; it can only lose the one already held. The floor — at least 2 of the 5 recovered that the diff-only arm missed — was not merely unreachable before a single token was generated, it was refuted: the number of recoverable rows is zero. That floor was chosen honestly, before anyone knew which way the numbers would fall, and it stands as written; what the measurement adds is why it could not be met, which is worth more than the miss.

Variance: settled before any conclusion was drawn, and it is zero

Stage 30 measured that llama.cpp's logits differ between batch widths, so "same seed, same answer" could not be assumed here — a single run of each arm would not have been a comparison. Measured rather than assumed: the diff-only arm was run twice over the first 4 corpus commits (54 files, 29% of the corpus), the second time through a different binary — the one carrying this PR's changes.

626 findings both times, 534 distinct both times, Jaccard 1.0000, not one finding different. So two things hold at once: greedy decoding at temperature 0.0 with speculation off is reproducible run to run here, and the refactor did not perturb the diff-only path. Stage 30's divergence needs ROTEIRO_SPECULATIVE, which no run in this stage set.

Stated limit: determinism was verified on 54 of 184 files, not all of them. The remaining 130 were not re-run, because the finding below made a second full pass the wrong place to spend three hours.

What the reviewer actually emits, which is the adoption verdict

diff-only arm
findings1,995 over 183 files — 10.9 per file
exact duplicate lines346 (17.3%)
labelled contract-drift1,412 (71%)
claiming compile failure0
files declared clean127 of 183
unadjudicated1,990 of 1,995

Two of those rows are results in their own right.

The reviewer says contract-drift about almost everything. 71% of its output carries that one label, and it recovered one of the five real contract-drift rows — by chance, on the permutation test above. A classifier that answers the same thing to nearly every question carries no information in its answer, and it explains why the class the graph arm was built to help is also the class the model over-claims: the arm was aimed at the one place the reviewer was already saturating.

The compile-claim filter never fires. 35a called it a free precision filter, on the strength of every false positive in the corpus being a compile claim (4 of 4). Against this model it is dead weight: not one of 1,995 findings claims the code will not compile. The filter is correct, cheap and idle. That is not an argument to remove it — a different model would claim differently, and the corpus says a human reviewer did — but it must not be counted as precision this reviewer is getting.

The diff-only baseline, and the discovery that it is not a measurement

The baseline this stage never had: qwen3-coder-30b-a3b, 15 of 15 commits, 183 files reviewed, one file refused (Cargo.lock, 64,389 tokens against the 49,152-token window) and one truncated. 4 of 22 real rows found, 1 of 4 known-false reproduced, and 1,995 findings emitted — 10.9 per file, of which 1,990 are unadjudicated.

Then the number was checked against a null, and it did not survive.

review_score::match_findings credits a finding to a row on (commit, path, line ±LINE_WINDOW) and never on what the finding says — not its class, not a word of its description. That is the correct rule for a scorer that must not reward eloquence. Its consequence had not been measured: a reviewer emitting 10.9 findings per file blankets the diff, and a "hit" is then explained by density rather than by insight.

Permutation null. Relocate every corpus row to a uniformly random line its own reconstructed diff actually shows; leave the run's findings byte-for-byte as emitted; rescore with the same greedy one-to-one rule. Over 2,000 trials:

observednull meanP(≥ observed)
real rows matched44.190.72
known-false reproduced10.490.49

The reviewer scored below chance. A tighter null that relocates only to added lines — where 23 of the 26 rows actually sit — gives 3.98 and P = 0.62, so the result is robust to the obvious objection. The reimplemented matcher reproduces the shipped scorer's counts exactly (4 real, 1 known-false), which is what licenses the comparison.

So 4/22 is not a recall figure, it is arithmetic, and the same is true of the 1/4 known-false "reproduction": inspected, the credited finding is a contract-drift claim about a Range: bytes= doc at line 2312, while the row at 2322 is a false compile claim. Different claims, ten lines apart. The scorer cannot tell them apart and was never built to.

This is the fourth silent-zero-shaped trap in this stage, and the first that is silently non-zero. Scoring against a PR head, against an empty diff, and against a truncated reasoning reply all report a clean zero that measures nothing; this reports a clean positive that measures nothing. It is now caught in code rather than in prose — Score::expected_by_position prints beside the recall, and caveats() states outright that the rate is not clearly above chance — on the same principle that made reasoning_truncated a reported outcome in PR 1: a number whose null is not stated is not yet a result.

The shipped estimate is honest about being an estimate. An exact null needs the diff, and review --score is pure by design so a published score recomputes on any machine, so it uses the candidate's own findings as the proxy for where it looked. Summing the ±10 neighbourhoods over-reads at 6.2; merging them under-reads at 3.0 against the permutation's 4.19. It is an order-of-magnitude guide, the caveat fires on a 2× margin because of that, and the docs name the permutation as the thing to run before believing any comparison.

What this costs the experiment. A recall comparison between two arms is only meaningful if either arm's recall is meaningful. At 10.9 findings per file neither is. The density sweep says how far that has to fall: thinning the run's own findings and re-running both the null and the observed match, the two stay within noise of each other at every rate measured — 4.19 vs 4.00 at 10.9 per file, 1.67 vs 2.08 at 2, 0.67 vs 0.94 at 0.5. The reviewer's recall never separates convincingly from chance at any density it was run at. PR 1 called the ~15-per-file rate "the number to beat"; it is worse than that — until it falls, the corpus cannot measure recall at all.

The premise, restated against the data. Stage 35's central claim was that contract-drift "puts the largest class squarely on the graph: per-file review cannot see a doc in another file contradicting the code under review." Measured against the rows rather than assumed: four of the five contract-drift rows have both halves inside the file under review, and three of those have both halves inside the diff. The class is not, on this corpus, cross-file. The premise was reasonable and it was wrong, and it was wrong in a way only a corpus could show.

The graph arm, scored — and the floor, settled

The graph arm completed and was scored against the same corpus with the same binary, so the two arms differ in exactly one variable.

diff-onlygraph
real rows found4 / 225 / 22
contract-drift1 / 51 / 5
known-false reproduced1 / 40 / 4
findings emitted1,9952,059
unadjudicated1,9902,054

contract-drift did not move, exactly as the power bound said it could not. That bound was computed before the arm ran: three of the four reachable rows receive no context at all, and the fourth is the one the diff-only arm already held. Recoverable rows were zero, and zero is what changed.

Against the floor pre-committed in PR 1:

The one row gained is missing-event, a class with a single real row. The scorer's own caveat says a singleton class is one bit rather than a rate and must not be read as a percentage. And it was gained while the finding count rose by 64 — so under the density explanation the baseline established, more findings producing one more hit is the expected direction of noise, not evidence of insight. A 4 → 5 move sits inside a null whose mean was 4.19.

Verdict: the graph arm is not adopted on this evidence, and the corpus cannot be made to say otherwise. Not because the graph is useless — because at ~11 findings per file neither arm's recall separates from chance, so a comparison between them has nothing to compare. What would change the answer is a reviewer whose finding rate falls far enough for recall to become a measurement; until then, more context is tuning against noise.

The chance baseline fires on both arms, which is what makes the comparison readable at all:

armA  4/22   chance baseline: ~3.0 would match by position alone
armB  5/22   chance baseline: ~2.6 would match by position alone
both  RECALL IS NOT CLEARLY ABOVE CHANCE AT THIS FINDING DENSITY

Read the estimator as the order-of-magnitude guide it says it is: it returns 3.0 where the exact permutation null returned 4.19, the ~30% under-shoot its own docs predict for merged neighbourhoods. No permutation was run for the graph arm, so its 5-against-~2.6 must not be read as separation — corrected for the same under-shoot its null sits near 3.6, and the caveat's own instruction is to confirm with a permutation before comparing two candidates.

(An earlier revision of this section recorded "the guard does not fire" as a defect. That was wrong, and wrong in the way this document keeps cataloguing: the scoring was run with a binary built 09:27, two and a half hours before the feature landed at 11:55 — strings finds no trace of it in that artifact. The claim was checked against the wrong build, not against the code. Recorded rather than quietly deleted, because a false gap in a document about false measurements is worth one paragraph.)


Stage 27 — v2.0 hardening & release → v2.0.0 · effort M ✅ hardening delivered; the release is held deliberately (#429)

The hardening audit ran; the release is held. Those are separate, and conflating them is how a stage looks unstarted when its work is done.

What the audit established, and it is the part worth keeping. All eight whole-graph surfaces are linear, measured to 72,509 nodes / 126,622 edges — 25× the graph Stage 26 measured, on a real repository rather than a synthetic one, worst absolute case duplicates at 0.89 s. Coverage went up to 89.77% while files measured grew 64 → 94. cargo deny --all-features clean. VENDORED_DEPENDENCIES.md's version rows exact, with zero advisories published since the vendored commit. That is a "fine at both, recorded so nobody re-measures it" result and it cost a 25×-scale benchmark to get.

What it found became issues, not paragraphs: #447 (recall --limit 0), #448 (Provenance closed?), #449 (advisory reachability), #450 (coverage), and the measured data on #431 — 66 exhaustive public enums across nine published crates, which is the real v2.0 semver decision.

The release is deferred on the owner's call, tracked at #429 with a completion condition rather than open-endedly, and the Release-plz workflow is disabled so it cannot cut by accident. The original deferral note follows, recorded so a stage nobody got to and a stage somebody decided to leave stay distinguishable.


Deferred deliberately, not merely unstarted. The owner's call, recorded so the two are distinguishable: a stage nobody has got to and a stage somebody decided to leave look identical six months later, and the second should not be picked up by whoever next has a free afternoon.

Nothing blocks it — Stages 21–25 and 28–32 are delivered and v1.15.0 is out. It is held because the hardening it describes is worth more once the work ahead of it has landed and been measured: the remaining A1 lenses (Stage 26) and Stages 33–35. And because v2.0.0 is a number worth spending once, deliberately.

Stage 34 in particular should land before this, not after: Stage 27 re-audits every "offline" claim, and adding a remote tier afterwards would reopen an audit that had just been closed.

One consequence to carry: the scope below grew during v1.10–v1.15. The offline claim it re-audits is now a real surface — docs/OFFLINE_SETUP.md, the security prefetch/status contract, digest-pinned assets, and per-file verification of the extracted boxlite runtime — so this is an audit against something concrete rather than a prose sweep.


7. Milestones → releases

ReleaseContainsGate
v1.10.0 ✅Stage 21 — analyzer contract + ingestArtifact byte-identical; ingest idempotent — met
v1.11.0 ✅Stage 22 — semgrep + cargo-audit (SAST axis, five languages)Offline warm-cache run; named cold-cache failure — met
v1.11.x ✅Stage 22b — osv-scanner (dependency axis: Python/Java/Node)Lockfile findings per ecosystem; Rust overlap resolved — met
v1.11.0 ✅Stage 23 — episodic memorySurvives rebuild; graph untouched — met (#317)
v1.13.0 ✅Stage 24 — boxlite backendParity with subprocess; cargo deny clean — met (#352): identical finding keys via both backends, differing only in isolation label and image digest
v1.12.0 ✅Stage 25 — recall + bounded cachedecay=none reproducible; no episodic eviction — met (#340). Shipped two releases ahead of its nominal target
v1.13.0 ✅Stage 26 — lenses Q3/Q1/S1All three, one PR each — Q3 (#346), Q1 (#372), S1. coupling --limit 0 over 2,887 nodes in 0.07 s; debt-density --limit 0 and config-secrets --limit 0 (372 keys) both 0.04 s, level with the debt baseline. No lens offers a CI gate, each saying why: Q3's cross-language call edges are name collisions (615/6,553 = 9.4%); Q1 inherits the marker scan's prose false positives and adds a file-length denominator; S1 is an inventory, and its one real finding indicts Roteiro's import layer, not the user's repo. EXTRACT_VERSION stayed 11 throughout. Cost: 11 files, 3,016 insertions — the corrected 195–500 LOC estimate is itself out by ~3–8×
v1.10.x ✅Stage 28 — generated media content moves out of derivedSilent clip cannot reach default search; media build restores searchability — met
v1.11.0 ✅Stage 29 — audio metadata as derived factsFormat read costs 1–100 µs and instantiates no decoder; duration exact/estimated/absent never guessed — met
v1.11.0 ✅Stage 30 — MTP speculative decodingOpt-in only; 1.22–1.50× on 27B — but output is not identical, so default-on is blocked on §9.6
v1.11.0 ✅Stage 31 — model lifecycle: resumable pulls, model rm, high tierInterrupted pull transfers only the remainder; checksum failure discards; pinned digest measured, not quoted
v1.12.0 ✅Stage 32 — guardrails: four confident wrong answers (#324, #321, #319, #330)Two ADRs on one id fail check naming both files; API and CLI debt agree, per repo; coverage measured (87.51% lines, 7/64 files under 85%) with no document claiming a gate that does not run; a new ADR on disk is never silently uncounted — met
v1.16.0 ✅Stage 33 — local model resolutionVision/audio/OCR pinnable per project; roteiro config answers why that model for every surface
v1.17.0 ✅Stage 34 — remote model tierADR-0019 Accepted. Cut in three, all delivered. 1 — the guard (#381): consent gate, payload allow-list, dry-run, egress ledger, in a build compiling no backend. 2a — the transport: ureq behind the off-by-default remote feature, the TTY invocation grant (status/dry-run never prompt), the response reader that refuses a truncated generation, and the README/website promise amendments on the commit that made them false. 2b — the surfaces: ModelSource::Remote { trust } with installed: None in the shared resolver (which forced ProducerTrust from rto-remote into rto-graph — one definition, or a ledger entry and a resolution could disagree about what "vendor-asserted" means), then spec draft --allow-remote and serve --allow-remote over it. Neither prompts — the flag is the only way on a surface whose default is local — and a refused --allow-remote stops the run rather than answering locally, which is the same silent downgrade a network failure would be. Both rebuild the request through the payload allow-list rather than forwarding a local prompt or a chat transcript, so the guard still assembles what leaves. Ask is wired by wrapping the served Engine in the binary, so rto-serve gains nothing. Project file may deny, never grant; no learned router on the local→remote edge; no reachability probe; no test can reach a network
v1.18.0 ✅Stage 35 — roteiro review LLM mode35a delivered (#380): roteiro review --score, per-class recall with denominators, and compile_claim's four axes — the instrument, with no verdict. 35b PR 1 delivered: ModelTask::Review on the shared generative key (a seventh task, not a fourth bespoke rule), the per-file reviewer, review --llm, and the review --replay harness. It overturned 35a's budget conclusion in the useful direction — per file the corpus is mean 3,275 / median 1,758 tokens as sent, and the largest reviewable source file-diff is 17,202, so ~28k of the single-call budget is free on a median file and the graph arm of PR 2 has room to be tested. It also found a vacuous measurement in the instrument itself: a reasoning model spent a 1,200-token cap entirely inside <think> on 4 files of 4 and the run reported "0 finding(s) over 4 file(s)" — zero recall that measured nothing, now detected, reported as not reviewed, and never scored. No per-class recall figure yet, deliberately: the honest remaining number is ~15 findings per file against 22 adjudicated rows over 184 files, unmoved by two precision-targeted prompt revisions
v2.0.0 ⏸️Stage 27 — hardeningFull gates; semver review complete — hardening delivered, findings filed as #431/#447–#450; the release itself is held on the owner's call (#429) and the Release-plz workflow is disabled so it cannot cut by accident

8. Risk register

RiskSeverityMitigation
A V2 record leaks into nodes/edges and breaks artifact purityHighNodeKind::Other("…") is mechanically possible — that is the trap. CI regression test asserting export_factset is byte-identical across ingest/memory writes.
Unreviewed memory acquires authored relevanceHighSeparate store, separate ranking channel; assert in tests that memory never scores through the authored path.
Memory captures secrets (tokens, stack traces, customer names)HighUncommitted .git/roteiro/ placement; explicit forget; documented that memory has no redaction chokepoint.
boxlite advisory lands and is missedMediumExact pin + deliberate advisory tracking as a standing duty (ADR-0014).
--all-features CI fails without /dev/kvmMediumRuntime capability probe; sandbox tests skip visibly.
Unbounded episodic growthMediumAccepted by design; explicit user reclamation only.
A single-vendor factual claim drives a designMediumThis plan already survived one: a "boxlite is unpublished, therefore unmergeable" blocker was refuted by direct crates.io checking. Verify checkable externals independently.
EXTRACT_VERSION bumped twice, forcing two full re-extractionsLowBatch all extraction-touching lenses behind one bump (Stage 26).
Speculative decoding silently changes a completionHighMeasured, not hypothetical (Stage 30). Off unless ROTEIRO_SPECULATIVE explicitly says on; unrecognised values are not consent. The risk is acceptance by default, so the mitigation is that there is no default.

8b. Beyond v2.0 — deferred work, and where it is written down

Stage 27 is the last scheduled stage. It is not the last work, and the difference has been invisible: everything below was decided somewhere in this document and then scattered across §9 and Stage 26's deferral list, where a reader planning the next quarter would not find it. This section is a map, not a new commitment — each item keeps its original reasoning at the reference given.

Already decided, deliberately not scheduled

ItemWhereWhy it is not in a stage
Semantic recall — vector index over memory§9.3Needs migration, model/dimension versioning, retention, rebuild and storage-size policy. Materially more than "persist embeddings".
Findings ↔ graph cross-surfacing§9.7Joining findings to graph facts is deliberately not free in this design. When wanted, it needs a designed join, not an implicit one.
code_interpreter§9.4, ADR-0014Rejected. The real question is is local code execution something Roteiro wants to be? — a product decision, not a backend one.
Q2 — LOC hotspotsStage 26Not a pure query: Node.span is byte offsets, so it needs net-new extraction metadata.
Q10 — dependency pinsStage 26Mis-scoped as written; existing pins are Docker image_ref and submodules, so package-manifest pins are extraction work.
Q7 — doc coverageStage 26Needs a language and a denominator; docs live mostly in symbol meta.content, not Doc nodes.
S2–S6 — the rest of the security lens seriesStage 26Taxonomy normalised (S1, S4 → GDS; S2, S3 → NNX; S5, S6 → EXT), but none is scoped.

The batching constraint that shapes all of it

Q2, Q10 and cross-language call-edge resolution each need extraction metadata, so each forces an EXTRACT_VERSION bump — and every bump is a full re-extraction for every user. The risk register already says to batch extraction-touching lenses behind a single bump. That makes these a cluster rather than three independent tickets: doing them one at a time is the expensive way to do the same work.

Cross-language call-edge resolution belongs in that cluster and is not yet recorded anywhere else. Stage 26's Q3 measured 615 of 6,553 call edges (9.4%) on this repository as cross-language name collisions — cross-file resolution binds a callee by simple name across every Fn node regardless of language, and no FFI is extracted. That is why Q3 offers no CI gate. Fixing it is extraction work, so it batches with Q2 and Q10 or it is paid for twice.

The one exception, and why it was made

The placeholder marker-needle correction (#384) spent a bump on its own (EXTRACT_VERSION 11 → 12), by the owner's decision. It is not a fourth member of the cluster above: it shipped alone, and Q2, Q10 and cross-language edge resolution remain a cluster with each other, still unpaid and still unscheduled — this exception does not release them.

The trade the batching rule is meant to prevent is paying twice for the same work. That is not what this was. The bare word placeholder was a stub needle scoring 0% precision — 36 of 36 findings on this repository, none a stub: the external-ref placeholder node (ADR-0009), the redaction placeholder (ADR-0015), S1's own sentence about not being able to tell a secret from a placeholder, a {tag} ref template, CSS ::placeholder. The lens was reporting the codebase's vocabulary as its debt, and roteiro check printed the inflated figure at a glance. Replacing the word with the two phrases that predicate incompleteness of an implementation — placeholder implementation, returns a placeholder — takes stub from 36 to 0, the true count, and reclassifies nothing else.

Holding that behind Q2, Q10 and cross-language edge resolution would have meant shipping a knowingly false count for the whole of an unscheduled, post-v2.0 horizon — indefinitely, since none of the three has a date. Stage 26's standard is that a lens which over-reports is worse than none; a 100%-noise category is the case that standard was written for. Correctness of a number users read now outweighed the cost of one re-extraction, so the batching rule was set aside deliberately rather than forgotten. It still governs the three items above: they each add new extraction metadata, they are genuinely one body of work, and nothing about this exception makes them cheaper to do separately.

What the bump cost, measured on a store extracted at version 11 and then opened by the version-12 binary: all 275 cached fact sets re-extracted (every tracked blob — the base version is unconditional, so nothing survives the key change), 3.2 s cold against 0.17 s warm on a debug build. The disk half of that price was open-ended when this was written: the object cache was write-and-keep with no eviction, so the superseded version-11 entries stayed on disk — .git/roteiro/objects went 4.8 MiB → 9.7 MiB and did not shrink again, per repository, per user, for every bump.

That half is now bounded (issue #387). sweep_superseded (crates/rto-graph/src/sync.rs) runs at the maintenance seam and deletes the generations a bump orphaned, keeping the current one and — by DEFAULT_KEEP_GENERATIONS — one behind it, so switching between a branch that bumped and the main it will merge into stays a cache hit. So a bump costs CPU once and disk once, not disk for ever, which is what the batching rule in this section already assumed it cost. Measured on a copy of this repository's own cache, where four generations had accumulated (9, 10, 11 and 12): 3,812 objects / 71.6 MB → 1,876 / 40.0 MB in one pass, and the live set verified intact by the only test that matters — a sync into an empty store afterwards reported 278 blobs, 0 extracted, 278 cached. Generation 12 alone is 594 objects / 10.0 MB, which is what keep_generations: 0 leaves; the 30 MB between the two is the retained generation 11, and is what the insurance costs. Two things are deliberately not reclaimed and stay a known cost: the environment tag (-e…) is a hash with no ordering, so no tag can be shown to supersede another, and the live set itself is still unbounded. .git/roteiro remains derived, so deleting the lot is still always safe.

Stages 33–35 — status

Formerly scoped-but-unrecorded, added to the roadmap by decision, and since largely delivered. Summarised here because §8b is where a reader looks for what outlives the current stage — the stages themselves carry the detail:


9. Open questions (decide before the stage that needs them)

  1. Cache bound value (Stage 25) — answered: 256 MB by default, configurable (decided by the owner). The unit was already settled as a byte budget following ModelCache; this fixes the number and makes it raisable for larger repositories.

    The scale that justifies it, measured on this repository: .git/roteiro is 49 MB (44 MB object cache over 2,395 entries, 4.6 MB graph.db) against a 91 MB .git. Roteiro's sidecar is already ~54% of the repository it describes, so a cache tier is not a new cost category — it is a bound on one that is currently unbounded in every direction. 256 MB is small against .git, trivial against an 18 GB model store, and large enough that an ordinary session never evicts.

    Erring small is deliberate and cheap: build_context is proven to reconstruct identically (context.rs asserts built == cached), so eviction costs cycles, never information. Erring large only costs disk. Neither error is expensive, which is precisely why this did not warrant more analysis than a measurement.

  2. Memory scope (Stage 23) — answered, ADR-0013 v1.1 §Scope. A lesson is valid in a tree only if the relevant association is present there in the same format, so the anchor is the scope test and scope is a coarse per-repo namespace, never a branch label. Shipped in #317; Stage 25's recall ranks on AnchorState::applies rather than inventing a second rule.

  3. Semantic recall (post-Stage 25): memory recall is lexical + anchor + decay in this plan. A vector index would need migration, model/dimension versioning, retention, rebuild and storage-size policy — materially more than "persist embeddings", and deferred deliberately.

  4. code_interpreter remains rejected (ADR-0014). The sharper question behind it — is local code execution something Roteiro wants to be? — is a product decision, not a backend one. If it ever becomes "yes", boxlite is the vehicle and Track A rides along; until then the answer stays "no".

  5. Is a faster decoder worth a different completion? Answered: yes (Stage 30, decided by the owner). Speculative decoding is measurably 1.22–1.50× on a 27B model and measurably does not reproduce plain decoding's text. Both halves are settled measurements; the judgement was whether Roteiro accepts the second to get the first, and it does.

    What that does not license is flipping the default. Generation was never a reproducible surface — sampling, quantisation and the served model all move the text already — so this changes how fast an already-variable answer arrives, not whether Roteiro keeps a promise it was making. The graph is where reproducibility is promised, and nothing in Stage 30 touches it. But the remaining honest reasons to keep ROTEIRO_SPECULATIVE opt-in stand on their own: the win is size-dependent (0.86–1.14× on a 9B — noise), it needs a draft head that most installed models do not ship, and the identity claim has never been observed to hold on any model. A default that helps one model class and silently changes output on the rest is a worse default than none. Revisit when a draft head is present on the common tier, not before.

  6. Findings ↔ graph cross-surfacing: joining findings to graph facts is deliberately not free in this design. When it is wanted, it needs a designed join, not an implicit one.