[!IMPORTANT] Archived on 2026-09-02. Kept for links and history; no longer current.
V2's stages were delivered and the work moved to tracked issues, which is what supersedes this document. Its predecessor is BUILD_PLAN.
Nothing here is maintained. In particular the Baseline (start of V2) table is a point-in-time snapshot — it still records "Released v1.15.0, all seven crates" against a project now past 5.x with nine — and is deliberately left as written. The
status: deprecatedabove is OKF §5.4's value, whose gloss is exactly this case: "kept for links and history; no longer current."
Status: Archived — delivered, then superseded by tracked issues · Owner: The Roteiro Project Team · Last-modified: 2026-08-19 Governing decisions: ADR-0001, ADR-0012, ADR-0013, ADR-0014, ADR-0018
This plan is finished, and it is a record rather than a roadmap
Every stage below is delivered. Work that was still outstanding when this was closed became tracked issues instead, and new work should be an issue from the start rather than a stage appended here.
Why it stops rather than continues. In a single day this document drifted from itself five times — stage headings marked delivered while the release table still showed the stage pending, and a summary section describing stages as blocked after they had shipped. Each time it was found by a person reading it, never by a gate. Its own standing rule says why that matters:
A plan that lags the tree is worse than no plan, because it is trusted.
The cause is structural rather than careless: the same facts live in three places here — the stage headings, the milestones table, and the dependency summary — and nothing holds them level. That is the shape this project has now removed six times in code (
[debt] ignoreacross three surfaces,limit == 0across five endpoints and then a sixth,[models]reporting versus resolving, the skill artifacts versus their template, the website versus the gate), and every one of those fixes worked by removing the possibility of divergence rather than the instances of it. An issue cannot drift the same way: it is open or it is closed, and it lives beside the work instead of describing it from a distance.What is deliberately kept. The measurements. Cost against estimate for the lenses (the corrected 195–500 LOC figure was itself out by 3–8×), the negative result on the review arm, the scale benchmark to 72,509 nodes, and the constraint each stage settled for the next. Those cost real work to establish and would have to be re-established if this were deleted.
This plan succeeds BUILD_PLAN.md, which took Roteiro from the v0.0.1 scaffold through Stage 20 to the released v1.9.0. V2 covers the next arc: Roteiro learns things that are not in the source tree — and stores them without compromising the promise that the graph is a pure function of the source tree.
Stage numbering continues from the v1 plan (which ended at Stage 20). As there,
stage numbers are labels, not execution order, and the v1.x/v2.0 headings
are nominal targets — release-plz cuts the real tags from conventional commits,
so a stage nominally marked v1.16.0 may actually ship in a v1.10.x. Each
delivered stage records the version it really shipped in.
Keep this document current. A stage is not finished when its code merges — it is finished when its entry here says what shipped, in which release, and what that settles for later stages. The stage entry is where the next person looks first; a plan that lags the tree is worse than no plan, because it is trusted. Update it in the same PR as the work wherever possible.
V1 built one thing well: a provenance-tagged graph, deterministically derived from git blobs, that humans and agents query through one surface. V2 adds three kinds of knowledge that do not fit that model and would corrupt it if forced in:
The organising rule for (1) and (2) is one sentence, and it is what makes V2 coherent rather than a list of features:
Knowledge that is not a derived/authored/inferred graph fact gets its own artifact store, and never borrows the graph's trust.
imports already works this way — it exists precisely because sync's rebuild
would destroy it, and is re-applied afterwards. V2 generalises that precedent
instead of inventing something new.
All seven principles of BUILD_PLAN.md §1 remain binding. V2 adds three invariants that constrain every stage below:
nodes/edges unless it is deterministically derived from (path, blob id, bytes). export_factset must remain byte-identical for a given tree.authored relevance boost, and none is exported in the GraphArtifact.Verified against main at the time of writing:
| Fact | Value | Consequence for V2 |
|---|---|---|
| Released | v1.15.0 on crates.io, all seven crates | V2 work is post-1.0 — semver is now real. |
| MSRV | rust-version = "1.94" | New deps must respect it. |
| Lints | unsafe_code = "forbid", clippy pedantic -D warnings | Native/FFI deps must be isolated behind a feature. |
| Coverage | measured in CI, not gated — cargo llvm-cov runs non-blocking; the 85% per-file floor is an aspiration (ADR-0001), never an enforced check (issue #319) | Every stage below still carries test cost, but a DoD may not cite "85% coverage" as if something verified it. |
| CI | Ubuntu-only; --all-features and the default set (the default-features job, added by #364 after the default set was found not to compile — issue #360) | /dev/kvm may be absent; Apple Silicon untested. Turning features on cannot find defects caused by code being cfg'd out. |
| Schema | migrations 1–13 applied (1–7 at V2's start) | V2 appends only; see §5. |
EXTRACT_VERSION | 11 (crates/rto-graph/src/extract.rs) — the Stage 28 bump landed in #316 | Bumping it forces full re-extraction for every user. No test pins the value. |
| Provenance | `Derived | Authored |
| Eviction idiom | in-memory byte-budget LRU (rto-llama ModelCache); nothing persisted is bounded | Stage 25 ports the existing policy to disk rather than inventing one — done, and tested against lru_evict_count's own numbers. |
| Crate | Change | Notes |
|---|---|---|
rto-exec | new | AnalyzerRunner trait + three backends (ADR-0014). Feature execution, subfeatures exec-boxlite, exec-subprocess. |
rto-graph | extended | Artifact-store tables + accessors; new query fns for lenses. Graph model untouched. |
roteiro (CLI) | extended | security {prefetch,status,run,ingest}, memory {add,list,recall,forget}, new lens subcommands. |
rto-serve | extended | New lenses surfaced to served-chat tools; memory recall exposed only behind explicit opt-in. |
rto-render | extended | Findings and lens renderers. |
Default install gains no new dependency. Everything in Stages 22/24 is feature-gated and off by default.
migrations.rs mandates append-only SQL: never edit a shipped migration. V2 adds
three tables across three migrations, deliberately not merged:
| Migration | Table | Lifetime | Evictable |
|---|---|---|---|
| 8 ✅ | analysis runs + findings (ADR-0012) | replaceable layer per (analyzer, worktree) | replaced wholesale, not aged out |
| 11 ✅ | agent_memory (ADR-0013 episodic) | durable, survives rebuild | never |
| 13 ✅ | agent_cache + agent_cache_clock (ADR-0013 transient) | bounded | yes, by capacity |
Numbers are assigned in landing order, not reserved in advance. Stage 21 landed
first and took 8; Stage 28 took 9 and 10; episodic memory took 11;
12 is sync_worktree, from the guardrails branch; and the cache tier took
13. Splitting memory across two migrations is intentional: different lifetimes
and guarantees, so the eviction tier can later be altered without touching durable
memory.
Check the other worktrees, not just the refs. Stage 25 was dispatched with
"12 is free" after git grep over every local and remote branch found nothing
claiming it — which was true of the refs and false of the tree, because the
guardrails work was sitting in an unpushed worktree. That is the third numbering
collision of the day and all three had the same cause, so the check that actually
works is:
git worktree list | awk '{print $1}' | xargs -I{} grep -ho 'version: [0-9]*' {}/crates/rto-graph/src/migrations.rs | sort -u
Two constants both declaring the same version merge cleanly in git and break
at runtime on whichever store applies them second, so this is not a conflict the
tooling will catch for you. migration_versions_are_unique_and_ascending now
fails the build if a merge ever produces one.
Stages 22 and 24 need no migration — RunnerKind shipped in migration 8 already
naming all three backends, with the schema CHECK accepting them. Stage 22 confirmed
this: two analyzers landed with no schema change at all, because FindingKey
takes each analyzer's own ordered identity components. Stage 22b confirmed it
again: osv-scanner landed with a five-component identity and no schema change,
and a test asserts exactly that.
EXTRACT_VERSION does not change in Stages 21–25. None of that work is
extraction output (Stage 21 shipped without touching it, as required). It does
change in Stage 26, once — see the note there.
Dependency shape — four tracks, only one hard chain:
Track A (findings): 21 ✅ ──► 22 ✅ ──► 22b ✅ ──► 24
Track B (memory): 23 ✅ ──────────────► 25
Track C (lenses): 26 (independent of A and B throughout)
Track D (media): 28 ✅ ──► 29
└──► 27 (v2.0 hardening)
Stages 21, 23 and 26 were the parallel-startable set; 21, 22, 22b, 23 and 28 have now landed, leaving 24, 25, 26 and 29 open. Nothing in Track C touches the artifact stores; nothing in Track B blocks Track A. 22b was sequenced after 22 and did not block 24: it added one adapter behind a seam 24 does not touch, and needed no migration.
Goal: land the whole value of the findings design with no analyzer and no sandbox — the seam, the schema, and a working ingest path. This is the stage that makes CI ingestion and local execution the same code path.
rto-exec crate; AnalyzerRunner trait (request: analyzer
id, read-only worktree, network: Deny, explicit consent → response: normalized
findings + evidence); IngestRunner as the first implementation; normalized
Finding + AnalysisRun types in rto-graph.finding:semgrep:<rule>:<path>:<start-byte>:<snippet-hash>,
finding:cargo-audit:<advisory>:<pkg>:<version>:<lockfile-blob>) and layer
replacement keyed security:<analyzer>:<worktree-id>.roteiro security ingest <normalized-json>, roteiro security list [--json].export_factset output is byte-identical before and
after ingest (regression test); no new nodes/edges rows exist; findings never
appear in search results ranked as authored.Delivered in v1.10.0 (PR #293). What shipped, and what it changes for later stages:
rto-exec with AnalyzerRunner + IngestRunner. The preflight (check_request)
deliberately sits outside the trait, so a later backend cannot quietly skip
the consent/worktree checks.analysis_runs + findings), appended; migrations 1–7 untouched.FindingKey = finding:<analyzer>:<analyzer's own ordered identity components>
with escaping. The two analyzers named above are therefore examples, not
schema — a new analyzer needs no migration.RunnerKind already names all three backends and the schema CHECK accepts
them, so Stages 22 and 24 need no further migration.execution ships as a default feature, reconciling this plan's "behind a
feature" with ADR-0014's "ingest is always available": a named seam for later
stages that is nonetheless present in a stock install. --no-default-features
builds with no analyzer surface.replace_findings_layer
deletes the previous run's finding rows explicitly (cascade is defence in depth),
and Store::orphan_finding_count() exists so tests assert zero orphans directly.nodes/edges counts and the exported artifact digest unchanged throughout.semgrep + cargo-audit → v1.11.0 · effort M + M ✅ deliveredGoal: two real analyzers behind the Stage 21 contract, via the subprocess runner, honestly labelled.
SubprocessRunner (feature exec-subprocess); per-analyzer
adapters normalising native output into Finding.roteiro security run <analyzer> [--allow-unsandboxed]. The flag is
required for subprocess execution; evidence records isolation=none.roteiro security prefetch / status land here — digest-pinned
advisory DB and rule sets, with assets-unavailable-offline as the cold-cache
failure (never an implicit fetch, never a silent host-tool fallback).advisory_db_published_at, fetched_at, age, and a possibly stale label.cargo deny clean.Delivered in v1.11.0 (PR #322). What shipped, and what it changes for later stages:
Coverage is a matrix, not an analyzer list (ADR-0018). The requirement is findings for Rust, Python, SQL, Java and Node. The two analyzers named in Stage 22 deliver that on the SAST axis only; Stage 22b closed the dependency axis:
| Language | SAST | Dependency vulnerabilities |
|---|---|---|
| Rust | semgrep (GA) | cargo-audit (RustSec) + osv-scanner — both kept, cross-referenced (22b) |
| Python | semgrep (GA) | osv-scanner (OSV PyPI) ✅ 22b |
| Java | semgrep (GA) | osv-scanner (OSV Maven) ✅ 22b |
| Node (JS/TS) | semgrep (GA) | osv-scanner (OSV npm) ✅ 22b |
| SQL | semgrep generic — token matching, no parser | n/a — no dependency ecosystem |
generic (token) engine: no AST, no dataflow, no types. It can say this
statement grants ALL PRIVILEGES; it cannot say this value reaches a query
unsanitised. The qualification is carried in the adapter's declared language
list (sql (generic mode)), in every SQL rule's engine-note metadata, and in
a test that asserts the note is present on every SQL finding.semgrep scan does against a lockfile, so it contributes nothing
to the second column.The adapter is the seam, so ingest and execution agree by construction. One
conversion (rto_exec::normalize_native) turns an analyzer's native output
into the normalized report Stage 21 already validates, and both paths call it: a
subprocess run hands it the analyzer's stdout, roteiro security ingest --analyzer <name> hands it a report file CI produced. Equality of Finding
values is a property of the code, not something a test establishes afterwards.
A new analyzer is a file in adapter/ and a row in ADAPTERS — no migration,
because FindingKey takes each analyzer's own ordered identity components.
Three tool behaviours found by running the tools, not by reading their docs. Each is a trap the next analyzer author would otherwise re-discover:
--config <file> it renames every
rule to <config.path.components>.<id>. The rule id is the first component of
a FindingKey, so this puts the local asset-cache directory into every stored
key — user-identifying data in a persisted record, and keys that differ
between two machines running the identical scan. --no-rewrite-rule-ids is
mandatory, and a test asserts no key contains a local path.extra.lines
and extra.fingerprint are the literal string "requires login" unless the
caller is authenticated to Semgrep's hosted platform. ADR-0012's identity
recipe ends in a snippet hash, so hashing that field makes the component a
constant and changes every stored key the day someone logs in. Snippets
are read from the worktree instead (rto_exec::snippet), which is a
function of the source rather than of the analyzer's auth state.cargo audit reports no advisory-database identity when you pin one.
Given --db <path> it returns last-commit: null and last-updated: null —
at a shallow clone and at its own managed checkout alike (0.22.2). Only the
unpinned, self-resolving configuration populates them, so the reproducible
offline configuration was the one losing its staleness evidence. Provisioning
records the publication date from the database's HEAD commit time and the
adapter falls back to it; the tool's own account still wins where it has one.Provisioning: writing and running are separate. prefetch is the only thing
that writes to the asset cache; a run never provisions, so "did this machine have
the pinned rules?" always has an answer. As shipped, prefetch verifies and
pins but fetches nothing — the rule set is vendored into the binary and the
RustSec advisory database is a git checkout with no digest-stable URL, so it is
refused with the exact clone command rather than obtained by shelling out to
git (which would be the host-tool fallback ADR-0014 forbids). ADR-0014 v1.1
records that clarification. An asset whose bytes change after provisioning is
refused, not warned about: a run would otherwise stamp a digest that does not
describe what it read.
Rules are ours, vendored and pinned. semgrep --config p/default resolves
against a network service, which makes an "offline" analyzer network-dependent
and its results irreproducible. Every shipped rule was written for this
repository under its own licence; no Semgrep Registry rule is vendored,
because those carry the Semgrep Rules License v1.0, which is not on
deny.toml's allow-list — and cargo deny governs crates, so it would never
have caught a rule file. The position is stated in the rule header, the adapter
docs and ADR-0018 rather than assumed.
Accepted gaps, recorded so they are decisions rather than oversights:
errors array is not converted to findings. A scan that failed
to parse some files reports fewer findings; only a fatal exit (≥2) is caught.related CVEs and categories
are preserved verbatim in meta; a number disagreeing with cargo audit's own
would be worse than carrying the vector unchanged.PATH splitting are written for it; nothing
has run on it.tests/, fixtures/ and similar, so a
user's repository is scanned with those excluded. That is semgrep's normal
behaviour, and it is why the live test copies the fixture tree somewhere
neutral first.No new dependencies — Cargo.lock is untouched; exec-subprocess is
std::process over the crates already present, and is off by default.
osv-scanner: the dependency axis for Python, Java and Node → effort M ✅ deliveredSplit out of Stage 22 because it is a different axis, not more of the same one, and because the SAST half is independently useful and independently reviewable.
prefetch/status and the finding schema are all
reused unchanged — which is the seam doing its job.https://osv-vulnerabilities.storage.googleapis.com/<ECOSYSTEM>/all.zip) are
single files at stable URLs, so they are the first asset that genuinely wants a
download-by-URL source. AssetSource is #[non_exhaustive] for that.osv-scanner also reads Cargo.lock and OSV ingests RustSec, so the same Rust
advisory arrives twice under two finding keys. ADR-0018
v1.1 resolves it: keep both findings and cross-reference them at the
reporting layer. Neither layer is filtered or made conditional on the other.
The join key needs no invention — OSV keys a RustSec-derived record by the
RUSTSEC id itself and carries aliases (RUSTSEC-2020-0071 →
CVE-2020-26235, GHSA-wcg3-cvx6-7396), and cargo-audit's adapter already
stores aliases verbatim in meta (crates/rto-exec/src/adapter/cargo_audit.rs:225).
Join on the RUSTSEC id, fall back to alias-set intersection. Render a duplicate
pair as one advisory confirmed by two analyzers, with both finding keys
still addressable — never a merged super-finding, and never a count that
silently halves.RUSTSEC-2024-0388,
RUSTSEC-2021-0139 and RUSTSEC-2026-0192 all resolve, each with
aliases: null — but whether osv-scanner the tool surfaces them by default
is unestablished, and ADR-0018 v1.0 conflated the two. Measure it and correct
the ADR. (2) RustSec→OSV ingestion lag: if it can trail by days the analyzers
will legitimately disagree for a window, and "present in one, absent in the
other" must render as a real state rather than a defect.cargo-audit is
explicitly resolved and recorded, rather than left to chance.Delivered. What shipped, and what it changes for later stages:
osv-scanner adapter, behind the Stage 21 seam. The subprocess runner,
asset cache, prefetch/status and finding schema were reused unchanged,
and — as Stage 21 predicted — it needed no migration: FindingKey already
takes each analyzer's own ordered identity components (advisory, ecosystem, package, version, manifest), and RunnerKind already names the backends. A
test asserts the new analyzer's keys are valid with no schema change.AssetSource::Download, the first asset fetched by URL — OSV's
per-ecosystem all.zip databases. rto-exec gained no network dependency:
it takes the fetcher as a function argument, so the code that can open a socket
lives in the CLI and is reachable only from prefetch. That flag is new and
required: roteiro security prefetch --analyzer osv-scanner --allow-download,
because the four databases are roughly 260 MB (npm alone is ~210 MB) and
that is not a reasonable surprise. The stale comment in assets.rs claiming an
unused fetch path is a security surface with no user was updated, not left to
contradict the code.rto_exec::cross_reference joins them at the reporting
layer on identifiers both upstreams publish. security list renders a
duplicate pair as one advisory confirmed by two analyzers with both finding
keys still addressable, and prints the finding total unchanged above it.
Real fixture data added a constraint the decision did not state: chrono's
cargo-audit advisory lists time's CVE under related, so the join also
requires the same package at the same version, and names a correspondence only
by an id an analyzer actually fired.osv-scanner 2.5.0 does report RustSec's unmaintained and
unsound by default — v1.0's options table conflated the database with the
tool. Only yanked is unavailable, and structurally: it is not an advisory,
cargo audit reads it from the crates.io index. (2) RustSec→OSV ingestion lag
is ~2.5 minutes, not days; what actually makes the two analyzers differ is
pin age, since each database is provisioned separately. "Present in one" is
rendered as a normal single-source row with its cause named.--offline-vulnerabilities
alone consults no database and reports a clean scan (--offline is what loads
it); reported paths are absolute even when the target is ., which would have
put the scanning machine's home directory into a persisted finding key; and the
same advisory is listed twice under its RUSTSEC and GHSA ids, with groups
already saying so.osv-scanner is on PATH.Goal: stop losing what sessions learn. Write path only — no retrieval ranking, no graph integration.
agent_memory accessors in rto-graph; anchor capture as
(anchor_key, anchor_blob, anchor_path); explicit superseded_by /
superseded_at. span is not an anchor — it is byte offsets and shifts on
any edit above it; blob_hash + node_key is the stable pair.INTEGER PRIMARY KEY AUTOINCREMENT supplies the monotonic
generation. created_at is written for humans and never read — matching how
imported_at already behaves. No wall-clock ranking, because the store is shared
across worktrees and branches and datetime('now') is second-granular..git/roteiro/ beside graph.db — per-clone, never
committed, never pushed. Privacy forces this: extraction redacts secret-looking
config values before persistence, and memory has no such chokepoint.roteiro memory add|list|forget.roteiro sync/rebuild (the imports property);
export_factset unchanged; nothing enters nodes/edges; supersession recorded
explicitly and superseded rows excluded from live listing.Delivered in #317. Migration 11 (agent_memory), the rto_graph::memory
store, and roteiro memory add|list|forget. Every DoD item above has a test;
memory cannot invalidate the fact cache, and is asserted so as a property of
memory writes (tests/sync.rs::memory_writes_do_not_invalidate_the_fact_cache),
because memory is not extraction output.
Four deviations from ADR-0013's proposed SQL, each deliberate:
kind is a closed CHECK … IN over the ADR's own five names
(lesson|attempt|decision|pattern|outcome), not the free TEXT proposed. Free
text makes lesson/Lesson/lessons three kinds, none findable by a filter,
and a vocabulary that cannot be filtered cannot later be ranked — which Stage 25
needs. Follows the analysis_runs.runner/isolation and media_content.kind
precedent. Cost: a sixth kind is an append-only migration, not a string.superseded_at is TEXT, not INTEGER. An integer here would hold the
generation of supersession — which is superseded_by, since the successor's
id is the generation — so it would duplicate the column beside it. As TEXT it
is a human timestamp on created_at's terms: written, displayed, never read.CHECKs making half-states unrepresentable, per migration 10's
precedent: superseded_by/superseded_at stand or fall together (a moment with
no successor is supersession inferred, the one thing the ADR rules out);
nothing supersedes itself; anchor evidence requires an anchor key; empty
scope/body refused.AUTOINCREMENT kept, and it is load-bearing — not decoration. A plain
INTEGER PRIMARY KEY is the rowid, and SQLite reuses the largest deleted one,
so forgetting the newest record would hand its number to the next write:
ORDER BY id DESC stops being newest-first and a surviving superseded_by
silently re-points at an unrelated record.Scope is settled, so Stage 25 inherits it rather than re-litigating it
(ADR-0013 v1.1 §Scope). The owner's rule: a lesson learned on a feature branch
is valid on main only if the relevant association is merged to main in the
same format — if not, then no. That needs no new machinery, because the anchor
is the scope test: a record applies to a tree when its anchor resolves there
with the same blob, or when it has no anchor at all (a general lesson, repo-wide).
Drifted, vanished or unverifiable ⇒ does not apply here, kept and marked.
"Same format" means the blob matches, strictly — a reformat breaks it, failing
toward marked rather than toward silently applying a lesson to code that moved.
Consequently scope is a coarse per-repo/project namespace and never a branch
label; no branch bookkeeping exists anywhere in the schema. Recall in Stage 25
should rank on this predicate (AnchorState::applies), not invent a second one.
Out of scope, still: the bounded cache tier, recall ranking, decay, and any
search integration — all Stage 25. Memory currently reaches search through no
channel at all, which is asserted rather than assumed.
Goal: the reproducible, offline-capable local run — one command, pinned inputs, digest-level evidence.
boxlite (Apache-2.0), pinned exactly, behind exec-boxlite.
Publication on crates.io was verified directly (17 versions, default 0.9.7, not
yanked), so this is a dependency addition, not a packaging problem.BoxliteRunner; digest-pinned OCI image; read-only worktree
mount, scrubbed environment, no ambient credentials, egress denied by default.--all-features must not fail on a runner without /dev/kvm — gate
sandbox tests on a runtime capability probe and skip with a visible message.
Apple Silicon microVM execution stays untested in CI, documented as an
accepted gap.cargo deny over the full resolved native/FFI closure.cargo deny clean on the
resolved tree.Delivered. BoxliteRunner behind exec-boxlite, boxlite pinned =0.9.7,
the DoD executed rather than argued: real semgrep 1.173.0 over one tree, once
as a host child process and once in a digest-pinned microVM, 4 identical
findings — PARITY OK … subprocess isolation=none image=none, boxlite isolation=microvm image=sha256:67319956…. 22 fault injections, 21 red on the
right message; the one that was not is recorded below, because it did not go
red and saying so is the point.
What the stage actually turned out to be about. The plan said publication was
verified "so this is a dependency addition, not a packaging problem". Publication
was real; the inference was not. boxlite from crates.io ships no hypervisor:
its three -sys crates each detect a published package (.cargo_vcs_info.json)
and disable themselves, and libkrun-sys excludes the sources they would build —
enabling its krun feature compiles and then fails to link with 26 undefined
symbols. What executes is a prebuilt runtime archive that boxlite's own build
script fetches with a bare curl -fsSL, include_bytes!s into the rlib, and
extracts and execs at run time. That fetch has no expected digest (searched
four ways, NOT FOUND) and an env-overridable URL, so two builds of the same
pinned version could embed different bytes undetectably.
So the stage's real work was governing that fetch, not adding a dependency:
AssetSource::PinnedArchive — the first asset source with a compile-time
digest, in runtime_pins.rs and shared with build.rs by include! so the
two cannot drift. It closes the gap Fetcher's contract has to leave open for
Download assets: a fetcher that reports success over a truncated body cannot
defeat a pin checked in-crate before install. Tested with a deliberately lying
fetcher.build.rs refuses to build unless BOXLITE_RUNTIME_URL names a local file
matching the pin. Unset, remote, or wrong bytes are hard failures with a
runnable recipe. boxlite's curl then never reaches the network, because
what it is asked to fetch is already on disk.deny.toml and discharged by
crates/rto-exec/NOTICE-boxlite-runtime.md, which is include_str!d and
printed by prefetch before installing. The archive embeds GPL-2.0 (mke2fs,
debugfs, libkrunfw) and LGPL-2.0-or-later (bwrap) binaries; they are
exec'd as separate processes, so this is aggregation and Roteiro stays
MIT OR Apache-2.0, but distributing a binary built this way carries GPL-2.0
§3 source-offer and LGPL relinking duties.The gate that could not see any of this is now closed.
crates/rto-exec/tests/build_script_fetch_audit.rs reads the build script of
every package in the --all-features graph and fails on anything that looks
like it fetches without a recorded pin. Measured: 613 packages, 89 of which have
a build script (96 script files), 2 flagged (both governed), 0 false
positives, in 0.3s — the audit prints those numbers on every run, so "it
flagged nothing" stays checkable rather than trusted. Review asked whether the
matcher was narrower than its docs claimed; it was, and the docs were the wrong
half: widening to a bare http(s):// was measured at 29 flagged with 27 false
positives (serde, quote, anyhow, winapi … all citing a docs URL in a
comment), so the claim was narrowed to what the code does and a test now pins the
two together. Its module docs state what it does not cover (helper crates,
include!d build modules, obfuscation, run-time fetches, already-vendored
bytes); read them before trusting it. This hole was general, not boxlite's:
cargo deny --all-features check reported licenses ok while 25 MB of GPL
binaries were being embedded, and was not wrong to — it governs crates.
Deviations and costs, stated rather than buried:
rusqlite moved =0.40.2 → 0.39. Not a preference: boxlite
requires rusqlite ^0.39, libsqlite3-sys declares links = "sqlite3", and
cargo forbids two versions of a links crate in one graph. It is a hard
resolution failure otherwise. It also restores what the pin's own comment
already described. Coordinate with #342 (store.rs/migrations.rs); the
workspace compiles unchanged on 0.39, and MSRV 1.94 still builds --all-features.protoc >= 3.12 is now a build requirement for exec-boxlite
(boxlite-shared's build script, no stub path around it). Documented in the
README before anyone meets it as a build failure; added to all three CI jobs.scripts/provision-sandbox-runtime.py, which
reads the digests out of runtime_pins.rs so it cannot drift). All three
--all-features jobs now also compile ~400 extra crates — a real cost on every
PR, in keeping with the llama.cpp build they already carry.unmaintained advisory ignores (term_size, bincode, adler,
atty), all arriving through boxlite, none a vulnerability, each with the
four-part rationale deny.toml demands. They are global — cargo-deny cannot
scope an ignore to a feature — and say so.roteiro/exec-boxlite implies exec-subprocess, because security prefetch|status|run are all gated on that feature today. CLI plumbing, not
policy: unsandboxed runs still need --allow-unsandboxed per invocation.semgrep has a pinned image. cargo-audit has no official one, and
publishing a security tool's container is not a job this project is taking on.SandboxError::Killed exists so
nobody diagnoses that twice.semgrep reports relative paths; fault injection
confirmed setting it to None leaves the parity test green. It is there for
osv-scanner, the next image candidate.What this settles for later stages. Analyzer execution now has one
provisioning contract with a real digest pin, so a future backend or analyzer
image inherits verification rather than re-inventing it; and the build-script
audit means the next dependency that fetches something is a failing test rather
than a discovery. Apple Silicon microVM execution remains untested in CI — the
parity proof above ran on Apple Silicon locally, and CI runners have no
/dev/kvm, so the sandbox tests skip there with a visible message and the ingest
and subprocess paths carry the functional coverage. That gap is unchanged and
still accepted.
Goal: make memory useful — recall that ranks by evidence, plus the bounded cache that stops sessions re-deriving what they already know.
build_context is proven to reconstruct identically
(context.rs asserts built == cached), which is what makes cache eviction cost
cycles rather than information.agent_cache with
bytes, generation, last_used, hits. No persisted access tracking
exists today, so the signal must be introduced with the table (the in-memory
ModelCache tracks recency by list order, which does not survive a process).AnchorState::applies (ADR-0013 v1.1 §Scope) is the whole rule, and it is what
anchor_penalty below should be built on. Do not add a branch or scope term
to recall: scope is a namespace, the anchor is the validity test, and a second
rule would give two answers to one question.ModelCache
(crates/rto-llama/src/llama.rs:120-137) rather than a new row-count cap —
evict oldest-first on (anchor_valid ASC, last_used ASC) until the tier fits,
always keeping at least the most-recently-used entry. Swept at the existing
maintenance seam where refresh_contexts is already called — not on the read
path, so reads never mutate. Never evict: anything episodic, or a
valid-anchored row written in the current generation.score = base_confidence × anchor_penalty × decay(current_generation − row.generation)
with decay ∈ {linear, exponential, none} and none guaranteeing reproducible
recall. A stored decaying score would rewrite the store on every read.search at all it needs a visually distinct
channel and its own score. It never takes the authored +40 boost.decay=none gives byte-identical recall for a fixed repo state across
runs; eviction never removes an episodic row; a superseded memory drops out of
recall immediately regardless of age; an unanchored memory is still retrievable
and clearly labelled.Delivered. Migration 13 (agent_cache + agent_cache_clock),
retrieval-time ranking in rto_graph::memory, the byte-budget sweep, and the
search memory channel. All four DoD items have a test, and each was
fault-injected: the guarded behaviour was broken, the guard watched go red,
and the source reverted byte-identically (15 injections, all red — two of them
only after the tests they exposed as weak were strengthened).
What shipped, and where it deviated:
Decay is none | linear[:span] | exponential[:half-life], and none is
the default. ADR-0013 offered the three modes without saying which one leads;
the reproducible answer is the default here, on the same terms as
SearchOptions defaulting generated content off. Age is counted in
generations — one per record written — so ranking never touches a clock.base_confidence defaults to 0.5 when a writer states none. Not in the
ADR. 1.0 would let every record that claimed nothing outrank one that
honestly claimed 0.9, pricing honesty; 0.0 would make the common case (the
CLI states no confidence unless asked) unrecallable. The midpoint is the only
value that makes stating one worth the trouble in both directions.anchor_penalty ranks drifted below vanished (valid 1.0,
unanchored 0.9, unverifiable 0.5, vanished 0.35, drifted 0.25). Drift
is the one state that can actively mislead about code still sitting under the
same key; a vanished anchor can mislead nobody, and ranking it lowest would
punish exactly the records the ADR says are worth keeping most. Two properties
are asserted rather than assumed: nothing is ever zero, and every state that
AnchorState::applies outranks every state that does not.agent_cache_clock beside the table (on sync_state's precedent).
ADR-0013 §3 rules out wall-clock and the ADR's proposed agent_cache names
generation/last_used without saying where the values come from. ticks
advances per access — ModelCache's list position, made durable, with no ties
for the sweep to break arbitrarily — and generation advances once per sweep,
which is what makes "written in the current generation" a window rather than a
single row, and what makes the pin lapse instead of becoming permanent.evict_count is a pure function tested against lru_evict_count's own
numbers. Porting an existing policy is worth nothing if the port quietly
behaves differently. The always-keep-the-MRU rule moved from a
len - evict > 1 guard into the caller's pinned set, because this tier pins
other rows too and one rule beats two. A sweep can therefore finish still
over budget when everything left is pinned; CacheSweep::over_budget says so
rather than leaving a bound that silently failed to bind.search through a third channel, --include-memory, off by
default — beside the graph and generated channels, never merged with either.
Its scorer shares no branch with the node scorer, so "memory never takes the
authored +40" is structural. Stage 23's total absence assertion is restated
rather than relaxed, and is now checked more sharply than absence could check
it: a memory hit's score must not exceed the ceiling its own lexical terms can
produce, which is what a leaked +40 would breach.roteiro memory recall [query] [--decay …] [--applicable-only] …,
which prints every term of the ranking and not just the product, and
roteiro memory cache [--sweep] [--budget-mb]. The sweep also runs at the
maintenance seam in roteiro context --refresh, never on a read path.
context --refresh --json keeps its long-standing shape: wrapping it would
break callers to pay for maintenance they did not ask about, so the sweep is
reported on stderr there and has its own --json under memory cache.--budget-mb, else ROTEIRO_CACHE_BUDGET_MB, else the 256 MB of
§9.1. An unreadable value is an error, not a fallback — running the default
under a name that says otherwise is how an operator ends up believing in a
bound that was never applied. The config-file layer was deliberately not
touched.Two absolute assertions on a shared constant disarmed, at store.rs:1709 and
:1880 — assert_eq!(store.schema_version()?, 11), which migration 13 breaks.
Both now read migrations::latest_version(), the idiom
a_later_migration_is_additive_on_a_populated_store already uses. The literal was
defended in a comment as making someone confirm a new migration is meant to apply
on open, but apply runs every migration newer than the recorded version, so
there was never a per-migration choice there to confirm — the literal asserted the
value of a shared constant and nothing else. Fault injection confirmed the
rewritten assertions still catch the thing they are for: an apply that stops one
migration short of latest_version() fails both.
EXTRACT_VERSION is unchanged, and
tests/sync.rs::memory_writes_do_not_invalidate_the_fact_cache still passes —
also fault-injected, by making a memory write clear sync_env.
Not in this stage, deliberately: the cache tier ships with its policy, its
seam and its API, but no producer — node_context is still the context cache.
Moving a live cache onto the bounded tier is a data migration, not a policy
change, and bundling it would have put a schema move and an eviction policy in one
reviewable unit.
Goal: deepen the graph itself — the on-brand work — with honest costs.
Cost correction, which this stage exists to respect: a fully surfaced lens is
~195–500 LOC across 6–8 files, not the ~20-line mirror previously assumed. That
figure describes only the internal query fn. There are seven surfacing stages,
not four: extraction (scan_markers + augment), the query fn, the query result
types, MCP (GraphServer) and served-chat (GraphToolRegistry) as
separate registries, Obsidian render, and CLI-side aggregation — plus tests and
docs.
Shortlist, in order:
All three shipped, as three separate PRs — a reviewable diff beats a complete one nobody can check.
Calls edges already retain
direction, and today's hotspot view throws that away by incrementing both
ends. Highest value per line in the set.Explicitly deferred out of this stage, with reasons:
Node.span is byte offsets, so it
needs net-new extraction metadata.image_ref
and submodules; package-manifest pins are extraction work. Split S / M.meta.content, not Doc nodes.EXTRACT_VERSION bump: required once, if and only if a lens adds derived
extraction metadata (Q2 and Q10 do; Q1/Q3/S1 as scoped do not). Bumping invalidates
every cached blob for every user and forces full re-extraction — so batch all
extraction-touching lenses behind a single bump rather than paying it twice.
Also in scope (documentation debt): normalise the security taxonomy — the prose
defines GDS/NNX/EXT/LLM while rows S1–S6 use undefined GPB/CVE/SAST labels — and
mark the "SmolVLM is too small to emit <tool_call>" claim as a hypothesis, as
it currently rests on no code or benchmark evidence.
roteiro check green; surfaced on all
applicable surfaces or explicitly documented as CLI-only; scale-benchmarked on
this repo (whole-graph lenses matter — search already scans all nodes); false
positives have a suppression story, a confidence signal and a baseline before any
CI-gating is offered.rto_graph::coupling reports, per node, fan-in (distinct callers) and
fan-out (distinct callees) over Calls edges, plus Martin's instability
fan_out / (fan_in + fan_out). Ranked by total | fan_in | fan_out, ties
broken by key, so identical input gives byte-identical output.
Surfaced on six of the seven stages; the seventh is a deliberate no-op.
| Stage | Q3 |
|---|---|
Extraction (scan_markers + augment) | Untouched by design — Q3 adds no derived metadata, so EXTRACT_VERSION stays 11. Confirmed live: sync after the change reported 237 of 239 blobs cached. |
| Query fn | rto_graph::coupling |
| Query result types | CouplingReport / CouplingItem / CouplingOrder |
MCP (GraphServer) | coupling tool (rto-render/src/mcp.rs) |
Served-chat (GraphToolRegistry) | coupling tool — a separate registry, with its own test |
| Obsidian render | _Home → "Most depended-on (call fan-in)" |
| CLI-side aggregation | roteiro coupling [--order] [--limit] [--json], GET /v1/graph/{project}/coupling |
Not surfaced: the explorer web app has no coupling panel. The /coupling
endpoint serves the data; wiring assets/app.js is follow-up, tracked here rather
than left to be discovered.
/hotspots is deliberately unchanged. Undirected degree over every edge kind
is a different, still-useful question; the explorer depends on its shape. Its doc
comment now says so and points at /coupling, rather than the discarded direction
being an unremarked accident.
Two counting rules that change the numbers, both reported rather than silent:
(src, dst, kind, provenance), which still admits parallel Calls edges at
two provenances. fan_in counts dependants, not layers that asserted a
dependency.sync.rs) binds a callee by simple name across every Fn node
regardless of language, and Roteiro extracts no FFI. On this repo that is
615 of 6553 call edges (9.4%) — enough that excluding them removes a
JavaScript clone from second place in the ranking.Confidence signal, and why there is no CI gate. fan_in is exactly as precise
as the edges beneath it, and the residual same-language case is not fixable here: a
lone Rust join helper still absorbs every .join(…) in the workspace, reading as
232 callers. That is a limit of call resolution — fixing it is extraction work,
and extraction work means bumping EXTRACT_VERSION. So the lens offers no CI
gate, says so, and carries the caveat on every surface including both tool
descriptions, so a model reporting a high fan_in passes it on. No suppression
mechanism was added: coupling is a measurement, not a finding, so there is
nothing to suppress — and a second exclusion vocabulary beside [debt] ignore
would not have helped anyway (this repo's [debt] ignore does not cover the
vendored cytoscape.min.js that dominates the ranking).
Scale benchmark, this repo (5501 nodes, 12036 edges, 6553 calls edges;
2887 coupled nodes), release build, warm, whole-process wall clock, best of 5:
| Command | Time |
|---|---|
roteiro coupling --limit 20 | 0.05 s |
roteiro coupling --limit 0 (all 2887) | 0.07 s |
roteiro debt (existing baseline) | 0.04 s |
roteiro search store (scans all nodes) | 0.05 s |
Ranking runs on the counts alone and only the nodes that survive the cap are read
back, so a top-N question costs one edge scan plus N node lookups — not the
whole-graph node scan /hotspots performs.
Documentation debt (both items), fixed in #288, which is where they live — not in any repo file:
GPB/CVE/SAST, which
the issue's own class key never defines. They now use the defined
GDS/NNX/EXT/LLM vocabulary, assigned to match the issue's own
A1–A4 tiering (S1, S4 → GDS; S2, S3 → NNX; S5, S6 → EXT), and S1 is
retitled to the inventory it can actually be.<tool_call>" rests on no code and no benchmark — NOT FOUND across four
queries: grep -rni smolvlm --include='*.rs' crates/ (10 hits, all registry /
media-producer / speculative-decoding plumbing), grep -rni tool_call filtered
to vision/VLM/mmproj/image terms (zero), and roteiro search for
"smolvlm tool_call" and "vision tool protocol" (no matches). Stage 31's DoD
required Qwen3 to be shown to emit a <tool_call>; no equivalent run exists
for SmolVLM. The describe-then-query recommendation built on it is relabelled as
the safe default, with the one-run experiment that would settle it.The A1 cost line in #288 ("~20-line mirror of debt") is struck through and
replaced with the ~195–500 LOC / 6–8 file figure this stage is built on.
rto_graph::debt_density ranks files by retained intent-debt markers per 1,000
lines. A raw marker count ranks by file size — the biggest file has the most lines
to put a marker on — so a 40-marker file of 4,000 lines and a 40-marker file of 200
lines are indistinguishable under debt and twenty-fold apart under this. Built on
debt's own output rather than re-walking markers, so the two lenses cannot
disagree about which markers exist, and [debt] ignore applies unchanged.
The denominator was the design decision, and it did not need extraction
metadata. lines is the file node's meta.lines — the count of \n bytes in
the blob, recorded at extraction since Stage 2 and therefore already in every
EXTRACT_VERSION 11 blob. Three alternatives were rejected: SLOC does not
exist in the graph and computing it is net-new derived metadata (which would have
moved Q1 into the Q2/Q10 batch below); per-symbol is what
Node.span's byte offsets cannot give, since a span is not a line range; and
the highest marker line is a lower bound on length, not the length. What
lines counts is stated on every surface: every line, blanks and comments
included — file length, not lines of code.
Two rules keep the arithmetic honest. A min_lines floor (default 50) keeps
the denominator's tail out of the ranking — one marker in a 3-line file is 333 per
kloc, true and useless — while leaving those files in files_with_markers and
total_markers, reported as short_files. And ranking cross-multiplies the exact
ratio in u128 rather than the rounded per_kloc, so densities differing in the
fourth decimal do not tie and silently reorder on path.
Surfaced on six of the seven stages; the seventh is a deliberate no-op.
| Stage | Q1 |
|---|---|
Extraction (scan_markers + augment) | Untouched by design — the denominator was already extracted, so EXTRACT_VERSION stays 11. Confirmed live: a rebuild from an empty store reported 257 of 258 blobs cached, 0 extracted. |
| Query fn | rto_graph::debt_density |
| Query result types | DebtDensityReport / DensityItem / DensityOrder / DEFAULT_MIN_LINES |
MCP (GraphServer) | debt_density tool (rto-render/src/mcp.rs) |
Served-chat (GraphToolRegistry) | debt_density tool — a separate registry, with its own test |
| Obsidian render | _Home → "Densest files (markers per 1,000 lines)", under the existing intent-debt section |
| CLI-side aggregation | roteiro debt-density [--kind] [--order] [--limit] [--min-lines] [--json], GET /v1/graph/{project}/debt/density |
Not surfaced: the explorer web app has no density panel — the same position as
Q3's coupling panel. roteiro debt is deliberately unchanged: a marker inventory is
a different question, and check depends on its shape.
No CI gate, and the caveat travels on every surface (both tool descriptions, the
_Home table, the CLI summary, the endpoint doc) — Q3's precedent, not one buried
mention. Density inherits the marker scan's prose false positives (for now,
placeholder, tbd fire on ordinary writing, so a design document ranks as dense
debt) and adds one of its own: the denominator is file length, so verbose,
generated or widely-indented files are systematically flattered and dense
languages penalised. A gate would fail builds on prose and on formatting. The
suppression story is the existing one — [debt] ignore globs and the
roteiro:ignore / roteiro:ignore-file directives, applied before anything is
counted.
One follow-up, found in review and fixed on the branch (fix(render): _Home scopes intent debt by [debt] ignore, both tables): the Obsidian _Home overview
computed debt with an empty ignore list, so the vault reported a different debt
for the same repository than debt, debt-density, check and the graph API —
the disagreement ADR-0007 v1.1 was amended to end. Both _Home calls were fixed,
not only the one Q1 added: leaving debt's wrong would have had the category
totals and the density table on one page disagreeing about which files exist,
which is worse than being consistently wrong. This was the third surface with
that defect (#321 records the first two) and it was missed for the same reason each
time — the earlier fix went to the surfaces that had been reported rather than to
this stage's own list of seven, which names the Obsidian render. The enumeration
that closes it: debt and debt_density are the only functions in the workspace
taking an ignore: &[String], and every call site of either now passes the target
project's list wherever one can be reached.
rto_graph::config_secrets reports secret-named config keys — paths, key
names, and whether each value was redacted before persistence — in three states,
because collapsing any two would misreport them:
| State | Meaning |
|---|---|
redacted | The value was read from a config file and replaced with the placeholder before anything was stored. The expected state. |
declared | The key carries no value at all — a @rto:config struct field, declared in Rust with no literal to redact. Neither a redaction nor a leak. |
present | A value that is not the placeholder. Extraction cannot produce this; Store::apply_import_layer can, so a non-zero count is a finding about this store, pointing at the importing tool rather than the source repository. |
redacted_not_secret_named counts values redacted for where they live (a
Kubernetes Secret's data) rather than what they are called, so the redaction
figures reconcile against the graph instead of leaving an unexplained surplus.
config_keys::REDACTED is now a single constant shared by the two redaction sites
and this one reader, so the lens cannot drift from the redactor by a spelling.
[debt] ignore deliberately does not apply here, which is worth stating given
the follow-up above. That list is defined as a debt-marker exclusion (ADR-0007), and
config_secrets takes no ignore parameter at all — so there is no call site that
could be passing the wrong thing. Giving it one would mean inventing a second
exclusion vocabulary for a different question, which is the mistake Q3 declined to
make for coupling.
The rename is the whole point, and it is enforced rather than merely documented.
The lens cannot detect a hardcoded credential in source (that produces no
config_key node), cannot judge validity (it never sees a value — no
ConfigSecretItem has a value field), and cannot tell a real secret from a
placeholder. An empty report means no secret-named config key — a statement
about naming, not a clean bill of health: a credential under dsn or endpoint
never appears. That limitation is carried in both tool descriptions in the
imperative ("state the limits when you report it"; "if asked to scan for secrets,
say plainly that this tool cannot do it"), in the CLI summary unconditionally
including on an empty report, in the _Home section, and in the endpoint doc — and
each of those is fault-injected. tests/config_secrets_cli.rs puts the same
token in a .env (where it becomes a redacted config key) and in a Rust function
body (where it becomes nothing this lens reads), so the boundary is a test.
Surfaced on six of the seven stages; the seventh is a deliberate no-op.
| Stage | S1 |
|---|---|
Extraction (scan_markers + augment) | Untouched by design — the redaction already happens there, so EXTRACT_VERSION stays 11. Confirmed live: 259 of 260 blobs cached, 0 extracted. |
| Query fn | rto_graph::config_secrets |
| Query result types | ConfigSecretReport / ConfigSecretItem / RedactionState |
MCP (GraphServer) | config_secrets tool (rto-render/src/mcp.rs) |
Served-chat (GraphToolRegistry) | config_secrets tool — a separate registry, with its own test |
| Obsidian render | _Home → "Config keys named like secrets" (counts and files, not key names — a vault note is browsed out of context) |
| CLI-side aggregation | roteiro config-secrets [--limit] [--json], GET /v1/graph/{project}/config-secrets |
Not surfaced: the explorer web app has no config-secret panel. No ordering
knob, deliberately: this is an inventory ordered by (path, name, key), and an
order would imply some keys are more secret than others. No CI gate and no
non-zero exit, even on an unredacted finding — that finding is about Roteiro's
own import layer, so failing a user's build over it would be the wrong response.
Scale benchmark, this repo (5,860 nodes, 13,048 edges; 372 config keys, 48 markers over 20 files, 32,444 ranked lines), release build, warm, whole-process wall clock, best of 5:
| Command | Time |
|---|---|
roteiro debt-density --limit 20 | 0.04 s |
roteiro debt-density --limit 0 (all 20 ranked) | 0.04 s |
roteiro debt-density --limit 0 --min-lines 0 (floor off) | 0.04 s |
roteiro config-secrets (default 50) | 0.04 s |
roteiro config-secrets --limit 0 (all 372 keys scanned) | 0.04 s |
roteiro debt (existing baseline) | 0.04 s |
roteiro coupling --limit 0 (Q3 reference) | 0.08 s |
roteiro search store (scans all nodes) | 0.05 s |
Both new lenses are indistinguishable from the debt baseline. Q1 reads back only
the files that actually carry a marker — one node lookup each, not a whole-graph
file-node scan, which matters because file nodes carry captured meta.content.
On this repository S1 finds 2 secret-named keys among 372 config keys, both redacted, none unredacted — which is also a fair illustration of its reach.
The cost estimate, corrected a second time. Actuals per lens, and the figure each was measured against:
| Lens | Files | Insertions |
|---|---|---|
| Q3 (#346) | 8 | 1,121 |
| Q1 (#372) | 8 | 1,509 |
| S1 | 10 | 1,523 |
| Stage total | 11 | 3,016 |
The RFC's original "~20-line mirror of debt" was out by ~75×. This stage's own
corrected 195–500 LOC / 6–8 files is out by ~3–8× on lines and right on files:
a fully surfaced lens is ~1,100–1,500 insertions across 8–10 files, and that
is the figure a future lens should be planned against. The line count is dominated
by tests and by doc comments defending the design decisions — Q1's denominator,
S1's three states and its limits — not by the query, which is under 120 lines in
both cases. S1 costs the same as Q1 despite a simpler query because the rename
obliged the limitation to be restated, and fault-injected, on every surface.
41 fault injections across the two lenses (20 for Q1, 21 for S1), one per new
behaviour on every surface, each caught by a named test with every file
byte-identical after revert. Two were retained after they exposed real test gaps
rather than being retargeted: Q1's ranking on the rounded ratio, and S1's CLI
warning path, which extraction cannot reach and so needed a unit test on
config_secrets_summary rather than a repository fixture.
EXTRACT_VERSION is still 11 and no lens in this stage touched extraction
output. The batched bump (§8b: Q2, Q10, cross-language edge resolution) remains
unpaid and unblocked.
derived (ADR-0015) → v1.10.x ✅ delivered (independent track)Goal: stop generative model output masquerading as deterministic extraction — without losing the ability to search it. Resolves #300.
ocrs-text) and PDF text stay
derived: they decode content that exists in the bytes, and their errors are
misreadings correctable against the source. ASR transcripts and VLM descriptions
move out: they invent fluent text when there is nothing to read.media_content store keyed by source blob id + producer identity
(model id + digest, quantisation, mmproj digest, prompt, sampling parameters).
Re-describing with a better model is a new record, not a mutation. Records
survive rebuild, following the imports precedent — they are expensive to
reproduce (a 715 MB projector load per blob, see #301) and not derivable from
source alone.roteiro media build [--audio] [--vision] [--force] (incremental — only blobs lacking a record for the current
producer), media status [--json], media clear [--producer <id>].roteiro search --include-generated, off by default; when on,
every hit is visibly marked as generated, ranked in its own channel, and never
given the authored boost. The explorer UI surfaces generated content on a media
node with its producer and a per-blob rebuild action.media_content record states the
reason and the measured value, so media status distinguishes not generated
from generated nothing. Conservative, configurable thresholds; --force
overrides. It raises the floor — quiet speech and subtly-textured images still
confabulate — so it complements the store rather than substituting for it.EXTRACT_VERSION bumps here — extraction output genuinely changes. This is the
one bump referenced in §5; batch it with Stage 26's extraction-touching lenses if
they land together, so users re-extract once rather than twice.nodes.meta.content, nothing is copied into the new store, and records are
produced on demand by media build. No shim, no dual-read, no deprecation window
— which is only true because it is being done now.search results; a silent
clip is refused before the model loads and the refusal is visible in media status with its measured value; generated text is attributable to a named
producer everywhere it surfaces; media build restores full searchability in one
command; export_factset is byte-identical across a media build; dropping a
producer's records leaves the graph untouched.Delivered across two PRs, both merged:
media_content store (migration 9), keyed by source blob +
producer identity; generated text stopped being written into nodes.meta.content;
EXTRACT_VERSION 9→10; media build|status|clear; search --include-generated
(off by default, always labelled, never the authored boost). A later fix
corrected the generation counter to read MAX(generation) before a --force
delete — a count is wrong under deletion, a max is not.media status distinguishes not
generated from generated nothing. Plus the explorer surfacing (attribution +
a copyable per-blob rebuild command; deliberately not a mutating endpoint, since
the explorer is llama-free per ADR-0010) and the deferred media CLI arg-shape
tests.Closes #300. Measured on a real repo: media build --audio refuses a silent
clip in 0.013 s with no model load, versus 12.9 s and ~2 KB of confabulated
prose under --force.
Known limit, recorded rather than papered over: MP3 and FLAC are not gated. They are entropy-coded, so measuring amplitude means decoding — behind the very load the gate avoids. The gate abstains, and abstention is a pass, so those formats still reach the model. Whether to close that gap with structural parsing (no dependency) is tracked separately.
derived facts (ADR-0016) → v1.11.0 · effort M ✅ delivered (independent track)Goal: the complement of Stage 28. That stage took generated content out of
derived because it is invented; this one puts extracted content in because it is
present in the bytes — codec, sample rate, bit depth, channels, duration, frame
count and tags, from a format read with no decoding and no model (measured
1–100 µs on this repo's own fixtures).
symphonia, default-features = false, codec/container features
plus the id3v1/id3v2/ape metadata readers — which are separate feature flags
and are not implied by flac/mp3/wav; without them every MP3 tag is
invisible. Adds MPL-2.0 to deny.toml (file-level copyleft; does not reach
Roteiro's own source), recorded with its rationale.derived fact is new here, so duration carries an
exact | estimated marker and absence is recorded as absence, never a guess.is_audio (symphonia does not support Opus at all), and
cross-container duplicate detection (duplicates matches on git blob hash, so it
could never pair).EXTRACT_VERSION bump;
export_factset unchanged in shape; tests need no model, so they run on CI
rather than self-skipping.Goal: spend the draft head a Qwen3.5+ GGUF already carries. Generation on Metal
is memory-bandwidth bound, so a decode that confirms three proposed tokens costs
about what a decode producing one costs. The drafter is the model's own nextn
block — no second model to install or keep in step — and on the vendored llama.cpp
(b10200, pre-upstream #26296) its tensors are loaded whether or not anything uses
them, so the memory is already paid for.
DRAFT_MAX = 3.
On Qwen3.5-9B the same sweep could not separate the widths from noise (0.86–1.14×,
both directions), so the win is size-dependent and the small-model case is honestly
no win.accept between), so the
distribution is untouched; but llama.cpp's logits differ between batch width 1
and width 4, and that is enough. tests/batch_numerics.rs measures the gap at
max |Δlogit| 0.115–0.282 on hybrid Qwen3.5 against 0.002–0.003 on a dense
llama model, once exceeding the top-two margin outright; tests/speculative.rs
found the text differing on every run it ever made ("about the approach" vs "about
the algorithm", and so on). Fixing the seed neither reveals nor prevents this.ROTEIRO_SPECULATIVE=1, and only an explicit recognised
"on" — an unrecognised value is not consent. A completion that changes because the
decoder got faster must be asked for, not inherited by upgrading.ggml-org/Qwen3.8-27B-GGUF — the shape the registry installs —
ships its head as a separate file and records no nextn_predict_layers in its
main GGUF. Found by convention (mtp.gguf beside model.gguf, where mmproj.gguf
already sits), confirmed rather than trusted, and charged to the residency budget
alongside its target.mtmd_eval_chunks decodes the
prompt itself, so the drafter never sees those batches), and shared-KV MTP
architectures (new_context_with_ctx_other would alias the target's memory across
a teardown order this deliberately declines).tests/speculative.rs
asserts what holds (speculation activated, proposals accepted, control arm plain)
and reports the divergence.Goal: make a multi-gigabyte model store survivable. Nothing here touches the
graph, the schema or EXTRACT_VERSION — it is the store and its CLI only, which
is why it rides an independent track.
download_verified deleted its temp file on any
early return, so a transport failure at 90% of an 18 GiB pull threw away every
byte. On an imperfect connection that is not "slow", it is never finishes:
each attempt must win the whole race from zero. download_resumable keeps the
partial and continues it with an HTTP Range request; the transport stays in
the roteiro binary, so rto-graph still never touches the network.200 where 206 was asked restarts
rather than appending a whole-file body onto a prefix — which would corrupt the
file and surface only as a checksum failure after another full transfer. A
.partial.json sidecar records the URL, pinned digest and total size the bytes
were started against, because anonymous bytes cannot be trusted as a prefix.io::copy returns Ok having moved less than the whole file.
Diagnosed naively that is a checksum failure — which discards the partial, and
so defeats resumption for the commonest failure there is. The transferred
length is checked explicitly.roteiro model rm. There was no supported way to remove a model; files
accumulated with nothing but rm -rf. Removal reports what it freed, takes the
whole directory (a model's file set changes between releases, and bytes
model list no longer mentions are exactly the accumulation this stops), and
clears orphaned partials. model list gained measured on-disk size, so "what
would this reclaim?" is answerable before removing.serve holding a model is not
something rm can discover. The help says so instead of implying a check.ggml-org/Qwen3.8-27B-GGUF Q4_K_M as
the high generative tier. ggml-org over unsloth because the latter bundles
MTP tensors into the main GGUF (allocated whether or not the head runs);
Q4_K_M over Q5_K_M because generation on Metal is bandwidth-bound. The tiered
matrix is not a leaderboard — qwen2.5-coder-3b and the rest stay for slower
hardware.<tool_call>,
because a served model that cannot reach the graph tools is much less useful.Goal: close four defects that share one shape — each let Roteiro, or its own documentation, state something false confidently. Every one of them was found during this project's own development, and every one passed CI. The fix in each case is to make the failure loud, not merely to make the output correct.
adr-id (#324). ADR nodes are keyed adr:NNNN, so two ADR files
declaring one id collapse into a single node: query adr:0016 answers for one
decision while the other is invisible, @rto:0016 binds to whichever won, and
the published artifact carries only the survivor. The two files merge cleanly
in git and check reported 0 violations. This happened here — two parallel
branches both authored ADR-0016. check now reports a duplicate-adr-id
violation naming both paths and the id. The collision class does not exist
for blueprints, lat.md, files or symbols: their ids are their paths, and a
tree cannot hold two files at one path.debt(s, &[], &[]) — the
ignore lists passed empty — so the explorer UI counted markers the CLI
excludes and browser and terminal disagreed about the same repository. Fixed on
three axes: pass the exclusions; resolve them per repo from the target's
own root (ADR-0009's rule extended to scanning — a repository's own config
governs how it is scanned, whoever is asking); and make [debt] ignore
merge across config layers instead of the project layer silently discarding
the user layer, with an explicit ignore_reset as the way to inherit nothing.
ADR-0007 amended to v1.1. The MCP debt tool had the same defect.cargo-llvm-cov ran with a per-file floor; the workflow contained no coverage
tooling at all, making every DoD citing "85% coverage" unverifiable. Coverage
is now measured, non-blocking, and every document saying otherwise was
corrected. Measured baseline: 87.51% lines workspace-wide, 7 of 64 files
below 85%. The workspace already clears 85%; a per-file gate would fail seven
files whose coverage is low for reasons about what the code does (CLI wiring,
and paths needing a loaded model, a GPU or a sandboxed subprocess). Choosing
the threshold is deliberately a separate change, now informed by real numbers.graph.db already lives under each worktree's own git
dir, not the shared common dir — verified with real linked worktrees. The
observed symptoms ("check said 17 ADRs while 18 files sat on disk";
"sync said up to date while the store lacked three ADRs") were reproduced,
and their real cause is unrelated to worktrees: sync_worktree deliberately
overlays untracked files into the derived layer, while the authored layer
read only the HEAD tree — so a brand-new ADR had its symbols extracted but
was never parsed as an ADR, in a single tree, silently. Fixed by making the two
layers agree. The per-worktree layout is now pinned by a test, and a
worktree stamp (migration 12) makes a store that does come to hold another
tree rebuild loudly rather than answer "up to date".apply ran migrations with
version > MAX(recorded), so a store stamped by a build that knew migration
13 but not 12 never got 12 — 12 > 13 is false, forever. The store
opened cleanly, reported a schema it did not have, and failed at run time on
the missing column. Reproduced against a copy of this repository's real
graph.db: sync died on no such column: worktree, i.e. the #330 tree
stamp added above was itself the thing silently absent. Selection is now by
set membership, so an unrecorded migration is repaired on the next open
wherever it sits; ordering comes from the migration list, which a const
assertion holds strictly ascending at compile time. schema_version() now
reports the highest gap-free version rather than the maximum, so it cannot
name a schema the store lacks. No gate could have caught this: CI always starts
from a fresh store, where both rules agree. It is the EXTRACT_VERSION
incident's shape exactly — two independently-correct branches, a failure that
exists only in the combination. Consequence: merging guardrails before
stage25 was load-bearing and a mistake would have been permanent for any store
that met stage25 first; it is now a preference.Known gap, not fixed here (separate work). The other direction is still
silent: a store at migration 13 opened by a binary that knows only 1..12 opens
without complaint. Reads stay sound — migrations are additive in effect, so the
columns an older binary reads still exist — but sync would re-extract under an
older EXTRACT_VERSION and rewrite the graph with worse content: a silent
downgrade, not a crash. A hard error in apply was considered and rejected as
the wrong granularity, since it would also block the reads that are provably
safe. The fix belongs on the write paths (sync/reconcile/rebuild),
refusing to rewrite a graph whose store is newer than the binary, and it needs a
StoreError variant — a semver-visible addition on a 1.x crate. Filed rather
than folded in.
Deliberately NOT done: per-worktree databases. findings, media_content,
agent_memory and imports all live inside graph.db, and ADR-0013
v1.1 depends on that store being shared — its scope rule (a memory applies
wherever its anchor resolves, with no branch bookkeeping) was demonstrated with
one row in one store giving opposite verdicts on two branches. Splitting the
database per worktree would silently reintroduce the branch-scoping that ADR
rejected, and would need the ADR amended, not extended. Note the distinction
the codebase already draws and this preserves: ObjectCache is content-addressed
by blob id, so sharing it across worktrees is correct and valuable; the
assembled graph is not, so sharing it would mean last-writer-wins.
EXTRACT_VERSION, schema_version, migration counts);
migration 12 is covered by the existing additive-migration property test (#329)
rather than a pinned version number.Goal: one place decides which model serves a task, and can say why.
The user-facing gap: [models] has keys for embedding and generative
only. Vision, audio and OCR are hard-coded string constants
(voxtral-mini-3b, smolvlm-500m-gguf, ocrs-text), so a project cannot pin
its ASR model today. Seven surfaces each pick a model by their own rule —
spec draft, infer --model, serve load, serve/Ask answer, media
generation, OCR during sync — and none knows the others exist.
Shape: one function in rto-graph — the crate that structurally cannot reach
the network — taking (task_kind, modality, config, host_platform) and returning
the model plus the rule that chose it, folding in the scattered call sites.
Deterministic rules over categorical signals, not a classifier: every reliably
observable signal here is low-cardinality (installed, modality, build feature,
task kind), and a table over categoricals is the correct model.
The seed already exists and is the right one: chat_capable_model_ids filters
models that cannot do the job, and exists because routing a BERT encoder
through /v1/chat/completions aborts llama.cpp with a GGML_ASSERT. Generalise
that, rather than starting from "which model is best".
No network, no new dependency, no ADR — it implements what ADR-0003 and ADR-0007 already document, so it takes amendments with version-history rows, not a new decision.
DoD: vision, audio and OCR models are configurable and pinned per project;
roteiro config answers why did it use that model? for every surface;
resolution is deterministic and unit-tested without loading a model.
rto_graph::model_choice — resolve_with(task, pins) -> Result<ModelChoice, ModelChoiceError>, plus a process-wide pin slot published once at startup beside
the existing [paths] model_store one. ModelChoice carries the model, the rule
(pinned / built-in default), and whether the weights are on disk; the error
type names the offending key.
[models] grew to five keys — embedding, generative, vision, audio,
ocr — one per model kind, not per command, so generative governs both
spec draft and Ask.
Signature, corrected. The plan said (task_kind, modality, config, host_platform). modality turned out to be the same axis as task_kind — a
transcribe task is the audio modality — so a separate parameter would have
admitted the meaningless pair (Ocr, Audio). Host platform is read inside, from
the registry's existing Platform::host(), rather than passed: making it an
argument would have let a caller ask about a machine that is not the one the model
must load on. The signature is (task, pins).
Nine call sites, not seven. The plan's seven are all real and all folded in:
spec draft (generative), infer (embedding — the config half only; a --model
flag still wins and is validated by the embedder), serve load (served_models'
kind filter), serve/Ask (chat_capable_model_ids, plus a startup check so a bad
pin fails before the listener opens rather than per request), media generation
audio, media generation vision, and OCR during sync. Two the plan did
not name turned up while enumerating:
media status, which told an operator to roteiro model pull the built-in
default even when the project had pinned another model — advice that would have
them download the wrong weights and still be unable to build.media_env_tag), which folded ocrs-text by
name. Left alone it would have made [models] ocr the one pin that changes
what is extracted without invalidating what was extracted before it, so a
repository would keep serving text read by a model it no longer uses.Two behaviours changed for a set config, deliberately. spec draft used to
filter out a [models] generative that was not a generative model and fall
through to the default — a silent fallback, and exactly the failure this stage
exists to remove; it now refuses, naming the key. And a pinned model that is not
installed is now a hard error there, matching what roteiro infer has always done
with a configured embedding model. Unset behaviour is unchanged on every surface.
The one exception to failing loudly is roteiro config itself, which reports
a bad key rather than refusing — it is the command an operator runs because a
pin is misbehaving, so it must not be the command the pin breaks.
Cost, measured: +1,667 / −106 lines across 12 files (1,635 of the
insertions are Rust across 9 files; the rest is the two ADR amendments, this
entry, and the website's config sample). Against an estimate of S–M. The bulk
is the resolver itself (769 lines, of which roughly half is the module's own
documentation and its 12 unit tests) and the new 355-line CLI test. As with Stage
26, the surfaced-everywhere work is what costs: 366 lines of main.rs are the
call sites, the roteiro config resolution table, and its --json twin.
Gates: fmt clean; clippy --all-targets and clippy --all-targets --all-features clean at -D warnings; cargo test --workspace --no-fail-fast
853 passed / 0 failed; --all-features --no-fail-fast 1,073 passed / 0 failed.
EXTRACT_VERSION unchanged at 11, no migration, no new dependency, no network.
Every new test was fault-injected — 12 unit tests and 6 CLI tests, each shown to
fail under a mutation of the behaviour it claims to check, with the tree
byte-identical afterwards.
Unblocked. ADR-0019 is Accepted (2026-08-17), so this stage has a settled contract to build against. It remains the largest posture change in the project: the first capability that sends repository content off the machine.
Cut in three, guard first, then the thing it guards, then the surfaces that use it. Part 1 — the consent gate, the payload allow-list, the dry-run and the egress record — landed in a build that compiled no backend and therefore could not send anything. Part 2a is the ureq transport, the TTY form of the invocation grant, the response reader, and the README/website promise amendments that the transport makes necessary. Part 2b — delivered here — is the model_choice amendment and the spec draft / Ask wiring: it was cut out of 2a because it lands in a shared resolver that seven surfaces already read, and merging that with the first code in the project that can open a socket would put two unrelated risks in one review. See What shipped below.
Goal: an optional, explicitly-consented remote model backend for work local models cannot do.
Why an ADR is a prerequisite rather than paperwork. This is the first capability that sends repository content off the machine, and three written promises currently forbid it:
roteiro.toml is
committed and shared by design, so a project file may deny but never grant.
Grant lives at the user layer plus the invocation — both required, neither
sufficient. A teammate must not inherit egress from a merged line.The framing that decides the design: mis-routing among local models wastes tokens; mis-routing outward sends source off the machine for a reason nobody can inspect. The local→remote edge is not a routing decision — it is a gate the user opened. So no learned router, at any model quality.
And the disclosure gap must be stated in the ADR, not deferred: extraction
redacts secret-named config keys before persistence, but that is name-matching
over ten needles and there is no redaction chokepoint on a prompt. Prompts
carry symbol names and prose; DATABASE_URL=postgres://user:pw@host matches none
of those needles.
Also unresolved by design: ADR-0015's Producer identity folds a
model_digest. A hosted model has no digest — a vendor model string is a
mutable pointer, and the weights behind it can change while the name does
not. If remote output is ever stored, it needs ProducerTrust::{PinnedDigest, VendorAsserted} so a record states on its face that its identity is a claim.
Sequencing note: Stage 27 re-audits every "offline" claim. Landing this before it converts a documentation task into a re-litigation of the product's identity.
The stage is cut at the seam ADR-0019's own structure suggests. Part 1 is everything that decides, shows and records; part 2a is the thing that sends. The reason is not PR size, though the size is real: an egress path whose guard lands in the same change as its transport is a guard nobody reviewed on its own, and ADR-0019 §4 names that failure explicitly — "deferring this is how an egress path ships before its guard". Landing the guard first inverts that, and leaves a build that cannot send anything at all to review it in.
rto-remote — a new crate holding the policy and no HTTP client.
call_with takes the transport as a caller-supplied closure, exactly as
rto-exec takes its Fetcher, so the code that decides whether bytes may leave
is not the code that can make them leave. The guarantee is checkable from a
Cargo.toml rather than promised in prose, and every test exercises the whole
path with no network — a test cannot accidentally become the first thing that
sends data. rto-graph gains nothing: its gix is still pinned
default-features = false, and rto-remote depends on it.
Five modules, one per clause:
consent — ADR-0019 §3's inversion. ConfigGrant::from_layers is the
workspace's single implementation of "a project may deny but never grant", and
the binary's config layering calls it, so the value roteiro config echoes and
the value the gate consults cannot drift apart. Seven named Reasons — six
until part 2a split the invocation's two denial forms apart (see below) — each
with a remedy, except ProjectDenied, whose honest remedy is "no flag
overrides this; take it up with the repository". A discarded project grant is
reported rather than swallowed: a committed setting that silently does
nothing is worse than one refused out loud.payload — the allow-list as a type. ContextItem::from_node reads five
named fields off a node — key, kind, name, path, and up to 1,500 characters of
meta.content; every other key in its free-form meta is unreachable, and the
test that proves it plants a credential in a sibling key. disclosure() says
what leaves and refuses to stop at the reassuring half, naming the
DATABASE_URL case that matches none of is_secret_key's ten needles.record — the egress ledger at $ROTEIRO_HOME/remote/egress.jsonl
(owner-only on Unix). Endpoint, model, ProducerTrust, timestamp and a copy of
the body, written before the transport runs, so a call that hung is still a
call you know about. An unwritable ledger refuses the call rather than
sending unrecorded — the one ordering decision in call_with that is a policy
rather than a convenience.escalation — the deterministic post-hoc check. LocalAttempt carries
nothing but measurements of a finished run, so it cannot be constructed
before the local attempt happened — a stronger guarantee than a comment
saying so. A trigger is an input to the gate, never a substitute for it.trust — ProducerTrust::{PinnedDigest, VendorAsserted}, with the caveat
a vendor-asserted record is displayed with.Config. [remote] grows three keys, of which exactly one inverts. enabled
goes through ConfigGrant; endpoint and model are ordinary keys and layer
ordinarily, so a project may choose where its gateway is without being able to
turn the tier on. roteiro config prints the section as layers rather than one
merged value, because a reader applying the general precedence here would be
wrong about the one key where being wrong means believing egress is off when it
is on. Without the feature the section still prints, saying the build has no
tier — an omitted section reads as "no such setting".
CLI. roteiro remote status | dry-run | log, behind an off-by-default
remote feature that adds no third-party dependency. status reports the gate
layer by layer then the decision; dry-run prints the exact bytes and sends
nothing; log reads the ledger and says "nothing has left this machine" rather
than leaving that to be inferred from silence.
Part 1's build could not send. This one can, and the change of posture is stated everywhere it is now false to say otherwise.
The transport, and where it is not. crates/roteiro/src/remote_transport.rs
— one ureq call, in the binary, reachable from one command. rto-remote still
holds no HTTP client, and call_with is unchanged: the transport is handed
to it as a closure, so the crate that decides whether bytes may leave still
cannot make them leave, and that stays checkable from
crates/rto-remote/Cargo.toml. Three things it does not do, each for a written
reason: no reachability probe (ADR-0019 §2 — a probe is egress, and a DNS
lookup leaks the query to a resolver), no retry (a retried call is a second
disclosure, and that decision belongs to whoever consented), and no redirect
(max_redirects(0) — the ledger records the endpoint that was consented to, and
a 302 would make that record a lie about where the bytes went). It is not a
new dependency: ureq was already in the tree for model pull and security prefetch, exactly as ADR-0019 anticipated.
The credential is an environment variable and cannot be a config key.
ROTEIRO_REMOTE_API_KEY, because roteiro.toml is committed by design — the
same fact that inverted the precedence for enabled — and a key that could be
set there is a key that gets committed. It cannot reach the ledger either, and
not by discipline: the ledger records Payload::body, and headers are not part
of it. remote status reports whether one is set, never what it is.
response — the receive side, in the crate with no socket.
rto_remote::response::parse is the mirror of Payload::body: pure, over a
&str, so a truncated body, a malformed body and a body the endpoint filled with
its own error are all string literals in a unit test. It refuses a generation
that stopped short (finish_reason: length, content_filter, anything outside a
short list of completion reasons) rather than returning it, because a completion
that stopped early reads as finished — handing it over is the silent downgrade
ADR-0019 §6 most needs to prevent, wearing the endpoint's own name. It also
reports the one identity check this machine can make: a vendor model string is a
mutable pointer, so the weights cannot be verified, but a name that answered
under a different name than the one requested is a discrepancy, and it is
printed rather than dropped.
The prompt, and the two commands that must never show it. remote call is
the only surface that asks. status and dry-run are the commands you run to
find out what would happen, and a command that asks permission in order to tell
you is useless — so the rule is asserted rather than assumed, with a test that
runs both in the exact gate state a prompt could resolve. And exactly one gate
state may be resolved by asking: InvocationUnset — the human opted in, this run
has not. A prompt may never stand in for the user layer, or the two grants
ADR-0019 §3 requires separately collapse into one keystroke; it may never
override a project denial; and a non-interactive stdin is refused, not
assumed to agree, because a pipe cannot consent.
Two review findings on #386, both about a message that was untrue rather than merely unhelpful — the class this project keeps finding, because such a message passes every gate.
Declining a prompt claimed you had passed a flag. The invocation's two forms
were collapsed into one Option<bool>, so answering no produced
InvocationDenied, whose text names --no-remote — a flag the person never
typed and would not find in their shell history. On the consent path especially,
a message that misreports how consent was withheld undermines the thing it
reports on. Fixed with Invocation::{Unset, Flag, Prompt} and a seventh
Reason::PromptDeclined; decide keeps its Option<bool> flag form and became
a thin wrapper over decide_with, so there is still one implementation of which
layer outranks which. A distinct variant rather than a payload on the existing
one, for a reason that decides it: Reason::as_str is the stable token in
remote status --json, so a variant adds a token where a payload would change
the shape of one readers already parse. The test asserts the rendered text,
not the variant — a Reason that classified right while still printing the flag
would be the same bug.
And that variant was a semver break, caught in review. rto-remote had been
published at 1.19.0 hours earlier, and Reason was not #[non_exhaustive], so
the seventh variant would stop a downstream exhaustive match from compiling.
Three options were weighed. Design around it, as StoreError did
(#342/#348), does not transfer: there the fact had a home outside the enum, where
here the fact is which reason to report, so moving it out would leave
--json emitting invocation_denied for a prompt and reinstate the defect one
layer down. Accept it and mark the commit ! would cut 2.0.0, the
version this plan reserves for Stage 27 — the mistake #341 already made by
accident. So: #[non_exhaustive] went on in the same change, along with the rest
of the crate's open enums, while the crate had nine downloads and no consumer
that could exist. That is the workspace convention rather than a departure from
it — ExecError, SubprocessError, AssetError, AssetSource and
NetworkPolicy all carry it, and AssetSource's docs record that the attribute
"is what made adding it a non-breaking change". StoreError is the one that
missed the convention and paid for it. Trigger and ProducerTrust stay
exhaustive on purpose — closed sets, where a new member is a redesign a
downstream match should be made to notice — and say so at their definitions.
The version number will not record the break; Reason's doc comment, the
crate README and the commit do.
An absent finish_reason was read as "it finished". parse only refused when
the field was present and outside the allow-list, so silence passed — a strictly
weaker reading than the length case it already refuses, since length at least
says something. This is #367's rule at the other end of the same wire: a length
that cannot be established is not a length that checks out. Now
ResponseError::Indeterminate, and the message says why (completeness could not
be established) rather than that a field was missing. Checked before changing it
that nothing this tier addresses legitimately omits the field: rto-serve's own
ChatChoice::finish_reason is a non-optional &'static str, and the one shape
that legitimately carries null is a streaming delta, which Payload::body's
pinned "stream": false means this tier never asks for — so a null here is an
endpoint streaming at a request that said not to, which is the least complete a
body can be.
Testing the untestable bit. Part 1 could say the binary compiled no backend.
That sentence is now false, so the guarantee is re-established on different
ground rather than dropped: the granted-path CLI tests point [remote] endpoint
at http://127.0.0.1:1/…. Loopback is not a network — the kernel cannot route
those bytes to another host — there is no DNS lookup (which would itself be
the egress §2 describes), and nothing listens, so the call is refused in
microseconds. Everything about a successful response is tested over string
literals in rto-remote. The failure paths all have tests: no network, an
endpoint refusing by status, a malformed body, a truncated generation, and — the
one that matters most — a response arriving after the ledger write, asserted
from inside a failing transport closure, which only the write-first ordering
can satisfy. Reading the ledger after the call returns would be satisfied by a
single line written at the end, and the calls worth knowing about are exactly the
ones that never returned.
The promises, amended on the commit that makes them false. Part 1 left
"nothing leaves the machine" alone, correctly: with no backend compiled it was
still literally true of every build it produced. This build can send, so the
README and the website now say the scoped thing — that the sentence is true
of Roteiro as shipped and as configured, not of the software as a whole — name
the tier in their feature tables, and flag that --all-features includes it.
ADR-0006 was already scoped to serving in v1.3. The website additionally carries
the terminology guard ADR-0019 asked for: "Online mode" keeps its existing
meaning (a one-time, consented download, after which inference is local), and
a note beside it says in as many words that enabling
inference-local-models did not enable this.
Parts 1 and 2a built a tier you had to name to use. This part puts it behind two
surfaces where a person could reach it without typing remote — and most of
the work is making that impossible to do unaware.
ModelSource::Remote { trust }, and the three things it forced. The variant
carries the trust grade rather than a model name, because that is the whole of
what it exists to say: a hosted model has no registry entry and no digest, a
vendor model string is a mutable pointer, and a resolution must state on its face
that its identity is a claim (ADR-0019 §5). Three consequences, none of them
predicted by the plan and all of them structural rather than stylistic:
ProducerTrust moved crate, from rto-remote to rto-graph, re-exported
so rto_remote::ProducerTrust still resolves. Forced: rto-remote depends on
rto-graph, so a variant in the resolver cannot name a type in the crate above
it. The alternative was a second, parallel enum — which would let a ledger
entry and a resolution disagree about what "vendor-asserted" means, and a
grade two types could disagree about is not a grade. One definition now, three
paths, held level by a test that only compiles if they are the same type.ModelChoice.installed became Option<bool>, exactly as the plan
proposed. None is a third answer rather than a negative one: there are no
weights on any disk, so Some(false) would point a reader at roteiro model pull for a model no registry lists.ModelChoice.model stays a registry name and is None for a remote
choice. A vendor string in that field would put a mutable pointer in the slot
reserved for digest-pinned names. The endpoint owns the string, together with
the grade that qualifies it.The blast radius, counted rather than estimated. model_choice::resolve
spelled literally has 3 non-test call sites; the resolver family — resolve,
resolve_model, resolve_model_with, resolve_models — has 10. Eight
needed no edit, because 2b adds resolve_with_remote as a new entry point rather
than changing resolve_with's signature; the other two are installed
comparisons. A RemoteTier argument that defaults to Unavailable is what
makes that safe: a caller who forgets the tier gets the local answer by
construction rather than by remembering to.
Precedence, decided in the open. The pin resolves first, and its errors
still fail, even on a run whose local answer is about to be discarded — a
--allow-remote that made a broken [models] generative stop being reported
would hide bugs as a side effect of granting egress. Then the tier wins, because
a per-run flag someone typed is more specific than a standing project default —
and it wins out loud: ModelChoice::why names [models] <key> as not
applying, so a displaced pin is reported rather than silently skipped. Only
Draft and Chat are eligible; the other four tasks are refused structurally,
because Transcribe, Describe and Ocr run inside extraction where there is
no invocation to grant anything, so consent there could only come from a config
value — the user layer again, which ADR-0019 §3 says never suffices alone.
spec draft: the flag is the only way, and a refusal stops the run.
--allow-remote / --no-remote, and no TTY prompt — deliberately unlike
roteiro remote call. That command prompts because sending is what it does; here
the default is local, and a prompt on a default path turns a habituated "y" into
consent-by-default, which is the thing a two-layer gate exists to prevent. The
harder half is the refusal: a --allow-remote the gate turns down fails,
naming the layer and its remedy, rather than drafting locally. Handing back a
local model's prose there is a different answer with no signal that anything
changed — ADR-0019 §6's named failure, arriving through the consent gate instead
of through a socket. It is the same silent downgrade wearing a different hat, and
it is the failure this part most had to get right.
The remote draft is not the local prompt on a wire, and could not have been.
rto_spec::draft_prompt interpolates grounded symbol and ADR names into a
string. Sending that string would put graph content on the wire without it ever
passing ContextItem::from_node — the allow-list would still exist and simply
have nothing to do, which is the "whatever the local path happened to build"
assembly §4 forbids by name. So the remote path rebuilds: the instruction
carries the task and no graph content, and the nodes travel as allow-listed
context items. The cost is stated rather than hidden — a remote draft is not the
same prompt as a local one — because keeping the allow-list load-bearing is worth
more than prompt parity on the one path that leaves the machine. One call per
unfilled section, so each disclosure is its own ledger line rather than several
hidden inside one.
Ask: the served Engine, wrapped in the binary. The explorer's panel POSTs
/v1/chat/completions, which rto_serve::server handles over an
Arc<dyn Engine> — so Ask is not a Roteiro function and the branch had two
possible homes. crates/roteiro/src/remote_engine.rs is the one that changes
nothing: rto-serve gains no code, no feature and no dependency, and the
socket stays in remote_transport. The hosted model becomes one more served id
and leads the Ask pool, so models[0] — what the UI sends — is the tier a server
was started with --allow-remote to use; every local model stays served and
addressable, so a default moved rather than anything being removed. The cost is
named rather than discovered: that id is in GET /v1/models, so any client on
the port may address it — bounded by loopback-by-default, by the user layer
having granted independently, and by the ledger.
A chat transcript cannot be wrapped in an allow-list, so it is reduced. With
graph tools on, chat_with_tools injects tool results — raw graph query
output — into the conversation before the model sees it. Proxying that array
would route graph content past the guard entirely. So the remote path takes the
user turns and only those, drops assistant, tool and system turns, grounds the
question itself through rto_spec::context, and does not run the tool loop
remotely at all. A remote Ask is therefore not the same conversation as a
local one. That is a trade, made in the open, in the same direction as the draft
path's.
serve --allow-remote is a process-scoped grant, per ADR-0019 v1.2. Decided
once, in the command dispatch, before anything is built or bound — so a refused
grant stops the server from starting rather than letting it come up and answer
locally. The Decision is then a field on the engine: never recomputed, never
persisted, never read back from anywhere, dropped with the process. There is no
code that could infer a grant from a previous session because there is nothing to
infer one from.
remote now implies models, and it is a resolver dependency. The tier does
not choose itself: ModelSource::Remote is a variant of the shared resolver,
which is gated on rto-graph/models, and both surfaces ask that resolver which
backend serves them before either can send. Without it a --features remote
build would hold a consent gate and a transport and nothing that decides when to
use them. It pulls no llama.cpp and enables nothing — and it is what makes
spec draft reachable under remote alone, which is the point of a tier meant
for work local models cannot do.
Two defects found in review, and one gap the merge opened. Recorded because each was a class rather than an instance, and the class is the part worth carrying forward.
[remote] model could squat on a local model id. Nothing stopped
model = "qwen3-0.6b", and under serve --allow-remote that id became a
served model: /v1/models listed it twice and every request naming it was
answered by the hosted endpoint. The duplicate listing is the symptom; the
defect is that repository content would leave under a name that reads as
local. The user layer granted, so it is not unauthorised — it is
unrecognisable, which is §1's knowing gate and §4's inspectable payload
defeated one step removed. Refused now at remote_endpoint, the single
constructor, so the check is one place and covers every surface including
remote status, where someone deciding whether to grant should find out
first. Refused rather than resolved in either direction: preferring the
local model would leave a granted tier silently inert, which is this ADR's own
failure mode in other clothes. Checked against the registry rather than the
served set, so a configuration cannot be legal until someone runs
roteiro model pull.RemoteError was folded
into EngineError::Inference, which rto-serve maps to 500 — and this file
already disagreed with itself, using InvalidRequest for its own refusals
three lines up. Beyond the 4xx/5xx contract it told the wrong story: a 500
says the server broke, so a reader goes looking for a fault when what they
need is to grant consent. Fixed as a class, not an instance:
NotConsented → InvalidRequest (400), NoTransport → Unsupported (501, a
capability gap), Transport and Ledger → Inference (500, genuinely
server-side). No new EngineError variant, though 403 is the honest
status: that enum is not #[non_exhaustive] and rto-llama/rto-serve are
published at 1.x, so this is the rto_graph::StoreError case rather than the
Reason case — the fact is expressed within the existing set instead of by
widening it.#[non_exhaustive] guard lost ProducerTrust when the type moved.
#391 added every_public_enum_either_is_non_exhaustive_or_says_why_not, which
scans rto-remote/src/; 2b moved ProducerTrust to rto-graph and out of the
scan. The seen >= 9 floor did catch it — but as "the scan stopped
matching", not as "a type escaped", and only because the count was exact.
The scan now follows the crate's re-exports, and asserts ProducerTrust by
name so the list cannot be emptied and replaced by an unrelated addition.
#391's reasoning for leaving the type exhaustive moved with the type, since it
is a property of the distinction and not of the directory.The promises, amended again on the commit that changes them. 2a's README and
website text described a tier reached through roteiro remote …. Two more
commands can send now, so both say so, name which one prompts and which two do
not, and state the serve exposure in the terms ADR-0019 v1.2 uses — including
that the person starting the server consents on behalf of every later request to
it.
Gates: fmt clean; clippy --all-targets and --all-targets --all-features
clean at -D warnings; cargo test --workspace --no-fail-fast 981 passed / 0
failed and --all-features --no-fail-fast 1,240 passed / 0 failed; cargo deny clean. EXTRACT_VERSION unchanged at 12, no migration, no new
dependency, and no network — in either the code or the tests, where the only
address any granted call may reach is http://127.0.0.1:1/… (loopback, a literal
so there is no DNS query, and nothing listening). Every new test was
fault-injected: 6 resolver, 1 re-export, 7 engine, 1 Ask-pool and 10 CLI, each
shown to fail under a mutation of the behaviour it claims to check, with the tree
byte-identical afterwards. The collision guard's test asserts the rendered
message names both the key and the local id, because "that name is taken" is
unactionable without saying by what; and the semver guard was shown to fail under
a mutation that empties its re-export list, which is the regression the type's
move actually caused. The Ask grounding budget is asserted at compile
time rather than in a test, so violating it fails the build that edits the
literal.
roteiro review LLM mode → v1.18.0 · effort M–L (independent track) ✅ delivered — the instrument, and a negative result on the reviewerDepends on Stage 33 (a reviewer must resolve a model without a fourth bespoke rule). Independent of Stage 34 — it can run wholly local.
Split into two PRs at the measurement seam. 35a — the scoring harness and the suppression filter — is delivered:
roteiro review --score, the corpus as a typed shipped asset, per-class recall, andrto_graph::compile_claim. It lands before any reviewer because it is what makes a reviewer's value a number instead of an impression, and because a "do not build" verdict needs the same harness a "build it" verdict does.35b is cut in two at the measurement seam, and PR 1 is delivered: the reviewer, its two surfaces (
review --llmand thereview --replayharness),ModelTask::Review, and the compile-claim suppression wired throughcompile_claim's four axes. PR 2 is the graph-context arm and the comparison that is the actual experiment. PR 1 ships no per-class recall figure — see below for why that is a result rather than an omission.The original 35b framing follows, and its second bullet has since been overturned by measurement. Two measurements from 35a constrain it, and both were taken on this repository rather than assumed:
- Whole-diff review is not the shape. Reconstructing the 15 review diffs at
-U3costs ~513k tokens total, ~34k mean, 103k worst (PR #339). Against the measured ~30k single-call budget, 9 of 15 do not fit in one call before any graph context is added. The ~79k per-file budget is the one to design to, which putscontract-drift— the largest class, 5 of 22 — squarely on the graph: per-file review cannot see a doc in another file contradicting the code under review unless something hands it that doc. That is the claim to test, and it is now testable.- The corpus can falsify a reviewer but not finely rank two. 22 real rows over 13 classes, 8 of them holding a single row. A 0-or-2-of-22 result is decisive; a 9-vs-11 difference is not. 35b's DoD should be a floor to clear, not a percentage to maximise.
35b also needs a resolver addition, not a workaround:
ModelTask::Reviewinrto_graph::model_choice, sharing thegenerativekey withDraft/Chat. Adding a seventh task is the Stage 33-sanctioned move; a bespoke selection rule in the reviewer would be the fourth one Stage 33 exists to have removed.
Goal: give the adjudicated review corpus a consumer, and put the graph to work on the one thing a diff-only reviewer structurally cannot see.
The asset that already exists: crates/rto-graph/tests/fixtures/review/ holds
26 adjudicated review comments with verdicts, defect classes, and the
reviewed_sha each was left on. That makes a reviewer measurable rather than
guessed at — which most projects cannot do. It currently has no consumer,
and that is how a fixture rots.
What the graph adds, stated honestly: not access — Copilot has been agentic
since March 2026 and reads repository context via tool calls, verified against
this corpus (it cited a file outside a PR's changed set, correctly). What
roteiro review has is pre-assembled, provenance-tagged context: governing
ADRs, authored drift, blast radius, intent debt. A weaker claim than a moat, and
the one to test.
Two constraints inherited from the investigation:
msrv job already refutes them ~60 s before a human reads the comment. So
withhold any finding claiming the code will not compile while the relevant
check is green at that commit and configuration — see
docs/REVIEW_CHECKLIST.md, which records why "green build" alone is too coarse
(ubuntu-only, --all-features; the macOS teardown abort of #291 was invisible
to it).DoD: scored against the corpus with per-defect-class recall — not an
average, which hides the only thing an implementer needs — at the
reviewed_sha of each comment, never the PR head (merged heads contain the fix
commits, so scoring against them measures recall on already-fixed code and
silently reports zero).
Delivered in 35a, so 35b inherits rather than rebuilds it: review --score,
per-class recall with the denominators printed beside the rates, an outright
refusal to score a commit the corpus does not know (the PR-head guard), and
compile_claim's coverage model. One thing 35a had to fix on the way: the
corpus README's own reconstruction recipe — merge-base <base> <reviewed_sha> —
produced an empty diff for 13 of the 15 review commits, because a merged PR
branch is an ancestor of main and the merge base is then the review commit
itself. A reviewer handed an empty diff also scores zero, silently, from the
opposite direction. The recipe is corrected and now has an executable form
(every_row_reconstructs_a_non_empty_reviewed_diff), which also checks that each
reconstructed diff touches the file its comment is anchored to.
The budget is not the constraint — the earlier conclusion understated the room. 35a established that whole-diff review does not fit and named ~79k per file as the figure to design to. Measured across all 190 changed paths of the corpus (of which 184 have a reviewable diff; six are binary audio fixtures): mean 2,704 tokens raw and 3,275 as sent, median 1,476/1,758, p90 5,621/6,711. Exactly one of 190 exceeds even the ~30k single-call budget, and it is a generated JSON fixture. The largest reviewable source file-diff in the corpus is 14,034 raw and 17,202 as sent — under two-thirds of the single-call budget. "As sent" carries the line-number column the reviewer adds, a measured 1.21× (9 characters per line, so ~1.2× on source and far worse on very short lines).
That is what makes PR 2 worth running at all: the median file leaves ~28k of the budget unused, so the graph arm has somewhere to put whatever it has to add. Had the budget been tight, the stage's central claim would have been untestable here whatever the graph contained.
The instrument had a vacuous measurement in it, and PR 1 walked into it. A
reasoning GGUF opens <think> and deliberates before answering. Run with a
1,200-token generation cap, qwen3.8-27b spent the entire budget inside that
block on 4 files of 4 of a held-out commit, and the harness reported "0
finding(s) over 4 file(s)". Nothing was wrong-looking about that output. Scored,
it would have produced zero recall across every class and read as a clean,
honest negative result about local reviewers — when in fact no review had
happened at all.
This is the same silent zero as scoring against a PR head or against an empty
diff, arriving from a third direction, and it is worse than both because the other
two are refused by construction. It is now detected
(reviewer::Parsed::reasoning_truncated), reported on both surfaces as not
reviewed, and pinned by a test; the cap is 4,096 and the context window is passed
explicitly rather than inheriting LlamaEngine's 4,096 default, which had already
killed an earlier run outright. A truncated file is never counted as clean and a
run containing one is never scored.
The general lesson, for anyone measuring a reasoning model on anything: a generation cap does not degrade a reasoning model's answer, it removes the answer while leaving the response well-formed. Silence and truncation must be distinguishable at the instrument, not inferred afterwards from a suspicious number.
Not yet scored, deliberately. PR 1 ships no per-class recall. The pass that
would produce one had not been run under the fixed harness at the time of writing,
and a recall figure from a run where truncation was possible is not a measurement.
The remaining observation is honest but not a score: with qwen3-coder-30b-a3b
(non-reasoning, so unaffected by the above) the reviewer emitted 60 findings
across 4 held-out files — ~15 per file, with visible duplicates — and two
prompt revisions written specifically to impose precision discipline changed that
number not at all. Against 22 adjudicated rows spread over 184 files, that rate
is the number to beat, and it is the one PR 2 has to move.
"Do not build this" remains an acceptable outcome. The corpus keeps its value either way: it is how any future reviewer, hosted or local, gets measured. Nothing in PR 1 is evidence for building PR 2 except the headroom finding; the finding rate is evidence against, and both go into the decision.
The pre-committed floor was not cleared, and the finding that matters is that it could not have been. The interesting half of this result is not a disappointing number; it is that three measurements taken before the graph arm ever ran bound what any run of it could possibly show, on this corpus, with any model. That bound is stated first because it is the durable part.
PR 1 shipped no per-class recall deliberately — a figure from a run where
truncation was possible is not a measurement — so there was no baseline, and
"the graph arm found N contract-drift rows" would have meant nothing beside it.
PR 2's first deliverable was therefore the diff-only arm under the corrected
harness: GraphContext::none, the truncation guard armed, scored per class at
each row's reviewed_sha. The graph arm followed as the same binary, the same
model, the same corpus and the same reconstruction, differing in exactly one
variable — which is now recorded in the run document (RunArm) rather than in
a filename, because a comparison a reader cannot audit from the artifacts is not
one.
1. The recall ceiling is 21 of 22, not 22 of 22. Scoring matches a finding to
a row within LINE_WINDOW (±10 lines), so a row whose anchored line is not in
the reconstructed diff cannot be found by any per-file diff reviewer, whatever
the context. Measured over every row: 20 of the 22 real rows have their anchored
line inside the shown -U3 diff, 21 fall within the scoring window, and one
does not. It is crates/roteiro/src/config.rs:634 — the ignore_reset
inherited via .or() — whose line sits in the gap between hunks covering 564–575
and 984–1180. It is a contract-drift row, so that class's real ceiling is 4 of
5, not 5 of 5. The reviewer that found it originally was agentic and read the
whole file; a per-file diff reviewer structurally cannot, and no amount of graph
context moves a line number into a hunk.
2. The arm's dose is real, not nominal. Over the whole corpus the assembler
emits 922 items on 85 of the ~184 reviewable files — 189 authored (governing
ADR and blueprint sections) and 733 derived (doc comments from elsewhere in the
file) — for ~152k estimated tokens, about 1.8k per file that carries any,
with 2,305 further items dropped by the cap. On a median file that roughly
doubles the prompt. So the treatment was administered; a null result here is not
a null dose.
3. And the one that decides it: on 3 of the 4 reachable contract-drift rows,
the two arms send a byte-identical prompt.
contract-drift row | reachable? | context the graph arm supplied |
|---|---|---|
rto-graph/src/engine_slot.rs:16 | yes | none — file added in that commit |
rto-graph/tests/fixtures/audio/README.md:11 | yes | none — markdown fixture, no symbols |
docs/adr/0005-image-ocr-vision-ingestion.md:16 | yes | none — an ADR under review |
rto-graph/src/query.rs:716 | yes | 19 items (0 authored, 19 derived) |
roteiro/src/config.rs:634 | no | 21 items (6 authored, 15 derived) |
Each zero has a specific and defensible cause, and none of them is a bug in the assembler:
adr_section node per heading with no body, and no
authored edge points into one. The row is "frontmatter bumped to 1.3 while
the summary table still reports 1.2": both halves live in the ADR, one is
outside the -U3 window, and the authored layer cannot describe the authored
layer. The corpus's clearest case of an ADR contradicting itself is the case
the graph is blindest to.So at most one contract-drift row could ever change — and the baseline
below makes that bound tighter still. The single reachable row that does receive
context, query.rs:716, is the one row the diff-only arm already matched. The
graph arm therefore cannot gain a contract-drift row from any starting point;
it can only lose the one already held. The floor — at least 2 of the 5 recovered
that the diff-only arm missed — was not merely unreachable before a single token
was generated, it was refuted: the number of recoverable rows is zero. That floor was chosen honestly,
before anyone knew which way the numbers would fall, and it stands as written;
what the measurement adds is why it could not be met, which is worth more than
the miss.
Stage 30 measured that llama.cpp's logits differ between batch widths, so "same seed, same answer" could not be assumed here — a single run of each arm would not have been a comparison. Measured rather than assumed: the diff-only arm was run twice over the first 4 corpus commits (54 files, 29% of the corpus), the second time through a different binary — the one carrying this PR's changes.
626 findings both times, 534 distinct both times, Jaccard 1.0000, not one
finding different. So two things hold at once: greedy decoding at
temperature 0.0 with speculation off is reproducible run to run here, and the
refactor did not perturb the diff-only path. Stage 30's divergence needs
ROTEIRO_SPECULATIVE, which no run in this stage set.
Stated limit: determinism was verified on 54 of 184 files, not all of them. The remaining 130 were not re-run, because the finding below made a second full pass the wrong place to spend three hours.
| diff-only arm | |
|---|---|
| findings | 1,995 over 183 files — 10.9 per file |
| exact duplicate lines | 346 (17.3%) |
labelled contract-drift | 1,412 (71%) |
| claiming compile failure | 0 |
| files declared clean | 127 of 183 |
| unadjudicated | 1,990 of 1,995 |
Two of those rows are results in their own right.
The reviewer says contract-drift about almost everything. 71% of its output
carries that one label, and it recovered one of the five real contract-drift
rows — by chance, on the permutation test above. A classifier that answers the
same thing to nearly every question carries no information in its answer, and it
explains why the class the graph arm was built to help is also the class the
model over-claims: the arm was aimed at the one place the reviewer was already
saturating.
The compile-claim filter never fires. 35a called it a free precision filter, on the strength of every false positive in the corpus being a compile claim (4 of 4). Against this model it is dead weight: not one of 1,995 findings claims the code will not compile. The filter is correct, cheap and idle. That is not an argument to remove it — a different model would claim differently, and the corpus says a human reviewer did — but it must not be counted as precision this reviewer is getting.
The baseline this stage never had: qwen3-coder-30b-a3b, 15 of 15 commits, 183
files reviewed, one file refused (Cargo.lock, 64,389 tokens against the
49,152-token window) and one truncated. 4 of 22 real rows found, 1 of 4
known-false reproduced, and 1,995 findings emitted — 10.9 per file, of which
1,990 are unadjudicated.
Then the number was checked against a null, and it did not survive.
review_score::match_findings credits a finding to a row on (commit, path, line ±LINE_WINDOW) and never on what the finding says — not its class, not a
word of its description. That is the correct rule for a scorer that must not
reward eloquence. Its consequence had not been measured: a reviewer emitting
10.9 findings per file blankets the diff, and a "hit" is then explained by
density rather than by insight.
Permutation null. Relocate every corpus row to a uniformly random line its own reconstructed diff actually shows; leave the run's findings byte-for-byte as emitted; rescore with the same greedy one-to-one rule. Over 2,000 trials:
| observed | null mean | P(≥ observed) | |
|---|---|---|---|
| real rows matched | 4 | 4.19 | 0.72 |
| known-false reproduced | 1 | 0.49 | 0.49 |
The reviewer scored below chance. A tighter null that relocates only to added lines — where 23 of the 26 rows actually sit — gives 3.98 and P = 0.62, so the result is robust to the obvious objection. The reimplemented matcher reproduces the shipped scorer's counts exactly (4 real, 1 known-false), which is what licenses the comparison.
So 4/22 is not a recall figure, it is arithmetic, and the same is true of
the 1/4 known-false "reproduction": inspected, the credited finding is a
contract-drift claim about a Range: bytes= doc at line 2312, while the row at
2322 is a false compile claim. Different claims, ten lines apart. The scorer
cannot tell them apart and was never built to.
This is the fourth silent-zero-shaped trap in this stage, and the first that is
silently non-zero. Scoring against a PR head, against an empty diff, and
against a truncated reasoning reply all report a clean zero that measures
nothing; this reports a clean positive that measures nothing. It is now caught
in code rather than in prose — Score::expected_by_position prints beside the
recall, and caveats() states outright that the rate is not clearly above
chance — on the same principle that made reasoning_truncated a reported outcome
in PR 1: a number whose null is not stated is not yet a result.
The shipped estimate is honest about being an estimate. An exact null needs the
diff, and review --score is pure by design so a published score recomputes on
any machine, so it uses the candidate's own findings as the proxy for where it
looked. Summing the ±10 neighbourhoods over-reads at 6.2; merging them
under-reads at 3.0 against the permutation's 4.19. It is an order-of-magnitude
guide, the caveat fires on a 2× margin because of that, and the docs name the
permutation as the thing to run before believing any comparison.
What this costs the experiment. A recall comparison between two arms is only meaningful if either arm's recall is meaningful. At 10.9 findings per file neither is. The density sweep says how far that has to fall: thinning the run's own findings and re-running both the null and the observed match, the two stay within noise of each other at every rate measured — 4.19 vs 4.00 at 10.9 per file, 1.67 vs 2.08 at 2, 0.67 vs 0.94 at 0.5. The reviewer's recall never separates convincingly from chance at any density it was run at. PR 1 called the ~15-per-file rate "the number to beat"; it is worse than that — until it falls, the corpus cannot measure recall at all.
The premise, restated against the data. Stage 35's central claim was that
contract-drift "puts the largest class squarely on the graph: per-file review
cannot see a doc in another file contradicting the code under review." Measured
against the rows rather than assumed: four of the five contract-drift rows
have both halves inside the file under review, and three of those have both
halves inside the diff. The class is not, on this corpus, cross-file. The
premise was reasonable and it was wrong, and it was wrong in a way only a corpus
could show.
The graph arm completed and was scored against the same corpus with the same binary, so the two arms differ in exactly one variable.
| diff-only | graph | |
|---|---|---|
| real rows found | 4 / 22 | 5 / 22 |
contract-drift | 1 / 5 | 1 / 5 |
| known-false reproduced | 1 / 4 | 0 / 4 |
| findings emitted | 1,995 | 2,059 |
| unadjudicated | 1,990 | 2,054 |
contract-drift did not move, exactly as the power bound said it could not.
That bound was computed before the arm ran: three of the four reachable rows
receive no context at all, and the fourth is the one the diff-only arm already
held. Recoverable rows were zero, and zero is what changed.
Against the floor pre-committed in PR 1:
contract-drift rows recovered — FAILED (0). Refuted in advance by
arithmetic, not by the run.contract-drift
claim at 2312 credited to a false compile claim at 2322), and the graph arm
does not reproduce it.The one row gained is missing-event, a class with a single real row. The
scorer's own caveat says a singleton class is one bit rather than a rate and
must not be read as a percentage. And it was gained while the finding count rose
by 64 — so under the density explanation the baseline established, more
findings producing one more hit is the expected direction of noise, not
evidence of insight. A 4 → 5 move sits inside a null whose mean was 4.19.
Verdict: the graph arm is not adopted on this evidence, and the corpus cannot be made to say otherwise. Not because the graph is useless — because at ~11 findings per file neither arm's recall separates from chance, so a comparison between them has nothing to compare. What would change the answer is a reviewer whose finding rate falls far enough for recall to become a measurement; until then, more context is tuning against noise.
The chance baseline fires on both arms, which is what makes the comparison readable at all:
armA 4/22 chance baseline: ~3.0 would match by position alone
armB 5/22 chance baseline: ~2.6 would match by position alone
both RECALL IS NOT CLEARLY ABOVE CHANCE AT THIS FINDING DENSITY
Read the estimator as the order-of-magnitude guide it says it is: it returns 3.0 where the exact permutation null returned 4.19, the ~30% under-shoot its own docs predict for merged neighbourhoods. No permutation was run for the graph arm, so its 5-against-~2.6 must not be read as separation — corrected for the same under-shoot its null sits near 3.6, and the caveat's own instruction is to confirm with a permutation before comparing two candidates.
(An earlier revision of this section recorded "the guard does not fire" as a
defect. That was wrong, and wrong in the way this document keeps cataloguing:
the scoring was run with a binary built 09:27, two and a half hours before the
feature landed at 11:55 — strings finds no trace of it in that artifact. The
claim was checked against the wrong build, not against the code. Recorded rather
than quietly deleted, because a false gap in a document about false measurements
is worth one paragraph.)
The hardening audit ran; the release is held. Those are separate, and conflating them is how a stage looks unstarted when its work is done.
What the audit established, and it is the part worth keeping. All eight whole-graph surfaces are linear, measured to 72,509 nodes / 126,622 edges — 25× the graph Stage 26 measured, on a real repository rather than a synthetic one, worst absolute case
duplicatesat 0.89 s. Coverage went up to 89.77% while files measured grew 64 → 94.cargo deny --all-featuresclean.VENDORED_DEPENDENCIES.md's version rows exact, with zero advisories published since the vendored commit. That is a "fine at both, recorded so nobody re-measures it" result and it cost a 25×-scale benchmark to get.What it found became issues, not paragraphs: #447 (
recall --limit 0), #448 (Provenanceclosed?), #449 (advisory reachability), #450 (coverage), and the measured data on #431 — 66 exhaustive public enums across nine published crates, which is the real v2.0 semver decision.The release is deferred on the owner's call, tracked at #429 with a completion condition rather than open-endedly, and the
Release-plzworkflow is disabled so it cannot cut by accident. The original deferral note follows, recorded so a stage nobody got to and a stage somebody decided to leave stay distinguishable.
Deferred deliberately, not merely unstarted. The owner's call, recorded so the two are distinguishable: a stage nobody has got to and a stage somebody decided to leave look identical six months later, and the second should not be picked up by whoever next has a free afternoon.
Nothing blocks it — Stages 21–25 and 28–32 are delivered and v1.15.0 is out. It is held because the hardening it describes is worth more once the work ahead of it has landed and been measured: the remaining A1 lenses (Stage 26) and Stages 33–35. And because v2.0.0 is a number worth spending once, deliberately.
Stage 34 in particular should land before this, not after: Stage 27 re-audits every "offline" claim, and adding a remote tier afterwards would reopen an audit that had just been closed.
One consequence to carry: the scope below grew during v1.10–v1.15. The offline claim it re-audits is now a real surface —
docs/OFFLINE_SETUP.md, thesecurity prefetch/statuscontract, digest-pinned assets, and per-file verification of the extracted boxlite runtime — so this is an audit against something concrete rather than a prose sweep.
docs/JSON_SCHEMA.md extended for findings + memory,
every "offline" claim re-audited to say offline-capable once provisioned
where that is the truth.cargo deny clean
with --all-features on the resolved native closure.| Release | Contains | Gate |
|---|---|---|
| v1.10.0 ✅ | Stage 21 — analyzer contract + ingest | Artifact byte-identical; ingest idempotent — met |
| v1.11.0 ✅ | Stage 22 — semgrep + cargo-audit (SAST axis, five languages) | Offline warm-cache run; named cold-cache failure — met |
| v1.11.x ✅ | Stage 22b — osv-scanner (dependency axis: Python/Java/Node) | Lockfile findings per ecosystem; Rust overlap resolved — met |
| v1.11.0 ✅ | Stage 23 — episodic memory | Survives rebuild; graph untouched — met (#317) |
| v1.13.0 ✅ | Stage 24 — boxlite backend | Parity with subprocess; cargo deny clean — met (#352): identical finding keys via both backends, differing only in isolation label and image digest |
| v1.12.0 ✅ | Stage 25 — recall + bounded cache | decay=none reproducible; no episodic eviction — met (#340). Shipped two releases ahead of its nominal target |
| v1.13.0 ✅ | Stage 26 — lenses Q3/Q1/S1 | All three, one PR each — Q3 (#346), Q1 (#372), S1. coupling --limit 0 over 2,887 nodes in 0.07 s; debt-density --limit 0 and config-secrets --limit 0 (372 keys) both 0.04 s, level with the debt baseline. No lens offers a CI gate, each saying why: Q3's cross-language call edges are name collisions (615/6,553 = 9.4%); Q1 inherits the marker scan's prose false positives and adds a file-length denominator; S1 is an inventory, and its one real finding indicts Roteiro's import layer, not the user's repo. EXTRACT_VERSION stayed 11 throughout. Cost: 11 files, 3,016 insertions — the corrected 195–500 LOC estimate is itself out by ~3–8× |
| v1.10.x ✅ | Stage 28 — generated media content moves out of derived | Silent clip cannot reach default search; media build restores searchability — met |
| v1.11.0 ✅ | Stage 29 — audio metadata as derived facts | Format read costs 1–100 µs and instantiates no decoder; duration exact/estimated/absent never guessed — met |
| v1.11.0 ✅ | Stage 30 — MTP speculative decoding | Opt-in only; 1.22–1.50× on 27B — but output is not identical, so default-on is blocked on §9.6 |
| v1.11.0 ✅ | Stage 31 — model lifecycle: resumable pulls, model rm, high tier | Interrupted pull transfers only the remainder; checksum failure discards; pinned digest measured, not quoted |
| v1.12.0 ✅ | Stage 32 — guardrails: four confident wrong answers (#324, #321, #319, #330) | Two ADRs on one id fail check naming both files; API and CLI debt agree, per repo; coverage measured (87.51% lines, 7/64 files under 85%) with no document claiming a gate that does not run; a new ADR on disk is never silently uncounted — met |
| v1.16.0 ✅ | Stage 33 — local model resolution | Vision/audio/OCR pinnable per project; roteiro config answers why that model for every surface |
| v1.17.0 ✅ | Stage 34 — remote model tier | ADR-0019 Accepted. Cut in three, all delivered. 1 — the guard (#381): consent gate, payload allow-list, dry-run, egress ledger, in a build compiling no backend. 2a — the transport: ureq behind the off-by-default remote feature, the TTY invocation grant (status/dry-run never prompt), the response reader that refuses a truncated generation, and the README/website promise amendments on the commit that made them false. 2b — the surfaces: ModelSource::Remote { trust } with installed: None in the shared resolver (which forced ProducerTrust from rto-remote into rto-graph — one definition, or a ledger entry and a resolution could disagree about what "vendor-asserted" means), then spec draft --allow-remote and serve --allow-remote over it. Neither prompts — the flag is the only way on a surface whose default is local — and a refused --allow-remote stops the run rather than answering locally, which is the same silent downgrade a network failure would be. Both rebuild the request through the payload allow-list rather than forwarding a local prompt or a chat transcript, so the guard still assembles what leaves. Ask is wired by wrapping the served Engine in the binary, so rto-serve gains nothing. Project file may deny, never grant; no learned router on the local→remote edge; no reachability probe; no test can reach a network |
| v1.18.0 ✅ | Stage 35 — roteiro review LLM mode | 35a delivered (#380): roteiro review --score, per-class recall with denominators, and compile_claim's four axes — the instrument, with no verdict. 35b PR 1 delivered: ModelTask::Review on the shared generative key (a seventh task, not a fourth bespoke rule), the per-file reviewer, review --llm, and the review --replay harness. It overturned 35a's budget conclusion in the useful direction — per file the corpus is mean 3,275 / median 1,758 tokens as sent, and the largest reviewable source file-diff is 17,202, so ~28k of the single-call budget is free on a median file and the graph arm of PR 2 has room to be tested. It also found a vacuous measurement in the instrument itself: a reasoning model spent a 1,200-token cap entirely inside <think> on 4 files of 4 and the run reported "0 finding(s) over 4 file(s)" — zero recall that measured nothing, now detected, reported as not reviewed, and never scored. No per-class recall figure yet, deliberately: the honest remaining number is ~15 findings per file against 22 adjudicated rows over 184 files, unmoved by two precision-targeted prompt revisions |
| v2.0.0 ⏸️ | Stage 27 — hardening | Full gates; semver review complete — hardening delivered, findings filed as #431/#447–#450; the release itself is held on the owner's call (#429) and the Release-plz workflow is disabled so it cannot cut by accident |
| Risk | Severity | Mitigation |
|---|---|---|
A V2 record leaks into nodes/edges and breaks artifact purity | High | NodeKind::Other("…") is mechanically possible — that is the trap. CI regression test asserting export_factset is byte-identical across ingest/memory writes. |
Unreviewed memory acquires authored relevance | High | Separate store, separate ranking channel; assert in tests that memory never scores through the authored path. |
| Memory captures secrets (tokens, stack traces, customer names) | High | Uncommitted .git/roteiro/ placement; explicit forget; documented that memory has no redaction chokepoint. |
| boxlite advisory lands and is missed | Medium | Exact pin + deliberate advisory tracking as a standing duty (ADR-0014). |
--all-features CI fails without /dev/kvm | Medium | Runtime capability probe; sandbox tests skip visibly. |
| Unbounded episodic growth | Medium | Accepted by design; explicit user reclamation only. |
| A single-vendor factual claim drives a design | Medium | This plan already survived one: a "boxlite is unpublished, therefore unmergeable" blocker was refuted by direct crates.io checking. Verify checkable externals independently. |
EXTRACT_VERSION bumped twice, forcing two full re-extractions | Low | Batch all extraction-touching lenses behind one bump (Stage 26). |
| Speculative decoding silently changes a completion | High | Measured, not hypothetical (Stage 30). Off unless ROTEIRO_SPECULATIVE explicitly says on; unrecognised values are not consent. The risk is acceptance by default, so the mitigation is that there is no default. |
Stage 27 is the last scheduled stage. It is not the last work, and the difference has been invisible: everything below was decided somewhere in this document and then scattered across §9 and Stage 26's deferral list, where a reader planning the next quarter would not find it. This section is a map, not a new commitment — each item keeps its original reasoning at the reference given.
| Item | Where | Why it is not in a stage |
|---|---|---|
| Semantic recall — vector index over memory | §9.3 | Needs migration, model/dimension versioning, retention, rebuild and storage-size policy. Materially more than "persist embeddings". |
| Findings ↔ graph cross-surfacing | §9.7 | Joining findings to graph facts is deliberately not free in this design. When wanted, it needs a designed join, not an implicit one. |
code_interpreter | §9.4, ADR-0014 | Rejected. The real question is is local code execution something Roteiro wants to be? — a product decision, not a backend one. |
| Q2 — LOC hotspots | Stage 26 | Not a pure query: Node.span is byte offsets, so it needs net-new extraction metadata. |
| Q10 — dependency pins | Stage 26 | Mis-scoped as written; existing pins are Docker image_ref and submodules, so package-manifest pins are extraction work. |
| Q7 — doc coverage | Stage 26 | Needs a language and a denominator; docs live mostly in symbol meta.content, not Doc nodes. |
| S2–S6 — the rest of the security lens series | Stage 26 | Taxonomy normalised (S1, S4 → GDS; S2, S3 → NNX; S5, S6 → EXT), but none is scoped. |
Q2, Q10 and cross-language call-edge resolution each need extraction metadata,
so each forces an EXTRACT_VERSION bump — and every bump is a full
re-extraction for every user. The risk register already says to batch
extraction-touching lenses behind a single bump. That makes these a cluster
rather than three independent tickets: doing them one at a time is the expensive
way to do the same work.
Cross-language call-edge resolution belongs in that cluster and is not yet
recorded anywhere else. Stage 26's Q3 measured 615 of 6,553 call edges (9.4%)
on this repository as cross-language name collisions — cross-file resolution
binds a callee by simple name across every Fn node regardless of language, and
no FFI is extracted. That is why Q3 offers no CI gate. Fixing it is extraction
work, so it batches with Q2 and Q10 or it is paid for twice.
The placeholder marker-needle correction (#384) spent a bump on its own
(EXTRACT_VERSION 11 → 12), by the owner's decision. It is not a fourth member of
the cluster above: it shipped alone, and Q2, Q10 and cross-language edge
resolution remain a cluster with each other, still unpaid and still
unscheduled — this exception does not release them.
The trade the batching rule is meant to prevent is paying twice for the same
work. That is not what this was. The bare word placeholder was a stub needle
scoring 0% precision — 36 of 36 findings on this repository, none a stub: the
external-ref placeholder node (ADR-0009), the redaction placeholder (ADR-0015),
S1's own sentence about not being able to tell a secret from a placeholder, a
{tag} ref template, CSS ::placeholder. The lens was reporting the codebase's
vocabulary as its debt, and roteiro check printed the inflated figure at a
glance. Replacing the word with the two phrases that predicate incompleteness of
an implementation — placeholder implementation, returns a placeholder — takes
stub from 36 to 0, the true count, and reclassifies nothing else.
Holding that behind Q2, Q10 and cross-language edge resolution would have meant shipping a knowingly false count for the whole of an unscheduled, post-v2.0 horizon — indefinitely, since none of the three has a date. Stage 26's standard is that a lens which over-reports is worse than none; a 100%-noise category is the case that standard was written for. Correctness of a number users read now outweighed the cost of one re-extraction, so the batching rule was set aside deliberately rather than forgotten. It still governs the three items above: they each add new extraction metadata, they are genuinely one body of work, and nothing about this exception makes them cheaper to do separately.
What the bump cost, measured on a store extracted at version 11 and then
opened by the version-12 binary: all 275 cached fact sets re-extracted (every
tracked blob — the base version is unconditional, so nothing survives the key
change), 3.2 s cold against 0.17 s warm on a debug build. The disk half of that
price was open-ended when this was written: the object cache was write-and-keep
with no eviction, so the superseded version-11 entries stayed on disk —
.git/roteiro/objects went 4.8 MiB → 9.7 MiB and did not shrink again, per
repository, per user, for every bump.
That half is now bounded (issue #387). sweep_superseded
(crates/rto-graph/src/sync.rs) runs at the maintenance seam and deletes the
generations a bump orphaned, keeping the current one and — by
DEFAULT_KEEP_GENERATIONS — one behind it, so switching between a branch that
bumped and the main it will merge into stays a cache hit. So a bump costs CPU
once and disk once, not disk for ever, which is what the batching rule in this
section already assumed it cost. Measured on a copy of this repository's own
cache, where four generations had accumulated (9, 10, 11 and 12): 3,812 objects
/ 71.6 MB → 1,876 / 40.0 MB in one pass, and the live set verified intact by
the only test that matters — a sync into an empty store afterwards reported 278
blobs, 0 extracted, 278 cached. Generation 12 alone is 594 objects / 10.0 MB,
which is what keep_generations: 0 leaves; the 30 MB between the two is the
retained generation 11, and is what the insurance costs. Two things are
deliberately not
reclaimed and stay a known cost: the environment tag (-e…) is a hash with no
ordering, so no tag can be shown to supersede another, and the live set itself is
still unbounded. .git/roteiro remains derived, so deleting the lot is still
always safe.
Formerly scoped-but-unrecorded, added to the roadmap by decision, and since largely delivered. Summarised here because §8b is where a reader looks for what outlives the current stage — the stages themselves carry the detail:
[models]
now takes vision, audio and ocr, so a project can pin its ASR model.
Enumeration found nine call sites, not the seven predicted — including the
extraction cache key, which folded the OCR model by name and would have made
[models] ocr the one pin that changes what is extracted without invalidating
what was extracted before it.ModelSource::Remote { trust } in the shared resolver and wired spec draft
and Ask over it. The thing 2b had to get right was not the sending but the
not sending: on a surface whose default is local there is no prompt, only
the flag, and a --allow-remote the gate refuses stops the run instead of
quietly producing a local answer.roteiro review LLM mode. 🔶 35a delivered; 35b cut in two,
PR 1 delivered. The corpus has an instrument — roteiro review --score,
per-class recall, compile_claim — and now a reviewer to point at it:
ModelTask::Review, the per-file review --llm, and the review --replay
harness. Still no verdict, and deliberately no recall figure. PR 1's two
results are that the per-file budget is far roomier than 35a concluded (median
1,758 tokens as sent against a ~30k single-call budget, so PR 2's graph arm has
space to be tested) and that the instrument could report a vacuous zero: a
reasoning model whose whole generation budget went inside <think> produced
"0 finding(s) over 4 file(s)", which would have scored as zero recall and read
as an honest negative. That is now caught and reported rather than scored.
"Do not build this" remains a legitimate outcome, and the ~15-findings-per-file
rate PR 1 measured is the evidence pointing that way.Cache bound value (Stage 25) — answered: 256 MB by default,
configurable (decided by the owner). The unit was already settled as a byte
budget following ModelCache; this fixes the number and makes it raisable for
larger repositories.
The scale that justifies it, measured on this repository: .git/roteiro is
49 MB (44 MB object cache over 2,395 entries, 4.6 MB graph.db) against a
91 MB .git. Roteiro's sidecar is already ~54% of the repository it
describes, so a cache tier is not a new cost category — it is a bound on one
that is currently unbounded in every direction. 256 MB is small against .git,
trivial against an 18 GB model store, and large enough that an ordinary session
never evicts.
Erring small is deliberate and cheap: build_context is proven to reconstruct
identically (context.rs asserts built == cached), so eviction costs cycles,
never information. Erring large only costs disk. Neither error is expensive,
which is precisely why this did not warrant more analysis than a measurement.
Memory scope (Stage 23) — answered, ADR-0013 v1.1 §Scope. A lesson
is valid in a tree only if the relevant association is present there in the
same format, so the anchor is the scope test and scope is a coarse
per-repo namespace, never a branch label. Shipped in #317; Stage 25's recall
ranks on AnchorState::applies rather than inventing a second rule.
Semantic recall (post-Stage 25): memory recall is lexical + anchor + decay in this plan. A vector index would need migration, model/dimension versioning, retention, rebuild and storage-size policy — materially more than "persist embeddings", and deferred deliberately.
code_interpreter remains rejected (ADR-0014). The sharper question behind
it — is local code execution something Roteiro wants to be? — is a product
decision, not a backend one. If it ever becomes "yes", boxlite is the vehicle and
Track A rides along; until then the answer stays "no".
Is a faster decoder worth a different completion? Answered: yes
(Stage 30, decided by the owner). Speculative decoding is measurably 1.22–1.50×
on a 27B model and measurably does not reproduce plain decoding's text. Both
halves are settled measurements; the judgement was whether Roteiro accepts the
second to get the first, and it does.
What that does not license is flipping the default. Generation was never a
reproducible surface — sampling, quantisation and the served model all move the
text already — so this changes how fast an already-variable answer arrives,
not whether Roteiro keeps a promise it was making. The graph is where
reproducibility is promised, and nothing in Stage 30 touches it. But the
remaining honest reasons to keep ROTEIRO_SPECULATIVE opt-in stand on their
own: the win is size-dependent (0.86–1.14× on a 9B — noise), it needs a
draft head that most installed models do not ship, and the identity claim has
never been observed to hold on any model. A default that helps one model class
and silently changes output on the rest is a worse default than none. Revisit
when a draft head is present on the common tier, not before.
Findings ↔ graph cross-surfacing: joining findings to graph facts is deliberately not free in this design. When it is wanted, it needs a designed join, not an implicit one.