ADR-0012: Analyzer findings — a separate artifact model, never a provenance class

StateAccepted
Architectural SignificanceHIGH
DomainKnowledge Graph
Document version1.3

Reference

Establishes how output from external security/quality analyzers (cargo-audit, semgrep, and successors) is persisted. Such output is not a fact extracted from source, and it does not get a provenance class: it is stored as its own analysis-run / finding artifact model, leaving derived | authored | inferred (crates/rto-graph/src/provenance.rs) untouched. Governed by the determinism and offline-first principles of ADR-0001 and the optional-capability posture of ADR-0003 and ADR-0006.

This ADR is the first of a pair. Its sibling, ADR-0013, applies the same principle to durable agent-learned knowledge. Together they state a structural rule: knowledge that is not a derived/authored/inferred graph fact gets its own artifact store, and never borrows the graph's trust.

Summary

Context

Roteiro's headline promise is provenance-tagged, deterministic, reproducible knowledge: the graph rebuilds identically from source, and AGENTS.md states derived extraction is a pure function of (path, blob id, bytes). Analyzer findings do not fit that contract:

  1. They are produced by an external tool at a point in time, against a pinned rule set and advisory database that both change independently of the source tree.
  2. They carry severity, which is a tool judgement, not the confidence score that inferred requires (migrations.rs — "a fuzzy suggestion without a score is a bug").
  3. They are re-derivable but not source-pure: re-running the same analyzer at the same commit with a newer advisory DB legitimately yields a different answer. That is a feature of security scanning and a violation of derivation.

ADR-0001 records that provenance exists precisely because the source tools had incompatible production models. Analyzer output is a fourth production model: asserted by a tool run, valid as of that run's inputs. The consistent response is a new artifact model, not a new label on an existing one.

Three tempting shortcuts were considered and rejected — see below. The most dangerous is the one that works mechanically: NodeKind::Other("security_finding") would compile and pass, while inheriting nodes.provenance NOT NULL DEFAULT 'derived' and being swept into export_factset, silently publishing tool output into an artifact that is supposed to be a pure function of the tree.

Decision makers

The Roteiro Project Team.

A separate persisted analysis-run/finding artifact model, with its own tables, its own retrieval surface, and no contact with graph provenance.

What is stored

An analysis run records the execution and everything needed to reproduce or distrust it: analyzer id and version, runner kind (sandboxed | subprocess | ingested), isolation label, image digest where applicable, rules digest, advisory-DB digest and publication date, command policy (network denied, read-only worktree, scrubbed environment), the source identity it ran against (commit / tree / lockfile blob), start and end timestamps, exit status, and a digest of the raw report.

A finding belongs to a run and carries a stable identity key so the same issue is recognisable across runs:

Findings are a replaceable layer keyed by security:<analyzer>:<worktree-id>: a successful re-run replaces the previous layer wholesale, so a fixed issue disappears rather than lingering. Note the known gap — existing import code deletes edges but not obsolete owned nodes, so owned-record cleanup is net-new work and must be implemented rather than assumed.

What is deliberately not done

Offline and degradation contract

Analyzers and their inputs are pre-provisioned, then pinned, mirroring the existing model-pull UX (roteiro model list/pull): roteiro security prefetch fetches and verifies pinned assets by digest; roteiro security status reports each digest, fetch time, and advisory-DB age.

This satisfies "mostly offline": pre-download is expected, degradation is explicit and informative, and no code path becomes quietly network-dependent.

Options considered + consequences

OptionVerdict
New Provenance::Analyzed variantRejected — broad semver/storage churn, and it dilutes a three-word vocabulary that is the project's headline promise.
Reuse Derived + metadataRejected — redefines derived from pure function of source to produced by some configured producer; a permanent conceptual smear.
NodeKind::Other("security_finding") in nodes/edgesRejected — works mechanically, which is the trap; inherits provenance DEFAULT 'derived' and leaks into export_factset.
Separate artifact model (chosen)Keeps the graph a pure function of source; gives findings the evidence chain they actually need.

Consequences

Positive

Negative / costs

Status

Accepted (2026-08-17), and implemented — Stage 21 (#293), released in v1.10.0. It was executed contract-first, as BUILD_PLAN_V2 Stage 21 sequenced it: the runner trait, the normalized finding schema, and roteiro security ingest landed before any sandboxed backend (the backend itself is ADR-0014).

Version history

VersionDateNotes
1.12026-08-17Accepted. No content changed. Status corrected: this ADR described shipped, released behaviour while still reading For Review.
1.22026-08-22A workspace vault publishes findings — deliberately, and this ADR is where that is recorded (issue #442 part 2). This document rejected NodeKind::Other("security_finding") on the grounds that it would leak into export_factset, "silently publishing tool output into an artifact". roteiro render obsidian --workspace-name now does publish it, into a different artifact, and the distinction that keeps this ADR intact is silently. The mechanism is unchanged: findings are still not nodes or edges, they still carry no provenance class, and export_factset still cannot see them — still asserted by export_factset_is_byte_identical_across_an_ingest, which this change does not touch. What changed is that one renderer reads the findings store on purpose, because the owner ruled that a hand-over document should answer "what is wrong with this workspace" as well as "what is in it". Three consequences are built rather than assumed. (1) The vault renders "an analyzer ran and reported none" and "no analyzer has ever run" as different sections, because they are opposite facts that an empty list renders identically — the failure roteiro security status records as no-analyzer-on-record, and a shareable artifact must refuse it harder than a CLI does, since its reader is the one person who cannot go and check. Coverage::NotRun is the Default, so a summary that never asked reads as unanalyzed rather than clean. (2) An analyzer's message is rendered verbatim inside a fence sized to beat any backtick run it contains, not as Markdown: it is tool output entering a document handed to people, and as Markdown it could open headings or links that restructure the note. The vault quotes it; the vault does not become it. (3) The _Home share-time section now warns instead of reassuring — it says the file lists unpatched weaknesses and where they are, and that unlike a local store it cannot be un-shared. Severity is rendered as the analyzer's own label, unrecognised levels included, because mapping one onto a known rung would invent a judgement the tool did not make.
1.32026-08-28The vault publication path is removed with the vault (issue #663). v1.2 recorded that a workspace vault publishes findings and says who was never analyzed — the Obsidian renderer is deleted in 4.0.0, so that surface no longer exists. The decision v1.2 made is untouched and still governs: coverage is reported per member, and no-analyzer-on-record is a distinct state from a clean result, because a surface that renders them identically is the vacuous zero this ADR exists to prevent. What changed is only where it is written. Not yet re-implemented: the OKF bundle carries per-member concepts but no cross-member findings summary. Recorded here rather than left implicit, because a capability that quietly stops being published is exactly what a version history is for — and the state this ADR most warns about is one nobody notices.