ADR-0019: Remote model tier — an explicitly consented egress path, and the promises it changes

StateAccepted
Architectural SignificanceVERY HIGH
DomainInference
Document version1.2

Reference

Decides whether Roteiro may call a hosted model, and under what conditions. It is the prerequisite for Build Plan V2 Stage 34 and it exists because that stage cannot be built without changing promises made elsewhere: ADR-0006 on what leaves the machine, ADR-0007 on how configuration layers, and ADR-0003 on consent for model acquisition. Storage of any remote output is governed by ADR-0015, whose producer identity this ADR qualifies.

This is the first capability in Roteiro that sends repository content off the machine. That is the whole of its significance, and it is why this is an ADR and not a feature flag.

Summary

Roteiro may call a hosted model, as an optional, default-off capability, but only through a gate the user has deliberately opened. Three specific promises are amended rather than quietly outgrown, and the amendments are the substance of this decision:

  1. ADR-0006's "nothing leaves the machine" is scoped, explicitly, to serving.
  2. ADR-0007's precedence inverts for one key: the committed project file may deny but never grant egress.
  3. Principle 10 is exempted for this one capability, because a remote call can be neither digest-pinned nor prefetched.

And one thing is decided that looks like a routing question and is not: the local→remote edge is a consent gate, not a model-selection decision.

Context

Every model Roteiro uses today is a local single-file GGUF, pinned by URL and SHA-256, acquired once with explicit consent (ADR-0003) and thereafter used offline. roteiro serve exposes only installed models and never downloads. rto-graph — where extraction and the graph live — structurally cannot reach the network: gix is pinned default-features = false specifically to exclude transports, and the model downloader takes its transport as a caller-supplied closure.

That is a coherent position, and this ADR does not abandon it. It carves a single, named exception and says exactly what it costs.

A terminology collision to avoid first

"Online mode" is already taken, and means nearly the opposite of what this capability does. The website defines it as:

Online mode — richer inference with local models. Pull a real embedding or generative model once (with consent), then run everything locally — the "online" is a one-time, explicit download, after which inference is offline again.

Reusing that phrase for "may call a hosted API at use time" would make the existing documentation actively misleading. This capability is called the remote model tier, and the words remote or hosted are used throughout. "Online mode" keeps its existing meaning.

Decision makers

The Roteiro Project Team.

Adopt an optional remote model tier, default-off, under the conditions below.

The cost asymmetry is not symmetric. Mis-routing among local models wastes tokens. Mis-routing outward sends source off the machine for a reason nobody can inspect afterwards.

Therefore consent must not be probabilistic. A classifier may not decide that a request leaves the machine, at any model quality. Model resolution may be as clever as it likes among local models (Stage 33); the edge is a boolean the user opened.

This also disposes of the "route to a frontier model when local can't help" framing. That sounds like a routing rule and is actually a consent boundary wearing one. If escalation is wanted, it must be a deterministic, recorded check of the local result — empty output, no tool call after MAX_ROUNDS, below a length floor — evaluated after the local attempt, with the measured value recorded. Never a prediction made before it.

2. Reachability must not be probed

Online-ness is not observable, and must not be made observable. A reachability probe is egress: a DNS lookup leaks the query to a resolver, and doing it to decide whether egress is permitted inverts the gate.

The correct proxy is policy, not measurement — the user said yes. That is also the deterministic answer, where a probe is not.

ADR-0007 establishes:

CLI flag > project roteiro.toml > user ~/.roteiro/config.toml > built-in default.

For every other key that is right. For this one it is inverted, and the inversion is the point: roteiro.toml is committed and shared by design — the ADR's own words are "committed — so a team shares the same, reproducible settings". A merged line in a shared file authorising egress on every teammate's machine is not consent; it is consent by pull request, granted by someone else, noticed by nobody.

So, for the remote-enable key only:

LayerMay denyMay grant
Built-in defaultdenied by default—
Project roteiro.tomlyesno
User ~/.roteiro/config.tomlyesyes — necessary, not sufficient
Invocation (flag, or a TTY prompt)yesyes — necessary, not sufficient

Both the user layer and the invocation must grant. Neither alone suffices. The user layer opts the human in; the invocation opts the run in.

A project may still switch it off for everyone — a locked-down repository is a legitimate thing to express, and denial has none of the problems of grant.

What "the invocation" means for a long-lived process (v1.2)

For a one-shot command the invocation is the command. For roteiro serve it is the server process: serve --allow-remote grants every Ask request handled for the life of that process, not one request at a time.

Decided deliberately rather than inherited. The alternative — re-asking per request — cannot work: there is no human at an HTTP request to ask, so a per-request grant would either be a config value (which is the user layer again, not an invocation) or a prompt nobody is present to answer. A gate that cannot be operated is not a stricter gate, it is a broken one.

State the exposure honestly, because it is real. A server started with --allow-remote and left running sends graph-derived context to the hosted model for every Ask it answers, for as long as it runs. That is a materially different profile from a one-shot remote call, and the person starting the server is consenting on behalf of every later request to it — including requests made by someone else who can reach the port.

Three things bound it, and all three are already required elsewhere in this ADR:

What this does not license: a remote grant surviving the process, being persisted anywhere, or being inferred from a previous session. It is scoped to one process and dies with it.

This deviation must be stated in ADR-0007 itself, not only here, because a reader of that ADR will otherwise apply the general rule and be wrong.

4. What may be sent — decided here, not deferred

Deferring this is how an egress path ships before its guard.

What the graph holds is narrower than it looks. Function bodies are not stored. meta.content is capped at 1,500 characters and is populated only from prose files, doc-comments, PDF/OCR text and audio summaries; extraction asserts that a .rs file node carries no meta.content. So a graph-derived prompt carries symbol names, headings and topics — not source text.

But there is no redaction chokepoint on a prompt. Extraction redacts secret-looking config values before persistence precisely because the store is exportable. That mechanism does not apply here, and it is weaker than it sounds even where it does apply: is_secret_key matches key names only, against ten needles (secret, password, passwd, passphrase, token, apikey, credential, privatekey, accesskey, pwd), with no inspection of values. DATABASE_URL=postgres://user:hunter2@host matches none of them.

Therefore:

5. Remote output is not a graph fact

It is not a pure function of (path, blob id, bytes), so it acquires no node, no edge, no Provenance variant, and never appears in export_factset. This is the fourth instance of the rule already applied to analyzer findings (ADR-0012), agent memory (ADR-0013) and generated media (ADR-0015).

One thing does not transfer, and must be handled. ADR-0015's Producer is not a label but a verifiable identity, folding a model_digest as pinned in the registry into a canonical id — which is what makes re-describing with different weights a new record rather than a silent overwrite. A hosted model has no digest that can be computed. A vendor model string is a mutable pointer: the weights behind it can change while the name does not, and Roteiro cannot detect it.

So if remote output is ever persisted, it carries an explicit ProducerTrust::{PinnedDigest, VendorAsserted}, and a VendorAsserted record states on its face that its identity is a claim. Doing that honestly is worth more than the feature.

6. Principle 10 is exempted, explicitly

Build Plan V2 principle 10 reads:

Offline-capable, not "offline". Optional capabilities may require pre-provisioned assets; they must be digest-pinned, explicitly prefetched, and must fail with a named, actionable error rather than fetching implicitly or silently degrading.

A remote call is fetching by definition. It cannot be digest-pinned and cannot be prefetched, so this is the first capability the principle can only exempt, never satisfy.

The second half still binds, and binds harder: with the tier enabled and no network, Roteiro must fail with a named, actionable error naming the endpoint — never silently degrade, and never fall back to a local model without saying so. An unannounced downgrade is the failure mode this ADR most needs to prevent, because it produces a different answer with no signal that anything changed.

Options considered + consequences

OptionVerdict
Adopt, gated as aboveRecommended. Capability without a silent egress path.
Adopt with a learned router deciding the edgeRejected. Consent must not be probabilistic, and a weight vector cannot answer why did this leave?
Adopt with project-level grant (normal ADR-0007 precedence)Rejected. Authorises egress for a whole team from a merged line.
Do not adoptViable. Costs nothing already promised; the local tier is unaffected.

Consequences

Status

Accepted (2026-08-17), and unbuilt — a departure from this repository's habit, worth stating rather than leaving to be noticed. Every ADR accepted before this one was accepted alongside working code. This is a decision about a capability that does not exist yet, accepted so Stage 34 has a settled contract to build against rather than discovering its consent model halfway through.

What that means in practice: the decision is not open for re-litigation, but nothing here has been proved by an implementation. Where building Stage 34 shows a clause to be unworkable, that is an amendment to this ADR with a version-history row — never a quiet deviation in code.

Version history

VersionDateChange
1.02026-08-17Initial. Written to unblock Stage 34.
1.12026-08-17Accepted. No content changed; Stage 34 is unblocked.
1.22026-08-17Scoped "the invocation" for long-lived processes: serve --allow-remote grants for the life of the server process, decided by the owner. Records why a per-request grant is unworkable (nobody is present at an HTTP request to ask), states the exposure plainly, and names the three existing bounds. No change to the grant/deny table.