| State | Accepted |
| Architectural Significance | HIGH |
| Domain | Developer Tooling |
| Document version | 1.4 |
Governs the embedding model used by the inference layer (Stage 8) of ADR-0001. Answers open question Q7 (the offline embedding model + binary-size budget) from docs/history/BUILD_PLAN.md.
Ship a tiny static (int8) model2vec-style embedding compiled into the binary as the default, so inference works offline with no download and adds only single-digit MB. Beyond that, models are pluggable at runtime: a small in-binary registry describes downloadable models (per-platform variants, url, sha256, licence, dimension), Roteiro selects the right variant for the host (Metal/Apple-oriented on macOS/Apple Silicon, standard GGUF elsewhere), and fetching is consent-gated ([y/N] prompt on a TTY; never in automation). The heavy inference backend (candle, to run GGUF/local models) sits behind a second feature flag so even the inference-enabled default build stays small. Chosen model file format for pluggables: GGUF.
Stage 8 emits inferred edges (doc/PDF/symbol similarity) with confidence scores. It needs a vector embedding of text. ADR-0001 mandates offline by default: the default build must work with no network and no separate model download. That collides with model size — a transformer embedding is tens-to-hundreds of MB to bundle; an API call breaks offline entirely.
Three forces to reconcile:
cargo install roteiro must stay small and never require a download to function.A single bundled model can satisfy at most two of these. Decoupling the default from what's possible resolves the tension.
Option 3 — tiny static default + pluggable local models (recommended).
model2vec-style), compiled into the binary. Pure-Rust, no runtime, fully offline, single-digit MB. Sufficient for the similarity that inferred edges need.{ name, dim, licence, variants: { "<platform>": { url, sha256, format } } }. Ships as data so pull can suggest/verify and load can checksum.target_os/target_arch:
candle with its metal feature runs ordinary weights on the Apple GPU) and the model variant (where a re-encoded build such as an mlx-community quantisation exists, the registry lists it as the macOS variant and it is preferred; otherwise standard GGUF + Metal backend).roteiro model pull <name> prints size + source + licence, then on a TTY prompts Download <name> (~N MB, <licence>) from <url>? [y/N]; on y it fetches to ~/.roteiro/models/ and verifies the sha256. In a non-TTY / CI context it defaults to No and prints the exact manual command instead — so nothing ever touches the network without an explicit human "yes".roteiro infer --model <name|path> loads a local model, falling back to the bundled static default if the named model is absent.inference — the tiny static model; offline; lean.inference-local-models — pulls in candle (+ metal on macOS), GGUF loading, the registry/fetch flow.roteiro binary; heavy even when feature-gated; still can't use the user's GPU or a different model. Rejected — violates the lean-binary force.inference build.inference-local-models (candle, GGUF, an HTTP client for pull), each subject to the cargo deny licence gate; PDF/image extraction crates (Stage 8 ingestion) are checked likewise. The default and inference builds pull none of them. (Update, v1.3: the HTTP client half no longer holds. The registry/pull machinery was split out into its own models feature, and models is now on by default — see v1.3 below. The heavy inference backend is still opt-in, which is what this bullet was protecting.)~/.roteiro/models/ becomes a user-level cache; model licences are surfaced at pull time and recorded in the registry.spec draft searched the registry, infer --model validated by hand, serve filtered the served set, serve/Ask ranked it, and the audio, vision and OCR paths held compiled-in string constants. The last of those is the part that reached users: [models] had keys for embedding and generative only, so a project could not pin its ASR model at all, and no configuration could change which model transcribed its audio. Now:
[models] grows to five keys — embedding, generative, vision, audio, ocr — one per model kind, not per command. generative governs both spec draft and Ask, because they want the same kind of model. Every key stays optional, and unset resolves to exactly the model that surface used before (qwen3-0.6b, voxtral-mini-3b, smolvlm-500m-gguf, ocrs-text, and the compiled-in hashing embedder for infer).rto_graph::model_choice::resolve, takes the task and the pins and returns the model plus the rule that chose it. The rule is not decoration: roteiro config has to answer why did it use that model? per surface, and it cannot do that from a bare string.rto-graph deliberately. That crate pins gix with default-features = false to exclude the network transports, so the code that decides which model runs structurally cannot grow a "check the hub for a newer one" call. The registry, the store and now the choice all sit behind the same wall.chat_capable_model_ids, which filters models that cannot do the job (an embedding model routed through /v1/chat/completions aborts llama.cpp with a GGML_ASSERT); the table generalises capability, and still ranks nothing.mtmd path a model of the wrong architecture does not mis-answer, it aborts the process. roteiro config is the single exception: it reports the error rather than refusing to run, because it is the command an operator reaches for when a pin is not doing what they expected.llama-cpp-2) as the serving engine — fastest, cargo deny-clean, GGUF tokenizer/template for free — and named it the target for the whole inference core. The migration is staged: generation is moved — spec draft now generates through llama.cpp (the serve feature's engine) over the same GGUFs, with the candle LocalGenerator kept only as a transitional fallback on an inference-local-models-without-serve build. Embeddings (infer) and image vision (sync) are scheduled next — embeddings move to GGUF embedding models (already served via llama.cpp, e.g. bge-small-en-v1.5-gguf), vision to mmproj (already served, e.g. smolvlm-500m-gguf); until those internal call-sites cut over, candle (inference-local-models / image-vision) remains their backend and the two coexist transitionally. End state: one llama.cpp inference core shared by serving and internal uses; candle is removed once embeddings + vision are migrated.Decision refined with the project team: (a) prefer models built/re-encoded for the host GPU architecture — Apple-oriented on macOS, standard elsewhere; (b) when the model source is known, offer an interactive Y/N fetch rather than only printing a command. Both are incorporated above.
| Version | Date | Notes |
|---|---|---|
| 1.0 | 2026-08-08 | Accepted. Tiny static int8 default compiled in; GGUF pluggable local models via an in-binary registry; platform-aware (Metal/Apple vs standard) variant selection; consent-gated fetch; candle behind inference-local-models. Answers ADR-0001 Q7. |
| 1.1 | 2026-08-09 | Amended (Stage 20) for the inference-core unification on llama.cpp (direction set in ADR-0006). Generation moved: spec draft generates via llama.cpp (the serve engine) over the same GGUFs; candle LocalGenerator is a transitional fallback only. Embeddings → GGUF and vision → mmproj scheduled next (both already served via llama.cpp); candle stays their backend until they cut over, then is removed. Adds a role label (instruct/coding/reasoning) and opt-in coding (qwen2.5-coder-3b) + reasoning (deepseek-r1-distill-qwen-1.5b) registry entries. |
| 1.2 | 2026-08-09 | Unification complete — candle removed. The engine was extracted into a shared rto-llama crate (no HTTP/async deps), and all three internal uses cut over to it: infer --model embeds via GGUF embedding models (bge-* re-listed as F16 GGUF; the safetensors all-MiniLM entry dropped), sync image understanding uses smolvlm-500m-gguf + mmproj (candle moondream removed), and spec draft generates via the shared engine directly. candle-core/nn/transformers + tokenizers are gone from the tree; inference-local-models/image-vision now mean "local llama.cpp models". One inference core, shared by serving and internal uses. |
| 1.3 | 2026-08-16 | models becomes a default feature. The registry and consent-gated pull ship in a stock cargo install roteiro; only running local models (inference-local-models, image-vision, audio-transcribe, serve) stays opt-in, so this ADR's title — "tiny static default, opt-in local models" — is unchanged in substance. The reason is that roteiro model pull is the prerequisite for the offline story ADR-0001 mandates, and gating it made that story unreachable from the shipped default: the clap variant is #[cfg(feature = "models")], so a stock install answered unrecognized subcommand rather than degrading. The consent gate is untouched — presence is not activity; nothing is fetched without an explicit [y/N] (or --yes). Cost, measured: ~2.3 MB of binary and 20 crates (ureq + rustls), pure Rust, no new host-toolchain class. serve was considered and deliberately not flipped (llama.cpp from source, a hard cmake/libclang build-script failure, and 13 unmonitored vendored advisories — see crates/roteiro/Cargo.toml). Licence disclosure in ADR-0017 v1.2. |
| 1.4 | 2026-08-17 | Amended (Stage 33) — local model resolution. [models] grows from two keys to five (vision, audio, ocr join embedding and generative), closing the gap that a project could not pin its ASR model at all: audio, vision and OCR were compiled-in string constants. One resolver in rto-graph — the crate whose gix pin excludes network transports — takes a task and the pins and returns the model plus the rule that chose it, and seven scattered call sites now ask it instead of deciding for themselves. Deterministic table over categoricals, not a classifier; it generalises chat_capable_model_ids's capability filter and ranks nothing. A pin that names an unknown model or the wrong modality is a named error quoting the key, never a silent fallback (llama.cpp aborts rather than errors on the wrong architecture). Unset is byte-identical to the previous behaviour, proven rather than asserted. No new dependency, no network, EXTRACT_VERSION unchanged at 11. Also corrects the frontmatter and header table, which both read 1.2 while the history below already carried a 1.3 row. |