ADR-0003: Pluggable embedding models — tiny static default, opt-in local models

StateAccepted
Architectural SignificanceHIGH
DomainDeveloper Tooling
Document version1.4

Reference

Governs the embedding model used by the inference layer (Stage 8) of ADR-0001. Answers open question Q7 (the offline embedding model + binary-size budget) from docs/history/BUILD_PLAN.md.

Summary

Ship a tiny static (int8) model2vec-style embedding compiled into the binary as the default, so inference works offline with no download and adds only single-digit MB. Beyond that, models are pluggable at runtime: a small in-binary registry describes downloadable models (per-platform variants, url, sha256, licence, dimension), Roteiro selects the right variant for the host (Metal/Apple-oriented on macOS/Apple Silicon, standard GGUF elsewhere), and fetching is consent-gated ([y/N] prompt on a TTY; never in automation). The heavy inference backend (candle, to run GGUF/local models) sits behind a second feature flag so even the inference-enabled default build stays small. Chosen model file format for pluggables: GGUF.

Context

Stage 8 emits inferred edges (doc/PDF/symbol similarity) with confidence scores. It needs a vector embedding of text. ADR-0001 mandates offline by default: the default build must work with no network and no separate model download. That collides with model size — a transformer embedding is tens-to-hundreds of MB to bundle; an API call breaks offline entirely.

Three forces to reconcile:

  1. Offline-by-default & lean binary — cargo install roteiro must stay small and never require a download to function.
  2. Quality ceiling — some users want better embeddings than a tiny static model gives, and want to use hardware acceleration (Apple GPU / CUDA) with larger local models.
  3. No silent network — fetching a model is a network action a human must authorise, per ADR-0001.

A single bundled model can satisfy at most two of these. Decoupling the default from what's possible resolves the tension.

Decision makers

Option 3 — tiny static default + pluggable local models (recommended).

Options considered + consequences

Option 1: Single bundled transformer model

Option 2: No bundled model; always fetch

Consequences

Advice Received

Decision refined with the project team: (a) prefer models built/re-encoded for the host GPU architecture — Apple-oriented on macOS, standard elsewhere; (b) when the model source is known, offer an interactive Y/N fetch rather than only printing a command. Both are incorporated above.

Document version history

VersionDateNotes
1.02026-08-08Accepted. Tiny static int8 default compiled in; GGUF pluggable local models via an in-binary registry; platform-aware (Metal/Apple vs standard) variant selection; consent-gated fetch; candle behind inference-local-models. Answers ADR-0001 Q7.
1.12026-08-09Amended (Stage 20) for the inference-core unification on llama.cpp (direction set in ADR-0006). Generation moved: spec draft generates via llama.cpp (the serve engine) over the same GGUFs; candle LocalGenerator is a transitional fallback only. Embeddings → GGUF and vision → mmproj scheduled next (both already served via llama.cpp); candle stays their backend until they cut over, then is removed. Adds a role label (instruct/coding/reasoning) and opt-in coding (qwen2.5-coder-3b) + reasoning (deepseek-r1-distill-qwen-1.5b) registry entries.
1.22026-08-09Unification complete — candle removed. The engine was extracted into a shared rto-llama crate (no HTTP/async deps), and all three internal uses cut over to it: infer --model embeds via GGUF embedding models (bge-* re-listed as F16 GGUF; the safetensors all-MiniLM entry dropped), sync image understanding uses smolvlm-500m-gguf + mmproj (candle moondream removed), and spec draft generates via the shared engine directly. candle-core/nn/transformers + tokenizers are gone from the tree; inference-local-models/image-vision now mean "local llama.cpp models". One inference core, shared by serving and internal uses.
1.32026-08-16models becomes a default feature. The registry and consent-gated pull ship in a stock cargo install roteiro; only running local models (inference-local-models, image-vision, audio-transcribe, serve) stays opt-in, so this ADR's title — "tiny static default, opt-in local models" — is unchanged in substance. The reason is that roteiro model pull is the prerequisite for the offline story ADR-0001 mandates, and gating it made that story unreachable from the shipped default: the clap variant is #[cfg(feature = "models")], so a stock install answered unrecognized subcommand rather than degrading. The consent gate is untouched — presence is not activity; nothing is fetched without an explicit [y/N] (or --yes). Cost, measured: ~2.3 MB of binary and 20 crates (ureq + rustls), pure Rust, no new host-toolchain class. serve was considered and deliberately not flipped (llama.cpp from source, a hard cmake/libclang build-script failure, and 13 unmonitored vendored advisories — see crates/roteiro/Cargo.toml). Licence disclosure in ADR-0017 v1.2.
1.42026-08-17Amended (Stage 33) — local model resolution. [models] grows from two keys to five (vision, audio, ocr join embedding and generative), closing the gap that a project could not pin its ASR model at all: audio, vision and OCR were compiled-in string constants. One resolver in rto-graph — the crate whose gix pin excludes network transports — takes a task and the pins and returns the model plus the rule that chose it, and seven scattered call sites now ask it instead of deciding for themselves. Deterministic table over categoricals, not a classifier; it generalises chat_capable_model_ids's capability filter and ranks nothing. A pin that names an unknown model or the wrong modality is a named error quoting the key, never a silent fallback (llama.cpp aborts rather than errors on the wrong architecture). Unset is byte-identical to the previous behaviour, proven rather than asserted. No new dependency, no network, EXTRACT_VERSION unchanged at 11. Also corrects the frontmatter and header table, which both read 1.2 while the history below already carried a 1.3 row.