| State | Accepted |
| Architectural Significance | MEDIUM |
| Domain | Developer Tooling |
| Document version | 1.4 |
Governs the image ingestion item of Stage 12 (see docs/history/BUILD_PLAN.md) — the last remaining piece of "make inferred edges meaningful by embedding real content". It extends the content-ingestion pattern already shipped for prose and PDF text (their bodies land in meta.content and are embedded), and reuses the tiered, offline-first, consent-gated model machinery decided in ADR-0003. It answers the item ADR-0001 and the Stage 12 plan flagged as "the one genuinely uncertain item in the backlog." See ADR-0001.
Ingest text and meaning from images (screenshots, scanned docs, diagrams) into meta.content, so inference relates an image to code by content — exactly as prose and PDF already do. Structure it in two tiers by purpose, not one spliced pipeline:
ocrs + rten. A detection→recognition OCR engine running on rten, a pure-Rust model runtime (no ONNX Runtime, no tesseract, no C/C++ FFI). Extracts literal text from an image into meta.content. This is the common case and keeps the pure-Rust stance (Principle 3, ADR-0001).candle. For images where literal OCR is insufficient (architecture diagrams, charts), a vision-language model — run on the candle stack we already vendor — produces a semantic description. Weights flow through ADR-0003's registry / consent-gated pull, exactly like the embedding and generative model tiers.Both are feature-gated and opt-in; the default and inference builds pull neither. This mirrors the tiering Roteiro already uses for inference (hashing → local GGUF → larger → agent) and authoring (ADR-0004).
We reject all three surveyed OCR crates (rusto-rs, yingkitw/ocr, oar-ocr) and we explicitly reject splicing Tier B's model into Tier A's pipeline (using candle-TrOCR as the recognizer behind rten detection).
Stage 12 already embeds real content for prose and PDF (meta.content → node_text → embedding). Images are the last modality. The Stage 12 plan called this out as needing "its own decision (like candle)": pure-Rust OCR is historically weak, while accurate OCR usually means a C/C++ inference engine (tesseract, ONNX Runtime, MNN) — against the pure-Rust stance — or a heavy vision model with large weights.
Forces to reconcile (from ADR-0001 and the project's stance):
unsafe_code = "forbid" in our crates; the build must stay cargo-only with no system libraries to install. A bundled ONNX Runtime / tesseract binary is the thing to avoid.roteiro binary must not grow, and must never make a network call implicitly.cargo deny licence gate. Every transitive dependency must satisfy the allow-list (MIT, Apache-2.0, BSD-2/3, Unicode-3.0, Zlib, ISC).derived-ish content fed to inference; a VLM description is clearly a suggestion. Neither should masquerade as authored fact.Three community Rust OCR projects were surveyed as candidates (the user's shortlist):
rusto-rs — uses the MNN C++ inference engine + PaddleOCR models converted to MNN. Accurate, but a native C++ engine and an offline model-conversion step. Immature (v0.1.2).yingkitw/ocr — the only pure-Rust one (ndarray), Apache-2.0, bundled models. But very immature (1★) and its accurate CRNN path is explicitly unvalidated ("blocked on hardware"); the default is pattern-matching for clean printed text only. Pulls tokio.oar-ocr — the most mature (142★, crates.io), but built on ONNX Runtime (ort, a bundled C++ binary) for its classic pipeline, with a candle VLM path alongside. The ONNX Runtime native dependency is the disqualifier.Option 4 — two-tier, pure-Rust-default image ingestion (recommended).
ocrs + rten (feature image-ocr, off by default): detect + recognise text, cap/whitespace-collapse into meta.content, embedded like prose/PDF. Panic-guarded and size-capped like the PDF path (ADR-adjacent to the pdf-text feature). Models are .rten files pulled through the ADR-0003 registry/consent gate into the model store, not ocrs's default ~/.cache auto-fetch — so the offline + consent invariants hold.candle document VLM (feature image-vision, off by default, implies inference-local-models): for richer understanding, produce a description into meta.content. Reuses ADR-0003's registry / platform-variant / consent pull / candle backend — adding a vision model kind alongside the embedding and generative kinds.image dependency is pinned to minimal codecs — default-features = false, features = ["png", "jpeg"] (± webp) — see Consequences for why this is load-bearing.A throwaway spike gated Tier A against the invariants (ocrs 0.12.2, rten 0.25.0, image 0.25.10):
| Gate | Result |
|---|---|
| MSRV 1.94 build | ✅ compiles clean |
| Pure-Rust (no FFI) | ✅ no ONNX Runtime / tesseract / *-sys in the tree |
cargo deny advisories | ✅ |
cargo deny licences | ✅ with minimal image codecs (see below) |
| Tree size | 73 crates (≈ the pdf-text feature's 71), opt-in |
rusto-rs (MNN engine)yingkitw/ocr (pure-Rust)tokio. Too immature to depend on. Rejected (it is the immature end of the very category Tier A occupies more maturely).oar-ocr (ONNX Runtime + candle VLM)ort + a bundled/downloaded C++ binary) — a native dependency and an implicit binary download, against the pure-Rust and offline stances. Adopting it drags ort into the tree even if only the VLM path is used. Rejected.rten detection + candle-TrOCR recognition (one pipeline)candle stack we already vendor for the recognizer.rten and candle) → double the deny/MSRV surface, larger binary, two model formats; ocrs already ships a recognizer tuned to its own detector (bridging the seam risks accuracy); and TrOCR is autoregressive-per-line and CPU-only here, far slower than rten's CTC recognizer. Explicitly rejected — candle earns its place as a separate understanding tier, not as the recognizer inside the OCR pipeline.cargo deny, and is a modest opt-in tree (~73 crates) — validated by the spike. Tier B adds genuine image understanding by reusing ADR-0003 machinery, and only when opted in. Matches Roteiro's established tiering; the default build is unchanged. ocrs/rten are the most mature pure-Rust OCR available (1.9k★, actively maintained), far ahead of Option 2.ocrs is Latin-script/early-preview (mitigated: it is opt-in and degrades to "no content", never blocks sync; non-Latin/handwriting is out of scope for v1); routing model downloads through our consent gate rather than ocrs's default cache is a small integration cost; Tier B weights are large (mitigated: separate opt-in feature + consent gate, and Tier A needs none).image dependency must be pinned to minimal codecs (default-features = false, features = ["png", "jpeg"]). With default features, image pulls the AVIF encoder chain ravif → rav1e → libfuzzer-sys, which carries the NCSA licence — not on our cargo deny allow-list. Pinning to the codecs OCR actually needs drops that chain entirely: the spike went from a licence rejection at 144 crates to a clean advisories ok, bans ok, licenses ok, sources ok at 73. No deny policy change is required. This constraint is recorded here because it is easy to reintroduce by adding a bare image dependency.image-ocr (Tier A, ocrs/rten) and image-vision (Tier B, candle VLM, implies inference-local-models). The default, inference, and pdf-text builds pull neither. Each is subject to the same deny gate and MSRV as every dependency.extract's content path beside prose and PDF: a decode/OCR failure degrades an image to a plain file node (no meta.content), never aborting sync. EXTRACT_VERSION gains a distinct namespace when image-ocr is enabled, so an OCR build and a default build never serve each other stale image facts from the shared cache (as pdf-text already does)..rten models and Tier B's VLM weights are registered in the ADR-0003 registry and pulled with consent into the model store — no implicit network fetch, offline-first preserved.inferred similarity edges (clearly labelled, confidence-scored) and participates in semantic dedup and the context cache. It is never authored fact.statics, so holding the engine in a static OnceLock left that set non-empty when libc's C++ finalizers tore ggml-metal down at exit(); ggml_metal_rsets_free asserted and aborted a successful run with SIGABRT — exit 134 for any subcommand that described at least one image (issue #291). Extraction therefore keeps each engine in a releasable slot that crates/rto-graph/src/extract.rs#release_media_engines empties, and the CLI owns that teardown for the length of a run via crates/rto-graph/src/extract.rs#MediaEngineGuard. Recorded here, like the codec pin above, because parking the engine in a static is the obvious-looking way to write this and silently reintroduces the abort.LlamaBackend::init() refuses a second call while a first backend is alive. Each engine used to initialise its own, so in a build with both media features the second engine a run needed failed to construct — and because the extractors resolve an engine with .ok(), the second modality was silently inert rather than reported (issue #296). The backend is therefore started once and shared: engines hold an Arc handle from crates/rto-llama/src/backend.rs#shared_backend rather than a backend of their own. That also carries the release ordering above out of the engine struct without losing it — llama.cpp frees models before the backend, and crates/rto-llama/src/backend.rs#release_shared_backend simply declines while any engine still holds a handle, so "engines first, backend last" is a property of ownership rather than of call order. Both live in the same build-once/release-deterministically holder, crates/rto-llama/src/slot.rs#EngineSlot. Recorded here beside the release consequence because per-engine backend initialisation is the obvious-looking way to write this, and its failure mode is a missing capability rather than an error.mtmd context built on top of it was not: chat_media called MtmdContext::init_from_file per blob, so a sync re-read the mmproj GGUF once per media file — 688 MB for the audio projector, plus a fresh clip context and its GPU buffers each time (issue #301). It is now built once and reused, keyed by the pair (loaded model, mmproj path), in crates/rto-llama/src/llama.rs#LlamaEngine::projector over crates/rto-llama/src/slot.rs#KeyedSlot. Both halves of that key are load-bearing. The path, because since #296 a vision and an audio projector are live at the same time and are not interchangeable — an unkeyed slot would hand an audio request the vision context, which reports support_audio() == false. The model, because mtmd_init_from_file keeps the llama_model * it is given and dereferences it on every tokenize and eval: a projector reused across models, or outliving a model the residency cache has evicted, is a use-after-free rather than merely a stale answer. That half is expressed structurally — a projector lives in the crates/rto-llama/src/llama.rs#Loaded entry of its own model and carries an Arc on it — which is also what keeps the teardown ordering above intact now that there is a third native object in it: projectors die with their model, models with their engine, the engine before the backend, all by ownership rather than by call order. Recorded here because "cache it in a static, keyed by the mmproj path" is the obvious-looking way to write this and is unsound in both of the ways this key exists to prevent. Measured effect, so the expectation is on record with the decision: over a six-clip sync the initialisations collapse from one per blob to one per projector and kernel CPU falls by about a third, but wall-clock is unchanged on a host where the mmaped projector stays in the page cache — the ~5 s per clip observed in #299 is that clip's encode-and-generate cost, not the projector load.image-ocr) is the committed deliverable and ships first, since it is the default/common case and the gate passed. Tier B (image-vision) is designed here but sequenced after, as an opt-in enrichment.Project direction incorporated above: keep the pure-Rust / no-C++-FFI stance (reject the ONNX-Runtime and MNN engines despite their maturity); prefer reusing the existing candle + ADR-0003 machinery for anything model-based; and — from the review of combining ocrs/rten with candle — treat candle as a separate understanding tier, not as a recognizer spliced into the OCR pipeline, because splicing combines the runtimes' costs rather than their benefits.
| Version | Date | Notes |
|---|---|---|
| 1.0 | 2026-08-09 | Accepted. Two-tier image ingestion: Tier A pure-Rust OCR (ocrs/rten, feature image-ocr) as the default text tier; Tier B optional candle document-VLM understanding (feature image-vision) reusing ADR-0003. Rejects rusto-rs (MNN C++), yingkitw/ocr (immature/unvalidated), oar-ocr (ONNX Runtime C++), and the spliced rten+candle-TrOCR pipeline. Go/no-go spike passed: MSRV 1.94 build, no FFI, cargo deny clean at ~73 crates — provided image is pinned to minimal codecs (default features pull an AVIF→libfuzzer-sys NCSA chain). |
| 1.1 | 2026-08-15 | Consequence added: the shared vision/audio engine must be released before process exit — a static-cached engine is never dropped and aborts a Metal build in ggml-metal's exit-time teardown (issue #291). No decision changed. |
| 1.2 | 2026-08-15 | Consequence added: the llama.cpp backend is a process-global, so it is initialised once and shared by every engine — a second engine used to fail to construct and go silently inert (issue #296). Release ordering (engines, then backend) is now enforced by Arc ownership rather than by call order. No decision changed. |
| 1.3 | 2026-08-15 | Consequence added: the mtmd projector is cached per (loaded model, mmproj path) instead of being re-read per blob — 688 MB for the audio projector (issue #301), with the measured effect (initialisations N→1 per projector, ~⅓ less kernel CPU, wall-clock unchanged on a page-cache-warm host) recorded beside it. Records why both halves of the key are required (two live projectors are not interchangeable; an mtmd context holds the llama_model * it was built over) and that projectors join the existing ownership-ordered teardown rather than a parallel one. No decision changed. |
| 1.4 | 2026-09-01 | Force 4 restated to MSRV 1.96, tracking the raise recorded in ADR-0001 v1.4 (driver: the OKF conformance stack's oxc_* crates). That clause is a restatement of ADR-0001's standing policy — it is not an ADR-0005 decision — so it moves with its source. The 1.94 figures elsewhere in this ADR are left as written: the go/no-go spike table, the options analysis and the 1.0 history row record what was measured in August 2026 on the toolchain in force at the time, and rewriting them would falsify the evidence rather than update it. No decision in this ADR changes; the ocrs/rten tree is unaffected. |