ADR-0005: Image ingestion — tiered OCR (pure-Rust) + optional vision understanding

StateAccepted
Architectural SignificanceMEDIUM
DomainDeveloper Tooling
Document version1.4

Reference

Governs the image ingestion item of Stage 12 (see docs/history/BUILD_PLAN.md) — the last remaining piece of "make inferred edges meaningful by embedding real content". It extends the content-ingestion pattern already shipped for prose and PDF text (their bodies land in meta.content and are embedded), and reuses the tiered, offline-first, consent-gated model machinery decided in ADR-0003. It answers the item ADR-0001 and the Stage 12 plan flagged as "the one genuinely uncertain item in the backlog." See ADR-0001.

Summary

Ingest text and meaning from images (screenshots, scanned docs, diagrams) into meta.content, so inference relates an image to code by content — exactly as prose and PDF already do. Structure it in two tiers by purpose, not one spliced pipeline:

Both are feature-gated and opt-in; the default and inference builds pull neither. This mirrors the tiering Roteiro already uses for inference (hashing → local GGUF → larger → agent) and authoring (ADR-0004).

We reject all three surveyed OCR crates (rusto-rs, yingkitw/ocr, oar-ocr) and we explicitly reject splicing Tier B's model into Tier A's pipeline (using candle-TrOCR as the recognizer behind rten detection).

Context

Stage 12 already embeds real content for prose and PDF (meta.content → node_text → embedding). Images are the last modality. The Stage 12 plan called this out as needing "its own decision (like candle)": pure-Rust OCR is historically weak, while accurate OCR usually means a C/C++ inference engine (tesseract, ONNX Runtime, MNN) — against the pure-Rust stance — or a heavy vision model with large weights.

Forces to reconcile (from ADR-0001 and the project's stance):

  1. Pure-Rust, no C/C++ FFI (strong preference). unsafe_code = "forbid" in our crates; the build must stay cargo-only with no system libraries to install. A bundled ONNX Runtime / tesseract binary is the thing to avoid.
  2. Offline-by-default & lean binary. Any model is opt-in and local, pulled through the ADR-0003 consent gate; the default roteiro binary must not grow, and must never make a network call implicitly.
  3. cargo deny licence gate. Every transitive dependency must satisfy the allow-list (MIT, Apache-2.0, BSD-2/3, Unicode-3.0, Zlib, ISC).
  4. MSRV 1.96. Any new dependency tree must build on 1.96.
  5. Correctness over fluency. OCR text is derived-ish content fed to inference; a VLM description is clearly a suggestion. Neither should masquerade as authored fact.

Three community Rust OCR projects were surveyed as candidates (the user's shortlist):

Decision makers

Option 4 — two-tier, pure-Rust-default image ingestion (recommended).

Go/no-go de-risk (completed before this ADR was accepted)

A throwaway spike gated Tier A against the invariants (ocrs 0.12.2, rten 0.25.0, image 0.25.10):

GateResult
MSRV 1.94 build✅ compiles clean
Pure-Rust (no FFI)✅ no ONNX Runtime / tesseract / *-sys in the tree
cargo deny advisories✅
cargo deny licences✅ with minimal image codecs (see below)
Tree size73 crates (≈ the pdf-text feature's 71), opt-in

Options considered + consequences

Option 1: rusto-rs (MNN engine)

Option 2: yingkitw/ocr (pure-Rust)

Option 3: oar-ocr (ONNX Runtime + candle VLM)

Option 3b: Splice — rten detection + candle-TrOCR recognition (one pipeline)

Consequences

Advice Received

Project direction incorporated above: keep the pure-Rust / no-C++-FFI stance (reject the ONNX-Runtime and MNN engines despite their maturity); prefer reusing the existing candle + ADR-0003 machinery for anything model-based; and — from the review of combining ocrs/rten with candle — treat candle as a separate understanding tier, not as a recognizer spliced into the OCR pipeline, because splicing combines the runtimes' costs rather than their benefits.

Document version history

VersionDateNotes
1.02026-08-09Accepted. Two-tier image ingestion: Tier A pure-Rust OCR (ocrs/rten, feature image-ocr) as the default text tier; Tier B optional candle document-VLM understanding (feature image-vision) reusing ADR-0003. Rejects rusto-rs (MNN C++), yingkitw/ocr (immature/unvalidated), oar-ocr (ONNX Runtime C++), and the spliced rten+candle-TrOCR pipeline. Go/no-go spike passed: MSRV 1.94 build, no FFI, cargo deny clean at ~73 crates — provided image is pinned to minimal codecs (default features pull an AVIF→libfuzzer-sys NCSA chain).
1.12026-08-15Consequence added: the shared vision/audio engine must be released before process exit — a static-cached engine is never dropped and aborts a Metal build in ggml-metal's exit-time teardown (issue #291). No decision changed.
1.22026-08-15Consequence added: the llama.cpp backend is a process-global, so it is initialised once and shared by every engine — a second engine used to fail to construct and go silently inert (issue #296). Release ordering (engines, then backend) is now enforced by Arc ownership rather than by call order. No decision changed.
1.32026-08-15Consequence added: the mtmd projector is cached per (loaded model, mmproj path) instead of being re-read per blob — 688 MB for the audio projector (issue #301), with the measured effect (initialisations N→1 per projector, ~⅓ less kernel CPU, wall-clock unchanged on a page-cache-warm host) recorded beside it. Records why both halves of the key are required (two live projectors are not interchangeable; an mtmd context holds the llama_model * it was built over) and that projectors join the existing ownership-ordered teardown rather than a parallel one. No decision changed.
1.42026-09-01Force 4 restated to MSRV 1.96, tracking the raise recorded in ADR-0001 v1.4 (driver: the OKF conformance stack's oxc_* crates). That clause is a restatement of ADR-0001's standing policy — it is not an ADR-0005 decision — so it moves with its source. The 1.94 figures elsewhere in this ADR are left as written: the go/no-go spike table, the options analysis and the 1.0 history row record what was measured in August 2026 on the toolchain in force at the time, and rewriting them would falsify the evidence rather than update it. No decision in this ADR changes; the ocrs/rten tree is unaffected.