ADR-0016: Audio metadata as derived facts — symphonia for format reads, and MPL-2.0 in the licence allowlist

StateAccepted
Architectural SignificanceMEDIUM
DomainKnowledge Graph
Document version1.2

Reference

Adds deterministic audio metadata — codec, sample rate, bit depth, channel layout, duration, tags — to the graph as ordinary derived facts, read from the container without decoding audio and without any model. Adopts symphonia for the format read, and therefore admits MPL-2.0 to the deny.toml licence allowlist.

This is the deliberate complement to ADR-0015: 0015 removed generated content (transcripts, descriptions) from derived because it is invented; this ADR adds extracted content because it is present in the bytes. Together they draw the line in both directions rather than only one.

Governed by ADR-0001; extends the ingestion story of ADR-0005.

Summary

Context

Today an audio blob is close to opaque in the graph: a file node with a path and a size. Everything Roteiro knew about its content came from an ASR model — and ADR-0015 correctly moved that out of derived, because a transcript is generated rather than extracted.

That left an asymmetry worth fixing. A .wav file genuinely contains facts: it is 16 kHz, it is mono, it is 256 ms long, and those statements are exactly as deterministic as "this Rust file declares fn extract". They satisfy the derived contract in full — same bytes, same answer, no clock, no sampling, no model — and the graph currently records none of them.

Reading them needs a container parser. Three routes were assessed:

  1. Reuse llama.cpp's bundled miniaudio — impossible. It is compiled #define MA_API static, so every symbol has internal linkage; nm over the built libmtmd.a finds no ma_decoder symbols at all, llama-cpp-sys-2's bindgen allowlist covers only ggml_*/gguf_*/llama_*/mtmd_*, and the one public entry point requires a live projector context — the very multi-gigabyte load we are trying to avoid.
  2. Hand-roll the parsers — viable for silence detection (see below), but for metadata it means MP3 frame-header walking plus Xing/VBRI, FLAC STREAMINFO, RIFF chunk walking, and then ID3v2 and Vorbis comment parsing. Tag parsers in particular are where hand-written code quietly goes wrong.
  3. Take a decoding library for its format layer — symphonia is the only maintained candidate; it was previously excluded not on merit but because MPL-2.0 is absent from the allowlist.

Decision makers

The Roteiro Project Team.

What becomes a derived fact

Per audio blob, from a format read only:

These are ordinary graph facts: cached by blob id like every other extraction, sorted deterministically, and carrying no clock or environment input beyond the media env tag already folded into the cache key.

Tag formats actually reached (v1.1)

Two corrections from implementing this, both recorded rather than quietly absorbed. Version 1.0 listed the tag formats aspirationally; this is what the chosen dependency actually delivers.

FormatReached?Why
ID3v1 / ID3v2 (MP3)yes, with id3v1/id3v2not implied by mp3
APE (MP3)yes, with apenot implied by mp3
Vorbis comments (FLAC)yesembedded in the container; arrives with flac
RIFF INFO (WAV)no — upstream bugparsed, then discarded (below)
Vorbis comments (OGG)noogg is not enabled; .ogg is not is_audio
MP4 atomsnoisomp4 is not enabled; .m4a is not is_audio

ID3v1/ID3v2/APE need their own feature flags. They are standalone metadata readers in symphonia-metadata, registered on the probe separately from any container. With features = ["flac", "mp3", "wav"] alone, register_enabled_formats registers no metadata reader at all, and every tag on an MP3 — the format most likely to carry one — is silently invisible. The three flags are therefore added. They pull no additional package: all three gate code inside symphonia-metadata, which flac and wav already require.

WAV RIFF INFO is lost inside symphonia 0.6.1. WavReader::try_new parses a LIST/INFO chunk into a local metadata binding, then constructs itself from opts.external_data.metadata.unwrap_or_default() instead — so the parsed revision is dropped. (symphonia-bundle-flac does the corresponding thing correctly, which is why FLAC's Vorbis comments do arrive.) Recovering them would mean re-parsing the RIFF chunk list ourselves, i.e. the hand-rolling this ADR declined; so the limitation is accepted, and pinned by a test that fails when a future symphonia fixes it — at which point this table should be revised. WAV stream facts are unaffected: rate, depth, channels and an exact duration all read correctly.

The two nos at the bottom of the table are consequences of the deliberately narrow scope below, not new decisions: neither .ogg nor .m4a is an extension is_audio admits, so no blob in scope can carry either.

Exact vs estimated — the one subtlety

MP3 duration is not always exact. It comes from a Xing/VBRI header when one is present, and is otherwise inferred from bitrate — and only when the source is seekable. On the degenerate 1 040-byte fixture the reader states no frame count at all, so there is no duration to record.

That is still extraction (deterministic given the bytes, no generation), so it stays derived. But a derived fact that is deterministic yet inexact is new here, and must not be laundered into a precise-looking number:

Asserting an approximate number under the graph's strongest provenance label would repeat, in miniature, the mistake ADR-0015 exists to correct.

Every MPEG-audio duration is marked estimated (v1.1) — including one whose frame count came from a Xing header. MpaReader reaches a frame count by three routes (Xing, VBRI, or an inference from bitrate and stream length) and exposes no flag saying which; distinguishing them would mean parsing the first MPEG frame for a Xing/Info/VBRI header ourselves, which is the hand-rolling this ADR rejected. A Xing count is in any case the encoder's claim about the stream rather than a property of it, unlike a WAV data chunk length or FLAC's STREAMINFO total-samples field, which are the two cases marked exact. The only error possible here is calling an estimate exact, and this construction cannot make it.

The absence is per fact, not per blob (v1.1). The degenerate MP3 still yields codec, sample rate and channel count — what its frame headers state outright — and simply carries no duration key. A blob the reader rejects outright yields no node at all. Both are tested.

Where the facts live (v1.1)

Each audio blob's facts become their own node — kind audio_stream, key audio:<path>, hung off the file node by a derived contains edge — rather than extra keys on the file node's meta. This is the shape config_key (ADR-0009) and image_ref already use, and it is chosen for three reasons:

AudioDuration is a struct ({ms, exactness}), not a bare number beside a qualifier — the same "sum type, not a nullable column" discipline MediaOutcome uses. No consumer can reach the milliseconds without the marker, which turns "every surface shows an estimate as an estimate" into a property of the type rather than a standing review obligation.

Licence: MPL-2.0 admitted, deliberately

All seven symphonia-* crates are MPL-2.0; every transitive dependency (lazy_static, log, bitflags, bytemuck, smallvec, num-complex, num-traits, regex-lite, extended) already falls inside the existing allowlist. So MPL-2.0 is the sole addition.

MPL-2.0 is file-level weak copyleft: it obliges publication of modifications to symphonia's own files, plus notice and source availability when binaries are shipped. It does not reach Roteiro's own source, and leaves the project's MIT/ Apache-2.0 dual licence untouched. Using the crate unmodified — the expected case — reduces the obligation to keeping the notice.

The allowlist entry must carry a comment saying this, so a future reader sees a decision rather than an oversight.

Scope — deliberately narrow

In: the format read and the facts above.

Out, with reasons:

Options considered + consequences

OptionVerdict
Reuse llama.cpp's miniaudioImpossible — static linkage, no exported symbols, public path needs a projector context.
Hand-roll metadata parsersRejected — tag parsing is disproportionate risk for the value.
claxon + puremp3Rejected — clear the licence gate but are abandoned (2020 and a 2019 0.1.0).
symphonia, decode features offChosen — small, modular, safe-Rust, and the only maintained option.
Do nothingLeaves the graph blind to facts it could hold for microseconds of work.

Consequences

Positive

Negative / costs

Status

Accepted (2026-08-17), and implemented — Stage 29 (#316), released in v1.11.0. Implemented in the same PR that carries this ADR, per the decision to land the rationale and the code together.

Document version history

VersionDateNotes
1.02026-08-15Audio metadata as derived facts via a symphonia format read; MPL-2.0 admitted to the deny.toml allow-list with a recorded rationale; duration carries an exact/estimated marker and absence is recorded as absence; scope narrowed to exclude ASR decoding, widening is_audio, cross-container duplicates and silence detection.
1.12026-08-15Implementation corrections, no decision changed. Feature list: id3v1/id3v2/ape added — they are standalone metadata readers not implied by mp3, so without them the probe registers no metadata reader and every MP3 tag is invisible; they add no package. RIFF INFO: recorded as unreachable in symphonia 0.6.1 (the WAV reader parses the LIST/INFO chunk and then discards it), pinned by a test that fails when upstream fixes it. Exactness: every MPEG-audio duration is estimated, Xing-backed ones included, because the reader exposes no flag saying which route it took. Absence: clarified as per-fact — the degenerate MP3 still yields codec, rate and channels. Placement: facts live on an audio_stream node keyed audio:<path> under a contains edge, not on the file node's meta, so ADR-0015's emptied meta.content slot stays empty. Package count corrected (16 locked, not 17).
1.22026-08-17Accepted. No content changed. Status corrected: this ADR described shipped, released behaviour while still reading For Review.