ADR-0006: Local model serving — a llama.cpp-backed, code-aware OpenAI-compatible endpoint

StateAccepted
Architectural SignificanceHIGH
DomainDeveloper Tooling
Document version1.12

Reference

Adds an opt-in local model server so tools other than Roteiro (e.g. an Omnigent agent, an editor) can call the models a user has already pulled — offline, with no second download. Reuses the model registry and consent-gated store from ADR-0003, and wires in the graph query tools from ADR-0002 so the served model is code-aware. Configured through ADR-0007's [serve] table. Rests on the offline-first principle of ADR-0001. Introduces llama.cpp as the high-performance inference backend, evaluated against candle and mistral.rs below.

Summary

Serve local models over an opt-in, loopback, OpenAI-compatible HTTP endpoint (roteiro serve --models), backed by llama.cpp (via the Rust binding llama-cpp-2) for performance, with Roteiro's own graph tools auto-registered so the served model can query the codebase.

Three decisions:

  1. Engine: llama.cpp (llama-cpp-2), opt-in. After a head-to-head de-risk (candle vs mistral.rs vs llama.cpp on this project's MSRV 1.94 + strict cargo deny), llama.cpp is the choice: it is the fastest (Metal ~129 tok/s on a 0.6B, ~2.75× CPU), the only candidate that passes our cargo deny unchanged (46 crates, all allow-listed), and it reads a plain GGUF's embedded tokenizer + chat template for free (no tokenizer.json, no quant plumbing). The price is a C/C++ toolchain (it compiles vendored llama.cpp via cmake) — accepted, deliberately, because performance is the priority (models run often in the background and developers should not wait) and the crate tree is ~10× smaller than the alternatives.
  2. Serving layer: our own thin /v1, not the stock llama-server. llama-cpp-2 exposes inference primitives only; we hand-roll a small axum /v1 (/v1/chat/completions, /v1/embeddings, /v1/models) over load → render the model's own chat template → tokenize → decode → sample (v1.7: apply_chat_template until then, which ran no Jinja). We own the loop so we can wire Roteiro's tools into it (see #3). (A passthrough to the standalone llama-server was considered and deferred — it is a separate process and makes the tool integration external.)
  3. Code-aware by default: auto-register Roteiro's MCP graph tools. On serve, the model is handed Roteiro's ADR-0002 tools (explain / debt / path / search) via OpenAI function-calling — so a locally-served model can query this codebase's graph out of the box. This is Roteiro's differentiator: not just a model server, a graph-grounded one.

Scope remains reuse + performance, not a general model server. Loopback-bound by default; serves only installed models; never downloads. candle stays the backend for the internal uses (infer, spec draft, image vision) for now; unifying the inference core on llama.cpp is the stated direction (§Consequences) — a follow-up amendment to ADR-0003, not a big-bang.

Context

Roteiro already downloads, verifies, and stores real GGUF/safetensors models for its own use. A separate local tool that only sometimes needs a model would otherwise re-download its own copy and ship its own runtime. Serving reuses the local store so nothing new is fetched and nothing leaves the machine. Scoped to serving (v1.3). That sentence is about roteiro serve and remains true of it. It is not a project-wide guarantee: ADR-0019 adds an optional, default-off remote model tier that does send content off the machine, under a consent gate described there. Nothing in this ADR changes — serve still exposes only installed models and still never downloads — but a reader quoting this line as a general promise would now be wrong. Two things sharpened the design after the first draft:

  1. Performance is a first-class requirement. These models run frequently in the background; slow inference is a real developer-experience cost. That tilts the engine choice toward the fastest viable option, and makes candle's modest quantized-Metal decode a genuine liability rather than a footnote.
  2. A local server is only interesting if it's ours. Anyone can run Ollama or llama-server. Roteiro's reason to serve is to hand the model the codebase graph — the one query surface from ADR-0001/0002 — so the served model is code-aware. That argues for owning the request loop (to inject tools), not delegating to a black-box server.

Forces to reconcile: offline-first & self-contained (ADR-0001); don't become a general model server; honest about the C++ trade (we have held a pure-Rust preference — this opt-in feature is where we consciously relax it for performance, while the default build stays pure-Rust); and a universal interface (OpenAI API, which every calling tool already speaks).

Decision makers

llama.cpp engine + our own /v1 layer + auto-registered graph tools (recommended).

Options considered + consequences

Engine — candle vs mistral.rs vs llama.cpp (de-risked on MSRV 1.94 + strict cargo deny)

llama.cpp (llama-cpp-2) — chosenmistral.rscandle (hand-roll)
Metal tok/s (0.6B)~129 (2.75× CPU)~125slower (quant-decode ties CPU)
cargo deny✅ passes unchanged❌ fails (MPL-2.0/CDLA/0BSD core deps)✅
Crates46 (mostly build-only)465large candle tree
OpenAI serverour /v1embeddedour /v1
GGUF tokenizer/templatefree (embedded)freewe hand-code it
CostC++ (cmake/clang/libclang)C/C++ + licence waiverspure-Rust; slower; more of our code

Serving layer — our /v1 vs the stock llama-server

Surface — OpenAI REST vs MCP-only (from v1.0 of this ADR)

Consequences

Advice Received

Project direction incorporated: prioritise performance (background use; developers shouldn't wait) — so use the fastest viable engine even at the cost of C++; since we're allowing C++ bindings, llama.cpp is the pick (and it passes deny cleanly, unlike mistral.rs); use our own internal serving layer (not the stock server) so we can auto-register Roteiro's MCP/agent tools and serve a code-aware model; keep it opt-in and offline; and treat unifying the inference core on llama.cpp as the direction.

Document version history

VersionDateNotes
1.02026-08-09Accepted. Opt-in loopback OpenAI-compatible endpoint reusing installed models, warm + serialised over the ADR-0002 stack; scoped to reuse; candle-implied engine; rejected Ollama-replacement and MCP-only as the front door.
1.12026-08-09Revised after a head-to-head engine de-risk. Engine → llama.cpp (llama-cpp-2) — fastest, and the only candidate passing cargo deny unchanged (mistral.rs fails on MPL-2.0/CDLA/0BSD; candle is slower). Serving via our own thin /v1 (not stock llama-server) so Roteiro's graph tools auto-register into the model (code-aware serving). Accepts a C/C++ toolchain for the opt-in serve feature; no deny change needed. States the candle→llama.cpp inference-core unify as the direction (follow-up ADR-0003 amendment).
1.22026-08-15Consequence added: llama.cpp's backend is a process-global, initialised once and shared by every engine, so a long-lived serve process can hold more than one engine instead of silently losing every engine after the first (issue #296). No decision changed.
1.32026-08-17Scoped, not changed. "Nothing leaves the machine" is stated of serving, which is what it always described; ADR-0019 adds an optional default-off remote model tier elsewhere in the product, so the sentence needed a boundary before it was read as project-wide. No decision in this ADR changed.
1.42026-08-18HTTP/2 recorded as a non-goal rather than left as an absence: axum is taken with http1 only, HTTP/2 is a recurring DoS-advisory surface (Rapid Reset, CONTINUATION flood, RUSTSEC-2026-0258 — to which this build's exposure was nil because http2 is off), it buys nothing over a loopback bind, and where it matters it is already delegated to the reverse proxy this ADR terminates TLS at. What would overturn it is stated: a concrete client that fails over HTTP/1.1. Also corrected two defects found while editing — an inline note cited (Update, v1.5), a version this document has never had (the change it describes landed 2026-08-14 while the document was at 1.1, and was never given a history row), now labelled by its date; and the history table listed 1.3 above 1.2, now ascending.
1.52026-08-19Amended (issue #486). The context window becomes a per-request, per-model quantity rather than one hardcoded 4,096 that no configuration key could reach. Three measurements decided the shape. (a) llama.cpp allocates the KV cache eagerly in the llama_kv_cache constructor, and LlamaEngine builds a context per generation — so a large fixed window is spent on every request, including a fifty-token one: 16,466 MiB on qwen3.8-27b at its trained 262,144. (b) The served models' trained windows span 512× (262,144 for qwen3.8-27b, 512 for bge-large-en-v1.5), so no single number is correct for the set. (c) Sizing is possible because tokenisation already precedes context creation on both the text and media paths, so the count is exact rather than estimated. A request therefore gets prompt + max_tokens + headroom, floored at the old 4,096 so nothing shrinks, and capped at the model's own n_ctx_train. New [serve] max_context_tokens lowers that cap; it is a value under ADR-0007 v1.4 by that ADR's default rule — the default already grants each model's full window, so the key can only spend less of the machine, and clause 4 is never reached. A ceiling above n_ctx_train is clamped with a warning (one number spans models differing 512×); a request that does not fit is refused as a 400, never truncated. n_ubatch is unchanged at 512, and #349's finding that n_batch may follow n_ctx for free was re-measured at the larger window rather than extrapolated: +2 MiB against 8,366 MiB at n_ctx = 131,072. KV-cache quantisation is reachable (with_type_k/with_type_v) and measured at 1.85× on this model, but is not adopted — it changes generated output, and per-request sizing removes the memory pressure that would have justified it. The ceiling is load-bearing, not merely a memory setting. Under a fixed window the allocation was independent of the prompt; under per-request sizing it follows it, so wherever a client may contribute to the prompt — supplying its own tool definitions, say — an outside party has a hand in how much is allocated, at 64 KiB/token up to the trained window. The bound on that belongs at the serving edge, which refuses an oversized tool surface with a 400; the engine deliberately adds no second clamp, because two independent bounds on one quantity drift apart and then neither can be trusted. [serve] max_context_tokens is what decides the worst case a single request can reach regardless, so raising it is a decision about exposure and not only about memory.
1.62026-08-21Scoped, not changed (issue #578). v1.5's memory figures are re-stated as what they measure rather than corrected: 429 MiB / 16,466 MiB is a ps RSS delta covering KV and recurrent state, not the whole cost of a context. Re-measured on the same instrument, llama.cpp reports allocating 256 MiB KV + 149.62 MiB recurrent + 509.02 MiB Metal compute + 24.02 MiB CPU compute at n_ctx = 4,096 — 938.66 MiB against a 429 MiB delta — and a real 2,001-token decode adds only 35 MiB more, so the compute buffers are not merely waiting to be faulted in. "Metal is invisible to ps" is not the explanation either: the KV and recurrent buffers are MTL0 allocations too, and they are counted. Why the compute buffers differ is unresolved, and ps cannot answer it. The v1.5 decision is untouched: KV is what scales with n_ctx (64 KiB/token exactly on this model — 16 full-attention layers of 64, full_attention_interval = 4, x 4 KV heads x (256+256) x 2 bytes), it is allocated eagerly, and a context is built per generation, so per-request sizing remains the answer and the 512x spread across served models is unchanged. What changes is only what may be inferred from those numbers: they cannot price anything that scales the compute buffer. n_ubatch is exactly that, which is why it stays at 512 — swept at n_ctx = 4,096, 1,024 and 2,048 cost 2.000x and 4.000x the compute buffer while running 6% and 13% slower, so v1.5's "unchanged at 512" is now a measured optimum rather than a memory-driven default. See speculative::base_params and the note in tests/context_window.rs.
1.72026-08-29Amended (issue #492). The prompt is rendered from the model's own chat template, by Roteiro, and each tool is stated to the model exactly once. apply_chat_template wraps llama_chat_apply_template, which llama.h:1197 states plainly does not use a jinja parser: it substring-matches for `<
1.82026-09-01Amended (issue #592). A call form's envelope is a property of the dialect that speaks it, so Dialect::ALL means reachable as well as consistent. v1.7 recorded that "read_markup keys on that wrapper"; that was the defect. The <tool_call> wrapper was searched for one layer above Dialect and a dialect was consulted only inside a wrapper already found, so the array guaranteed that every dialect was handled consistently — parser and parity test drive from it — and guaranteed nothing about any of them being reachable. Adding an entry for a non-ChatML form would have been inert, because the miss was above where Dialect is read. Measured on voxtral-mini-3b's real embedded template, which v1.7 made renderable: it writes a call as "[TOOL_CALLS]" + name + "[ARGS]" + arguments and contains no <tool_call> anywhere, so such a model now produces a well-formed call the parser could not see and the loop hands the raw markup to the user as prose — #489's failure mode by a new route. Envelope now says where a dialect's calls begin and end and what proves one arrived; widening the old search to more literals was rejected as rebuilding the same layering one literal taller. The completeness rule is restated rather than relaxed. v1.7's rule could only be spoken in a wrapper's vocabulary — without </tool_call> the call did not arrive — and a self-delimiting form has no closing marker to be missing, so the rule becomes a call arrived only when its own dialect can prove it whole. A delimited envelope proves it with the closing marker and never with the body, for the reason that made the old rule uniform: parse_xml_body tolerates a missing </function>, so a truncated XML call reads as a whole one, a different question silently answered. A self-closing envelope must prove arrival from the body's own grammar, which is legitimate only for a grammar that rejects a truncated body — enforced by feeding every dialect in ALL every proper prefix of a call it renders itself, so declaring the lenient XML grammar self-closing fails rather than passing quietly. Dialect::Mistral lands with it so the second envelope shape is exercised in the shipped binary rather than only under #[cfg(test)]; the wrapped dialects share one envelope and so do not move.
1.92026-09-01Amended (no issue; implemented directly at the owner's request). An argument key a tool does not declare is refused, never dropped — on both tool surfaces. The two surfaces spell one of debt's arguments differently: kind on the ADR-0002 MCP surface, categories on this one's /v1 registry. Neither rejected the other's name, and the measured effect was not a missing filter but a wrong answer: sending {"categories":["todo"]} to the MCP debt deserialised to kind: [], and an empty filter means every category — so a model that asked for one kind of marker was handed the whole repository's debt, presented as the filtered set it asked for, with nothing in the result to tell the two apart. Every mistyped, hallucinated or cross-surface key has that shape, so the rule is stated over the class rather than over the one argument. Each surface enforces it through the declaration it already publishes, so the advertised schema and the behaviour cannot drift: MCP's argument structs are #[serde(deny_unknown_fields)], which serde refuses with a message naming the keys that would have worked and schemars renders as additionalProperties: false in the advertised inputSchema; the served definitions declare additionalProperties: false on the composed schema (after with_project splices the project selector in, or project itself would be refused), and tools::unknown_argument refuses against that at the one execution funnel, before the registry — which is also what keeps a route that pre-binds an argument of its own (ScopedTools fills in project) from being judged as if a model had sent it. A registry that declares nothing is unaffected, so this is a property a tool opts into rather than a rule imposed on every ToolRegistry. Verified safe against the MCP protocol before adopting: rmcp's Parameters<P> deserialises the call's arguments object and nothing else — _meta, input_responses and request_state are siblings of arguments in CallToolRequestParams, never members of it — and rmcp inserts nothing of its own, so nothing protocol-level is now rejected. Two tools that took no Parameters at all (list_projects, list_tool_classes) got rmcp's empty-input schema, which forbids nothing, and now declare empty argument structs: no argument means the empty set, not any set. The divergent spelling is deliberately left alone. Unifying kind and categories is a separate decision, and it was unsafe to take before this landed — while unknown keys were dropped, a rename would have changed existing callers' results silently instead of telling them.
1.102026-09-01Amended (no issue; implemented directly at the owner's request). The two tool surfaces now spell debt's category filter the same way: categories, on both — reversing v1.9's “the divergent spelling is deliberately left alone”, which deferred this rather than rejecting it. The MCP argument structs' kind field becomes categories; the served /v1 registry is unchanged, because it already had the better name. Three things settled the direction rather than a preference for one word: kind was already taken on the very surface that used it for this — list_kind's kind is a node kind token (fn, struct, adr, file), so one token meant two unrelated things across neighbouring tools, which is a genuine ambiguity for a model choosing arguments rather than an aesthetic complaint; the MCP doc comment already contradicted its own field name, reading “Restrict to these categories” directly above kind: Vec<String>; and the values are marker categories, which is what crates/rto-graph/src/query.rs#debt has always called that parameter. The sequencing is the substance of this row. This rename was unsafe to take before v1.9 and safe immediately after, and the difference is not a matter of degree. While an unrecognised key was dropped, a caller sending {"kind":["todo"]} would have gone on deserialising cleanly into categories: [] — and an empty filter means every category — so every existing MCP caller would have silently begun receiving the whole repository's debt as the filtered set it asked for, which is the same defect v1.9 was written to close, re-introduced by the fix for it. With the refusal in place the same call returns unknown field `kind`, expected one of `categories`, `project` as a JSON-RPC invalid_params error: the break is loud, names its own remedy, and costs a correcting client one round. That is why C was taken before B. Agreement is now guarded rather than asserted. A test that merely checked debt declares categories would pass on a tree where the surfaces had drifted apart again, so crates/rto-render/src/mcp.rs#tool_argument_names exposes each MCP tool's argument names read off the advertised inputSchema schemars derives, and roteiro's both_surfaces_name_a_shared_tools_arguments_the_same_way compares that against the served registry's declared properties for every tool present on both, failing with both spellings named. Neither side is privileged, so renaming either one alone fails it; verified by injecting each direction in turn. It is the argument-level twin of both_tool_surfaces_describe_a_tool_the_same_way, and the gap between those two tests is exactly where this defect lived: debt was advertised on both surfaces under the same name with byte-identical prose while its filter was called two different things, so every existing cross-surface check passed throughout. Deliberately not changed. The roteiro debt --kind / roteiro debt-density --kind CLI flags keep their spelling. A CLI flag is breaking in the ordinary way under AGENTS.md, it has no equivalent of the refusal that makes the tool-surface rename self-announcing, and it is a separate decision on a separate surface; the flags already carry value_name = "CATEGORY" and describe themselves as categories, so the same latent ambiguity is recorded here rather than fixed in passing. The shared tool descriptions in crates/rto-render/src/tool_text.rs still do not name the argument: #695 removed it precisely because one shared string could not name it correctly for both surfaces, and now that it could, reintroducing it would spend advertised bytes on a fact the inputSchema states beside the prose on both surfaces — the exact duplication #675's rule cuts.
1.112026-09-07Amended (no issue; implemented directly at the owner's request). The serving-edge bound on a client's tools array becomes an operator setting — new [serve] max_client_tool_bytes, default the 32,768 it was hardcoded to, so nothing moves for anyone who does not set it. v1.5 put this bound at the serving edge deliberately and gave the engine no second clamp; that division is unchanged, and this only makes the edge's number reachable. It was hardcoded because an unbounded tools array lets a caller choose the allocation, and that argument survives intact: a request still cannot raise it, and the refusal now says so and names the key rather than claiming the limit "is not raised" at all. What the argument never covered is the machine's own owner, who is not an outside party and who — running a real agent bundle against local models — is the person the 32 KiB actually binds. max_context_tokens remains the backstop that bounds the allocation directly, so where it is set the worst case is bounded whatever arrives at the edge. A capability, not a value, under ADR-0007 v1.4 — the mirror image of max_context_tokens, whose built-in already grants each model's whole window and which is therefore a value: this key's built-in denies, so a project file raising it reaches clause 4 and may only ever lower it. Raising it makes the bound reachable, not the surface affordable, and #578 lists raising it as an explicit non-goal. There is no prefix cache, so the whole advertised surface is re-prefilled every turn — measured there on qwen3.8-27b at 4.94 bytes per token and 3.13 ms per prompt token, which puts the 32 KiB default at ~6,600 tokens and ~21 s of prefill per turn, and 128 KiB at ~83 s. So a raised bound trades a hard refusal for a slow session; cutting the surface is what makes a large client affordable, and #578's prefix cache is what would make it unnecessary. The key is justified by whose decision it is — an operator could not reach their own bound — and not by the trade being a good one at every size.
1.122026-09-07Amended (issue #578). A prompt's unchanged preamble is prefilled once and restored thereafter, where v1.5 said a context is built per generation and "a window is paid for in full on every request". That is still true of the window; it is no longer true of the prompt. New [serve] prefix_cache_mb, default 0 — off — so nothing changes for anyone who does not set it. Three measurements decided the shape, and each had to hold or the feature was impossible on the primary model. (a) state_seq_get/set round-trips the recurrent half: qwen3.8-27b is hybrid, 16 of 64 layers carrying KV and 48 carrying Gated Delta Net state, and a restore that dropped the latter would be silently three-quarters wrong rather than an error. (b) A snapshot is not welded to the n_ctx it was taken at — it restores into a larger, smaller and equal window alike, all exact — which is what lets this coexist with v1.5's per-request sizing instead of forcing a choice between them; the binding's own docs say the opposite, so it was tested rather than assumed. (c) Every entry costs a fixed ~149.6 MiB plus 64 KiB/token, because the recurrent module serialises whole whatever the prompt length — so there is no cheap small entry, the budget is denominated in whole preambles, and a 512-token floor stops the cache spending 150 MiB on a prefix that cannot repay it. Recurrent state cannot be rewound (kv_cache_seq_rm reaches 16 of 64 layers), so an entry is restored whole, at position 0, or not at all: no trimming and no partial credit for a shorter shared prefix. The boundary is learned, not declared — it is the longest common prefix of two consecutive prompts, which for an agent client is exactly its system prompt and tools array — so no client cooperation is needed and prompt_cache_key stays Dropped. Measured end to end: a third turn on a 1,462-token prompt falls from 5.80 s to 2.11 s with byte-identical output. Reuse begins at the third turn, since the first two are what reveal the boundary. Inert under speculative decoding, where target and draft states are not independent.