| State | Accepted |
| Architectural Significance | HIGH |
| Domain | Developer Tooling |
| Document version | 1.12 |
Adds an opt-in local model server so tools other than Roteiro (e.g. an Omnigent agent, an editor) can call the models a user has already pulled — offline, with no second download. Reuses the model registry and consent-gated store from ADR-0003, and wires in the graph query tools from ADR-0002 so the served model is code-aware. Configured through ADR-0007's [serve] table. Rests on the offline-first principle of ADR-0001. Introduces llama.cpp as the high-performance inference backend, evaluated against candle and mistral.rs below.
Serve local models over an opt-in, loopback, OpenAI-compatible HTTP endpoint (roteiro serve --models), backed by llama.cpp (via the Rust binding llama-cpp-2) for performance, with Roteiro's own graph tools auto-registered so the served model can query the codebase.
Three decisions:
llama-cpp-2), opt-in. After a head-to-head de-risk (candle vs mistral.rs vs llama.cpp on this project's MSRV 1.94 + strict cargo deny), llama.cpp is the choice: it is the fastest (Metal ~129 tok/s on a 0.6B, ~2.75× CPU), the only candidate that passes our cargo deny unchanged (46 crates, all allow-listed), and it reads a plain GGUF's embedded tokenizer + chat template for free (no tokenizer.json, no quant plumbing). The price is a C/C++ toolchain (it compiles vendored llama.cpp via cmake) — accepted, deliberately, because performance is the priority (models run often in the background and developers should not wait) and the crate tree is ~10× smaller than the alternatives./v1, not the stock llama-server. llama-cpp-2 exposes inference primitives only; we hand-roll a small axum /v1 (/v1/chat/completions, /v1/embeddings, /v1/models) over load → render the model's own chat template → tokenize → decode → sample (v1.7: apply_chat_template until then, which ran no Jinja). We own the loop so we can wire Roteiro's tools into it (see #3). (A passthrough to the standalone llama-server was considered and deferred — it is a separate process and makes the tool integration external.)serve, the model is handed Roteiro's ADR-0002 tools (explain / debt / path / search) via OpenAI function-calling — so a locally-served model can query this codebase's graph out of the box. This is Roteiro's differentiator: not just a model server, a graph-grounded one.Scope remains reuse + performance, not a general model server. Loopback-bound by default; serves only installed models; never downloads. candle stays the backend for the internal uses (infer, spec draft, image vision) for now; unifying the inference core on llama.cpp is the stated direction (§Consequences) — a follow-up amendment to ADR-0003, not a big-bang.
Roteiro already downloads, verifies, and stores real GGUF/safetensors models for its own use. A separate local tool that only sometimes needs a model would otherwise re-download its own copy and ship its own runtime. Serving reuses the local store so nothing new is fetched and nothing leaves the machine. Scoped to serving (v1.3). That sentence is about roteiro serve and remains true of it. It is not a project-wide guarantee: ADR-0019 adds an optional, default-off remote model tier that does send content off the machine, under a consent gate described there. Nothing in this ADR changes — serve still exposes only installed models and still never downloads — but a reader quoting this line as a general promise would now be wrong. Two things sharpened the design after the first draft:
llama-server. Roteiro's reason to serve is to hand the model the codebase graph — the one query surface from ADR-0001/0002 — so the served model is code-aware. That argues for owning the request loop (to inject tools), not delegating to a black-box server.Forces to reconcile: offline-first & self-contained (ADR-0001); don't become a general model server; honest about the C++ trade (we have held a pure-Rust preference — this opt-in feature is where we consciously relax it for performance, while the default build stays pure-Rust); and a universal interface (OpenAI API, which every calling tool already speaks).
llama.cpp engine + our own /v1 layer + auto-registered graph tools (recommended).
CLI: roteiro serve --models [--addr 127.0.0.1:PORT] behind an opt-in serve feature (pulls llama-cpp-2). Off by default; loopback bind; warns on a non-loopback address (no auth — a localhost dev tool; TLS/authn terminate at a reverse proxy, as ADR-0002 frames for MCP). (Update, 2026-08-14: this is now simply roteiro serve [--addr …] — the model endpoint is the default for serve, and --models is a redundant deprecated flag. A serve build with no model installed, or a build without the serve feature, degrades to the llama-free /v1/graph API + web UI (ADR-0010) instead of erroring.)
Endpoints (grow across PRs): /v1/chat/completions (generation, over the same installed GGUFs), /v1/models (installed only), then /v1/embeddings. (Embeddings note: llama.cpp embeds via GGUF embedding models; our current embedders are BERT safetensors, so embedding-serving either adds GGUF embedding entries to the registry or is served via the existing candle LocalEmbedder — resolved at implementation.)
Execution: models loaded lazily on first request and kept warm in a memory-bounded LRU — the resident set is capped by a byte budget ([serve] memory_budget_mb, GGUF size as the footprint proxy) and the least-recently-used model is unloaded past it, so several models can stay warm and swap in real time on a machine that can afford them, while a memory-limited host keeps one. The context window is sized per request and bounded per model (v1.5, issue #486): llama.cpp allocates the KV cache eagerly when a context is created and a context is created per generation, so a window is paid for in full on every request whether or not it is used — measured on qwen3.8-27b, a context's KV and recurrent state cost 429 MiB at 4,096 tokens and 16,466 MiB at its trained 262,144 (v1.6 qualifies what that instrument measures). The prompt is already tokenised before the context exists, so each context is instead sized to that request's own prompt + max_tokens, floored at the 4,096 every request used to get and capped at the model's GGUF n_ctx_train. [serve] max_context_tokens lowers that cap across every served model; unset, each model offers the whole window it was trained for. Requests are serialised through the single engine mutex (so all requests are mutually exclusive, across models as well as within one) — llama.cpp batching / per-model concurrency is a later enhancement. Serves only installed models; never downloads.
Graph tools: Roteiro's MCP tools auto-registered into the served model's function-calling; the model calls a tool → Roteiro executes it against the graph → result is fed back.
HTTP/1.1 only — HTTP/2 is a non-goal. axum is taken with http1 and
without http2, so serve speaks no HTTP/2 and this is a decision rather than
an omission. Three reasons, and the first is the one that would otherwise be
rediscovered the expensive way. HTTP/2 is a recurring source of
denial-of-service advisories — Rapid Reset, CONTINUATION flood, and
RUSTSEC-2026-0258 (h2 unbounded empty DATA frames, adopted here on an
approved cooldown bypass). Roteiro's exposure to that last one was nil
precisely because http2 is off; enabling it converts that class of advisory
from theoretical to reachable, permanently. Second, over loopback it buys
nothing — multiplexing, HPACK and connection reuse pay off over
high-latency links, and the bottleneck here is model inference, not transport;
SSE streaming works over HTTP/1.1 and is universally supported. Third, where
it would matter it is already delegated: HTTP/2 in practice requires TLS
(browsers do not speak h2c), and this ADR already terminates TLS at a reverse
proxy — so HTTP/2 belongs exactly where TLS already belongs, in software whose
job is surviving hostile traffic.
What would overturn this: a concrete client that fails over HTTP/1.1. No such client is known, and this repository records no client-compatibility matrix, so the absence is unproven rather than established. If one appears, that is a requirement to design against — not a reason to enable a protocol speculatively.
Acceleration: llama.cpp's Metal backend (enabled by default on macOS in the vendored build) — this is the acceleration story for served models, and moots candle's quantized-Metal weakness on the serving path.
cargo deny)llama.cpp (llama-cpp-2) — chosen | mistral.rs | candle (hand-roll) | |
|---|---|---|---|
| Metal tok/s (0.6B) | ~129 (2.75× CPU) | ~125 | slower (quant-decode ties CPU) |
cargo deny | ✅ passes unchanged | ❌ fails (MPL-2.0/CDLA/0BSD core deps) | ✅ |
| Crates | 46 (mostly build-only) | 465 | large candle tree |
| OpenAI server | our /v1 | embedded | our /v1 |
| GGUF tokenizer/template | free (embedded) | free | we hand-code it |
| Cost | C++ (cmake/clang/libclang) | C/C++ + licence waivers | pure-Rust; slower; more of our code |
option-ext via hf-hub), CDLA-Permissive-2.0 (webpki-roots), and 0BSD (interprocess) — unavoidable, so it fails our cargo deny allow-list without waivers, and it drags 465 crates + candle 0.10 (a version split from our 0.11). Disqualified under current policy.deny-clean unchanged, minimal Rust surface, GGUF tokenizer/template for free. Trade accepted: a C/C++ build for the opt-in serve feature./v1 vs the stock llama-serverllama-server (stock) — deferred. llama-cpp-2 does not build it; using it means a separate process, and it makes wiring our graph tools external. A future optional passthrough is possible, but it is not where the Roteiro value (code-awareness) lives./v1 — chosen. ~5 primitives; we own the loop, so tool-registration is natural.serve feature pulls llama-cpp-2 — a C/C++ toolchain (cmake, clang, libclang) is required to build that feature; the default and other opt-in builds stay as they are. No cargo deny change is needed — llama-cpp-2's 46-crate tree is fully allow-listed with no advisories. This is the point at which the project consciously accepts C++ FFI for an opt-in performance path, while holding pure-Rust for the default build.explain/search/path/debt over this repo — dogfooding the one query surface (ADR-0001) for external agents.spec draft, infer), so llama.cpp — now proven fast and deny-clean — is the stated target for the whole inference core: a staged migration off candle (generation first; embeddings move to GGUF models; vision to mmproj), recorded as a follow-up amendment to ADR-0003, not a big-bang. Until then candle remains the internal backend and the two coexist only transitionally (serving reads the same GGUFs candle does, so generation stays consistent).serve shares the process's one llama.cpp backend, so a second engine is possible at all. llama.cpp's backend is a process-global and refuses a second initialisation, so while each engine initialised its own, the long-lived serve process was limited to whichever engine it built first: a second modality arriving alongside chat could not be served without a restart, and — because callers resolve an engine with .ok() — the failure surfaced as a quietly missing capability rather than an error, which is far worse in a server than in a one-shot CLI run (issue #296). The backend is now started once and handed out as an Arc by crates/rto-llama/src/backend.rs#shared_backend; the server's engine and the extractors' engines are peers on it, and it is freed only once none of them holds a handle (crates/rto-llama/src/backend.rs#release_shared_backend). This is a property of the shared engine core, so it applies identically to serve, infer --model and spec draft.[serve] sets defaults (enable, addr, which models, tool-registration on/off), overridable by CLI flags.Project direction incorporated: prioritise performance (background use; developers shouldn't wait) — so use the fastest viable engine even at the cost of C++; since we're allowing C++ bindings, llama.cpp is the pick (and it passes deny cleanly, unlike mistral.rs); use our own internal serving layer (not the stock server) so we can auto-register Roteiro's MCP/agent tools and serve a code-aware model; keep it opt-in and offline; and treat unifying the inference core on llama.cpp as the direction.
| Version | Date | Notes |
|---|---|---|
| 1.0 | 2026-08-09 | Accepted. Opt-in loopback OpenAI-compatible endpoint reusing installed models, warm + serialised over the ADR-0002 stack; scoped to reuse; candle-implied engine; rejected Ollama-replacement and MCP-only as the front door. |
| 1.1 | 2026-08-09 | Revised after a head-to-head engine de-risk. Engine → llama.cpp (llama-cpp-2) — fastest, and the only candidate passing cargo deny unchanged (mistral.rs fails on MPL-2.0/CDLA/0BSD; candle is slower). Serving via our own thin /v1 (not stock llama-server) so Roteiro's graph tools auto-register into the model (code-aware serving). Accepts a C/C++ toolchain for the opt-in serve feature; no deny change needed. States the candle→llama.cpp inference-core unify as the direction (follow-up ADR-0003 amendment). |
| 1.2 | 2026-08-15 | Consequence added: llama.cpp's backend is a process-global, initialised once and shared by every engine, so a long-lived serve process can hold more than one engine instead of silently losing every engine after the first (issue #296). No decision changed. |
| 1.3 | 2026-08-17 | Scoped, not changed. "Nothing leaves the machine" is stated of serving, which is what it always described; ADR-0019 adds an optional default-off remote model tier elsewhere in the product, so the sentence needed a boundary before it was read as project-wide. No decision in this ADR changed. |
| 1.4 | 2026-08-18 | HTTP/2 recorded as a non-goal rather than left as an absence: axum is taken with http1 only, HTTP/2 is a recurring DoS-advisory surface (Rapid Reset, CONTINUATION flood, RUSTSEC-2026-0258 — to which this build's exposure was nil because http2 is off), it buys nothing over a loopback bind, and where it matters it is already delegated to the reverse proxy this ADR terminates TLS at. What would overturn it is stated: a concrete client that fails over HTTP/1.1. Also corrected two defects found while editing — an inline note cited (Update, v1.5), a version this document has never had (the change it describes landed 2026-08-14 while the document was at 1.1, and was never given a history row), now labelled by its date; and the history table listed 1.3 above 1.2, now ascending. |
| 1.5 | 2026-08-19 | Amended (issue #486). The context window becomes a per-request, per-model quantity rather than one hardcoded 4,096 that no configuration key could reach. Three measurements decided the shape. (a) llama.cpp allocates the KV cache eagerly in the llama_kv_cache constructor, and LlamaEngine builds a context per generation — so a large fixed window is spent on every request, including a fifty-token one: 16,466 MiB on qwen3.8-27b at its trained 262,144. (b) The served models' trained windows span 512× (262,144 for qwen3.8-27b, 512 for bge-large-en-v1.5), so no single number is correct for the set. (c) Sizing is possible because tokenisation already precedes context creation on both the text and media paths, so the count is exact rather than estimated. A request therefore gets prompt + max_tokens + headroom, floored at the old 4,096 so nothing shrinks, and capped at the model's own n_ctx_train. New [serve] max_context_tokens lowers that cap; it is a value under ADR-0007 v1.4 by that ADR's default rule — the default already grants each model's full window, so the key can only spend less of the machine, and clause 4 is never reached. A ceiling above n_ctx_train is clamped with a warning (one number spans models differing 512×); a request that does not fit is refused as a 400, never truncated. n_ubatch is unchanged at 512, and #349's finding that n_batch may follow n_ctx for free was re-measured at the larger window rather than extrapolated: +2 MiB against 8,366 MiB at n_ctx = 131,072. KV-cache quantisation is reachable (with_type_k/with_type_v) and measured at 1.85× on this model, but is not adopted — it changes generated output, and per-request sizing removes the memory pressure that would have justified it. The ceiling is load-bearing, not merely a memory setting. Under a fixed window the allocation was independent of the prompt; under per-request sizing it follows it, so wherever a client may contribute to the prompt — supplying its own tool definitions, say — an outside party has a hand in how much is allocated, at 64 KiB/token up to the trained window. The bound on that belongs at the serving edge, which refuses an oversized tool surface with a 400; the engine deliberately adds no second clamp, because two independent bounds on one quantity drift apart and then neither can be trusted. [serve] max_context_tokens is what decides the worst case a single request can reach regardless, so raising it is a decision about exposure and not only about memory. |
| 1.6 | 2026-08-21 | Scoped, not changed (issue #578). v1.5's memory figures are re-stated as what they measure rather than corrected: 429 MiB / 16,466 MiB is a ps RSS delta covering KV and recurrent state, not the whole cost of a context. Re-measured on the same instrument, llama.cpp reports allocating 256 MiB KV + 149.62 MiB recurrent + 509.02 MiB Metal compute + 24.02 MiB CPU compute at n_ctx = 4,096 — 938.66 MiB against a 429 MiB delta — and a real 2,001-token decode adds only 35 MiB more, so the compute buffers are not merely waiting to be faulted in. "Metal is invisible to ps" is not the explanation either: the KV and recurrent buffers are MTL0 allocations too, and they are counted. Why the compute buffers differ is unresolved, and ps cannot answer it. The v1.5 decision is untouched: KV is what scales with n_ctx (64 KiB/token exactly on this model — 16 full-attention layers of 64, full_attention_interval = 4, x 4 KV heads x (256+256) x 2 bytes), it is allocated eagerly, and a context is built per generation, so per-request sizing remains the answer and the 512x spread across served models is unchanged. What changes is only what may be inferred from those numbers: they cannot price anything that scales the compute buffer. n_ubatch is exactly that, which is why it stays at 512 — swept at n_ctx = 4,096, 1,024 and 2,048 cost 2.000x and 4.000x the compute buffer while running 6% and 13% slower, so v1.5's "unchanged at 512" is now a measured optimum rather than a memory-driven default. See speculative::base_params and the note in tests/context_window.rs. |
| 1.7 | 2026-08-29 | Amended (issue #492). The prompt is rendered from the model's own chat template, by Roteiro, and each tool is stated to the model exactly once. apply_chat_template wraps llama_chat_apply_template, which llama.h:1197 states plainly does not use a jinja parser: it substring-matches for `< |
| 1.8 | 2026-09-01 | Amended (issue #592). A call form's envelope is a property of the dialect that speaks it, so Dialect::ALL means reachable as well as consistent. v1.7 recorded that "read_markup keys on that wrapper"; that was the defect. The <tool_call> wrapper was searched for one layer above Dialect and a dialect was consulted only inside a wrapper already found, so the array guaranteed that every dialect was handled consistently — parser and parity test drive from it — and guaranteed nothing about any of them being reachable. Adding an entry for a non-ChatML form would have been inert, because the miss was above where Dialect is read. Measured on voxtral-mini-3b's real embedded template, which v1.7 made renderable: it writes a call as "[TOOL_CALLS]" + name + "[ARGS]" + arguments and contains no <tool_call> anywhere, so such a model now produces a well-formed call the parser could not see and the loop hands the raw markup to the user as prose — #489's failure mode by a new route. Envelope now says where a dialect's calls begin and end and what proves one arrived; widening the old search to more literals was rejected as rebuilding the same layering one literal taller. The completeness rule is restated rather than relaxed. v1.7's rule could only be spoken in a wrapper's vocabulary — without </tool_call> the call did not arrive — and a self-delimiting form has no closing marker to be missing, so the rule becomes a call arrived only when its own dialect can prove it whole. A delimited envelope proves it with the closing marker and never with the body, for the reason that made the old rule uniform: parse_xml_body tolerates a missing </function>, so a truncated XML call reads as a whole one, a different question silently answered. A self-closing envelope must prove arrival from the body's own grammar, which is legitimate only for a grammar that rejects a truncated body — enforced by feeding every dialect in ALL every proper prefix of a call it renders itself, so declaring the lenient XML grammar self-closing fails rather than passing quietly. Dialect::Mistral lands with it so the second envelope shape is exercised in the shipped binary rather than only under #[cfg(test)]; the wrapped dialects share one envelope and so do not move. |
| 1.9 | 2026-09-01 | Amended (no issue; implemented directly at the owner's request). An argument key a tool does not declare is refused, never dropped — on both tool surfaces. The two surfaces spell one of debt's arguments differently: kind on the ADR-0002 MCP surface, categories on this one's /v1 registry. Neither rejected the other's name, and the measured effect was not a missing filter but a wrong answer: sending {"categories":["todo"]} to the MCP debt deserialised to kind: [], and an empty filter means every category — so a model that asked for one kind of marker was handed the whole repository's debt, presented as the filtered set it asked for, with nothing in the result to tell the two apart. Every mistyped, hallucinated or cross-surface key has that shape, so the rule is stated over the class rather than over the one argument. Each surface enforces it through the declaration it already publishes, so the advertised schema and the behaviour cannot drift: MCP's argument structs are #[serde(deny_unknown_fields)], which serde refuses with a message naming the keys that would have worked and schemars renders as additionalProperties: false in the advertised inputSchema; the served definitions declare additionalProperties: false on the composed schema (after with_project splices the project selector in, or project itself would be refused), and tools::unknown_argument refuses against that at the one execution funnel, before the registry — which is also what keeps a route that pre-binds an argument of its own (ScopedTools fills in project) from being judged as if a model had sent it. A registry that declares nothing is unaffected, so this is a property a tool opts into rather than a rule imposed on every ToolRegistry. Verified safe against the MCP protocol before adopting: rmcp's Parameters<P> deserialises the call's arguments object and nothing else — _meta, input_responses and request_state are siblings of arguments in CallToolRequestParams, never members of it — and rmcp inserts nothing of its own, so nothing protocol-level is now rejected. Two tools that took no Parameters at all (list_projects, list_tool_classes) got rmcp's empty-input schema, which forbids nothing, and now declare empty argument structs: no argument means the empty set, not any set. The divergent spelling is deliberately left alone. Unifying kind and categories is a separate decision, and it was unsafe to take before this landed — while unknown keys were dropped, a rename would have changed existing callers' results silently instead of telling them. |
| 1.10 | 2026-09-01 | Amended (no issue; implemented directly at the owner's request). The two tool surfaces now spell debt's category filter the same way: categories, on both — reversing v1.9's “the divergent spelling is deliberately left alone”, which deferred this rather than rejecting it. The MCP argument structs' kind field becomes categories; the served /v1 registry is unchanged, because it already had the better name. Three things settled the direction rather than a preference for one word: kind was already taken on the very surface that used it for this — list_kind's kind is a node kind token (fn, struct, adr, file), so one token meant two unrelated things across neighbouring tools, which is a genuine ambiguity for a model choosing arguments rather than an aesthetic complaint; the MCP doc comment already contradicted its own field name, reading “Restrict to these categories” directly above kind: Vec<String>; and the values are marker categories, which is what crates/rto-graph/src/query.rs#debt has always called that parameter. The sequencing is the substance of this row. This rename was unsafe to take before v1.9 and safe immediately after, and the difference is not a matter of degree. While an unrecognised key was dropped, a caller sending {"kind":["todo"]} would have gone on deserialising cleanly into categories: [] — and an empty filter means every category — so every existing MCP caller would have silently begun receiving the whole repository's debt as the filtered set it asked for, which is the same defect v1.9 was written to close, re-introduced by the fix for it. With the refusal in place the same call returns unknown field `kind`, expected one of `categories`, `project` as a JSON-RPC invalid_params error: the break is loud, names its own remedy, and costs a correcting client one round. That is why C was taken before B. Agreement is now guarded rather than asserted. A test that merely checked debt declares categories would pass on a tree where the surfaces had drifted apart again, so crates/rto-render/src/mcp.rs#tool_argument_names exposes each MCP tool's argument names read off the advertised inputSchema schemars derives, and roteiro's both_surfaces_name_a_shared_tools_arguments_the_same_way compares that against the served registry's declared properties for every tool present on both, failing with both spellings named. Neither side is privileged, so renaming either one alone fails it; verified by injecting each direction in turn. It is the argument-level twin of both_tool_surfaces_describe_a_tool_the_same_way, and the gap between those two tests is exactly where this defect lived: debt was advertised on both surfaces under the same name with byte-identical prose while its filter was called two different things, so every existing cross-surface check passed throughout. Deliberately not changed. The roteiro debt --kind / roteiro debt-density --kind CLI flags keep their spelling. A CLI flag is breaking in the ordinary way under AGENTS.md, it has no equivalent of the refusal that makes the tool-surface rename self-announcing, and it is a separate decision on a separate surface; the flags already carry value_name = "CATEGORY" and describe themselves as categories, so the same latent ambiguity is recorded here rather than fixed in passing. The shared tool descriptions in crates/rto-render/src/tool_text.rs still do not name the argument: #695 removed it precisely because one shared string could not name it correctly for both surfaces, and now that it could, reintroducing it would spend advertised bytes on a fact the inputSchema states beside the prose on both surfaces — the exact duplication #675's rule cuts. |
| 1.11 | 2026-09-07 | Amended (no issue; implemented directly at the owner's request). The serving-edge bound on a client's tools array becomes an operator setting — new [serve] max_client_tool_bytes, default the 32,768 it was hardcoded to, so nothing moves for anyone who does not set it. v1.5 put this bound at the serving edge deliberately and gave the engine no second clamp; that division is unchanged, and this only makes the edge's number reachable. It was hardcoded because an unbounded tools array lets a caller choose the allocation, and that argument survives intact: a request still cannot raise it, and the refusal now says so and names the key rather than claiming the limit "is not raised" at all. What the argument never covered is the machine's own owner, who is not an outside party and who — running a real agent bundle against local models — is the person the 32 KiB actually binds. max_context_tokens remains the backstop that bounds the allocation directly, so where it is set the worst case is bounded whatever arrives at the edge. A capability, not a value, under ADR-0007 v1.4 — the mirror image of max_context_tokens, whose built-in already grants each model's whole window and which is therefore a value: this key's built-in denies, so a project file raising it reaches clause 4 and may only ever lower it. Raising it makes the bound reachable, not the surface affordable, and #578 lists raising it as an explicit non-goal. There is no prefix cache, so the whole advertised surface is re-prefilled every turn — measured there on qwen3.8-27b at 4.94 bytes per token and 3.13 ms per prompt token, which puts the 32 KiB default at ~6,600 tokens and ~21 s of prefill per turn, and 128 KiB at ~83 s. So a raised bound trades a hard refusal for a slow session; cutting the surface is what makes a large client affordable, and #578's prefix cache is what would make it unnecessary. The key is justified by whose decision it is — an operator could not reach their own bound — and not by the trade being a good one at every size. |
| 1.12 | 2026-09-07 | Amended (issue #578). A prompt's unchanged preamble is prefilled once and restored thereafter, where v1.5 said a context is built per generation and "a window is paid for in full on every request". That is still true of the window; it is no longer true of the prompt. New [serve] prefix_cache_mb, default 0 — off — so nothing changes for anyone who does not set it. Three measurements decided the shape, and each had to hold or the feature was impossible on the primary model. (a) state_seq_get/set round-trips the recurrent half: qwen3.8-27b is hybrid, 16 of 64 layers carrying KV and 48 carrying Gated Delta Net state, and a restore that dropped the latter would be silently three-quarters wrong rather than an error. (b) A snapshot is not welded to the n_ctx it was taken at — it restores into a larger, smaller and equal window alike, all exact — which is what lets this coexist with v1.5's per-request sizing instead of forcing a choice between them; the binding's own docs say the opposite, so it was tested rather than assumed. (c) Every entry costs a fixed ~149.6 MiB plus 64 KiB/token, because the recurrent module serialises whole whatever the prompt length — so there is no cheap small entry, the budget is denominated in whole preambles, and a 512-token floor stops the cache spending 150 MiB on a prefix that cannot repay it. Recurrent state cannot be rewound (kv_cache_seq_rm reaches 16 of 64 layers), so an entry is restored whole, at position 0, or not at all: no trimming and no partial credit for a shorter shared prefix. The boundary is learned, not declared — it is the longest common prefix of two consecutive prompts, which for an agent client is exactly its system prompt and tools array — so no client cooperation is needed and prompt_cache_key stays Dropped. Measured end to end: a third turn on a 1,462-token prompt falls from 5.80 s to 2.11 s with byte-identical output. Reuse begins at the third turn, since the first two are what reveal the boundary. Inert under speculative decoding, where target and draft states are not independent. |