uteke v0.19.0: an API that describes itself, and memory that keeps its history

uteke v0.19.0 re-anchors its benchmark headline to 98.2% from committed raw results, and ships a self-describing API, audit-grade memory history, live index repair, and budgeted context packing.

uteke v0.19.0: an API that describes itself, and memory that keeps its history

We published a 98.2% retrieval score today, and the raw per-question results for it are sitting in the repo. Not a screenshot. Not a claim. A JSONL file with 500 ranked answer lists, next to the raw files from the two releases before it. If you doubt the number, ~20 lines of Python and one minute get you from "trust me" to your own count.

That release discipline is the context for everything else in v0.19.0. The theme of this one is not retrieval quality (the strict numbers have not moved across three releases, and we will show you exactly what did move). The theme is honesty in the surfaces around retrieval: the API can now describe itself, memory history survives deletion, a degraded index can be audited and repaired live, and the remember command tells you when it could not do the full job.

The benchmark, anchored to the release

Full 500-question LongMemEval-S run on the v0.19.0 tree (image built from the exact release SHA, 10-way Modal fan-out, strategy default; a tree diff against the v0.19.0 release tag shows only version-string bumps, so the run measures the released binary):

Metricv0.16.0v0.17.0v0.19.0
recall_any@598.2%98.4%98.2%
recall_any@1098.8%98.8%98.8%
strict recall_all@588.0%88.0%88.0%
strict recall_all@1095.4%95.4%95.4%

Compare v0.19.0 against v0.17.0 question by question and you get 250 identical rankings, zero flips at @10 and @50, and exactly one flip at @5: a two-gold-session question where one gold session moved from rank 5 to rank 6. That is a near-tie boundary crossing, the same class of noise we documented in the cross-architecture reproduction (107 of 108 questions with identical rankings on different CPUs). We think a headline that survives recomputation is worth more than the best single number in the band, so 98.2% is the anchor and the full table ships next to it.

GET /routes: the API describes itself

The failure this closes was embarrassing and common: an agent reads stale docs, concludes a feature does not exist, and files a confident "missing feature" report. We had seen the GET-vs-POST and wrong-param-name class of mismatch more than once.

uteke-serve now exposes GET /routes, generated from the same registry that generates the API reference. Every route carries its method, path, tier (CORE or LAB), description, and request/response types. A typed client can validate itself against the live server instead of a README snapshot. If the registry lies, the docs gate in CI goes red, so the introspection endpoint cannot drift from the reference.

Memory that keeps its history

Two changes make the store audit-grade.

First, the timeline no longer dies with the row. Before v0.19.0, the timeline_events table had an ON DELETE CASCADE on memory ID, so hard-deleting a memory erased the record that it ever existed. For anything doing compliance or forensics on agent behavior, that is the wrong shape. The table is rebuilt without the cascade (schema v20), forget() writes a forgot tombstone before the delete, deprecate_with_reason records its reason, and the migration backfills created events for older memories. There is a new MCP tool, uteke_timeline, that reads the trail, tombstones included. The owner's framing when this was decided: the timeline is the audit layer, history is never deleted.

Second, remember stops lying by omission. When embedding generation fails after retries, the memory is still stored in SQLite and FTS5, so it is keyword-searchable but invisible to vector recall. Older versions either failed hard or passed silently. Now every surface (CLI, HTTP, MCP) returns embedding_written plus a warning, and uteke repair rebuilds the vector entries later. Silent partial success is the worst failure mode for a memory system, because you find out months later, by absence.

An index that heals itself

The stale-vector-index failure motivated both changes above: rows exist in SQLite and FTS5, but the vector index missed them, so memories went missing under fusion and hybrid recall until someone restarted the server into a rebuild. v0.19.0 turns that into a runtime operation. POST /verify compares the SQLite row count against the vector index size and reports a match flag. POST /repair rebuilds the index from stored embeddings. Both are also MCP tools. No restart, no downtime, and the runbook after any server upgrade is now: verify, then repair if mismatched.

Context that fits a budget

Agents rarely want a ranked list; they want the best content that fits the prompt budget. uteke recall --pack takes the same fusion ranking and greedily fills a character budget, rank order intact, returning an envelope: which items were selected, which were skipped and why (caller-excluded IDs or budget), and how much budget was used. It is deterministic and LLM-free. It is a selection primitive, not a re-ranker; the retrieval quality numbers above are untouched by it. It ships across CLI flags, HTTP POST /recall, and the MCP tool with the same parameters.

The correctness fixes worth knowing about

Three namespace fixes landed. doc list honors --namespace now on CLI, HTTP, and MCP (it accepted one nowhere, despite the help text advertising it). Unified recall with --type all no longer floods scoped results with cross-namespace documents, and its document scores report the raw hybrid sum instead of a rank-derived fake 1.000 that outranked every real memory. And GET /room/memories now honors the namespace query param; before, it parsed it and threw it away, returning a 200 with the room's full cross-namespace list. The issue reproduction was blunt: 72 entries returned when the requested namespace held 19. A wrong answer with a success status code is the kind of bug you build infrastructure to catch, which is what the CORE/LAB contract gate added this release is for.

Run it yourself

# verify the committed results (no embedder needed, ~1 min)
cd benchmarks/longmemeval
python3 - <<'EOF'
import json
rows = {}
for line in open("results/default-500q-v0.19.0.jsonl"):
    r = json.loads(line)
    rows[r["question_id"]] = r["retrieval_results"]["retrieved_session_ids"]
ds = json.load(open("data/longmemeval_s_cleaned.json"))
gold = {q["question_id"]: set(q["answer_session_ids"]) for q in ds}
non_abs = [q for q in rows if "_abs" not in q]
any5 = sum(1 for q in non_abs if gold[q] & set(rows[q][:5]))
print(f"recall_any@5 = {any5}/{len(non_abs)} = {any5/len(non_abs):.4f}")
EOF

Expect 462/470 on the answerable basis, 0.9820 on the full 500. Full instructions, including a ~$10 full re-run on Modal, are in benchmarks/longmemeval/REPRODUCING.md.

Upgrading: uteke upgrade --yes, or pin with UTEKE_VERSION=v0.19.0. After upgrading a running uteke-serve, run POST /verify and POST /repair if needed; that is now the supported path instead of a restart.