Chapter 15: tests/fixtures, money spent once
A passing test suite proves the extension does what the tests say. It does
not prove Jev answers the way the design assumes. Some claims can only be
settled against the real model: that two layouts give the same answers, how
far repeats of one question spread, whether a label-tree search lands on the
right leaf, whether ORDER BY jev_prob … LIMIT ranks well. Those are the
release gates, and each needs live answers, which cost real money.
So each gate's live run was recorded once, on 2026-09-26, against
jev-1.13.0, by tools/record-gates --send (chapter 24), with the spend
capped by --max-cost. The four files here are those recordings: every
request exactly as sent, and every response exactly as it came back,
headers included. From then on the gate_* tests replay them through the
extension, and spend nothing.
flowchart LR R["tools/record-gates --send<br/>(once, on the live key)"] -->|"writes"| F["tests/fixtures/*.json"] F -->|"read by"| M["jev-mock in replay mode"] E["the extension, as it is today"] -->|"its requests"| M M -->|"same bytes: the recorded answer<br/>other bytes: fail, naming the request"| E
A replay matches by the exact request bytes. If a change to the extension alters a single byte of what it sends (a key order, a space, a question's wording, a lookahead the search now asks), the mock has no recorded answer for it, and the test fails naming the request. That makes these files a tripwire as much as a measurement.
Aside: what the money bought. Read layout_parity.json and you will meet four people, each asked "Could this person work from home?" on their own: Ana the backend engineer 0.88, Rui the nurse 0.27, Inês the accountant 0.81, João the bus driver 0.05. Then the same four in one request, each row inside its own question: 0.84, 0.21, 0.78, 0.06. Close, but not the same, and four rows prove nothing either way, which is why that layout stays switched off until a proper measurement says it is safe. And in ranking.json, a separate recording, Ana is 0.87 and Rui 0.28: the same bytes sent, a hundredth apart. That is why the contract calls a cached answer "one recorded draw, not a fixed point".
Try it.
nix develop -c buck2 test //tests:gate_rankingreplays the ranking gate through a throwaway Postgres, with no network at all.
For the people who maintain it
| File | What |
|---|---|
| layout_parity.json | Row as state (one request per row) and row in instructions (every row a question in one request), for the same rows and question. The extension sends only the first; the second is checked for an answer per row. |
| unsure_band.json | One question about one row; its repeats are one distinct request, so it holds one answer. |
| label_tree.json | The whole label-tree search for one row: both rounds, lookahead included, and the "does any label fit" question. |
| ranking.json | ORDER BY jev_prob … LIMIT over two rows, one request each. |
Each file is {"model", "exchanges": [{"request", "responses": [{"status", "headers", "body"}]}]}, the format tests/replay.rs pins. The rows, the
question and the SQL that produce these requests are in
../support/gates.rs.
← Previous: Chapter 14, support/ · Up: tests · Next: Chapter 16, nix/ →