postjevsql.git / tests / fixtures / README.md

Chapter 15: tests/fixtures, money spent once

A passing test suite proves the extension does what the tests say. It does not prove Jev answers the way the design assumes. Some claims can only be settled against the real model: that two layouts give the same answers, how far repeats of one question spread, whether a label-tree search lands on the right leaf, whether ORDER BY jev_prob … LIMIT ranks well. Those are the release gates, and each needs live answers, which cost real money.

So each gate's live run was recorded once, on 2026-09-26, against jev-1.13.0, by tools/record-gates --send (chapter 24), with the spend capped by --max-cost. The four files here are those recordings: every request exactly as sent, and every response exactly as it came back, headers included. From then on the gate_* tests replay them through the extension, and spend nothing.

flowchart LR
  R["tools/record-gates --send<br/>(once, on the live key)"] -->|"writes"| F["tests/fixtures/*.json"]
  F -->|"read by"| M["jev-mock in replay mode"]
  E["the extension, as it is today"] -->|"its requests"| M
  M -->|"same bytes: the recorded answer<br/>other bytes: fail, naming the request"| E

A replay matches by the exact request bytes. If a change to the extension alters a single byte of what it sends (a key order, a space, a question's wording, a lookahead the search now asks), the mock has no recorded answer for it, and the test fails naming the request. That makes these files a tripwire as much as a measurement.

Aside: what the money bought. Read layout_parity.json and you will meet four people, each asked "Could this person work from home?" on their own: Ana the backend engineer 0.88, Rui the nurse 0.27, Inês the accountant 0.81, João the bus driver 0.05. Then the same four in one request, each row inside its own question: 0.84, 0.21, 0.78, 0.06. Close, but not the same, and four rows prove nothing either way, which is why that layout stays switched off until a proper measurement says it is safe. And in ranking.json, a separate recording, Ana is 0.87 and Rui 0.28: the same bytes sent, a hundredth apart. That is why the contract calls a cached answer "one recorded draw, not a fixed point".

Try it. nix develop -c buck2 test //tests:gate_ranking replays the ranking gate through a throwaway Postgres, with no network at all.

For the people who maintain it

FileWhat
layout_parity.jsonRow as state (one request per row) and row in instructions (every row a question in one request), for the same rows and question. The extension sends only the first; the second is checked for an answer per row.
unsure_band.jsonOne question about one row; its repeats are one distinct request, so it holds one answer.
label_tree.jsonThe whole label-tree search for one row: both rounds, lookahead included, and the "does any label fit" question.
ranking.jsonORDER BY jev_prob … LIMIT over two rows, one request each.

Each file is {"model", "exchanges": [{"request", "responses": [{"status", "headers", "body"}]}]}, the format tests/replay.rs pins. The rows, the question and the SQL that produce these requests are in ../support/gates.rs.

← Previous: Chapter 14, support/ · Up: tests · Next: Chapter 16, nix/ →