Chapter 9: postjevsql-core/src, six decisions and a gate
Eight files, and the best way to read them is as answers to questions the extension asks itself on the way to sending a request.
1. batch.rs: how does this row reach the model?
The planner looks at the distinct questions asked about a row and chooses
its layout. A row with two or more questions is sent as the request's
state, with each question a branch: one request, the row billed once. A row
with exactly one question could travel inside its question's instructions,
beside other rows' questions in the same request, which would be far
cheaper in requests. But nobody has measured whether Jev answers as well that
way. So that layout waits on a release gate, and until then it cannot be
built: its evidence type, LayoutParity, is an enum with no values, so no
Planner can hold one, and every row is row as state.
Aside. An enum with no values is Rust's way of writing "this cannot happen yet" so that the compiler, not a code reviewer, enforces it. There is no
ifto forget and no flag to flip by accident. The day the gate passes, someone givesLayoutParitya value, and every place that needed one starts compiling.
2. cache_key.rs: have we asked exactly this before?
sha256(model ‖ prompt version ‖ namespace ‖ layout ‖ state ‖ question),
over the exact bytes of the request, each field length-prefixed so that no
two different inputs run together into the same bytes. The question's id is
left out, because the API says it "is not sent to the underlying model": two
ids for one question are one judgment. The key is computed from the
Placement the planner returned, so the layout in the key is the one sent.
3. token_ratio.rs: how many tokens is this?
Jev's tokenizer is unpublished, and the terms forbid working it out. So the
ratio of characters to tokens is learned, per model, from what answers
reported in usage.input_tokens, falling back to a conservative 2.6
characters per token when nothing is known. A request costs 267 tokens of
fixed overhead plus its characters over the ratio, and the worst case is
three billed attempts, which is what jev.max_cost budgets.
4. gcra.rs: may this request go now?
The account's two rate limits (tokens per second, requests per minute) as
GCRA: one atomic word per limit holding a theoretical arrival time. A
request is admitted when that time is not too far in the future, and moves
it on by its cost. One compare-and-swap loop per limit, where a token bucket
would need two words and a lock. Adaptive halves the rate on a 429 and
climbs back as answers arrive.
5. label_tree.rs: which of ten thousand labels?
A Choice takes at most 255 options. For more, arrange the labels as a tree, and ask one Choice per node, best-first. A node's score is the product of the probabilities down to it, which can only fall with depth, so a partial path bounds every leaf below it, and pruning on that bound is exact. Each round asks the three best unexpanded nodes (K = 3) in one request, with lookahead (their likeliest children's Choices, when the token budget allows) and an order twin (a close call asked again in reverse order, the two averaged). Neither changes the answer, only how fast it arrives.
flowchart TD R["Root: Office · Care · Transport"] -->|"1.00"| O["Office: Engineering · Finance"] R -->|"0.00"| C["Care"] R -->|"0.00"| T["Transport"] O -->|"1.00"| E["Engineering: Backend · Frontend"] O -->|"0.00"| F["Finance"] E -->|"1.00"| B["Backend ✓"] E -->|"0.00"| FE["Frontend"]
(The recorded label-tree gate, from tests/fixtures/label_tree.json: Ana, the backend engineer, placed in two requests. The first also carried Office's and Care's Choices ahead of time, as lookahead, and the "does any label fit" question, which came back 0.77. The second settled Backend, and asked the tree's other open nodes too, because the result names its runner-up exactly, and that takes knowing every leaf that could be second.)
6. shortlist.rs: which of a thousand labels with no tree?
Two rounds at any size: ⌈N/254⌉ Choices over even chunks of your list, labels only, all in one request, keeping the top 3 of each; then one Choice over those finalists, with their descriptions. Nothing is reordered, ever. The probability returned is among the finalists, which is exactly what was measured, and not a probability over all N, which nothing measured.
And the gate: does any label fit at all?
A Choice's probabilities always add up to 1, so some label always wins, even
when none is right. gate.rs builds the yes/no question both searches send
beside their first round, "Does any of these labels fit?", whose answer does
not depend on the others and can fall near zero. The _full SQL functions
return it as fit. It never ranks labels: a Noul's probability and a
Choice's are not comparable.
Try it.
cargo test -p postjevsql-core label_treeruns the search's tests,exact_against_exhaustiveamong them: every combination of answers, searched both ways, and the best-first result must equal the exhaustive one.
For the people who maintain it
| File | What |
|---|---|
| lib.rs | The modules, and the crate's re-exports. |
| batch.rs | Planner, Planned, Placement, and the uninhabited LayoutParity. |
| cache_key.rs | CacheKey, KeyScope, Layout, PROMPT_VERSION. |
| token_ratio.rs | TokenRatio (learned, or FALLBACK), Sample, REQUEST_TOKENS (267) and BILLED_ATTEMPTS (3). |
| gcra.rs | Limit, Adaptive, Limiter, Ceilings, Rate, Admission. |
| label_tree.rs | Branch::from_paths, Tree, Search, Params (K, Tau, Lookahead, TwinBelow, Describe), Outcome. |
| shortlist.rs | Shortlist, Search, Params (keep, default 3), CHUNK (254), Outcome. |
| gate.rs | The "does any label fit" Noul, built by Tree::gate and Shortlist::gate. |
← Previous: Chapter 8, postjevsql-core/ · Up: postjevsql-core · Next: Chapter 10, postjevsql-sidecar/ →