README.mdpreviewREADME.mdsource116 lines · 6.1 KB · raw

Chapter 9: postjevsql-core/src, six decisions and a gate

Eight files, and the best way to read them is as answers to questions the extension asks itself on the way to sending a request.

1. batch.rs: how does this row reach the model?

The planner looks at the distinct questions asked about a row and chooses its layout. A row with two or more questions is sent as the request's state, with each question a branch: one request, the row billed once. A row with exactly one question could travel inside its question's instructions, beside other rows' questions in the same request, which would be far cheaper in requests. But nobody has measured whether Jev answers as well that way. So that layout waits on a release gate, and until then it cannot be built: its evidence type, LayoutParity, is an enum with no values, so no Planner can hold one, and every row is row as state.

Aside. An enum with no values is Rust's way of writing "this cannot happen yet" so that the compiler, not a code reviewer, enforces it. There is no if to forget and no flag to flip by accident. The day the gate passes, someone gives LayoutParity a value, and every place that needed one starts compiling.

2. cache_key.rs: have we asked exactly this before?

sha256(model ‖ prompt version ‖ namespace ‖ layout ‖ state ‖ question), over the exact bytes of the request, each field length-prefixed so that no two different inputs run together into the same bytes. The question's id is left out, because the API says it "is not sent to the underlying model": two ids for one question are one judgment. The key is computed from the Placement the planner returned, so the layout in the key is the one sent.

3. token_ratio.rs: how many tokens is this?

Jev's tokenizer is unpublished, and the terms forbid working it out. So the ratio of characters to tokens is learned, per model, from what answers reported in usage.input_tokens, falling back to a conservative 2.6 characters per token when nothing is known. A request costs 267 tokens of fixed overhead plus its characters over the ratio, and the worst case is three billed attempts, which is what jev.max_cost budgets.

4. gcra.rs: may this request go now?

The account's two rate limits (tokens per second, requests per minute) as GCRA: one atomic word per limit holding a theoretical arrival time. A request is admitted when that time is not too far in the future, and moves it on by its cost. One compare-and-swap loop per limit, where a token bucket would need two words and a lock. Adaptive halves the rate on a 429 and climbs back as answers arrive.

5. label_tree.rs: which of ten thousand labels?

A Choice takes at most 255 options. For more, arrange the labels as a tree, and ask one Choice per node, best-first. A node's score is the product of the probabilities down to it, which can only fall with depth, so a partial path bounds every leaf below it, and pruning on that bound is exact. Each round asks the three best unexpanded nodes (K = 3) in one request, with lookahead (their likeliest children's Choices, when the token budget allows) and an order twin (a close call asked again in reverse order, the two averaged). Neither changes the answer, only how fast it arrives.

flowchart TD
  R["Root: Office · Care · Transport"] -->|"1.00"| O["Office: Engineering · Finance"]
  R -->|"0.00"| C["Care"]
  R -->|"0.00"| T["Transport"]
  O -->|"1.00"| E["Engineering: Backend · Frontend"]
  O -->|"0.00"| F["Finance"]
  E -->|"1.00"| B["Backend ✓"]
  E -->|"0.00"| FE["Frontend"]

(The recorded label-tree gate, from tests/fixtures/label_tree.json: Ana, the backend engineer, placed in two requests. The first also carried Office's and Care's Choices ahead of time, as lookahead, and the "does any label fit" question, which came back 0.77. The second settled Backend, and asked the tree's other open nodes too, because the result names its runner-up exactly, and that takes knowing every leaf that could be second.)

6. shortlist.rs: which of a thousand labels with no tree?

Two rounds at any size: ⌈N/254⌉ Choices over even chunks of your list, labels only, all in one request, keeping the top 3 of each; then one Choice over those finalists, with their descriptions. Nothing is reordered, ever. The probability returned is among the finalists, which is exactly what was measured, and not a probability over all N, which nothing measured.

And the gate: does any label fit at all?

A Choice's probabilities always add up to 1, so some label always wins, even when none is right. gate.rs builds the yes/no question both searches send beside their first round, "Does any of these labels fit?", whose answer does not depend on the others and can fall near zero. The _full SQL functions return it as fit. It never ranks labels: a Noul's probability and a Choice's are not comparable.

Try it. cargo test -p postjevsql-core label_tree runs the search's tests, exact_against_exhaustive among them: every combination of answers, searched both ways, and the best-first result must equal the exhaustive one.

For the people who maintain it

FileWhat
lib.rsThe modules, and the crate's re-exports.
batch.rsPlanner, Planned, Placement, and the uninhabited LayoutParity.
cache_key.rsCacheKey, KeyScope, Layout, PROMPT_VERSION.
token_ratio.rsTokenRatio (learned, or FALLBACK), Sample, REQUEST_TOKENS (267) and BILLED_ATTEMPTS (3).
gcra.rsLimit, Adaptive, Limiter, Ceilings, Rate, Admission.
label_tree.rsBranch::from_paths, Tree, Search, Params (K, Tau, Lookahead, TwinBelow, Describe), Outcome.
shortlist.rsShortlist, Search, Params (keep, default 3), CHUNK (254), Outcome.
gate.rsThe "does any label fit" Noul, built by Tree::gate and Shortlist::gate.

← Previous: Chapter 8, postjevsql-core/ · Up: postjevsql-core · Next: Chapter 10, postjevsql-sidecar/ →