For agents, on top of README.md, which they read first.
Design contract
Each rule below names the project it comes from: pg-jev, jevQL or
pg_typesafe. Rules marked new are this repo's own. Evidence and
the full comparison are in the brain at
~/brains/personal/technology/artificial-intelligence/jev/postjevsql.md
and the ecosystem/ pages next to it. When a rule changes, update that
page too.
SQL surface
-
Functions (pg-jev / jevQL names, kept compatible):
jev(row, q [, threshold])returnsbool.jev_probreturns the Noul probability.jev_choiceandjev_score/jev_score_normreturn the top label and the level.jev_score(row, q, levels text[])built 2026-09-26 (tests/score.rs): the vendor's probability-weighted level, 0 to n − 1, andjev_score_normthat level over n − 1. Levels are sent lowest first in the caller's order, so reordering them is another question and another cache key. Fewer than 2 or more than 10 levels, or a NULL level, raises 22023 before anything is sent. The two share one judgment and onejev_cacherow, which stores the score, confidence, per-level probabilities and legend.jev_choice(row, q, options text[])built 2026-09-26 (tests/choice.rs): the top label. Options are sent in the caller's order as label keys with no description, so reordering them is another question and another cache key. A single option is returned without a request or a cache row. No options, more than 255, or a NULL, empty or repeated option raises 22023 before anything is sent. Thejev_cacherow stores the choice, confidence and per-option probabilities as ordered pairs, since a jsonb object would sort them.jev_confidence(row, q, kind, options text[])returns the vendor's confidence in a Choice or Score, stored exactly as returned and never recomputed. Built 2026-09-26 (tests/confidence.rs) with pg-jev's signature:kindis'choice'or'score', and it asks the same question asjev_choiceorjev_scoreon the same options, so they share one judgment, one request and onejev_cacherow. A single option is 1.0, asjev_choice_fullreports it.- A Noul carries no confidence, and none is derived (settled
2026-09-26).
'noul'(or any other kind) raises 22023 before anything is sent. The evidence:- The vendor says "(Noul answers don't carry one.)"
(docs.typesafe.ai/confidence, read 2026-09-26). The wire contract
agrees: in
typesafe-sdk0.7.0 and 0.7.2,_schemas/models.pyis generated fromhttps://api.typesafe.ai/openapi.json, andNoulAnswerhas onlytypeandnoul, whileChoiceAnswerandScoreAnswerrequireconfidence(pypi.org/project/typesafe-sdk/ 0.7.2). The Noul page's response example has none either (docs.typesafe.ai/primitives/noul). For a Noul the probability is the answer and the certainty in one. - jevQL derives one as
max(p, 1 − p), its own addition (https://github.com/kylemclaren/jevql/blob/274532af852e8edfb7715ec6dca1113e589cb191/internal/typesafe/client.go#L59-L68). That is a recomputation, and it disagrees with the vendor's own formula (below) at n = 2, which is2 · max(p) − 1: at p = 0.8 the vendor's form gives 0.6 and jevQL's 0.8, and jevQL's never falls below 0.5, even at p = 0.5. A threshold tuned on Choice confidences would mean something else on it. It is not taken. - pg-jev reads the answer's
confidencekey, so itskind = 'noul'silently returns NULL (https://github.com/realZachi/pg-jev/blob/afd11fa856d7a2b831a1bfd8ee7f869ce8efcd62/sql/jev--0.2.0.sql#L527-L531). postjevsql raises 22023 instead: refuse loudly, not guess. - A caller who wants jevQL's figure writes
greatest(p, 1 - p)overjev_probin their own query.
- The vendor says "(Noul answers don't carry one.)"
(docs.typesafe.ai/confidence, read 2026-09-26). The wire contract
agrees: in
jev_evalreturns the Noul in full, as a composite (Typed results), not raw JSON.- Threshold precedence: argument, then the
jev.thresholdGUC, then 0.5. Built 2026-09-26 (tests/threshold.rs):jevis two overloads,(anyelement, text)and(anyelement, text, float8), claimed by the scan likejev_prob. True when the probability is>=the threshold (pg-jev's comparison).jev.thresholdis a Userset GUC in [0, 1] whose default is 0.5, soRESETfalls back to it; an argument outside [0, 1] raises 22023 before anything is sent. The threshold is applied to the judgment, never sent, sojevandjev_probon the same question share one request and one cache row. A NULL threshold argument gives NULL, since every function is STRICT (pg-jev'sDEFAULT NULLfalls through to the setting instead).
-
Operational functions and modes (from the newer Jev extensions: prasanthj/duckdb-jev, parable-work/jev-datafusion):
jev_stats()returns this backend's counters since it started: requests that reached the server, retries, redials (never sent), connections opened, cache hits and misses, input and output tokens, and cost (input tokens atjev.price_per_mtokwhen answered; output is not billed), andin_flight, the HTTP/2 streams open now. Per backend, not cluster-wide. Built 2026-09-26 (postjevsql-pg/src/stats.rs). A stream is counted by a guard held across the transport's send, so a timed-out, cancelled or failed request leaves the count onDrop.dedupe_hitscounts calls that waited on an equal judgment instead of sending.- Every retry, redial, connection and answer is logged at DEBUG1 with
its request id, through
jev-client'sObserverport (review D4). jev.on_error=error(default) |unsure.unsureturns a failed judgment into NULL:jev()is then false and the probability NULL, so a network fault skips a row and never passes one. A failed judgment is the remote or the path failing on this row: a 408, 429, 529, 5xx, context overflow, connect or interrupted exchange, timeout, oversized response, or the failure of the request an equal call waited on. Settings, the key (401/403), a request we built wrong, TLS, a moved model, the budget andjev.cache_onlystill raise, since they would fail every row alike or guard spend. Nothing is cached for a NULL. Built 2026-09-26 (Failure::is_failed_judgment,tests/on_error.rs); Userset.
-
rowis a table alias (the whole row, pg-jev) or a column listjev((a, b), q), which sends only those columns (jevQL). Prefer the column list. It sends less data and costs fewer tokens. -
Typed results (pg_typesafe).
jev_evaland the*_fullvariants return composites, not bare text or jsonb:jev_choice_result(choice text, confidence float8, probabilities jsonb, model text, input_tokens int, output_tokens int)- The same shape for Noul and Score.
Built 2026-09-26 (
tests/eval.rs):jev_eval(row, q)returnsjev_noul_result(probability, model, input_tokens, output_tokens), the Noul having no confidence;jev_score_full(row, q, levels)returnsjev_score_result(score, confidence, probabilities, legend, model, input_tokens, output_tokens);jev_choice_full(row, q, options)returnsjev_choice_result. They are claimed by the scan and ask the same question as their scalar function, so they share its judgment, request andjev_cacherow; a hit reads the row'sanswered_modeland tokens.probabilitiesis a JSON array in the order asked ([label, p] pairs for a Choice). The tokens are the whole request's, as stored. A single option has a NULL model and 0 tokens, since nothing was sent. -
A Choice sends 2–255 options; a Score sends 2–10 levels. The API reference caps Choice at "a maximum of 255 options per Choice", and Score "should have at least two levels; the API accepts up to 10" (docs.typesafe.ai/api, read 2026-09-22). The Rust type makes any other count unconstructible. A one-option Choice is answered locally with probability 1.0 and never sent, as the vendor's own hierarchical cookbook does.
-
Option names are sent to the model, and their order is part of the question. The docs say "the option names and their descriptions are both sent to the model". Hume's
option-position.jsonmeasured order bias:- Reversing a list moved a support classification from 0.84–0.89 to 0.93–0.96.
- The correct answer scored 0.34 when listed first and 0.83 when listed last.
So:
- Options keep the caller's order end to end and are never shuffled. Randomizing would make the same row give different answers and break the cache key.
- Options travel as an ordered list of
(label, description)pairs (SQLtext[]or a composite array). They never pass throughjsonb, which does not preserve key order.serde_jsonis built withpreserve_order, and the request serializer writescriteriain that order. - The label itself is the option key. Do not use opaque
c0…cNkeys: at 200 options they cost +67% input tokens (4,244 against 2,546) for the same accuracy (Hume,option-position.json). Siblings are unique already. - Offer an
other/none of the aboveoption when the list may not cover every input (vendor guidance).
-
More than 255 options has two modes. Both are exact about what they return.
- Label tree (default, when the labels have a meaningful
hierarchy). This is the vendor's documented method
(docs.typesafe.ai/cookbooks/hierarchical_classification), corrected
as follows:
- Group siblings by meaning. Never split on the bytes or bits of IDs, because Jev cannot judge an arbitrary bucket. Jev also reads hex and numeric encodings worse than names (jaggedness page).
- Keep fanout well below the cap: the vendor reports "a Choice works reliably up to roughly 240 options" (classification_using_confidence cookbook, jev-1.12).
- Each option describes its subtree: its direct children and a sample of leaves. The instructions name the parent path. This is the vendor's own guidance ("each option's value is the child's tree… lets the model see what lives under a branch"); its hierarchical cookbook does neither.
- A node's score is the product of its edge probabilities, kept as
Σ log max(p, 1e-9). It is a real probability that can only fall with depth, so a partial path bounds every leaf below it, and best-first search with pruning is exact. Do not use the cookbook's geometric meanexp(mean(log p)). It rises with depth and favours deep paths of confident edges (five edges at 0.9 score 0.90 against one edge at 0.8, while the real probabilities are 0.59 and 0.8). A one-child edge is ×1 and needs no special case. - Beam with K = 3 by default, all frontier Choices in one request. K = 3 is the vendor's figure and has not been measured by us.
- Lookahead: each request also carries the child Choices of each frontier node's top-m likely children, up to the batch token budget. That scores two levels per round trip and roughly halves the round trips (CPC runs to 12 levels). Questions in one request are independent (Hume), and the vendor recommends speculative questions.
- Close calls get an order-debiasing twin. When a node's separation is near 1×, a reversed-order copy of that Choice is asked, and the log-probabilities are averaged (permutation self-consistency, arXiv 2310.07712). The separation is known only from the answer, so the twin cannot ride in the same request: the node is queued again at its score and the twin takes a K slot in a later round. It costs one copy of that node's option tokens, a round trip only when nothing else is left to ask, and nothing when the node is pruned first.
- Return the leaf, its path probability, and the separation
exp(score_top − score_second)(near 1× means ambiguous). - The search is built 2026-09-26 (
postjevsql-core/src/label_tree.rs), sans-I/O;jev_choice_tree(below) drives it from SQL. - Lookahead and the twin are built 2026-09-26 (
Lookahead,TwinBelow, on by default at m = 2, a 60,000-token round and 1.5×, all unmeasured and part of the release gate). Lookahead's "likely" is a prior, since a child's Choice rides with its parent's: the children with the most leaves, ties in the caller's order. An answer held before its node is scored is applied without a request once it is, so neither changes the result (exact_against_exhaustiveruns every combination). K is the Choices asked per round, not a width that drops nodes: a node is dropped only when its bound cannot change the top two leaves (or, under τ, the deepest node at or above τ), so the result and the separation are exact. A one-child node is passed through at ×1 without a request. - Each node's Choice is built by
Tree::choice(2026-09-26): the instructions are the question and, below the root,Within: A > B; an internal child's description isContains:its first 20 direct children (, and N more) andFor example:its first 5 leaves not already listed, depth first; a leaf child has none. The 20 and 5 (Describe) are unmeasured and belong to the release gate. The scan side is built (2026-09-26,postjevsql-pg/src/scan/backend.rs): a judgement awaitsscan::on_backend(closure), the scan runs the closure on the backend between waits (so a round's cache lookups are SPI there) and the judgement resumes with the result. A closure whose judgement was dropped is skipped. jev_choice_tree(row, q, paths text[] [, tau])built 2026-09-26 (TreeCallincrates/postjevsql/src/lib.rs;tests/label_tree.rs). Each path is a leaf's labels joined by>, the separatorWithin:uses, and the result is the chosen node's path in the same form, so it is unique where a bare label need not be. Undertaua stop at the root (nothing at or above τ) is NULL. The tree is built and refused (22023) before anything is sent. Each round's Choices are looked up throughon_backendand the misses sent in one request with the row as the state; each node's Choice is its own judgment, deduplicated and cached like ajev_choice. Every round that sends counts as a row againstjev.max_rows. A label containing>cannot be expressed;jsonbinput is not built.jev_choice_tree_full(row, q, paths [, tau])built 2026-09-26 (tree_outincrates/postjevsql/src/lib.rs;tests/label_tree.rs) returnsjev_tree_result(path text, probability float8, separation float8, depth int, leaf bool, fit float8)from the coreOutcome, sharing every node's judgment withjev_choice_tree. Without τ it is the leaf, its depth,leaftrue and the separation (NULL when the tree has one leaf). Under τ the separation is NULL, since a stop proves no runner-up, and a stop at the root has a NULL path, probability 1 and depth 0.Branch::from_paths(2026-09-26) builds the tree from label paths, top down, siblings in first-appearance order. A path given twice, or one that is both a leaf and the parent of others, is refused: without τ only leaves are returned, so such a label could never be an answer.- Optional early stop with
tau => …: return the deepest node whose path probability is ≥ τ, with its depth and a leaf/internal flag. The default (τ unset) always returns a leaf. The vendor measured the trade: forced leaves were right on 39/60, and reporting the uncertain half one level up made 48/60 useful. Any τ is measured against the pinned model version.
- Chunk and shortlist (opt-in, for label sets with no meaningful
hierarchy and N up to about 1,000). This is the vendor's shape in the
skill-suggestion and line-by-line-search cookbooks:
- Request 1 sends ⌈N/254⌉ chunk Choices together and keeps the top few from each.
- Request 2 is one Choice over the finalists, with full descriptions.
- That is two round trips at any N. The cost is tokens: about 8.6 tokens per option (Hume: 20 × 200 options came to 34,590 tokens), so it is roughly 10–40× the input of a tree.
- The search is built 2026-09-26 (
postjevsql-core/src/shortlist.rs), sans-I/O. Chunks are contiguous runs of the caller's list, as even as ⌈N/254⌉ allows, sent as labels only; the topkeepof each (3, ours and unmeasured, part of the release gate) go to the finalist Choice in the caller's order with their descriptions. Its probability is among the finalists, not over all N. No labels, an empty or repeated label, or akeepthat leaves fewer than 2 or more than 255 finalists is refused (22023). One label is returned without a request. jev_choice_shortlist(row, q, options text[] [, descriptions text[]])built 2026-09-26 (ShortlistCallincrates/postjevsql/src/lib.rs;tests/shortlist.rs), opt-in besidejev_choice_treeand driven the same way: each round's Choices are looked up throughon_backend, the misses sent in one request with the row as the state, and each Choice is its own judgment, deduplicated and cached. Every round that sends counts as a row againstjev.max_rows. Descriptions are one per option (a NULL one sends none) and ride with the finalists only; a NULL option, or a descriptions array of another length, is refused (22023) with the core's refusals before anything is sent. It returns the label.jev_choice_shortlist_full(row, q, options [, descriptions])built 2026-09-26 (ShortlistCallincrates/postjevsql/src/lib.rs;tests/shortlist.rs) returnsjev_shortlist_result(jev_choice_resultplusfit), sharing every Choice's judgment andjev_cacherow withjev_choice_shortlist.choiceis the label the search chose (the vendor's may differ on a tie). Confidence,probabilities([label, p] pairs, the finalists in the caller's order, among the finalists and not over all N), model and tokens are the finalist Choice's, as stored; round 1's tokens are injev_stats()and EXPLAIN ANALYZE. A single option is 1.0 with a NULL model and 0 tokens, asjev_choice_fullreports it.
- "Does any label fit at all" is a separate Noul in the same request
(the vendor's
gate/existspattern). Nouls never rank labels, because Noul and Choice scores are not comparable. Built 2026-09-26 (postjevsql-core/src/gate.rs,Tree::gate,Shortlist::gate;tests/label_tree.rs,tests/shortlist.rs): the_fullfunctions ask it in their first round's request and return its raw probability asfit, the last column ofjev_tree_resultand ofjev_shortlist_result(jev_choice_resultplusfit). No threshold is applied, since thresholds do not transfer. The scalar functions do not ask it, as they could not return it. Its instructions are the question andDoes any of these labels fit?followed by the tree as its root is described (Contains:/For example:), or every shortlist label in the caller's order; it is its own judgment, deduplicated and cached.fitis NULL when the search sends nothing (one leaf or one option). - The cookbook's "beam 4/4 against greedy 2/4" is one hand-picked example per hierarchy, so it is anecdote, not accuracy evidence. The literature picks no single winner: Yoshimura & Kashima (arXiv 2508.04219) find top-down and direct prompting each win on different datasets, and LATTICE (arXiv 2510.13217) finds shortlist reranking wins at low token budgets and tree search at high ones. This is why both modes exist and the label-tree harness is a release gate (Testing).
- This is hierarchical classification. It is not a radix tree.
- Label tree (default, when the labels have a meaningful
hierarchy). This is the vendor's documented method
(docs.typesafe.ai/cookbooks/hierarchical_classification), corrected
as follows:
-
Thresholds do not transfer between question types or model versions. The jaggedness page shows a Noul and a yes/no Choice on the same question disagreeing (0.22 against 0.01).
jev.thresholdapplies to Noul only, and any tuned threshold is recorded against the pinned model version. -
Refuse loudly instead of guessing (jevQL's discipline). jevQL's refusal list is the set of cases this extension must eventually support, not the set it may skip:
jev_*inORHAVING- window functions
- DML (handled, with RETURNING still refused; below)
- CTEs and subqueries
DISTINCT- comparisons like
jev_prob(...) > 0.7 - a
jev((a.x, b.y), q)over two relations. That is a join condition, whichset_rel_pathlist_hooknever sees; it needs a join or upper-path hook.
Comparisons,
ORandANDbatch in the relation scan: it claims a call anywhere in a restriction clause, and judges only where Postgres reaches it (Execution).Aggregates,
HAVING,DISTINCTand window functions are handled (2026-09-27,postjevsql-pg/src/scan/lift.rs;tests/batching.rs): aplanner_hookpost-pass, afterstandard_planner, puts a jev scan between an Agg, WindowAgg, Group or Result and its child wherever that node evaluates a call reading only the child's columns (OUTER_VAR, no Param, SubPlan or aggregate), laid out asplan_custom_pathlays it out, and rewrites the call to an OUTER_VAR column of the new scan. The scan'slefttreeis also the child, for EXPLAIN's deparser only. The relation scan no longer claims calls under an Aggref or WindowFunc.CTEs and subqueries are handled (2026-09-27,
tests/batching.rs, on 17 and 18 against a per-row reference) with no code of their own: each is planned by its ownPlannerInfo, so the relation scan wraps its relations there, and the post-pass and theExecutorStartcheck walkPlannedStmt->subplansand every plan state'sinitPlan/subPlan. Covered: aMATERIALIZEDCTE and one read twice (judged once), a scalar InitPlan,IN(pulled up to a semi join), a hashedNOT INSubPlan, and a correlated scalar SubPlan.EXISTSwith a jev qual stays a correlated SubPlan, becauseconvert_EXISTS_sublink_to_joinrefuses a volatile WHERE. A correlated SubPlan runs once per outer row, so each run's scan judges that run's rows and repeats are cache hits; it is batched per run, never across outer rows. A call reading an outer column (jev_prob((t.body, o.id), q)in a correlated subquery) gets it as a Param: the scan evaluates it on every run, each run sends its own value, and a repeated one is a cache hit, whether or not the child depends on the Param. UnderWITH RECURSIVE, a condition over the worktable (jev(r, q)) and one over a table the recursive term joins are judged per iteration, the recursion stopping where the condition fails, and no row twice (tests/batching.rs, 2026-09-27).Partitioned and inherited tables are handled (2026-09-27,
wrap_pathsinpostjevsql-pg/src/scan/plan.rs;tests/batching.rs, on 17 and 18): the appendrel parent is not wrapped, each member is, with its translated quals and the parent's select list translated byadjust_appendrel_attrs_multilevel. An Append pulls members in turn, so a member's rows are in flight together; the members share the statement's dedupe and budget, and an equal judgment in two partitions is one request.DML and row locks are handled (2026-09-27,
tests/dml.rs, on 17 and 18 against the mock's per-row answers;tests/sidecar.rsfor UPDATE through a foreign table):wrap_paths,wrap_projectionsand thelift.rspost-pass no longer skip non-SELECT statements or rowMarks, so UPDATE's SET and WHERE, DELETE's WHERE, INSERT … SELECT and SELECT … FOR UPDATE are batched, and a call in both SET and WHERE is one judgment. An EvalPlanQual recheck never sends (settled 2026-09-27). The scan'sscanrelidis 0, so ExecScan's recheck asks it to fill the slot: it pulls the child, whose scan substitutes the row's current version, and judges it as any row. The recheck's executor has its ownEState;Statement::newfollowses_epq_activeto the parent's, so the recheck shares the statement's dedupe registry and budget, andStatement::recheckinggives it no client. An unchanged version has the same key and is answered from the statement's judgments or the cache. A version whose judged columns a concurrent update changed is a question the statement never asked, so it is refused with 40001 (serialization failure, retry the statement) rather than sent or answered with the old version's answer. Underjev.on_error = unsure, a failed judgment inSETwrites NULL.Still refused: a call over two relations (a join qual or a join's select list); a call in RETURNING, ON CONFLICT or a MERGE action, which ModifyTable evaluates row by row above the scan (the
ExecutorStartcheck reads those lists); and a whole-row argument of a foreign table being updated, which setrefs matches to postgres_fdw's own whole-row identity column (both varattno 0) and which fails with 42804. Use a column list there.Until one of the rest is handled, it raises a specific error. The error comes before anything is sent: an
ExecutorStartcheck (postjevsql-pg/src/scan/guard.rs) refuses any plan in which a call is evaluated outside a jev scan, so a scan elsewhere in the same plan never judges a window first. -
Volatility.
jev*functions areVOLATILE(they call a remote model whose answers vary between calls) andPARALLEL RESTRICTED, so the relation gets no partial paths. Otherwise a parallel seq scan could evaluate them one row at a time in the workers, becauseset_rel_pathlist_hookruns before Gather paths are generated (allpaths.c:542-561).
Execution
-
Batching belongs to the plan, not to a function (new). A CustomScan node, added via
set_rel_pathlist_hook, consumes child tuples, builds requests and emits results. Do not use pg-jev's side-channel read-ahead inside a scalar function, which falls back to one request per row for subqueries and CTEs (pg-jevAGENTS.md:70-80). Do not make the caller build atext[](pg_typesafe's*_many). See Implementation §1 for how the node is written over pgrx.- The node claims every
jev*expression, not just the WHERE quals.set_rel_pathlist_hooksees only the relation's restriction clauses. Ajev_prob(...) AS pin the target list or ORDER BY would otherwise run once per row in the projection, outside any batch. The node readsroot->processed_tlist, computesjev*there throughcustom_scan_tlist, and setsCUSTOMPATH_SUPPORT_PROJECTION. ParadeDB does the same for its score functions (pg_search/src/postgres/customscan/basescan/mod.rs:1109-1170, AGPL, design reference only). - It wraps every path in
rel->pathlist, parameterized ones included, so index nested loops survive. - LIMIT is honoured by demand, not by a bound.
ExecSetTupleBoundnever reaches a CustomScan (execProcnode.c), so the node sends requests only as rows are pulled, keeping at most the in-flight window ahead. - Settings are read in
begin, not at plan time. Generic plans for prepared statements outlive aSET. - Rows the SQL filters out are never judged by construction: the
child plan runs all non-
jevquals before the node sees a row. - A call Postgres would not reach is never judged (built
2026-09-27,
postjevsql-pg/src/scan/reach.rs;tests/batching.rs). The plan post-pass derives each call's reach from the scan's final quals and projection: earlierORarmsIS NOT TRUE, earlierANDarmsIS NOT FALSE, theCASEtests before itIS NOT TRUEand its ownIS TRUE, earlier qualsIS TRUE, and every qual for the projection; OR over a call's occurrences. It is the last per-call entry ofcustom_exprs. The scan evaluates it on each child row and leaves an unreached call NULL, sending nothing. A condition that reads another call's answer, a volatile function, a SubPlan or aCASE x WHENplaceholder is dropped, which only widens the reach (jev(a) OR jev(b)still judges both).
- The node claims every
-
A planner support function sets the cost (
SupportRequestCost), so row estimates and join order account for the expense. It does not order quals by selectivity.order_qual_clausessorts by security level and per-tuple cost only, and deliberately ignores selectivity (createplan.c:5416-5535, REL_18_STABLE).SupportRequestSelectivityapplies only to a bare booleanjev(...);jev_prob(...) > 0.7is an operator expression and gets the default inequality estimate (~⅓). -
Request layout: each row is judged alone (new; this replaces "20 rows per shared state"). pg-jev's layout packs rows into one
stateas{"condition", "rows": [...]}. Its "100% at 1–20 rows" measured thresholded labels on three easy questions, not probabilities. On the probabilities that layout is measurably wrong:- jev-orderby-bench (commit 5239795, 2026-09-20, pre-registered, same text at 40 rows per state against 1) found rank correlation with the single-row baseline fell to 0.579, and 77 of 360 decisions flipped at 0.5. It is a position effect: rows in slots 0–7 moved 0.049, slots 16–23 moved 0.31, slots 24–39 about 0.42. A 20-row batch reaches into the bad band.
- colliber/duckdb-jev issue #4 (2026-09-18) measured
scoreanswers shifting by −0.29 levels on a 0–3 scale depending on which rows shared the state, whilechoicelabels agreed 269/270. - The vendor's own re-ranking cookbook sends "one request per
candidate · no request sees another". MotherDuck's
prompt_jev()packs 32 rows per request by default; this contract rejects that on the evidence above.
So a request is built one of two ways, and no row ever shares context with another:
- Row as state when a row has two or more distinct questions in the query. The state is the row, and each question is a branch. This is the measured single-row baseline, and the state is billed once for all of that row's questions.
- Row in instructions when a row has exactly one question. The
stateholds only context shared by every row (often empty). Each (row, question) becomes its own question whose structuredinstructionscarry the row. The vendor documents this form ("loop over the potential records and build one of these questions per field, all sent in a single call", primitives/advanced). Hume measured that questions are isolated: a secret placed in a sibling question scored 0.00. So many rows share a request without seeing each other. Throughput is then bounded by tokens, not requests: at ~175 tokens per row, 250k tokens/s is ~1,400 rows/s against 20 requests/s.
Release gate: row-in-instructions has no published accuracy measurement against row-as-state, and Hume found facts placed in options scored worse than in the state (0.49–0.89 against 1.00). Before it is enabled, the jev-orderby-bench harness must compare row-as-state against row-in-instructions on mean |Δp|, Spearman and position effect (Testing). If it fails, the fallback is row as state for every row, which is measured correct and caps throughput at the account's request limit (20 rows/s).
-
Deduplicate before sending (new). Within a statement, identical cache keys share one request: the first sends, and the rest wait on its answer. The column-list form makes duplicates common. JevSQL and LOTUS both deduplicate before calling. Built 2026-09-26 (
crates/postjevsql/src/lib.rs,Shared;tests/cache.rs): the first call of a key registers it, and an equal call waits on it without a cache lookup or a budget charge. The registry is one per statement, held in the sameStatementstate as the spend guards (postjevsql-pg/src/scan/statement.rs), so the scans of a self-join share requests too. A request that fails fails its waiters. One that is dropped unanswered (a rescan, a cancel) unregisters the key, and each waiter then asks again itself, so another scan's rescan never fails a call. -
Network I/O stays in the backend, non-blocking, and waits on the latch (pg_typesafe's principle; the mechanism is in Implementation §2):
- Ctrl-C and
statement_timeoutmust work while a request is in flight. - Never call a Postgres API from a spawned thread.
- Ctrl-C and
-
Concurrency is bounded by
jev.concurrency, which counts HTTP/2 streams on the backend's one connection, not connections. It is clamped to the server'sMAX_CONCURRENT_STREAMS(100, measured; Implementation §2). -
Retries follow the vendor SDKs, checked in source: Python
typesafe-sdk0.7.1 (_core/retry.py) and JS@typesafe-ai/sdk0.6.0 (dist/index.mjs:73-119), fetched 2026-09-22.- Retry 408, 429 and 500–599 (529 is inside that range), plus connection errors and timeouts.
- At most 2 retries after the first attempt.
- Backoff starts at 0.5 s, doubles to a 5 s cap, and subtracts up to 25% jitter.
- A server delay is read from
retry-after-msfirst, thenRetry-After, and capped at 60 s (the JS SDK's cap). - Each attempt times out after 10 s (both SDKs' default). The whole
retry budget is at most 30 s (Python) and never beyond
statement_timeout. - Retries send
X-TypeSafe-Retry-Count: n, as both SDKs do. - Any other 4xx fails immediately. A missing key returns 403
(measured 2026-09-22); an invalid key returns 401
authentication_error(measured 2026-09-23). Error bodies are{"detail": {"error_type", "message"}}. - Context overflow is a 400 with
detail.error_type = "max_tokens_exceeded", not the 422 the vendor table implies (Hume'scontext-limit.json, 20 trials; jev-axisrc/client.ts:311). That 400 halves the request and retries. A single row that still overflows fails with a clear error. - Every error message includes the
x-typesafe-request-idresponse header, which is present even on auth errors. - Responses are capped at 8 MB (pg_typesafe). An oversized answer was already produced and billed, so it is never requested again.
- Never sent is not a retry. When hyper hands the request back unsent, or h2 reports REFUSED_STREAM or a GOAWAY on it, the request is resent at once on a fresh connection. There is no backoff and no retry-count, at most twice.
- A TLS refusal (certificate, no h2) is never retried.
-
Errors carry a SQLSTATE a caller can branch on, with the request id and attempt count in DETAIL (
crates/postjevsql/src/failure.rs). The classes follow postgres_fdw's use: 08 when the remote cannot be reached, 38 when the remote failed.Failure SQLSTATE settings, invalid request 22023 no or two API keys, HTTP 401/403 28000 DNS, connect, TLS 08001 interrupted, timed out 08006 429/529 53000 a miss under jev.cache_only55000
| a changed row version in an EvalPlanQual recheck | 40001 | | context overflow, response over 8 MB | 54000 | | other API errors, wrong answer | 38000 |
-
The account's rate limit is shared cluster-wide, and the extension enforces it rather than discovering it through 429s (new). jev-1.13.0 allows 250,000 tokens/s and 1,200 requests/min per account, and the vendor warns the limits "can change without notice". No
x-ratelimit-*headers are documented, and none appear on error responses.- GCRA, one
PgAtomic<AtomicU64>per limit, each holding a theoretical arrival time. Admission is one compare-and-swap loop per limit, and a request is admitted only if both limits pass. The token cost is estimated, then corrected fromusage.input_tokensafter the response. A token bucket would need two words and a lock; GCRA needs one atomic. - The rate adapts to 429s: a shared effective rate in shmem halves
on each 429 and rises gradually after successes. The GUCs
jev.max_tokens_per_secondandjev.max_requests_per_minuteare hard ceilings. - The wait for admission happens inside the
WaitEventSet, so it stays cancellable. - Built 2026-09-26 (
postjevsql-core/src/gcra.rs,postjevsql-pg/src/ratelimit.rs;tests/ratelimit.rs). The three words (two TATs and the adaptive scale) live in a PG17+ named DSM segment (GetNamedDSMSegment), attached inbegin, soshared_preload_librariesis not needed; they are stdAtomicU64s timed onCLOCK_MONOTONIC, which every process shares. The ceilings areSighup(one account, one limit), default 250,000 and 1,200, and 0 is no limit.Transport::admitwaits before every attempt, retries included, on an executor timer, outside the attempt timeout and the 30 s retry budget, so a queue is never read as a slow server. The estimate is the body at the learned characters per token (below; read once inbegin), corrected fromusage.input_tokens; a never-sent request returns its tokens. A 429 halves the rate and re-prices the throttled request's slot at the halved rate, so its retry already waits the new interval.
- GCRA, one
-
A request fits the model's context, and tokens are the only per-request bound. jev-1.13.0 takes 64k tokens per request, and 32k for the state plus any one question. Hume's measurements:
- 6,893 questions in one request were accepted (65,402 tokens); 8,000 were refused, for tokens only.
- The largest accepted request was 65,756 tokens; the largest single branch 33,002.
- A request's fixed overhead is ~267 tokens, plus ~8 per extra minimal question.
The tokenizer is unpublished and matches none of 192 public ones, and deriving it is barred by the terms (MCA §2.3(c)). So the planner estimates with a ratio learned from cached
usage.input_tokensper model, falling back tochars / 2.6(jev-axi's conservative figure). It validates option and level counts before sending. Overflow is handled by the 400 rule above. Built 2026-09-26 (postjevsql-core/src/token_ratio.rs,learned_ratioincrates/postjevsql/src/lib.rs;tests/budget.rs): per request injev_cacheforjev.model(rows sharing a request id, the state once and each question), characters over the input tokens above the 267-token overhead. Statement pricing, the spend guards and EXPLAIN'sEstimated Input Tokensuse it, read once per statement, and so do the rate limiter's admission estimate and the label tree's lookahead budget, which take the statement's ratio inbegin(tests/ratelimit.rs,tests/label_tree.rs). -
The connection is reused across queries (new), not rebuilt per query. A reconnect happens only after more than 300 s idle, because Cloudflare closes idle connections at 400 s (Implementation §2). Measured here: a cold call costs 44.5 ms of our own (library load, TLS setup, TCP/TLS/h2 handshake), and a warm one about 0.17 ms (
tests/latency.rs, 2026-09-23).
Cache
-
Durable table, not per session (new; pg-jev caches per session, jevQL in a local SQLite file, and pg_typesafe not at all). TypeSafe keeps nothing for you: "TypeSafe will be under no obligation to store or retain Customer Data and may delete Customer Data at any time" (MCA §10.3, "Last updated Sep 19, 2026"). Storing outputs is allowed: "TypeSafe … assigns to Customer all of its right, title, and interest" in Output (MCA §4.2).
-
The key is the hash of what was actually sent, so an incomplete key cannot be written:
sha256(pinned model ID ‖ prompt version ‖ namespace ‖ layout ‖ the exact serialized bytes of the state and of this judgment's question object, with the question id removed)- The question id is left out because it is "not sent to the underlying model" (api.md). The vendor's cookbooks do the same: "What goes on the wire and into the cache key: no id, no bookkeeping."
- Layout (row as state or row in instructions) is in the key. Answers are not assumed equal across layouts.
- Namespace is a GUC (
jev.cache_namespace, from JevSQL). Changing it on purpose re-judges every input. - Lookup uses
jev.model, because the answering model is unknown before the call. The check that the response'smodelequals the pin is what keeps that honest.
-
jev.modelaccepts only versioned IDs. A GUC check hook rejects the aliasesjev-latestandjev-preview. The vendor says an alias "moves when a new release ships, so the answers behind it can change without a change on your side", and advises pinning once thresholds are tuned. Versioned IDs are accepted "whether or not they appear" inGET /v1/models. If the response'smodeldiffers from the pin, the statement fails, so a moved model is never cached under the wrong version. No deprecation policy is published, and MCA §2.5 promises only "commercially reasonable efforts" of notice. Sojev.cache_onlyserves hits and fails on misses; the cache stays usable after a pinned version is withdrawn. -
Serialization is defined, not "sorted keys":
- Rows are canonicalized with RFC 8785 (JCS): sorted keys, no
whitespace, RFC 3339 UTC timestamps, NULL columns written as
null.numericvalues beyond f64 precision are written as JSON strings, so no digits are lost. Built 2026-09-23:jev-protocol/src/jcs.rskeeps each number's lexeme. It emits the ECMAScript form only when that form has exactly the original value. Otherwise it emits a string of the exact value in the same ECMAScript layout, so equal values give equal bytes whatever their scale (…890.10000and…890.1). Generic canonicalizers round through f64: serde_json_canonicalizer sent 9007199254740993 as …992.- Dates and times are written by the scan, not
to_json(postjevsql-pg/src/row.rs,RowJson;tests/canonical.rs).to_jsonwrites BC and five-digit years outside RFC 3339 ("0044-03-15T12:00:00+00:00 BC","10000-01-01T00:00:00") and keeps atimetz's own offset ("12:00:00-05"), measured on PG 18, 2026-09-26. Sotimestamptz,timestamp,dateandtimetzuse ECMAScript'stoISOStringyears: astronomical (1 BC is0000), four digits for 0000–9999, otherwise a sign and at least six (-000043,+010000,+5874897).timestamptzends+00:00andtimetzis shifted to UTC; the infinities stay"infinity"/"-infinity".RowJsonwalks records, arrays and domains itself so nested values get the same form, and writes every other leaf withto_json. - The scan evaluates those leaves under a GUC nest level with
TimeZone=UTC, IntervalStyle=iso_8601, extra_float_digits=1,
bytea_output=hex and lc_monetary=C (
postjevsql-pg/src/row.rs). This is the mechanism of a function's SET clause, undone on return and on abort.
- Choice options keep the caller's order (SQL surface) and are never passed through JCS.
- The request body and the cache key both come from this one serializer.
- Rows are canonicalized with RFC 8785 (JCS): sorted keys, no
whitespace, RFC 3339 UTC timestamps, NULL columns written as
-
A cached answer is one recorded draw, not a fixed point. Repeats spread by up to 0.06 on identical input (Hume, 30 calls × 24 questions) and by up to 0.10 when only irrelevant state bytes change (the vendor's Noul self-consistency cookbook:
covered0.43–0.53). The vendor says plainly: "The model is no more deterministic for it." So a threshold's unsure band is ±0.10 until it is measured on this estate's own questions. The cache is still right to have, because it makes reruns reproducible, but it does not rest on a determinism argument. -
confidenceis stored exactly as the vendor returned it, never recomputed, and only a Choice or Score has one (SQL surface,jev_confidence). The docs call it "a statistic computed from the probability distribution" without giving a formula in the prose (docs.typesafe.ai/confidence, read 2026-09-26). The page's interactive widget, visible only in the raw page source, computes(n · max(p) − 1) / (n − 1)clamped to [0, 1], where n is the number of options or levels. The API reference's Choice example (0.88/0.12/0) reports 0.81 where that formula gives 0.82 (docs.typesafe.ai/api), so the formula is a close description, not a contract, and recomputing it would not reproduce the vendor's number. The full distribution is stored too. -
Built 2026-09-26 (
crates/postjevsql/src/lib.rs,tests/cache.rs): thejev_cachetable, one row per (state, question) judgment. The scan looks each call up (SPI) before building requests, sends only the misses, and stores answers inJudge::settle, which runs on the backend between waits and at scan end, since a judgement future may not call Postgres.jev_stats()counts hits and misses per judgment, and EXPLAIN ANALYZE per scan (Cost and safety).jev.cache_only(Userset) serves hits and fails a miss with 55000 before anything is sent; it needs no API key and is exempt from the spend guards, since it spends nothing (tests/cache.rs, 2026-09-26). The table isREVOKEd from PUBLIC, since a writer could forge answers; the scan reaches it as its owner (postjevsql-pg/src/owner.rs), as a SECURITY DEFINER function would. -
Each cache row is an audit receipt:
- the key and the request bytes
- the requested and the answered model
- the
x-typesafe-request-id(the only handle for a billing dispute) usage- the full answer
- the time and the namespace
Cost and safety
-
EXPLAIN shows the price before it is paid (jevQL's
--explain, moved intoExplainCustomScan):- candidate rows after SQL filters
- cache hits (cost $0)
- requests
- estimated input tokens (Execution's estimator)
- dollars
Cache hits are shown only under ANALYZE (decided 2026-09-26). Plain EXPLAIN runs no rows, and the hits depend on the rows' canonical JSON, so counting them would mean running the child scan and hashing every row: a full read of the data from a command that promises to read none. Plain EXPLAIN therefore prices every candidate as a miss, which is the bound
jev.max_costrefuses on, and labels itWorst-Case Cost. Built 2026-09-26 (Judge::explain,postjevsql-pg/src/scan/exec.rsexplain;tests/cache.rs):- without
COSTS OFF:Candidate Rows,Estimated Requests,Estimated Input Tokens(one attempt) andWorst-Case Cost(three attempts,Budget::price), from the child's row estimate; that estimate is taken before the jev conditions (2026-09-27,rows_beforeinpostjevsql-pg/src/scan/plan.rs), since the child no longer runs them and every row it emits is judged: a table's is its tuples times the selectivity of the other clauses, as the planner computes it, and any other child's is divided by the jev conditions' selectivity (tests/batching.rs,tests/budget.rs); - under ANALYZE, per scan:
Cache Hits,Cache Misses,Shared Judgments(deduplicated),Requestssent,Input Tokensreported by the answers, andCostatjev.price_per_mtok, counted asjev_stats()counts them.
The price is a per-model GUC (
jev.price_per_mtok, default 0.042 for jev-1.13.0), not a constant: MCA §8.2 says rates "may vary based on … the model". There is no batch endpoint, async discount, cached-input discount or idempotency key (searched docs and both SDKs, 2026-09-22). -
Spend guards (jevQL):
jev.max_rowsandjev.max_cost.- Both are checked before the first request, against the estimate.
jev.max_costbudgets 3 attempts per request. With no idempotency key, a retry after a timeout can be billed twice.- Actual
usageis checked while the statement runs, and the statement aborts once spent plus in flight would exceed the budget. - Built 2026-09-26: both are
Userset, with -1 meaning no limit (astemp_file_limit), so 0 means "send nothing".ExecutorStart(postjevsql-pg/src/scan/guard.rs) sums the jev scans' child row estimates and refuses with 54000 before the first request. While running, one budget per statement (postjevsql-pg/src/scan/ statement.rs, keyed by theEStateand dropped when its query context resets) is shared by every jev scan of the plan. Before a row is sent, its worst case (3 attempts, at the learned ratio) is held in flight, and the row that would take the rows sent, or the dollars spent plus in flight, over a guard is refused. When the answers arrive, the worst case is replaced by theirusage.input_tokenstimes each one's attempts, since a timed-out attempt may have been billed. A row dropped unanswered (an error, a cancel, a LIMIT) keeps its worst case as spent. Built 2026-09-26 (tests/budget.rs).
-
Terms that bind the design (MCA, Sep 19 2026):
- §2.3(b): output may not be used to "train a model to imitate the output of the Services". Cached distributions must never train a stand-in for Jev. Using them as features downstream is allowed; the vendor suggests it.
- §2.3(a): no offering the Services "as a standalone service". One API key must never serve third parties through this extension.
- §2.3(c): no deriving the model's "algorithms, structure". So no tokenizer fingerprinting.
- §16.7: terms change on 60 days' notice.
-
Privileges (pg_typesafe):
REVOKE ALL … FROM PUBLICon every function at install. Spending quota needs an explicitGRANT. -
The API key never appears in SQL or logs (pg_typesafe):
- Read it from
TYPESAFE_API_KEYin the server environment, or from the superuser-onlyjev.api_key_file. Exactly one source: both set is an error rather than a precedence rule, which would silently bill the wrong account. Errors name the sources, never the key. - A session
SETis superuser-only and documented as appearing in logs.
- Read it from
-
The endpoint must be
https://.http://is allowed only to localhost for tests. -
jev.mock_responseis superuser-only (pg_typesafe), so a granted role cannot forge answers.
Implementation: Rust's four weak spots and how they are closed
Checked 2026-09-22. Evidence is in the brain page.
0. Workspace: unsafe lives in one edge crate
jev-protocol, jev-client and jev-mock moved to ~/jevcrates on
2026-10-01, shared with jevsnes and jevhooks, and come in as the git
submodule third-party/jevcrates (pinned by its gitlink, buckified by
reindeer as path dependencies). Change them there, then move the pin.
-
jev-protocol: the System One protocol, with no I/O, no runtime and no Postgres, so any Jev client can use it. It holds:- question types (Choice 2–255 unique labels in caller order, Score 2–10 levels, enforced by construction)
ModelId, which accepts only versioned idsJson, with RFC 8785 canonical form for row state- the exact request bytes
- response verification against the questions asked, read through typed keys
- structured API errors, including the
max_tokens_exceededkind - the retry policy
It is our own because no published crate fits (surveyed 2026-09-23). kunobi-jev requires reqwest and tokio. typesafe-client's types-only mode sends state re-serialized through
serde_json::Value, so its bytes would differ from the canonical bytes a cache key hashes. Its design points are credited inlib.rs. -
jev-client: the policy, with no runtime or HTTP stack of its own. It owns:- retries and their budget, and per-attempt timeouts
- "never sent" redials, which are not retries
- request ids on every error, and error classification
It reaches the world through two ports:
Transport: send one request, and classify the failure as NotSent, Connect, Interrupted, Tls, TooLarge or ConfigRuntime: time, sleep, timeout and jitter
postjevsql-pgimplements both. The tests implement them with a script and a virtual clock, so every policy case is exact and instant (review D1–D3, 2026-09-23). -
crates/postjevsql-core(added with the cache): no pgrx and no I/O. It holds:- the batch planner
- the cache key
- the rate limiter's GCRA
- the label-tree search
- chunk-and-shortlist
The batch planner (
src/batch.rs, built 2026-09-26) chooses each row's layout from its distinct questions and builds its state, and the cache key hashes thePlacementit returns, so a key's layout is the one sent. Row in instructions is unreachable until the layout-parity gate passes: its evidence type,LayoutParity, has no values, so every row is row as state. A row is placed only from aPlannedthe planner returned, never from a bareLayout, and row in instructions carries the evidence, so no caller can ask for it.crates/postjevsql/src/lib.rscalls it for every request. It has#![forbid(unsafe_code)]and is tested with plainrust_test. -
jev-mock: a mock System One endpoint for tests. -
crates/postjevsql-pg: the only crate that may containunsafe. It holds:- the CustomScan wrapper (§1)
- the WaitEventSet executor (§2)
- the memory-context reset glue
- every hook registration:
set_rel_pathlist_hook, and the GUC check hooks, which pgrx 0.19.2 madeunsafe(#2348)
It exports safe types only. It may depend on the pure crates (
jev-protocol), never the reverse. -
crates/postjevsql: the pgrx extension. It gluesjev-protocoltopostjevsql-pgthrough their safe APIs, under#![forbid(unsafe_code)]. That is compatible with pgrx's macros:- pgrx emits its
unsafecode with call-site ormixed_site().located_at(..)spans (pgrx-sql-entity-graph/src/ pg_extern/mod.rs:49-56,pgrx-macros/src/rewriter.rs:92). - A rustc 1.98.1 proc-macro experiment confirmed that those spans do not trip the lint; only user-hygiene spans do (2026-09-22).
No published project was found doing this, so CI builds the crate with the lint on to keep the claim true.
- pgrx emits its
1. pgrx has no CustomScan wrapper, so write one, once
pgrx v0.19.2 exposes CustomScan and the planner hooks only as raw
pg_sys bindings. All of that unsafe lives in postjevsql-pg behind a
safe trait. The Jev scan then implements that trait and never touches
pg_sys:
create_custom_pathplan_custom_pathbeginexec(returns the next slot)explainend
No safe Rust CustomScan crate is published: crates.io has none, and
pgrx issue #1405 ("Add ExtensionNode support") has been open since
2024. Model the trait on lagodb-core's LagodbCustomScanProvider
(github.com/lagodb/lagodb, lagodb-core/src/customscan, Apache-2.0,
@db35f59). It has begin, next_slot, rescan, end, DSM, re-checks
for rows locked after a change, and a custom_private envelope that
survives copyObject. It may be copied with attribution. xataio/deltax
(Apache-2.0) and darthunix/pg_fusion (BSD-2) are further permissive
references. ParadeDB's pg_search (trait CustomScan) proves the
shape, but ParadeDB is AGPL-3.0: read it for the design, and do not
copy its code into this MIT-intended repo. Every unsafe block carries a
// SAFETY: comment naming the Postgres invariant it relies on.
The wrapper is sound only if it enforces these. Each one is a trap:
- Every callback in
CustomScanMethods/CustomExecMethodsis#[pg_guard] extern "C-unwind". A PostgresERRORis alongjmp. Jumping across Rust frames that ownDropvalues is undefined behaviour.pg_guardturns anERRORinto a Rust panic and turns a panic at the boundary back into anERROR. One unguarded callback voids the whole wrapper. - Never put a Rust pointer in
custom_private. Plans are copied (copyObject), cached in prepared statements, and serialized to parallel workers. Plan-time state must be a List of Nodes. Runtime state belongs in the scan state, built inbegin. EndCustomScanis not called when the query errors. Anything that owns an OS resource (the connection, sockets) must also be freed by aMemoryContextRegisterResetCallbackon the executor's context. Otherwise every cancelled query leaks sockets.- The scan state lives in palloc'd memory, which never runs
Drop. EmbedCustomScanStateas the first field of a#[repr(C)]struct. Drop the Rust part in place, exactly once: inend, or in the reset callback, whichever runs first. - References into Postgres memory carry the context's lifetime.
Use pgrx's
MemCx<'mcx>. Nothing borrowed from a per-tuple context may outlive the nextexec. - The safe types are
!Sendand!Sync, so safe code cannot move them to another thread. rescanis required, not optional.ReScanCustomScanis listed under "Required executor methods", and any scan on the inner side of a nested loop is rescanned. It resets the batch state; repeats are answered from the cache.- Path flags cover only BACKWARD_SCAN, MARK_RESTORE and PROJECTION
(PG18
extensible.h). Don't set the first two. Set PROJECTION (Execution). Parallelism is controlled throughparallel_safeand partial paths, and thejev*functions arePARALLEL RESTRICTED.
2. Async inside the backend, on Postgres's own wait
Async runs in the backend on a single-threaded executor whose reactor
is a Postgres WaitEventSet. There is one design and no fallback.
-
Executor (in
postjevsql-pg): single-threaded, with noSendbound on futures. The ready queue lives on the backend thread. Wakers push to that queue. There are no other threads, so nothing wakes from elsewhere; debug builds assert the calling thread. -
Reactor: a
WaitEventSetbuilt for each wait from the currently registered fds, holding:WL_LATCH_SETonMyLatchWL_EXIT_ON_PM_DEATHWL_SOCKET_READABLE/WL_SOCKET_WRITEABLEfor every registered socket
It is freed by a
Dropguard. It is not kept long-lived, because there is noRemoveWaitEventin any supported major or master. A long-lived set could never drop a closed socket after a reconnect or a DNS UDP exchange, and its size is fixed at creation.CreateWaitEventSettakes aResourceOwnerin every supported major (17 and later), so there is no version shim. With the usual single h2 socket this isWaitLatchOrSocket, and the extra epoll syscalls cost nothing against 70–500 ms requests.The wait timeout is the executor's earliest timer. After each wake:
ResetLatch,CHECK_FOR_INTERRUPTS()(guarded), then poll the woken tasks. There is one blocking point, with no polling slices and no threads. -
Task scopes. The connection task is scoped to the backend and held in a thread-local. Each scan's tasks belong to a scan guard, and the guard's
Dropremoves them from the executor. The guard is dropped byendor by the memory-context reset callback. Otherwise a cancel either kills the shared connection (if the executor is dropped) or leaves the cancelled query's futures in the ready queue (if it is not). -
Cancellation is
Drop. Cancel orstatement_timeoutraises anERRORfromCHECK_FOR_INTERRUPTS().pg_guardturns it into a panic, which unwinds out of the executor and drops the in-flight futures. That closes their streams, which is async's native cancel semantics. -
Stack (runtime-agnostic crates only):
hyper1.x client, with our socket type implementinghyper::rt::Read/Writeover a non-blockingstd::net::TcpStreamregistered in the reactor, and our executor as hyper's spawner.rustls, which needs no runtime, with the system trust store loaded once.hickory-resolveron the same executor, verified 2026-09-22 against 0.26.3:- It accepts custom runtimes through
hickory_net::runtime:: RuntimeProvider(crates/net/src/runtime.rs:245at tagv0.26.3), plugged in withResolver::builder_with_config(config, provider). - Its I/O traits are
futures-io, not tokio. - Built (2026-09-23):
postjevsql-pg/src/dns.rs. Its sockets areSend + Syncbare fds, and debug builds assert the backend thread.jev.dns_servers(superuser;ip[:port], …) overrides/etc/resolv.conf. That is what lets the tests use a mock DNS server on a high port. Resolved addresses are tried in the resolver's order (IPv4 first). - hickory's cache is moka's
sync::Cache, whosethread::spawncalls are all in tests or doc comments (checked in moka 0.12.16), so it adds no threads. - It is kept over
getaddrinfobecause DNS is the one step where plain libc would block without seeing a cancel (glibc waits 5 s per attempt, per nameserver), and reconnects now happen after every 300 s idle gap. The resolved address is cached for its TTL across reconnects. - Stated cost: hickory reads
/etc/resolv.confand/etc/hosts(crates/resolver/src/hosts.rs:226) but not nsswitch, so nscd, nss-resolve, LDAP and mDNS names do not resolve.
- It accepts custom runtimes through
-
HTTP/2, one connection per backend.
api.typesafe.ainegotiatesh2over ALPN (checked 2026-09-22 withopenssl s_client -alpn h2). Every batch of every query is a stream on that one connection, so there is no per-request connection setup and no pool.-
Server limits (raw h2 handshake, 2026-09-22; the path is a Cloudflare edge in front of an Envoy origin):
SETTINGS_MAX_CONCURRENT_STREAMS= 100INITIAL_WINDOW_SIZE= 65536MAX_FRAME_SIZE= 16777215
jev.concurrencyis clamped to the peer'sMAX_CONCURRENT_STREAMS. Above that limit, hyper'spoll_readyqueues without saying so, and the setting would not mean what it says. hyper keeps h2's view of the peer's SETTINGS private, sonet.rs'sPeerSettingsreads the value off the server's frames as hyper reads them, and the scan asks for its window before each row it pulls. Until a backend's first connection has sent SETTINGS, the limit is taken as 100 (RFC 9113 §6.5.2's recommended minimum); a reconnect starts from the last value seen. Request bodies near the 64k-token limit (~256 KB) exceed the server's 64 KiB stream window and pay window-update round trips; that is a known cost, not a bug. -
Our receive windows are adaptive (hyper's
adaptive_window, set inhttps.rs). hyper's defaults (5 MB per connection, 2 MB per stream,src/proto/h2/client.rs:48-49) are below8 MB × concurrency. -
The connection is reused across queries, not held for the backend's lifetime. Cloudflare closes idle client HTTP/2 connections after 400 s, not configurable (developers.cloudflare .com/fundamentals/reference/connection-limits, updated 2026-07-23). An idle backend runs no executor, so it cannot ping between queries. At the start of each scan, if the connection has been idle for more than 300 s, drop it and reconnect before sending. Pings run only while a scan is active: hyper's keep-alive on the executor's timers, every 20 s without a frame, closing the connection when one goes unanswered for 10 s. Both are superuser GUCs (
jev.keepalive_interval,jev.keepalive_timeout) and part of theEndpoint, so changing one opens a new connection.tests/keepalive.rsshows against jev-mock that pings go out while a scan waits and not between statements, and that an unanswered one closes the connection and the request is retried on a new one (2026-09-26). -
On
GOAWAY, retry only streams abovelast-stream-idor those refused withREFUSED_STREAM. A stream cut mid-flight goes through the normal retry policy. -
No HTTP proxy support comes for free. hyper does not read
HTTPS_PROXY. If a deployment needs an egress proxy, CONNECT tunnelling is built explicitly.
-
Hickory constraints (verified in source):
-
default-features = false, features = ["system-config"]. The default features includetokio, andhickory-net'stokiofeature enablestokio/rt-multi-thread. Every DNS-over-TLS, DNS-over-HTTPS and QUIC feature forcestokiothrough__tls. So the backend resolves with plain DNS from/etc/resolv.confonly. -
Its traits demand
Send + Sync:RuntimeProvider: Clone + Send + Sync + UnpinSpawn::spawn_bgtakesFuture + SendDnsUdpSocket,DnsTcpStreamandTimeareSend + Sync
So the executor must accept
Sendfutures as well as local ones. The socket and handle types must beSend + Syncby construction: they hold a raw fd and a token, and reach the reactor through a backend-thread-local. They are never a pointer into Postgres memory. Debug builds assert that they are used on the backend thread.
Process hazards:
- Create nothing in
_PG_init. Undershared_preload_librariesit runs in the postmaster, before the fork. The rustlsClientConfig, RNG, TLS session state and connection are all built lazily in the backend on first use. Do not rely on aws-lc's fork detection. - No Rust dependency may install a signal handler. Postgres's own
SIGINT/SIGALRM handlers set
MyLatch, which is how cancel reaches the wait. Dropof futures, streams and connections must never panic. Cancellation unwinds (panic = "unwind", from the cargo-pgrx template), and a panic during unwind aborts the backend.- Incompatible with extensions that call Postgres from their own threads. pgrx #2228: pg_duckdb does this, so another extension's thread can run our planner and executor hooks, where pgrx's thread check aborts. Loading both is unsupported; the README says so.
Traps (each looks like a simplification and is a regression):
-
tokio in the backend. Its multi-thread runtime runs futures on worker threads, where Postgres APIs corrupt the backend. Even its current-thread runtime sleeps in its own
epoll, which retries afterEINTRand never sees the latch, so a cancel waits for the request to finish. reqwest, and hyper-util's default resolver, bringspawn_blockingthreads. -
Running tokio in 10 ms slices with interrupt checks in between. That is polling: it adds latency and burns CPU.
This is a project rule, not a pgrx rule. pgrx's own
docs/src/design-decisions.mdallows threads that never touch Postgres. We use none, so no Postgres call can ever happen off-thread. pgrx #1067 (async design, still open) concedes that tokiocurrent_threaddoes not handle interrupts. -
A proxy daemon for the HTTP. It adds a process, an IPC hop and a serialization step to every request, and buys nothing that one multiplexed connection per backend does not already give.
-
Threads in a bgworker. Never.
Considered and not chosen:
- libcurl's
multi_socketinterface. It is a correct alternative. It multiplexes over HTTP/2 by default (CURLMOPT_PIPELININGdefaults toCURLPIPE_MULTIPLEX).CURLMOPT_SOCKETFUNCTIONandTIMERFUNCTIONlet it register sockets in the WaitEventSet, and curl-rust exposes them (src/multi.rs:157,256,515,378). It also honoursHTTPS_PROXYand resolves through nsswitch. Supabase pg_net drives exactly this, from a bgworker with its own epoll. hyper is chosen so that the edge crate's types stay in Rust. The price is building proxy support explicitly and accepting hickory's DNS limits.
3. Packaging per major is nix's job
Supported majors are PostgreSQL 17 and later (17 and 18 today;
PG_MAJORS in build/defs.bzl). The floor is 17 because the rate
limiter's only path is GetNamedDSMSegment (PG17+) and
CreateWaitEventSet takes a ResourceOwner from 17 on, so no code
path carries a version shim. Supporting 16 and earlier is a deliberate
later pass, not an omission: it would add a second rate-limiter path
(shared_preload_libraries shmem) and the WaitEventSet shim back.
Both packages build (2026-09-26). The test harness serves either
major: tests/support/postgres.rs uses extension_control_path on 18,
and on 17, which lacks it, runs the server from a symlink prefix whose
extension directory links the built tree (nixpkgs' relative-to-symlinks
patch makes the server read that prefix's sharedir; measured
2026-09-26). The devshell writes postgres_bin_NN and pg_config_NN
for every major into .buckconfig.local.
One buck graph builds every major (built 2026-09-26). The major is
a configuration: //platforms:pgNN is the host platform plus the
pg_major constraint, and everything that differs by major branches on
it through pg_major_select in build/defs.bzl (unconstrained
configurations get the newest):
- pgrx's and pgrx-pg-sys's
pgNNfeatures, and pgrx-pg-sys's bindgen run'sPGRX_PG_CONFIG_PATH(frompg_config_NN), are set by the wrappers inthird-party/pg_major.bzl, whichreindeer.toml'sbuckfile_importsloads over the prelude'scargoandbuildscript_run, so areindeer buckifykeeps them. - First-party crates'
pgNNcfg, and the tests' server (POSTJEVSQL_POSTGRES_BIN) andPOSTJEVSQL_PG_MAJOR, select the same way (tests/BUCK,tools/record-gates/BUCK). first_party_testalso emits<test>-pgNNfor every older major, withdefault_target_platform = //platforms:pgNN, sobuck2 test //...runs the whole suite on 17 and 18.
Built 2026-09-26 through buck, not cargo-pgrx (nix/package.nix,
flake legacyPackages.<system>.postgresqlNNPackages.postjevsql). The
first-party crates exist only as BUCK targets, so buildPgrxExtension
would need a second, hand-kept Cargo graph beside them. The package
builds //crates/postjevsql:ext in the sandbox, the tree the tests
install, and puts it in nixpkgs' layout (lib/,
share/postgresql/extension/), so withPackages and
services.postgresql.extensions take it. The traps it closes:
- Every
http_archiveinthird-party/BUCKis fetched by nix from the url and sha256 read out of that file at eval, andnix/offline_archive.bzlis loaded over the builtin in the sandbox: buck2's downloader refusesfile://(measured 2026-09-26). - buck2 loads a trust store at startup, so
SSL_CERT_FILEis set though nothing is fetched. - The prelude's wrapper scripts start
#!/usr/bin/env bash, which the sandbox lacks, so buck runs under bwrap with that one path added. - Release builds turn
-Cdebug-assertionsoff intoolchains/BUCK. - The majors built are
PG_MAJORSinbuild/defs.bzl, read bynix/majors.nixfor the flake and the package, which builds with--target-platforms //platforms:pgNNand writes that major'spostgres_bin_NNandpg_config_NN. Any major not listed fails at eval with the supported list, rather than compiling against the wrong headers.
The cargo-pgrx bullets below are the reasoning for the pin, which buck
carries in third-party/Cargo.toml; no cargo-pgrx is built.
- Pin
pgrx = "=0.19.2", and buildcargo-pgrx0.19.2 in this repo's flake. Build it withrustPlatform.buildRustPackage, the same shape as nixpkgs'generichelper; it becomes a one-liner if NixOS/nixpkgs#526051 merges. Pass it explicitly tobuildPgrxExtension. - Never use nixpkgs' default
cargo-pgrx. nixpkgs' owncargo-pgrx/default.nixsays the default is "Not to be used with buildPgrxExtension, where it should be pinned". Every in-tree pgrx extension takes a pinned attribute. - Why not 0.18.x (this reverses an earlier pin to 0.18.1):
- There is no
cargo-pgrx_0_18_1attribute in nixpkgs, so the pin breaks the day the default moves. - pgrx 0.18.0 and 0.18.1 wipe the DETAIL line of every ERROR that pgrx
catches and rethrows (#2262, fixed by #2361 in 0.19.2). Every
pg_guardboundary in this design would lose error detail. - 0.18.1 supports pg13 to pg18; 0.19.2 adds pg19. nixpkgs already
ships
postgresql_19(beta), and GA is expected around the end of October 2026. - 0.19.2 adds safe
PgNodecasting (#2347) and extra executor headers (#2353), both useful for the CustomScan wrapper. - No nixpkgs PR bumps cargo-pgrx to 0.19 (searched 2026-09-22), so carrying our own is the only route.
- There is no
- Per-major builds are ours to wire. Only packages inside nixpkgs'
ext/directory land inpostgresqlNNPackagesautomatically, throughpackagesFromDirectoryRecursive. The flake callspostgresql_NN.pkgs.callPackage ./nix/package.nix {}for each supported major inPG_MAJORS, as the nixpkgs PostgreSQL manual documents. The sidecar module usesservices.postgresql.extensions = ps: [ (ps.callPackage ./nix/package.nix {}) ]. - The test suite is a separate flake check. The package sets
doCheck = false, as every in-tree pgrx package does, because "pgrx tests try to install the extension into the postgresql nix store". nixpkgs'postgresqlTestExtensionprovides the smoke test. - Outside nix, a CI matrix runs
cargo pgrx packageonce per major. - The Cargo workspace is generated from buck (built 2026-09-26,
tools/cargo-gen). Eachfirst_party_*macro records its kind, crate root and deps; the tool writes the rootCargo.toml(members,[workspace.dependencies]fromthird-party/Cargo.toml, pgrx's=0.19.2pin with its major feature stripped, pgrx'spanic = "unwind"profiles) and one manifest per package, with apgNNfeature for each ofPG_MAJORSinbuild/defs.bzl(the last is the default) on every crate that reaches pgrx. Every manifest isMIT OR Apache-2.0,authors = ["The postjevsql Authors"], norepository.//tools/cargo-gen:driftfails when the committed manifests differ. Integration tests (tests/*.rs) stay buck-only: they need buck's location of the built extension. - Licences are checked, not assumed.
deny.toml(built 2026-09-26) is pgrx 0.19.2's allowlist plusBSD-2-ClauseandCDLA-Permissive-2.0; cargo-deny rejects anything unlisted, so AGPL is denied (verified by clarifying a crate as AGPL-3.0-only). The repo isMIT OR Apache-2.0(LICENSE-MIT,LICENSE-APACHE, "The postjevsql Authors"). THIRD-PARTYcarries the linked crates' notices (built 2026-09-26,tools/third-party-notices): every crates.io crate linked into the.so(the target deps of//crates/postjevsql:postjevsql, so build scripts, proc-macros and test crates are left out), its SPDX expression and every licence text, identical texts once. Among them ring is Apache-2.0 and ISC and webpki-roots is CDLA-Permissive-2.0, both compatible with MIT distribution. aws-lc is not linked: rustls is built on ring. The nix package installs it inshare/doc/postjevsql.THIRD-PARTY-sidecardoes the same for the sidecar CLI (built 2026-09-27): the target deps of//crates/postjevsql-sidecar:cli(tokio, tokio-postgres, toml, …), which links no pgrx. Each binary ships only its own set:nix/sidecar-cli.nixinstalls it asshare/doc/postjevsql-sidecar/THIRD-PARTY. Both files come from one run of the tool (:notices[extension],:notices[sidecar]), because onetexts/serves both and its staleness check is against the union.- The archives are read through
$(query_outputs …), which reaches the privatehttp_archivetargets inthird-party/BUCKwhere a dep cannot (athird-party/PACKAGEvisibility is overridden by the targets' ownvisibility = [], and reindeer hard-codes it). So no second list of urls and hashes exists. - A crate with no licence expression or no text fails the build. pgrx,
pgrx-pg-sys, pgrx-sql-entity-graph and seahash ship no text in their
archives, so theirs is in
tools/third-party-notices/texts/, fetched from upstream at the version and named inSOURCE; an entry there for a crate linked into neither binary, or one that ships its own, fails too. //tools/third-party-notices:driftfails when either committed file is stale;buck2 run //tools/third-party-notices:updaterewrites both.
- The archives are read through
4. Everything outside Rust is generated or thin
-
Install SQL: generated by pgrx from
#[pg_extern],PostgresTypeandextension_sql!. Under buck,tools/pgrx-schemareads the.pgrxscsection of the built.so(whatcargo pgrx schemadoes). Do not hand-writepostjevsql--x.y.sql. Grants and revokes go inextension_sql!blocks. -
The version is written once, in
build/defs.bzl. The control file carries pgrx's@CARGO_VERSION@placeholder, filled at build. -
No Perl TAP. Rust integration tests (Testing) start a throwaway instance and connect with
tokio-postgres. They cover what pg_typesafe's TAP tests cover:- retries on 429/529
pg_cancel_backendfrom a second connection mid-request- the 8 MB cap
- https enforcement
The details are in Testing.
-
What remains in SQL is the user interface itself and the expected output, which is correct: SQL is the product's surface.
Deployment: in-database or sidecar
The same extension ships two ways. Choose per target and never fork the code.
- In-database. Install the extension into the Postgres that owns the data. This is the default whenever you control that server.
- Sidecar (new). A local Postgres runs the extension and exposes the
target's tables as
postgres_fdwforeign tables. The target installs nothing. Clients connect to the sidecar with the ordinary protocol: psql, JDBC and GUI tools all work.
Managed hosts cannot run it in-database (vendor docs, 2026-09-22):
- No custom compiled extensions at all: RDS, Aurora, Cloud SQL, AlloyDB, Azure, Supabase and PlanetScale. pg_tle takes "JavaScript, Perl, Tcl, PL/pgSQL, and SQL", no C. Trusted PL/Rust cannot open sockets.
- Possible on request: Neon, Xata and Crunchy may host a custom extension if you ask their support. Xata also offers bring-your-own-cloud.
- Their built-in HTTP calls are per row:
http,pg_net(async; results land after the query ends),aws_lambda,google_ml.predict_rowandazure_ml. They lose the batch node, the cache and the cost guard.
So the sidecar is the supported mode for managed hosts. Proxies were
checked and rejected as hosts: PgDog's plugin API has only
init/route/fini and cannot rewrite queries or see results, and
PgCat is unmaintained.
How sidecar mode splits a query
postgres_fdw ships a function or operator to the remote only if it is
built in, or belongs to an extension listed in the server's extensions
option, and is IMMUTABLE (deparse.c:285,
contain_mutable_functions). jev* is VOLATILE, so it never ships.
A jev* qual that stays on the sidecar is a local condition, and
postgres_fdw then refuses to push down (PG18 postgres_fdw.c):
- joins involving that table (L5837-5842)
- aggregates over it (L6527-6532)
- LIMIT/OFFSET (L7196-7200)
On a foreign table that carries a jev* qual, only the other WHERE
clauses and ORDER BY run on the target. Everything else runs on the
sidecar. So count(*) … WHERE jev(…) pulls every row that passes the
target's filters. A qual referencing two tables is a join qual, so that
join can still push down. LIMIT still stops early, because postgres_fdw
reads through a cursor fetch_size rows at a time.
Everything else in this contract works unchanged, including queries
jevQL refuses: CTEs, subqueries, and UPDATE … WHERE jev(…) through the
foreign table.
The CustomScan wraps whatever scan the planner chose, ForeignScan
included. The child plan is built with the jev* quals removed, and the
node applies them. This holds in both modes: a child that kept them would
judge rows one at a time, outside any request.
Sidecar invariants
- Never list this extension in a foreign server's
extensionsoption. Volatility already stopsjev*from shipping, so this is a guard for the day someone marks a helper IMMUTABLE. At plan time the CustomScan raises anERRORif the foreign server of anyForeignScanit wraps lists this extension. Built 2026-09-26 (refuse_shipping_serversinpostjevsql-pg/src/scan/plan.rs;tests/sidecar.rs, a postgres_fdw loopback):plan_custom_pathwalks the child plan, so aForeignScanunder a Sort (ORDER BY) is found too, and splits the option as postgres_fdw does. The error is 22023, raised while planning, so EXPLAIN is refused as well and nothing is sent. fetch_sizeis at least the node's in-flight window (default 100 rows). Otherwise each window costs several network round trips.batch_sizedoes the same for INSERT.async_capable,parallel_commitandparallel_aborthelp only with several foreign scans or servers, so they stay off for a single target.- The target's credentials never live in the repo.
- On PG18, prefer
use_scram_passthrough. No password is stored, but both servers need the same SCRAM secret, clients must log into the sidecar with SCRAM, and the target must requirescram-sha-256. - Otherwise, the user mapping's password is written from a secret when the module activates (the estate's opnix / systemd-creds path).
- The connection to the target uses
sslmode=verify-full.
- On PG18, prefer
- The planner gets real remote statistics. Set
use_remote_estimate 'true'on the server, orANALYZEthe foreign tables on a timer. Without either, the high cost ofjev*is weighed against guessed row counts, and the batch plan is wrong. - The cache lives in the sidecar (a local table). The target is never written to for bookkeeping.
Costs to state, not hide
- Every row that passes the target-side filters crosses the network to the sidecar, as with jevQL.
- postgres_fdw is not two-phase commit ("it is currently not supported
by postgres_fdw to prepare the remote transaction for two-phase
commit"). The remote COMMIT runs before the local one
(
connection.c:1072-1089). So a failure between the two leaves the target committed and loses only the sidecar's cache rows, which is harmless.PREPARE TRANSACTIONfails on the sidecar once a transaction has touched a foreign table. - Pushdown is narrower than in-database (see above): joins, aggregates
and LIMIT involving a
jev-filtered foreign table run on the sidecar. - There is one extra network hop for every query,
jevor not, that goes through the sidecar.
Launching it
The Rust CLI postjevsql-sidecar is the single engine
(crates/postjevsql-sidecar). It reads one TOML config and:
converge: createspostgres_fdwandpostjevsql, declares the foreign server (target host,sslmodedefaulting to verify-full,use_remote_estimate) with exactly the declared options, and the user mappings (password from a secret, oruse_scram_passthroughon PG18 when none is given)sync: compares the target's catalog with the foreign tables and plans CREATE, ALTER FOREIGN TABLE (added, altered, renamed or dropped columns) and DROP, never drop-and-reimport. It prints the plan, and applies it in one transaction only with--apply. A change that would break a sidecar view, grant or function (pg_depend) is refused, naming each dependent, and nothing is changed.
Each secret (the target password, the Jev API key) comes from a file or an env var, exactly one; both or neither is refused, naming the sources, never the value.
Everything that launches a sidecar calls it: the NixOS module
(nix/sidecar.nix, whose oneshot runs converge then sync --apply,
with opnix or systemd-creds writing the secret files), the container
image, and a Postgres the user already runs. The module keeps only what
nix owns: services.postgresql, this extension from
postgresqlNNPackages.postjevsql, postgres_fdw (contrib) and the unit.
This reverses an earlier rejection (owner, planning interview, 2026-09-26). The contract used to say a Rust launcher would be a second configuration engine beside nix. That held only while nix was the one route. The audience is now any Postgres user, including managed-host users with no nix, served by a container image and by setup for their own Postgres, and the owner ruled that setup tooling is a Rust CLI, never Bash. With three launchers, the configuration engine has to live below all of them, so the CLI is the one engine and nix is one caller; declaring the same state in both would be the second engine.
Status (2026-09-27): the binary and its pg_depend refusal are built
(cli/main.rs, postjevsql cac7966; tests/sidecar_cli.rs), and so is
the module around it. So is the container image (2026-09-27,
nix/sidecar-image.nix, flake packages.<system>.postjevsql-sidecar-image,
a docker load-able tarball): PostgreSQL 18 with the per-major package
and postgres_fdw, dash as /bin/sh only (initdb runs postgres -V
through popen), uid 999, the cluster in the volume
/var/lib/postgresql. Its entrypoint is the CLI's serve, not a
script: resolve every secret (so file-and-env is refused before
anything starts), initdb if $PGDATA is empty, start postgres, wait
until it accepts a connection, converge and sync --apply, then stop
it with pg_ctl stop and exec postgres in its place, so the server is
PID 1 and the runtime's signals reach it; a refusal stops the server and
exits 1 with the CLI's message. Postgres options go after --. Its VM
test (nix/sidecar-image-test.nix, flake check sidecar-image) runs it
in docker against nix/test-target.nix's TLS- and SCRAM-only target.
The module (nix/sidecar.nix, flake nixosModules.sidecar) takes the
extension as extension = ps: … (default nix/package.nix for the
server's major; the flake's module passes it the pinned buck2) and the
CLI as package (default nix/sidecar-cli.nix, the same derivation
building //crates/postjevsql-sidecar:cli; also flake
packages.<system>.postjevsql-sidecar). A oneshot after postgresql
writes one TOML per declared server with pkgs.formats.toml and runs
converge then sync --apply on it. converge also creates
postgres_fdw and postjevsql. The config holds paths only: each
mapping's password file is loaded with LoadCredential and named by
its /run/credentials/postjevsql-sidecar.service/… path, and
apiKeyFile (required, since the unit has no other source) is the
server's jev.api_key_file. Schemas import into same-named local
schemas, as the CLI does. The nixosTest (nix/sidecar-test.nix, flake
check sidecar) boots it against a target cluster on loopback port
5433, TLS and SCRAM only, and checks CREATE EXTENSION postjevsql, the
imported tables read over verify-full with the mapping's password, a
jev call planned as the scan over a ForeignScan, that a restart of the
unit picks up a new target column and drops a hand-added extensions
option, and that neither the unit script nor the config holds the
password.
Testing
- Build and test through buck2 from the nix devshell
(
nix develop -c buck2 test //...). nix provisions toolchains and Postgres; buck owns the graph. Third-party crates come fromthird-party/Cargo.tomlvia reindeer (vendor = false), withcargo_env = trueso build scripts see the full Cargo environment. Toolchain paths reach buck through.buckconfig.local, which the devshell writes, because the prelude's cc shim execs without a PATH search. tests/run-check.shis the repo's one check entry and only delegates tonix develop -c buck2 test //.... mefi-studio runs it to verify a finished task (itsbaseCheckForProjectfindstest/ortests/run-check.shoff Windows, mefi-studio94350fc). Without a check, every finished run was recorded as "unverified" and re-dispatched, so each cache plan ran four times (2026-09-26). Do not add apackage.json: Studio prefers its scripts over this file, so it would become a second entry point.- One buck daemon per checkout, shared by every session in it. Never
buck2 killit and never run concurrent agents in one checkout. Each agent works in its own git worktree, which gets its own daemon andbuck-out. On 2026-09-23 two agents in the main checkout killed each other's daemon ("Forkserver is unavailable"), both stalled, and the cache-key work sat uncommitted for three days. - Warnings fail tests. buck prints no compiler output for actions
that succeed, so a warning is invisible. Every first-party target is
declared through
build/defs.bzl'sfirst_party_*macros, which add a<name>-linttest. That test fails with the report when clippy (which includes every rustc warning) says anything. Dev and test builds keep debug assertions on. - Unit tests:
jev-protocol(in ~/jevcrates,cargo test) has plain tests, plus proptests asserting that every parser of network bytes rejects rather than panics. - Integration tests run a throwaway Postgres (17 and 18,
PG_MAJORS) from the devshell (tests/support/postgres.rs):initdbinto a temp dir,postgresas a child process, the built extension served in place throughextension_control_path/dynamic_library_path. The API is testcontainers-shaped (builder,start, teardown onDrop). There is no docker (ruled 2026-09-23, replacing a testcontainers setup that needed a nix-built image, host networking and a container reaper).- Teardown cannot leak. The postmaster runs with
PR_SET_PDEATHSIG= SIGQUIT, so the kernel shuts it down when the spawning test thread dies, SIGKILL included (verified). - Every test sets
statement_timeout, so a hang fails the test instead of reaching buck's 10-minute kill. tests/latency.rsmeasures the extension's own cost per call against the instant mock (2026-09-23: warm p50 310 µs against a 143 µs floor for the bare query; cold 44.5 ms).
- Teardown cannot leak. The postmaster runs with
- One mock endpoint, our own (
jev-mock): hyper h2- tokio-rustls + an rcgen CA, trusted through
jev.ca_fileand answering from a closure. It serves every HTTP case (429/529 withRetry-After, the 400max_tokens_exceededsplit, the 403 on a missing key, the 8 MB cap, https enforcement). It can also observeRST_STREAMthrough a handlerDropguard, which httpmock cannot, so httpmock is not added as a second mock.
- tokio-rustls + an rcgen CA, trusted through
- Cancellation proves the reset, not merely that the query returned:
- Cancel through both
pg_cancel_backendand tokio-postgres'CancelToken, from a second connection to the same instance. - Assert that no socket leaks after 100 cancels.
- Cancel through both
- Offline regression runs against
jev.mock_response(pg_typesafe's pattern). - Release gates (ground-truth runs, not just a passing suite; each is
recorded against the pinned model version):
- Layout parity: the jev-orderby-bench harness compares row as state against row in instructions on mean |Δp|, Spearman and position effect. Row in instructions stays disabled until it passes.
- Label tree: beam width K, lookahead, the order twin and τ are measured on a labelled hierarchy (the vendor's CPC, Shopify and MeSH sets are pinned and public), against chunk-and-shortlist.
- Unsure band: repeat-spread on this estate's own questions replaces the provisional ±0.10.
- Ranking:
ORDER BY jev_prob … LIMITpasses a pairwise-inversion gate (jev-orderby-bench's ≤ 0.15). Probabilities come back at two decimals, and ties are common (53 of 360 rows at 0.99 in that bench). The README tells users to add their own tie-break column. - Recording is approved; do not ask again. The owner decided (planning
interview, 2026-09-26) that this pass records each gate ONCE on the
live key with
tools/record-gates --send, capped by--max-cost, and commits the result as replay fixtures; tests replay them from then on with no further spend. Running that recording needs no further owner step. Recorded 2026-09-26 against jev-1.13.0 (tests/fixtures/) and replayed bytests/gate_*.rs. What stays owed to a later pass: the measuring benchmarks themselves, and a replay that stores N responses per request.