jevsnes.git / research / jev-api-digest.md
1# TypeSafe AI "Jev" API — implementation digest
2
3Sources read in full: vendor docs `llms-full-2026-09-19.txt` (20,229 lines,
4authoritative), `jev-plays-pokemon/agent.py` + `README.md`, and the brain
5pages `jev.md`, `choice.md`, `score.md`, `noul.md`,
6`speculative-fan-out.md`, `jev-fast-llm-deterministic-tools.md`,
7`ecosystem/{typesafe-mario,jevclient,jev-doom-agent}.md`. Line numbers below
8refer to `llms-full-2026-09-19.txt` unless stated otherwise.
9
10---
11
12## 1. Endpoint, auth, headers, model field
13
14**Evaluation endpoint** (line 117-123, repeated at 12469-12473):
15```http
16POST https://api.typesafe.ai/v1/systemone
17Authorization: Bearer <API_KEY>
18Content-Type: application/json
19```
20
21**Models list endpoint** (line 12888-12896):
22```
23GET https://api.typesafe.ai/v1/models
24Authorization: Bearer <API_KEY>
25```
26
27- Auth header: exactly `Authorization: Bearer <API_KEY>`. No other header is
28  documented as required beyond `Content-Type: application/json` for POST.
29- Base URL default (Python SDK constant, line 18134-18140):
30  `DEFAULT_BASE_URL = 'https://api.typesafe.ai'` (note: **no** `/v1` in the
31  SDK's base URL constant — the SDK presumably appends `/v1/systemone`
32  itself; not spelled out further in the docs).
33- Request-ID response header: `x-typesafe-request-id` (lines 14530, 18230,
34  19314, etc.) — read via `response.request_id` on the Python client, or
35  `error.request_id` on a raised exception.
36- `model` field: string, e.g. `"jev-latest"`. Selects which model/alias
37  handles the request (line 133-135, 12837).
38
39---
40
41## 2. Request JSON schema
42
43### Top-level fields (line 125-158)
44
45```json
46{
47  "state": "Help! My payouts have been failing for 3 days.",
48  "model": "jev-latest",
49  "questions": {
50    "is_urgent": {
51      "type": "noul",
52      "instructions": "Does this convey urgency?"
53    }
54  }
55}
56```
57
58- `state` (`string | object | array`, required): the content to evaluate.
59  See §5 for shaping guidance.
60- `model` (`string`, required in examples but presumably has an SDK-side
61  default `jev-latest` — vendor docs don't state the HTTP API has a
62  server-side default if `model` is omitted; not in vendor docs whether a
63  raw HTTP POST without `model` is accepted).
64- `questions` (`map<string, Question>`, required): caller-chosen keys →
65  typed question objects. "The key is not sent to the underlying model and
66  is not used in inference" (line 142-143).
67
68### Question types — shared fields
69
70All three share `type` and `instructions`; each adds its own `criteria`
71(line 162).
72
73#### Noul (yes/no, line 164-203)
74
75| Field | Required | Shape |
76|---|---|---|
77| `type` | yes | `"noul"` |
78| `instructions` | yes | `string \| object \| array` — the yes/no question |
79| `criteria` | **no** | optional `{true, false}` — string/object/array descriptions of what a yes and a no mean |
80
81```json
82{
83  "state": "Help! My payouts have been failing for 3 days.",
84  "model": "jev-latest",
85  "questions": {
86    "is_urgent": {
87      "type": "noul",
88      "instructions": "Does this convey urgency?",
89      "criteria": {
90        "true": "Explicitly time-sensitive",
91        "false": "No urgency expressed"
92      }
93    }
94  }
95}
96```
97
98#0 Noul does **not** have a minimum/maximum on criteria (it's just an
99optional true/false pair, not a list).
100
101#### Choice (pick one option, line 205-241, 13469-13513)
102
103| Field | Required | Shape |
104|---|---|---|
105| `type` | yes | `"choice"` |
106| `instructions` | yes | `string \| object \| array` |
107| `criteria` | yes | `map<string, string \| null>` — option name → description (`null` = no extra detail) |
108
109**Limit: up to 255 options per Choice question** (line 13552: "A Choice
110question accepts up to 255 options, and adding options costs a few tokens
111each"). No stated minimum, but two is the practical floor (a single-option
112Choice is degenerate).
113
114```json
115{
116  "state": "Help! My payouts have been failing for 3 days.",
117  "model": "jev-latest",
118  "questions": {
119    "department": {
120      "type": "choice",
121      "instructions": "Which team should handle this?",
122      "criteria": {
123        "billing": "Payments, invoicing, refunds",
124        "technical": "Bugs, outages, integrations",
125        "sales": "Pricing, upgrades, new accounts"
126      }
127    }
128  }
129}
130```
131
132#### Score (position on an ordered rubric, line 243-269, 13875-13927)
133
134| Field | Required | Shape |
135|---|---|---|
136| `type` | yes | `"score"` |
137| `instructions` | yes | `string \| object \| array` |
138| `criteria` | yes | ordered `array` of level descriptions, low→high. **"Needs at least two levels and takes up to 10."** (line 13881) |
139
140A level's number is its zero-based position in the array (line 13891). The
141model is never shown the level's number or neighbours — only the
142description text (line 14044), so numeric-only level descriptions ("0",
143"1", "2") score badly (worked example, line 14046-14052: same report scores
1440.0/confidence 1.0 with descriptive levels vs. 0.57/confidence 0.35 with
145bare numeral levels).
146
147```json
148{
149  "state": "Help! My payouts have been failing for 3 days.",
150  "model": "jev-latest",
151  "questions": {
152    "frustration": {
153      "type": "score",
154      "instructions": "How frustrated is the customer?",
155      "criteria": ["Calm", "Frustrated", "Very angry"]
156    }
157  }
158}
159```
160
161### Structured `instructions`/`criteria` (line 13367-13446)
162
163Every one of `instructions`, Choice's per-option description, Score's
164per-level description, and Noul's `criteria.true`/`criteria.false` accepts
165`string | object | array | null` (the JS SDK type alias `EntryType`, line
16613376-13383). Field names inside a structured object (e.g. `question`,
167`focus`, `what`, `not_for`, `examples`) are **not part of the API and none
168are reserved** — "You choose them... use short names that label what
169follows" (line 13769). Use this when two options/levels are easily
170confused, or a level needs "what it covers" + "example situations" fields
171(worked examples at line 13740-13767 and 14199-14263 both show measurable
172accuracy/confidence improvement from adding an `examples` array).
173
174Reference a specific nested `state` field by dot-and-index path **inside
175backticks** in the instructions text, e.g. `` `ticket.messages[0].text` ``
176(line 530-536, 13227-13269) — this is a documented convention, not a
177separate API field.
178
179---
180
181## 3. Response JSON schema
182
183### Top-level (line 271-310)
184
185```json
186{
187  "model": "jev-latest",
188  "answers": {
189    "is_urgent": { "type": "noul", "noul": 0.92 }
190  },
191  "usage": { "input_tokens": 312, "output_tokens": 48 }
192}
193```
194
195- `model` (string): the model that actually answered (the resolved
196  versioned ID, not necessarily the alias sent).
197- `answers` (`map<string, Answer>`): one answer per question id you chose.
198- `usage.input_tokens` / `usage.output_tokens` (integers).
199
200### Answer shapes (line 312-414; JSON Schemas at line 19086-19219 from the
201Python SDK's pydantic models — these are the authoritative field-level
202schemas)
203
204**Noul answer** — `{"type": "noul", "noul": <float 0..1>}`. No confidence
205field. "`noul` ranges from 0 to 1, representing the probability that the
206answer is **yes**" (line 13819).
207
208**Choice answer**:
209```json
210{
211  "type": "choice",
212  "choice": "technical",
213  "probabilities": { "billing": 0.08, "technical": 0.85, "sales": 0.07 },
214  "confidence": 0.82
215}
216```
217- `choice` (string): highest-probability option name.
218- `probabilities` (`map<string, number>`): every option → probability,
219  "floats that sum to 1" (approximately — the pydantic schema description
220  says "values sum to approximately 1", line 19115).
221- `confidence` (number 0-1): derived from the shape of `probabilities`.
222
223**Score answer**:
224```json
225{
226  "type": "score",
227  "score": 1.6,
228  "legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
229  "probabilities": { "0": 0.05, "1": 0.3, "2": 0.65 },
230  "confidence": 0.78
231}
232```
233- `score` (number): probability-weighted mean of level indices — "each
234  level number multiplied by its probability, added up" (line 13963),
235  e.g. `0×0.0 + 1×0.70 + 2×0.30 = 1.30`. Can land between integer levels.
236- `legend` (`map<string level-index, string|object|array>`): echoes each
237  level's description back, keyed by the level's index as a **string**.
238- `probabilities` (`map<string level-index, number>`), string-keyed like
239  `legend`.
240- `confidence` (number 0-1).
241- **Python SDK note (line 13969):** `ScoreAnswer` re-keys `probabilities`
242  and `legend` by **integer** level in the typed SDK object, even though
243  the wire JSON keys them as strings.
244
245### Full worked request/response pair (Quick Start, line 12495-12564)
246
247Request:
248```json
249{
250  "state": "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
251  "model": "jev-latest",
252  "questions": {
253    "department": {
254      "type": "choice",
255      "instructions": "Which team should handle this",
256      "criteria": {
257        "billing": "Payment or subscription issues",
258        "technical": "Bugs or integration problems",
259        "sales": "Pricing or account questions"
260      }
261    },
262    "frustration": {
263      "type": "score",
264      "instructions": "How frustrated the customer appears",
265      "criteria": [
266        "Calm, just stating facts",
267        "Frustrated but civil",
268        "Very angry, strong language"
269      ]
270    },
271    "is_urgent": {
272      "type": "noul",
273      "instructions": "The message conveys urgency or time-sensitivity"
274    }
275  }
276}
277```
278
279Response:
280```json
281{
282  "model": "jev-latest",
283  "answers": {
284    "department": {
285      "type": "choice",
286      "choice": "billing",
287      "probabilities": { "billing": 0.84, "technical": 0.159, "sales": 0.001 },
288      "confidence": 0.596
289    },
290    "frustration": {
291      "type": "score",
292      "score": 1.035,
293      "legend": {
294        "0": "Calm, just stating facts",
295        "1": "Frustrated but civil",
296        "2": "Very angry, strong language"
297      },
298      "confidence": 0.842
299    },
300    "is_urgent": { "type": "noul", "noul": 0.999 }
301  },
302  "usage": { "input_tokens": 312, "output_tokens": 48 }
303}
304```
305(Note: this particular published example response for `frustration` omits
306`probabilities` in the docs' own JSON, even though the schema requires it —
307likely a docs-authoring trim, not a real absent-field case; every other
308Score example in the docs includes `probabilities`.)
309
310### `GET /v1/models` response (line 12888-12933)
311
312```json
313{ "models": [ { "name": "jev-latest", "description": "...", "release_date": "..." } ] }
314```
315- `models[].name` (string): the model ID/alias, valid in the `model`
316  request field.
317- `models[].description` (string).
318- `models[].release_date` (string).
319- "It currently lists the aliases. Versioned IDs such as `jev-1.13.0` are
320  accepted by the `model` field whether or not they appear in the list."
321
322---
323
324## 4. Errors, rate limits, retry semantics
325
326### HTTP API error table (line 416-425 — the only error table in the HTTP
327API reference)
328
329| Status | Meaning |
330|---|---|
331| `401 Unauthorized` | Missing or invalid API key. Check the `Authorization` header. |
332| `422 Unprocessable Entity` | Request body failed validation — missing required field or malformed question. "The body details the offending field." (exact shape of that body is **not in vendor docs**.) |
333| `429 Too Many Requests` | Exceeded rate limit. Back off and retry after a short delay. |
334| `529 Overloaded` | TypeSafe is temporarily overloaded. Retry after a short delay. |
335
336"When you receive a `429` or `529`... retry the request with exponential
337backoff instead of retrying immediately. Our client SDKs handle this
338automatically... with its default retry policy." (line 427-429)
339
340**Disagreement/gap noted:** the HTTP API reference documents only
341401/422/429/529, but the **Python SDK's exception hierarchy** (line
34218165-18339) is broader — it defines dedicated exceptions for `400`
343(`TypeSafeBadRequestError`), `403` (`TypeSafePermissionDeniedError`), `404`
344(`TypeSafeNotFoundError`), `422` (`TypeSafeUnprocessableEntityError`), `429`
345(`TypeSafeRateLimitError`), and generic `5xx`
346(`TypeSafeInternalServerError` — which would presumably catch `529` too,
347though `529` is never named in the SDK docs). The HTTP reference page is
348the narrower, curated list; the SDK exception surface is the fuller
349(implicit) one. Treat both as authoritative for their own scope.
350
351### Python SDK exception classes (line 18165-18432)
352
353- `TypeSafeError(Exception)` — base for all SDK failures.
354- `TypeSafeAPIError(TypeSafeError)` — unsuccessful HTTP response. Fields:
355  `status` (int), `body` (parsed JSON error body / plain text / `None` for
356  empty), `headers`, `endpoint` (method+URL, no creds/query/fragment),
357  `request_id` (property, from `x-typesafe-request-id` or `None`).
358  - `TypeSafeBadRequestError` (400)
359  - `TypeSafeAuthenticationError` (401)
360  - `TypeSafePermissionDeniedError` (403)
361  - `TypeSafeNotFoundError` (404)
362  - `TypeSafeUnprocessableEntityError` (422)
363  - `TypeSafeRateLimitError` (429) — adds `retry_after_ms` (`parse_retry_after(headers)`, `None` if unavailable)
364  - `TypeSafeInternalServerError` (5xx)
365- `TypeSafeAPIConnectionError(TypeSafeError, ConnectionError)` — no HTTP
366  response at all (network failure).
367  - `TypeSafeAPITimeoutError(TypeSafeAPIConnectionError, TimeoutError)` —
368    adds `timeout` (seconds or `httpx2.Timeout`).
369- `TypeSafeAPIResponseValidationError(TypeSafeAPIError)` — 2xx response
370  whose body was structurally invalid; adds `field_path` (dotted path to
371  the offending field, e.g. `answers.tone.confidence`).
372
373Rate-limit headers honored: **`Retry-After`** and **`retry-after-ms`**
374(line 16758, 18398) — "Whether to honor `Retry-After` and `retry-after-ms`
375response headers", up to `maxRetryAfterMs` in the JS SDK's RetryPolicy.
376
377### `RetryPolicy` (Python SDK, line 18346-18430)
378
379```python
380from typesafe_sdk import RetryPolicy, TypeSafeClient
381client = TypeSafeClient(
382    retry=RetryPolicy(max_retries=3, timeout=10.0, http_statuses={429, 500, 502, 503, 504})
383)
384```
385Fields: `max_retries` (0 disables), `backoff_initial` (seconds, doubled
386each attempt up to `backoff_max`; 0 disables backoff), `backoff_max`,
387`backoff_jitter` (fraction randomly subtracted, 0-1), `http_statuses`
388(retried set), `respect_retry_after` (bool), `api_connection_error` (bool,
389retry on connection failure), `api_timeout_error` (bool), `exceptions`
390(extra exception types to retry), `predicate` (callable on the raised
391exception → bool), `timeout` (total retry budget in seconds across the
392whole call including delays — "Stops before a retry whose delay would
393reach or exceed the budget, re-raising the last error").
394
395### Rate limits and context/pricing (Models page, line 12839-12934 — the
396authoritative numbers table)
397
398| Jev 1.13 | `jev-1.13.0` |
399|---|---|
400| Price (per Btok / per Mtok) | **$42 / $0.042** — input tokens only; **output tokens are free** |
401| Rate limits | **250,000 tokens/second** / **1,200 requests/minute** |
402| Context length | **64k tokens per request** (state + all questions combined); **32k tokens for state + the single longest question** |
403| Input | Text only: string, JSON object, array of text values. No image/audio/video. |
404
405"**Rate limits are adjusting dynamically**... can change without notice
406while we [serve growing demand]... Higher limits are available on custom
407and enterprise plans." (line 12853-12855) — i.e. these are not guaranteed
408stable numbers.
409
410**Internal inconsistency in the vendor docs, flagged not resolved:** the
411Primitives page (line 13328) separately states "The number of questions in
412one request is limited only by the request's token budget, which the state
413and the questions share. The budget is around **32,000 tokens, roughly
414150,000 characters** of English text." This describes the *whole request's*
415budget as ~32k, which conflicts with the Models page's more precise split
416(64k total request / 32k for state + longest single question). Vendor docs
417win per the task's instruction, but the two vendor passages themselves
418disagree on which number (32k or 64k) is the ceiling for a many-question
419request — treat 64k (Models page, more specific and more recently
420authored-sounding) as the more load-bearing number for total request
421budget, and 32k as a per-state-plus-single-question sub-limit.
422
423### Aliases (line 12857-12870)
424
425| Alias | Points to | Meaning |
426|---|---|---|
427| `jev-latest` | `jev-1.13.0` | Most recent stable/official release. SDK default. |
428| `jev-preview` | `jev-1.13.0` | Most recent release whether or not official; currently identical to `jev-latest` (no preview build live). |
429
430"If you have tuned confidence thresholds against a specific version, pin
431that version's ID instead of the alias."
432
433---
434
435## 5. State shaping, calibration, sampling — what the docs say to do
436
437### State (line 843-902)
438
439- `state` is `string | JSON object | array`; text only, no images/audio/
440  video (repeated at line 871, 916-918, 12846, 12851).
441- Prefer an **object** for most requests "so each part of the state has a
442  descriptive name and its relationships remain clear." A plain string is
443  fine only for one simple piece of text (line 868).
444- Every question in a request sees the **same** state and is evaluated
445  independently — "one primitive's result does not become hidden context
446  that changes another primitive's result" (line 481, 850).
447- Keep only the context relevant to the current questions — "This helps
448  the model avoid distractions and context rot" (line 522).
449- Point questions at specific nested values with a backticked dot/index
450  path, e.g. `` `support.tickets[0].message` `` (line 530-536).
451- English is the primary training language; other languages including CJK
452  are accepted but currently lower accuracy (line 871, 12882).
453
454### Model jaggedness — `jev-1.13` failure modes and mitigations (line
45512677-12830, table at 12690-12700)
456
4571. **Literal reading** — answers the words written, not the intent. Put
458   boundary cases explicitly in `criteria`; split ambiguous judgments into
459   two literal questions and combine in code.
4602. **Math and numbers** — "Jev is not a calculator." Don't ask it to count
461   (characters, occurrences, list items) — count in code (worked example:
462   ask one `Noul` per item, sum `> threshold` results in code, line
463   12720-12737). Numeric representations (hex/RGB, low-level code) score
464   worse than named/semantic equivalents — convert in code, pass semantic
465   labels. **Score outputs must not be used to reconstruct an exact number
466   by interpolating between levels** — only threshold the expectation.
4673. **Date/time comparison** — dates are read as text, not ordered
468   quantities; ordering/duration/window checks are unreliable, worse with
469   mixed formats. Extract components as bounded `Choice`s (month/day/year
470   as closed sets, with an explicit "not stated" option), assemble and do
471   all arithmetic in code.
4724. **Indirection** — double negatives / multi-hop "property of a property"
473   questions lose accuracy. Write instructions directly; name relevant
474   state parts explicitly.
4755. **Large state full of irrelevant detail** — accuracy falls as
476   unrelated content grows in `state` ("context rot" reiterated). Filter/
477   retrieve in code first; when that's not possible, use a `Noul` as a
478   relevance filter first.
4796. **Adversarial content** — `state` is treated as data, not hostile, by
480   default; injected instructions / misleading framing can move the
481   answer. Be explicit in criteria; test adversarial inputs before
482   deploying.
4837. **Contradictory instructions/criteria** — e.g. a Noul where `true` maps
484   to "no" semantically will perform worse. Align criteria as an extension
485   of instructions, plain everyday phrasing.
4868. **Common-sense structural invariants do NOT hold** — worked numeric
487   example: the same yes/no judgment asked as `Noul` vs. `Choice` gave
488   `noul=0.22` vs. `probabilities["yes"]=0.01` — not comparable. Two
489   Nouls for "X" and "not X" summed to 1.19, not 1.0. **Don't carry a
490   threshold tuned on one question type over to another; don't expect
491   arithmetic identities between separate questions.**
4929. **Generation** — Jev is not trained to generate text; forcing free text
493   via chained choices "will not work well and will be very slow." Turn
494   bounded extraction into a `Choice`; use a real generative model for
495   actual text generation.
496
497Summary reminders (line 12818-12825): avoid asking Jev something code can
498compute exactly, hiding multiple judgments in one question, multi-hop
499"System Two" tasks, or feeding more `state` context than a question needs.
500
501### Confidence / calibration (line 1149-1230)
502
503- `confidence` (0-1) is present on **Choice and Score answers only**, not
504  Noul. It is a single-number summary of how peaked/flat the
505  `probabilities` distribution is (line 1156, 1160).
506- "We provide `confidence` as a convenient measure that fits most
507  use-cases, but you are never locked into our definition... a different
508  measure may serve you better... which is exactly why we give you the
509  full `probabilities`" (line 1163) — i.e. compute your own statistic from
510  `probabilities` if `confidence` doesn't fit.
511- Calibration is trained via **RLCD** ("Reinforcement Learning for
512  Calibrated Decisions" per the brain notes; vendor docs link it as
513  "[RLCD](/introduction/machine-learning-primer)" without spelling out the
514  acronym inline at line 493). "Calibration is measured across groups of
515  predictions; it does not guarantee that an individual answer is
516  correct." (line 922)
517- Recommended pattern: **three confidence bands** — high (act
518  automatically), medium (confirm/flag/gather more info), low (route to a
519  human/fallback) — with threshold values that scale with the stakes of
520  the specific action, not one global number (line 1174-1229, worked
521  banking-command example).
522- No documented `temperature` or sampling parameter exists on the API
523  itself — Jev always returns a full deterministic-per-call probability
524  distribution; there's no server-side "sampling temperature" request
525  field anywhere in the vendor docs. (**Not in vendor docs**: any request
526  field to control determinism/temperature/top-k on the Jev call itself.)
527  Client-side sampling *from* the returned distribution — as `agent.py`
528  does — is a caller-side technique, not a documented API feature.
529
530---
531
532## 6. From `agent.py` (jev-plays-pokemon) — concrete harness patterns
533
534Full source read; path:
535`/home/nixos/brains/personal/raw/code/jev-plays-pokemon/agent.py`
536(+ `README.md` beside it).
537
538### Imports / SDK usage confirmed
539
540```python
541from typesafe_sdk import Choice, Noul, RetryPolicy, TypeSafeClient
542```
543This matches the **official** `typesafe-sdk` Python package from the
544vendor docs exactly (not `jevclient`, which is an unrelated third-party
545package — see §8).
546
547### Environment variables it reads
548
549- `TYPESAFE_API_KEY` — checked directly with `os.getenv`, raises
550  `RuntimeError("TYPESAFE_API_KEY is not set. Create a .env file with your key.")`
551  if absent, **before** ever constructing `TypeSafeClient` (the SDK itself
552  would also raise `TypeSafeError` for a missing key, but the agent
553  pre-checks with its own message).
554- `TYPESAFE_MAX_RETRIES` (default `4`)
555- `TYPESAFE_BACKOFF_INITIAL` (default `1.0` seconds)
556- `TYPESAFE_BACKOFF_MAX` (default `8.0` seconds)
557- `TYPESAFE_TIMEOUT` (default `45.0` seconds) — this is the **retry
558  budget** (`RetryPolicy.timeout`), not the SDK's default 10s
559  per-HTTP-operation timeout.
560- `TYPESAFE_TEMPERATURE` (default `2.0`) — **agent-defined**, not an SDK/API
561  concept; used only in `_sample_goal` (client-side).
562- `TYPESAFE_GOAL_FLOOR` (default `0.1`) — same, agent-defined.
563- Loaded via `python-dotenv`'s `load_dotenv()` at import time, from a
564  project-root `.env` file (README: `cp .env.example .env`).
565
566### Retry policy construction (verbatim)
567
568```python
569def retry_policy() -> RetryPolicy:
570    return RetryPolicy(
571        max_retries=_int("TYPESAFE_MAX_RETRIES", 4),
572        backoff_initial=_float("TYPESAFE_BACKOFF_INITIAL", 1.0),
573        backoff_max=_float("TYPESAFE_BACKOFF_MAX", 8.0),
574        backoff_jitter=0.2,
575        timeout=_float("TYPESAFE_TIMEOUT", 45.0),
576    )
577```
578`backoff_jitter=0.2` is hardcoded (not env-overridable). `RetryPolicy` is
579passed once at client construction (`TypeSafeClient(retry=retry_policy(), ...)`),
580not per-call.
581
582### How it structures state and questions for a game loop
583
584Two call shapes, both against the same `state: dict[str, Any]` built
585elsewhere (`state.py`, not requested for this digest) and the **same
586shared set of per-button `Noul` questions** (`self.action_questions`,
587built once in `__init__` via `build_questions()`):
588
589```python
590def build_questions() -> dict[str, Any]:
591    return {
592        action: Noul(instructions=f"Is the best single action right now to {ACTION_NAMES[action]}?")
593        for action in ACTION_FRAMES
594    } | {
595        "menu_open": Noul(
596            instructions=(
597                "Is a selectable menu open on screen (start menu, POKEMON/BAG/ITEM list, "
598                "battle FIGHT/PKMN/BAG/RUN, or a yes/no question)? A plain text dialog is NOT a menu."
599            )
600        ),
601    }
602```
603— 8 action Nouls (`press_a`, `press_b`, `press_start`, `walk_up/down/left/right`,
604`wait`) + 1 `menu_open` Noul, all fired in **one call**, every turn. Comment
605at file top: "Uses one atomic Noul per candidate action (asked in parallel
606in a single call)... Code then picks the strongest answer with a margin
607requirement and sanity-checks it against the deterministic walkability
608map." This is the vendor docs' "speculative fan-out" pattern applied
609directly.
610
611`decide_goal()` adds **one `Choice`** (`next_goal`, options = the
612currently-available high-level goals, e.g. `talk_to_Mom`, `explore`,
613`reach_exit`) to the same 9 action Nouls, **all in a single
614`client.system_one(state=state, questions=questions)` call**:
615
616```python
617questions = {
618    "next_goal": Choice(
619        instructions={
620            "question": "Which goal should RED pursue next?",
621            "focus": "Pick the ONE goal that best advances your long-term goal (badges/Champion) "
622            "given the screen text, recent hints, nearby objects, walkability and room map. "
623            "Code will execute the walking.",
624        },
625        criteria=criteria,  # {goal_id: goal_desc, ...}
626    )
627}
628questions.update(self.action_questions)
629response = self._interruptible(self.client.system_one, state=state, questions=questions)
630```
631Note the structured `instructions` object (`{"question": ..., "focus": ...}`)
632for the Choice — exactly the vendor docs' "structured instructions" pattern
633(§2 above), with caller-chosen, non-reserved field names.
634
635`decide_action()` (menu/battle micro-decisions) sends **only** the 9 action
636Nouls, no Choice.
637
638### How it samples from the probability distribution (`_sample_goal`)
639
640The vendor docs never describe server-side sampling — this is 100%
641client-side post-processing of the returned `choice.probabilities`:
642
643```python
644def _sample_goal(
645    probabilities: dict[str, float],
646    temperature: float = 2.0,
647    floor: float = 0.05,
648) -> tuple[str, float]:
649    labels = list(probabilities)
650    ...
651    probs = [max(0.0, probabilities[g]) for g in labels]
652    total = sum(probs)
653    probs = [p / total for p in probs]
654    # Floor: guarantee every option keeps some mass (p=1 is never 100%).
655    floor = max(0.0, floor)
656    probs = [p + floor for p in probs]
657    total = sum(probs)
658    probs = [p / total for p in probs]
659    # Softmax flattening: logits = log(p)/T. T>1 pulls the peak down and
660    # lifts the tail, keeping a real chance of doing something else.
661    if temperature != 1.0 and temperature > 0:
662        logits = [math.log(p + 1e-12) / temperature for p in probs]
663        m = max(logits)
664        exp = [math.exp(x - m) for x in logits]
665        exp_total = sum(exp)
666        probs = [e / exp_total for e in exp]
667    label = random.choices(labels, weights=probs, k=1)[0]
668    return label, probs[labels.index(label)]
669```
670Two composed transforms before `random.choices`: (1) an additive **floor**
671so no option ever reaches literal 0% or literal 100% mass; (2) **softmax
672flattening** of `log(p)/T` with `T=2.0` default so a confidently-wrong
673p=0.99 pick fires only "~80%" per the docstring, ~60-70% per the README —
674i.e. deliberately de-sharpening Jev's own calibrated distribution to avoid
675looping on a single wrong-but-confident goal every turn. The Noul actions
676(`decide_action`) are **not** sampled this way — they use `argmax` with a
677margin gate instead (below), because action selection wants determinism +
678a safety fallback, not exploration.
679
680### Anti-stuck / safety tricks (`pick_action`, deterministic, no model call)
681
682- **Margin + confidence gate**: only trust the model's top Noul pick if
683  `decision.margin >= 0.08 and decision.confidence >= 0.45` (`margin` =
684  winning `noul` minus runner-up `noul`). Below that, fall back to
685  advancing dialog / open menu (`press_a`) / walking toward an
686  `explore_hint` / `wait`, in that priority order.
687- **Menu override**: if `menu_open` Noul ≥ 0.5 and the winning action's
688  confidence < 0.7, force `press_a` (confirm highlighted option) —
689  "reliably selects the starter (and most story choices)."
690- **Deterministic wall check**: for any `walk_*` action, cross-check
691  against a ground-truth `walkable: dict[str, bool]` map read from
692  emulator RAM (not from the model). If the chosen direction is blocked
693  and not every direction is blocked and not explicitly probing for an
694  exit, override to `press_a` (if a dialog is active) or `wait` — i.e. the
695  model's spatial judgment is never trusted over the deterministic
696  collision map except when probing for exits.
697- **Anti-pacing**: never immediately reverse the previous walk
698  (`_REVERSE` map) if another open direction exists — picks the
699  next-best-scoring open direction instead, gated at `noul >= 0.3`.
700- Higher-level (README, not in `agent.py` itself): stuck detection
701  (press B → wait → re-route) and exit-probing when a room is fully
702  explored live in `play.py`/`navigation.py`, not shown in this file.
703
704### Threading / interruptibility (`_interruptible`)
705
706Every blocking `client.system_one` call is run on a **daemon worker
707thread** while the main thread polls a `stop_requested()` flag every 50ms
708and raises `KeyboardInterrupt` itself — because "the HTTP stack swallows
709SIGINT while it is running." This is a harness-level workaround, not
710documented anywhere in the vendor SDK docs (**not in vendor docs**: any
711statement that the SDK/httpx2 swallows SIGINT — this is the agent
712author's own empirical finding, not a cited vendor fact).
713
714### Measured cost/latency (README, not vendor docs — first-party but
715project-specific measurement)
716
717- "Jev is fast (tens–hundreds of ms per decision)."
718- `--turbo` mode note: "the Jev API call per decision (~0.5 s) becomes the
719  pacing factor" — i.e. ~500ms observed round-trip in this specific
720  harness/network conditions, at the high end of "tens-hundreds of ms."
721- Token estimate per decision: shared instructions ~850 tok + questions
722  (8 Nouls + `menu_open`) ~350 tok + per-turn state ~400-900 tok ≈
723  **1,600-2,100 input tokens/decision**.
724- Cost at $0.042/MTok: ≈$0.27/hour at 1 decision/sec, ≈$0.76 per 10,000
725  decisions.
726
727---
728
729## 7. Official Python SDK vs. raw HTTP
730
731**Yes, an official SDK exists**: package `typesafe-sdk` on PyPI (line
73212573, 17492-17500), requires Python ≥ 3.10 (line 12570) — note `agent.py`'s
733own project requires Python ≥ 3.14 per its README, which is the harness's
734own constraint, not the SDK's floor.
735
736```bash
737pip install typesafe-sdk
738# or
739uv add typesafe-sdk
740```
741
742Minimal usage (sync, line 17539-17561 / 12583-12617):
743```python
744from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
745
746with TypeSafeClient() as client:  # reads TYPESAFE_API_KEY from env
747    response = client.system_one(
748        state={"document": "I was charged twice. Please fix this ASAP."},
749        questions={
750            "billing": Noul(instructions="Is this ticket about billing?"),
751            "tone": Choice(
752                instructions="What is the customer's tone?",
753                criteria={"calm": None, "frustrated": None, "angry": None},
754            ),
755            "urgency": Score(
756                instructions="How urgent is this ticket?",
757                criteria=["can wait", "this week", "today"],
758            ),
759        },
760    )
761
762print(response.nouls["billing"].noul)
763print(response.choices["tone"].choice)
764print(response.scores["urgency"].score)
765```
766
767Async variant: `AsyncTypeSafeClient` / `await client.system_one(...)`,
768identical parameter shape, requires `async with` and `await`.
769
770`TypeSafeClient` / `AsyncTypeSafeClient` constructor params (line
77117600-17635): `api_key` (→ `TYPESAFE_API_KEY`), `model` (→
772`TYPESAFE_DEFAULT_MODEL`), `retry` (`RetryPolicy`), `timeout` (float
773seconds or `httpx2.Timeout`; SDK default `DEFAULT_TIMEOUT = 10.0`),
774`headers` (extra), `transport` / `http_client` (mutually exclusive custom
775httpx2 transport), `base_url` (→ `TYPESAFE_BASE_URL`). Raises
776`TypeSafeError` if API key missing or timeout invalid; `ValueError` if both
777`transport` and `http_client` given.
778
779`client.system_one(...)` params (line 17685-17735): `state`
780(`JSONContent`), `questions` (`Mapping[str, Question]`, non-empty — raises
781`TypeSafeError` if empty, or if a Score question's `criteria` list is
782empty), `model` (override), `retry` (per-call override), `timeout`
783(per-call override), `extra_headers`, `extra_body` (forward-compat, shallow
784last-write-wins merge over the body), `response_model` (optional Pydantic
785`BaseModel` subclass to type the response — see forward-compat section
786below).
787
788`client.models.list(...)` → `ListModelsResponse` with `.models[]` of
789`ModelMetadata{name, description, release_date}` (line 19770-19925).
790
791Environment variables the SDK itself reads (line 18094-18172, restated at
79220161-20172 — **this table is the authoritative env-var reference**):
793
794| Variable | Configures | Default |
795|---|---|---|
796| `TYPESAFE_API_KEY` | API key (required) | — |
797| `TYPESAFE_BASE_URL` | API root URL | `https://api.typesafe.ai` |
798| `TYPESAFE_DEFAULT_MODEL` | Default model | `jev-latest` |
799| `TYPESAFE_LOG_LEVEL` | `typesafe_sdk` logger level, applied once at import | unset (off) |
800
801"Explicit options take precedence over environment variables; empty or
802whitespace-only environment values are ignored." (line 17592)
803
804Logging: SDK logs to Python logger name `typesafe_sdk`. `info` = one
805summary line/request; `debug` = also headers+bodies. **Secret headers**
806(authorization, API keys, cookies, any header with `token`/`secret` in the
807name) are redacted; **request/response bodies are NOT redacted** (line
80820159).
809
810Forward-compatibility escape hatches (line 20174-20228):
811- `extra_body={"beam_width": 4}` — send request fields the SDK predates.
812- Raw dict questions instead of `Noul`/`Choice`/`Score` objects, including
813  extra unknown keys (e.g. `{"type": "noul", "instructions": "...", "weight": 2}`)
814  — "Ignore their type-checking errors and prefer upgrading the SDK
815  instead."
816- Unknown answer kinds: SDK logs a warning and skips them; use
817  `result.raw_http_response.json()["answers"]` to see everything including
818  unrecognized kinds.
819- Unknown extra fields on recognized responses are silently ignored.
820
821**JavaScript/TypeScript SDK also exists** (`@typesafe-ai/sdk`,
822`TypeSafeClient.systemOne(...)`, line 14286-17471) — not needed for a
823Python harness, only noted for completeness; its error class names
824(`APIConnectionError`, `RateLimitError` with `retryAfterMs`, etc.) mirror
825the Python SDK's one-for-one.
826
827**Raw HTTP is fully viable** — the cURL example (line 12477-12493) and the
828plain-dict "raw question dictionaries" pattern above show the wire format
829is a stable, documented contract independent of any SDK; a hand-rolled
830`requests`/`httpx` client needs only: POST to
831`https://api.typesafe.ai/v1/systemone`, `Authorization: Bearer` header,
832JSON body per §2, and to implement retry/backoff itself for 429/529 (with
833`Retry-After`/`retry-after-ms` header awareness) since raw HTTP gets none
834of the SDK's automatic retry behavior.
835
836---
837
838## 7b. Measured against the live API, 2026-09-20 (`apps/jevprobe`, `apps/player`)
839
840Everything in sections 1-7 was read, not run. This section is the opposite:
841what the API actually did on this machine, from a hand-rolled raw-HTTP client
842(`packages/jev-http`) rather than an SDK.
843
844**Nothing in the digest needed correcting.** The request shape in §2 was
845accepted as documented, and the response matched §3 field for field - no
846unknown fields, `probabilities` over every option, `confidence` on the Choice,
847`usage` with both token counts. Raw response to a one-Choice request:
848
849```json
850{"model":"jev-1.13.0","answers":{"move":{"type":"choice","choice":"walk_south",
851 "confidence":0.98,"probabilities":{"press_a":0.01,"walk_south":0.99,
852 "walk_north":0.0,"walk_west":0.0}}},"usage":{"input_tokens":544,"output_tokens":54}}
853```
854
855- **`model` in the response is the resolved versioned id**, `jev-1.13.0`, for a
856  request that sent the alias `jev-latest` - exactly as §3 says, now seen.
857- **Latency.** §8 collects five conflicting figures and notes none is in the
858  API reference. Measured here, from a NixOS host over residential fibre, on a
859  request of ~700 input tokens carrying one Choice, one Score and one Noul:
860  **cold connection ~400 ms; warm 110-280 ms, typically ~160 ms**, over 755
861  consecutive requests. Keeping one client alive is therefore worth more than
862  twice the cost of the answer itself, which is the argument for a client with
863  a connection pool rather than one TLS handshake per decision.
864- **Cost.** 755 three-question decisions came to **$0.0277** - about
865  **$0.000037 each**, at 550-900 input tokens per request. At one decision a
866  second that is $0.13 an hour. The budget cap in `packages/brain` is a stop,
867  not a budget.
868- **Every request succeeded on the first attempt.** No 429, no 529, no
869  connection failure in 755 requests plus a few hundred more across earlier
870  runs, so the retry policy in §4/§6 is implemented and has never yet fired.
871  That is the shape of the absence, not evidence it is unnecessary.
872- **Browser origins are still refused** (recorded in the brain page, retested
873  by construction here): the native client uses the API directly, and a web
874  build will need a relay.
875
876---
877
878## 8. Cross-source disagreements (vendor docs win; noted for awareness)
879
8801. **Latency figures vary by source and none is in the API reference
881   itself:**
882   - Vendor docs prose (not the API reference table): "Most queries
883     complete in about 100 ms" (line 489) and "Frontier intelligence at
884     real-time speeds (150ms)" (line 974) — both **marketing/concepts
885     pages**, not a guaranteed SLA number.
886   - `jev-plays-pokemon` README (first-party measurement, not vendor):
887     "tens–hundreds of ms per decision," ~0.5s observed in turbo mode.
888   - Brain page `jev.md` (secondary, citing TechCrunch/flaviocopes, not
889     vendor docs): "70ms-500ms" end-to-end.
890   - Brain page `ecosystem/jev-doom-agent.md` (secondary, citing
891     TypeSafe's own launch blog — a different document from the docs site
892     that was the assigned source): "roughly 10 decisions per second"
893     (~100ms) at "~$7/hour."
894   - Brain page `ecosystem/jevclient.md` (third-party unofficial SDK's own
895     measured table, Netherlands, cold-connection): 690-1,336ms for 3-400
896     questions in one call; ~250-580ms warm-connection.
897   - **None of these is inside `llms-full-2026-09-19.txt`'s API reference
898     or Models page** — the authoritative vendor doc gives no numeric
899     latency SLA at all, only rate limits (250k tok/s, 1,200 req/min) and
900     price. Treat any specific millisecond figure as informal/measured,
901     not contractual.
902
9032. **Token budget for "how many questions fit in one call"**: internal
904   vendor-doc inconsistency between the Models page's precise 64k/32k
905   split and the Primitives page's rounder "~32,000 tokens total budget" —
906   see §4 above. Not resolvable from the docs as given; use the Models
907   page's more specific numbers for capacity planning and treat 32k as a
908   conservative single-call ceiling if being cautious.
909
9103. **HTTP status code coverage**: the HTTP API reference documents only
911   401/422/429/529; the Python SDK's exception hierarchy additionally
912   covers 400/403/404/generic-5xx. Not a contradiction so much as the HTTP
913   page being an incomplete subset — code defensively for the SDK's fuller
914   list if writing a raw HTTP client.
915
9164. **Score level count**: vendor docs say "at least two levels... up to
917   10" (line 13881); brain page `score.md` independently says "2-10 level
918   scale" — these agree, no real disagreement, included here only because
919   it was cross-checked.
920
9215. **`jevclient` (PyPI) is a different, unofficial, third-party SDK**, not
922   `typesafe-sdk` — do not confuse the two. `agent.py` uses the official
923   `typesafe-sdk`. `jevclient` has its own error classes
924   (`JevAuthError`/`JevValidationError`/`JevRateLimitError`/
925   `JevOverloadedError`/`JevConnectionError`) that do **not** appear
926   anywhere in the vendor docs and should not be assumed to exist on the
927   real API — they are that unofficial package's own wrapper naming.
928
929---
930
931## 9. Quick-reference: minimal working Python harness shape
932
933```python
934import os
935from typesafe_sdk import Choice, Noul, Score, RetryPolicy, TypeSafeClient
936
937# env: TYPESAFE_API_KEY (required), optionally TYPESAFE_BASE_URL,
938# TYPESAFE_DEFAULT_MODEL, TYPESAFE_LOG_LEVEL
939
940client = TypeSafeClient(
941    retry=RetryPolicy(max_retries=4, backoff_initial=1.0, backoff_max=8.0,
942                       backoff_jitter=0.2, timeout=45.0),
943)
944
945response = client.system_one(
946    state={"document": "..."},          # string | dict | list, text only
947    questions={
948        "is_urgent": Noul(instructions="Does this convey urgency?"),
949        "department": Choice(
950            instructions="Which team should handle this?",
951            criteria={"billing": "...", "technical": "...", "sales": "..."},
952        ),
953        "frustration": Score(
954            instructions="How frustrated is the customer?",
955            criteria=["Calm", "Frustrated", "Very angry"],  # 2-10 levels
956        ),
957    },
958    # model="jev-latest",  # default
959)
960
961response.nouls["is_urgent"].noul                # float 0..1
962response.choices["department"].choice           # str
963response.choices["department"].probabilities    # {opt: float}
964response.choices["department"].confidence       # float 0..1
965response.scores["frustration"].score            # float, e.g. 1.6
966response.scores["frustration"].confidence       # float 0..1
967response.usage.input_tokens / .output_tokens    # int
968client.close()
969```
970
971Error handling:
972```python
973from typesafe_sdk import TypeSafeAPIError, TypeSafeRateLimitError
974
975try:
976    client.system_one(state, questions)
977except TypeSafeRateLimitError as e:
978    time.sleep((e.retry_after_ms or 1000) / 1000)  # if not letting RetryPolicy handle it
979except TypeSafeAPIError as e:
980    print(e.status, e.request_id, e.body)
981```
982
983---
984
985## 10. What the vendor does not expose - checked directly for the budget-safety
986work, 2026-09-20
987
988The earlier sections cite the vendor docs closely enough that an absence
989could still be an oversight in the digest rather than in the docs. This
990section is a direct re-check of the primary source itself
991(`~/brains/personal/raw/vendor-docs/typesafe/llms-full-2026-09-19.txt`, all
99220,229 lines) plus one live, billed request made specifically to look at
993headers, run for `packages/jev-http`'s spend ledger and rate-limit handling
994(`packages/jev-http/src/ledger.rs`, `apps/jevprobe --headers`).
995
996### No usage, balance or credit endpoint exists (vendor docs, grepped directly)
997
998Searching the whole vendor doc file for `balance`, `credit`, `/usage`,
999`/v1/usage`, `/v1/billing` and `402` turns up only unrelated worked examples:
1000a banking-agent demo whose fictional tool is named `check_balance`, an
1001invoice-processing example that asks Jev to classify money as a `credit` or a
1002`charge`, and the JS/Python SDK's `Usage` type - which is `{input_tokens,
1003output_tokens}` on one API response (§3 above), not an account-level
1004usage/spend endpoint. There is no vendor equivalent of a `GET
1005/v1/account/usage` or `/v1/billing` route anywhere in the reference. **`GET
1006/v1/models` (§1) is the only account-adjacent `GET` endpoint that exists, and
1007it lists model aliases, not spend.** So `packages/jev-http`'s ledger cannot
1008reconcile against a vendor-reported figure - there is nothing to reconcile
1009against - and the `budget` MCP tool's "vendor-reported balance" field is
1010always `null`, permanently, unless a future version of the API adds one.
1011
1012### No rate-limit headers are documented, and none arrived on a real response
1013
1014Searching the same file for `x-ratelimit`, `ratelimit-`, `rate-limit-` finds
1015no header of that shape anywhere in the HTTP reference, the Python SDK docs,
1016or the JS SDK docs. The **only** rate-limit-adjacent headers named anywhere
1017in the vendor docs are `Retry-After` and `retry-after-ms` (§4), and both are
1018described only in the context of a `429`/`529` response, never as present on
1019an ordinary `2xx`.
1020
1021Confirmed empirically, not just from the docs' silence: `apps/jevprobe
1022--headers` (added for this work) made one real, billed request and printed
1023every response header verbatim:
1024
1025```
1026date                           Mon, 21 Sep 2026 02:37:59 GMT
1027server                         istio-envoy
1028content-length                 231
1029content-type                   application/json
1030x-typesafe-request-id          req_01a0c1d3d6187b088bd31bc35d654f10
1031x-envoy-upstream-service-time  81
1032```
1033
1034Six headers, none of them about rate limits or quota. `server: istio-envoy`
1035is new information not in any digest section above - the API sits behind an
1036Envoy/Istio ingress - and is otherwise unremarkable. `packages/jev-http`
1037still parses the conventional `x-ratelimit-*`/`ratelimit-*` spellings
1038opportunistically into a typed `RateLimit` snapshot on every response
1039(`parse_rate_limit`), so if the vendor ever starts sending one, every caller
1040starts seeing it with no code changed; today that snapshot is `None` on every
1041response, confirmed rather than assumed.
1042
1043**Consequence for the client:** since there is no live signal to react to,
1044`packages/jev-http::Jev` enforces the vendor's own *documented static*
1045ceilings (§4: 250,000 tokens/second, 1,200 requests/minute) with a client-side
1046token-bucket, rather than reading a header. In practice this bucket almost
1047never engages, because the app's own circuit breaker (30 questions/minute
1048default) is far stricter than the vendor's 1,200/minute - the vendor bucket
1049exists for the case where that default is loosened well past its own default,
1050not because it is expected to bind day to day.
1051
1052### No out-of-credit response shape is documented anywhere
1053
1054Searched for `402`, "payment required", and every billing-sounding word
1055above: nothing. The HTTP reference's error table (§4) stops at
1056401/422/429/529; the Python SDK's broader exception hierarchy (§4) adds
1057400/403/404/generic-5xx but never a billing-specific status or exception
1058class. **There is no way to know, from the docs, what an
1059account with no credit left gets back.** `packages/jev::Failure::OutOfCredit`
1060is therefore an inferred, not a documented, classification: a `402` (the
1061conventional "payment required" code, even though the vendor never writes
1062it down) or a `403` whose body mentions credit, balance or quota. Both are
1063guesses about a case this project has never hit (measured spend as of
10642026-09-20, this task included, is a few cents against a five-dollar
1065lifetime credit) and cannot safely go looking for by spending down the
1066account to find out. Should a real out-of-credit response ever arrive with
1067a different shape, `Failure::of`'s catch-all still sorts it as `Invalid`
1068rather than panicking or silently retrying it - it is just not labelled
1069`OutOfCredit` until this classification is corrected against a real example.
1070
1071### Summary table
1072
1073| Vendor exposes | Checked how | Result |
1074| --- | --- | --- |
1075| A usage/balance/credit endpoint | Full-text search of the vendor docs | No - `GET /v1/models` is the only account-adjacent read, and it lists aliases, not spend |
1076| Rate-limit headers on a normal response | Full-text search + one live request (`jevprobe --headers`) | No - six headers arrived, none about limits or quota |
1077| `Retry-After`/`retry-after-ms` on `429`/`529` | Vendor docs (§4) | Yes - the only rate-limit-adjacent headers documented at all |
1078| A documented out-of-credit response shape | Full-text search | No - `402` is never mentioned; `Failure::OutOfCredit` is an inferred guess, not a cited fact |
1079| `x-typesafe-request-id` | Vendor docs (§1) + confirmed live | Yes, on every response checked |