TypeSafe AI "Jev" API — implementation digest
Sources read in full: vendor docs llms-full-2026-09-19.txt (20,229 lines,
authoritative), jev-plays-pokemon/agent.py + README.md, and the brain
pages jev.md, choice.md, score.md, noul.md,
speculative-fan-out.md, jev-fast-llm-deterministic-tools.md,
ecosystem/{typesafe-mario,jevclient,jev-doom-agent}.md. Line numbers below
refer to llms-full-2026-09-19.txt unless stated otherwise.
1. Endpoint, auth, headers, model field
Evaluation endpoint (line 117-123, repeated at 12469-12473):
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer <API_KEY>
Content-Type: application/json
Models list endpoint (line 12888-12896):
GET https://api.typesafe.ai/v1/models
Authorization: Bearer <API_KEY>
- Auth header: exactly
Authorization: Bearer <API_KEY>. No other header is documented as required beyondContent-Type: application/jsonfor POST. - Base URL default (Python SDK constant, line 18134-18140):
DEFAULT_BASE_URL = 'https://api.typesafe.ai'(note: no/v1in the SDK's base URL constant — the SDK presumably appends/v1/systemoneitself; not spelled out further in the docs). - Request-ID response header:
x-typesafe-request-id(lines 14530, 18230, 19314, etc.) — read viaresponse.request_idon the Python client, orerror.request_idon a raised exception. modelfield: string, e.g."jev-latest". Selects which model/alias handles the request (line 133-135, 12837).
2. Request JSON schema
Top-level fields (line 125-158)
{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"is_urgent": {
"type": "noul",
"instructions": "Does this convey urgency?"
}
}
}
state(string | object | array, required): the content to evaluate. See §5 for shaping guidance.model(string, required in examples but presumably has an SDK-side defaultjev-latest— vendor docs don't state the HTTP API has a server-side default ifmodelis omitted; not in vendor docs whether a raw HTTP POST withoutmodelis accepted).questions(map<string, Question>, required): caller-chosen keys → typed question objects. "The key is not sent to the underlying model and is not used in inference" (line 142-143).
Question types — shared fields
All three share type and instructions; each adds its own criteria
(line 162).
Noul (yes/no, line 164-203)
| Field | Required | Shape |
|---|---|---|
type | yes | "noul" |
instructions | yes | string | object | array — the yes/no question |
criteria | no | optional {true, false} — string/object/array descriptions of what a yes and a no mean |
{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"is_urgent": {
"type": "noul",
"instructions": "Does this convey urgency?",
"criteria": {
"true": "Explicitly time-sensitive",
"false": "No urgency expressed"
}
}
}
}
#0 Noul does not have a minimum/maximum on criteria (it's just an optional true/false pair, not a list).
Choice (pick one option, line 205-241, 13469-13513)
| Field | Required | Shape |
|---|---|---|
type | yes | "choice" |
instructions | yes | string | object | array |
criteria | yes | map<string, string | null> — option name → description (null = no extra detail) |
Limit: up to 255 options per Choice question (line 13552: "A Choice question accepts up to 255 options, and adding options costs a few tokens each"). No stated minimum, but two is the practical floor (a single-option Choice is degenerate).
{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoicing, refunds",
"technical": "Bugs, outages, integrations",
"sales": "Pricing, upgrades, new accounts"
}
}
}
}
Score (position on an ordered rubric, line 243-269, 13875-13927)
| Field | Required | Shape |
|---|---|---|
type | yes | "score" |
instructions | yes | string | object | array |
criteria | yes | ordered array of level descriptions, low→high. "Needs at least two levels and takes up to 10." (line 13881) |
A level's number is its zero-based position in the array (line 13891). The model is never shown the level's number or neighbours — only the description text (line 14044), so numeric-only level descriptions ("0", "1", "2") score badly (worked example, line 14046-14052: same report scores 0.0/confidence 1.0 with descriptive levels vs. 0.57/confidence 0.35 with bare numeral levels).
{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]
}
}
}
Structured instructions/criteria (line 13367-13446)
Every one of instructions, Choice's per-option description, Score's
per-level description, and Noul's criteria.true/criteria.false accepts
string | object | array | null (the JS SDK type alias EntryType, line
13376-13383). Field names inside a structured object (e.g. question,
focus, what, not_for, examples) are not part of the API and none
are reserved — "You choose them... use short names that label what
follows" (line 13769). Use this when two options/levels are easily
confused, or a level needs "what it covers" + "example situations" fields
(worked examples at line 13740-13767 and 14199-14263 both show measurable
accuracy/confidence improvement from adding an examples array).
Reference a specific nested state field by dot-and-index path inside
backticks in the instructions text, e.g. `ticket.messages[0].text`
(line 530-536, 13227-13269) — this is a documented convention, not a
separate API field.
3. Response JSON schema
Top-level (line 271-310)
{
"model": "jev-latest",
"answers": {
"is_urgent": { "type": "noul", "noul": 0.92 }
},
"usage": { "input_tokens": 312, "output_tokens": 48 }
}
model(string): the model that actually answered (the resolved versioned ID, not necessarily the alias sent).answers(map<string, Answer>): one answer per question id you chose.usage.input_tokens/usage.output_tokens(integers).
Answer shapes (line 312-414; JSON Schemas at line 19086-19219 from the
Python SDK's pydantic models — these are the authoritative field-level schemas)
Noul answer — {"type": "noul", "noul": <float 0..1>}. No confidence
field. "noul ranges from 0 to 1, representing the probability that the
answer is yes" (line 13819).
Choice answer:
{
"type": "choice",
"choice": "technical",
"probabilities": { "billing": 0.08, "technical": 0.85, "sales": 0.07 },
"confidence": 0.82
}
choice(string): highest-probability option name.probabilities(map<string, number>): every option → probability, "floats that sum to 1" (approximately — the pydantic schema description says "values sum to approximately 1", line 19115).confidence(number 0-1): derived from the shape ofprobabilities.
Score answer:
{
"type": "score",
"score": 1.6,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.05, "1": 0.3, "2": 0.65 },
"confidence": 0.78
}
score(number): probability-weighted mean of level indices — "each level number multiplied by its probability, added up" (line 13963), e.g.0×0.0 + 1×0.70 + 2×0.30 = 1.30. Can land between integer levels.legend(map<string level-index, string|object|array>): echoes each level's description back, keyed by the level's index as a string.probabilities(map<string level-index, number>), string-keyed likelegend.confidence(number 0-1).- Python SDK note (line 13969):
ScoreAnswerre-keysprobabilitiesandlegendby integer level in the typed SDK object, even though the wire JSON keys them as strings.
Full worked request/response pair (Quick Start, line 12495-12564)
Request:
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}
Response:
{
"model": "jev-latest",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"probabilities": { "billing": 0.84, "technical": 0.159, "sales": 0.001 },
"confidence": 0.596
},
"frustration": {
"type": "score",
"score": 1.035,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language"
},
"confidence": 0.842
},
"is_urgent": { "type": "noul", "noul": 0.999 }
},
"usage": { "input_tokens": 312, "output_tokens": 48 }
}
(Note: this particular published example response for frustration omits
probabilities in the docs' own JSON, even though the schema requires it —
likely a docs-authoring trim, not a real absent-field case; every other
Score example in the docs includes probabilities.)
GET /v1/models response (line 12888-12933)
{ "models": [ { "name": "jev-latest", "description": "...", "release_date": "..." } ] }
models[].name(string): the model ID/alias, valid in themodelrequest field.models[].description(string).models[].release_date(string).- "It currently lists the aliases. Versioned IDs such as
jev-1.13.0are accepted by themodelfield whether or not they appear in the list."
4. Errors, rate limits, retry semantics
HTTP API error table (line 416-425 — the only error table in the HTTP
API reference)
| Status | Meaning |
|---|---|
401 Unauthorized | Missing or invalid API key. Check the Authorization header. |
422 Unprocessable Entity | Request body failed validation — missing required field or malformed question. "The body details the offending field." (exact shape of that body is not in vendor docs.) |
429 Too Many Requests | Exceeded rate limit. Back off and retry after a short delay. |
529 Overloaded | TypeSafe is temporarily overloaded. Retry after a short delay. |
"When you receive a 429 or 529... retry the request with exponential
backoff instead of retrying immediately. Our client SDKs handle this
automatically... with its default retry policy." (line 427-429)
Disagreement/gap noted: the HTTP API reference documents only
401/422/429/529, but the Python SDK's exception hierarchy (line
18165-18339) is broader — it defines dedicated exceptions for 400
(TypeSafeBadRequestError), 403 (TypeSafePermissionDeniedError), 404
(TypeSafeNotFoundError), 422 (TypeSafeUnprocessableEntityError), 429
(TypeSafeRateLimitError), and generic 5xx
(TypeSafeInternalServerError — which would presumably catch 529 too,
though 529 is never named in the SDK docs). The HTTP reference page is
the narrower, curated list; the SDK exception surface is the fuller
(implicit) one. Treat both as authoritative for their own scope.
Python SDK exception classes (line 18165-18432)
TypeSafeError(Exception)— base for all SDK failures.TypeSafeAPIError(TypeSafeError)— unsuccessful HTTP response. Fields:status(int),body(parsed JSON error body / plain text /Nonefor empty),headers,endpoint(method+URL, no creds/query/fragment),request_id(property, fromx-typesafe-request-idorNone).TypeSafeBadRequestError(400)TypeSafeAuthenticationError(401)TypeSafePermissionDeniedError(403)TypeSafeNotFoundError(404)TypeSafeUnprocessableEntityError(422)TypeSafeRateLimitError(429) — addsretry_after_ms(parse_retry_after(headers),Noneif unavailable)TypeSafeInternalServerError(5xx)
TypeSafeAPIConnectionError(TypeSafeError, ConnectionError)— no HTTP response at all (network failure).TypeSafeAPITimeoutError(TypeSafeAPIConnectionError, TimeoutError)— addstimeout(seconds orhttpx2.Timeout).
TypeSafeAPIResponseValidationError(TypeSafeAPIError)— 2xx response whose body was structurally invalid; addsfield_path(dotted path to the offending field, e.g.answers.tone.confidence).
Rate-limit headers honored: Retry-After and retry-after-ms
(line 16758, 18398) — "Whether to honor Retry-After and retry-after-ms
response headers", up to maxRetryAfterMs in the JS SDK's RetryPolicy.
RetryPolicy (Python SDK, line 18346-18430)
from typesafe_sdk import RetryPolicy, TypeSafeClient
client = TypeSafeClient(
retry=RetryPolicy(max_retries=3, timeout=10.0, http_statuses={429, 500, 502, 503, 504})
)
Fields: max_retries (0 disables), backoff_initial (seconds, doubled
each attempt up to backoff_max; 0 disables backoff), backoff_max,
backoff_jitter (fraction randomly subtracted, 0-1), http_statuses
(retried set), respect_retry_after (bool), api_connection_error (bool,
retry on connection failure), api_timeout_error (bool), exceptions
(extra exception types to retry), predicate (callable on the raised
exception → bool), timeout (total retry budget in seconds across the
whole call including delays — "Stops before a retry whose delay would
reach or exceed the budget, re-raising the last error").
Rate limits and context/pricing (Models page, line 12839-12934 — the
authoritative numbers table)
| Jev 1.13 | jev-1.13.0 |
|---|---|
| Price (per Btok / per Mtok) | $42 / $0.042 — input tokens only; output tokens are free |
| Rate limits | 250,000 tokens/second / 1,200 requests/minute |
| Context length | 64k tokens per request (state + all questions combined); 32k tokens for state + the single longest question |
| Input | Text only: string, JSON object, array of text values. No image/audio/video. |
"Rate limits are adjusting dynamically... can change without notice while we [serve growing demand]... Higher limits are available on custom and enterprise plans." (line 12853-12855) — i.e. these are not guaranteed stable numbers.
Internal inconsistency in the vendor docs, flagged not resolved: the Primitives page (line 13328) separately states "The number of questions in one request is limited only by the request's token budget, which the state and the questions share. The budget is around 32,000 tokens, roughly 150,000 characters of English text." This describes the whole request's budget as ~32k, which conflicts with the Models page's more precise split (64k total request / 32k for state + longest single question). Vendor docs win per the task's instruction, but the two vendor passages themselves disagree on which number (32k or 64k) is the ceiling for a many-question request — treat 64k (Models page, more specific and more recently authored-sounding) as the more load-bearing number for total request budget, and 32k as a per-state-plus-single-question sub-limit.
Aliases (line 12857-12870)
| Alias | Points to | Meaning |
|---|---|---|
jev-latest | jev-1.13.0 | Most recent stable/official release. SDK default. |
jev-preview | jev-1.13.0 | Most recent release whether or not official; currently identical to jev-latest (no preview build live). |
"If you have tuned confidence thresholds against a specific version, pin that version's ID instead of the alias."
5. State shaping, calibration, sampling — what the docs say to do
State (line 843-902)
stateisstring | JSON object | array; text only, no images/audio/ video (repeated at line 871, 916-918, 12846, 12851).- Prefer an object for most requests "so each part of the state has a descriptive name and its relationships remain clear." A plain string is fine only for one simple piece of text (line 868).
- Every question in a request sees the same state and is evaluated independently — "one primitive's result does not become hidden context that changes another primitive's result" (line 481, 850).
- Keep only the context relevant to the current questions — "This helps the model avoid distractions and context rot" (line 522).
- Point questions at specific nested values with a backticked dot/index
path, e.g.
`support.tickets[0].message`(line 530-536). - English is the primary training language; other languages including CJK are accepted but currently lower accuracy (line 871, 12882).
Model jaggedness — jev-1.13 failure modes and mitigations (line
12677-12830, table at 12690-12700)
- Literal reading — answers the words written, not the intent. Put
boundary cases explicitly in
criteria; split ambiguous judgments into two literal questions and combine in code. - Math and numbers — "Jev is not a calculator." Don't ask it to count
(characters, occurrences, list items) — count in code (worked example:
ask one
Noulper item, sum> thresholdresults in code, line 12720-12737). Numeric representations (hex/RGB, low-level code) score worse than named/semantic equivalents — convert in code, pass semantic labels. Score outputs must not be used to reconstruct an exact number by interpolating between levels — only threshold the expectation. - Date/time comparison — dates are read as text, not ordered
quantities; ordering/duration/window checks are unreliable, worse with
mixed formats. Extract components as bounded
Choices (month/day/year as closed sets, with an explicit "not stated" option), assemble and do all arithmetic in code. - Indirection — double negatives / multi-hop "property of a property" questions lose accuracy. Write instructions directly; name relevant state parts explicitly.
- Large state full of irrelevant detail — accuracy falls as
unrelated content grows in
state("context rot" reiterated). Filter/ retrieve in code first; when that's not possible, use aNoulas a relevance filter first. - Adversarial content —
stateis treated as data, not hostile, by default; injected instructions / misleading framing can move the answer. Be explicit in criteria; test adversarial inputs before deploying. - Contradictory instructions/criteria — e.g. a Noul where
truemaps to "no" semantically will perform worse. Align criteria as an extension of instructions, plain everyday phrasing. - Common-sense structural invariants do NOT hold — worked numeric
example: the same yes/no judgment asked as
Noulvs.Choicegavenoul=0.22vs.probabilities["yes"]=0.01— not comparable. Two Nouls for "X" and "not X" summed to 1.19, not 1.0. Don't carry a threshold tuned on one question type over to another; don't expect arithmetic identities between separate questions. - Generation — Jev is not trained to generate text; forcing free text
via chained choices "will not work well and will be very slow." Turn
bounded extraction into a
Choice; use a real generative model for actual text generation.
Summary reminders (line 12818-12825): avoid asking Jev something code can
compute exactly, hiding multiple judgments in one question, multi-hop
"System Two" tasks, or feeding more state context than a question needs.
Confidence / calibration (line 1149-1230)
confidence(0-1) is present on Choice and Score answers only, not Noul. It is a single-number summary of how peaked/flat theprobabilitiesdistribution is (line 1156, 1160).- "We provide
confidenceas a convenient measure that fits most use-cases, but you are never locked into our definition... a different measure may serve you better... which is exactly why we give you the fullprobabilities" (line 1163) — i.e. compute your own statistic fromprobabilitiesifconfidencedoesn't fit. - Calibration is trained via RLCD ("Reinforcement Learning for Calibrated Decisions" per the brain notes; vendor docs link it as "RLCD" without spelling out the acronym inline at line 493). "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." (line 922)
- Recommended pattern: three confidence bands — high (act automatically), medium (confirm/flag/gather more info), low (route to a human/fallback) — with threshold values that scale with the stakes of the specific action, not one global number (line 1174-1229, worked banking-command example).
- No documented
temperatureor sampling parameter exists on the API itself — Jev always returns a full deterministic-per-call probability distribution; there's no server-side "sampling temperature" request field anywhere in the vendor docs. (Not in vendor docs: any request field to control determinism/temperature/top-k on the Jev call itself.) Client-side sampling from the returned distribution — asagent.pydoes — is a caller-side technique, not a documented API feature.
6. From agent.py (jev-plays-pokemon) — concrete harness patterns
Full source read; path:
/home/nixos/brains/personal/raw/code/jev-plays-pokemon/agent.py
(+ README.md beside it).
Imports / SDK usage confirmed
from typesafe_sdk import Choice, Noul, RetryPolicy, TypeSafeClient
This matches the official typesafe-sdk Python package from the
vendor docs exactly (not jevclient, which is an unrelated third-party
package — see §8).
Environment variables it reads
TYPESAFE_API_KEY— checked directly withos.getenv, raisesRuntimeError("TYPESAFE_API_KEY is not set. Create a .env file with your key.")if absent, before ever constructingTypeSafeClient(the SDK itself would also raiseTypeSafeErrorfor a missing key, but the agent pre-checks with its own message).TYPESAFE_MAX_RETRIES(default4)TYPESAFE_BACKOFF_INITIAL(default1.0seconds)TYPESAFE_BACKOFF_MAX(default8.0seconds)TYPESAFE_TIMEOUT(default45.0seconds) — this is the retry budget (RetryPolicy.timeout), not the SDK's default 10s per-HTTP-operation timeout.TYPESAFE_TEMPERATURE(default2.0) — agent-defined, not an SDK/API concept; used only in_sample_goal(client-side).TYPESAFE_GOAL_FLOOR(default0.1) — same, agent-defined.- Loaded via
python-dotenv'sload_dotenv()at import time, from a project-root.envfile (README:cp .env.example .env).
Retry policy construction (verbatim)
def retry_policy() -> RetryPolicy:
return RetryPolicy(
max_retries=_int("TYPESAFE_MAX_RETRIES", 4),
backoff_initial=_float("TYPESAFE_BACKOFF_INITIAL", 1.0),
backoff_max=_float("TYPESAFE_BACKOFF_MAX", 8.0),
backoff_jitter=0.2,
timeout=_float("TYPESAFE_TIMEOUT", 45.0),
)
backoff_jitter=0.2 is hardcoded (not env-overridable). RetryPolicy is
passed once at client construction (TypeSafeClient(retry=retry_policy(), ...)),
not per-call.
How it structures state and questions for a game loop
Two call shapes, both against the same state: dict[str, Any] built
elsewhere (state.py, not requested for this digest) and the same
shared set of per-button Noul questions (self.action_questions,
built once in __init__ via build_questions()):
def build_questions() -> dict[str, Any]:
return {
action: Noul(instructions=f"Is the best single action right now to {ACTION_NAMES[action]}?")
for action in ACTION_FRAMES
} | {
"menu_open": Noul(
instructions=(
"Is a selectable menu open on screen (start menu, POKEMON/BAG/ITEM list, "
"battle FIGHT/PKMN/BAG/RUN, or a yes/no question)? A plain text dialog is NOT a menu."
)
),
}
— 8 action Nouls (press_a, press_b, press_start, walk_up/down/left/right,
wait) + 1 menu_open Noul, all fired in one call, every turn. Comment
at file top: "Uses one atomic Noul per candidate action (asked in parallel
in a single call)... Code then picks the strongest answer with a margin
requirement and sanity-checks it against the deterministic walkability
map." This is the vendor docs' "speculative fan-out" pattern applied
directly.
decide_goal() adds one Choice (next_goal, options = the
currently-available high-level goals, e.g. talk_to_Mom, explore,
reach_exit) to the same 9 action Nouls, all in a single
client.system_one(state=state, questions=questions) call:
questions = {
"next_goal": Choice(
instructions={
"question": "Which goal should RED pursue next?",
"focus": "Pick the ONE goal that best advances your long-term goal (badges/Champion) "
"given the screen text, recent hints, nearby objects, walkability and room map. "
"Code will execute the walking.",
},
criteria=criteria, # {goal_id: goal_desc, ...}
)
}
questions.update(self.action_questions)
response = self._interruptible(self.client.system_one, state=state, questions=questions)
Note the structured instructions object ({"question": ..., "focus": ...})
for the Choice — exactly the vendor docs' "structured instructions" pattern
(§2 above), with caller-chosen, non-reserved field names.
decide_action() (menu/battle micro-decisions) sends only the 9 action
Nouls, no Choice.
How it samples from the probability distribution (_sample_goal)
The vendor docs never describe server-side sampling — this is 100%
client-side post-processing of the returned choice.probabilities:
def _sample_goal(
probabilities: dict[str, float],
temperature: float = 2.0,
floor: float = 0.05,
) -> tuple[str, float]:
labels = list(probabilities)
...
probs = [max(0.0, probabilities[g]) for g in labels]
total = sum(probs)
probs = [p / total for p in probs]
# Floor: guarantee every option keeps some mass (p=1 is never 100%).
floor = max(0.0, floor)
probs = [p + floor for p in probs]
total = sum(probs)
probs = [p / total for p in probs]
# Softmax flattening: logits = log(p)/T. T>1 pulls the peak down and
# lifts the tail, keeping a real chance of doing something else.
if temperature != 1.0 and temperature > 0:
logits = [math.log(p + 1e-12) / temperature for p in probs]
m = max(logits)
exp = [math.exp(x - m) for x in logits]
exp_total = sum(exp)
probs = [e / exp_total for e in exp]
label = random.choices(labels, weights=probs, k=1)[0]
return label, probs[labels.index(label)]
Two composed transforms before random.choices: (1) an additive floor
so no option ever reaches literal 0% or literal 100% mass; (2) softmax
flattening of log(p)/T with T=2.0 default so a confidently-wrong
p=0.99 pick fires only "~80%" per the docstring, ~60-70% per the README —
i.e. deliberately de-sharpening Jev's own calibrated distribution to avoid
looping on a single wrong-but-confident goal every turn. The Noul actions
(decide_action) are not sampled this way — they use argmax with a
margin gate instead (below), because action selection wants determinism +
a safety fallback, not exploration.
Anti-stuck / safety tricks (pick_action, deterministic, no model call)
- Margin + confidence gate: only trust the model's top Noul pick if
decision.margin >= 0.08 and decision.confidence >= 0.45(margin= winningnoulminus runner-upnoul). Below that, fall back to advancing dialog / open menu (press_a) / walking toward anexplore_hint/wait, in that priority order. - Menu override: if
menu_openNoul ≥ 0.5 and the winning action's confidence < 0.7, forcepress_a(confirm highlighted option) — "reliably selects the starter (and most story choices)." - Deterministic wall check: for any
walk_*action, cross-check against a ground-truthwalkable: dict[str, bool]map read from emulator RAM (not from the model). If the chosen direction is blocked and not every direction is blocked and not explicitly probing for an exit, override topress_a(if a dialog is active) orwait— i.e. the model's spatial judgment is never trusted over the deterministic collision map except when probing for exits. - Anti-pacing: never immediately reverse the previous walk
(
_REVERSEmap) if another open direction exists — picks the next-best-scoring open direction instead, gated atnoul >= 0.3. - Higher-level (README, not in
agent.pyitself): stuck detection (press B → wait → re-route) and exit-probing when a room is fully explored live inplay.py/navigation.py, not shown in this file.
Threading / interruptibility (_interruptible)
Every blocking client.system_one call is run on a daemon worker
thread while the main thread polls a stop_requested() flag every 50ms
and raises KeyboardInterrupt itself — because "the HTTP stack swallows
SIGINT while it is running." This is a harness-level workaround, not
documented anywhere in the vendor SDK docs (not in vendor docs: any
statement that the SDK/httpx2 swallows SIGINT — this is the agent
author's own empirical finding, not a cited vendor fact).
Measured cost/latency (README, not vendor docs — first-party but
project-specific measurement)
- "Jev is fast (tens–hundreds of ms per decision)."
--turbomode note: "the Jev API call per decision (~0.5 s) becomes the pacing factor" — i.e. ~500ms observed round-trip in this specific harness/network conditions, at the high end of "tens-hundreds of ms."- Token estimate per decision: shared instructions ~850 tok + questions
(8 Nouls +
menu_open) ~350 tok + per-turn state ~400-900 tok ≈ 1,600-2,100 input tokens/decision. - Cost at $0.042/MTok: ≈$0.27/hour at 1 decision/sec, ≈$0.76 per 10,000 decisions.
7. Official Python SDK vs. raw HTTP
Yes, an official SDK exists: package typesafe-sdk on PyPI (line
12573, 17492-17500), requires Python ≥ 3.10 (line 12570) — note agent.py's
own project requires Python ≥ 3.14 per its README, which is the harness's
own constraint, not the SDK's floor.
pip install typesafe-sdk
# or
uv add typesafe-sdk
Minimal usage (sync, line 17539-17561 / 12583-12617):
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
with TypeSafeClient() as client: # reads TYPESAFE_API_KEY from env
response = client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(
instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None},
),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["can wait", "this week", "today"],
),
},
)
print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)
Async variant: AsyncTypeSafeClient / await client.system_one(...),
identical parameter shape, requires async with and await.
TypeSafeClient / AsyncTypeSafeClient constructor params (line
17600-17635): api_key (→ TYPESAFE_API_KEY), model (→
TYPESAFE_DEFAULT_MODEL), retry (RetryPolicy), timeout (float
seconds or httpx2.Timeout; SDK default DEFAULT_TIMEOUT = 10.0),
headers (extra), transport / http_client (mutually exclusive custom
httpx2 transport), base_url (→ TYPESAFE_BASE_URL). Raises
TypeSafeError if API key missing or timeout invalid; ValueError if both
transport and http_client given.
client.system_one(...) params (line 17685-17735): state
(JSONContent), questions (Mapping[str, Question], non-empty — raises
TypeSafeError if empty, or if a Score question's criteria list is
empty), model (override), retry (per-call override), timeout
(per-call override), extra_headers, extra_body (forward-compat, shallow
last-write-wins merge over the body), response_model (optional Pydantic
BaseModel subclass to type the response — see forward-compat section
below).
client.models.list(...) → ListModelsResponse with .models[] of
ModelMetadata{name, description, release_date} (line 19770-19925).
Environment variables the SDK itself reads (line 18094-18172, restated at 20161-20172 — this table is the authoritative env-var reference):
| Variable | Configures | Default |
|---|---|---|
TYPESAFE_API_KEY | API key (required) | — |
TYPESAFE_BASE_URL | API root URL | https://api.typesafe.ai |
TYPESAFE_DEFAULT_MODEL | Default model | jev-latest |
TYPESAFE_LOG_LEVEL | typesafe_sdk logger level, applied once at import | unset (off) |
"Explicit options take precedence over environment variables; empty or whitespace-only environment values are ignored." (line 17592)
Logging: SDK logs to Python logger name typesafe_sdk. info = one
summary line/request; debug = also headers+bodies. Secret headers
(authorization, API keys, cookies, any header with token/secret in the
name) are redacted; request/response bodies are NOT redacted (line
20159).
Forward-compatibility escape hatches (line 20174-20228):
extra_body={"beam_width": 4}— send request fields the SDK predates.- Raw dict questions instead of
Noul/Choice/Scoreobjects, including extra unknown keys (e.g.{"type": "noul", "instructions": "...", "weight": 2}) — "Ignore their type-checking errors and prefer upgrading the SDK instead." - Unknown answer kinds: SDK logs a warning and skips them; use
result.raw_http_response.json()["answers"]to see everything including unrecognized kinds. - Unknown extra fields on recognized responses are silently ignored.
JavaScript/TypeScript SDK also exists (@typesafe-ai/sdk,
TypeSafeClient.systemOne(...), line 14286-17471) — not needed for a
Python harness, only noted for completeness; its error class names
(APIConnectionError, RateLimitError with retryAfterMs, etc.) mirror
the Python SDK's one-for-one.
Raw HTTP is fully viable — the cURL example (line 12477-12493) and the
plain-dict "raw question dictionaries" pattern above show the wire format
is a stable, documented contract independent of any SDK; a hand-rolled
requests/httpx client needs only: POST to
https://api.typesafe.ai/v1/systemone, Authorization: Bearer header,
JSON body per §2, and to implement retry/backoff itself for 429/529 (with
Retry-After/retry-after-ms header awareness) since raw HTTP gets none
of the SDK's automatic retry behavior.
7b. Measured against the live API, 2026-09-20 (apps/jevprobe, apps/player)
Everything in sections 1-7 was read, not run. This section is the opposite:
what the API actually did on this machine, from a hand-rolled raw-HTTP client
(packages/jev-http) rather than an SDK.
Nothing in the digest needed correcting. The request shape in §2 was
accepted as documented, and the response matched §3 field for field - no
unknown fields, probabilities over every option, confidence on the Choice,
usage with both token counts. Raw response to a one-Choice request:
{"model":"jev-1.13.0","answers":{"move":{"type":"choice","choice":"walk_south",
"confidence":0.98,"probabilities":{"press_a":0.01,"walk_south":0.99,
"walk_north":0.0,"walk_west":0.0}}},"usage":{"input_tokens":544,"output_tokens":54}}
modelin the response is the resolved versioned id,jev-1.13.0, for a request that sent the aliasjev-latest- exactly as §3 says, now seen.- Latency. §8 collects five conflicting figures and notes none is in the API reference. Measured here, from a NixOS host over residential fibre, on a request of ~700 input tokens carrying one Choice, one Score and one Noul: cold connection ~400 ms; warm 110-280 ms, typically ~160 ms, over 755 consecutive requests. Keeping one client alive is therefore worth more than twice the cost of the answer itself, which is the argument for a client with a connection pool rather than one TLS handshake per decision.
- Cost. 755 three-question decisions came to $0.0277 - about
$0.000037 each, at 550-900 input tokens per request. At one decision a
second that is $0.13 an hour. The budget cap in
packages/brainis a stop, not a budget. - Every request succeeded on the first attempt. No 429, no 529, no connection failure in 755 requests plus a few hundred more across earlier runs, so the retry policy in §4/§6 is implemented and has never yet fired. That is the shape of the absence, not evidence it is unnecessary.
- Browser origins are still refused (recorded in the brain page, retested by construction here): the native client uses the API directly, and a web build will need a relay.
8. Cross-source disagreements (vendor docs win; noted for awareness)
-
Latency figures vary by source and none is in the API reference itself:
- Vendor docs prose (not the API reference table): "Most queries complete in about 100 ms" (line 489) and "Frontier intelligence at real-time speeds (150ms)" (line 974) — both marketing/concepts pages, not a guaranteed SLA number.
jev-plays-pokemonREADME (first-party measurement, not vendor): "tens–hundreds of ms per decision," ~0.5s observed in turbo mode.- Brain page
jev.md(secondary, citing TechCrunch/flaviocopes, not vendor docs): "70ms-500ms" end-to-end. - Brain page
ecosystem/jev-doom-agent.md(secondary, citing TypeSafe's own launch blog — a different document from the docs site that was the assigned source): "roughly 10 decisions per second" (100ms) at "$7/hour." - Brain page
ecosystem/jevclient.md(third-party unofficial SDK's own measured table, Netherlands, cold-connection): 690-1,336ms for 3-400 questions in one call; ~250-580ms warm-connection. - None of these is inside
llms-full-2026-09-19.txt's API reference or Models page — the authoritative vendor doc gives no numeric latency SLA at all, only rate limits (250k tok/s, 1,200 req/min) and price. Treat any specific millisecond figure as informal/measured, not contractual.
-
Token budget for "how many questions fit in one call": internal vendor-doc inconsistency between the Models page's precise 64k/32k split and the Primitives page's rounder "~32,000 tokens total budget" — see §4 above. Not resolvable from the docs as given; use the Models page's more specific numbers for capacity planning and treat 32k as a conservative single-call ceiling if being cautious.
-
HTTP status code coverage: the HTTP API reference documents only 401/422/429/529; the Python SDK's exception hierarchy additionally covers 400/403/404/generic-5xx. Not a contradiction so much as the HTTP page being an incomplete subset — code defensively for the SDK's fuller list if writing a raw HTTP client.
-
Score level count: vendor docs say "at least two levels... up to 10" (line 13881); brain page
score.mdindependently says "2-10 level scale" — these agree, no real disagreement, included here only because it was cross-checked. -
jevclient(PyPI) is a different, unofficial, third-party SDK, nottypesafe-sdk— do not confuse the two.agent.pyuses the officialtypesafe-sdk.jevclienthas its own error classes (JevAuthError/JevValidationError/JevRateLimitError/JevOverloadedError/JevConnectionError) that do not appear anywhere in the vendor docs and should not be assumed to exist on the real API — they are that unofficial package's own wrapper naming.
9. Quick-reference: minimal working Python harness shape
import os
from typesafe_sdk import Choice, Noul, Score, RetryPolicy, TypeSafeClient
# env: TYPESAFE_API_KEY (required), optionally TYPESAFE_BASE_URL,
# TYPESAFE_DEFAULT_MODEL, TYPESAFE_LOG_LEVEL
client = TypeSafeClient(
retry=RetryPolicy(max_retries=4, backoff_initial=1.0, backoff_max=8.0,
backoff_jitter=0.2, timeout=45.0),
)
response = client.system_one(
state={"document": "..."}, # string | dict | list, text only
questions={
"is_urgent": Noul(instructions="Does this convey urgency?"),
"department": Choice(
instructions="Which team should handle this?",
criteria={"billing": "...", "technical": "...", "sales": "..."},
),
"frustration": Score(
instructions="How frustrated is the customer?",
criteria=["Calm", "Frustrated", "Very angry"], # 2-10 levels
),
},
# model="jev-latest", # default
)
response.nouls["is_urgent"].noul # float 0..1
response.choices["department"].choice # str
response.choices["department"].probabilities # {opt: float}
response.choices["department"].confidence # float 0..1
response.scores["frustration"].score # float, e.g. 1.6
response.scores["frustration"].confidence # float 0..1
response.usage.input_tokens / .output_tokens # int
client.close()
Error handling:
from typesafe_sdk import TypeSafeAPIError, TypeSafeRateLimitError
try:
client.system_one(state, questions)
except TypeSafeRateLimitError as e:
time.sleep((e.retry_after_ms or 1000) / 1000) # if not letting RetryPolicy handle it
except TypeSafeAPIError as e:
print(e.status, e.request_id, e.body)
10. What the vendor does not expose - checked directly for the budget-safety
work, 2026-09-20
The earlier sections cite the vendor docs closely enough that an absence
could still be an oversight in the digest rather than in the docs. This
section is a direct re-check of the primary source itself
(~/brains/personal/raw/vendor-docs/typesafe/llms-full-2026-09-19.txt, all
20,229 lines) plus one live, billed request made specifically to look at
headers, run for packages/jev-http's spend ledger and rate-limit handling
(packages/jev-http/src/ledger.rs, apps/jevprobe --headers).
No usage, balance or credit endpoint exists (vendor docs, grepped directly)
Searching the whole vendor doc file for balance, credit, /usage,
/v1/usage, /v1/billing and 402 turns up only unrelated worked examples:
a banking-agent demo whose fictional tool is named check_balance, an
invoice-processing example that asks Jev to classify money as a credit or a
charge, and the JS/Python SDK's Usage type - which is {input_tokens, output_tokens} on one API response (§3 above), not an account-level
usage/spend endpoint. There is no vendor equivalent of a GET /v1/account/usage or /v1/billing route anywhere in the reference. GET /v1/models (§1) is the only account-adjacent GET endpoint that exists, and
it lists model aliases, not spend. So packages/jev-http's ledger cannot
reconcile against a vendor-reported figure - there is nothing to reconcile
against - and the budget MCP tool's "vendor-reported balance" field is
always null, permanently, unless a future version of the API adds one.
No rate-limit headers are documented, and none arrived on a real response
Searching the same file for x-ratelimit, ratelimit-, rate-limit- finds
no header of that shape anywhere in the HTTP reference, the Python SDK docs,
or the JS SDK docs. The only rate-limit-adjacent headers named anywhere
in the vendor docs are Retry-After and retry-after-ms (§4), and both are
described only in the context of a 429/529 response, never as present on
an ordinary 2xx.
Confirmed empirically, not just from the docs' silence: apps/jevprobe --headers (added for this work) made one real, billed request and printed
every response header verbatim:
date Mon, 21 Sep 2026 02:37:59 GMT
server istio-envoy
content-length 231
content-type application/json
x-typesafe-request-id req_01a0c1d3d6187b088bd31bc35d654f10
x-envoy-upstream-service-time 81
Six headers, none of them about rate limits or quota. server: istio-envoy
is new information not in any digest section above - the API sits behind an
Envoy/Istio ingress - and is otherwise unremarkable. packages/jev-http
still parses the conventional x-ratelimit-*/ratelimit-* spellings
opportunistically into a typed RateLimit snapshot on every response
(parse_rate_limit), so if the vendor ever starts sending one, every caller
starts seeing it with no code changed; today that snapshot is None on every
response, confirmed rather than assumed.
Consequence for the client: since there is no live signal to react to,
packages/jev-http::Jev enforces the vendor's own documented static
ceilings (§4: 250,000 tokens/second, 1,200 requests/minute) with a client-side
token-bucket, rather than reading a header. In practice this bucket almost
never engages, because the app's own circuit breaker (30 questions/minute
default) is far stricter than the vendor's 1,200/minute - the vendor bucket
exists for the case where that default is loosened well past its own default,
not because it is expected to bind day to day.
No out-of-credit response shape is documented anywhere
Searched for 402, "payment required", and every billing-sounding word
above: nothing. The HTTP reference's error table (§4) stops at
401/422/429/529; the Python SDK's broader exception hierarchy (§4) adds
400/403/404/generic-5xx but never a billing-specific status or exception
class. There is no way to know, from the docs, what an
account with no credit left gets back. packages/jev::Failure::OutOfCredit
is therefore an inferred, not a documented, classification: a 402 (the
conventional "payment required" code, even though the vendor never writes
it down) or a 403 whose body mentions credit, balance or quota. Both are
guesses about a case this project has never hit (measured spend as of
2026-09-20, this task included, is a few cents against a five-dollar
lifetime credit) and cannot safely go looking for by spending down the
account to find out. Should a real out-of-credit response ever arrive with
a different shape, Failure::of's catch-all still sorts it as Invalid
rather than panicking or silently retrying it - it is just not labelled
OutOfCredit until this classification is corrected against a real example.
Summary table
| Vendor exposes | Checked how | Result |
|---|---|---|
| A usage/balance/credit endpoint | Full-text search of the vendor docs | No - GET /v1/models is the only account-adjacent read, and it lists aliases, not spend |
| Rate-limit headers on a normal response | Full-text search + one live request (jevprobe --headers) | No - six headers arrived, none about limits or quota |
Retry-After/retry-after-ms on 429/529 | Vendor docs (§4) | Yes - the only rate-limit-adjacent headers documented at all |
| A documented out-of-credit response shape | Full-text search | No - 402 is never mentioned; Failure::OutOfCredit is an inferred guess, not a cited fact |
x-typesafe-request-id | Vendor docs (§1) + confirmed live | Yes, on every response checked |