jevsnes.git / research / jev-api-digest.md

TypeSafe AI "Jev" API — implementation digest

Sources read in full: vendor docs llms-full-2026-09-19.txt (20,229 lines, authoritative), jev-plays-pokemon/agent.py + README.md, and the brain pages jev.md, choice.md, score.md, noul.md, speculative-fan-out.md, jev-fast-llm-deterministic-tools.md, ecosystem/{typesafe-mario,jevclient,jev-doom-agent}.md. Line numbers below refer to llms-full-2026-09-19.txt unless stated otherwise.


1. Endpoint, auth, headers, model field

Evaluation endpoint (line 117-123, repeated at 12469-12473):

POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer <API_KEY>
Content-Type: application/json

Models list endpoint (line 12888-12896):

GET https://api.typesafe.ai/v1/models
Authorization: Bearer <API_KEY>
  • Auth header: exactly Authorization: Bearer <API_KEY>. No other header is documented as required beyond Content-Type: application/json for POST.
  • Base URL default (Python SDK constant, line 18134-18140): DEFAULT_BASE_URL = 'https://api.typesafe.ai' (note: no /v1 in the SDK's base URL constant — the SDK presumably appends /v1/systemone itself; not spelled out further in the docs).
  • Request-ID response header: x-typesafe-request-id (lines 14530, 18230, 19314, etc.) — read via response.request_id on the Python client, or error.request_id on a raised exception.
  • model field: string, e.g. "jev-latest". Selects which model/alias handles the request (line 133-135, 12837).

2. Request JSON schema

Top-level fields (line 125-158)

{
  "state": "Help! My payouts have been failing for 3 days.",
  "model": "jev-latest",
  "questions": {
    "is_urgent": {
      "type": "noul",
      "instructions": "Does this convey urgency?"
    }
  }
}
  • state (string | object | array, required): the content to evaluate. See §5 for shaping guidance.
  • model (string, required in examples but presumably has an SDK-side default jev-latest — vendor docs don't state the HTTP API has a server-side default if model is omitted; not in vendor docs whether a raw HTTP POST without model is accepted).
  • questions (map<string, Question>, required): caller-chosen keys → typed question objects. "The key is not sent to the underlying model and is not used in inference" (line 142-143).

Question types — shared fields

All three share type and instructions; each adds its own criteria (line 162).

Noul (yes/no, line 164-203)

FieldRequiredShape
typeyes"noul"
instructionsyesstring | object | array — the yes/no question
criterianooptional {true, false} — string/object/array descriptions of what a yes and a no mean
{
  "state": "Help! My payouts have been failing for 3 days.",
  "model": "jev-latest",
  "questions": {
    "is_urgent": {
      "type": "noul",
      "instructions": "Does this convey urgency?",
      "criteria": {
        "true": "Explicitly time-sensitive",
        "false": "No urgency expressed"
      }
    }
  }
}

#0 Noul does not have a minimum/maximum on criteria (it's just an optional true/false pair, not a list).

Choice (pick one option, line 205-241, 13469-13513)

FieldRequiredShape
typeyes"choice"
instructionsyesstring | object | array
criteriayesmap<string, string | null> — option name → description (null = no extra detail)

Limit: up to 255 options per Choice question (line 13552: "A Choice question accepts up to 255 options, and adding options costs a few tokens each"). No stated minimum, but two is the practical floor (a single-option Choice is degenerate).

{
  "state": "Help! My payouts have been failing for 3 days.",
  "model": "jev-latest",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "billing": "Payments, invoicing, refunds",
        "technical": "Bugs, outages, integrations",
        "sales": "Pricing, upgrades, new accounts"
      }
    }
  }
}

Score (position on an ordered rubric, line 243-269, 13875-13927)

FieldRequiredShape
typeyes"score"
instructionsyesstring | object | array
criteriayesordered array of level descriptions, low→high. "Needs at least two levels and takes up to 10." (line 13881)

A level's number is its zero-based position in the array (line 13891). The model is never shown the level's number or neighbours — only the description text (line 14044), so numeric-only level descriptions ("0", "1", "2") score badly (worked example, line 14046-14052: same report scores 0.0/confidence 1.0 with descriptive levels vs. 0.57/confidence 0.35 with bare numeral levels).

{
  "state": "Help! My payouts have been failing for 3 days.",
  "model": "jev-latest",
  "questions": {
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": ["Calm", "Frustrated", "Very angry"]
    }
  }
}

Structured instructions/criteria (line 13367-13446)

Every one of instructions, Choice's per-option description, Score's per-level description, and Noul's criteria.true/criteria.false accepts string | object | array | null (the JS SDK type alias EntryType, line 13376-13383). Field names inside a structured object (e.g. question, focus, what, not_for, examples) are not part of the API and none are reserved — "You choose them... use short names that label what follows" (line 13769). Use this when two options/levels are easily confused, or a level needs "what it covers" + "example situations" fields (worked examples at line 13740-13767 and 14199-14263 both show measurable accuracy/confidence improvement from adding an examples array).

Reference a specific nested state field by dot-and-index path inside backticks in the instructions text, e.g. `ticket.messages[0].text` (line 530-536, 13227-13269) — this is a documented convention, not a separate API field.


3. Response JSON schema

Top-level (line 271-310)

{
  "model": "jev-latest",
  "answers": {
    "is_urgent": { "type": "noul", "noul": 0.92 }
  },
  "usage": { "input_tokens": 312, "output_tokens": 48 }
}
  • model (string): the model that actually answered (the resolved versioned ID, not necessarily the alias sent).
  • answers (map<string, Answer>): one answer per question id you chose.
  • usage.input_tokens / usage.output_tokens (integers).

Answer shapes (line 312-414; JSON Schemas at line 19086-19219 from the

Python SDK's pydantic models — these are the authoritative field-level schemas)

Noul answer — {"type": "noul", "noul": <float 0..1>}. No confidence field. "noul ranges from 0 to 1, representing the probability that the answer is yes" (line 13819).

Choice answer:

{
  "type": "choice",
  "choice": "technical",
  "probabilities": { "billing": 0.08, "technical": 0.85, "sales": 0.07 },
  "confidence": 0.82
}
  • choice (string): highest-probability option name.
  • probabilities (map<string, number>): every option → probability, "floats that sum to 1" (approximately — the pydantic schema description says "values sum to approximately 1", line 19115).
  • confidence (number 0-1): derived from the shape of probabilities.

Score answer:

{
  "type": "score",
  "score": 1.6,
  "legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
  "probabilities": { "0": 0.05, "1": 0.3, "2": 0.65 },
  "confidence": 0.78
}
  • score (number): probability-weighted mean of level indices — "each level number multiplied by its probability, added up" (line 13963), e.g. 0×0.0 + 1×0.70 + 2×0.30 = 1.30. Can land between integer levels.
  • legend (map<string level-index, string|object|array>): echoes each level's description back, keyed by the level's index as a string.
  • probabilities (map<string level-index, number>), string-keyed like legend.
  • confidence (number 0-1).
  • Python SDK note (line 13969): ScoreAnswer re-keys probabilities and legend by integer level in the typed SDK object, even though the wire JSON keys them as strings.

Full worked request/response pair (Quick Start, line 12495-12564)

Request:

{
  "state": "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
  "model": "jev-latest",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this",
      "criteria": {
        "billing": "Payment or subscription issues",
        "technical": "Bugs or integration problems",
        "sales": "Pricing or account questions"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated the customer appears",
      "criteria": [
        "Calm, just stating facts",
        "Frustrated but civil",
        "Very angry, strong language"
      ]
    },
    "is_urgent": {
      "type": "noul",
      "instructions": "The message conveys urgency or time-sensitivity"
    }
  }
}

Response:

{
  "model": "jev-latest",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "billing",
      "probabilities": { "billing": 0.84, "technical": 0.159, "sales": 0.001 },
      "confidence": 0.596
    },
    "frustration": {
      "type": "score",
      "score": 1.035,
      "legend": {
        "0": "Calm, just stating facts",
        "1": "Frustrated but civil",
        "2": "Very angry, strong language"
      },
      "confidence": 0.842
    },
    "is_urgent": { "type": "noul", "noul": 0.999 }
  },
  "usage": { "input_tokens": 312, "output_tokens": 48 }
}

(Note: this particular published example response for frustration omits probabilities in the docs' own JSON, even though the schema requires it — likely a docs-authoring trim, not a real absent-field case; every other Score example in the docs includes probabilities.)

GET /v1/models response (line 12888-12933)

{ "models": [ { "name": "jev-latest", "description": "...", "release_date": "..." } ] }
  • models[].name (string): the model ID/alias, valid in the model request field.
  • models[].description (string).
  • models[].release_date (string).
  • "It currently lists the aliases. Versioned IDs such as jev-1.13.0 are accepted by the model field whether or not they appear in the list."

4. Errors, rate limits, retry semantics

HTTP API error table (line 416-425 — the only error table in the HTTP

API reference)

StatusMeaning
401 UnauthorizedMissing or invalid API key. Check the Authorization header.
422 Unprocessable EntityRequest body failed validation — missing required field or malformed question. "The body details the offending field." (exact shape of that body is not in vendor docs.)
429 Too Many RequestsExceeded rate limit. Back off and retry after a short delay.
529 OverloadedTypeSafe is temporarily overloaded. Retry after a short delay.

"When you receive a 429 or 529... retry the request with exponential backoff instead of retrying immediately. Our client SDKs handle this automatically... with its default retry policy." (line 427-429)

Disagreement/gap noted: the HTTP API reference documents only 401/422/429/529, but the Python SDK's exception hierarchy (line 18165-18339) is broader — it defines dedicated exceptions for 400 (TypeSafeBadRequestError), 403 (TypeSafePermissionDeniedError), 404 (TypeSafeNotFoundError), 422 (TypeSafeUnprocessableEntityError), 429 (TypeSafeRateLimitError), and generic 5xx (TypeSafeInternalServerError — which would presumably catch 529 too, though 529 is never named in the SDK docs). The HTTP reference page is the narrower, curated list; the SDK exception surface is the fuller (implicit) one. Treat both as authoritative for their own scope.

Python SDK exception classes (line 18165-18432)

  • TypeSafeError(Exception) — base for all SDK failures.
  • TypeSafeAPIError(TypeSafeError) — unsuccessful HTTP response. Fields: status (int), body (parsed JSON error body / plain text / None for empty), headers, endpoint (method+URL, no creds/query/fragment), request_id (property, from x-typesafe-request-id or None).
    • TypeSafeBadRequestError (400)
    • TypeSafeAuthenticationError (401)
    • TypeSafePermissionDeniedError (403)
    • TypeSafeNotFoundError (404)
    • TypeSafeUnprocessableEntityError (422)
    • TypeSafeRateLimitError (429) — adds retry_after_ms (parse_retry_after(headers), None if unavailable)
    • TypeSafeInternalServerError (5xx)
  • TypeSafeAPIConnectionError(TypeSafeError, ConnectionError) — no HTTP response at all (network failure).
    • TypeSafeAPITimeoutError(TypeSafeAPIConnectionError, TimeoutError) — adds timeout (seconds or httpx2.Timeout).
  • TypeSafeAPIResponseValidationError(TypeSafeAPIError) — 2xx response whose body was structurally invalid; adds field_path (dotted path to the offending field, e.g. answers.tone.confidence).

Rate-limit headers honored: Retry-After and retry-after-ms (line 16758, 18398) — "Whether to honor Retry-After and retry-after-ms response headers", up to maxRetryAfterMs in the JS SDK's RetryPolicy.

RetryPolicy (Python SDK, line 18346-18430)

from typesafe_sdk import RetryPolicy, TypeSafeClient
client = TypeSafeClient(
    retry=RetryPolicy(max_retries=3, timeout=10.0, http_statuses={429, 500, 502, 503, 504})
)

Fields: max_retries (0 disables), backoff_initial (seconds, doubled each attempt up to backoff_max; 0 disables backoff), backoff_max, backoff_jitter (fraction randomly subtracted, 0-1), http_statuses (retried set), respect_retry_after (bool), api_connection_error (bool, retry on connection failure), api_timeout_error (bool), exceptions (extra exception types to retry), predicate (callable on the raised exception → bool), timeout (total retry budget in seconds across the whole call including delays — "Stops before a retry whose delay would reach or exceed the budget, re-raising the last error").

Rate limits and context/pricing (Models page, line 12839-12934 — the

authoritative numbers table)

Jev 1.13jev-1.13.0
Price (per Btok / per Mtok)$42 / $0.042 — input tokens only; output tokens are free
Rate limits250,000 tokens/second / 1,200 requests/minute
Context length64k tokens per request (state + all questions combined); 32k tokens for state + the single longest question
InputText only: string, JSON object, array of text values. No image/audio/video.

"Rate limits are adjusting dynamically... can change without notice while we [serve growing demand]... Higher limits are available on custom and enterprise plans." (line 12853-12855) — i.e. these are not guaranteed stable numbers.

Internal inconsistency in the vendor docs, flagged not resolved: the Primitives page (line 13328) separately states "The number of questions in one request is limited only by the request's token budget, which the state and the questions share. The budget is around 32,000 tokens, roughly 150,000 characters of English text." This describes the whole request's budget as ~32k, which conflicts with the Models page's more precise split (64k total request / 32k for state + longest single question). Vendor docs win per the task's instruction, but the two vendor passages themselves disagree on which number (32k or 64k) is the ceiling for a many-question request — treat 64k (Models page, more specific and more recently authored-sounding) as the more load-bearing number for total request budget, and 32k as a per-state-plus-single-question sub-limit.

Aliases (line 12857-12870)

AliasPoints toMeaning
jev-latestjev-1.13.0Most recent stable/official release. SDK default.
jev-previewjev-1.13.0Most recent release whether or not official; currently identical to jev-latest (no preview build live).

"If you have tuned confidence thresholds against a specific version, pin that version's ID instead of the alias."


5. State shaping, calibration, sampling — what the docs say to do

State (line 843-902)

  • state is string | JSON object | array; text only, no images/audio/ video (repeated at line 871, 916-918, 12846, 12851).
  • Prefer an object for most requests "so each part of the state has a descriptive name and its relationships remain clear." A plain string is fine only for one simple piece of text (line 868).
  • Every question in a request sees the same state and is evaluated independently — "one primitive's result does not become hidden context that changes another primitive's result" (line 481, 850).
  • Keep only the context relevant to the current questions — "This helps the model avoid distractions and context rot" (line 522).
  • Point questions at specific nested values with a backticked dot/index path, e.g. `support.tickets[0].message` (line 530-536).
  • English is the primary training language; other languages including CJK are accepted but currently lower accuracy (line 871, 12882).

Model jaggedness — jev-1.13 failure modes and mitigations (line

12677-12830, table at 12690-12700)

  1. Literal reading — answers the words written, not the intent. Put boundary cases explicitly in criteria; split ambiguous judgments into two literal questions and combine in code.
  2. Math and numbers — "Jev is not a calculator." Don't ask it to count (characters, occurrences, list items) — count in code (worked example: ask one Noul per item, sum > threshold results in code, line 12720-12737). Numeric representations (hex/RGB, low-level code) score worse than named/semantic equivalents — convert in code, pass semantic labels. Score outputs must not be used to reconstruct an exact number by interpolating between levels — only threshold the expectation.
  3. Date/time comparison — dates are read as text, not ordered quantities; ordering/duration/window checks are unreliable, worse with mixed formats. Extract components as bounded Choices (month/day/year as closed sets, with an explicit "not stated" option), assemble and do all arithmetic in code.
  4. Indirection — double negatives / multi-hop "property of a property" questions lose accuracy. Write instructions directly; name relevant state parts explicitly.
  5. Large state full of irrelevant detail — accuracy falls as unrelated content grows in state ("context rot" reiterated). Filter/ retrieve in code first; when that's not possible, use a Noul as a relevance filter first.
  6. Adversarial content — state is treated as data, not hostile, by default; injected instructions / misleading framing can move the answer. Be explicit in criteria; test adversarial inputs before deploying.
  7. Contradictory instructions/criteria — e.g. a Noul where true maps to "no" semantically will perform worse. Align criteria as an extension of instructions, plain everyday phrasing.
  8. Common-sense structural invariants do NOT hold — worked numeric example: the same yes/no judgment asked as Noul vs. Choice gave noul=0.22 vs. probabilities["yes"]=0.01 — not comparable. Two Nouls for "X" and "not X" summed to 1.19, not 1.0. Don't carry a threshold tuned on one question type over to another; don't expect arithmetic identities between separate questions.
  9. Generation — Jev is not trained to generate text; forcing free text via chained choices "will not work well and will be very slow." Turn bounded extraction into a Choice; use a real generative model for actual text generation.

Summary reminders (line 12818-12825): avoid asking Jev something code can compute exactly, hiding multiple judgments in one question, multi-hop "System Two" tasks, or feeding more state context than a question needs.

Confidence / calibration (line 1149-1230)

  • confidence (0-1) is present on Choice and Score answers only, not Noul. It is a single-number summary of how peaked/flat the probabilities distribution is (line 1156, 1160).
  • "We provide confidence as a convenient measure that fits most use-cases, but you are never locked into our definition... a different measure may serve you better... which is exactly why we give you the full probabilities" (line 1163) — i.e. compute your own statistic from probabilities if confidence doesn't fit.
  • Calibration is trained via RLCD ("Reinforcement Learning for Calibrated Decisions" per the brain notes; vendor docs link it as "RLCD" without spelling out the acronym inline at line 493). "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." (line 922)
  • Recommended pattern: three confidence bands — high (act automatically), medium (confirm/flag/gather more info), low (route to a human/fallback) — with threshold values that scale with the stakes of the specific action, not one global number (line 1174-1229, worked banking-command example).
  • No documented temperature or sampling parameter exists on the API itself — Jev always returns a full deterministic-per-call probability distribution; there's no server-side "sampling temperature" request field anywhere in the vendor docs. (Not in vendor docs: any request field to control determinism/temperature/top-k on the Jev call itself.) Client-side sampling from the returned distribution — as agent.py does — is a caller-side technique, not a documented API feature.

6. From agent.py (jev-plays-pokemon) — concrete harness patterns

Full source read; path: /home/nixos/brains/personal/raw/code/jev-plays-pokemon/agent.py (+ README.md beside it).

Imports / SDK usage confirmed

from typesafe_sdk import Choice, Noul, RetryPolicy, TypeSafeClient

This matches the official typesafe-sdk Python package from the vendor docs exactly (not jevclient, which is an unrelated third-party package — see §8).

Environment variables it reads

  • TYPESAFE_API_KEY — checked directly with os.getenv, raises RuntimeError("TYPESAFE_API_KEY is not set. Create a .env file with your key.") if absent, before ever constructing TypeSafeClient (the SDK itself would also raise TypeSafeError for a missing key, but the agent pre-checks with its own message).
  • TYPESAFE_MAX_RETRIES (default 4)
  • TYPESAFE_BACKOFF_INITIAL (default 1.0 seconds)
  • TYPESAFE_BACKOFF_MAX (default 8.0 seconds)
  • TYPESAFE_TIMEOUT (default 45.0 seconds) — this is the retry budget (RetryPolicy.timeout), not the SDK's default 10s per-HTTP-operation timeout.
  • TYPESAFE_TEMPERATURE (default 2.0) — agent-defined, not an SDK/API concept; used only in _sample_goal (client-side).
  • TYPESAFE_GOAL_FLOOR (default 0.1) — same, agent-defined.
  • Loaded via python-dotenv's load_dotenv() at import time, from a project-root .env file (README: cp .env.example .env).

Retry policy construction (verbatim)

def retry_policy() -> RetryPolicy:
    return RetryPolicy(
        max_retries=_int("TYPESAFE_MAX_RETRIES", 4),
        backoff_initial=_float("TYPESAFE_BACKOFF_INITIAL", 1.0),
        backoff_max=_float("TYPESAFE_BACKOFF_MAX", 8.0),
        backoff_jitter=0.2,
        timeout=_float("TYPESAFE_TIMEOUT", 45.0),
    )

backoff_jitter=0.2 is hardcoded (not env-overridable). RetryPolicy is passed once at client construction (TypeSafeClient(retry=retry_policy(), ...)), not per-call.

How it structures state and questions for a game loop

Two call shapes, both against the same state: dict[str, Any] built elsewhere (state.py, not requested for this digest) and the same shared set of per-button Noul questions (self.action_questions, built once in __init__ via build_questions()):

def build_questions() -> dict[str, Any]:
    return {
        action: Noul(instructions=f"Is the best single action right now to {ACTION_NAMES[action]}?")
        for action in ACTION_FRAMES
    } | {
        "menu_open": Noul(
            instructions=(
                "Is a selectable menu open on screen (start menu, POKEMON/BAG/ITEM list, "
                "battle FIGHT/PKMN/BAG/RUN, or a yes/no question)? A plain text dialog is NOT a menu."
            )
        ),
    }

— 8 action Nouls (press_a, press_b, press_start, walk_up/down/left/right, wait) + 1 menu_open Noul, all fired in one call, every turn. Comment at file top: "Uses one atomic Noul per candidate action (asked in parallel in a single call)... Code then picks the strongest answer with a margin requirement and sanity-checks it against the deterministic walkability map." This is the vendor docs' "speculative fan-out" pattern applied directly.

decide_goal() adds one Choice (next_goal, options = the currently-available high-level goals, e.g. talk_to_Mom, explore, reach_exit) to the same 9 action Nouls, all in a single client.system_one(state=state, questions=questions) call:

questions = {
    "next_goal": Choice(
        instructions={
            "question": "Which goal should RED pursue next?",
            "focus": "Pick the ONE goal that best advances your long-term goal (badges/Champion) "
            "given the screen text, recent hints, nearby objects, walkability and room map. "
            "Code will execute the walking.",
        },
        criteria=criteria,  # {goal_id: goal_desc, ...}
    )
}
questions.update(self.action_questions)
response = self._interruptible(self.client.system_one, state=state, questions=questions)

Note the structured instructions object ({"question": ..., "focus": ...}) for the Choice — exactly the vendor docs' "structured instructions" pattern (§2 above), with caller-chosen, non-reserved field names.

decide_action() (menu/battle micro-decisions) sends only the 9 action Nouls, no Choice.

How it samples from the probability distribution (_sample_goal)

The vendor docs never describe server-side sampling — this is 100% client-side post-processing of the returned choice.probabilities:

def _sample_goal(
    probabilities: dict[str, float],
    temperature: float = 2.0,
    floor: float = 0.05,
) -> tuple[str, float]:
    labels = list(probabilities)
    ...
    probs = [max(0.0, probabilities[g]) for g in labels]
    total = sum(probs)
    probs = [p / total for p in probs]
    # Floor: guarantee every option keeps some mass (p=1 is never 100%).
    floor = max(0.0, floor)
    probs = [p + floor for p in probs]
    total = sum(probs)
    probs = [p / total for p in probs]
    # Softmax flattening: logits = log(p)/T. T>1 pulls the peak down and
    # lifts the tail, keeping a real chance of doing something else.
    if temperature != 1.0 and temperature > 0:
        logits = [math.log(p + 1e-12) / temperature for p in probs]
        m = max(logits)
        exp = [math.exp(x - m) for x in logits]
        exp_total = sum(exp)
        probs = [e / exp_total for e in exp]
    label = random.choices(labels, weights=probs, k=1)[0]
    return label, probs[labels.index(label)]

Two composed transforms before random.choices: (1) an additive floor so no option ever reaches literal 0% or literal 100% mass; (2) softmax flattening of log(p)/T with T=2.0 default so a confidently-wrong p=0.99 pick fires only "~80%" per the docstring, ~60-70% per the README — i.e. deliberately de-sharpening Jev's own calibrated distribution to avoid looping on a single wrong-but-confident goal every turn. The Noul actions (decide_action) are not sampled this way — they use argmax with a margin gate instead (below), because action selection wants determinism + a safety fallback, not exploration.

Anti-stuck / safety tricks (pick_action, deterministic, no model call)

  • Margin + confidence gate: only trust the model's top Noul pick if decision.margin >= 0.08 and decision.confidence >= 0.45 (margin = winning noul minus runner-up noul). Below that, fall back to advancing dialog / open menu (press_a) / walking toward an explore_hint / wait, in that priority order.
  • Menu override: if menu_open Noul ≥ 0.5 and the winning action's confidence < 0.7, force press_a (confirm highlighted option) — "reliably selects the starter (and most story choices)."
  • Deterministic wall check: for any walk_* action, cross-check against a ground-truth walkable: dict[str, bool] map read from emulator RAM (not from the model). If the chosen direction is blocked and not every direction is blocked and not explicitly probing for an exit, override to press_a (if a dialog is active) or wait — i.e. the model's spatial judgment is never trusted over the deterministic collision map except when probing for exits.
  • Anti-pacing: never immediately reverse the previous walk (_REVERSE map) if another open direction exists — picks the next-best-scoring open direction instead, gated at noul >= 0.3.
  • Higher-level (README, not in agent.py itself): stuck detection (press B → wait → re-route) and exit-probing when a room is fully explored live in play.py/navigation.py, not shown in this file.

Threading / interruptibility (_interruptible)

Every blocking client.system_one call is run on a daemon worker thread while the main thread polls a stop_requested() flag every 50ms and raises KeyboardInterrupt itself — because "the HTTP stack swallows SIGINT while it is running." This is a harness-level workaround, not documented anywhere in the vendor SDK docs (not in vendor docs: any statement that the SDK/httpx2 swallows SIGINT — this is the agent author's own empirical finding, not a cited vendor fact).

Measured cost/latency (README, not vendor docs — first-party but

project-specific measurement)

  • "Jev is fast (tens–hundreds of ms per decision)."
  • --turbo mode note: "the Jev API call per decision (~0.5 s) becomes the pacing factor" — i.e. ~500ms observed round-trip in this specific harness/network conditions, at the high end of "tens-hundreds of ms."
  • Token estimate per decision: shared instructions ~850 tok + questions (8 Nouls + menu_open) ~350 tok + per-turn state ~400-900 tok ≈ 1,600-2,100 input tokens/decision.
  • Cost at $0.042/MTok: ≈$0.27/hour at 1 decision/sec, ≈$0.76 per 10,000 decisions.

7. Official Python SDK vs. raw HTTP

Yes, an official SDK exists: package typesafe-sdk on PyPI (line 12573, 17492-17500), requires Python ≥ 3.10 (line 12570) — note agent.py's own project requires Python ≥ 3.14 per its README, which is the harness's own constraint, not the SDK's floor.

pip install typesafe-sdk
# or
uv add typesafe-sdk

Minimal usage (sync, line 17539-17561 / 12583-12617):

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

with TypeSafeClient() as client:  # reads TYPESAFE_API_KEY from env
    response = client.system_one(
        state={"document": "I was charged twice. Please fix this ASAP."},
        questions={
            "billing": Noul(instructions="Is this ticket about billing?"),
            "tone": Choice(
                instructions="What is the customer's tone?",
                criteria={"calm": None, "frustrated": None, "angry": None},
            ),
            "urgency": Score(
                instructions="How urgent is this ticket?",
                criteria=["can wait", "this week", "today"],
            ),
        },
    )

print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)

Async variant: AsyncTypeSafeClient / await client.system_one(...), identical parameter shape, requires async with and await.

TypeSafeClient / AsyncTypeSafeClient constructor params (line 17600-17635): api_key (→ TYPESAFE_API_KEY), model (→ TYPESAFE_DEFAULT_MODEL), retry (RetryPolicy), timeout (float seconds or httpx2.Timeout; SDK default DEFAULT_TIMEOUT = 10.0), headers (extra), transport / http_client (mutually exclusive custom httpx2 transport), base_url (→ TYPESAFE_BASE_URL). Raises TypeSafeError if API key missing or timeout invalid; ValueError if both transport and http_client given.

client.system_one(...) params (line 17685-17735): state (JSONContent), questions (Mapping[str, Question], non-empty — raises TypeSafeError if empty, or if a Score question's criteria list is empty), model (override), retry (per-call override), timeout (per-call override), extra_headers, extra_body (forward-compat, shallow last-write-wins merge over the body), response_model (optional Pydantic BaseModel subclass to type the response — see forward-compat section below).

client.models.list(...) → ListModelsResponse with .models[] of ModelMetadata{name, description, release_date} (line 19770-19925).

Environment variables the SDK itself reads (line 18094-18172, restated at 20161-20172 — this table is the authoritative env-var reference):

VariableConfiguresDefault
TYPESAFE_API_KEYAPI key (required)—
TYPESAFE_BASE_URLAPI root URLhttps://api.typesafe.ai
TYPESAFE_DEFAULT_MODELDefault modeljev-latest
TYPESAFE_LOG_LEVELtypesafe_sdk logger level, applied once at importunset (off)

"Explicit options take precedence over environment variables; empty or whitespace-only environment values are ignored." (line 17592)

Logging: SDK logs to Python logger name typesafe_sdk. info = one summary line/request; debug = also headers+bodies. Secret headers (authorization, API keys, cookies, any header with token/secret in the name) are redacted; request/response bodies are NOT redacted (line 20159).

Forward-compatibility escape hatches (line 20174-20228):

  • extra_body={"beam_width": 4} — send request fields the SDK predates.
  • Raw dict questions instead of Noul/Choice/Score objects, including extra unknown keys (e.g. {"type": "noul", "instructions": "...", "weight": 2}) — "Ignore their type-checking errors and prefer upgrading the SDK instead."
  • Unknown answer kinds: SDK logs a warning and skips them; use result.raw_http_response.json()["answers"] to see everything including unrecognized kinds.
  • Unknown extra fields on recognized responses are silently ignored.

JavaScript/TypeScript SDK also exists (@typesafe-ai/sdk, TypeSafeClient.systemOne(...), line 14286-17471) — not needed for a Python harness, only noted for completeness; its error class names (APIConnectionError, RateLimitError with retryAfterMs, etc.) mirror the Python SDK's one-for-one.

Raw HTTP is fully viable — the cURL example (line 12477-12493) and the plain-dict "raw question dictionaries" pattern above show the wire format is a stable, documented contract independent of any SDK; a hand-rolled requests/httpx client needs only: POST to https://api.typesafe.ai/v1/systemone, Authorization: Bearer header, JSON body per §2, and to implement retry/backoff itself for 429/529 (with Retry-After/retry-after-ms header awareness) since raw HTTP gets none of the SDK's automatic retry behavior.


7b. Measured against the live API, 2026-09-20 (apps/jevprobe, apps/player)

Everything in sections 1-7 was read, not run. This section is the opposite: what the API actually did on this machine, from a hand-rolled raw-HTTP client (packages/jev-http) rather than an SDK.

Nothing in the digest needed correcting. The request shape in §2 was accepted as documented, and the response matched §3 field for field - no unknown fields, probabilities over every option, confidence on the Choice, usage with both token counts. Raw response to a one-Choice request:

{"model":"jev-1.13.0","answers":{"move":{"type":"choice","choice":"walk_south",
 "confidence":0.98,"probabilities":{"press_a":0.01,"walk_south":0.99,
 "walk_north":0.0,"walk_west":0.0}}},"usage":{"input_tokens":544,"output_tokens":54}}
  • model in the response is the resolved versioned id, jev-1.13.0, for a request that sent the alias jev-latest - exactly as §3 says, now seen.
  • Latency. §8 collects five conflicting figures and notes none is in the API reference. Measured here, from a NixOS host over residential fibre, on a request of ~700 input tokens carrying one Choice, one Score and one Noul: cold connection ~400 ms; warm 110-280 ms, typically ~160 ms, over 755 consecutive requests. Keeping one client alive is therefore worth more than twice the cost of the answer itself, which is the argument for a client with a connection pool rather than one TLS handshake per decision.
  • Cost. 755 three-question decisions came to $0.0277 - about $0.000037 each, at 550-900 input tokens per request. At one decision a second that is $0.13 an hour. The budget cap in packages/brain is a stop, not a budget.
  • Every request succeeded on the first attempt. No 429, no 529, no connection failure in 755 requests plus a few hundred more across earlier runs, so the retry policy in §4/§6 is implemented and has never yet fired. That is the shape of the absence, not evidence it is unnecessary.
  • Browser origins are still refused (recorded in the brain page, retested by construction here): the native client uses the API directly, and a web build will need a relay.

8. Cross-source disagreements (vendor docs win; noted for awareness)

  1. Latency figures vary by source and none is in the API reference itself:

    • Vendor docs prose (not the API reference table): "Most queries complete in about 100 ms" (line 489) and "Frontier intelligence at real-time speeds (150ms)" (line 974) — both marketing/concepts pages, not a guaranteed SLA number.
    • jev-plays-pokemon README (first-party measurement, not vendor): "tens–hundreds of ms per decision," ~0.5s observed in turbo mode.
    • Brain page jev.md (secondary, citing TechCrunch/flaviocopes, not vendor docs): "70ms-500ms" end-to-end.
    • Brain page ecosystem/jev-doom-agent.md (secondary, citing TypeSafe's own launch blog — a different document from the docs site that was the assigned source): "roughly 10 decisions per second" (100ms) at "$7/hour."
    • Brain page ecosystem/jevclient.md (third-party unofficial SDK's own measured table, Netherlands, cold-connection): 690-1,336ms for 3-400 questions in one call; ~250-580ms warm-connection.
    • None of these is inside llms-full-2026-09-19.txt's API reference or Models page — the authoritative vendor doc gives no numeric latency SLA at all, only rate limits (250k tok/s, 1,200 req/min) and price. Treat any specific millisecond figure as informal/measured, not contractual.
  2. Token budget for "how many questions fit in one call": internal vendor-doc inconsistency between the Models page's precise 64k/32k split and the Primitives page's rounder "~32,000 tokens total budget" — see §4 above. Not resolvable from the docs as given; use the Models page's more specific numbers for capacity planning and treat 32k as a conservative single-call ceiling if being cautious.

  3. HTTP status code coverage: the HTTP API reference documents only 401/422/429/529; the Python SDK's exception hierarchy additionally covers 400/403/404/generic-5xx. Not a contradiction so much as the HTTP page being an incomplete subset — code defensively for the SDK's fuller list if writing a raw HTTP client.

  4. Score level count: vendor docs say "at least two levels... up to 10" (line 13881); brain page score.md independently says "2-10 level scale" — these agree, no real disagreement, included here only because it was cross-checked.

  5. jevclient (PyPI) is a different, unofficial, third-party SDK, not typesafe-sdk — do not confuse the two. agent.py uses the official typesafe-sdk. jevclient has its own error classes (JevAuthError/JevValidationError/JevRateLimitError/ JevOverloadedError/JevConnectionError) that do not appear anywhere in the vendor docs and should not be assumed to exist on the real API — they are that unofficial package's own wrapper naming.


9. Quick-reference: minimal working Python harness shape

import os
from typesafe_sdk import Choice, Noul, Score, RetryPolicy, TypeSafeClient

# env: TYPESAFE_API_KEY (required), optionally TYPESAFE_BASE_URL,
# TYPESAFE_DEFAULT_MODEL, TYPESAFE_LOG_LEVEL

client = TypeSafeClient(
    retry=RetryPolicy(max_retries=4, backoff_initial=1.0, backoff_max=8.0,
                       backoff_jitter=0.2, timeout=45.0),
)

response = client.system_one(
    state={"document": "..."},          # string | dict | list, text only
    questions={
        "is_urgent": Noul(instructions="Does this convey urgency?"),
        "department": Choice(
            instructions="Which team should handle this?",
            criteria={"billing": "...", "technical": "...", "sales": "..."},
        ),
        "frustration": Score(
            instructions="How frustrated is the customer?",
            criteria=["Calm", "Frustrated", "Very angry"],  # 2-10 levels
        ),
    },
    # model="jev-latest",  # default
)

response.nouls["is_urgent"].noul                # float 0..1
response.choices["department"].choice           # str
response.choices["department"].probabilities    # {opt: float}
response.choices["department"].confidence       # float 0..1
response.scores["frustration"].score            # float, e.g. 1.6
response.scores["frustration"].confidence       # float 0..1
response.usage.input_tokens / .output_tokens    # int
client.close()

Error handling:

from typesafe_sdk import TypeSafeAPIError, TypeSafeRateLimitError

try:
    client.system_one(state, questions)
except TypeSafeRateLimitError as e:
    time.sleep((e.retry_after_ms or 1000) / 1000)  # if not letting RetryPolicy handle it
except TypeSafeAPIError as e:
    print(e.status, e.request_id, e.body)

10. What the vendor does not expose - checked directly for the budget-safety

work, 2026-09-20

The earlier sections cite the vendor docs closely enough that an absence could still be an oversight in the digest rather than in the docs. This section is a direct re-check of the primary source itself (~/brains/personal/raw/vendor-docs/typesafe/llms-full-2026-09-19.txt, all 20,229 lines) plus one live, billed request made specifically to look at headers, run for packages/jev-http's spend ledger and rate-limit handling (packages/jev-http/src/ledger.rs, apps/jevprobe --headers).

No usage, balance or credit endpoint exists (vendor docs, grepped directly)

Searching the whole vendor doc file for balance, credit, /usage, /v1/usage, /v1/billing and 402 turns up only unrelated worked examples: a banking-agent demo whose fictional tool is named check_balance, an invoice-processing example that asks Jev to classify money as a credit or a charge, and the JS/Python SDK's Usage type - which is {input_tokens, output_tokens} on one API response (§3 above), not an account-level usage/spend endpoint. There is no vendor equivalent of a GET /v1/account/usage or /v1/billing route anywhere in the reference. GET /v1/models (§1) is the only account-adjacent GET endpoint that exists, and it lists model aliases, not spend. So packages/jev-http's ledger cannot reconcile against a vendor-reported figure - there is nothing to reconcile against - and the budget MCP tool's "vendor-reported balance" field is always null, permanently, unless a future version of the API adds one.

No rate-limit headers are documented, and none arrived on a real response

Searching the same file for x-ratelimit, ratelimit-, rate-limit- finds no header of that shape anywhere in the HTTP reference, the Python SDK docs, or the JS SDK docs. The only rate-limit-adjacent headers named anywhere in the vendor docs are Retry-After and retry-after-ms (§4), and both are described only in the context of a 429/529 response, never as present on an ordinary 2xx.

Confirmed empirically, not just from the docs' silence: apps/jevprobe --headers (added for this work) made one real, billed request and printed every response header verbatim:

date                           Mon, 21 Sep 2026 02:37:59 GMT
server                         istio-envoy
content-length                 231
content-type                   application/json
x-typesafe-request-id          req_01a0c1d3d6187b088bd31bc35d654f10
x-envoy-upstream-service-time  81

Six headers, none of them about rate limits or quota. server: istio-envoy is new information not in any digest section above - the API sits behind an Envoy/Istio ingress - and is otherwise unremarkable. packages/jev-http still parses the conventional x-ratelimit-*/ratelimit-* spellings opportunistically into a typed RateLimit snapshot on every response (parse_rate_limit), so if the vendor ever starts sending one, every caller starts seeing it with no code changed; today that snapshot is None on every response, confirmed rather than assumed.

Consequence for the client: since there is no live signal to react to, packages/jev-http::Jev enforces the vendor's own documented static ceilings (§4: 250,000 tokens/second, 1,200 requests/minute) with a client-side token-bucket, rather than reading a header. In practice this bucket almost never engages, because the app's own circuit breaker (30 questions/minute default) is far stricter than the vendor's 1,200/minute - the vendor bucket exists for the case where that default is loosened well past its own default, not because it is expected to bind day to day.

No out-of-credit response shape is documented anywhere

Searched for 402, "payment required", and every billing-sounding word above: nothing. The HTTP reference's error table (§4) stops at 401/422/429/529; the Python SDK's broader exception hierarchy (§4) adds 400/403/404/generic-5xx but never a billing-specific status or exception class. There is no way to know, from the docs, what an account with no credit left gets back. packages/jev::Failure::OutOfCredit is therefore an inferred, not a documented, classification: a 402 (the conventional "payment required" code, even though the vendor never writes it down) or a 403 whose body mentions credit, balance or quota. Both are guesses about a case this project has never hit (measured spend as of 2026-09-20, this task included, is a few cents against a five-dollar lifetime credit) and cannot safely go looking for by spending down the account to find out. Should a real out-of-credit response ever arrive with a different shape, Failure::of's catch-all still sorts it as Invalid rather than panicking or silently retrying it - it is just not labelled OutOfCredit until this classification is corrected against a real example.

Summary table

Vendor exposesChecked howResult
A usage/balance/credit endpointFull-text search of the vendor docsNo - GET /v1/models is the only account-adjacent read, and it lists aliases, not spend
Rate-limit headers on a normal responseFull-text search + one live request (jevprobe --headers)No - six headers arrived, none about limits or quota
Retry-After/retry-after-ms on 429/529Vendor docs (§4)Yes - the only rate-limit-adjacent headers documented at all
A documented out-of-credit response shapeFull-text searchNo - 402 is never mentioned; Failure::OutOfCredit is an inferred guess, not a cited fact
x-typesafe-request-idVendor docs (§1) + confirmed liveYes, on every response checked