Chapter 8: llm, the eloquent one on a short leash
Sometimes a question needs writing before Jev can judge it. "What is the best pizza topping?" is a pick, and a pick needs options; Jev cannot make them up. "How hot is the sun?" needs a scale in kelvin, not "Slightly" to "Extremely". "Is Rust fast and is it hard to learn?" is two questions wearing one coat. For these, and only these, an LLM is called.
It is told plainly what it is for. This is the start of its system prompt:
You sit between a person and Jev. Jev is a model that cannot write text. It only judges, in three ways [...] Turn the person's input into the tool call that answers it. Always call a tool. Never answer the question yourself, and write no other text.
It gets three tools, one per Jev type, and it answers by calling them:
| Tool | Jev type | What the LLM writes |
|---|---|---|
jev_noul | Noul | A yes-or-no question, with what yes and no mean. |
jev_choice | Choice | 2 to 8 real, specific options, each with a one-line description. |
jev_score | Score | 3 to 7 levels, lowest first (Jev takes 2 to 10). |
At most 4 calls per input. What it writes is then judged by Jev, and only Jev's answer is ever shown as the answer: never the LLM's own words.
Aside: the leash. When the rules have already worked out what kind of question it is (say, a pick), the LLM is told so and given only that tool. Otherwise it might helpfully also write a yes-or-no version, Jev would answer that too, and the page would show two answers that disagree. That happened once (2026-10-02), which is why
takesnow refuses any call of a kind the rules did not ask for.
Try it. Ask https://lmjtfy.fun/?q=what+is+the+best+pizza+topping and open the LLM's panel under the answer: the request it was sent, and the tool call it wrote back, byte for byte.
Which model is the LLM? The cheapest Workers AI model that gets every test
case right, with Claude Haiku's price as the ceiling. Chapter 14 is how that
is decided; today it is @cf/qwen/qwen3-30b-a3b-fp8.
For the people who maintain it
request(input, wants)is the Workers AI request body as JSON text: the system prompt, the input, and the tools.wantsis what the rules want written (rules::Network::wants); when it settles the kind, only those tools are given and the prompt says so.takes(wants, draft)says whether a tool call is of a kind that was asked for. One that is not is not sent.parse(body)reads a chat-completion reply intoToolCalls. Each carries the arguments as the model wrote them and aDraft(the question for Jev) or the reason the arguments do not make one. It repairs exactly one thing: arguments encoded as JSON twice.CANDIDATESis the price list of the models that could do the job, andModelturns usage into neurons and a request into its worst case.FREE_NEURONS_PER_DAYis Workers AI's free daily allocation.
The Worker sends the request through the AI binding and the eval sends it
over REST. Both read the reply with parse.
In this folder
| Path | What |
|---|---|
| src/ | The prompt, the tools, the parser, the price list. |
| Cargo.toml | The crate: rules, serde and serde_json. |
← Previous: Chapter 7½, ask/src/ · Up: packages · Next: Chapter 8½, llm/src/ →