Chapter 8: llm, the eloquent one on a short leash

Sometimes a question needs writing before Jev can judge it. "What is the best pizza topping?" is a pick, and a pick needs options; Jev cannot make them up. "How hot is the sun?" needs a scale in kelvin, not "Slightly" to "Extremely". "Is Rust fast and is it hard to learn?" is two questions wearing one coat. For these, and only these, an LLM is called.

It is told plainly what it is for. This is the start of its system prompt:

You sit between a person and Jev. Jev is a model that cannot write text. It only judges, in three ways [...] Turn the person's input into the tool call that answers it. Always call a tool. Never answer the question yourself, and write no other text.

It gets three tools, one per Jev type, and it answers by calling them:

ToolJev typeWhat the LLM writes
jev_noulNoulA yes-or-no question, with what yes and no mean.
jev_choiceChoice2 to 8 real, specific options, each with a one-line description.
jev_scoreScore3 to 7 levels, lowest first (Jev takes 2 to 10).

At most 4 calls per input. What it writes is then judged by Jev, and only Jev's answer is ever shown as the answer: never the LLM's own words.

Aside: the leash. When the rules have already worked out what kind of question it is (say, a pick), the LLM is told so and given only that tool. Otherwise it might helpfully also write a yes-or-no version, Jev would answer that too, and the page would show two answers that disagree. That happened once (2026-10-02), which is why takes now refuses any call of a kind the rules did not ask for.

Try it. Ask https://lmjtfy.fun/?q=what+is+the+best+pizza+topping and open the LLM's panel under the answer: the request it was sent, and the tool call it wrote back, byte for byte.

Which model is the LLM? The cheapest Workers AI model that gets every test case right, with Claude Haiku's price as the ceiling. Chapter 14 is how that is decided; today it is @cf/qwen/qwen3-30b-a3b-fp8.

For the people who maintain it

  • request(input, wants) is the Workers AI request body as JSON text: the system prompt, the input, and the tools. wants is what the rules want written (rules::Network::wants); when it settles the kind, only those tools are given and the prompt says so.
  • takes(wants, draft) says whether a tool call is of a kind that was asked for. One that is not is not sent.
  • parse(body) reads a chat-completion reply into ToolCalls. Each carries the arguments as the model wrote them and a Draft (the question for Jev) or the reason the arguments do not make one. It repairs exactly one thing: arguments encoded as JSON twice.
  • CANDIDATES is the price list of the models that could do the job, and Model turns usage into neurons and a request into its worst case. FREE_NEURONS_PER_DAY is Workers AI's free daily allocation.

The Worker sends the request through the AI binding and the eval sends it over REST. Both read the reply with parse.

In this folder

PathWhat
src/The prompt, the tools, the parser, the price list.
Cargo.tomlThe crate: rules, serde and serde_json.

← Previous: Chapter 7½, ask/src/ · Up: packages · Next: Chapter 8½, llm/src/ →