Chapter 14: eval, choosing an LLM with a test instead of a hunch
There are a dozen Workers AI models that can call tools, at prices that differ by more than tenfold. Which one should write Jev's questions? The rule (the owner, 2026-10-02): the cheapest model that writes a correct Jev tool call every time, with Claude Haiku's price as the ceiling. Not the smartest, not the newest: the cheapest that passes.
So this program asks each candidate, cheapest first, to do the job on a set
of cases, and stops at the first that gets them all right. It uses the
Worker's own three steps (llm::request, llm::parse, ask::check), so a
pass here is a pass on the site.
Aside. The cases have opinions. "How likely is X" counts as a yes-or-no question, because the probability of yes is the likelihood. It was first written as a how-much, and failed a model for being right. The case was fixed, not the model.
nix develop .#owner
lmjtfy-eval # cheapest first, stop at the first that passes every case
lmjtfy-eval --all # every candidate
lmjtfy-eval @cf/qwen/qwen3-30b-a3b-fp8 # one model, printing each reply
The candidates and their prices are llm::CANDIDATES (chapter 8). The cases
are CASES in src/main.rs: four yes-or-no questions, four with a best
answer, three of degree, one "how likely", one with quotes and symbols, and
one that asks two things. A case passes when every tool call is well formed,
Jev's protocol accepts it, and the tools called are the ones the case
expects. A run spends real neurons from the account's daily 10,000: 14
requests a model.
Result, 2026-10-02
| Model | Passed | Neurons a call | Seconds a call |
|---|---|---|---|
@cf/ibm-granite/granite-4.0-h-micro | 11 / 14 | 2.4 | 3.6 |
@cf/qwen/qwen3-30b-a3b-fp8 | 14 / 14 | 16.0 | 3.1 |
granite called jev_choice for "should I rewrite it in rust", left
yes_means out of a jev_noul, and made one call for the two-part question.
qwen3 is the Worker's LLM_MODEL. At 16 neurons a call the free allocation
covers about 625 questions a day. The more expensive candidates were not run.
The facts: lmjtfy-eval facts
The eight questions Jev is asked first (chapter 7) decide everything: what is refused, what Jev answers alone, what reaches the LLM, what goes on the feed. Their wording is behaviour, and this checks it.
op-env-run -- lmjtfy-eval facts
src/facts.rs has 25 inputs, each with what the facts should be: answerable
or not, which kinds it may be read as, whether it is several questions,
whether it needs a scale of its own, whether it is fit to show. The Worker's
own request (ask::wanted with every Jev fact) goes to Jev through
jev-http, so the estate's shared spend ledger admits it, and ask::learned
reads the answers as the Worker does.
Run it after changing a fact question's wording (packages/ask), the split
threshold (ask::ALSO) or the Jev model. A run costs under a tenth of a
cent. The key is LMJTFY_TYPESAFE_API_KEY, from op-env-run.
Result, 2026-10-02: 25 / 25, about 920 tokens a request.
In this folder
| Path | What |
|---|---|
| src/ | The two evals. |
| Cargo.toml | The program: llm, ask and rules, ureq for Cloudflare's REST API, and jev-http for the facts eval. |
← Previous: Chapter 13, tools/ · Up: tools · Next: Chapter 14½, eval/src/ →