1# Chapter 14: eval, choosing an LLM with a test instead of a hunch 2 3There are a dozen Workers AI models that can call tools, at prices that 4differ by more than tenfold. Which one should write Jev's questions? The 5rule (the owner, 2026-10-02): **the cheapest model that writes a correct Jev 6tool call every time, with Claude Haiku's price as the ceiling.** Not the 7smartest, not the newest: the cheapest that passes. 8 9So this program asks each candidate, cheapest first, to do the job on a set 10of cases, and stops at the first that gets them all right. It uses the 11Worker's own three steps (`llm::request`, `llm::parse`, `ask::check`), so a 12pass here is a pass on the site. 13 14> **Aside.** The cases have opinions. "How likely is X" counts as a 15> yes-or-no question, because the probability of yes *is* the likelihood. It 16> was first written as a how-much, and failed a model for being right. The 17> case was fixed, not the model. 18 19 nix develop .#owner 20 lmjtfy-eval # cheapest first, stop at the first that passes every case 21 lmjtfy-eval --all # every candidate 22 lmjtfy-eval @cf/qwen/qwen3-30b-a3b-fp8 # one model, printing each reply 23 24The candidates and their prices are `llm::CANDIDATES` (chapter 8). The cases 25are `CASES` in `src/main.rs`: four yes-or-no questions, four with a best 26answer, three of degree, one "how likely", one with quotes and symbols, and 27one that asks two things. A case passes when every tool call is well formed, 28Jev's protocol accepts it, and the tools called are the ones the case 29expects. A run spends real neurons from the account's daily 10,000: 14 30requests a model. 31 32## Result, 2026-10-02 33 34| Model | Passed | Neurons a call | Seconds a call | 35| --- | --- | --- | --- | 36| `@cf/ibm-granite/granite-4.0-h-micro` | 11 / 14 | 2.4 | 3.6 | 37| `@cf/qwen/qwen3-30b-a3b-fp8` | 14 / 14 | 16.0 | 3.1 | 38 39granite called `jev_choice` for "should I rewrite it in rust", left 40`yes_means` out of a `jev_noul`, and made one call for the two-part question. 41qwen3 is the Worker's `LLM_MODEL`. At 16 neurons a call the free allocation 42covers about 625 questions a day. The more expensive candidates were not run. 43 44## The facts: `lmjtfy-eval facts` 45 46The eight questions Jev is asked first (chapter 7) decide everything: what 47is refused, what Jev answers alone, what reaches the LLM, what goes on the 48feed. Their wording is behaviour, and this checks it. 49 50 op-env-run -- lmjtfy-eval facts 51 52`src/facts.rs` has 25 inputs, each with what the facts should be: answerable 53or not, which kinds it may be read as, whether it is several questions, 54whether it needs a scale of its own, whether it is fit to show. The Worker's 55own request (`ask::wanted` with every Jev fact) goes to Jev through 56`jev-http`, so the estate's shared spend ledger admits it, and `ask::learned` 57reads the answers as the Worker does. 58 59Run it after changing a fact question's wording (`packages/ask`), the split 60threshold (`ask::ALSO`) or the Jev model. A run costs under a tenth of a 61cent. The key is `LMJTFY_TYPESAFE_API_KEY`, from `op-env-run`. 62 63Result, 2026-10-02: 25 / 25, about 920 tokens a request. 64 65## In this folder 66 67| Path | What | 68| --- | --- | 69| [src/](src/) | The two evals. | 70| [Cargo.toml](Cargo.toml) | The program: `llm`, `ask` and `rules`, `ureq` for Cloudflare's REST API, and `jev-http` for the facts eval. | 71 72← Previous: [Chapter 13, tools/](../) · Up: [tools](../) · Next: [Chapter 14½, eval/src/](src/) →