lmjtfy.git / tools / eval / README.md
1# Chapter 14: eval, choosing an LLM with a test instead of a hunch
2
3There are a dozen Workers AI models that can call tools, at prices that
4differ by more than tenfold. Which one should write Jev's questions? The
5rule (the owner, 2026-10-02): **the cheapest model that writes a correct Jev
6tool call every time, with Claude Haiku's price as the ceiling.** Not the
7smartest, not the newest: the cheapest that passes.
8
9So this program asks each candidate, cheapest first, to do the job on a set
10of cases, and stops at the first that gets them all right. It uses the
11Worker's own three steps (`llm::request`, `llm::parse`, `ask::check`), so a
12pass here is a pass on the site.
13
14> **Aside.** The cases have opinions. "How likely is X" counts as a
15> yes-or-no question, because the probability of yes *is* the likelihood. It
16> was first written as a how-much, and failed a model for being right. The
17> case was fixed, not the model.
18
19    nix develop .#owner
20    lmjtfy-eval              # cheapest first, stop at the first that passes every case
21    lmjtfy-eval --all        # every candidate
22    lmjtfy-eval @cf/qwen/qwen3-30b-a3b-fp8   # one model, printing each reply
23
24The candidates and their prices are `llm::CANDIDATES` (chapter 8). The cases
25are `CASES` in `src/main.rs`: four yes-or-no questions, four with a best
26answer, three of degree, one "how likely", one with quotes and symbols, and
27one that asks two things. A case passes when every tool call is well formed,
28Jev's protocol accepts it, and the tools called are the ones the case
29expects. A run spends real neurons from the account's daily 10,000: 14
30requests a model.
31
32## Result, 2026-10-02
33
34| Model | Passed | Neurons a call | Seconds a call |
35| --- | --- | --- | --- |
36| `@cf/ibm-granite/granite-4.0-h-micro` | 11 / 14 | 2.4 | 3.6 |
37| `@cf/qwen/qwen3-30b-a3b-fp8` | 14 / 14 | 16.0 | 3.1 |
38
39granite called `jev_choice` for "should I rewrite it in rust", left
40`yes_means` out of a `jev_noul`, and made one call for the two-part question.
41qwen3 is the Worker's `LLM_MODEL`. At 16 neurons a call the free allocation
42covers about 625 questions a day. The more expensive candidates were not run.
43
44## The facts: `lmjtfy-eval facts`
45
46The eight questions Jev is asked first (chapter 7) decide everything: what
47is refused, what Jev answers alone, what reaches the LLM, what goes on the
48feed. Their wording is behaviour, and this checks it.
49
50    op-env-run -- lmjtfy-eval facts
51
52`src/facts.rs` has 25 inputs, each with what the facts should be: answerable
53or not, which kinds it may be read as, whether it is several questions,
54whether it needs a scale of its own, whether it is fit to show. The Worker's
55own request (`ask::wanted` with every Jev fact) goes to Jev through
56`jev-http`, so the estate's shared spend ledger admits it, and `ask::learned`
57reads the answers as the Worker does.
58
59Run it after changing a fact question's wording (`packages/ask`), the split
60threshold (`ask::ALSO`) or the Jev model. A run costs under a tenth of a
61cent. The key is `LMJTFY_TYPESAFE_API_KEY`, from `op-env-run`.
62
63Result, 2026-10-02: 25 / 25, about 920 tokens a request.
64
65## In this folder
66
67| Path | What |
68| --- | --- |
69| [src/](src/) | The two evals. |
70| [Cargo.toml](Cargo.toml) | The program: `llm`, `ask` and `rules`, `ureq` for Cloudflare's REST API, and `jev-http` for the facts eval. |
71
72← Previous: [Chapter 13, tools/](../) · Up: [tools](../) · Next: [Chapter 14½, eval/src/](src/) →