lmjtfy.git / tools / eval / README.md

Chapter 14: eval, choosing an LLM with a test instead of a hunch

There are a dozen Workers AI models that can call tools, at prices that differ by more than tenfold. Which one should write Jev's questions? The rule (the owner, 2026-10-02): the cheapest model that writes a correct Jev tool call every time, with Claude Haiku's price as the ceiling. Not the smartest, not the newest: the cheapest that passes.

So this program asks each candidate, cheapest first, to do the job on a set of cases, and stops at the first that gets them all right. It uses the Worker's own three steps (llm::request, llm::parse, ask::check), so a pass here is a pass on the site.

Aside. The cases have opinions. "How likely is X" counts as a yes-or-no question, because the probability of yes is the likelihood. It was first written as a how-much, and failed a model for being right. The case was fixed, not the model.

nix develop .#owner
lmjtfy-eval              # cheapest first, stop at the first that passes every case
lmjtfy-eval --all        # every candidate
lmjtfy-eval @cf/qwen/qwen3-30b-a3b-fp8   # one model, printing each reply

The candidates and their prices are llm::CANDIDATES (chapter 8). The cases are CASES in src/main.rs: four yes-or-no questions, four with a best answer, three of degree, one "how likely", one with quotes and symbols, and one that asks two things. A case passes when every tool call is well formed, Jev's protocol accepts it, and the tools called are the ones the case expects. A run spends real neurons from the account's daily 10,000: 14 requests a model.

Result, 2026-10-02

ModelPassedNeurons a callSeconds a call
@cf/ibm-granite/granite-4.0-h-micro11 / 142.43.6
@cf/qwen/qwen3-30b-a3b-fp814 / 1416.03.1

granite called jev_choice for "should I rewrite it in rust", left yes_means out of a jev_noul, and made one call for the two-part question. qwen3 is the Worker's LLM_MODEL. At 16 neurons a call the free allocation covers about 625 questions a day. The more expensive candidates were not run.

The facts: lmjtfy-eval facts

The eight questions Jev is asked first (chapter 7) decide everything: what is refused, what Jev answers alone, what reaches the LLM, what goes on the feed. Their wording is behaviour, and this checks it.

op-env-run -- lmjtfy-eval facts

src/facts.rs has 25 inputs, each with what the facts should be: answerable or not, which kinds it may be read as, whether it is several questions, whether it needs a scale of its own, whether it is fit to show. The Worker's own request (ask::wanted with every Jev fact) goes to Jev through jev-http, so the estate's shared spend ledger admits it, and ask::learned reads the answers as the Worker does.

Run it after changing a fact question's wording (packages/ask), the split threshold (ask::ALSO) or the Jev model. A run costs under a tenth of a cent. The key is LMJTFY_TYPESAFE_API_KEY, from op-env-run.

Result, 2026-10-02: 25 / 25, about 920 tokens a request.

In this folder

PathWhat
src/The two evals.
Cargo.tomlThe program: llm, ask and rules, ureq for Cloudflare's REST API, and jev-http for the facts eval.

← Previous: Chapter 13, tools/ · Up: tools · Next: Chapter 14½, eval/src/ →