Chapter 14½: inside eval/src

Two files, one per question the eval answers.

  • main.rs: which model is the LLM. It opens with what a run costs, then CASES (each input with the tools a careful person would accept), then the loop: build the Worker's request, send it over Cloudflare's REST API with the token read from 1Password for that run, parse the reply as the Worker does, and score it.
  • facts.rs: whether Jev's first request reads inputs the way the rules need. 25 cases, each with the facts it should come back with, sent through jev-http so the shared spend ledger admits it.

Both spend real money or neurons, so neither runs in cargo test.

FileWhat
main.rsThe LLM eval, and the entry point for lmjtfy-eval.
facts.rsThe facts eval, lmjtfy-eval facts.

← Previous: Chapter 14, eval/ · Up: eval · Next: Chapter 15, ds-bundle/ →