lmjtfy.git / tools / eval / CLAUDE.md
1@README.md
2
3- **The eval must use `llm::request`, `llm::parse` and `ask::check`
4  unchanged.** A second prompt or a more forgiving parser here would pass
5  models the Worker then fails with.
6- **Changing the system prompt or the tools invalidates the result.** Rerun,
7  update the table in the README, and set `LLM_MODEL` in
8  `apps/lmjtfy/wrangler.toml` to whatever now passes.
9- **A case's `accept` lists every reading a careful person would allow, and
10  no more.** "How likely is X" is a Noul, because a Noul's probability is the
11  likelihood; it was first written as a Score and failed a model for being right.