lmjtfy.git / tools / eval

For agents, on top of README.md, which they read first.

  • The eval must use llm::request, llm::parse and ask::check unchanged. A second prompt or a more forgiving parser here would pass models the Worker then fails with.
  • Changing the system prompt or the tools invalidates the result. Rerun, update the table in the README, and set LLM_MODEL in apps/lmjtfy/wrangler.toml to whatever now passes.
  • A case's accept lists every reading a careful person would allow, and no more. "How likely is X" is a Noul, because a Noul's probability is the likelihood; it was first written as a Score and failed a model for being right.