jevstrudel.git / tools / critic / README.md
1# tools/critic
2
3Jev's art critic. `critic.mjs` asks Jev whether each revision of each song is
4art, and records the answer in the song's `SPEC.md` frontmatter:
5- `art:` on every revision, with `artRuns:` under it;
6- `critic:` (date, model, `method`, `art`, `artRuns`, and each criterion's
7  score) for the song as it stands.
8
9Each revision is asked several times (`--runs`, default 3) and the mean is
10recorded: `art` and every criterion are means over the runs, rounded as one
11run's are (art to the hundredth, criteria to the tenth), and `artRuns` lists
12each run's own art in the order asked. One run moves by up to 0.06 on
13unchanged code (measured 2026-09-25 across 14 songs), enough to reorder the
14song list, which is sorted by art. Only the recorded scores are averaged:
15the REPL's live "is it art?" asks once.
16
17The song's card shows `critic:` as "N% art". The spec view on the welcome
18tab lists each revision's score, and the change from the first scored
19revision to the last.
20
21Every score comes from the same rubric in `website/src/jev/critic.mjs`:
22six criteria scored 0-5 on absolute ladders in one call, with the song's
23title, spec body, code and facts measured from the code as the state; `art`
24is their mean out of 5. The REPL's own art check uses it too, so all the
25numbers are comparable. Its header says why it replaced the single yes/no
26question.
27
28## Running it
29
30```sh
31op-env-run -- node tools/critic/critic.mjs            # dry run: lists what it would score
32op-env-run -- node tools/critic/critic.mjs --write    # scores unscored revisions and records them
33op-env-run -- node tools/critic/critic.mjs --write --song jev/dial-up   # one song
34op-env-run -- node tools/critic/critic.mjs --write --all                # rescore everything
35op-env-run -- node tools/critic/critic.mjs --write --runs 5             # five runs per revision
36```
37
38It calls TypeSafe directly with `JEVSTRUDEL_TYPESAFE_API_KEY`, which
39`op-env-run` puts in its environment. Each call is a few thousand input
40tokens, about $0.0001, and a revision costs `--runs` calls.
41
42## Which code a revision is
43
44- **`output` names a file in the song folder** (`song.js`, `glitch.js`): that
45  file is scored, against the current spec.
46- **`output: null`** (overwritten): the code is recovered from git, as
47  `song.js` and `SPEC.md` stood just before the commit that recorded the next
48  revision. If that snapshot isn't the revision's own (for example, it
49  predates this repo), the revision gets `art: null` with a comment saying
50  why.
51
52Even averaged, scores can move by a point or so between rescores, and after
53code fixes that aren't revisions, so a rescore can change a song's card by a
54point. `artRuns` shows how far one revision's runs disagreed.