tools/critic

Jev's art critic. critic.mjs asks Jev whether each revision of each song is art, and records the answer in the song's SPEC.md frontmatter:

  • art: on every revision, with artRuns: under it;
  • critic: (date, model, method, art, artRuns, and each criterion's score) for the song as it stands.

Each revision is asked several times (--runs, default 3) and the mean is recorded: art and every criterion are means over the runs, rounded as one run's are (art to the hundredth, criteria to the tenth), and artRuns lists each run's own art in the order asked. One run moves by up to 0.06 on unchanged code (measured 2026-09-25 across 14 songs), enough to reorder the song list, which is sorted by art. Only the recorded scores are averaged: the REPL's live "is it art?" asks once.

The song's card shows critic: as "N% art". The spec view on the welcome tab lists each revision's score, and the change from the first scored revision to the last.

Every score comes from the same rubric in website/src/jev/critic.mjs: six criteria scored 0-5 on absolute ladders in one call, with the song's title, spec body, code and facts measured from the code as the state; art is their mean out of 5. The REPL's own art check uses it too, so all the numbers are comparable. Its header says why it replaced the single yes/no question.

Running it

op-env-run -- node tools/critic/critic.mjs            # dry run: lists what it would score
op-env-run -- node tools/critic/critic.mjs --write    # scores unscored revisions and records them
op-env-run -- node tools/critic/critic.mjs --write --song jev/dial-up   # one song
op-env-run -- node tools/critic/critic.mjs --write --all                # rescore everything
op-env-run -- node tools/critic/critic.mjs --write --runs 5             # five runs per revision

It calls TypeSafe directly with JEVSTRUDEL_TYPESAFE_API_KEY, which op-env-run puts in its environment. Each call is a few thousand input tokens, about $0.0001, and a revision costs --runs calls.

Which code a revision is

  • output names a file in the song folder (song.js, glitch.js): that file is scored, against the current spec.
  • output: null (overwritten): the code is recovered from git, as song.js and SPEC.md stood just before the commit that recorded the next revision. If that snapshot isn't the revision's own (for example, it predates this repo), the revision gets art: null with a comment saying why.

Even averaged, scores can move by a point or so between rescores, and after code fixes that aren't revisions, so a rescore can change a song's card by a point. artRuns shows how far one revision's runs disagreed.