postjevsql.git / tests / CLAUDE.md

For agents, on top of README.md, which they read first.

  • Always SET statement_timeout in a test, so a hang fails fast.

  • Start instances only through TestPostgres; its PDEATHSIG is what makes teardown leak-proof. Never spawn postgres another way. It listens on a Unix socket in its own temp dir only; don't give it a TCP port again (a released port-0 pick can be taken before it binds).

  • Never assert an upper bound on wall-clock time. The overseer's check shares the machine with agents' buck builds, and "took under 2 × delay" failed there on 2026-09-26 while passing alone. Measure the thing itself at the mock instead: peak_in_flight() for concurrency, accepts()/connections() for redials, request headers for retries, the error message for which cancel fired. Lower bounds (a backoff waited at least its delay) are safe and stay.

  • A failed run that names no test is buck's, not ours. Three such failures had one cause: buck2 2026-08-22 is built with h2 0.4.18, whose small-DATA-frame budget is a fixed 25,600 bytes. A burst of tiny gRPC messages that the receiver has not read yet exhausts it, and h2 closes that connection with GOAWAY(ENHANCE_YOUR_CALM, "too_many_data_frames"). Upstream traced the event-bus and forkserver errors to it (facebook/buck2 1cc3e05f85, 2026-09-11) and fixed the test-executor channels in the same commit; the second one below is the same close seen from a channel that cannot redial:

    • "Buck daemon event bus encountered an error … too_many_data_frames" is the client's connection to the daemon.
    • "Failed to notify test executor of end-of-tests … Cannot reconnect after connection loss" is the daemon's connection to the test executor. That channel wraps the executor's pipes, so tonic's reconnect has nothing to redial (buck2_grpc/src/channel.rs make_channel), and the GOAWAY itself is never logged.
    • "Forkserver is unavailable … Forkserver exited exit status: 0" on every action is the forkserver's connection. spawn_oneshot then shuts the forkserver down cleanly, and it is never respawned, so every later run fails until the daemon restarts. A worktree's daemon is that worktree's alone, so buck2 kill there is safe; the main checkout's is shared (root CLAUDE.md).

    The fix raises every buck-internal connection window to 64 MiB on h2 0.4.19, whose budget is half the window. It shipped in buck2 2026-09-15, which the flake takes from nix-pkgs (the estate-wide copy, shared/pkgs/by-name/bu/buck2 in nixos-config) until nixpkgs#560478 lands; that package fails to evaluate once nixpkgs catches up. All three happened on 2026-09-26 on 2026-08-22 while the suite passed in full on the same commit. If one recurs on 2026-09-15, read the cause in buck-out/v2/log/<trace>/command_report.json and rerun; don't change a test for it.

  • Two things that were suspected of the same failures, checked on 2026-09-26:

    • The devshell rewriting .buckconfig.local on every entry did not restart the daemon or its forkserver (same PIDs before and after), but buck logged every write as a changed file. flake.nix now writes it by rename, and only when its content changes (same mtime and inode across two entries).
    • A second buck2 test //... on the daemon while one was running broke the later one's test executor, on 2026-08-22, with or without a config rewrite between them. That fits the frame-budget cause above, since two runs double the burst. run-check.sh still queues checks in one checkout on buck-out/run-check.lock: Studio's overseer and a worker both run it, and a shared daemon gains nothing from racing. Don't run buck2 test beside it by hand.
    • The lock must not reach buck. run-check.sh first exec'd buck while holding it, and a daemon that run started inherited the descriptor (buck closes none it inherits) and held the lock for its whole life: every later check in that checkout waited forever, and a Studio agent was killed at its run budget waiting (2026-09-26). buck now runs as a child with fd 9 closed; reproduced both ways with a probe daemon.
    • To wait for another checkout's run, don't loop on pgrep -f 'buck2 test //': the pattern is in the waiting shell's own command line, so it always matches itself. A worker lost 10 of its 25 minutes that way and was killed at the budget right after finishing (2026-09-26). Wait on that checkout's lock instead, flock <checkout>/buck-out/run-check.lock true; pgrep -x buck2 would match its idle daemon too. A cold run-check.sh in a fresh worktree builds every major and took about 14 minutes on 3 cores.
  • Run //... from the main checkout only with .mefi in .buckconfig's project.ignore. Studio's worktrees nest there, and a worktree deleted mid-walk failed the check with exit 3.