postjevsql.git / tests / CLAUDE.md
1@README.md
2
3- Always `SET statement_timeout` in a test, so a hang fails fast.
4- Start instances only through `TestPostgres`; its PDEATHSIG is what
5  makes teardown leak-proof. Never spawn `postgres` another way. It
6  listens on a Unix socket in its own temp dir only; don't give it a TCP
7  port again (a released port-0 pick can be taken before it binds).
8- Never assert an upper bound on wall-clock time. The overseer's check
9  shares the machine with agents' buck builds, and "took under 2 × delay"
10  failed there on 2026-09-26 while passing alone. Measure the thing
11  itself at the mock instead: `peak_in_flight()` for concurrency,
12  `accepts()`/`connections()` for redials, request headers for retries,
13  the error message for which cancel fired. Lower bounds (a backoff
14  waited at least its delay) are safe and stay.
15- A failed run that names no test is buck's, not ours. Three such failures
16  had one cause: buck2 2026-08-22 is built with h2 0.4.18, whose
17  small-DATA-frame budget is a fixed 25,600 bytes. A burst of tiny gRPC
18  messages that the receiver has not read yet exhausts it, and h2 closes
19  that connection with `GOAWAY(ENHANCE_YOUR_CALM, "too_many_data_frames")`.
20  Upstream traced the event-bus and forkserver errors to it
21  (facebook/buck2 1cc3e05f85, 2026-09-11) and fixed the test-executor
22  channels in the same commit; the second one below is the same close
23  seen from a channel that cannot redial:
24  - "Buck daemon event bus encountered an error … too_many_data_frames"
25    is the client's connection to the daemon.
26  - "Failed to notify test executor of end-of-tests … Cannot reconnect
27    after connection loss" is the daemon's connection to the test
28    executor. That channel wraps the executor's pipes, so tonic's
29    reconnect has nothing to redial (`buck2_grpc/src/channel.rs`
30    `make_channel`), and the GOAWAY itself is never logged.
31  - "Forkserver is unavailable … Forkserver exited exit status: 0" on
32    every action is the forkserver's connection. `spawn_oneshot` then
33    shuts the forkserver down cleanly, and it is never respawned, so
34    every later run fails until the daemon restarts. A worktree's daemon
35    is that worktree's alone, so `buck2 kill` there is safe; the main
36    checkout's is shared (root CLAUDE.md).
37
38  The fix raises every buck-internal connection window to 64 MiB on h2
39  0.4.19, whose budget is half the window. It shipped in buck2
40  2026-09-15, which the flake takes from nix-pkgs (the estate-wide copy,
41  shared/pkgs/by-name/bu/buck2 in nixos-config) until nixpkgs#560478
42  lands; that package fails to evaluate once nixpkgs catches up. All three happened on
43  2026-09-26 on 2026-08-22 while the suite passed in full on the same
44  commit. If one recurs on 2026-09-15, read the cause in
45  `buck-out/v2/log/<trace>/command_report.json` and rerun; don't change a
46  test for it.
47- Two things that were suspected of the same failures, checked on
48  2026-09-26:
49  - The devshell rewriting `.buckconfig.local` on every entry did not
50    restart the daemon or its forkserver (same PIDs before and after), but
51    buck logged every write as a changed file. `flake.nix` now writes it
52    by rename, and only when its content changes (same mtime and inode
53    across two entries).
54  - A second `buck2 test //...` on the daemon while one was running broke
55    the later one's test executor, on 2026-08-22, with or without a config
56    rewrite between them. That fits the frame-budget cause above, since two
57    runs double the burst. `run-check.sh` still queues checks in one
58    checkout on `buck-out/run-check.lock`: Studio's overseer and a worker
59    both run it, and a shared daemon gains nothing from racing. Don't run
60    `buck2 test` beside it by hand.
61  - The lock must not reach buck. `run-check.sh` first `exec`'d buck while
62    holding it, and a daemon that run started inherited the descriptor
63    (buck closes none it inherits) and held the lock for its whole life:
64    every later check in that checkout waited forever, and a Studio agent
65    was killed at its run budget waiting (2026-09-26). buck now runs as a
66    child with fd 9 closed; reproduced both ways with a probe daemon.
67  - To wait for another checkout's run, don't loop on `pgrep -f 'buck2
68    test //'`: the pattern is in the waiting shell's own command line, so
69    it always matches itself. A worker lost 10 of its 25 minutes that way
70    and was killed at the budget right after finishing (2026-09-26). Wait on
71    that checkout's lock instead, `flock <checkout>/buck-out/run-check.lock
72    true`; `pgrep -x buck2` would match its idle daemon too. A cold `run-check.sh` in a fresh worktree
73    builds every major and took about 14 minutes on 3 cores.
74- Run `//...` from the main checkout only with `.mefi` in `.buckconfig`'s
75  `project.ignore`. Studio's worktrees nest there, and a worktree deleted
76  mid-walk failed the check with exit 3.