mirror of
https://github.com/dtzp555-max/ocp.git
synced 2026-07-27 16:05:07 +00:00
92787dccedf67757008a6dded4b50d5f7af4a000
1
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
92787dcced |
ci: manual flake hunt for #203 (Linux × Node version) (#211)
* ci: add a manual flake hunt for #203 (Linux x Node version) #203 has been seen four times, always on Linux CI, never once on macOS across 1000+ runs in two independent experiments. The documented next step is a Linux + Node 24 reproduction, and there is currently no way to run one without reddening an unrelated PR and waiting for chance. workflow_dispatch only — it never runs on push or pull_request, so it costs nothing until someone asks for it. Inputs are the two suspected variables: - `node` is a choice (24 / 22 / 26) rather than pinned, because 22-vs-24 is a comparison worth running deliberately. A previous Linux VM attempt was invalidated by Node 22's `node:sqlite` ExperimentalWarning landing on stderr, which the boot gate read as "server did not start" — ~23 of 50 runs failed spuriously. The workflow says so, so the next person interprets a 22 result against that confound instead of rediscovering it. - `concurrency` is the knob that actually reproduces, not round count: test() is fire-and-forget for async bodies, so ONE suite run already spawns 15 concurrent server.mjs children across 11 ltBoot tests. Several suites at once is what multiplies cross-process contention. Classification is anchored on the failure marker AND the test name, which is not cosmetic. ltDiag samples the first 900B of child stdout (#204) — the boot banner — so a log where only the GATE test failed still contains "Local tools: ON". Validated against real logs rather than reasoned about: 3 clean runs + 1 gate-mutation run gave unanchored 'Local tools' -> 1 hit, attributed to the WRONG category anchored -> gate=1, announce=0 (correct) totals -> total=4 clean=3 (correct) My first draft had the unanchored form with a comment claiming the failure mode was "matches every log". That was wrong — the real defect is cross-category attribution, caused by the diagnostic quoting the child's own output. Found by running the logic on real logs; the comment now states the measured reason. The job does NOT fail on a reproduction: catching the flake is the goal, and a red X would read as "the hunt is broken" rather than "the flake was caught". Results go to the step summary; logs upload as an artifact. server.mjs: unchanged (CI-only; ALIGNMENT.md requires no cli.js citation). Verified: YAML parses, all three run blocks pass `bash -n`, classify logic exercised against real suite output. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8 * ci: fix the hunt's fatal shell bug and retract a wrong root-cause story Review found two HIGH defects. Both confirmed independently before fixing. 1. The Classify step would have failed on EVERY run and produced no summary. I wrote `set -uo pipefail # NOT -e`. GitHub's default shell is `bash -e {0}`, and that line does not turn errexit OFF — it only ADDS pipefail. So any category counting zero makes grep exit 1, pipefail propagates it through the pipeline, the command substitution inherits it, and errexit kills the step. The all-clean case — the result this workflow most wants to report — dies on the FIRST counter. bash -e -c 'set -uo pipefail; x=$(echo hi | grep -c nomatch); echo REACHED' -> exit=1, "REACHED" never printed bash -e -c 'set +e -u -o pipefail; ...; echo REACHED' -> REACHED, x=0 Now `set +e -u -o pipefail`, with the reason in the file so nobody "tidies" it back. Verified by EXECUTING the extracted step under `bash -e` against real logs: exit=0, 2836 bytes of summary. My stated verification was `bash -n`, which is a pure syntax check and structurally cannot catch this. And because workflow_dispatch requires the file on the default branch, the workflow could not have been run end-to-end before merge — so nothing else would have caught it either. 2. The Node 22 story was a misattribution, and I had propagated it four places. I claimed Node 22's `node:sqlite` ExperimentalWarning was read by the boot gate as "server did not start", invalidating an earlier Linux run. The warning is real; the causal claim is false. The predicate is ltWait(() => buf.out.includes("listening on") || buf.exit != null) stdout only. `buf.err` appears in the assertion MESSAGE, never in the condition, so a stderr warning cannot fail it. Review also ran the full suite on Node 22.23.1: 462 passed, 0 failed, with the warning present in the logs. What actually produced that noise floor is something I had already measured and then failed to connect: pre-#204 fixed ports gave 246 EADDRINUSE and only 42/200 clean runs on unmodified main. The warning was merely VISIBLE in the failure text — via buf.err.slice(0,200) — and I read presence in the error message as causation. That is the same error I have been correcting in others' findings all week. The cost was not cosmetic: the input description told the next person that Node 22 was confounded, which would have made them discard a perfectly usable arm. Node 22 is now offered plainly, and the file states the correction so the wrong story does not survive in the artifact that outlives this PR. Also from review: - `ref` input (MED-1): a null result on current main is uninterpretable, because #204 may already have fixed #203. Hunting `7f15921^` is the positive control. - inputs go through `env:` (MED-2): free-text inputs were interpolated straight into the script body. GitHub documents `inputs.*` as untrusted. Added `type: number`. - #203's SIGNATURE, not its test name (MED-3): a CPU-starved runner blows the 9s ltWait and emits the same "✗ boot gate REFUSES" line. #203 is closed=true + non-zero exit + EMPTY stderr. Both counters are reported; the difference is contention. Demonstrated live — a gate-mutation run scored gate=1, sig=0. - full histogram instead of a hand-maintained category table: an enumerated list silently drops the failure nobody thought of, which on a hunt is the interesting one. Bucketed on a 72-char name prefix, NOT `sed 's/:.*//'` as suggested — test names contain colons, so that collapses every `localToolsSafetyError: <case>` into one bucket (verified). - per-run `timeout 300` and step-level timeout (MED-4), `permissions: contents: read`, a `concurrency:` group, Node 25 in the choices. - "15 concurrent server.mjs children" was wrong: 15 is TOTAL spawns; peak is 11-12 (review measured 11 x3, I measured 12) because the gate test's 3 cases and the 2 epoch boots are awaited serially. Corrected. - dropped the ephemeral-port note: after #204 ports come from ltFreePort(), so it is stale and the inference now runs backwards. server.mjs: unchanged (CI-only; ALIGNMENT.md requires no cli.js citation). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8 * ci: fix the ref the control arm depends on, and a signature that matched the opposite bug Delta review of |