Compare commits

..
Author SHA1 Message Date
taodengandClaude Opus 5 2fb1954f97 test: fix the adjacent race the first pass missed, plus four review corrections
The review's most valuable finding is that my SEARCH was incompetent, not just
inconclusive. test-features.mjs:39-58 makes `test()` fire-and-forget for async
bodies, so ~15 ltBoot integration tests run CONCURRENTLY inside one suite run.
My "150 targeted runs in isolation" therefore removed the only condition that
produces the race. Under full-suite x 4-way concurrency the reviewer measured
#199 at 6/200 pre-PR and 0/400 post-PR (Fisher p ~ 0.001).

HIGH — I fixed one race and left its twin one line away (:1092).
`Local tools: ON` is written 14 console.log calls AFTER `listening on`, so
gating on the boot marker and then asserting the announcement reads a buffer
that may hold only the first chunk. Identical shape to #199, same pipe this
time. Measured 9/200 before AND 9/200 after my first pass — my PR did not touch
it, leaving it as the suite's top flake, and it reddens with a message that
looks exactly like a real regression. Now waits for the line under assertion;
review measured the same fix at 0/200.

MED — ENOTEMPTY (4/200). child.kill("SIGKILL") kills server.mjs but not the
fake `claude` grandchildren, which can still be writing into `dir` while rmSync
walks it. New _ltRmRetry uses Node's maxRetries/retryDelay and never lets a
leaked temp dir fail a test. Applied to all 10 lt call sites.

LOW — boot-gate wait was `buf.closed` while the TUI one was
`... || buf.spawnErr`; made symmetric so a spawn failure cannot become a 9s
hang later. ltFreePort moved above its first use (it worked only via hoisting).

Two claims of mine that were wrong, corrected in the PR body:
  - "ltDiag is used in EVERY assertion message" — it is used in the two target
    tests, 7 sites. Overstatement in a PR whose whole argument is scope honesty.
  - "39330+i was shared across a 3-iteration loop" — it was not; i is 0/1/2 so
    the ports differ. The real fixed-port risk is CROSS-PROCESS, which the
    review's data confirms: 21/200 EADDRINUSE, all on 39360/39361/39364.

Linux experiment for #203, run on the Oracle VM against pre-PR main:
  #199 reproduced at 25/50 (vs ~3% on macOS) — platform-dependent stdout/stderr
  delivery, exactly the mechanism the review identified.
  #203 did NOT reproduce, but that result is NOT usable: Oracle runs Node 22,
  whose SQLite ExperimentalWarning broke ~23/50 runs with spurious "server did
  not start" failures. The noise floor was too high to conclude anything, and
  CI is Node 24. Reporting it as contaminated rather than as evidence.

Tests: 457 passed, 0 failed; 7 consecutive clean runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8
2026-07-27 09:26:06 +10:00
taodengandClaude Opus 5 7eeb2230ea test: harden and instrument the ltBoot integration tests (#203, #199)
I could NOT reproduce either flake, and this commit does not claim to have
root-caused them. Stating that plainly up front, because a "fix" for an
unreproduced fault is a placebo unless you say what it actually is.

What I tried, all negative:
  - 150 targeted runs of the #203 boot-gate scenario in isolation (25 rounds x
    3 cases, with and without 8 CPU hogs)  -> 0 failures
  - 6 consecutive full-suite runs                                -> 0 failures
  - a direct test of the exit-before-stdio-drain hypothesis (40 spawns writing
    a stderr burst then exiting)                                 -> 0 races
  - a direct test of the port-clash hypothesis: with the port already held, the
    boot gate still fires FIRST and stderr still carries the FATAL, so a clash
    cannot produce the observed empty-stderr signature

So instead of guessing a cause, this removes the mechanisms that were
individually defensible anyway, and makes the next occurrence self-diagnosing.

1. Wait for 'close', not 'exit', wherever a test reads buf.err/buf.out after
   the child terminates. 'exit' fires on process termination while its stdio
   pipes may still hold unread data; 'close' is the event that guarantees both
   are drained. I could not get this to race locally, but the ordering is
   simply correct and costs nothing — and it exactly matches #203's signature
   (non-zero exit, EMPTY stderr).

2. Stop gating on a proxy signal. The #199 test waited for "listening on" and
   then asserted a DIFFERENT line ("ignored in TUI mode") that is written
   independently — a race by construction. It now waits for the line actually
   under assertion, still bounded by the same timeout, with the boot markers
   kept in the predicate so a failed boot ends the wait immediately.

3. Every fixed port is gone. All 10 ltBoot sites and their 8 matching ltPost
   calls now use ltFreePort(). #193 already hit a fixed-port variant of this,
   and 39330+i in particular was shared across a 3-iteration loop.

4. ltBoot now tracks signal/closed/spawnErr, and an 'error' listener is
   attached — without one, a spawn error is re-thrown as an uncaught exception
   and takes down the whole runner instead of failing one test.

5. New ltDiag(buf) is used in every assertion message. The historical failure
   text was `expected a local-tools FATAL, got: ` — an empty string, which
   distinguishes nothing. Both mutations below now produce an immediately
   readable cause:

     gate neutered   -> "[non-loopback] process never closed — exit=undefined
                         signal=undefined closed=false | stderr(92B)=... |
                         stdout(1147B)=..."   (it booted; the gate is gone)
     warning renamed -> "no inert-flag warning appeared — ... stderr(295B)=
                         \"⚠ OCP_LOCAL_TOOLS=1 is SILENCED-FOR-MUTATION ...\""
                         (the text changed, and you can see the new text)

Both target tests remain mutation-proven: neutering the non-loopback branch of
localToolsSafetyError fails the boot-gate test, and renaming the TUI inert
warning fails the inert test.

Stability after the change: 4 sequential + 3 concurrent full-suite runs, all
457/0. That shows no regression was introduced; it does NOT show the flakes are
gone, since I never reproduced them. #203 and #199 should stay open until a
run passes through this instrumentation and reports what it saw.

Tests: 457 passed, 0 failed. No production code touched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8
2026-07-27 08:43:42 +10:00
3 changed files with 98 additions and 77 deletions
-16
View File
@@ -52,22 +52,6 @@ Runtime: Node.js (ESM, `.mjs` throughout). No build step. No bundler. `server.mj
---
## Testing: reaching faults inside `server.mjs`
`test-features.mjs` cannot `import` `server.mjs` (it calls `server.listen()` at top level), and that has twice led to the wrong conclusion that a class of bug is untestable. It isn't. Read this before writing "no regression test is possible here".
**There is a real live-server fixture.** `ltBoot(env, dir, nodeArgs)` (around `test-features.mjs:990`) spawns the actual `server.mjs` as a child with a **fake `claude` binary**, so integration tests cost no quota. `ltPost` / `ltPostStatus` / `ltWait` / `ltFreePort` round it out. It already covers boot gates, cache-epoch invalidation across two boots sharing one SQLite store, and system-prompt capture.
**`--stack-size` is a fault lever.** `ltBoot`'s third argument passes V8 flags to the child, which puts recursion- and argument-count-limited failures in reach at a much smaller input. `#193` needed a *synchronous throw* deep inside `spawnClaudeProcess`; `buildCliArgs` does `args.push("--allowedTools", ...ALLOWED_TOOLS)`, and under `--stack-size=200` that spread throws at ~24k elements instead of ~124k — which is what brings the trigger under Linux's `MAX_ARG_STRLEN` (131072 bytes for a single env string) so the test runs in CI rather than skipping. **No production fault hook was needed.**
Three rules that made it hold up, all learned the hard way:
- **Discover the threshold in a child under the same flags**, never in the test process — the parent's stack is not the one that matters.
- **Assert that the fault actually fired**, not just that the outcome looks right. `#193` asserts HTTP 500 *and* that the body carries `call stack size exceeded`; a control mutation (trigger neutered, bug still present) proves the test fails rather than passing vacuously.
- **Wait for the thing you are about to assert**, not a proxy for it. Waiting on `listening on` and then asserting a different line is a race (`#199`); waiting for the process to *exit* and then reading its `stderr` is another, because a terminated child's pipes may still hold unread data (`#203` — wait for the stdio to close, not for the exit).
Allocate ports with `ltFreePort()`. Fixed ports have caused at least one flake here.
## Release protocol
OCP follows the machine-readable `release_kit:` overlay in `CLAUDE.md` (Iron Rule 5.5). Before any version bump or tag push, re-read that YAML block and walk every item in `new_feature_doc_expectations` and `bootstrap_quirk_policy`. Tag push triggers `.github/workflows/release.yml`, which creates the GitHub Release automatically — do not create the release manually.
-12
View File
@@ -2923,16 +2923,6 @@ async function handleChatCompletions(req, res) {
const t0s = Date.now();
const promptCharsS = messages.reduce((a, m) => a + contentToText(m.content).length, 0);
let structuredHash = null;
// DO NOT collapse this with `dedupKey` below (#200). The two cacheHash calls take IDENTICAL
// arguments and look like obvious duplicate work — they are not interchangeable, because
// their GUARDS differ: this one additionally requires CACHE_TTL > 0. CLAUDE_CACHE_TTL
// DEFAULTS TO 0, so in the default configuration structuredHash is null while dedupKey must
// still be computed — it drives #153's single-flight stampede protection, which is
// deliberately independent of whether response caching is on. `dedupKey = structuredHash`
// would therefore silently disable stampede protection by default, in exactly the
// concurrent-AI-Task case it exists to bound. The duplicate call is the honest price of the
// asymmetry. If you do deduplicate it, compute once under the WEAKER guard and derive the
// cache lookup under the stronger one — and add a stampede test before you do.
if (CACHE_TTL > 0 && !conversationId && !hasCacheControl(messages)) {
structuredHash = cacheHash(cacheModel, messages, { keyId: req._authKeyId, temperature: parsed.temperature, max_tokens: parsed.max_tokens, top_p: parsed.top_p, structured, configEpoch: CONFIG_EPOCH });
try {
@@ -2952,8 +2942,6 @@ async function handleChatCompletions(req, res) {
// firing several AI Tasks at once) must NOT each pay N× — they share one flight. We dedup every
// one-off structured request (not stateful sessions / client-side prompt caching), independent of
// whether OCP response caching is enabled; when caching IS on, the same key gates cache read/write.
// Note the guard here is deliberately WEAKER than structuredHash's — no CACHE_TTL check. See the
// do-not-collapse comment above (#200).
const dedupKey = (!conversationId && !hasCacheControl(messages))
? cacheHash(cacheModel, messages, { keyId: req._authKeyId, temperature: parsed.temperature, max_tokens: parsed.max_tokens, top_p: parsed.top_p, structured, configEpoch: CONFIG_EPOCH })
: null;
+98 -49
View File
@@ -991,17 +991,48 @@ function ltBoot(env, dir, nodeArgs = []) {
CLAUDE_BIND: "127.0.0.1", CLAUDE_AUTH_MODE: "none", CLAUDE_CACHE_TTL: "0", CLAUDE_TIMEOUT: "4000", ...env },
stdio: ["ignore", "pipe", "pipe"],
});
const buf = { out: "", err: "", exit: undefined };
const buf = { out: "", err: "", exit: undefined, signal: undefined, closed: false, spawnErr: null };
child.stdout.on("data", d => { buf.out += d; });
child.stderr.on("data", d => { buf.err += d; });
child.on("exit", code => { buf.exit = code; });
// 'exit' fires when the process terminates, but its stdio pipes may still hold unread data —
// 'close' is the one that guarantees both are drained. A test that terminates the child and
// then asserts on buf.err/buf.out must wait for `closed`, not `exit != null`, or it can read
// an empty buffer. See ltAssertBoot below.
child.on("exit", (code, signal) => { buf.exit = code; buf.signal = signal; });
child.on("close", () => { buf.closed = true; });
// Without a listener, a spawn 'error' is re-thrown as an uncaught exception and takes down the
// whole runner instead of failing one test.
child.on("error", e => { buf.spawnErr = e; });
return { child, buf };
}
// child.kill("SIGKILL") kills server.mjs but NOT the fake `claude` grandchildren it spawned, and
// those can still be writing sp.txt / spawns.txt into `dir` while rmSync walks it — which surfaced
// as an intermittent ENOTEMPTY (4/200 in review). Node's own retry loop handles the window.
function _ltRmRetry(dir) {
try { _ltRm(dir, { recursive: true, force: true, maxRetries: 5, retryDelay: 50 }); }
catch { /* a leaked temp dir must never fail a test — the OS reaps it */ }
}
// Every ltBoot assertion failure should be self-diagnosing. The historical failure text was
// `expected a local-tools FATAL, got: ` — an empty string, which says nothing about whether the
// child never wrote, wrote to the other stream, died on a signal, or was never spawned.
function ltDiag(buf) {
return `exit=${buf.exit} signal=${buf.signal} closed=${buf.closed}` +
(buf.spawnErr ? ` spawnErr=${buf.spawnErr.code || buf.spawnErr.message}` : "") +
` | stderr(${buf.err.length}B)=${JSON.stringify(buf.err.slice(0, 200))}` +
` | stdout(${buf.out.length}B)=${JSON.stringify(buf.out.slice(-200))}`;
}
async function ltWait(cond, ms = 9000) {
const start = Date.now();
while (Date.now() - start < ms) { if (cond()) return true; await new Promise(r => setTimeout(r, 40)); }
return false;
}
async function ltFreePort() {
const srv = _ltNetServer();
await new Promise(r => srv.listen(0, "127.0.0.1", r));
const p = srv.address().port;
await new Promise(r => srv.close(r));
return p;
}
async function ltPost(port, body) {
try {
await fetch(`http://127.0.0.1:${port}/v1/chat/completions`, {
@@ -1015,29 +1046,31 @@ console.log("\nOCP_LOCAL_TOOLS integration (boot server.mjs):");
test("integration: OCP_LOCAL_TOOLS=1 → the -p spawn receives the POSITIVE wrapper (kills the no-op mutation)", async () => {
if (!LT_POSIX) return; // sh fake — skip on Windows CI
const dir = ltMkdir(); const cap = join(dir, "sp.txt"); const fake = ltFake(dir);
const { child, buf } = ltBoot({ OCP_LOCAL_TOOLS: "1", CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: "39321", SP_CAPTURE: cap }, dir);
const port = await ltFreePort();
const { child, buf } = ltBoot({ OCP_LOCAL_TOOLS: "1", CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(port), SP_CAPTURE: cap }, dir);
try {
assert.ok(await ltWait(() => buf.out.includes("listening on") || buf.exit != null), `server did not start: ${buf.err.slice(0,200)}`);
await ltPost(39321, { model: "sonnet", messages: [{ role: "user", content: "hi" }] });
await ltPost(port, { model: "sonnet", messages: [{ role: "user", content: "hi" }] });
assert.ok(await ltWait(() => _ltExists(cap)), "fake claude was spawned and captured --system-prompt");
const sp = _ltRead(cap, "utf8");
assert.ok(sp.includes(LT_POS_MARK), `expected POSITIVE wrapper in --system-prompt, got: ${sp.slice(0,90)}`);
assert.ok(!sp.includes(LT_NEG_MARK), "positive wrapper must REPLACE the negative one, not append");
} finally { child.kill("SIGKILL"); _ltRm(dir, { recursive: true, force: true }); }
} finally { child.kill("SIGKILL"); _ltRmRetry(dir); }
});
test("integration: flag OFF → the -p spawn receives the EXACT negative wrapper (default path byte-for-byte)", async () => {
if (!LT_POSIX) return;
const dir = ltMkdir(); const cap = join(dir, "sp.txt"); const fake = ltFake(dir);
const { child, buf } = ltBoot({ CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: "39322", SP_CAPTURE: cap }, dir); // OCP_LOCAL_TOOLS unset
const port = await ltFreePort();
const { child, buf } = ltBoot({ CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(port), SP_CAPTURE: cap }, dir); // OCP_LOCAL_TOOLS unset
try {
assert.ok(await ltWait(() => buf.out.includes("listening on") || buf.exit != null), `server did not start: ${buf.err.slice(0,200)}`);
await ltPost(39322, { model: "sonnet", messages: [{ role: "user", content: "hi" }] });
await ltPost(port, { model: "sonnet", messages: [{ role: "user", content: "hi" }] });
assert.ok(await ltWait(() => _ltExists(cap)), "fake claude captured --system-prompt");
const sp = _ltRead(cap, "utf8");
// No system messages + no CLAUDE_SYSTEM_PROMPT → the wrapper is passed verbatim.
assert.equal(sp, `You are accessed via the OCP HTTP proxy. You do NOT have access to any local filesystem, working directory, shell, git status, or machine environment. Do not infer or invent such information from any context you observe. Respond only based on the conversation provided.`);
} finally { child.kill("SIGKILL"); _ltRm(dir, { recursive: true, force: true }); }
} finally { child.kill("SIGKILL"); _ltRmRetry(dir); }
});
test("integration: boot gate REFUSES each unsafe config (multi / non-loopback / anon key)", async () => {
@@ -1049,25 +1082,35 @@ test("integration: boot gate REFUSES each unsafe config (multi / non-loopback /
{ label: "anon", env: { PROXY_ANONYMOUS_KEY: "pub" } },
];
try {
for (const [i, c] of cases.entries()) {
const { child, buf } = ltBoot({ OCP_LOCAL_TOOLS: "1", CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(39330 + i), ...c.env }, dir);
for (const c of cases) {
const port = await ltFreePort();
const { child, buf } = ltBoot({ OCP_LOCAL_TOOLS: "1", CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(port), ...c.env }, dir);
try {
assert.ok(await ltWait(() => buf.exit != null), `[${c.label}] expected the process to exit`);
assert.notEqual(buf.exit, 0, `[${c.label}] must exit non-zero`);
assert.ok(/FATAL[\s\S]*OCP_LOCAL_TOOLS/.test(buf.err), `[${c.label}] expected a local-tools FATAL, got: ${buf.err.slice(0,160)}`);
// Wait for `closed`, not `exit`: the assertion below reads buf.err, and stderr is only
// guaranteed drained at 'close'. This is the ordering #203 was filed for.
assert.ok(await ltWait(() => buf.closed || buf.spawnErr), `[${c.label}] process never closed — ${ltDiag(buf)}`);
assert.notEqual(buf.exit, 0, `[${c.label}] must exit non-zero — ${ltDiag(buf)}`);
assert.ok(/FATAL[\s\S]*OCP_LOCAL_TOOLS/.test(buf.err), `[${c.label}] expected a local-tools FATAL — ${ltDiag(buf)}`);
} finally { child.kill("SIGKILL"); }
}
} finally { _ltRm(dir, { recursive: true, force: true }); }
} finally { _ltRmRetry(dir); }
});
test("integration: safe single-user config BOOTS past the gate and announces local tools", async () => {
if (!LT_POSIX) return;
const dir = ltMkdir(); const fake = ltFake(dir);
const { child, buf } = ltBoot({ OCP_LOCAL_TOOLS: "1", CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: "39340" }, dir); // loopback + none
const port = await ltFreePort();
const { child, buf } = ltBoot({ OCP_LOCAL_TOOLS: "1", CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(port) }, dir); // loopback + none
try {
assert.ok(await ltWait(() => buf.out.includes("listening on")), `safe config must boot, got: ${buf.err.slice(0,200)}`);
assert.ok(buf.out.includes("Local tools: ON"), "startup must announce local tools when active");
} finally { child.kill("SIGKILL"); _ltRm(dir, { recursive: true, force: true }); }
// Same race as #199, one line over: "Local tools: ON" is written 14 console.log calls AFTER
// "listening on", so gating on the boot marker and then asserting the announcement can read a
// buffer holding only the first chunk. Wait for the line actually under assertion. Measured by
// review at 9/200 before this change and 0/200 after — it was the suite's top flake.
assert.ok(await ltWait(() => buf.out.includes("Local tools: ON") || buf.closed || buf.spawnErr),
`startup must announce local tools when active — ${ltDiag(buf)}`);
assert.ok(buf.out.includes("Local tools: ON"),
`startup must announce local tools when active — ${ltDiag(buf)}`);
} finally { child.kill("SIGKILL"); _ltRmRetry(dir); }
});
test("integration: TUI mode → flag is announced INERT (not 'ON'), boot not refused", async () => {
@@ -1075,12 +1118,22 @@ test("integration: TUI mode → flag is announced INERT (not 'ON'), boot not ref
const dir = ltMkdir(); const fake = ltFake(dir);
// Non-loopback would normally trip the local-tools gate; under TUI the flag is inert so the
// gate must NOT fire on its behalf. Use loopback here to isolate TUI's own guards from ours.
const { child, buf } = ltBoot({ OCP_LOCAL_TOOLS: "1", CLAUDE_TUI_MODE: "true", CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: "39341" }, dir);
const port = await ltFreePort();
const { child, buf } = ltBoot({ OCP_LOCAL_TOOLS: "1", CLAUDE_TUI_MODE: "true", CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(port) }, dir);
try {
assert.ok(await ltWait(() => buf.out.includes("listening on") || buf.exit != null), `did not start: ${buf.err.slice(0,200)}`);
assert.ok(!buf.out.includes("Local tools: ON"), "must NOT claim local tools are ON in TUI mode (the wrapper is unused there)");
assert.ok(/ignored in TUI mode/.test(buf.out + buf.err), "must warn that OCP_LOCAL_TOOLS is inert under TUI");
} finally { child.kill("SIGKILL"); _ltRm(dir, { recursive: true, force: true }); }
// Wait for the line actually under assertion, not for a proxy signal. "listening on" and the
// inert-flag warning are written independently, so gating on the former and then asserting
// the latter is a race — the flake #199 was filed for. Still bounded by the same timeout, and
// the boot markers are kept in the predicate so a failed boot ends the wait immediately
// rather than burning it.
const ready = await ltWait(() => /ignored in TUI mode/.test(buf.out + buf.err)
|| buf.closed || buf.spawnErr);
assert.ok(ready, `no inert-flag warning appeared — ${ltDiag(buf)}`);
assert.ok(/ignored in TUI mode/.test(buf.out + buf.err),
`must warn that OCP_LOCAL_TOOLS is inert under TUI — ${ltDiag(buf)}`);
assert.ok(!buf.out.includes("Local tools: ON"),
`must NOT claim local tools are ON in TUI mode (the wrapper is unused there) — ${ltDiag(buf)}`);
} finally { child.kill("SIGKILL"); _ltRmRetry(dir); }
});
test("integration: toggling OCP_LOCAL_TOOLS invalidates the standard response cache (epoch fold)", async () => {
@@ -1098,11 +1151,11 @@ test("integration: toggling OCP_LOCAL_TOOLS invalidates the standard response ca
} finally { child.kill("SIGKILL"); }
};
try {
const off = await bootOnce({}, 39350); // caches "OK" under epoch(negative)
const on = await bootOnce({ OCP_LOCAL_TOOLS: "1" }, 39351); // same DB, epoch(positive) → must MISS → re-spawn
const off = await bootOnce({}, await ltFreePort()); // caches "OK" under epoch(negative)
const on = await bootOnce({ OCP_LOCAL_TOOLS: "1" }, await ltFreePort()); // same DB, epoch(positive) → must MISS → re-spawn
assert.equal(off, 1, "first request (cache empty) must spawn claude");
assert.equal(on, 1, "after toggling the flag the identical request must NOT be served from the old cache (epoch differs → re-spawn)");
} finally { _ltRm(dir, { recursive: true, force: true }); }
} finally { _ltRmRetry(dir); }
});
// ── active-request counter is paired to the process lifecycle (#180 / #193) ──
@@ -1138,13 +1191,6 @@ function ltSpreadThrowCount(stackKb) {
return Number(String(_ltExecFile(process.execPath, [`--stack-size=${stackKb}`, "-e", src], { encoding: "utf8" })).trim()) || 0;
} catch { return 0; }
}
async function ltFreePort() {
const srv = _ltNetServer();
await new Promise(r => srv.listen(0, "127.0.0.1", r));
const p = srv.address().port;
await new Promise(r => srv.close(r));
return p;
}
async function ltPostStatus(port, body) {
try {
const r = await fetch(`http://127.0.0.1:${port}/v1/chat/completions`, {
@@ -1191,7 +1237,7 @@ test("integration: a synchronous pre-spawn throw must not leak stats.activeReque
const active = (await r.json()).requests.active;
assert.equal(active, 0,
`3 requests threw before their spawn; the counter must be back to 0, got ${active} (this is the #180 leak)`);
} finally { child.kill("SIGKILL"); _ltRm(dir, { recursive: true, force: true }); }
} finally { child.kill("SIGKILL"); _ltRmRetry(dir); }
});
// ── Cache keys hash the RESOLVED model, not the alias string (#194) ──────────
@@ -1219,36 +1265,38 @@ console.log("\nCache key resolves the model alias (#194):");
test("integration: an alias and its canonical target share ONE cache slot (normal path)", async () => {
if (!LT_POSIX) return;
const dir = ltMkdir(); const fake = ltFake(dir); const counter = join(dir, "spawns.txt");
const { child, buf } = ltBoot({ CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: "39360", CLAUDE_CACHE_TTL: "60000", SP_COUNTER: counter }, dir);
const port = await ltFreePort();
const { child, buf } = ltBoot({ CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(port), CLAUDE_CACHE_TTL: "60000", SP_COUNTER: counter }, dir);
try {
assert.ok(await ltWait(() => buf.out.includes("listening on")), `did not start: ${buf.err.slice(0, 200)}`);
_ltWrite(counter, "0");
const msgs = [{ role: "user", content: "alias-resolution-probe" }];
await ltPost(39360, { model: "sonnet", messages: msgs }); // miss → spawn
await ltPost(port, { model: "sonnet", messages: msgs }); // miss → spawn
await ltWait(() => (Number(_ltRead(counter, "utf8")) || 0) >= 1, 3000);
await ltPost(39360, { model: "claude-sonnet-5", messages: msgs }); // same resolved model → HIT
await ltPost(port, { model: "claude-sonnet-5", messages: msgs }); // same resolved model → HIT
await new Promise(r => setTimeout(r, 600));
assert.equal(Number(_ltRead(counter, "utf8")) || 0, 1,
"the canonical id must hit the slot the alias populated — a 2nd spawn means the key still hashes the raw alias");
} finally { child.kill("SIGKILL"); _ltRm(dir, { recursive: true, force: true }); }
} finally { child.kill("SIGKILL"); _ltRmRetry(dir); }
});
test("integration: an alias and its canonical target share ONE cache slot (STRUCTURED path)", async () => {
if (!LT_POSIX) return;
const dir = ltMkdir(); const fake = ltFakeJson(dir); const counter = join(dir, "spawns.txt");
const { child, buf } = ltBoot({ CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: "39361", CLAUDE_CACHE_TTL: "60000", SP_COUNTER: counter }, dir);
const port = await ltFreePort();
const { child, buf } = ltBoot({ CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(port), CLAUDE_CACHE_TTL: "60000", SP_COUNTER: counter }, dir);
try {
assert.ok(await ltWait(() => buf.out.includes("listening on")), `did not start: ${buf.err.slice(0, 200)}`);
_ltWrite(counter, "0");
const rf = { type: "json_schema", json_schema: { name: "probe", schema: LT_SCHEMA } };
const msgs = [{ role: "user", content: "structured-alias-probe" }];
await ltPost(39361, { model: "sonnet", messages: msgs, response_format: rf });
await ltPost(port, { model: "sonnet", messages: msgs, response_format: rf });
await ltWait(() => (Number(_ltRead(counter, "utf8")) || 0) >= 1, 4000);
await ltPost(39361, { model: "claude-sonnet-5", messages: msgs, response_format: rf });
await ltPost(port, { model: "claude-sonnet-5", messages: msgs, response_format: rf });
await new Promise(r => setTimeout(r, 600));
assert.equal(Number(_ltRead(counter, "utf8")) || 0, 1,
"structured cache key must resolve the alias too — this is the path the epoch-only fix missed");
} finally { child.kill("SIGKILL"); _ltRm(dir, { recursive: true, force: true }); }
} finally { child.kill("SIGKILL"); _ltRmRetry(dir); }
});
// MODEL_MAP is models[] + aliases + legacyAliases, so resolving covers legacyAliases for free.
@@ -1257,18 +1305,19 @@ test("integration: an alias and its canonical target share ONE cache slot (STRUC
test("integration: a legacyAlias shares ONE cache slot with its canonical target", async () => {
if (!LT_POSIX) return;
const dir = ltMkdir(); const fake = ltFake(dir); const counter = join(dir, "spawns.txt");
const { child, buf } = ltBoot({ CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: "39364", CLAUDE_CACHE_TTL: "60000", SP_COUNTER: counter }, dir);
const port = await ltFreePort();
const { child, buf } = ltBoot({ CLAUDE_BIN: fake, CLAUDE_PROXY_PORT: String(port), CLAUDE_CACHE_TTL: "60000", SP_COUNTER: counter }, dir);
try {
assert.ok(await ltWait(() => buf.out.includes("listening on")), `did not start: ${buf.err.slice(0, 200)}`);
_ltWrite(counter, "0");
const msgs = [{ role: "user", content: "legacy-alias-probe" }];
await ltPost(39364, { model: "claude-haiku-4-5", messages: msgs }); // legacyAlias
await ltPost(port, { model: "claude-haiku-4-5", messages: msgs }); // legacyAlias
await ltWait(() => (Number(_ltRead(counter, "utf8")) || 0) >= 1, 3000);
await ltPost(39364, { model: "claude-haiku-4-5-20251001", messages: msgs }); // canonical
await ltPost(port, { model: "claude-haiku-4-5-20251001", messages: msgs }); // canonical
await new Promise(r => setTimeout(r, 600));
assert.equal(Number(_ltRead(counter, "utf8")) || 0, 1,
"legacyAliases live in MODEL_MAP too — resolving must collapse them onto the canonical slot");
} finally { child.kill("SIGKILL"); _ltRm(dir, { recursive: true, force: true }); }
} finally { child.kill("SIGKILL"); _ltRmRetry(dir); }
});
test("integration: a config change invalidates the STRUCTURED cache too (closes the #177 gap)", async () => {
@@ -1287,11 +1336,11 @@ test("integration: a config change invalidates the STRUCTURED cache too (closes
} finally { child.kill("SIGKILL"); }
};
try {
const off = await bootOnce({}, 39362); // caches under epoch(negative wrapper)
const on = await bootOnce({ OCP_LOCAL_TOOLS: "1" }, 39363); // same DB, epoch differs → must re-spawn
const off = await bootOnce({}, await ltFreePort()); // caches under epoch(negative wrapper)
const on = await bootOnce({ OCP_LOCAL_TOOLS: "1" }, await ltFreePort()); // same DB, epoch differs → must re-spawn
assert.equal(off, 1, "first structured request (cache empty) must spawn claude");
assert.equal(on, 1, "structured cache must honor CONFIG_EPOCH — before #194 it omitted the epoch entirely and served the stale answer");
} finally { _ltRm(dir, { recursive: true, force: true }); }
} finally { _ltRmRetry(dir); }
});
// ── Upgrade Tests ──