Compare commits

..
Author SHA1 Message Date
taodengandClaude Fable 5 da204ba1fc chore(release): v3.21.1 — concurrency queue + spawn-token + TUI session-scope fixes
Patch release bundling three merged bug fixes (no new cli.js wire behavior,
no new endpoint/header/env var):

- #148 fix(tui): session prefix + reap/kill-server scoped per-instance by port
- #150 fix(server): serialize -p real-HOME token fallback behind a mutex +
  30s TTL keychain read cache + de-staled isolation decision (new lib/spawn-auth.mjs)
- #149 fix: semaphore honors runtime-lowered maxConcurrent, queued requests
  cancelled on client disconnect, singleflight follower retry on leader
  disconnect, exact queued accounting, quiet disconnect handling

Release-kit walk (CLAUDE.md Iron Rule 5.5): package.json version bump,
CHANGELOG.md entry, one-sentence Troubleshooting note for the genuinely
operator-visible upgrade-overlap caveat from #148. No changes to
models.json / Available Models / API Endpoints / Environment Variables
tables — verified by diffing all three merged commits (2922d68..d96da46);
none add an endpoint, header, or env var.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 23:07:58 +10:00
d96da46fa0 fix: honor runtime-lowered concurrency limit + cancel queued waiters on disconnect (#149)
* fix: honor runtime-lowered concurrency limit + cancel queued waiters on disconnect

Fixes three findings from an independent concurrency audit of the -p/stream-json
wait-queue (lib/tui/semaphore.mjs, reused by server.mjs as `claudeSemaphore`) and
its acquireClaudeSlot()/callClaudeTui() callers in server.mjs:

F1 (MEDIUM) — release() handed a freed slot straight to the next queued waiter
without re-checking `this.limit`, so a PATCH /settings maxConcurrent decrease was
silently ignored until every already-inflight task happened to finish on its own.
release() now only re-grants when post-decrement inflight is still under the
current limit, and a new setLimit() wakes queued waiters immediately when the
limit is raised instead of only on the next incidental release().

F2 (MEDIUM) — a request queued behind the concurrency limit had no link to its
HTTP connection, so a client that disconnected while still queued would still
get a claude process spawned for it once a slot freed — burning subscription
quota for a dead socket. acquire() now accepts an optional AbortSignal; server.mjs
derives one from the client's res "close" event (closeSignalFor) and passes it
into claudeSemaphore.acquire() / tuiSemaphore.acquire() while queued. On abort the
waiter is spliced out of the queue (not just flagged), so `queued` accounting
stays exact; the same "close" signal is wired into acquireClaudeSlot() (-p path,
non-streaming + streaming + singleflight-wrapped) and callClaudeTui() (TUI path).
If the response is already destroyed by the time we try to queue, we reject
immediately without ever entering the queue.

F8 (cosmetic) — acquireClaudeSlot() set `stats.queued = claudeSemaphore.queued + 1`
BEFORE calling acquire(), over-reporting /health's queued count by 1 whenever a
slot was granted immediately (the common, non-queued case). acquire() already
updates its internal queue synchronously before returning a Promise, so reading
claudeSemaphore.queued right AFTER calling it (instead of guessing "+1" before)
is exact. No /health field was added, removed, or renamed.

ALIGNMENT.md: this PR touches request-handler code (callClaude, callClaudeStreaming,
callClaudeTui, acquireClaudeSlot) but is local concurrency-control/queue-accounting
infrastructure with no cli.js wire analogue — it does not add, rename, or change any
endpoint, header, request field, or response field, and does not touch the /v1/messages
forwarding path or the OAuth bearer machinery (the two Class A surfaces this repo
governs). The /health response shape is unchanged (same field set, same nesting;
only the *value* of the pre-existing `stats.queued` field is corrected). Per
CLAUDE.md hard-requirement #1, a cli.js citation is therefore declared ABSENT:
there is no corresponding cli.js operation to cite because this is not a
cli.js-mirror (Class A) change and not a Class B endpoint-contract change either.

Tests: added 6 unit tests to test-features.mjs against the shared TuiSemaphore
(lowering the limit mid-load does not over-admit; raising the limit wakes queued
waiters up to the new headroom, FIFO; a queued waiter cancelled via AbortSignal
is spliced out and never later acquires; an already-aborted signal never touches
the queue; cancelling one of several queued waiters preserves FIFO for the rest).
238 pre-existing tests remain green; suite is now 244/244.

Verification: `node --check server.mjs && node --check lib/tui/semaphore.mjs && npm test` — 244 passed, 0 failed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: singleflight follower retry on leader disconnect + quiet disconnect handling (review M1/L1/L2)

Addresses the independent reviewer's APPROVE-WITH-CHANGES findings on PR #149:

M1 (MEDIUM, F2 regression) — when a singleflight LEADER disconnected while queued,
its RequestDisconnectedError rejected the SHARED promise, so live followers fell
into respondUpstreamError's generic branch and got a spurious 500 on a healthy
socket. Fix: keys.mjs singleflight() gains an optional follower-side `retryIf`
predicate. When a follower joins an existing flight and the shared promise rejects
with an error retryIf() accepts, the follower does NOT inherit the rejection — it
re-enters singleflight with its OWN fn (the map entry is guaranteed already deleted:
the delete-finally is attached upstream of the promise followers await), becoming
the new leader or joining a retrying sibling's fresh flight. The leader's own
rejection is never retried (it IS that client's disconnect). server.mjs passes
retryIf = (err) => err instanceof RequestDisconnectedError && !res.destroyed, so a
follower whose own client is also gone still propagates quietly. Callers without
retryIf keep byte-for-byte pre-existing share-everything semantics (pinned by the
existing failure-fan-out test).

L1 (LOW) — a disconnect-while-queued on the non-streaming paths was recorded as a
usage FAILURE row and logged as a [proxy] error: metric noise for a non-error.
Both non-streaming catch blocks now early-return on RequestDisconnectedError
without recordUsage(success:false) and without console.error — mirroring the
streaming path, which returns silently. The disconnect remains observable at info
level (concurrency_wait_cancelled, now also emitted with path:"tui" from
callClaudeTui for parity with acquireClaudeSlot's -p log).

L2 (LOW, test gap) — added a unit test for the abort-after-grant race: a waiter
granted its slot whose signal aborts afterward must see no rejection, no queue
corruption, and its slot released exactly once via the normal path (the semaphore
detaches the abort listener at grant; the onAbort idx===-1 guard is the in-dispatch
backstop).

Tests: +3 (2× M1 in the singleflight section, 1× L2 in the F2 section) — suite is
now 247/247 green.

ALIGNMENT.md: unchanged declaration — still local concurrency/dedup infrastructure
with no cli.js wire analogue; no endpoint, header, request field, or response field
added or changed; /health shape untouched. cli.js citation declared ABSENT per
CLAUDE.md hard-requirement #1 (not a Class A mirror change, not a Class B
contract change).

Verification: node --check server.mjs lib/tui/semaphore.mjs keys.mjs && npm test
— 247 passed, 0 failed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 23:04:13 +10:00
2538233059 fix(server): serialize -p real-HOME token fallback + TTL-cache keychain + de-stale isolation decision (F3/F5/F6) (#150)
* fix(server): serialize -p real-HOME token fallback + TTL-cache keychain + de-stale isolation decision (F3/F5/F6)

Three audit findings in the -p spawn-token resolution + HOME-isolation layer. All are infra/
process changes to how OCP READS and GATES an OAuth token it already holds; none touch the OAuth
wire machinery.

F3 (MEDIUM) — expiry-window fallback herds concurrent -p spawns into real HOME.
When the keychain token is within 5 min of expiry, resolveSpawnToken() returns null and every
concurrent spawn simultaneously falls back to the real HOME; each spawned claude then races a
refresh_token grant against the SAME single-use refresh token — rotating it out from under the
others and the operator's real claude (the credential-fork hazard, #112/#146 class). Fix: a
promise-chain mutex (createSerialMutex) serializes ONLY the real-HOME fallback — one such spawn at
a time. When a serialized waiter is admitted (prior holder torn down → its claude has refreshed the
keychain), it re-runs resolveSpawnToken(): a now-fresh token means it proceeds ISOLATED instead of
real-HOME, so the queue drains to the fast path. Isolated spawns never touch the mutex.

F5 (LOW-MED) — per-spawn double keychain exec on the hot path.
getOAuthCredentials() sync-exec'd `security find-generic-password` up to twice (wrong label first),
worst case 5s×2, blocking the event loop and stalling in-flight SSE streams. Fix: (a) memoize the
last-good keychain label and try it first (orderLabelsLastGoodFirst); (b) a 30s TTL cache of the
read (createTtlCache). This does NOT reintroduce the #146 forever-memoized regression: the TTL
bounds only how often we re-READ the keychain; resolveSpawnToken() still applies the 5-min expiry
gate (isTokenExpiring) to the CACHED creds on EVERY use, so a token expiring within the window is
still rejected → real-HOME fallback. Call sites stay synchronous (no async conversion).

F6 (LOW) — memoized isolation decision goes stale; /health could misreport.
getSpawnHomeMode() memoized the isolated/real-home decision forever: credentials appearing after
startup never enabled isolation; deleting ~/.ocp/spawn-home at runtime ENOENT'd every isolated
spawn until restart; during an expiry stint /health reported isolated:true while spawns ran real-
HOME. Fix: re-evaluate the decision per spawn (cheap now that F5 caches the keychain read);
ensureSpawnHome() re-verifies + re-prepares the scratch dir per isolated spawn; and /health now
reports the EFFECTIVE decision (token presence AND expiry gate). The /health field set is
UNCHANGED — no field added/removed/renamed — only the values are made truthful.

Alignment:
- Class: Not a wire/endpoint change for the spawn-token layer + Class B (B.2) for /health.
- cli.js citation: DECLARED ABSENT. cli.js does NOT perform OCP's spawn-token resolution, HOME
  isolation, keychain caching, or fallback serialization — these are proxy-internal process
  concerns with no cli.js analogue, so no Class A cli.js:NNNN citation exists or is required
  (ALIGNMENT.md Rule 2: this is proxy infra, not an invented forwarded endpoint/header/body).
- The /health change is authorized by ADR 0006 (grandfathered B.2 as of v3.16.4) and is a
  behaviour-preserving contract change: same fields, truthful values.
- OCP still NEVER performs a refresh_token grant itself — that property is preserved; the fix only
  serializes/gates reads of a token refreshed by the spawned or real claude (#112).
- No new endpoint, header, or request/response field. alignment.yml blacklist unaffected.

Tests: extracted the pure primitives to lib/spawn-auth.mjs and added 11 unit tests (mutex
serialization order + idempotent release; TTL cache freshness + null-miss; expiry gate; label
ordering; and the combined invariant that the TTL cache respects the expiry gate). node --check
clean; 249 tests pass (238 + 11).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(server): drain F3 fallback queue immediately — invalidate F5 keychain cache before the serialized re-check

Follow-up to the F3/F5/F6 fix (independent-review observation). F5's 30s keychain TTL cache could
make F3's post-refresh re-check see the stale (expiring) cached creds for up to ~30s, so a waiter
admitted right after the prior real-HOME holder's claude refreshed the keychain would needlessly
fall back to real HOME again instead of proceeding ISOLATED. Serialization safety was never at risk
(still one real-HOME spawn at a time, no double-refresh); only the drain-to-fast-path optimization
lagged.

Fix: invalidateKeychainReadCache() clears the F5 TTL cache; resolveSpawnDecision() calls it under
the fallback mutex, immediately before the re-check, so the admitted waiter reads FRESH keychain
state and drains to the isolated fast path at once. The extra keychain read happens only on the rare
real-HOME fallback path and only under the mutex (serialized, one at a time).

Alignment: unchanged from the parent commit — proxy-internal keychain/HOME-isolation process logic,
no cli.js analogue (cli.js citation DECLARED ABSENT, ALIGNMENT.md Rule 2), no endpoint/header/body,
/health shape unchanged, OCP still never performs a refresh_token grant itself (#112). 249 tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 22:50:52 +10:00
31e5a44099 fix(tui): scope session prefix + reap/kill-server to this instance's port (F7) (#148)
Audit finding F7 (LOW): lib/tui/session.mjs hardcoded SESSION_PREFIX =
"ocp-tui-" as a bare, host-wide constant. The boot-reap and periodic
idle-reap in server.mjs used it to decide which tmux sessions to
kill-session and whether to kill-server (which flushes defunct <claude>
zombies but tears down the WHOLE tmux server, including any live pane).
The coexistence guard only ever spared foreign product prefixes
(olp-tui-*); a SECOND OCP instance on the same host — e.g. a temporary
verification instance stood up alongside production, a real pattern
used during PR #144/#146 verification — was indistinguishable from
"ours" and could have its LIVE sessions reaped/kill-server'd by the
other instance's boot or periodic sweep.

Fix: scope the session-name prefix to this instance's own listen port
(the natural stable per-instance discriminator on one host — two OCP
instances cannot share a port): `ocp-tui-<port>-`. A sibling instance's
`ocp-tui-<otherPort>-*` sessions now fail the own-prefix startsWith
check and fall into the same "othersRemain" bucket as olp-tui-*,
so they are never touched and never used to justify kill-server.

lib/tui/session.mjs:
  - sessionPrefixForPort(port) replaces the bare SESSION_PREFIX export.
  - reapStaleTuiSessions({ tmux, port, includeLegacy }) now requires
    port and computes its own prefix from it.
  - runTuiTurn({ ..., port }) builds the tmux session name from
    sessionPrefixForPort(port) instead of the old bare constant.
  - LEGACY_SESSION_PREFIX / LEGACY_SESSION_NAME_RE (exact
    "ocp-tui-<8-hex>" shape, no port segment) describe the OLD
    pre-fix session-name shape, retained only for the migration below.

Legacy migration rule (chosen + reasoning): a bare-prefix legacy
session cannot be created by any post-fix OCP process, so if one is
seen it is presumed to be an orphaned zombie from THIS instance's own
PRE-fix process generation (left behind across an in-place upgrade),
not a stranger's. reapStaleTuiSessions() therefore accepts an
includeLegacy flag: server.mjs's one-time BOOT reap passes
includeLegacy: true (claims exact-legacy-shape sessions as its own,
enabling cleanup right after an upgrade); the periodic 15-min idle
sweep does NOT set it, so a lingering legacy-shaped session during
steady-state is conservatively treated as foreign and cannot trigger
kill-server on a routine tick. Residual (documented, accepted): a
genuinely-still-running PRE-fix OCP instance coexisting on the host at
the exact moment a new instance boots could have its live legacy
session reaped — the same class of residual risk the audit finding
itself accepts ("no live instance of the new version creates them");
this PR does not regress that scenario, it only removes the far more
common same-version collision that is the actual F7 finding.
LEGACY_SESSION_NAME_RE (`^ocp-tui-[0-9a-f]{8}$`) can never match the
new shape: the new shape always inserts a literal "-" between the
port digits and the 8-hex suffix, which the anchored 8-hex-only legacy
regex cannot satisfy.

server.mjs changes are local TUI session-lifecycle infrastructure
(tmux session naming, boot/periodic reap, kill-server) with no cli.js
wire analogue — verified via `strings` against the compiled claude
CLI 2.1.198 binary (this machine ships cli.js as a Mach-O binary per
ALIGNMENT.md's "OAuth token-host verification" precedent): cli.js
contains only `env.TMUX` detection (whether IT is running inside a
tmux pane) and an unrelated `--remote-control-session-name-prefix`
flag for its own remote-control feature — no session-prefix/reap/
kill-server mechanism of any kind. Per ALIGNMENT.md Rule 2, this is
declared absent: no endpoint, header, request, or response shape
changed; only OCP's own local process-lifecycle bookkeeping. No PORT
literal was hardcoded (CI port-SPOT check) — PORT is threaded through
from the existing server.mjs SPOT (lib/constants.mjs DEFAULT_PORT via
CLAUDE_PROXY_PORT).

test-features.mjs: rewrote the reaper suite's fixture session names to
the new port-scoped shape, added tests for sessionPrefixForPort(),
LEGACY_SESSION_NAME_RE's non-collision with the new shape, a sibling
same-host OCP instance being treated as foreign (F7 regression test),
and the includeLegacy boot-migration behavior (claims legacy zombies,
still spares a sibling instance's port-scoped session).

Verified: node --check server.mjs && node --check lib/tui/session.mjs
&& npm test → 243 passed, 0 failed.

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 22:43:19 +10:00
8 changed files with 689 additions and 84 deletions
+10
View File
@@ -1,5 +1,15 @@
# Changelog # Changelog
## v3.21.1 — 2026-07-07
Patch release: three bug fixes from an independent concurrency/session-lifecycle audit, each its own PR with a fresh-context reviewer (Iron Rule 10). No new `cli.js` wire behavior, no new endpoint, header, or env var; the `/health` field set is unchanged (only value truthfulness improved).
### Fixed
- **TUI session-scope / boot-reap (#148)** — `lib/tui/session.mjs`'s tmux session prefix is now scoped per-instance by listen port (`ocp-tui-<port>-`) instead of a bare host-wide `ocp-tui-` constant, so a second OCP instance on the same host (e.g. a temporary verification instance) can no longer have its live TUI sessions reaped or `kill-server`'d by another instance's boot/periodic sweep. The one-time boot reap also claims exact-shape legacy `ocp-tui-<8hex>` sessions (pre-fix naming) once, to clean up zombies left behind across an in-place upgrade.
- **`-p` spawn-token mutex + keychain caching (#150)** — the real-HOME token fallback used when the keychain token is within its 5-minute expiry window is now serialized behind a mutex, so concurrent `-p` spawns no longer race the same single-use refresh token against each other (the credential-fork hazard). Added a 30s TTL cache + last-good-label memoization for the keychain read, cutting per-spawn event-loop blocking. The isolation decision (`/health` isolated/real-home reporting) is now re-evaluated per spawn instead of memoized forever, so `/health` no longer misreports a stale decision. New module `lib/spawn-auth.mjs` extracts the pure, unit-testable primitives (mutex, TTL cache, expiry gate, label ordering).
- **Concurrency queue / disconnect handling (#149)** — the shared semaphore now honors a runtime-lowered `maxConcurrent` immediately (previously a decrease was silently ignored until in-flight tasks finished on their own) and wakes queued waiters right away when the limit is raised. Queued `-p`/TUI requests are now linked to the client's HTTP connection via `AbortSignal`; a client that disconnects while queued is spliced out of the queue instead of still spawning `claude` once a slot frees. A singleflight follower whose leader disconnected now retries instead of inheriting a spurious 500, and a queued-then-disconnected request is no longer recorded as a usage failure or logged as an error (quiet disconnect handling).
## v3.21.0 — 2026-06-25 ## v3.21.0 — 2026-06-25
Cleanup + docs release: TUI dead-code removal, docs honesty, and release prep. No new `cli.js` wire behavior; the default path (`CLAUDE_TUI_MODE` unset) is byte-for-byte unchanged. Cleanup + docs release: TUI dead-code removal, docs honesty, and release prep. No new `cli.js` wire behavior; the default path (`CLAUDE_TUI_MODE` unset) is byte-for-byte unchanged.
+4
View File
@@ -894,6 +894,10 @@ node ~/ocp/scripts/sync-openclaw.mjs
This is read-only at startup; the warning never blocks the gateway from running. This is read-only at startup; the warning never blocks the gateway from running.
### A TUI session vanished right after upgrading OCP
If you ran a pre-3.21.1 OCP instance and a post-3.21.1 instance on the same host at the same time during an upgrade, the new instance's one-time boot reap can, once, kill an old-format (`ocp-tui-<8hex>`) live TUI session belonging to the still-running old instance — restart the affected session (`ocp restart` or re-run your TUI turn) and it will come back under the new instance's port-scoped naming.
### OpenClaw shows old models after `ocp update` (v3.10→v3.11 only) ### OpenClaw shows old models after `ocp update` (v3.10→v3.11 only)
One-time bootstrap quirk for the v3.10.0 → v3.11.0 jump only — the running shell had the old `cmd_update` cached. Run once manually: One-time bootstrap quirk for the v3.10.0 → v3.11.0 jump only — the running shell had the old `cmd_update` cached. Run once manually:
+16 -2
View File
@@ -382,11 +382,25 @@ export function getCacheStats() {
// Per ADR 0005 / spec D4: in-process scope only (single Node process per host). // Per ADR 0005 / spec D4: in-process scope only (single Node process per host).
const inflightMap = new Map(); const inflightMap = new Map();
export function singleflight(hash, fn) { // `retryIf` (optional, audit finding M1): a predicate applied on the FOLLOWER path only.
// When a follower joins an existing flight and the shared promise rejects with an error for
// which retryIf(err) is true (in practice: the LEADER's client disconnected while queued —
// an error that is personal to the leader, not a verdict about the upstream), the follower
// does NOT inherit that rejection. Instead it re-enters singleflight with its OWN fn: it
// either becomes the new leader (the map entry is already deleted — see the finally below,
// which runs before any follower's catch because it is attached upstream of the promise the
// followers await) or joins a flight another retrying follower just created. The leader's
// own rejection is never retried here — its error belongs to it (leader path returns the
// bare promise). Callers that pass no retryIf get the exact pre-M1 share-everything behavior.
export function singleflight(hash, fn, retryIf) {
const existing = inflightMap.get(hash); const existing = inflightMap.get(hash);
if (existing) { if (existing) {
existing.requesters++; existing.requesters++;
return existing.promise; if (!retryIf) return existing.promise;
return existing.promise.catch((err) => {
if (!retryIf(err)) throw err;
return singleflight(hash, fn, retryIf);
});
} }
// Wrap fn() in Promise.resolve().then() so synchronous throws don't escape. // Wrap fn() in Promise.resolve().then() so synchronous throws don't escape.
const promise = Promise.resolve().then(fn).finally(() => { const promise = Promise.resolve().then(fn).finally(() => {
+61 -11
View File
@@ -20,6 +20,15 @@
// //
// Pure + importable so test-features.mjs can assert the bound directly (no server boot). // Pure + importable so test-features.mjs can assert the bound directly (no server boot).
// Thrown by acquire() when the caller-supplied AbortSignal fires before a slot was granted
// (audit finding F2 — a client that disconnects while queued must never receive a slot; the
// queue entry is spliced out, not just flagged, so `queued` accounting stays exact). Distinct
// `name` lets callers (server.mjs acquireClaudeSlot) tell "client went away" apart from
// "queue is full" without string-matching the message.
export class SemaphoreAbortError extends Error {
constructor(message) { super(message); this.name = "SemaphoreAbortError"; }
}
export class TuiSemaphore { export class TuiSemaphore {
// limit: max concurrent slots. maxQueue: max waiters before run() rejects with backpressure. // limit: max concurrent slots. maxQueue: max waiters before run() rejects with backpressure.
constructor(limit, { maxQueue } = {}) { constructor(limit, { maxQueue } = {}) {
@@ -34,9 +43,30 @@ export class TuiSemaphore {
get inflight() { return this._inflight; } get inflight() { return this._inflight; }
get queued() { return this._waiters.length; } get queued() { return this._waiters.length; }
// Runtime-adjust the concurrency limit (audit finding F1 — a PATCH /settings maxConcurrent
// change must actually take effect, not just be ignored until every currently-inflight task
// happens to finish). Lowering the limit is handled lazily by release() (see below) — it
// simply stops re-granting until inflight drains under the new, lower limit. Raising the
// limit has immediate headroom, so we wake as many queued waiters as now fit.
setLimit(limit) {
this.limit = Math.max(1, parseInt(limit, 10) || 1);
while (this._inflight < this.limit && this._waiters.length > 0) {
const next = this._waiters.shift();
this._inflight++;
next();
}
}
// Acquire a slot. Resolves once a slot is free (immediately if under the limit, otherwise // Acquire a slot. Resolves once a slot is free (immediately if under the limit, otherwise
// when an in-flight task releases). Rejects synchronously-ish if the wait queue is full. // when an in-flight task releases). Rejects synchronously-ish if the wait queue is full.
acquire() { // `signal` (optional AbortSignal, F2) lets the caller cancel a QUEUED wait — e.g. wired to
// a client's socket "close" event so a request that disconnects before a slot is granted
// is removed from the queue instead of eventually being handed a slot for a dead socket.
// If `signal` is already aborted, reject immediately without ever touching the queue.
acquire(signal) {
if (signal?.aborted) {
return Promise.reject(new SemaphoreAbortError("acquire aborted before requesting a slot"));
}
if (this._inflight < this.limit) { if (this._inflight < this.limit) {
this._inflight++; this._inflight++;
return Promise.resolve(); return Promise.resolve();
@@ -46,24 +76,44 @@ export class TuiSemaphore {
`tui_queue_full: TUI concurrency limit (${this.limit}) reached and wait queue ` + `tui_queue_full: TUI concurrency limit (${this.limit}) reached and wait queue ` +
`(${this.maxQueue}) is full`)); `(${this.maxQueue}) is full`));
} }
return new Promise((resolve) => { this._waiters.push(resolve); }); return new Promise((resolve, reject) => {
let waiter; // the FIFO entry — captured so onAbort can find + splice exactly this one
const onAbort = () => {
const idx = this._waiters.indexOf(waiter);
if (idx === -1) return; // already granted a slot (shifted out by release()/setLimit) — too late to cancel
this._waiters.splice(idx, 1); // remove, not just flag — keeps `queued` accounting exact
reject(new SemaphoreAbortError("acquire aborted while queued"));
};
waiter = () => {
signal?.removeEventListener("abort", onAbort);
resolve();
};
signal?.addEventListener("abort", onAbort, { once: true });
this._waiters.push(waiter);
});
} }
// Release a slot. If a waiter is queued, hand the slot directly to it (inflight stays // Release a slot. Always frees the caller's own slot first, then re-grants it to the next
// constant across the handoff); otherwise decrement. // waiter ONLY if the (post-decrement) inflight count is still under the current limit (F1
// fix). This is what makes a runtime-lowered limit actually bite: if the limit was lowered
// while over-subscribed, releases stop re-granting and inflight drains toward the new limit
// instead of a freed slot being handed straight back out at the old, higher occupancy.
release() { release() {
const next = this._waiters.shift(); if (this._inflight > 0) this._inflight--;
if (next) { if (this._inflight < this.limit) {
next(); // the woken waiter already "owns" the slot — inflight unchanged const next = this._waiters.shift();
} else if (this._inflight > 0) { if (next) {
this._inflight--; this._inflight++;
next();
}
} }
} }
// Run fn() under one slot. Releases in a finally so a throw (PR-A's honesty gates, // Run fn() under one slot. Releases in a finally so a throw (PR-A's honesty gates,
// wallclock truncation, paste-not-landed, tmux spawn failure) NEVER leaks a slot. // wallclock truncation, paste-not-landed, tmux spawn failure) NEVER leaks a slot.
async run(fn) { // `signal` (optional, F2) is forwarded to acquire() so a queued run() can be cancelled.
await this.acquire(); async run(fn, signal) {
await this.acquire(signal);
try { try {
return await fn(); return await fn();
} finally { } finally {
+66 -12
View File
@@ -15,14 +15,48 @@ import { tmpdir } from "node:os";
import { randomUUID } from "node:crypto"; import { randomUUID } from "node:crypto";
import { readTuiTranscript } from "./transcript.mjs"; import { readTuiTranscript } from "./transcript.mjs";
export const SESSION_PREFIX = "ocp-tui-"; // per-proxy namespace (coexistence rule) // F7 fix (audit finding, LOW): the prefix used to be a bare, host-wide constant
// ("ocp-tui-"), so a SECOND OCP instance on the same host (e.g. a temporary
// verification instance stood up alongside production — a real pattern used during
// PR #144/#146 verification) would boot-reap and potentially kill-server the OTHER
// instance's LIVE sessions: the coexistence guard below only ever spared foreign
// PRODUCT prefixes (olp-tui-*), never a second ocp-tui-* instance on a different port.
//
// Fix: scope the prefix to the instance's own listen port. The port is the natural
// stable per-instance discriminator on one host (two OCP instances cannot share a
// port), so `ocp-tui-<port>-` uniquely namespaces this instance's sessions and makes
// a same-host sibling OCP instance look exactly like a foreign product (olp-tui-*) to
// the coexistence guard — its `ocp-tui-<otherPort>-*` sessions never match our own
// prefix and are therefore never reaped/kill-server'd by us.
//
// LEGACY_SESSION_PREFIX / LEGACY_SESSION_NAME_RE describe the OLD bare-prefix shape
// (pre-this-fix), retained ONLY for the boot-time legacy-zombie migration handled in
// reapStaleTuiSessions (see comment there). No code path in this version ever CREATES
// a legacy-shaped session name again — sessionPrefixForPort() is the only session-name
// prefix constructor used going forward.
export const LEGACY_SESSION_PREFIX = "ocp-tui-";
// Exact legacy shape: LEGACY_SESSION_PREFIX + sessionId.slice(0, 8), where sessionId is
// a randomUUID() — so the suffix is always exactly 8 lowercase hex characters with NO
// further separator. The new port-scoped shape always inserts a "-" between the port
// digits and the 8-hex suffix (see sessionPrefixForPort), so this regex can never match
// a new-shape name: a new-shape suffix is `<port digits>-<8 hex>` (contains a literal
// "-"), which `[0-9a-f]{8}$` anchored immediately after the prefix cannot satisfy.
export const LEGACY_SESSION_NAME_RE = /^ocp-tui-[0-9a-f]{8}$/;
// Build this instance's own session-name prefix, scoped by its listen port so a
// second OCP instance on the same host (different port) is never mistaken for "ours".
export function sessionPrefixForPort(port) {
return `ocp-tui-${port}-`;
}
const TMUX = process.env.OCP_TUI_TMUX_BIN || "tmux"; const TMUX = process.env.OCP_TUI_TMUX_BIN || "tmux";
const defaultTmux = (args, opts = {}) => const defaultTmux = (args, opts = {}) =>
spawnSync(TMUX, args, { encoding: "utf8", ...opts }); spawnSync(TMUX, args, { encoding: "utf8", ...opts });
// Kill ONLY our own stale sessions. Scoped to SESSION_PREFIX so a co-hosted // Kill ONLY our own stale sessions. Scoped to sessionPrefixForPort(port) so a co-hosted
// OLP test instance's `olp-tui-*` sessions are never touched. // OLP test instance's `olp-tui-*` sessions — AND a co-hosted second OCP instance's
// `ocp-tui-<otherPort>-*` sessions — are never touched (F7 fix).
// //
// Defunct-reaping (PI231 incident): the pane's `claude` process is a child of the // Defunct-reaping (PI231 incident): the pane's `claude` process is a child of the
// long-lived tmux SERVER daemon, NOT of the OCP node process — `tmux new-session -d` // long-lived tmux SERVER daemon, NOT of the OCP node process — `tmux new-session -d`
@@ -36,23 +70,40 @@ const defaultTmux = (args, opts = {}) =>
// merely re-signalling — is to stop the tmux server: when the server exits, the kernel // merely re-signalling — is to stop the tmux server: when the server exits, the kernel
// reparents its surviving children to init (PID 1), which reaps them immediately. // reparents its surviving children to init (PID 1), which reaps them immediately.
// //
// So after killing our own sessions, if the server has NO sessions left of ANY prefix // `port` (required) is this instance's own listen port (server.mjs's PORT / lib/constants.mjs
// (i.e. nothing we could disrupt — no co-hosted `olp-tui-*` or other instance), we // DEFAULT_PORT resolution) — the SPOT for "which sessions are ours."
// `kill-server` to flush the defunct backlog. If ANY non-ocp session remains we leave the //
// server running (coexistence rule, ADR 0007) and let the next boot/periodic sweep retry // `includeLegacy` (default false): when true, sessions matching the exact OLD bare-prefix
// once the server is otherwise idle. // shape (LEGACY_SESSION_NAME_RE) are ALSO treated as ours for kill-session purposes. This is
export function reapStaleTuiSessions({ tmux = defaultTmux } = {}) { // the boot-time legacy migration: an operator upgrading past this fix could otherwise be left
// with orphaned bare-prefix zombie sessions from the PREVIOUS (pre-fix) process generation of
// this SAME instance, since no live instance of the new version ever creates that shape again
// — a legacy-shaped session found at boot is therefore presumed to be this instance's own
// leftover, not a stranger's. Passed true ONLY from the one-time boot-reap call site in
// server.mjs; the periodic idle-reap sweep does NOT set it, so a lingering legacy session
// during steady-state is conservatively treated as foreign (correctly blocking kill-server)
// rather than assumed to be ours on every 15-minute tick. Residual (accepted, documented):
// if a genuinely-still-running PRE-FIX OCP instance is coexisting on the same host at the
// exact moment a new instance boots, its live legacy-shaped session could be reaped — the
// same class of residual risk the audit finding itself accepts ("no live instance of the new
// version creates them"); this PR does not regress that scenario, it only removes the far
// more common same-version collision (the actual F7 finding).
export function reapStaleTuiSessions({ tmux = defaultTmux, port, includeLegacy = false } = {}) {
const r = tmux(["list-sessions", "-F", "#{session_name}"]); const r = tmux(["list-sessions", "-F", "#{session_name}"]);
if (!r || r.status !== 0) return 0; // no tmux server / no sessions if (!r || r.status !== 0) return 0; // no tmux server / no sessions
const names = String(r.stdout || "").split("\n").map((s) => s.trim()).filter(Boolean); const names = String(r.stdout || "").split("\n").map((s) => s.trim()).filter(Boolean);
const ownPrefix = sessionPrefixForPort(port);
let killed = 0; let killed = 0;
let othersRemain = false; let othersRemain = false;
for (const name of names) { for (const name of names) {
if (name.startsWith(SESSION_PREFIX)) { const isOwn = name.startsWith(ownPrefix);
const isLegacyOwn = includeLegacy && LEGACY_SESSION_NAME_RE.test(name);
if (isOwn || isLegacyOwn) {
tmux(["kill-session", "-t", name]); tmux(["kill-session", "-t", name]);
killed++; killed++;
} else { } else {
othersRemain = true; // a session we do NOT own (e.g. olp-tui-*) — never kill-server othersRemain = true; // a session we do NOT own (olp-tui-*, a sibling ocp-tui-<otherPort>-*,
// or — outside includeLegacy — a legacy-shaped name) — never kill-server
} }
} }
// Reap defunct `claude` zombies: safe ONLY when the server is now ours-only/empty. // Reap defunct `claude` zombies: safe ONLY when the server is now ours-only/empty.
@@ -358,12 +409,15 @@ export async function runTuiTurn({
home, home,
realHome, realHome,
cwd, cwd,
port,
wallclockMs = 120000, wallclockMs = 120000,
entrypointMode = "cli", entrypointMode = "cli",
tmux = defaultTmux, tmux = defaultTmux,
}) { }) {
const sessionId = randomUUID(); const sessionId = randomUUID();
const tmuxName = SESSION_PREFIX + sessionId.slice(0, 8); // Port-scoped session name (F7 fix) — see sessionPrefixForPort / reapStaleTuiSessions
// for why this instance's own listen port is the namespace discriminator.
const tmuxName = sessionPrefixForPort(port) + sessionId.slice(0, 8);
const ehome = home || process.env.HOME; // HOME claude runs under (scratch or real) const ehome = home || process.env.HOME; // HOME claude runs under (scratch or real)
const rhome = realHome || process.env.HOME; // real home (OAuth + onboarded config source) const rhome = realHome || process.env.HOME; // real home (OAuth + onboarded config source)
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "open-claude-proxy", "name": "open-claude-proxy",
"version": "3.21.0", "version": "3.21.1",
"description": "OCP (Open Claude Proxy) — use your Claude Pro/Max subscription as an OpenAI-compatible API for any IDE. Works with Cline, OpenCode, Aider, Continue.dev, OpenClaw, and more.", "description": "OCP (Open Claude Proxy) — use your Claude Pro/Max subscription as an OpenAI-compatible API for any IDE. Works with Cline, OpenCode, Aider, Continue.dev, OpenClaw, and more.",
"type": "module", "type": "module",
"bin": { "bin": {
+170 -43
View File
@@ -43,7 +43,7 @@ import { DEFAULT_PORT } from "./lib/constants.mjs";
import { isLoopbackBind } from "./lib/net.mjs"; import { isLoopbackBind } from "./lib/net.mjs";
import { runTuiTurn, reapStaleTuiSessions, resolveTuiHome } from "./lib/tui/session.mjs"; import { runTuiTurn, reapStaleTuiSessions, resolveTuiHome } from "./lib/tui/session.mjs";
import { detectTuiUpstreamError } from "./lib/tui/transcript.mjs"; import { detectTuiUpstreamError } from "./lib/tui/transcript.mjs";
import { TuiSemaphore, recordTuiEntrypoint, buildTuiHealthBlock } from "./lib/tui/semaphore.mjs"; import { TuiSemaphore, SemaphoreAbortError, recordTuiEntrypoint, buildTuiHealthBlock } from "./lib/tui/semaphore.mjs";
import { createSerialMutex, createTtlCache, isTokenExpiring, orderLabelsLastGoodFirst } from "./lib/spawn-auth.mjs"; import { createSerialMutex, createTtlCache, isTokenExpiring, orderLabelsLastGoodFirst } from "./lib/spawn-auth.mjs";
const __dirname = dirname(fileURLToPath(import.meta.url)); const __dirname = dirname(fileURLToPath(import.meta.url));
@@ -507,16 +507,61 @@ class ConcurrencyOverflowError extends Error {
constructor(message) { super(message); this.name = "ConcurrencyOverflowError"; this.httpStatus = 429; this.retryAfter = CLAUDE_QUEUE_RETRY_AFTER; } constructor(message) { super(message); this.name = "ConcurrencyOverflowError"; this.httpStatus = 429; this.retryAfter = CLAUDE_QUEUE_RETRY_AFTER; }
} }
// Tagged error for audit finding F2: the client disconnected while queued (or was already gone
// before we even tried to queue it). Distinct from ConcurrencyOverflowError so callers never send
// a response on this path — there is no socket left to write to.
class RequestDisconnectedError extends Error {
constructor(message) { super(message); this.name = "RequestDisconnectedError"; }
}
// Build an AbortSignal that fires when `res` (an http.ServerResponse) closes — i.e. the client
// disconnected. Used to cancel a QUEUED concurrency-slot wait (F2) so a client that gives up
// before a slot is granted is spliced out of the wait queue instead of eventually spawning a
// claude process for a dead socket. If `res` has already closed by the time we get here (its
// underlying stream already torn down), the signal is returned pre-aborted so acquire() rejects
// immediately without ever touching the queue — the "close already fired before we attach" case.
// `detach()` MUST be called once the wait settles (granted or rejected) to avoid a listener leak.
function closeSignalFor(res) {
const controller = new AbortController();
if (!res || typeof res.on !== "function") return { signal: controller.signal, detach() {} };
if (res.destroyed) {
controller.abort();
return { signal: controller.signal, detach() {} };
}
const onClose = () => controller.abort();
res.on("close", onClose);
return { signal: controller.signal, detach() { res.removeListener("close", onClose); } };
}
// Acquire a -p concurrency slot, queuing if all are busy (up to CLAUDE_MAX_QUEUE). Resolves to a // Acquire a -p concurrency slot, queuing if all are busy (up to CLAUDE_MAX_QUEUE). Resolves to a
// release() fn that MUST be called exactly once on every exit path (wired into ctx.cleanup()). // release() fn that MUST be called exactly once on every exit path (wired into ctx.cleanup()).
// Rejects with ConcurrencyOverflowError when the wait-queue is full. Increments stats.queued while // Rejects with ConcurrencyOverflowError when the wait-queue is full, or with
// waiting (decremented on acquire) and stats.queueRejections on overflow. // RequestDisconnectedError when `res` closes before a slot is granted (F2) — the caller must not
async function acquireClaudeSlot() { // spawn claude in that case. `res` is optional (back-compat for any caller without a live response
stats.queued = claudeSemaphore.queued + 1; // reflect this waiter before we (maybe) block // object); omitting it just means a queued wait can't be cancelled early.
//
// F8 fix: stats.queued is set from claudeSemaphore.queued AFTER calling acquire() (not before) —
// acquire() synchronously updates _inflight/_waiters before its Promise ever resolves, so reading
// .queued right after the call already reflects reality. The old code set `queued + 1` BEFORE
// calling acquire() to account for "this waiter", which over-reported by 1 whenever the slot was
// granted immediately (the common case, not a queue at all).
async function acquireClaudeSlot(res) {
const { signal, detach } = closeSignalFor(res);
const slot = claudeSemaphore.acquire(signal);
stats.queued = claudeSemaphore.queued; // accurate: acquire() already updated the queue synchronously
try { try {
await claudeSemaphore.acquire(); await slot;
} catch (e) { } catch (e) {
detach();
stats.queued = claudeSemaphore.queued; stats.queued = claudeSemaphore.queued;
if (e instanceof SemaphoreAbortError) {
// Client-driven cancellation, not backpressure — do NOT count it as a queueRejection or
// log it as concurrency_queue_full (that log/counter means "the queue itself is full").
logEvent("info", "concurrency_wait_cancelled", {
reason: "client_disconnected", inflight: claudeSemaphore.inflight, queued: claudeSemaphore.queued,
});
throw new RequestDisconnectedError("client disconnected while waiting for a concurrency slot");
}
stats.queueRejections++; stats.queueRejections++;
logEvent("warn", "concurrency_queue_full", { logEvent("warn", "concurrency_queue_full", {
limit: claudeSemaphore.limit, maxQueue: claudeSemaphore.maxQueue, limit: claudeSemaphore.limit, maxQueue: claudeSemaphore.maxQueue,
@@ -526,6 +571,7 @@ async function acquireClaudeSlot() {
`backpressure: concurrency limit (${claudeSemaphore.limit}) reached and wait queue ` + `backpressure: concurrency limit (${claudeSemaphore.limit}) reached and wait queue ` +
`(${claudeSemaphore.maxQueue}) is full — retry shortly`); `(${claudeSemaphore.maxQueue}) is full — retry shortly`);
} }
detach();
stats.queued = claudeSemaphore.queued; stats.queued = claudeSemaphore.queued;
let released = false; let released = false;
return function releaseClaudeSlot() { return function releaseClaudeSlot() {
@@ -728,7 +774,13 @@ const TUI_REAP_INTERVAL_MS = 15 * 60 * 1000;
const tuiReapInterval = TUI_MODE ? setInterval(() => { const tuiReapInterval = TUI_MODE ? setInterval(() => {
if (tuiSemaphore.inflight > 0 || tuiSemaphore.queued > 0) return; // a turn is live — defer if (tuiSemaphore.inflight > 0 || tuiSemaphore.queued > 0) return; // a turn is live — defer
try { try {
const n = reapStaleTuiSessions(); // F7 fix: scope to THIS instance's own port; a sibling ocp-tui-<otherPort>-* session
// (a second OCP instance on the same host) is treated as foreign, same as olp-tui-*.
// includeLegacy is NOT set here — see reapStaleTuiSessions' comment: the periodic sweep
// conservatively treats any lingering bare-prefix legacy session as foreign so it can
// never trigger kill-server on a steady-state tick; only the one-time boot reap below
// claims legacy-shaped zombies.
const n = reapStaleTuiSessions({ port: PORT });
if (n) logEvent("info", "tui_reaped_stale_sessions", { count: n, trigger: "periodic" }); if (n) logEvent("info", "tui_reaped_stale_sessions", { count: n, trigger: "periodic" });
} catch (e) { logEvent("error", "tui_periodic_reap_failed", { error: e.message }); } } catch (e) { logEvent("error", "tui_periodic_reap_failed", { error: e.message }); }
}, TUI_REAP_INTERVAL_MS) : null; }, TUI_REAP_INTERVAL_MS) : null;
@@ -1129,14 +1181,29 @@ function spawnClaudeProcess(model, messages, conversationId, keyName, releaseSlo
// We accumulate full text across all content_block_delta events plus the // We accumulate full text across all content_block_delta events plus the
// assistant-aggregate fallback, then resolve with the assembled string. // assistant-aggregate fallback, then resolve with the assembled string.
// Reference: OLP ADR 0009 Amendment 1 + commit 97e7d16. // Reference: OLP ADR 0009 Amendment 1 + commit 97e7d16.
async function callClaude(model, messages, conversationId, keyName) { // `res` (optional, F2) is the client's http.ServerResponse — passed through so a queued wait
// can be cancelled the moment the client disconnects, instead of spawning claude for a dead
// socket once a slot finally frees up.
async function callClaude(model, messages, conversationId, keyName, res) {
// FIX ⑥: acquire a concurrency slot first (queues up to CLAUDE_MAX_QUEUE; rejects with a // FIX ⑥: acquire a concurrency slot first (queues up to CLAUDE_MAX_QUEUE; rejects with a
// ConcurrencyOverflowError → 429 when the queue is full). The release fn is passed into the // ConcurrencyOverflowError → 429 when the queue is full, or a RequestDisconnectedError (F2)
// spawn so the idempotent cleanup() frees it on every exit path. If the spawn itself throws // if the client goes away first). The release fn is passed into the spawn so the idempotent
// synchronously (before cleanup is wired), release here so the slot never leaks. // cleanup() frees it on every exit path. If the spawn itself throws synchronously (before
const releaseSlot = await acquireClaudeSlot(); // cleanup is wired), release here so the slot never leaks.
// F3: resolve the per-spawn HOME/token decision (may serialize on the real-HOME fallback mutex). // F2×F3 composition: the slot acquire comes FIRST and is the cancellable step — a client
const spawnDecision = await resolveSpawnDecision(); // that disconnects while queued rejects here, BEFORE resolveSpawnDecision() runs, so a
// cancelled request can never acquire (or briefly hold) the real-HOME fallback mutex.
const releaseSlot = await acquireClaudeSlot(res);
// F3: resolve the per-spawn HOME/token decision (may serialize on the real-HOME fallback
// mutex). If it throws, release the just-acquired slot before propagating — cleanup() is
// not wired yet at this point.
let spawnDecision;
try {
spawnDecision = await resolveSpawnDecision();
} catch (err) {
releaseSlot();
throw err;
}
return new Promise((resolve, reject) => { return new Promise((resolve, reject) => {
let ctx; let ctx;
try { try {
@@ -1215,26 +1282,50 @@ async function callClaude(model, messages, conversationId, keyName) {
// flag that could perturb cc_entrypoint classification. // flag that could perturb cc_entrypoint classification.
// Authority: claude CLI v2.1.158 interactive mode (cc_entrypoint=cli). // Authority: claude CLI v2.1.158 interactive mode (cc_entrypoint=cli).
// SECURITY: A-path single-user ONLY — home is NOT isolation (see ADR 0007). // SECURITY: A-path single-user ONLY — home is NOT isolation (see ADR 0007).
function callClaudeTui(model, messages, _conversationId, _keyName) { // `res` (optional, F2) is the client's http.ServerResponse — see closeSignalFor.
async function callClaudeTui(model, messages, _conversationId, _keyName, res) {
const cliModel = MODEL_MAP[model] || model; const cliModel = MODEL_MAP[model] || model;
const prompt = messagesToPrompt(messages); // includes system as [System] inline const prompt = messagesToPrompt(messages); // includes system as [System] inline
recordModelRequest(cliModel, prompt.length); recordModelRequest(cliModel, prompt.length);
// C-4: gate the heavy interactive boot behind the TUI semaphore. run() acquires a slot // C-4: gate the heavy interactive boot behind the TUI semaphore (queuing if all slots are
// (queuing if all are busy, up to maxQueue), then releases in a finally so any throw from // busy, up to maxQueue). F2: `signal` (tied to `res` "close") cancels a QUEUED wait the
// runTuiTurn (tmux spawn failure, paste-not-landed) OR from the honesty gates below // instant the client disconnects, so a dead socket never triggers a cold-boot tmux+claude
// (truncation / error banner) can NEVER leak a slot. tuiSemaphore.inflight feeds /health. // spawn; detach() drops the "close" listener as soon as the wait settles rather than
return tuiSemaphore.run(() => runTuiTurn({ // holding it for the whole (up to 120s) turn.
prompt, const { signal, detach } = closeSignalFor(res);
model: cliModel, try {
claudeBin: CLAUDE, await tuiSemaphore.acquire(signal);
home: TUI_HOME, } catch (err) {
realHome: process.env.HOME, detach();
cwd: TUI_CWD, if (err instanceof SemaphoreAbortError) {
wallclockMs: TUI_WALLCLOCK_MS, // L1: client-driven cancellation, not an upstream failure — info, not error (mirrors
entrypointMode: TUI_ENTRYPOINT, // acquireClaudeSlot's concurrency_wait_cancelled on the -p path).
}).then(({ text, entrypoint, truncated }) => { logEvent("info", "concurrency_wait_cancelled", {
reason: "client_disconnected", path: "tui", inflight: tuiSemaphore.inflight, queued: tuiSemaphore.queued,
});
throw new RequestDisconnectedError("client disconnected while waiting for a TUI concurrency slot");
}
throw err;
}
detach();
// release() runs in a finally so any throw from runTuiTurn (tmux spawn failure,
// paste-not-landed) OR from the honesty gates below (truncation / error banner) can NEVER
// leak a slot. tuiSemaphore.inflight feeds /health.
try {
const { text, entrypoint, truncated } = await runTuiTurn({
prompt,
model: cliModel,
claudeBin: CLAUDE,
home: TUI_HOME,
realHome: process.env.HOME,
cwd: TUI_CWD,
port: PORT, // F7 fix: port-scopes the tmux session name so a sibling OCP instance on a
// different port never collides with this instance's reap/kill-server logic.
wallclockMs: TUI_WALLCLOCK_MS,
entrypointMode: TUI_ENTRYPOINT,
});
// ── Honesty gates (issue #133) ─ run BEFORE recordModelSuccess / cache write-back. // ── Honesty gates (issue #133) ─ run BEFORE recordModelSuccess / cache write-back.
// A throw here propagates to the .catch below (recordModelError + reject), so the // A throw here propagates to the catch below (recordModelError + reject), so the
// result never reaches the downstream setCachedResponse / singleflight / SUCCESS path. // result never reaches the downstream setCachedResponse / singleflight / SUCCESS path.
// C-2: the wall-clock cap hit with partial text and NO terminal marker — the turn // C-2: the wall-clock cap hit with partial text and NO terminal marker — the turn
@@ -1269,10 +1360,12 @@ function callClaudeTui(model, messages, _conversationId, _keyName) {
logEvent("warn", "tui_entrypoint_mismatch", { expected: "cli", got: entrypoint, model: cliModel }); logEvent("warn", "tui_entrypoint_mismatch", { expected: "cli", got: entrypoint, model: cliModel });
} }
return text; return text;
}).catch((err) => { } catch (err) {
recordModelError(cliModel, false); recordModelError(cliModel, false);
throw err; throw err;
})); } finally {
tuiSemaphore.release();
}
} }
// ── SSE heartbeat (opt-in idle watchdog) ──────────────────────────────── // ── SSE heartbeat (opt-in idle watchdog) ────────────────────────────────
@@ -1318,18 +1411,30 @@ async function callClaudeStreaming(model, messages, conversationId, res, authInf
// FIX ⑥: acquire a concurrency slot first (queues up to CLAUDE_MAX_QUEUE). On overflow, surface // FIX ⑥: acquire a concurrency slot first (queues up to CLAUDE_MAX_QUEUE). On overflow, surface
// HTTP 429 + Retry-After (NOT 500). Release is wired into cleanup() for every exit path; if the // HTTP 429 + Retry-After (NOT 500). Release is wired into cleanup() for every exit path; if the
// spawn throws synchronously before cleanup is wired, release here. // spawn throws synchronously before cleanup is wired, release here.
// F2: pass `res` so a queued wait is cancelled the instant this client disconnects — the client
// is already gone in that case, so there is no response to send back.
let releaseSlot; let releaseSlot;
try { try {
releaseSlot = await acquireClaudeSlot(); releaseSlot = await acquireClaudeSlot(res);
} catch (err) { } catch (err) {
if (err instanceof RequestDisconnectedError) return; // client gone — nothing to write to
if (err instanceof ConcurrencyOverflowError) { if (err instanceof ConcurrencyOverflowError) {
return jsonResponse(res, 429, { error: { message: sanitizeError(err.message), type: "rate_limit_error" } }, { "Retry-After": String(err.retryAfter) }); return jsonResponse(res, 429, { error: { message: sanitizeError(err.message), type: "rate_limit_error" } }, { "Retry-After": String(err.retryAfter) });
} }
return jsonResponse(res, 500, { error: { message: sanitizeError(err.message), type: "proxy_error" } }); return jsonResponse(res, 500, { error: { message: sanitizeError(err.message), type: "proxy_error" } });
} }
// F3: resolve the per-spawn HOME/token decision (may serialize on the real-HOME fallback mutex). // F3: resolve the per-spawn HOME/token decision (may serialize on the real-HOME fallback
const spawnDecision = await resolveSpawnDecision(); // mutex). F2×F3 composition: this runs strictly AFTER the (cancellable) slot acquire, so a
// request cancelled while queued never touches the fallback mutex. If it throws, release
// the just-acquired slot before responding — cleanup() is not wired yet at this point.
let spawnDecision;
try {
spawnDecision = await resolveSpawnDecision();
} catch (err) {
releaseSlot();
return jsonResponse(res, 500, { error: { message: sanitizeError(err.message), type: "proxy_error" } });
}
let ctx; let ctx;
try { try {
ctx = spawnClaudeProcess(model, messages, conversationId, authInfo.keyName, releaseSlot, spawnDecision); ctx = spawnClaudeProcess(model, messages, conversationId, authInfo.keyName, releaseSlot, spawnDecision);
@@ -2001,9 +2106,12 @@ function applySettingUpdate(key, value) {
switch (key) { switch (key) {
case "timeout": TIMEOUT = value; break; case "timeout": TIMEOUT = value; break;
// FIX ⑥: keep the -p wait-queue semaphore's limit in sync with the runtime MAX_CONCURRENT // FIX ⑥ + F1: keep the -p wait-queue semaphore's limit in sync with the runtime MAX_CONCURRENT
// so a /settings change to maxConcurrent actually changes how many claude procs run at once. // so a /settings change to maxConcurrent actually changes how many claude procs run at once
case "maxConcurrent": MAX_CONCURRENT = value; claudeSemaphore.limit = Math.max(1, value); break; // in BOTH directions. setLimit() (not a bare `.limit =` assignment) is required: lowering
// needs release() to stop over-granting until inflight drains under the new cap, and raising
// needs queued waiters woken immediately to use the new headroom. See lib/tui/semaphore.mjs.
case "maxConcurrent": MAX_CONCURRENT = value; claudeSemaphore.setLimit(value); break;
case "sessionTTL": SESSION_TTL = value; break; case "sessionTTL": SESSION_TTL = value; break;
case "maxPromptChars": MAX_PROMPT_CHARS = value; break; case "maxPromptChars": MAX_PROMPT_CHARS = value; break;
case "cacheTTL": CACHE_TTL = value; break; case "cacheTTL": CACHE_TTL = value; break;
@@ -2160,7 +2268,7 @@ async function handleChatCompletions(req, res) {
const t0TuiStream = Date.now(); const t0TuiStream = Date.now();
const promptCharsTuiStream = messages.reduce((a, m) => a + contentToText(m.content).length, 0); const promptCharsTuiStream = messages.reduce((a, m) => a + contentToText(m.content).length, 0);
try { try {
const content = await callClaudeTui(model, messages, conversationId, req._authKeyName); const content = await callClaudeTui(model, messages, conversationId, req._authKeyName, res);
if (CACHE_TTL > 0 && req._cacheHash) { if (CACHE_TTL > 0 && req._cacheHash) {
try { setCachedResponse(req._cacheHash, model, content); } catch (e) { logEvent("error", "cache_write_failed", { error: e.message }); } try { setCachedResponse(req._cacheHash, model, content); } catch (e) { logEvent("error", "cache_write_failed", { error: e.message }); }
} }
@@ -2197,15 +2305,27 @@ async function handleChatCompletions(req, res) {
// will re-read the freshly-populated cache entry here rather than spawning. // will re-read the freshly-populated cache entry here rather than spawning.
const recheck = getCachedResponse(req._cacheHash, CACHE_TTL); const recheck = getCachedResponse(req._cacheHash, CACHE_TTL);
if (recheck) return recheck.response; if (recheck) return recheck.response;
const c = await upstreamCall(model, messages, conversationId, req._authKeyName); const c = await upstreamCall(model, messages, conversationId, req._authKeyName, res);
try { setCachedResponse(req._cacheHash, model, c); } catch (e) { logEvent("error", "cache_write_failed", { error: e.message }); } try { setCachedResponse(req._cacheHash, model, c); } catch (e) { logEvent("error", "cache_write_failed", { error: e.message }); }
return c; return c;
}); },
// M1: if the LEADER disconnected while queued (F2), its RequestDisconnectedError is
// personal to the leader — a live follower must not inherit it as a spurious 500.
// retryIf makes this follower re-enter singleflight with its OWN fn (own res, own
// disconnect signal), becoming the new leader or joining a retrying sibling's flight —
// but only while OUR client is still connected. If our client is also gone, the
// rejection propagates and the RDE early-return in the catch below ends it quietly.
(err) => err instanceof RequestDisconnectedError && !res.destroyed);
const id = `chatcmpl-${randomUUID()}`; const id = `chatcmpl-${randomUUID()}`;
completionResponse(res, id, model, content); completionResponse(res, id, model, content);
try { recordUsage({ keyId: req._authKeyId, keyName: req._authKeyName, model, promptChars, responseChars: content.length, elapsedMs: Date.now() - t0Usage, success: true }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); } try { recordUsage({ keyId: req._authKeyId, keyName: req._authKeyName, model, promptChars, responseChars: content.length, elapsedMs: Date.now() - t0Usage, success: true }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); }
return; return;
} catch (err) { } catch (err) {
// L1: a client disconnect while queued is NOT an upstream failure — mirror the
// streaming path (which returns without recording anything): no usage-failure row,
// no [proxy] error log, no error response (the socket is gone). The disconnect is
// already logged at info level (concurrency_wait_cancelled) by acquireClaudeSlot.
if (err instanceof RequestDisconnectedError) { try { res.end(); } catch {} return; }
try { recordUsage({ keyId: req._authKeyId, keyName: req._authKeyName, model, promptChars, responseChars: 0, elapsedMs: Date.now() - t0Usage, success: false }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); } try { recordUsage({ keyId: req._authKeyId, keyName: req._authKeyName, model, promptChars, responseChars: 0, elapsedMs: Date.now() - t0Usage, success: false }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); }
console.error(`[proxy] error: ${err.message}`); console.error(`[proxy] error: ${err.message}`);
if (res.headersSent || res.writableEnded || res.destroyed) { if (res.headersSent || res.writableEnded || res.destroyed) {
@@ -2218,11 +2338,14 @@ async function handleChatCompletions(req, res) {
// Fallback: cache disabled (CACHE_TTL=0) or no _cacheHash — original path untouched. // Fallback: cache disabled (CACHE_TTL=0) or no _cacheHash — original path untouched.
try { try {
const content = await upstreamCall(model, messages, conversationId, req._authKeyName); const content = await upstreamCall(model, messages, conversationId, req._authKeyName, res);
const id = `chatcmpl-${randomUUID()}`; const id = `chatcmpl-${randomUUID()}`;
completionResponse(res, id, model, content); completionResponse(res, id, model, content);
try { recordUsage({ keyId: req._authKeyId, keyName: req._authKeyName, model, promptChars, responseChars: content.length, elapsedMs: Date.now() - t0Usage, success: true }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); } try { recordUsage({ keyId: req._authKeyId, keyName: req._authKeyName, model, promptChars, responseChars: content.length, elapsedMs: Date.now() - t0Usage, success: true }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); }
} catch (err) { } catch (err) {
// L1: disconnect-while-queued — same quiet non-error outcome as the singleflight
// path above and the streaming path (see acquireClaudeSlot's info-level log).
if (err instanceof RequestDisconnectedError) { try { res.end(); } catch {} return; }
try { recordUsage({ keyId: req._authKeyId, keyName: req._authKeyName, model, promptChars, responseChars: 0, elapsedMs: Date.now() - t0Usage, success: false }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); } try { recordUsage({ keyId: req._authKeyId, keyName: req._authKeyName, model, promptChars, responseChars: 0, elapsedMs: Date.now() - t0Usage, success: false }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); }
console.error(`[proxy] error: ${err.message}`); console.error(`[proxy] error: ${err.message}`);
if (res.headersSent || res.writableEnded || res.destroyed) { if (res.headersSent || res.writableEnded || res.destroyed) {
@@ -2744,7 +2867,11 @@ server.listen(PORT, BIND_ADDRESS, () => {
: "credentials.json (no CLAUDE_CODE_OAUTH_TOKEN — see Troubleshooting #401)"; : "credentials.json (no CLAUDE_CODE_OAUTH_TOKEN — see Troubleshooting #401)";
console.log(` TUI-mode: ON home=${TUI_HOME} cwd=${TUI_CWD} auth=${tuiAuth} wallclock=${TUI_WALLCLOCK_MS}ms maxConcurrent=${TUI_MAX_CONCURRENT}`); console.log(` TUI-mode: ON home=${TUI_HOME} cwd=${TUI_CWD} auth=${tuiAuth} wallclock=${TUI_WALLCLOCK_MS}ms maxConcurrent=${TUI_MAX_CONCURRENT}`);
try { try {
const n = reapStaleTuiSessions(); // F7 fix: scope to THIS instance's own port (see reapStaleTuiSessions). includeLegacy:
// true ONLY here — the one-time boot reap is the designated point to claim orphaned
// bare-prefix ("ocp-tui-<uuid8>") zombie sessions left by a PRE-fix process generation
// of this same instance (no live post-fix instance ever creates that shape again).
const n = reapStaleTuiSessions({ port: PORT, includeLegacy: true });
if (n) logEvent("info", "tui_reaped_stale_sessions", { count: n }); if (n) logEvent("info", "tui_reaped_stale_sessions", { count: n });
} catch {} } catch {}
} }
+361 -15
View File
@@ -463,6 +463,52 @@ async function runSingleflightTests() {
assert.equal(r1, 1); assert.equal(r1, 1);
assert.equal(r2, 2); assert.equal(r2, 2);
}); });
// 7. M1: leader disconnect while queued must not poison live followers. server.mjs passes
// retryIf = (err) => err instanceof RequestDisconnectedError && !res.destroyed — here we
// model that with a tagged error class. The leader (no retryIf on its own promise — the
// rejection is ITS OWN disconnect) sees the error; the live follower re-executes its OWN
// fn and gets a real result instead of a spurious inherited failure.
await asyncTest("M1: leader disconnects while queued → live follower re-executes and gets a real result", async () => {
class FakeDisconnectError extends Error {}
const leaderGate = Promise.withResolvers();
let leaderRuns = 0;
let followerRuns = 0;
const leaderFn = async () => { leaderRuns++; await leaderGate.promise; throw new FakeDisconnectError("leader client gone"); };
const followerFn = async () => { followerRuns++; return "real-execution"; };
const retryIf = (err) => err instanceof FakeDisconnectError;
const leaderP = singleflight("sf-m1-leader-dc", leaderFn); // becomes leader
const followerP = singleflight("sf-m1-leader-dc", followerFn, retryIf); // joins as follower
leaderGate.resolve(); // leader "disconnects" while holding the flight
await assert.rejects(leaderP, FakeDisconnectError, "the leader itself still sees its own disconnect");
assert.equal(await followerP, "real-execution", "follower got a REAL execution, not the leader's disconnect");
assert.equal(leaderRuns, 1, "leader fn ran once");
assert.equal(followerRuns, 1, "follower re-executed exactly once (as the new leader)");
assert.equal(getInflightStats().inflight, 0, "map fully cleaned up after the retry flight settles");
});
// 8. M1 guard: a follower whose retryIf returns false (server.mjs: its OWN client is also
// gone) inherits the rejection unchanged — no retry, no masked error. And a follower with
// NO retryIf keeps the exact pre-M1 share-everything behavior (test 2 pins the fan-out;
// this pins the predicate=false path specifically for the disconnect error).
await asyncTest("M1: follower with retryIf=false (own client also gone) inherits the leader's rejection, no retry", async () => {
class FakeDisconnectError extends Error {}
const gate = Promise.withResolvers();
let followerRuns = 0;
const leaderFn = async () => { await gate.promise; throw new FakeDisconnectError("leader client gone"); };
const followerFn = async () => { followerRuns++; return "should-never-run"; };
const leaderP = singleflight("sf-m1-both-dc", leaderFn);
const followerP = singleflight("sf-m1-both-dc", followerFn, () => false); // own client dead → no retry
gate.resolve();
await assert.rejects(leaderP, FakeDisconnectError);
await assert.rejects(followerP, FakeDisconnectError, "rejection propagates unchanged when retryIf says no");
assert.equal(followerRuns, 0, "follower fn never executed — no wasted spawn for a dead client");
assert.equal(getInflightStats().inflight, 0);
});
} }
await runSingleflightTests(); await runSingleflightTests();
@@ -1664,12 +1710,23 @@ await asyncTest("readTuiTranscript throws when no text and cap elapses", async (
}); });
// ── TUI session reaper ─────────────────────────────────────────────────── // ── TUI session reaper ───────────────────────────────────────────────────
import { reapStaleTuiSessions, SESSION_PREFIX, buildTuiCmd } from "./lib/tui/session.mjs"; import { reapStaleTuiSessions, sessionPrefixForPort, LEGACY_SESSION_PREFIX, LEGACY_SESSION_NAME_RE, buildTuiCmd } from "./lib/tui/session.mjs";
console.log("\nTUI session reaper:"); console.log("\nTUI session reaper:");
test("SESSION_PREFIX is ocp-tui-", () => { // F7 fix: the session prefix is instance-scoped by listen port so a second OCP
assert.equal(SESSION_PREFIX, "ocp-tui-"); // instance on the same host (different port) is never mistaken for "ours".
test("sessionPrefixForPort embeds the port (F7 instance scoping)", () => {
assert.equal(sessionPrefixForPort(3456), "ocp-tui-3456-");
assert.equal(sessionPrefixForPort(4000), "ocp-tui-4000-");
assert.notEqual(sessionPrefixForPort(3456), sessionPrefixForPort(4000));
});
test("LEGACY_SESSION_NAME_RE matches only the exact old bare-prefix shape, never the new shape", () => {
assert.ok(LEGACY_SESSION_NAME_RE.test(`${LEGACY_SESSION_PREFIX}a1b2c3d4`), "legacy 8-hex shape matches");
assert.ok(!LEGACY_SESSION_NAME_RE.test("ocp-tui-3456-a1b2c3d4"), "new port-scoped shape must NOT match legacy regex");
assert.ok(!LEGACY_SESSION_NAME_RE.test("ocp-tui-a1b2c3"), "too-short suffix must not match");
assert.ok(!LEGACY_SESSION_NAME_RE.test("ocp-tui-a1b2c3d4extra"), "trailing extra chars must not match");
}); });
console.log("\nTUI command construction (proxy-purity / #4):"); console.log("\nTUI command construction (proxy-purity / #4):");
@@ -1776,22 +1833,40 @@ test("buildTuiCmd OCP_TUI_FULL_TOOLS=1 grants -p-equivalent tool surface (single
} }
}); });
test("reaper kills ONLY ocp-tui- sessions, never olp-tui-", () => { test("reaper kills ONLY this instance's own port-scoped sessions, never olp-tui-", () => {
const killed = []; const killed = [];
const fakeTmux = (args) => { const fakeTmux = (args) => {
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-aaaa\nolp-tui-bbbb\nmisc\nocp-tui-cccc\n" }; if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-aaaa\nolp-tui-bbbb\nmisc\nocp-tui-3456-cccc\n" };
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; } if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" }; return { status: 0, stdout: "" };
}; };
const n = reapStaleTuiSessions({ tmux: fakeTmux }); const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.equal(n, 2); assert.equal(n, 2);
assert.equal(killed.join(","), "ocp-tui-aaaa,ocp-tui-cccc"); assert.equal(killed.join(","), "ocp-tui-3456-aaaa,ocp-tui-3456-cccc");
assert.ok(!killed.includes("olp-tui-bbbb"), "olp-tui-bbbb must never be killed"); assert.ok(!killed.includes("olp-tui-bbbb"), "olp-tui-bbbb must never be killed");
}); });
// F7 fix: a second OCP instance on the same host (different port) must be treated exactly
// like a foreign product prefix — never reaped, never allowed to trigger kill-server.
test("reaper treats a sibling OCP instance on a DIFFERENT port as foreign (F7)", () => {
const killed = [];
const calls = [];
const fakeTmux = (args) => {
calls.push(args.join(" "));
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-aaaa\nocp-tui-9999-bbbb\n" };
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" };
};
const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.equal(n, 1, "killed only the own-port session");
assert.equal(killed.join(","), "ocp-tui-3456-aaaa");
assert.ok(!killed.includes("ocp-tui-9999-bbbb"), "sibling instance's session (port 9999) must NEVER be killed");
assert.ok(!calls.includes("kill-server"), "kill-server MUST NOT fire — sibling instance's session still live");
});
test("reaper returns 0 when tmux status !== 0 (no server)", () => { test("reaper returns 0 when tmux status !== 0 (no server)", () => {
const fakeTmux = (_args) => ({ status: 1, stdout: "" }); const fakeTmux = (_args) => ({ status: 1, stdout: "" });
const n = reapStaleTuiSessions({ tmux: fakeTmux }); const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.equal(n, 0); assert.equal(n, 0);
}); });
@@ -1802,7 +1877,7 @@ test("reaper returns 0 for empty session list", () => {
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; } if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" }; return { status: 0, stdout: "" };
}; };
const n = reapStaleTuiSessions({ tmux: fakeTmux }); const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.equal(n, 0); assert.equal(n, 0);
assert.equal(killed.length, 0); assert.equal(killed.length, 0);
}); });
@@ -1815,10 +1890,10 @@ test("reaper kill-servers when the server is ours-only (flush defunct claude zom
const calls = []; const calls = [];
const fakeTmux = (args) => { const fakeTmux = (args) => {
calls.push(args.join(" ")); calls.push(args.join(" "));
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-aaaa\nocp-tui-bbbb\n" }; if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-aaaa\nocp-tui-3456-bbbb\n" };
return { status: 0, stdout: "" }; return { status: 0, stdout: "" };
}; };
const n = reapStaleTuiSessions({ tmux: fakeTmux }); const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.equal(n, 2, "killed both of our sessions"); assert.equal(n, 2, "killed both of our sessions");
assert.ok(calls.includes("kill-server"), "kill-server fired — reaps the defunct backlog"); assert.ok(calls.includes("kill-server"), "kill-server fired — reaps the defunct backlog");
}); });
@@ -1827,10 +1902,10 @@ test("reaper does NOT kill-server when a foreign (non-ocp) session remains (coex
const calls = []; const calls = [];
const fakeTmux = (args) => { const fakeTmux = (args) => {
calls.push(args.join(" ")); calls.push(args.join(" "));
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-aaaa\nolp-tui-bbbb\n" }; if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-aaaa\nolp-tui-bbbb\n" };
return { status: 0, stdout: "" }; return { status: 0, stdout: "" };
}; };
const n = reapStaleTuiSessions({ tmux: fakeTmux }); const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.equal(n, 1, "killed only our own session"); assert.equal(n, 1, "killed only our own session");
assert.ok(!calls.includes("kill-server"), "kill-server MUST NOT fire — would disrupt olp-tui-*"); assert.ok(!calls.includes("kill-server"), "kill-server MUST NOT fire — would disrupt olp-tui-*");
}); });
@@ -1838,10 +1913,61 @@ test("reaper does NOT kill-server when a foreign (non-ocp) session remains (coex
test("reaper does NOT kill-server when there is no server (status !== 0)", () => { test("reaper does NOT kill-server when there is no server (status !== 0)", () => {
const calls = []; const calls = [];
const fakeTmux = (args) => { calls.push(args.join(" ")); return { status: 1, stdout: "" }; }; const fakeTmux = (args) => { calls.push(args.join(" ")); return { status: 1, stdout: "" }; };
reapStaleTuiSessions({ tmux: fakeTmux }); reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.ok(!calls.includes("kill-server"), "no server → no kill-server (early return)"); assert.ok(!calls.includes("kill-server"), "no server → no kill-server (early return)");
}); });
// Legacy migration (F7): pre-fix versions created bare-prefix `ocp-tui-<uuid8>` sessions with
// no port segment. includeLegacy is the boot-only opt-in that claims these as our own leftover
// zombies; the periodic sweep never sets it, so a lingering legacy session cannot trigger
// kill-server on a routine 15-minute tick.
console.log("\nTUI legacy-prefix migration (boot-only reap, F7):");
test("reaper leaves legacy bare-prefix sessions untouched by default (includeLegacy unset)", () => {
const killed = [];
const calls = [];
const fakeTmux = (args) => {
calls.push(args.join(" "));
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-aaaa\nocp-tui-deadbeef\n" };
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" };
};
const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.equal(n, 1, "killed only the own-port session");
assert.ok(!killed.includes("ocp-tui-deadbeef"), "legacy session must NOT be reaped without includeLegacy");
assert.ok(!calls.includes("kill-server"), "legacy session blocks kill-server when not claimed");
});
test("reaper claims legacy bare-prefix sessions when includeLegacy=true (boot-time migration)", () => {
const killed = [];
const calls = [];
const fakeTmux = (args) => {
calls.push(args.join(" "));
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-aaaa\nocp-tui-deadbeef\n" };
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" };
};
const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456, includeLegacy: true });
assert.equal(n, 2, "both own-port and legacy sessions reaped");
assert.ok(killed.includes("ocp-tui-deadbeef"), "legacy session claimed as our own leftover");
assert.ok(calls.includes("kill-server"), "kill-server fires once no foreign/unclaimed session remains");
});
test("reaper with includeLegacy=true still spares a sibling instance's port-scoped session", () => {
const killed = [];
const calls = [];
const fakeTmux = (args) => {
calls.push(args.join(" "));
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-aaaa\nocp-tui-deadbeef\nocp-tui-9999-zzzz\n" };
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" };
};
const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456, includeLegacy: true });
assert.equal(n, 2, "own-port + legacy reaped, sibling instance untouched");
assert.ok(!killed.includes("ocp-tui-9999-zzzz"), "sibling instance session must never be claimed as legacy");
assert.ok(!calls.includes("kill-server"), "sibling instance's live session still blocks kill-server");
});
// ── TUI home preparation (scratch vs real) ─────────────────────────────── // ── TUI home preparation (scratch vs real) ───────────────────────────────
import { prepareTuiHome, ensureTuiCwdTrusted } from "./lib/tui/session.mjs"; import { prepareTuiHome, ensureTuiCwdTrusted } from "./lib/tui/session.mjs";
import { mkdtempSync as hMkdtemp, mkdirSync as hMkdir, writeFileSync as hWrite, readFileSync as hRead, existsSync as hExists, readlinkSync as hReadlink } from "node:fs"; import { mkdtempSync as hMkdtemp, mkdirSync as hMkdir, writeFileSync as hWrite, readFileSync as hRead, existsSync as hExists, readlinkSync as hReadlink } from "node:fs";
@@ -1925,7 +2051,7 @@ test("resolveTuiHome: explicit OCP_TUI_HOME wins regardless of env token (back-c
}); });
// ── TUI concurrency limiter + drift observability (PR-B: audit C-4 / C-5) ── // ── TUI concurrency limiter + drift observability (PR-B: audit C-4 / C-5) ──
import { TuiSemaphore, recordTuiEntrypoint, buildTuiHealthBlock } from "./lib/tui/semaphore.mjs"; import { TuiSemaphore, SemaphoreAbortError, recordTuiEntrypoint, buildTuiHealthBlock } from "./lib/tui/semaphore.mjs";
console.log("\nTUI concurrency limiter (C-4):"); console.log("\nTUI concurrency limiter (C-4):");
@@ -2030,6 +2156,226 @@ await asyncTest("FIX ⑥: slot released on normal completion is immediately reus
assert.equal(sem.inflight, 0); assert.equal(sem.inflight, 0);
}); });
// ── Audit F1 — runtime-lowered/raised limit must actually bite ──────────────
// server.mjs reuses this same TuiSemaphore as `claudeSemaphore`; a PATCH /settings
// maxConcurrent update now calls `claudeSemaphore.setLimit(value)` (see applySettingUpdate's
// "maxConcurrent" case). These tests pin the semaphore-level contract that fix depends on.
console.log("\nF1 — runtime concurrency-limit changes (setLimit / release honoring the current limit):");
await asyncTest("F1: lowering the limit mid-load — release() stops re-granting until inflight drains under the new limit", async () => {
const sem = new TuiSemaphore(3, { maxQueue: 16 });
const g = [deferred(), deferred(), deferred()];
const held = g.map((d) => sem.run(async () => { await d.p; }));
await new Promise((r) => setImmediate(r));
assert.equal(sem.inflight, 3, "3 tasks hold the 3 slots");
// A 4th arrives while at capacity — it queues.
const g4 = deferred();
const queued4 = sem.run(async () => { await g4.p; });
await new Promise((r) => setImmediate(r));
assert.equal(sem.queued, 1, "4th request queued");
// Operator lowers maxConcurrent from 3 to 1 while all 3 original slots are still inflight
// (mirrors a PATCH /settings maxConcurrent=1 hitting server.mjs mid-burst).
sem.setLimit(1);
assert.equal(sem.limit, 1);
// Releasing one of the 3 original holders must NOT hand the freed slot to the queued 4th
// request — before the F1 fix, release() handed slots off unconditionally, so inflight
// would have stayed pinned at the OLD higher occupancy forever.
g[0].resolve();
await held[0];
await new Promise((r) => setImmediate(r));
assert.equal(sem.inflight, 2, "inflight drains toward the new limit, not re-granted");
assert.equal(sem.queued, 1, "4th request is STILL queued — not over-admitted");
g[1].resolve();
await held[1];
await new Promise((r) => setImmediate(r));
assert.equal(sem.inflight, 1, "inflight now exactly at the new limit (1)");
assert.equal(sem.queued, 1, "still queued — inflight(1) is not < limit(1), so no grant yet");
// Releasing the LAST original holder finally drops inflight under the new limit — only
// now does the queued 4th request get granted.
g[2].resolve();
await held[2];
await new Promise((r) => setImmediate(r));
assert.equal(sem.inflight, 1, "queued 4th request now holds the single slot");
assert.equal(sem.queued, 0, "queue drained");
g4.resolve();
await queued4;
assert.equal(sem.inflight, 0);
});
await asyncTest("F1: raising the limit wakes queued waiters immediately, up to the new headroom", async () => {
const sem = new TuiSemaphore(1, { maxQueue: 16 });
const g1 = deferred();
const t1 = sem.run(async () => { await g1.p; }); // holds the only slot
await new Promise((r) => setImmediate(r));
const started = [];
const g2 = deferred(), g3 = deferred();
const t2 = sem.run(async () => { started.push(2); await g2.p; });
const t3 = sem.run(async () => { started.push(3); await g3.p; });
await new Promise((r) => setImmediate(r));
assert.equal(sem.queued, 2, "both queue behind the single holder");
assert.deepEqual(started, [], "neither queued task has started");
// Operator raises maxConcurrent from 1 to 3 (2 units of new headroom) — BOTH queued
// waiters must be woken immediately, without waiting for t1 to release.
sem.setLimit(3);
await new Promise((r) => setImmediate(r));
assert.equal(sem.inflight, 3, "t1 + both newly-woken waiters now hold slots");
assert.equal(sem.queued, 0, "queue drained by the limit raise");
assert.deepEqual(started.sort(), [2, 3], "both queued tasks started without waiting for t1's release");
g1.resolve(); g2.resolve(); g3.resolve();
await Promise.all([t1, t2, t3]);
assert.equal(sem.inflight, 0);
});
await asyncTest("F1: raising the limit wakes only as many waiters as the new headroom allows (FIFO)", async () => {
const sem = new TuiSemaphore(1, { maxQueue: 16 });
const g1 = deferred();
const t1 = sem.run(async () => { await g1.p; });
await new Promise((r) => setImmediate(r));
const started = [];
const g2 = deferred(), g3 = deferred();
const t2 = sem.run(async () => { started.push(2); await g2.p; });
const t3 = sem.run(async () => { started.push(3); await g3.p; });
await new Promise((r) => setImmediate(r));
assert.equal(sem.queued, 2);
sem.setLimit(2); // only 1 unit of new headroom (1 -> 2) — exactly one queued waiter wakes
await new Promise((r) => setImmediate(r));
assert.equal(sem.inflight, 2);
assert.equal(sem.queued, 1, "one waiter still queued — only one slot of headroom existed");
assert.deepEqual(started, [2], "FIFO: the earlier-queued waiter (t2) wakes, not t3");
// Freeing t1's slot afterward still honors the (now current) limit of 2 via release()'s
// normal path — the still-queued t3 gets in once a slot actually frees.
g1.resolve();
await t1;
await new Promise((r) => setImmediate(r));
assert.deepEqual(started, [2, 3], "t3 granted once a slot frees, honoring the raised limit");
assert.equal(sem.queued, 0);
g2.resolve(); g3.resolve();
await t2; await t3;
assert.equal(sem.inflight, 0);
});
// ── Audit F2 — queued waiters must be cancellable on client disconnect ──────
// server.mjs wires an AbortSignal derived from the client's res "close" event into
// claudeSemaphore.acquire()/tuiSemaphore.acquire() (see closeSignalFor + acquireClaudeSlot /
// callClaudeTui). These tests pin the semaphore-level cancellation contract that depends on.
console.log("\nF2 — queued-wait cancellation via AbortSignal (client disconnect while queued):");
await asyncTest("F2: aborting a QUEUED waiter rejects with SemaphoreAbortError and SPLICES it out (queued drops immediately, not just flagged)", async () => {
const sem = new TuiSemaphore(1, { maxQueue: 16 });
const g1 = deferred();
const t1 = sem.run(async () => { await g1.p; }); // holds the only slot
await new Promise((r) => setImmediate(r));
const controller = new AbortController();
const acquire2 = sem.acquire(controller.signal); // queues behind t1
await new Promise((r) => setImmediate(r));
assert.equal(sem.queued, 1, "second acquire queued");
controller.abort(); // simulates the client disconnecting while still queued
await assert.rejects(acquire2, SemaphoreAbortError, "cancelled waiter rejects with SemaphoreAbortError");
assert.equal(sem.queued, 0, "cancelled waiter is REMOVED — queue length drops immediately");
assert.equal(sem.inflight, 1, "t1's slot is untouched by the cancellation");
// Prove the cancelled waiter never later acquires a slot: free t1's slot and confirm
// nobody is waiting to receive it (the queue is genuinely empty, not just decremented).
g1.resolve();
await t1;
assert.equal(sem.inflight, 0, "slot freed with nobody queued — the cancelled waiter never got it");
});
await asyncTest("F2: an already-aborted signal rejects acquire() immediately, never touching the wait queue", async () => {
const sem = new TuiSemaphore(1, { maxQueue: 16 });
const g1 = deferred();
const t1 = sem.run(async () => { await g1.p; }); // holds the only slot
await new Promise((r) => setImmediate(r));
const controller = new AbortController();
controller.abort(); // client already gone before this request ever tries to acquire
await assert.rejects(sem.acquire(controller.signal), SemaphoreAbortError);
assert.equal(sem.queued, 0, "never entered the wait queue at all");
g1.resolve(); await t1;
});
await asyncTest("F2: cancelling one queued waiter preserves FIFO order for the others", async () => {
const sem = new TuiSemaphore(1, { maxQueue: 16 });
const g1 = deferred();
const t1 = sem.run(async () => { await g1.p; });
await new Promise((r) => setImmediate(r));
const started = [];
const cA = new AbortController();
const cB = new AbortController();
const accA = sem.acquire(cA.signal).then(() => started.push("A"));
const accB = sem.acquire(cB.signal).then(() => started.push("B"));
const g3 = deferred();
const t3 = sem.run(async () => { started.push("C"); await g3.p; });
await new Promise((r) => setImmediate(r));
assert.equal(sem.queued, 3, "A, B, C all queued behind t1");
cB.abort(); // B (the middle waiter) disconnects
await assert.rejects(accB, SemaphoreAbortError);
assert.equal(sem.queued, 2, "B removed; A and C remain, in original relative order");
g1.resolve();
await t1;
await new Promise((r) => setImmediate(r));
assert.deepEqual(started, ["A"], "A (queued first, still present) is granted next — FIFO preserved after B's removal");
assert.equal(sem.inflight, 1);
assert.equal(sem.queued, 1, "C still waiting");
sem.release(); // A was acquired directly (not via run()) — free its slot manually
await new Promise((r) => setImmediate(r));
assert.deepEqual(started, ["A", "C"], "C granted next");
g3.resolve();
await t3;
assert.equal(sem.inflight, 0);
});
await asyncTest("F2/L2: abort AFTER grant is a no-op — waiter keeps its slot, no rejection, slot released exactly once", async () => {
const sem = new TuiSemaphore(1, { maxQueue: 16 });
const g1 = deferred();
const t1 = sem.run(async () => { await g1.p; }); // holds the only slot
await new Promise((r) => setImmediate(r));
const controller = new AbortController();
let granted = false;
const acq = sem.acquire(controller.signal).then(() => { granted = true; });
await new Promise((r) => setImmediate(r));
assert.equal(sem.queued, 1, "waiter queued behind t1");
// t1 finishes → release() shifts the waiter out and grants it the slot (waiter() detaches
// the abort listener before resolving).
g1.resolve();
await t1;
await acq;
assert.equal(granted, true, "waiter was granted the slot");
assert.equal(sem.inflight, 1, "granted waiter holds the slot");
assert.equal(sem.queued, 0);
// The client disconnects AFTER the grant — the abort-after-grant race. onAbort must be a
// no-op (the waiter is no longer in _waiters; idx===-1 guard): no rejection materializes,
// the queue is untouched, and the slot is still owned by the (already-resolved) acquirer.
controller.abort();
await new Promise((r) => setImmediate(r));
assert.equal(sem.inflight, 1, "abort after grant did NOT revoke or double-free the slot");
assert.equal(sem.queued, 0, "abort after grant did not corrupt queue accounting");
// The slot is released exactly once via the normal path and is immediately reusable.
sem.release();
assert.equal(sem.inflight, 0, "slot released exactly once via the normal path");
await sem.run(async () => {}); // prove the semaphore is fully healthy afterward
assert.equal(sem.inflight, 0);
});
console.log("\nTUI drift observability (C-5):"); console.log("\nTUI drift observability (C-5):");
test("recordTuiEntrypoint: observed 'cli' is NOT a mismatch and sets lastEntrypoint", () => { test("recordTuiEntrypoint: observed 'cli' is NOT a mismatch and sets lastEntrypoint", () => {