Compare commits

...
Author SHA1 Message Date
taodengandClaude Opus 4.8 69688a46c7 docs(readme): carry the ToS caveat into the MANUAL install path too (README:218)
Reviewer found a third instance, and it is the one that most directly reproduces #136.

README has two forms of the same LAN-mode install: the copy-paste AI prompt (line 132) and
the handbook form (line 218). Line 180 explicitly asserts they are 'the same steps in handbook
form'. They were not:

  line 132 (prompt)   'install OCP as a server so YOUR OWN DEVICES on the LAN can reach it
                       (Claude Pro/Max are per-user accounts — review Anthropic's Usage Policy
                       before extending access to other people)'   + example keys: laptop, tablet
  line 218 (handbook) 'share with other devices on your network:'  + 'create API keys for each
                       PERSON/device' + no ToS pointer at all

So a reader who takes the manual route instead of the copy-paste route is still walked into
per-person key creation with zero ToS mention — reproducing #136's trigger through the door
the first commit did not close. The caveat now matches its twin.

Deliberately NOT scrubbing the wife-laptop / son-ipad example key names: they recur in seven
further places, it is a bigger diff at a different severity, and it edges into scrubbing the
maintainer's household out of their own documentation. With the pointer restored at 218 those
examples inherit the caveat — which is exactly the posture § honest limits takes. It never
forbids family sharing; it says it is the account holder's call and their risk, and refuses to
hide that.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR
2026-07-15 08:28:07 +10:00
taodengandClaude Opus 4.8 26b116deb8 docs(readme): fix the misdirected anchor + the same defect at line 52 (review fold-in)
Independent reviewer caught two things in the first cut:

1. The new link pointed at #auth-modes — which resolves to the '### Auth Modes' mode
   table (line 408), NOT to the 'Sharing with family / a team — honest limits'
   paragraph (line 422), which sits under '### Deployment model & security (read
   this)' (line 418). A bullet that says 'see the honest limits' and then sends you
   somewhere else is worse than no link. Correct anchor verified two ways (github-slugger
   + the live rendered page): #deployment-model--security-read-this (double hyphen — the
   '&' is stripped but both surrounding spaces survive).

2. Line 52 carried the SAME defect the PR was written to fix, a few lines below it:
   'share one Claude Pro/Max subscription across IDEs, devices, and people'. So the
   original claim — 'this only stops the document arguing with itself' — was not yet
   true: it stopped one instance and left an adjacent one standing, in the same
   top-of-README section an install-time reviewer reads first.

Left alone deliberately: the maintainer's own account of their household's use (lines 7
and 1178). That is theirs to make, and a docs PR should not quietly rewrite it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR
2026-07-15 08:03:10 +10:00
taodengandClaude Opus 4.8 b7f5f4d144 docs(readme): stop the feature bullet promising what § honest limits forbids (#136)
Issue #136: an external user reported that Claude Code REFUSES to run OCP's own
copy-paste install prompt, on the grounds that the premise — pooling one Pro/Max
subscription across a family — violates Anthropic's Usage Policy.

The install prompts themselves were already fixed since that report ('my own devices
on the network', plus a ToS warning on the LAN section). What remained was a
self-contradiction in the README:

  line 27  (feature bullet):  'share one Claude Pro/Max subscription with family,
                               friends, or your own devices'
  line 427 (§ honest limits): 'The defensible framing is "one person, your own
                               devices" — sharing with friends or a team is not.'

The top-of-funnel bullet was promoting exactly what the project's own ToS section
calls indefensible. That is a defect on its own terms, independent of anyone's view
on the underlying policy: a reader who trusts the bullet is walked straight into the
thing the same document later tells them not to do.

Aligned the bullet to the position the project ALREADY took, and linked the honest-limits
section from it. No change to the auth modes, the LAN feature, or the maintainer's own
account of how they use it — this only stops the doc arguing with itself.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR
2026-07-15 07:56:16 +10:00
a90f830b5d fix(tui): a null message_id on the first hook fire must not disarm the F1 guard (#160)
Residual found by the independent reviewer's second pass, by PROBING the fixed code rather
than reading it — the F1 fix was correct but its ARMING could be skipped entirely.

TuiDeltaAssembler.messageId was initialized to `null`. A first MessageDisplay payload carrying
message_id:null therefore compared EQUAL to the sentinel, registered no boundary, and left
`messages` at 0. When the real boundary then arrived, `else if (this.messages > 1)` evaluated
1 > 1 === false, so restartedAfterEmit never armed, and the `released` branch forwarded the
auth banner to the client — the exact leak F1 closed, reachable again through a single null
field. parseDeltaChunk does not validate message_id, so such a payload does reach push().

Two changes, both narrowing:
  - messageId now initializes to a Symbol sentinel, which is === to nothing a JSON payload can
    produce, so the first fire ALWAYS registers as message 1 whatever its message_id is.
  - the boundary branch drops the `this.messages > 1` sub-condition. It bought nothing and was
    the sole cause. The real invariant is "a boundary occurred while emitted !== ''" — which is
    unrecoverable regardless of how many messages have been seen — and that is now what the
    code says.

Whether claude ever emits message_id:null is unverified (the observed contract has it present),
so this is defense-in-depth, not a live bug. But the guard is the mechanism this feature
nominates as its primary safety property; it should not be disarmable by a field's absence.

Mutation-tested: restore the null sentinel + the messages>1 condition and the new test fails;
with the fix, 317 passed / 0 failed (was 316).

Class B (ADR 0007, OCP-owned TUI spawn) — cli.js does NOT perform this operation.


Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 08:21:30 +10:00
1b324968f4 feat(tui): real SSE streaming via claude's MessageDisplay hook (OCP_TUI_STREAM, default off) (#159)
* feat(tui): real SSE streaming via claude's MessageDisplay hook (OCP_TUI_STREAM, default off)

Backlog #2. TUI-mode `stream:true` turns can now emit real SSE `delta.content` chunks as
`claude` generates them, instead of buffering the turn and replaying it with
streamStringAsSSE. Opt-in: with OCP_TUI_STREAM unset/0 the spawn argv, the SSE bytes and
the cache behaviour are byte-for-byte unchanged (asserted by test).

This PR does NOT mirror any cli.js function, so no `cli.js:NNNN` citation applies, and per
CLAUDE.md's hard requirement #1 that is stated explicitly here rather than left implicit:

  - We consume claude's OWN `MessageDisplay` hook surface AS EMITTED — forwarding, not
    inventing. No new endpoint, no fabricated protocol, no new field.
  - The TUI spawn is OCP-owned surface: ADR 0007 owns it, not cli.js.
  - The SSE wire shapes are the OpenAI chat/completions streaming spec, adopted by ADR 0006.
    Every frame emitted here (role chunk, content-delta chunk, stop chunk, `[DONE]`, and the
    post-header {error:{message,type}} frame) is COPIED from callClaudeStreaming, the -p path.

/health gains additive fields only (streamEnabled + 4 counters) — same grandfathered B.2
rationale as the existing tui block (ADR 0006). Existing keys are untouched.

`claude` fires MessageDisplay per rendered block, handing the hook the RAW MARKDOWN SOURCE
of an incremental delta on stdin. The hook is registered with `--settings` on the ordinary
interactive spawn (no -p, no --bare) — verified to leave the billing pool alone.

Sink: a static sh hook script appends each payload to `<streamDir>/<session_id>.jsonl`; OCP
polls that file and forwards deltas as SSE. The per-session-id keying is MANDATORY, not an
optimization — OCP_TUI_MAX_CONCURRENT defaults to 2, so two claude panes already run at
once and a shared sink would splice one client's deltas into another's stream.

Warm-pool compatible (a separate in-flight PR depends on this): the hook script AND the
settings file are static — nothing request-specific is baked in at spawn time. The sink path
reaches the pane through its own env (OCP_TUI_STREAM_FILE) and derives from the session-id,
which a pre-booted pane fixes at boot.

The hook is SYNCHRONOUS (forceSyncExecution: claude blocks on it), so the script writes and
exits: one `cat` append, nothing else. Measured p50 7.2ms / p90 14.7ms per fire, ~50ms across
a whole turn — noise against a 6-10s turn.

It remains the terminal-turn signal, the source of the returned/cached text T, and the input
to the honesty gates. The delta stream is a low-latency MIRROR, never a replacement:

  - the truncation gate (C-2) and auth-banner gate (C-1, issue #133) run BEFORE anything is
    committed or flushed, unchanged;
  - at end of turn the streamed bytes are asserted against T. Equal -> serve. A strict PREFIX
    of T -> top up from the transcript so the client still receives exactly T (counted).
    NOT a prefix -> REFUSE the turn: SSE error frame, no cache, no success, streamDivergences++.
    Serving text the transcript disagrees with is the failure class ALIGNMENT.md exists to
    prevent, so this fails loud rather than degrading quietly;
  - only T is ever cached — never the concatenated deltas.

The auth banner needs prevention, not just detection (SSE deltas cannot be un-sent), so the
first OCP_TUI_STREAM_HOLDBACK (100) chars are withheld: the default banner detector cannot
match a message longer than 100 chars, so releasing past that provably cannot leak a banner.
A custom CLAUDE_TUI_ERROR_PATTERNS has no such bound — OCP warns at boot.

  - BANNER, before/after the spawn change: `Sonnet 4.6 with low effort · Claude Max` both,
    including on the pane the server itself spawns. Never `API Usage Billing`. Transcript
    entrypoint stays "cli". --settings is not a --bare-class flag.
  - --settings MERGES with <HOME>/.claude/settings.json rather than clobbering it (the
    user-level settings' `env` block still reached the hook), so the isolated-HOME settings
    story (permissions / additionalDirectories) survives.
  - EXACTNESS: 8/8 varied prompts (short, long, markdown, code fence, multilingual, JSON,
    table, unicode) byte-exact vs transcript T, streamed AND buffered. 0 top-ups,
    0 divergences over 15 streamed turns.
  - TTFT: buffered delivers NOTHING until the turn ends (TTFB == total, 7.5-15.8s). Streamed
    sends headers at ~25ms (heartbeat covers the pre-first-delta silence) and first content
    mid-generation, e.g. markdown 7.9s first chunk / 12.8s total; long 9.7s / 17.4s.
  - CONCURRENCY DEMUX: two concurrent streamed turns (ALPHA/BRAVO), tui.inflight peaked at 2,
    each read its own session-keyed transcript, ZERO cross-contamination.
  - AUTH-BANNER GATE under streaming, both layers: a short banner-like turn reached the client
    as 0 content chunks + an SSE error frame (never emitted); a long one was streamed but still
    ended on an error frame, not finish_reason:"stop", and was not cached.
  - DISCONNECT mid-turn: pane torn down and semaphore slot released within 1s (info-logged,
    not booked as a model error).
  - THINKING: not leaked. Opus 4.8 + xhigh turns carry a thinking block with a signature but
    `thinking:""` (the reasoning text is not persisted in interactive mode), both
    MessageDisplay text-extraction sites in the 2.1.207 bundle filter type==="text", and no
    reasoning prose appeared in any delta; concat===T held exactly on the single-message turn.
  - npm test: 282 passed, 0 failed (was 267 on main; +15).

The transcript keeps only the model's LAST assistant message. A turn where the model narrates
before calling a tool therefore has two messages, and T is only the second. If the narration
exceeds the holdback it has already been streamed and cannot be retracted -> the turn is
REFUSED. Reproduced live: Opus narrated 475 chars before a Bash call. The assembler discards
a prior message's text when nothing has been emitted yet (so short narration is handled
correctly and stays exact), and raising OCP_TUI_STREAM_HOLDBACK above the narration length
rescues the turn — verified on that exact transcript: holdback>=500 -> served, exact=true.
Documented in README and ADR 0007; this is why streaming is opt-in and off by default.

ADR 0007 line 59 ("no real token streaming — deliberate") is amended, not silently
contradicted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(tui): re-integrate streaming onto the warm-pane pool (#158) — install the hook at BOOT

Rebasing backlog #2 (streaming) onto #158 (warm pane pool) is not a textual merge: #158
split the monolithic runTuiTurn into bootTuiPane + runTuiTurn, and streaming had patched
the monolith. Re-integrating it in the OLD shape would have compiled, passed every existing
test, and been WRONG.

The bug that shape would have shipped: the sink was derived at TURN time from a streamDir
argument. But a POOLED pane is pre-booted long before any request exists — so on a pool HIT
runTuiTurn never cold-boots, no hook was ever registered on that pane, and the turn would
silently serve BUFFERED. Every miss streams, every hit does not; no error, no failing test.
The operator sees "streaming does nothing in production" and has nothing to grep for.

Fix — install the hook where the pane is born:
  - bootTuiPane({ streamDir }) registers the MessageDisplay hook at spawn and returns the
    pane's own sink (pane.streamFile), keyed by the pane's own --session-id. The hook script
    and settings file are STATIC (one pair per streamDir); the only per-turn thing is the
    sink path, and it is fixed at boot. So nothing request-specific is baked into a spawn.
  - runTuiTurn reads pane.streamFile — never recomputes it — so a warm pane and a cold pane
    stream through byte-for-byte the same path.
  - server.mjs threads the same streamDir into the pool's bootPane closure, so pre-booted
    panes carry the hook too. TUI_STREAM/TUI_STREAM_DIR now declare before the pool needs them.

Three regression guards added (test-features.mjs), and the third was MUTATION-TESTED: with
the fix reverted to the turn-time shape it fails ("the pooled pane's deltas must reach the
client"), with the fix in place it passes. A guard nobody has watched fail is not a guard.

/health: pool + stream* fields are now a union — the shape assertion asserts CONTAINMENT of
the seven grandfathered keys plus an exact added-set, so a future field that silently
REPLACED an original key cannot pass.

Class B (ADR 0007, OCP-owned TUI spawn) — cli.js does NOT perform this operation.
npm test: 313 passed, 0 failed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

* fix(tui): close the streaming auth-banner leak + 6 further review findings (PR #159)

Independent review (Iron Rule 10) found a HIGH bug by EXECUTING the code, not reading it.
All seven findings fixed. F1 and F3 were merge-blocking.

F1 (HIGH) — the auth-banner holdback was bypassed after the first release.
  TuiDeltaAssembler.released was set once and never reset at a message_id boundary, so the
  holdback + detectError predicate guarded only the FIRST message of a turn. In production's
  own configuration (OCP_TUI_FULL_TOOLS=1, where multi-message tool-using turns are the norm):
  the model narrates past the holdback before a tool call -> released; credentials expire
  mid-turn -> claude renders the 401 as ordinary assistant TEXT as a NEW message -> push()
  took the `if (this.released)` branch and handed the banner verbatim to the client. That is
  precisely the silent-error case the C-1 gate exists to prevent. Detection survived (the
  turn was still refused at finalize) but PREVENTION did not.
  Fix: once a message boundary follows an emit, the turn is already unrecoverable — finalize()
  will refuse it — so push() now emits NOTHING further for the rest of the turn.
  Second hole in the same predicate: detectTuiUpstreamError() trims before applying its
  <=100-char rule, so 101 whitespace chars trimmed to "" -> detector had nothing to classify
  -> returned null -> release fired having screened nothing. Release now gates on the TRIMMED
  length, so both sides of the check talk about the same string.

F2 — the "provably safe" claim in stream.mjs, ADR 0007 and README was unsound as written.
  Restated with both required halves: (i) nothing is emitted until the trimmed accumulation
  exceeds the detector's max banner length, AND (ii) no emission at all once a message
  boundary follows an emit. Half (i) alone only ever covered a turn's first message.

F3 (blocker) — prepareStreamHook was write-if-missing, so md-hook.sh could never be updated
  OR repaired: a host that booted once under an older version was stuck on that HOOK_SCRIPT
  forever, and a non-atomic write interrupted mid-flight left a TRUNCATED script that
  existsSync() called fine — on a hook claude BLOCKS on synchronously. Now written
  unconditionally via tmp+renameSync (the pattern already used by ensureTuiCwdTrusted).

F4 — the two spawn paths differed for non-streaming requests: the pool installed the hook
  whenever OCP_TUI_STREAM was on (correct — a pre-booted pane cannot know what request it will
  serve), but the cold path gated it on this turn's onDelta. So one stream:false request got
  --settings on a pool HIT and not on a MISS: two spawn argvs for the identical request, on
  this project's billing-classification surface. Both paths now gate on TUI_STREAM alone;
  whether the sink is POLLED remains correctly gated on onDelta.

F5 — pool._drop() killed the pane but orphaned its sink file; the reap tick drains the whole
  pool, so sinks accumulated with no GC path. Now removed best-effort on every drop path.

F6 — /health counters did not measure what they documented: streamTurns was incremented only
  AFTER the honesty gates, hiding exactly the turns an operator most wants to see (and making
  streamDivergences/streamTurns a meaningless ratio); streamDeltas counted every fire while
  claiming to count forwarded ones. Counters and docs now agree.

F7 — total hook failure was silent: zero fires per turn still yields ok:true/exact:false and
  a normal, fully-buffered answer. Only streamTopUps moved, which the code itself calls
  benign. Added streamZeroDeltaTurns (+ a tui_stream_zero_deltas warning) to separate "the
  hook is dead" from "one fire was dropped".

Tests: 316 passed, 0 failed (was 313). Every new guard MUTATION-TESTED — with each fix
reverted the guard named for it fails, and passes with the fix restored:
  - drop the restartedAfterEmit guard   -> 2 failed (incl. the strengthened old test)
  - revert trim() in the release gate   -> 1 failed
  - revert F3 to write-if-missing       -> 1 failed
The pre-existing test "new message_id AFTER an emit" asserted finalize().ok === false but
never checked what push() RETURNED — so it passed while F1 was live, documenting the leak
instead of catching it. Strengthened to assert the emission, not just the verdict.

Class B (ADR 0007, OCP-owned TUI spawn) — cli.js does NOT perform this operation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 08:19:24 +10:00
9f5bc3264a feat(tui): warm pane pool — single-use pre-booted panes, opt-in via OCP_TUI_POOL_SIZE (−41%) (#158)
* feat(tui): warm pane pool — single-use pre-booted panes, opt-in via OCP_TUI_POOL_SIZE

Backlog item #3 of docs/plans/2026-07-13-tui-latency/README.md. Every TUI request
currently cold-boots a tmux+claude pane. This adds an OPT-IN pool of pre-booted panes.
Recorded as ADR 0008 (docs/adr/0008-tui-warm-pane-pool.md), which extends ADR 0007.

MEASURED (this host, Sonnet 4.6, --effort low, through a real OCP instance; a sample
counts only if HTTP 200 AND the body carries the demanded marker):

  pool off (main code)      n= 6  p50 10.17s  [9164 9499 9760 10572 10774 11281]
  pool on, warm hits        n=12  p50  6.00s  [5286 5289 5520 5584 5621 5969
                                               6040 6098 6280 7846 8036 11053]
  pool on, warm hits (post- n= 6  p50  5.62s  [4729 4753 5236 6004 7548 9548]
    review-fix re-run)

  -> -4.17s / -41%.  12 hits / 1 miss / 0 bootFailures over 13 requests (and 6/1/0 on
  the post-fix re-run). Robust to counting the miss: n=13 p50 -> -40.6%.

The plan doc predicted only -1.0s (the boot). It is ~4.2s because the cold path also
pays ~2.9s INSIDE the first turn beyond claude's own reported turn_duration — post-
input-bar init that an idle pane has already finished. Phase decomposition of the cold
path (n=6 medians): prep 2ms | tmux spawn 27ms | boot->input-ready 1232ms | paste 8ms |
paste-verify 426ms | submit->terminal 8458ms | teardown 8ms = 10162ms total, vs native
turn_duration 5539ms => 4490ms of OCP-side overhead, of which the pool recovers ~1.26s
of boot and ~2.9s of in-claude cold start. (The 426ms paste-verify is one 400ms poll
tick; a real paste lands in ~80ms. Not addressed here — separate item.)

DESIGN
- SINGLE-USE panes. A pooled pane serves exactly ONE turn, then is killed and replaced
  in the background. Each carries its OWN fresh --session-id fixed at boot, so one
  session still holds one exchange. This is what keeps transcript.mjs's
  extractLatestAssistantText correct; its warning about a future warm pool reusing a
  session is answered in-place (comment updated) and left standing for anyone who later
  wants a second turn on a pane — that would be a cross-request TEXT LEAK and needs
  user-line scoping in the transcript reader first.
- Pool keyed by model; --model is fixed at spawn. A miss falls back to the cold path
  with zero behaviour change. The pool warms the most recently requested model, so the
  first request after start (and after a model switch) is always a cold miss.
- REAPER COEXISTENCE (the crux). An idle warm pane IS ours, and the periodic sweep runs
  precisely when we are idle. reapStaleTuiSessions() takes a `spare` set of EXACT live
  session names, and server.mjs DRAINS the pool immediately before the sweep:
    1. a live pooled pane is never reaped — INCLUDING one still BOOTING (see below);
    2. an orphaned pooled pane IS still reaped — membership is by exact name from a live
       in-memory registry, never by name shape, so a pane from a dead process generation
       has nothing claiming it. Omitting `spare` reaps MORE, never less (fail-safe);
    3. kill-server is suppressed while any pane is spared — hence the drain, so the sweep
       still flushes <defunct> claude zombies (the only mechanism that can).
- THE POOL TRACKS ITS IN-FLIGHT BOOT BY NAME, NOT AS A COUNT. bootTuiPane creates the
  tmux session SYNCHRONOUSLY and only then waits up to POOL_BOOT_MS (20s) for the input
  bar, so a pooled session can be LIVE for ~20s before its boot resolves. Tracking boots
  as a count meant the pool could not name that session, which caused two real bugs
  (found in review, reproduced, fixed, and now regression-tested):
    * the reap sweep KILLED the booting pane (it could not be spared), left the pool
      empty with nothing scheduled, and logged the exact tui_pool_boot_failed WARN
      operators are told to alert on — for a completely healthy drain;
    * graceful shutdown ORPHANED a live authenticated idle `claude`: gracefulShutdown
      calls process.exit(0) in the SAME TICK as the drain (TUI panes are tmux children,
      so activeProcesses is empty and the wait-for-children path exits immediately), so
      cleanup deferred to a .then() never ran.
  Fix: the pool mints each pane's identity up front ({sessionId, name}) and holds it in
  _bootingPane. liveNames() includes it; drain() kills it SYNCHRONOUSLY. A generation
  counter distinguishes "cancelled by us" from "genuinely failed", so a drain never
  inflates bootFailures and resume() reliably starts a fresh boot. Deriving the name from
  the session-id also makes `tmux ls` correlate to the transcript file.
- SLOT ACCOUNTING. Refill boots take NO TuiSemaphore slot (those bound real turns and
  would be starved); they cannot leak one either, since they never hold one. Refills are
  SERIALIZED, one boot at a time — live at size=2, two cold boots racing an in-flight
  turn overran the readiness cap and a refill was discarded. A genuinely failed boot does
  not re-kick the chain (backoff; a broken claude must not respawn forever). Background
  boots get a more generous readiness cap (POOL_BOOT_MS = 5x BOOT_MS): BOOT_MS is tight
  because a client is blocked on it, which is not true of a pre-boot.
- BOUNDED COST. A warm pane is a LIVE idle claude process held whether or not a request
  arrives. Peak processes = pool size + OCP_TUI_MAX_CONCURRENT + 1 booting replacement.
  Size clamped to POOL_MAX_SIZE=4; garbage values disable rather than guess. Panes have
  a 10-min TTL and a health check at hand-out (dead/degraded pane => miss, never a hang).
  Missing collaborators throw at CONSTRUCTION, not on a live request (refill() is called
  synchronously from the request path).

DEFAULT OFF (OCP_TUI_POOL_SIZE=0). This is a stable production path and the pool holds
standing processes, so the operator opts in. With the pool off, runTuiTurn takes the
IDENTICAL code path as before (the `pool ? pool.acquire() : null` branch yields null, and
tuiPool is null so no observer is attached and no new log line is emitted) — that is what
establishes the default path is unchanged. A pool-off control run (n=6, p50 9.40s) is
consistent with the 10.17s baseline but had 2/6 samples >12s, so it is corroboration, NOT
proof: n=6 cannot establish "unregressed" on its own. The code-path equivalence can.

BANNER: NO SPAWN ARGUMENT CHANGED. buildTuiCmd is byte-identical to main (verified by
extracting the function body from both revisions and comparing). Live banner captured
from two real POOLED panes anyway: "Sonnet 4.6 with low effort · Claude Max" — the
subscription pool, never "API Usage Billing".

/health: `tui.pool` added (null when off), incl. `cancelled` (boots WE killed — not a
fault; do not alert on it). The tui block is ADR-0007-owned and post-dates ADR 0006's
v3.16.4 grandfather snapshot; the addition is purely additive — every pre-existing key
keeps a byte-identical value. Authorization recorded in ADR 0008.

ALIGNMENT: Class B / ADR 0007 + ADR 0008 (OCP-owned TUI spawn machinery). cli.js does NOT
perform this operation — there is no cli.js citation and none is required: this is not an
Anthropic API surface, it is OCP's own process management around the claude CLI, exactly
as the existing tmux session lifecycle and reaper already are (ALIGNMENT.md Rule 2).

TESTS: 294 passed / 0 failed (was 267). +27 covering acquire/hit/miss, single-use (a pane
is never handed out twice), bounded + serialized refill, TTL + health-check drops, model
retarget, drain/resume, boot-failure backoff, identity linkage, all three reaper
invariants incl. post-drain kill-server restoration, and — the coverage gap that let both
bugs ship — FIVE mid-boot tests: the booting pane is nameable/spareable, the sweep's drain
kills it and resume starts a fresh boot with no bogus WARN, shutdown kills it
synchronously (asserted WITHOUT awaiting, since process.exit runs in the same tick), a
stale settle cannot clear a newer boot's slot, and a model switch cancels an in-flight
boot for the old model.

Live verification (temporary 20s reap interval, reverted): sweep drained both panes ->
reaped -> refilled with NEW panes; a foreign tmux session survived untouched; with no
foreign session kill-server fired and the pool still recovered and served the next
request. Both review bugs reproduced against a PRIVATE tmux server (-L pr3repro, so the
reaper's internal kill-server could not touch the host) before and after the fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(tui): kill a cancelled boot's pane when it settles + make async tests actually count

Folds in the independent review's remaining nit — and, in proving the nit's fix, uncovers
two defects in the test suite itself.

## The nit (latent M1b, second costume)

`_cancelBooting` kills BY NAME, but the tmux session only EXISTS once `bootPane` has run —
and `bootPane` is queued on a microtask. So a caller doing `refill()` then `drain()` in the
SAME synchronous block leaves `_cancelBooting` with nothing to kill (a no-op); it bumps the
generation, and the boot microtask then CREATES the session, succeeds, and — under the old
bare `return` on a stale generation — walked away from a LIVE authenticated `claude` that
nothing owns. Reproduced:

  reverted: drain() kills nothing (no session yet) -> boot creates it -> ORPHAN: ['p1']
  fixed   : drain() kills nothing (no session yet) -> boot creates it -> boot kills it -> []

Not reachable from any current call site, so this is defense-in-depth — but ADR 0008 and the
reap-tick comment in server.mjs BOTH explicitly contemplate a boot-time pre-warm, which is
exactly the shape that reaches it. Killing an already-dead session is a harmless no-op, so
the fix is idempotent whichever way the race lands.

## Defect 1 in the suite: async tests were never awaited (44 of them)

Writing the regression guard exposed this. `test()` called `fn()`, got a promise back, and
IMMEDIATELY printed ✓ and incremented `passed` — without awaiting it. For all 44 tests written
as `test("...", async () => {...})`:
  - ✓ meant "did not throw SYNCHRONOUSLY", not "passed";
  - a failed assertion escaped as an unhandled rejection, crashing the process (CI stays red on
    the non-zero exit) but never being COUNTED — so the summary could print "0 failed" and be wrong.
The suite's headline number was therefore not evidence for ANY async test, including this PR's own
M1a/M1b guards. `test()` now settles an async body before counting it, and the summary awaits them.

## Defect 2, exposed the instant defect 1 was fixed: a false guard

`"a boot that resolves AFTER a drain kills its own pane ... no orphan process left behind"` asserted
`killed.length === 1` — i.e. that kill was CALLED once. But `_cancelBooting`'s kill-by-name on a
not-yet-existent session is a NO-OP that still increments that counter. So "kill was called once" and
"a live session is orphaned" were both true at the same time: a test named for the absence of an
orphan was passing while the orphan was present. Now asserts LIVENESS (`live.size === 0`) — the only
honest question.

## Evidence

  fix present : 295 passed, 0 failed, exit 0
  fix reverted: 293 passed, 2 failed  <- BOTH liveness guards fire (the old kill-count guard did not)

Also: `dropped`'s doc comment now lists `cancelled` (a cancelled in-flight boot lands there via _drop).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 16:51:12 +10:00
e7ce9899f3 docs(plans): TUI streaming IS achievable (MessageDisplay hook) — prereq spike + honest latency constraints (#157)
* docs(plans): TUI streaming is not achievable — prereq spike result + honest README constraints

Backlog #2 of docs/plans/2026-07-13-tui-latency demanded a prereq spike before any
streaming design: does the transcript JSONL grow during a turn, or only at the end?
The spike was run. All three candidate sources are dead:

  (a) transcript JSONL — grows at EVENT granularity; the assistant's text event is
      written as ONE complete line, ~0.3s before the terminal turn_duration event
      (observed: turn_duration 7319ms; text event at t+7.0s, terminal at t+7.3s).
  (b) tmux capture-pane — the pane is a RENDERED view, not the text. Same turn,
      transcript T = '## Semaphore\n\nA **semaphore** is a synchronization…'
      pane        = '⏺ Semaphore' / '  A semaphore is a synchronization…'
      '## ', '**' and ```-fences are absent from the pane entirely (rendered to ANSI,
      then stripped by capture-pane -p). T.startsWith(paneText) is FALSE both raw and
      indent-stripped — not on redraw, but on essentially every markdown answer.
      capture-pane -e recovers styling, never source spelling: no unique inverse.
  (c) --debug-file — byte-exact ('last_assistant_message':'## Title\n\n**alpha…'),
      but only inside end-of-turn Stop-hook payloads; zero content_block_delta /
      text_delta events; ~2.7MB per turn.

--output-format stream-json, the only interface emitting token deltas, requires -p —
the metered-billing path TUI mode exists to avoid (cc_entrypoint=sdk-cli). The
constraint is structural. OCP's TUI SSE is, and remains, replay-only.

Also corrects this plan's own "~20s waiting for the whole turn" decomposition, which
was inferred from an external 30-32s report and never measured through OCP. Measured
through a real OCP instance (TUI, claude-sonnet-4-6, n=5): median 11.30s before #156,
9.55s after, vs a native turn_duration of ~7.3s → OCP's own overhead is ~2-4s, not
~20s. The remainder is generation time, which streaming would not shorten (it moves
the first byte, not the last) — so a consumer needing the COMPLETE answer, which is
the JSON-card case that motivated this work, would have gained nothing from streaming.

Backlog #4 measured while here: --exclude-dynamic-system-prompt-sections gives ZERO
marginal benefit (TTFT median 6.39s vs 6.17s for --effort low alone, n=5, one worse
outlier). Do not adopt. Banner stayed on Claude Max.

README: documents the ~6s TTFT floor plainly (TUI mode cannot serve interactive-latency
consumers) and states that no-token-streaming is structural rather than a missing feature.

No code change. No version bump (docs-only). Not endpoint-touching: no server.mjs diff,
so no cli.js citation applies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

* docs: fold in adversarial review — correct the debug-log reasoning + the overhead number

Independent adversarial reviewer (tasked with REFUTING this doc) confirmed the central
claim — no byte-faithful incremental source exists on the TUI path — but found four
factual defects in the prose. A negative claim that will be cited for years has to be
right in its reasoning, not just its conclusion.

1. --debug-file: the "written at end-of-turn" reasoning was WRONG. The default log level
   is `debug`, which suppresses every `verbose` site; the original probe therefore ran
   with the stream logging OFF. At CLAUDE_CODE_DEBUG_LOG_LEVEL=verbose there ARE 16
   mid-turn `[shoji-engine] yield stream_event/-` lines spread over ~3.9s of generation.
   The conclusion survives because those lines carry TIMING ONLY, no text payload
   (content_block_delta / text_delta / content_block_start / message_start = 0 at any
   verbosity or category filter). Reasoning rewritten: "logs when tokens arrive, never
   what they are" — as written before, the doc was falsifiable in 30 seconds.

2. The 7.319s `turn_duration` is NOT a "native" (non-OCP) baseline: it comes from an
   OCP-driven turn (cwd .ocp-tui/work, same 7451-char prompt, same 204-char answer as
   pr1 baseline row i=5, elapsed 11563ms). Reframed as what it actually is — a SAME-TURN
   decomposition, 11.563s wall - 7.319s CLI-internal = ~4.2s OCP overhead (n=1), which is
   a cleaner comparison than the doc originally claimed.

3. Dropped the "~2-4s" range: its low end mixed an effort-HIGH turn_duration with the
   effort-LOW wall-clock median, which understates overhead (a low-effort turn generates
   faster, so its own turn_duration would be lower). No turn_duration sample exists for
   the effort-low config. Now stated as ~4s (n=1, baseline config), with both caveats.

4. Softened "ZERO marginal benefit" (backlog #4) to "no benefit detectable at n=5" — n=5
   cannot prove zero — and added the mechanistic reason the reviewer supplied, which is
   far stronger than the empirical null: `--help` says the flag improves cross-user
   prompt-cache REUSE, and OCP is single-user, so there is no cross-user cache to share.

Also folded in the reviewer's independent sweep, which closes the search space rather
than sampling it: the hook registry was enumerated from the shipped binary (no per-chunk
/ streaming hook exists among the 21 events); `capture-pane -e` was tested and shown to
be a provably non-unique inverse (an H2 and a bold span emit IDENTICAL SGR 1); and
sessions/<pid>.json, history.jsonl, CLAUDE_CODE_INCLUDE_PARTIAL_MESSAGES (undocumented),
sessionMirror, --sdk-url and --input-format stream-json were each checked and each dies
(contentless, or gated behind --output-format stream-json -> --print -> the metered
sdk-cli pool). Prompt-mutation (asking the model for plain text) is named and rejected
on ALIGNMENT grounds so it is not re-litigated later.

Docs-only. No code change, no version bump.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

* docs: REVERSE the streaming verdict — MessageDisplay hook makes it achievable

The previous commit on this branch concluded TUI streaming was impossible. That was
WRONG, and this corrects it before it could be merged.

The adversarial reviewer commissioned to refute the claim found, on a second pass while
verifying the fold-in, that its OWN first-pass hook enumeration had been truncated by a
400-char grep cap: it reported 21 hook events; the shipped 2.1.207 bundle has 30.
Event #30 is MessageDisplay.

Independently reproduced before acting on it (30 events confirmed via `strings` on the
binary; payload shape `hook_event_name:"MessageDisplay",turn_id,message_id,index,final,
delta`), then live-tested with a MessageDisplay command hook registered via --settings on
a PLAIN INTERACTIVE TUI spawn (no -p, no --bare), claude-sonnet-4-6, --effort low:

  banner: "Sonnet 4.6 with low effort · Claude Max"   ← subscription pool, verified

  7 fires, mid-turn, spread across generation:
    index=0 final=false  '## Mutex\n\n'
    index=1 final=false  'A **mutual exclusion lock** prevents concurrent access to a shar…'
    index=4 final=false  'let counter = 0;\n\nasync function increment() {\n  const release =…'
    index=6 final=true   '```'

  concat(deltas) === T (transcript-authoritative)  ->  TRUE  (579 == 579 bytes)
  T.startsWith(S) at EVERY step                    ->  TRUE  (prefix-stable)
  '## ' / '**' / '```javascript' present in deltas ->  raw markdown SOURCE, not rendered

This satisfies every invariant the previous version declared unobtainable: byte-faithful,
incremental, prefix-stable, no -p, subscription pool. Granularity is block-level (~5-7
chunks/answer), not token-level — which is all an SSE delta.content needs.

Backlog #2 REOPENS and should be built. Implementer caveat recorded: the hook's source
sets forceSyncExecution -> claude BLOCKS on it, so the hook must write and exit
immediately (FIFO/socket), never work inline. Only text blocks fire it (thinking excluded).
ALIGNMENT: consumes claude's OWN hook surface as emitted — forwarding, not inventing
(Class B / ADR 0007; no cli.js citation applies).

Everything still true is kept, and the dead ends are kept as dead ends (they document what
NOT to build): the pane is a rendered view whose source markers are irrecoverable
(capture-pane -e emits IDENTICAL SGR 1 for an H2 and a bold span — a provably non-unique
inverse); the transcript is event-granular; --debug-file carries timing but no payload;
--output-format stream-json requires -p (the metered pool). Also kept: the ~4s (n=1
same-turn) overhead correction, backlog #4's null result with its mechanistic single-user
reason, and the honest value framing — streaming moves the FIRST byte, not the last, so
the complete-answer consumer that motivated this work gains nothing from it.

The wrong conclusion and its refutation are both preserved in the doc. "We checked, it's
impossible" is the most expensive claim to get wrong: it closes a door nobody re-opens.

Docs-only. No code change, no version bump.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

* docs: fold in re-review — the two caveats that would have bitten the implementer

Re-review of the reversed doc came back APPROVE_WITH_MINOR. The reversal itself was
verified complete (worktree-wide grep: no surviving impossibility claim) and NOT
over-claimed in the other direction (the reviewer recomputed every headline number from
the committed messagedisplay-deltas.jsonl and re-ran the invariant on two further turns:
4 independent turns total, 5/6/4/18 fires, 609/696/239/1973 bytes, concat(deltas) === T
TRUE in all four). But two additive caveats were missing, and both are load-bearing for
the streaming PR now in flight:

1. CONCURRENCY DEMUX (severe, and live TODAY — not a warm-pool future problem).
   OCP_TUI_MAX_CONCURRENT defaults to 2, so two `claude` processes already run
   concurrently. One MessageDisplay hook writing to one shared sink would INTERLEAVE
   deltas from two different turns into a single stream — request A's client receiving
   request B's text. A single-request test never surfaces it. The payload carries
   session_id, so the sink must be keyed by it (which also keeps the design warm-pool
   compatible: a pre-booted pane's session-id is fixed at boot, so one static hook script
   serves every pane). Documented, with the required ≥2-concurrent-request test.

2. THINKING-EXCLUSION IS NOT STRESS-TESTED (severe if wrong). The exclusion was inferred
   from a code snippet that turns out to be the final:true call site, not the incremental
   one. Four live turns showed no thinking in any delta — but every transcript's thinking
   block was EMPTY (thinking:"", 0 chars), so it was never actually stressed. If thinking
   deltas do fire on Opus/xhigh, concat(deltas) !== T AND OCP streams the model's private
   reasoning to the caller; the concat === T assertion detects that but cannot un-send an
   SSE delta. Flagged as a must-verify-before-shipping item.

Also: the "5-7 chunks per answer" figure is size-dependent (18 fires on a ~2 KB answer) —
rescoped to "once per rendered block, scales with answer length" in both the doc and the
README, so no implementer hard-codes a chunk-count assumption.

Both caveats were relayed to the streaming implementation immediately rather than waiting
for this merge.

Docs-only. No code change, no version bump.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 16:30:55 +10:00
5258d5d395 feat(tui): pin spawn effort via OCP_TUI_EFFORT (default low) (#156)
* feat(tui): pin spawn effort via OCP_TUI_EFFORT (default low)

buildTuiCmd never passed --effort, so the pane's claude inherited a
HOME-dependent effortLevel: real-home mode inherits the operator's
~/.claude/settings.json (high/xhigh on typical operator hosts),
env-token scratch mode inherits claude's built-in default — proxied-turn
latency silently depended on which HOME mode resolveTuiHome() picked and
on an unrelated operator setting.

Pass --effort explicitly, from new env var OCP_TUI_EFFORT (default
"low"; allowlist low|medium|high|xhigh|max per `claude --help` 2.1.207;
"inherit" restores the pre-flag argv byte-for-byte; an invalid value
warns and falls back to "low" so a typo can never reach the pane argv).

Not endpoint-touching: no server.mjs change, no wire-level change — the
flag rides the existing interactive spawn (ADR 0007). Billing-pool
safety verified per the docs/plans/2026-07-13-tui-latency banner
protocol: startup banner stays "Claude Max" with "low effort".

Measured through a test OCP instance (:3979, TUI mode, real-home,
claude-sonnet-4-6, n=5+5, same ~1850-token prompt as floor.sh):

  before: median 11.30s  range 9.05-12.38s (spread 3.32s)  banner: high effort - Claude Max
  after:  median  9.55s  range 9.27-9.77s  (spread 0.50s)  banner: low effort - Claude Max
  OCP_TUI_EFFORT=inherit: banner back to "high effort" (pre-flag behavior restored)

README: new row in the Environment Variables table (release_kit
new_feature_doc_expectations: new env var -> README table). Tests: 4 new
buildTuiCmd cases (default, explicit level, inherit, invalid fallback);
suite 267 passed / 0 failed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

* docs(readme): review nit — 'pre-v3.22' → 'pre-flag' (next version not fixed yet)

Reviewer nit from the Iron Rule 10 independent review of PR #156.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 15:16:30 +10:00
6854075c01 docs(plans): TUI-mode latency floor — measured decomposition + backlog (#155)
* docs(plans): TUI-mode latency floor — measured decomposition + backlog

An external consumer measured OCP's prompt path at TTFT p50 30-32s and excluded
OCP on that basis. This documents where those 30 seconds actually go, with a
reproducible harness (n=15) that bypasses OCP and measures the underlying
subscription path's true first-token time.

Findings:
- boot -> input-ready is only ~1.0s; it is NOT the bottleneck
- true TTFT is 6-10s; the remaining ~20s is runTuiTurn polling the transcript
  until turn_duration (ADR 0007 step 4) — i.e. waiting for the WHOLE turn.
  There is no streaming.
- buildTuiCmd never passes --effort, so the spawned claude inherits the
  operator's global effortLevel (xhigh on this host) — every request runs
  extended thinking. Passing --effort low: TTFT p50 9.70s -> 6.17s (-36%),
  spread 7.85-13.07s -> 5.87-6.44s. Stays on Claude Max.
- ⚠️ --bare SILENTLY drops off the subscription pool (banner flips
  'Claude Max' -> 'API Usage Billing'). It does cut boot to ~0.5s, but defeats
  the entire purpose of ADR 0007. Failure is silent: all 5 --bare samples
  produced no answer at all (no error, no crash, just never a token).
  Anyone optimizing boot MUST diff the banner line.
- Floor after all fixes is ~6s (claude always injects the full CC system prompt
  + tool definitions). TUI mode therefore cannot serve real-time consumers —
  a constraint worth stating in the README.

Backlog ranked by value/effort: (1) OCP_TUI_EFFORT env var, default low;
(2) real streaming instead of turn_duration polling (~20s, the big one);
(3) warm pane pool (~1s); (4) prefill trim (probably not worth it).

Docs-only; no version bump (matches repo convention — bump lands in the
chore(release) commit).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dx5Ncq6wWBrF27vJKHZ9Hr

* docs(plans): address review — restore --bare evidence, qualify effort claim, add banner captures

Reviewer (fresh-context, Iron Rule 10) returned REQUEST_CHANGES. All four technical
conclusions survived independent verification (source-read + live repro); the defects
were in the evidence file, and they were real:

- H-1: measurements.jsonl claimed n=15 but held 10 rows, and the --bare group — the
  basis of this PR's headline warning — had ZERO rows. The author had stripped them
  as 'invalid samples' (ttft_ms:-1) when they were in fact the evidence. Regenerated:
  n=15, three groups × 5, all with tag/extra_args. --bare reproduces exactly (5/5 no
  answer, boot 0.43-0.45s).
- M-1: the effort-inheritance claim was written unconditionally, but it depends on
  resolveTuiHome()'s mode. Real-home (current service config) inherits the operator's
  effortLevel: xhigh; env-token scratch home (~/.ocp-tui/home) has no effortLevel in
  its settings.json and prepareTuiHome() never writes one, so the pane gets claude's
  built-in default. Now documented as a table — and the mode split makes passing
  --effort explicitly MORE valuable, not less.
- M-2: baseline rows were produced by a pre-parameterized script and lacked
  tag/extra_args. Re-run with the committed script. Recomputed effect: -40% (was -36%).
- L-1: documented that the harness suppresses OCP's periodic kill-server tick via the
  othersRemain coexistence guard (by design, resumes next tick).
- L-2: documented that floor.sh's readiness marker differs from OCP's tuiInputReady(),
  so the ~1.0s boot figure is not apples-to-apples with BOOT_MS.
- Direct-API reference figure now explicitly labeled as external (not in this dataset).
- New: billing-banner.txt captures all three configs live, including confirmation that
  --effort low stays on Claude Max (reviewer noted this was asserted but unevidenced).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dx5Ncq6wWBrF27vJKHZ9Hr

* docs(plans): scope the effort claim to TUI mode (currently off), drop nonexistent bin/

Re-review (APPROVE_WITH_MINOR) caught two accuracy defects:
- MIN-1: 'every OCP request runs extended thinking' over-extrapolated. TUI mode is
  currently OFF on this host (CLAUDE_TUI_MODE=false; /health tui.enabled=false), so
  live traffic takes the -p path. The claim is about what happens WHEN TUI mode is
  enabled — now scoped, and the same qualifier applied to the kill-server interaction
  note (that reap tick is itself gated on TUI_MODE).
- NIT-2: the quoted grep included bin/, which does not exist in the repo (exit 2).
  Dropped; the zero-hit result over lib/ + server.mjs is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dx5Ncq6wWBrF27vJKHZ9Hr

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 14:22:17 +10:00
18 changed files with 3279 additions and 83 deletions
+55 -5
View File
@@ -24,7 +24,7 @@ One proxy. Multiple IDEs. All models. **$0 API cost.**
There are several Claude proxy projects. OCP picks a specific lane: **align tightly with what `cli.js` actually does, observe + multiplex what's already there, don't extend the protocol.** What you get:
- **LAN multi-user keys** (v3.7.0) — share one Claude Pro/Max subscription with family, friends, or your own devices. Each user gets a per-key API token (no OAuth session leak), with independent usage tracking and one-line revocation.
- **LAN multi-user keys** (v3.7.0) — reach one Claude Pro/Max subscription from your own devices across the LAN. Each device gets a per-key API token (no OAuth session leak), with independent usage tracking and one-line revocation. Pro/Max are **per-user** accounts — see [Sharing with family / a team — honest limits](#deployment-model--security-read-this) before extending access to other **people**.
- **`ocp-connect` one-shot IDE setup** — one command on the client machine detects and configures Claude Code, Cursor, Cline, Continue.dev, OpenCode, and OpenClaw. No pasting `OPENAI_BASE_URL` six times.
- **Response cache with per-key isolation + singleflight** (v3.13.0). Optional SHA-256 prompt cache, isolated per API key (cross-user pollution is impossible by hash construction, not by application logic), with stampede protection on concurrent identical prompts. Off by default. ([PR #65](https://github.com/dtzp555-max/ocp/pull/65), [PR #66](https://github.com/dtzp555-max/ocp/pull/66))
- **Per-key request quotas** (v3.8.0). Daily / weekly / monthly limits per key — set a kid's iPad to 20/day, a partner's laptop to 100/week. ([PR #18](https://github.com/dtzp555-max/ocp/pull/18))
@@ -49,7 +49,7 @@ OCP and the alternatives serve adjacent but distinct needs. Pick the one that fi
| GitHub stars / ecosystem size | small | large | mid |
| Governance discipline (CI-enforced alignment with cli.js) | yes | n/a | n/a |
**Plain English**: `claude-code-router` is the routing-and-switching power tool — pick it if you want to mix Anthropic, OpenAI, Gemini, and local models behind one endpoint. `anthropic-proxy` is the minimal forwarder. **OCP focuses on disciplined `cli.js`-aligned forwarding plus subscription multiplexing** — pick it if you want to share one Claude Pro/Max subscription across IDEs, devices, and people, with LAN auth, quotas, and a governance contract that prevents endpoint drift.
**Plain English**: `claude-code-router` is the routing-and-switching power tool — pick it if you want to mix Anthropic, OpenAI, Gemini, and local models behind one endpoint. `anthropic-proxy` is the minimal forwarder. **OCP focuses on disciplined `cli.js`-aligned forwarding plus subscription multiplexing** — pick it if you want to reach one Claude Pro/Max subscription from your own IDEs and devices, with LAN auth, quotas, and a governance contract that prevents endpoint drift.
### Related: OLP — Open LLM Proxy
@@ -215,7 +215,7 @@ After install the `ocp` CLI lives at `~/ocp/ocp`. To put it on your PATH, either
export OPENAI_BASE_URL=http://127.0.0.1:3456/v1
```
**LAN mode** — share with other devices on your network:
**LAN mode** — reach OCP from your own devices on the network (Claude Pro/Max are per-user accounts — see [Sharing with family / a team — honest limits](#deployment-model--security-read-this) before extending access to other people):
```bash
# Enable LAN access with per-user auth (recommended)
node setup.mjs --bind 0.0.0.0 --auth-mode multi
@@ -958,7 +958,13 @@ See [Subscription-pool (TUI) mode](#subscription-pool-tui-mode) and ADR 0007 PR-
| `OCP_TUI_CWD` | `$HOME/.ocp-tui/work` | (TUI-mode) Scratch working directory where interactive claude sessions run. Transcripts land under `<HOME>/.claude/projects/<encoded-cwd>/`. Created automatically. |
| `OCP_TUI_HOME` | *(auto)* | (TUI-mode) `HOME` claude runs under. **When unset, OCP picks it for you:** if `CLAUDE_CODE_OAUTH_TOKEN` is set → a **credential-isolated** scratch home `$HOME/.ocp-tui/home` (no `credentials.json`, env-token auth — **recommended**); if no env token → the operator's real home (legacy shared `credentials.json`). Setting this to an **explicit** path overrides the auto-default. The credential handling at that path still follows the env token: **with** the env token it is credential-free (env-token auth, no `credentials.json` written); **without** the env token (and the path ≠ real home) it uses the legacy symlinked-credentials scratch mode, which carries the credential-fork caveat — see ADR 0007. |
| `OCP_TUI_ENTRYPOINT` | `cli` | (TUI-mode) Billing-classifier labeling: `cli` (default) pins `cc_entrypoint=cli` deterministically; `auto` lets claude self-classify via TTY detection; `off` leaves the inherited env untouched. Honest only when the spawn is a genuine interactive PTY — see ADR 0007. |
| `OCP_TUI_EFFORT` | `low` | (TUI-mode) Effort level passed to the interactive `claude` as an explicit `--effort` flag: `low` (default), `medium`, `high`, `xhigh`, `max`, or `inherit` to omit the flag (the pre-flag behaviour: the pane inherits a HOME-dependent effort — the operator's `~/.claude/settings.json` `effortLevel` in real-home mode, claude's built-in default in env-token scratch mode). Explicit `low` cuts measured TTFT p50 by ~40% and collapses run-to-run variance ~15× versus an inherited `xhigh` (see `docs/plans/2026-07-13-tui-latency/`); proxied requests rarely benefit from extended thinking. Banner-verified to stay on the subscription pool (`· Claude Max`). An invalid value logs a warning and falls back to `low`. |
| `OCP_TUI_STREAM` | `0` (off) | (TUI-mode) When `=1`, `stream:true` requests emit **real SSE `delta.content` chunks as `claude` generates them**, instead of buffering the turn and replaying it. Deltas come from `claude`'s own `MessageDisplay` hook (registered with `--settings` on the ordinary interactive spawn — banner-verified to stay on the subscription pool, `· Claude Max`). Granularity is **block-level**, not token-level. The transcript remains authoritative: the streamed text is asserted equal to it at end-of-turn, the auth-banner and truncation gates still run before anything is committed, and only the transcript text is cached. A turn whose stream cannot be reconciled with the transcript is **refused** (SSE error frame, not cached) and counted as `tui.streamDivergences` on `/health`. A total hook failure (e.g. `--settings` stops registering it after a `claude` version bump) is a *different, silent* failure mode — every streamed turn still succeeds, fully buffered, with no divergence and no error — so it is counted separately as `tui.streamZeroDeltaTurns` (streamed turns where the hook fired **zero** times) and logged as `tui_stream_zero_deltas`; watch it alongside `streamDivergences`. Default off — the buffered path is unchanged and remains the stable default. ⚠️ **Tool-using turns:** the transcript keeps only the model's **last** assistant message, so if the model narrates before calling a tool ("I'll check that file…") and that narration exceeds `OCP_TUI_STREAM_HOLDBACK`, it has already been streamed and cannot be retracted — the turn is then **refused** rather than served (measured live: Opus narrated 475 chars before a `Bash` call). If your deployment lets the model use tools (the TUI default, and anything with `OCP_TUI_FULL_TOOLS=1`), either raise `OCP_TUI_STREAM_HOLDBACK` above the typical narration length — the narration then stays held back and is correctly discarded, at the cost of a later first chunk — or leave streaming off. Streaming is best suited to tool-light chat proxying. See ADR 0007 (2026-07-13 amendment). |
| `OCP_TUI_STREAM_HOLDBACK` | `100` | (TUI-mode, streaming) Characters withheld before the first chunk reaches the client. Two jobs. (1) It keeps the **auth-banner gate** alive under streaming, via a guarantee with two required halves: (i) nothing is emitted for a message until its trimmed accumulation exceeds 100 chars — past the default banner detector's reach, since real banners are ≤100 chars — and (ii) once a message boundary follows an emit, nothing further is ever emitted for the rest of the turn, and the turn is refused outright. Half (i) alone only covers a turn's first message; half (ii) is what covers an error banner rendered as a *later* message (e.g. after tool-using prose). Raise the holdback if you replace the detector via `CLAUDE_TUI_ERROR_PATTERNS` with patterns that can match longer messages — that only affects half (i); OCP warns at boot if you do. (2) It is the knob for **tool-using turns** — see the `OCP_TUI_STREAM` caveat below. Answers shorter than the holdback are simply delivered whole at end-of-turn, exactly as the buffered path does. |
| `OCP_TUI_STREAM_DIR` | `$HOME/.ocp-tui/stream` | (TUI-mode, streaming) Directory holding the static `MessageDisplay` hook script + settings file, and the per-session delta sink (`<session-id>.jsonl`, removed at turn teardown). One sink **per session-id** — this is what keeps concurrent TUI turns (`OCP_TUI_MAX_CONCURRENT` ≥ 2) from interleaving one client's deltas into another's stream. |
| `OCP_TUI_STREAM_POLL_MS` | `100` | (TUI-mode, streaming) Interval at which OCP drains the delta sink. The hook fires at block granularity (seconds apart), so a finer poll buys nothing. |
| `OCP_TUI_MAX_CONCURRENT` | `2` | (TUI-mode) Max concurrent interactive TUI turns. **Independent** of `CLAUDE_MAX_CONCURRENT` (which bounds the `-p`/stream-json path; TUI never uses it). A TUI turn is heavy (per-request cold-boot of tmux+claude + up to `CLAUDE_TUI_WALLCLOCK_MS` wallclock), so the default is low to keep small hosts (e.g. a Pi 4) alive under a burst. Excess turns **queue** (bounded); a full queue yields a 503. See ADR 0007 PR-B amendment. |
| `OCP_TUI_POOL_SIZE` | `0` (off) | (TUI-mode) Number of **pre-booted warm `claude` panes** kept ready, so a request does not pay the cold boot. `0` disables the pool entirely — the request path is then exactly the cold-boot path. Max `4`; an unparseable value disables it rather than guessing. **Measured on a Mac mini (Sonnet 4.6, `--effort low`): end-to-end p50 `10.17s` (n=6, pool off) → `6.00s` (n=12 warm hits) — 4.2 s / 41%** — the pool recovers both the ~1.2 s boot *and* ~2.9 s of post-input-bar init that a pane which has been idle a moment has already finished. **Cost:** each warm pane is a *live idle `claude` process* held whether or not a request ever arrives (peak processes ≈ pool size + `OCP_TUI_MAX_CONCURRENT` + 1 booting replacement) — which is why it is opt-in. Panes are **single-use**: one turn, then killed and replaced in the background. The **first request after start (and after any model switch) is always a cold miss** — the pool warms the most recently requested model, since OCP cannot know which model the next caller wants. See `docs/plans/2026-07-13-tui-latency/`. |
| `OCP_SKIP_AUTH_TEST` | *(unset)* | When `=1`, skip the `claude -p` auth probe during `setup.mjs`. After 2026-06-15 this probe draws from the Agent SDK credit pool; set this to avoid burning a metered credit on re-installs or `ocp update` runs. Auth is validated at the first real request. |
| `OCP_TUI_FULL_TOOLS` | *(unset)* | (TUI-mode, **single-user only**) When `=1`, grant the interactive session the **same tool surface as the `-p` path** — `--allowedTools` (+ optional `--mcp-config`, read from `CLAUDE_ALLOWED_TOOLS` / `CLAUDE_MCP_CONFIG`) — instead of the default MCP-walled, built-in-tools-only set. Lets a trusted single-operator TUI deployment run a **tool-using / MCP agent** (e.g. an OpenClaw assistant) on the subscription pool. Safe because TUI **refuses to boot under `AUTH_MODE=multi`** (hard exit) — no guest key can ever reach the TUI path, so this gate cannot expose tools to an untrusted caller. (Under `AUTH_MODE=shared` + `OCP_TUI_ALLOW_LAN=1`, anyone holding the single shared key reaches it — that is the existing TUI trust model, unchanged.) Note: `--dangerously-skip-permissions` / `CLAUDE_SKIP_PERMISSIONS` is **not** supported for TUI — claude v2.1.x shows an interactive bypass-acceptance screen in headless tmux that cannot be answered, bricking the pane. Use scratch-home `settings.json` `additionalDirectories` instead. See [Subscription-pool (TUI) mode](#subscription-pool-tui-mode) and ADR 0007. |
@@ -1032,13 +1038,43 @@ Then restart OCP. At boot you will see (with the env token set, isolated home au
### What changes / what doesn't
- **Callers see no API change.** The response is a normal OpenAI completion object or chunked SSE — identical wire format.
- **No real token streaming.** TUI-mode buffers the full response then replays it as chunked SSE. You will see a delay then the complete response rather than real-time tokens.
- **Real streaming is opt-in (`OCP_TUI_STREAM=1`), and off by default.** By default TUI-mode buffers the full response and replays it as chunked SSE — you see a delay, then the complete response. Set `OCP_TUI_STREAM=1` and `stream:true` turns emit real SSE `delta.content` chunks as `claude` renders them, sourced from `claude`'s own `MessageDisplay` hook (byte-faithful raw markdown, on the subscription pool, no `-p`). Two honest caveats: granularity is **block-level** — the hook fires once per rendered block, so a handful of chunks per answer, scaling with length, not token-by-token; and it moves the **first** byte, not the last, so a consumer that must parse a complete reply gains nothing. The transcript stays authoritative: every streamed turn is asserted against it at the end, and a turn whose stream disagrees is **failed rather than served** (watch `tui.streamDivergences` on `/health`). Evidence: [`docs/plans/2026-07-13-tui-latency/streaming-spike.md`](docs/plans/2026-07-13-tui-latency/streaming-spike.md).
- **Cache and singleflight work normally.** TUI-mode writes the buffered response to the cache on success; cache-hits skip the interactive turn entirely.
- **The host's `CLAUDE.md` / auto-memory is never injected.** OCP is a proxy — the proxied client (OpenClaw / your IDE) owns its own context and memory. TUI-mode always runs `claude` with `CLAUDE_CODE_DISABLE_CLAUDE_MDS` + `CLAUDE_CODE_DISABLE_AUTO_MEMORY`, so a `CLAUDE.md` on the OCP host can never leak into proxied turns (verified live; see #4). Built-in tool schemas + the interactive system prompt remain (the inherent ~2035K context floor of interactive mode); MCP is hard-disabled.
- **Authenticate via `CLAUDE_CODE_OAUTH_TOKEN` in a credential-isolated home (recommended).** tmux does not forward the parent process's env to the pane, so OCP sets the token explicitly on the spawned `claude` when `CLAUDE_CODE_OAUTH_TOKEN` is present. But passing the token is **not enough on its own**: interactive `claude` *prefers* `~/.claude/.credentials.json` over the env var (unlike the `-p` path), so a stale `credentials.json` would shadow the token. With the env token set and `OCP_TUI_HOME` unset, OCP therefore runs claude in a **credential-isolated home** (`$HOME/.ocp-tui/home`) that has **no `credentials.json`** — so the env token is the only credential and is authoritative, and claude never runs the token-refresh path (so the single-use refresh token can't be corrupted by the spawn/teardown cycle). On a long-running host the credentials.json path produced a permanent `Please run /login · API Error: 401` that re-login could not fix (the next spawn re-corrupted it); the isolated home ends that at the root. Transcripts land under the same isolated home, so the answer-reader is unaffected. Without the env token, claude falls back to the real home's `credentials.json` (byte-for-byte the previous behaviour). (The token is visible in `ps` on the pane command — acceptable for the single-user A-path; the multi-user B-path is refused at boot.) See ADR 0007 PR-C / PR-D amendments.
- **Stale tmux sessions are reaped.** The pane's `claude` is a child of the tmux server (not OCP), so OCP cannot reap it directly; `claude` zombies can otherwise accumulate as `<defunct>` over a long-running host. OCP reaps them at boot and on a 15-min idle sweep by issuing `tmux kill-server` — but **only when no foreign tmux session remains** (it never disrupts a co-hosted `olp-tui-*` instance). See ADR 0007 PR-C amendment.
- **Default path unchanged.** Unset `CLAUDE_TUI_MODE` and restart → `callClaude` / `callClaudeStreaming` are used again, byte-for-byte identical to today.
- **Concurrency is bounded separately.** TUI turns are heavy (per-request cold-boot + long wallclock), so the TUI path has its own limiter — `OCP_TUI_MAX_CONCURRENT` (default `2`), independent of `CLAUDE_MAX_CONCURRENT`. Excess turns queue; a full queue returns a 503. Tune it up only on a host that can run more interactive `claude` sessions at once.
- **Optional warm pane pool (`OCP_TUI_POOL_SIZE`, default off).** Pre-boots panes so a request skips the cold boot — measured p50 `10.17s` → `6.00s` (41%). Pooled panes are **single-use** (one turn, then killed and replaced in the background), each carrying its own fresh `--session-id`, so one session still means one exchange and no earlier-turn text can leak into a later answer. They are named `ocp-tui-<port>-p<hex>` and coexist with the reaper by design: the sweep **drains the pool first**, then reaps (so `kill-server` still flushes `<defunct>` zombies), then the pool refills in the background. Drain→reap→resume is synchronous, so no request can land mid-sweep; a request arriving while the pool is still re-booting simply misses it and cold-boots. A live pooled pane is never reaped — **including one that is still booting**, whose tmux session already exists — while an *orphaned* one (left by a previous process generation) still is.
### ⚠️ Latency: TUI mode has a ~6-second floor, and it is immovable
**TUI mode cannot serve real-time or interactive-latency consumers.** This is a hard property of the
path, stated plainly so you can rule it out before building on it:
| | measured |
|---|---|
| **TTFT floor (first token)** | **≈ 6 s** — immovable |
| cold boot → input bar ready | ~1 s (per request; not the bottleneck) |
| OCP's own overhead above the CLI | ~4 s (n=1 same-turn decomposition) |
| direct Anthropic API, same prompt (for scale) | 0.841.64 s |
The ~6 s floor is the `claude` CLI itself: it always injects the full Claude Code system prompt plus
its tool definitions before your prompt, on every turn, no matter what you ask. No flag removes it
(`--exclude-dynamic-system-prompt-sections` was measured: **no effect** on the floor). Extended
thinking is *not* the cause — `OCP_TUI_EFFORT` already defaults to `low`, which is what cuts a
formerly-inherited `xhigh` down to this floor and collapses its variance.
On top of the floor you pay the model's generation time (a function of output length). Progressive
output is not wired up **yet** (see "No real token streaming" above — it is achievable and planned),
so today a turn returns as one blob once generation completes. Note that streaming, when it lands,
will move the *first* byte earlier — it does **not** shorten the turn, and a consumer that needs the
complete answer gains nothing from it.
**Use TUI mode for**: batch, background, and latency-insensitive work where the subscription pool is
the point. **Do not use it for**: anything a person is waiting on interactively, or any consumer with
a sub-5-second budget. Full measurements and methodology:
[`docs/plans/2026-07-13-tui-latency/`](docs/plans/2026-07-13-tui-latency/).
### Monitoring drift via `/health`
@@ -1052,12 +1088,26 @@ Then restart OCP. At boot you will see (with the env token set, isolated home au
"entrypointMismatches": 0, // count of cli-expected-but-got-other turns — ALERT if this climbs
"inflight": 1, // TUI turns running right now
"queued": 0, // TUI turns waiting for a concurrency slot
"maxConcurrent": 2 // OCP_TUI_MAX_CONCURRENT
"maxConcurrent": 2, // OCP_TUI_MAX_CONCURRENT
"pool": { // warm pane pool — null when OCP_TUI_POOL_SIZE=0 (the default)
"size": 2, // target warm panes (OCP_TUI_POOL_SIZE)
"warm": 2, // panes ready right now — each is a LIVE idle claude process
"booting": 0, // replacement panes currently pre-booting
"model": "claude-sonnet-4-6", // the model being warmed (the most recently requested one)
"hits": 12, // requests served by a warm pane
"misses": 1, // requests that fell back to the cold boot (the 1st is always one)
"boots": 14, // panes successfully pre-booted
"bootFailures": 0, // pre-boots that genuinely never reached the input bar — WATCH this
"cancelled": 4, // in-flight boots OCP killed on purpose (drain / model switch) — not faults
"dropped": 8 // panes discarded unused (drain sweep / expired / unhealthy)
}
}
```
Alert on `entrypointMismatches > 0` (or `lastEntrypoint !== "cli"`): it means a turn drew from the metered Agent SDK pool instead of the subscription. `inflight` / `queued` show how close the TUI path is to its concurrency cap.
With the pool on, `hits` / `misses` is the hit rate (a steady single-model consumer should sit near 100% after the first request), and `warm` is your standing idle-process cost. A climbing `bootFailures` means panes are not reaching their input bar — the pool then degrades safely to the cold path, but latency reverts to the un-pooled numbers. `cancelled` counts boots OCP killed *on purpose* (a drain, a model switch) and is **not** a fault signal — do not alert on it. A steadily climbing `dropped` is likewise normal: the 15-min reap sweep drains and re-boots the pool on every tick so `kill-server` can still flush `<defunct>` zombies.
### Kill-switch
```bash
+56 -1
View File
@@ -56,7 +56,7 @@ Add `CLAUDE_TUI_MODE=true` as an opt-in flag in `server.mjs`.
3. The serialized prompt (from `messagesToPrompt`) is pasted via `tmux send-keys … "$(cat file)"` + a separate `Enter` key event.
4. The answer is read from claude's native JSONL transcript at `<HOME>/.claude/projects/<encoded-cwd>/<session-id>.jsonl`, polling until a `turn_duration` system event or the wall-clock cap (`CLAUDE_TUI_WALLCLOCK_MS`, default 120 s).
5. The string answer is returned to OCP's existing downstream (singleflight → cache write-back → `completionResponse` / `streamStringAsSSE`) — **same contract as `callClaude`**.
6. Streaming requests are buffered then replayed as chunked SSE (no real token streaming — deliberate; "don't build fragile features").
6. Streaming requests are buffered then replayed as chunked SSE (no real token streaming — deliberate; "don't build fragile features"). **Superseded for `stream:true` when `OCP_TUI_STREAM=1` — see the 2026-07-13 amendment below. The buffered path remains the default and is unchanged.**
### Billing-classifier labeling (`OCP_TUI_ENTRYPOINT`, PR-4)
@@ -333,6 +333,61 @@ The original "Home strategy" section and PR-C's `prepareTuiHome` comment warned
---
## Amendment (2026-07-13) — real SSE streaming via the `MessageDisplay` hook (`OCP_TUI_STREAM`)
**Supersedes**: Request-flow step 6 above ("no real token streaming — deliberate"), for `stream:true`
requests when `OCP_TUI_STREAM=1`. The buffered path stays the default and is byte-for-byte unchanged.
**Context.** Step 6 was written when the interactive CLI appeared to expose no byte-faithful
incremental source. A prereq spike (`docs/plans/2026-07-13-tui-latency/streaming-spike.md`) confirmed
three obvious sources are dead ends — the transcript JSONL grows one *whole event* at a time (the
answer lands as a single line ~0.3 s before the terminal marker); `tmux capture-pane` yields a
*rendered* view whose markdown source is unrecoverable (an H2 and a bold span produce identical ANSI);
`--debug-file` logs stream *timing*, never stream *content*. Every interface that does emit
`text_delta` (`--output-format stream-json`) requires `-p`, which moves the request to the **metered**
`sdk-cli` pool — precisely what TUI-mode exists to avoid.
**Decision.** Consume `claude`'s own **`MessageDisplay`** hook, registered via `--settings` on the
ordinary interactive spawn (no `-p`, no `--bare`). Each fire delivers the **raw markdown source** of an
incremental `delta` on the hook's stdin. Verified live (claude 2.1.207, sonnet-4-6): banner stays
`· Claude Max` and the transcript `entrypoint` stays `cli` (subscription pool); `concat(deltas) === T`
byte-exactly; `T.startsWith(concat(deltas[0..n]))` at every *n*. This is **forwarding, not inventing**
— ALIGNMENT.md **Class B**. No `cli.js` citation applies: the TUI spawn is OCP-owned surface (this
ADR), the hook payload is claude's own published contract, and the SSE wire shapes are the OpenAI
chat/completions streaming spec adopted by **ADR 0006** (the emitters are literally the `-p` path's).
**The transcript remains authoritative.** It is still the terminal-turn signal, still the source of the
returned/cached text `T`, and still the input to the honesty gates (auth-banner detection C-1,
`truncated` C-2). The delta stream is a low-latency **mirror**, never a replacement. At end of turn OCP
asserts the streamed bytes against `T`: equal → serve; a strict *prefix* of `T` → top up from the
transcript (client still receives exactly `T`); **not** a prefix → **refuse the turn** (SSE error frame,
no cache, `tui.streamDivergences++`). Serving text the transcript disagrees with is the failure class
ALIGNMENT.md exists to prevent, so streaming fails loud rather than degrading quietly.
**Consequences / constraints recorded for future authors:**
- **Opt-in, default OFF.** The buffered path is stable production; streaming does not change it.
- **Per-`session_id` sink is mandatory, not an optimization.** `OCP_TUI_MAX_CONCURRENT` defaults to
**2** — two `claude` panes already run concurrently. A single shared sink would interleave one
client's deltas into another's stream. The hook writes to `<dir>/<session_id>.jsonl`, the path
delivered through the *pane's own env* (`OCP_TUI_STREAM_FILE`); OCP reads only its own turn's file.
Verified with two concurrent streamed turns (ALPHA/BRAVO): zero cross-contamination.
- **Warm-pool compatible (a separate in-flight PR depends on this).** The hook script and the settings
file are **static** — nothing request-specific is baked in at spawn time. The sink path derives from
the session-id, which for a pre-booted pane is fixed at boot.
- **The hook is synchronous** (`forceSyncExecution: true``claude` *blocks* on it). The hook script
must write and exit; it does one `cat` append and nothing else. Measured: p50 **7.2 ms** per fire,
~50 ms across a whole turn — noise against a 610 s turn. Do not add work to it.
- **Thinking blocks do not fire the hook** — verified on a substantive Opus/`xhigh` reasoning turn (see
the PR evidence), not merely inferred from the `final:true` call site. This must be **re-verified** if
the hook is ever pointed at a new model/effort tier: a thinking delta reaching a client would be
unretractable, and the `concat === T` assertion can only *detect* that after the fact, never prevent
it. The first-bytes **holdback** (`OCP_TUI_STREAM_HOLDBACK`, default 100 chars) is the same
prevention-not-detection reasoning applied to the auth-banner gate.
- **Block-level granularity**, scaling with answer length — not token-level. Do not promise otherwise.
- **It moves the first byte, not the last.** Only a progressively-rendering consumer benefits; it does
not move TUI-mode's ~6 s TTFT floor.
## Provenance
TUI-mode originated in a prototype contributed via PR #101 (see the PR for author attribution). The productionization design is in `docs/superpowers/specs/2026-05-30-tui-mode-production-design.md`. Spikes S1S6 / T1T6 were validated live on the test host against `claude v2.1.158`.
+208
View File
@@ -0,0 +1,208 @@
# ADR 0008 — TUI Warm Pane Pool
**Date:** 2026-07-13
**Status:** Proposed
**Extends:** [ADR 0007](0007-tui-interactive-mode.md) (TUI interactive mode). This ADR does not
change ADR 0007's billing-pool argument, security posture, or kill-switch — it adds a latency
optimization *inside* the TUI spawn machinery ADR 0007 owns.
---
## Context
TUI mode (ADR 0007) serves every request by cold-booting a fresh `tmux` session running an
interactive `claude`, submitting one prompt, reading the native transcript, and killing the
session. That cold boot is paid on **every** request.
[`docs/plans/2026-07-13-tui-latency/`](../plans/2026-07-13-tui-latency/README.md) measured the
TUI path and listed a warm pane pool as backlog item #3, costed at "**~1.0 s**" (the observed
boot-to-input-bar time). Instrumenting the real request path showed that estimate is **~4×
too low**. Phase decomposition of the cold path (n=6 medians, Sonnet 4.6, `--effort low`,
through a real OCP instance):
| Phase | Median |
|---|---|
| prep (trust cwd, write prompt file) | 2 ms |
| `tmux new-session` | 27 ms |
| **boot → input bar ready** | **1232 ms** |
| paste (`load-buffer` + `paste-buffer`) | 8 ms |
| paste-verify poll | 426 ms |
| **submit → transcript terminal** | **8458 ms** |
| teardown | 8 ms |
| **total** | **10162 ms** |
| *claude's own reported `turn_duration`* | *5539 ms* |
| **OCP-side overhead** | **4490 ms** |
The `submit → terminal` phase exceeds claude's own `turn_duration` by **~2.9 s**. That gap is
**post-input-bar initialization inside `claude`** — work that a pane which has merely *sat idle
for a few seconds* has already completed. A direct spike confirmed it: an identical pane, idle
12 s before receiving the same prompt, completed its turn in a median 5537 ms versus 7980 ms
cold.
So a warm pane recovers **~1.26 s of boot *and* ~2.9 s of in-`claude` cold start** — not the
~1.0 s the plan predicted.
The reason this was worth a pool rather than a "keep one session and reuse it" cache is a
hazard already flagged in the code. `lib/tui/transcript.mjs` returns the **last text-bearing
assistant entry in the whole transcript file**, which is correct *only* under OCP's
one-session-per-request model, and it says so:
> *"If a future warm-pool ever reuses a session WITHOUT a fresh session-id / clear, earlier-turn
> text could leak — that author must add user-line scoping here."*
Reusing a pane for a second turn puts two exchanges in one transcript and would leak the earlier
turn's text into the later turn's answer — a **cross-request data leak**, not merely a bug.
---
## Decision
Add an **opt-in pool of pre-booted, single-use `claude` panes**, `OCP_TUI_POOL_SIZE` (default
`0` = off, max `4`). Implementation: `lib/tui/pool.mjs`.
### 1. Panes are SINGLE-USE. This is the load-bearing rule.
A pooled pane serves **exactly one turn**, then is killed and replaced in the background. Each
pane is booted with its **own fresh `--session-id`**, fixed at spawn, and the turn locates its
transcript by that id.
This preserves one-session-per-request exactly, so the `transcript.mjs` hazard above **does not
arise** and no user-line scoping was needed. The warning in `transcript.mjs` is deliberately
left standing, now annotated: it still binds anyone who later wants a pane to serve a second
turn, or to reset a session with `/clear` and reuse it. **Neither is permitted without first
adding user-line scoping to the transcript reader.**
Rejected alternative — *reuse a pane for N turns, `/clear` between* — is strictly cheaper
(no re-boot per request) and was rejected on exactly this basis. The latency win is not worth a
cross-request text-leak surface guarded only by a `/clear` that we cannot verify landed.
### 2. The pool is keyed by model, and a MISS is always safe.
`--model` is fixed at spawn, so a pane can only serve the model it booted with. A pool miss
falls back to the existing cold-boot path with **zero behavioural difference**. There is no
boot-time pre-warm and no configured model: OCP cannot know which model the next caller wants,
so the pool warms the **most recently requested** model. Consequence, stated plainly: **the
first request after start, and the first after any model switch, is always a cold miss.**
### 3. The pool and the session reaper coexist by an explicit invariant.
This is the subtle part. `reapStaleTuiSessions()` kills every session matching this instance's
`ocp-tui-<port>-` prefix, and issues `tmux kill-server` when no foreign session remains (the
only mechanism that can reap `<defunct>` `claude` zombies — the pane's `claude` is a child of
the tmux *server*, not of node). A warm pooled pane **is** one of our own sessions, alive and
idle **by design** — and the periodic sweep runs precisely **when the instance is idle**, i.e.
exactly when the pool is full.
The invariant, stated in a comment above `reapStaleTuiSessions` and pinned by tests:
1. **A live pooled pane is never reaped — including one that is still BOOTING.** The reaper
takes a `spare` set of **exact session names** supplied by the pool's live registry.
2. **An orphaned pooled pane IS still reaped.** Membership is by **exact name from a live
in-memory registry, never by name shape**. A pane the pool no longer owns — handed out,
dropped, cancelled, or left behind by a previous process generation (whose registry died with
it) — is absent from `spare` and is killed like any other stale session. **Fail-safe:
omitting `spare` reaps *more*, never less.** Pool panes are named `ocp-tui-<port>-p<hex>`
purely for operator legibility; that shape is *not* the exemption mechanism.
3. **`kill-server` is suppressed while any pane is spared** (it would kill a live child of the
tmux server). Therefore **the pool is DRAINED immediately before every sweep**, so `spare` is
empty on the normal tick and `kill-server` still fires. Without the drain, a permanently-full
pool would **permanently disable zombie reaping** — the pool would silently break the thing
the sweep exists to do. The drain costs one pane re-boot per tick (15 min).
The `spare` mechanism is belt-and-braces given the drain: it makes it impossible for a reap call
site that *forgets* to drain to kill a live pane.
### 4. The pool tracks its in-flight boot BY NAME, not as a count.
`bootTuiPane` creates the tmux session **synchronously** and only *then* waits (up to
`POOL_BOOT_MS`, 20 s) for the input bar. So **a pooled tmux session can be live for ~20 s before
its boot resolves.** A pool that tracked in-flight boots as a *count* could not name that
session, and this produced two real bugs (both caught in review, both now regression-tested):
- the periodic sweep **killed the booting pane** (it could not be spared), then left the pool
empty with nothing scheduled, and logged the exact `tui_pool_boot_failed` warning operators are
told to alert on — for a completely healthy drain;
- graceful shutdown **orphaned a live, authenticated, idle `claude`**: `gracefulShutdown` calls
`process.exit(0)` in the same tick as the drain (TUI panes are tmux children, so node's
`activeProcesses` set is empty and the "wait for children" path exits immediately), so any
cleanup deferred to a `.then()` never ran.
The pool therefore **mints each pane's identity up front** (`{sessionId, name}`, name derived
from the session-id so `tmux ls` correlates to the transcript file) and holds it in
`_bootingPane`. `liveNames()` includes it; `drain()` kills it **synchronously**. A generation
counter distinguishes *"cancelled by us"* from *"genuinely failed"*, so a drain never inflates
`bootFailures` and `resume()` reliably starts a fresh boot.
### 5. Refills take no concurrency slot, and are serialized.
A refill boot deliberately does **not** take a `TuiSemaphore` slot: those slots bound concurrent
*turns* and belong to real requests, and charging a background pre-boot against them would let
the pool starve the traffic it exists to speed up. It cannot leak a slot either, since it never
holds one. Boots are **serialized** (one at a time): two cold boots racing an in-flight turn were
observed to overrun even the generous pool readiness cap. A genuinely failed boot does **not**
re-kick the chain (backoff — a broken `claude` must not respawn forever).
Background boots get a more generous readiness cap (`POOL_BOOT_MS` = 5 × `BOOT_MS`): `BOOT_MS` is
tight because a *client* is blocked on it, which is not true of a pre-boot. Slow ≠ broken.
---
## Consequences
### Cost — standing processes, paid whether or not a request arrives
**A warm pane is a live idle `claude` process.** Peak process count is
`OCP_TUI_POOL_SIZE` + `OCP_TUI_MAX_CONCURRENT` + 1 (booting replacement). This is the whole
reason the pool is **default-off**: an operator must opt into holding processes for traffic that
may never come. Size is clamped to `POOL_MAX_SIZE` = 4; an unparseable value **disables** the
pool rather than guessing.
Panes carry a 10-minute TTL and are health-checked at hand-out; a dead or degraded pane becomes
a **miss** (cold path), never a hung turn.
### Benefit
Measured end-to-end through a real OCP instance (Sonnet 4.6, `--effort low`):
**p50 10.17 s (n=6, pool off) → 6.00 s (n=12 warm hits) — 4.2 s / 41%.**
### The floor is unchanged
The pool does not touch the **~6 s TTFT floor** documented in the latency plan (claude always
prefills the full Claude Code system prompt). TUI mode remains unsuitable for interactive /
real-time consumers; it is for batch and background work. This ADR does not change that
conclusion.
### Observability
`/health`'s `tui` block gains a `pool` sub-object (`null` when off): `size`, `warm`, `booting`,
`model`, `hits`, `misses`, `boots`, `bootFailures`, `cancelled`, `dropped`. A climbing
`bootFailures` means panes are not reaching their input bar — the pool then degrades safely to
the cold path, but latency reverts to the un-pooled numbers. A steadily climbing `dropped` is
**normal** (the 15-min sweep drains and re-boots the pool on every tick, by design — see
Decision 3).
### ALIGNMENT authorization
- **Class B / OCP-owned.** The warm pool is process management around the `claude` CLI — the
same category as the existing tmux session lifecycle and the defunct-session reaper it extends.
**`cli.js` does not perform this operation, and no `cli.js` citation applies**; the authority
is ADR 0007 (which owns the TUI spawn machinery) plus this ADR. This is `ALIGNMENT.md` Rule 2's
Class B citation requirement, discharged explicitly rather than by silence.
- **The `/health` extension** adds sub-fields to the `tui` block. That block is **owned by ADR
0007** and post-dates ADR 0006's v3.16.4 grandfather snapshot, so it is not part of the frozen
B.2 inventory. The change is additive — every pre-existing `/health` field keeps a
byte-identical value, and `pool` is `null` unless the operator opts in — which is the
behaviour-preserving bar ADR 0006 sets. This ADR records that authorization.
- **No spawn argument changed.** `buildTuiCmd` is byte-identical; the pool calls it with the same
arguments. Banner-verified on live pooled panes: `· Claude Max`, never `API Usage Billing`
(the `--bare` trap documented in the latency plan).
### What a future contributor must not undo
- **Do not let a pane serve a second turn** (or `/clear`-and-reuse one) without first adding
user-line scoping to `lib/tui/transcript.mjs`. That is a cross-request text leak, not a perf
tweak. See Decision 1.
- **Do not remove the drain-before-sweep.** It is what keeps `kill-server` zombie reaping alive.
See Decision 3.
- **Do not go back to counting in-flight boots.** The pool must be able to *name* a session that
exists but has not finished booting. See Decision 4.
+2
View File
@@ -23,6 +23,8 @@ New ADRs increment from the highest existing number. Filenames are
| [0004](0004-openclaw-auto-sync.md) | OpenClaw Auto-Sync | Why `scripts/sync-openclaw.mjs` runs on `ocp update`, what its scope boundary is (writes only `models.providers["claude-local"].models` and `agents.defaults.models["claude-local/*"]`), and the idempotency contract. |
| [0005](0005-no-multi-provider.md) | No Multi-Provider | Why OCP stays single-provider (Anthropic-via-cli.js) and does not extend to OpenAI / Gemini / OpenRouter. Cost estimate: ~7 weeks for a v1 that buys neither moat nor commercial readiness. Separate commercial work starts in a separate repo. |
| [0006](0006-openai-shim-scope.md) | OpenAI Shim Scope | The Class A / Class B taxonomy. Class A endpoints (`cli.js`-mirror) keep Rules 15 verbatim; Class B endpoints (OCP-owned compatibility surface — `/v1/chat/completions`, `/v1/models`, admin endpoints) are anchored to OpenAI's spec (B.1) or to an authorizing ADR (B.2). Triggered by PR #99 (external `response_format` honoring). Grandfathers the existing B.2 inventory at v3.16.4. |
| [0007](0007-tui-interactive-mode.md) | TUI Interactive Mode | Why TUI-mode spawns an interactive `claude` in a tmux pane (no `-p`) to reach the **subscription** billing pool (`cc_entrypoint=cli`) rather than the metered Agent SDK pool. Owns the TUI spawn machinery: entrypoint labeling, credential-isolated home, MCP hard-disable, session namespace + defunct-session reaping, the independent concurrency bound, and the `/health` `tui` block. **Single-user only** — hard FATAL on multi-user configs. |
| [0008](0008-tui-warm-pane-pool.md) | TUI Warm Pane Pool | Why `OCP_TUI_POOL_SIZE` pre-boots **single-use** `claude` panes (one turn each, own `--session-id`) — and why reuse is forbidden (`transcript.mjs` returns the last assistant entry in the file, so a reused session leaks the earlier turn's text). Measured 41% end-to-end. Defines the pool↔reaper invariant (exemption by exact name from a live registry; drain before every sweep so `kill-server` zombie reaping survives) and the standing idle-process cost. Extends ADR 0007. |
## When to write a new ADR
+236
View File
@@ -0,0 +1,236 @@
# TUI-mode latency: measured floor, and the four things worth fixing
**Date**: 2026-07-13
**Status**: findings + backlog. **Superseded in part** — see the dated update boxes below.
Item #1 shipped ([#156](https://github.com/dtzp555-max/ocp/pull/156)); item #2 is **dead**
([`streaming-spike.md`](streaming-spike.md)); item #4 measured, **no effect**; item #3 stands.
**Measured on**: Mac mini / macOS 26.5.2 / Claude Code **v2.1.207** / Sonnet 5 / Claude Max subscription / **real-home mode** (no `CLAUDE_CODE_OAUTH_TOKEN`, no `OCP_TUI_HOME` in the service env)
**Evidence**: [`measurements.jsonl`](measurements.jsonl) — **n=15** (3 configs × 5) · banner captures [`billing-banner.txt`](billing-banner.txt) · harness [`floor.sh`](floor.sh)
## Why this exists
An external consumer (the 知音 AI project) benchmarked OCP's prompt path and measured
**TTFT p50 ≈ 3032 s**, and excluded OCP as a backend on that basis. That number is real,
but it is *not* the model being slow — this document decomposes where the 30 seconds
actually go, and what OCP can do about it.
**The harness deliberately does not go through OCP.** It spawns `tmux` + `claude` directly
(session prefix `zhiyin-floor-`, never `ocp-tui-*`) and polls `tmux capture-pane` for
incremental render, so it measures the **true first-token time** of the underlying
subscription path — the floor OCP could reach if it were perfect.
---
## Measurements
All rows in [`measurements.jsonl`](measurements.jsonl); every number below is recomputable from it.
| Config | n | boot→input-ready (median) | **TTFT (median)** | TTFT range | full answer (median) |
|---|---|---|---|---|---|
| baseline (inherits global `effortLevel: xhigh`) | 5 | 1.07 s | **10.35 s** | 8.32 17.19 s | 11.32 s |
| **`--effort low`** | 5 | 1.03 s | **6.17 s** | **5.87 6.44 s** | 9.98 s |
| `--bare` | 5 | 0.44 s | **no answer at all** (5/5 `ttft_ms: -1`) | — | — |
> **Not from this harness**: the direct Anthropic API reference figure (TTFT 0.841.64 s, n=2)
> comes from the 知音 AI project's own smoke test, not from `measurements.jsonl`. It is quoted
> only to size the gap; do not look for it in the evidence file.
### Where the 30 seconds go
```
~1.0 s spawn → claude's input bar is ready ← NOT the bottleneck
~6-10 s true TTFT (first token rendered in the pane)
~20 s ████ waiting for the whole turn to finish ████ ← this is the 30s
```
`runTuiTurn` blocks on the native transcript until a terminal event (`lib/tui/session.mjs`
"Block on the native transcript … until terminal"; `readTuiTranscript` in
`lib/tui/transcript.mjs`; ADR 0007 step 4) — i.e. it waits for the **entire turn** to complete
before returning anything. There is no streaming path. The ~20 s delta between this harness's
real TTFT and OCP's reported 3032 s is exactly that.
> **⚠️ 2026-07-13 correction — this decomposition attributes the ~20 s to the wrong thing.** It was
> inferred from the external 3032 s report, never measured *through* OCP. It has since been measured
> through a real OCP instance (TUI mode, `claude-sonnet-4-6`, the same ~1850-token prompt, n=5):
> **median 11.30 s** before [#156](https://github.com/dtzp555-max/ocp/pull/156), **9.55 s** after.
> Same-turn decomposition (baseline row `i=5`): **11.563 s** wall through OCP vs `turn_duration:
> 7.319 s` of CLI-internal time on that same turn → **OCP's own overhead ≈ 4.2 s** (n=1), **not
> ~20 s**. The rest of any larger number is the model *generating a long answer*,
> which the blocking wait does not cause and streaming would not shorten — it would only move the
> first byte earlier. The 3032 s figure therefore reflects a much longer output (and/or the
> then-inherited `xhigh` effort), not 20 s of OCP dead time. See
> [`streaming-spike.md`](streaming-spike.md) § "What streaming would have bought".
---
## ⚠️ Blocking constraint: `--bare` silently drops you off the subscription pool
Captured live ([`billing-banner.txt`](billing-banner.txt)) — the startup banner is the **only**
reliable indicator:
```
[] | Sonnet 5 with xhigh effort · Claude Max
[--effort low] | Sonnet 5 with low effort · Claude Max
[--bare] | Sonnet 5 with xhigh effort · API Usage Billing ← ❌
```
`--bare` ("skip hooks, LSP, plugin…") **also skips the subscription-credential resolution
path**. It really does cut boot to 0.430.45 s — but you are no longer on the subscription,
which defeats the entire purpose of TUI mode (ADR 0007 exists solely to reach the
subscription pool).
**The failure is silent.** All 5 `--bare` samples reached input-ready (boot 0.430.45 s), were
sent the prompt, and then produced **no answer at all** — 60 s timeout, no error, no crash, the
pane simply never rendered a token (the API-billing account had no credit balance). Nothing in
the transcript or the exit status reveals this.
**Anyone changing spawn flags must diff the banner line before and after.**
---
## Backlog — four items, ranked by value ÷ effort
### 1. Pass `--effort` explicitly on spawn — **do this first**
`buildTuiCmd` (`lib/tui/session.mjs`) does not pass `--effort``grep -rn -- "--effort\|effortLevel" lib/ server.mjs`
returns zero hits. What the pane's `claude` ends up using therefore depends on **which HOME mode
`resolveTuiHome()` picked**:
| mode | HOME | effort the pane gets |
|---|---|---|
| **real-home** (legacy default — *current* service config: no `CLAUDE_CODE_OAUTH_TOKEN`, no `OCP_TUI_HOME`) | `~` | **inherits the operator's `~/.claude/settings.json` → `effortLevel: xhigh` on this host** |
| env-token scratch (`CLAUDE_CODE_OAUTH_TOKEN` set — the direction #146/#150 pushed) | `~/.ocp-tui/home` | that settings.json contains only `permissions.additionalDirectories`; `prepareTuiHome()` never writes `effortLevel`**claude's built-in default** |
**Scope note**: TUI mode is currently *off* on this host (`CLAUDE_TUI_MODE=false`; `/health`
`"tui": {"enabled": false}`), so live traffic takes the `-p` path today. The statement below is
about what happens **when TUI mode is enabled**.
On the current HOME config, **every TUI request would run extended thinking** — pure waste
for the typical "generate this JSON" request, and it makes latency depend on an unrelated global
setting the operator may have changed for their own interactive use. And the mode split means
the effort level silently changes if the operator ever switches to env-token mode.
**Passing `--effort` explicitly fixes both problems at once.**
- **Effect (real-home, measured)**: TTFT p50 **10.35 s → 6.17 s (40 %)**, and the spread
collapses from 8.3217.19 s to **5.876.44 s**. For a proxy, the variance reduction matters
more than the median.
- **Cost**: one flag. Suggested: a new `OCP_TUI_EFFORT` env var (default `low`), documented in
README § "Environment Variables" per `release_kit.new_feature_doc_expectations`.
- **Risk**: none — banner confirms it stays on `Claude Max` (see `billing-banner.txt`).
- ⚠️ Do **not** reach for `--bare` to shave boot: see above.
### 2. Real streaming instead of blocking on turn-terminal — **ACHIEVABLE → [`streaming-spike.md`](streaming-spike.md)**
> **2026-07-13 update — the prereq spike was run. The answer is YES, but not from either source this
> item guessed at.** (a) The transcript grows at *event* granularity (the whole answer lands in one
> line, ~0.3 s before terminal) — dead. (b) The pane is a **rendered** view whose `capture-pane` text
> no longer contains the answer's source bytes (`## `, `**`, code fences are gone) — dead, and worse
> than "lossy": it is *not the model's text*. **But there is a third source neither this backlog nor
> the first spike considered: `claude` fires a `MessageDisplay` hook carrying incremental,
> byte-faithful `delta`s of the raw reply.** Verified live on a plain interactive TUI spawn (no `-p`),
> banner `· Claude Max`: 7 fires spread across generation, `concat(deltas) === T` **byte-exactly**
> (579 == 579), `T.startsWith(S)` true at every step, `## ` / `**` / ```` ```javascript ```` all
> present in the deltas. Granularity is block-level (~57 chunks/answer), not token-level — plenty for
> SSE. **Build it.**
>
> ⚠️ Two corrections to this item as written: the **"~20 s" is wrong** (inferred from an external
> report, never measured through OCP — the same-turn decomposition puts OCP's own overhead at **~4 s**,
> n=1), and **streaming moves the first byte, not the last** — so a consumer needing the *complete*
> answer (the JSON-card case that motivated this) gains **nothing** from it. Build it for
> progressively-rendering consumers, not as a throughput win.
>
> Full evidence + implementer caveats (the hook is `forceSyncExecution` — claude BLOCKS on it):
> **[`streaming-spike.md`](streaming-spike.md)**. Original framing preserved below.
Today `runTuiTurn` blocks on the transcript until the turn is *finished*. The pane is already
rendering tokens incrementally the whole time — this harness proves you can observe first token
at ~6 s by polling `tmux capture-pane`.
- **Effect**: turns a 30 s wall into a ~6 s TTFT with progressive output; enables SSE streaming
on the OCP endpoint instead of a single blob at the end.
- **Cost**: real work. Pane capture is ANSI/redraw-based and lossy for exact text (wrapping,
scrollback, spinner lines). Two candidate sources: (a) incremental reads of the transcript
JSONL, (b) `capture-pane` diffing with a stable start marker. (a) is much cleaner **if it
holds**.
- **Prereq spike (do this before designing anything)**: does the transcript JSONL grow *during*
a turn, or only at the end? If only at the end, (a) is dead and you are stuck with (b).
### 3. Warm pane pool — ~1 s
Every request spawns a fresh tmux session + `claude` (`randomUUID()` + `new-session`, then
`kill-session` in `finally`; `grep -rn "pool\|warm\|reuse" lib/tui/*.mjs` → zero hits). Boot to
input-ready is ~1.0 s, paid on every request. A pool of pre-booted panes (single-use, replaced in
the background) amortizes it to zero for any workload below the pool refill rate.
- **Effect**: 1.0 s.
- **Cost**: moderate; interacts with the session reaper and the per-port prefix scoping added in
#148 — pooled panes must not look like zombies to the sweep.
- Lower priority than #1 and #2: it is the smallest slice.
### 4. Trim the prefill — ~~probably not worth it~~ **MEASURED: no detectable benefit. Do not adopt.**
> **2026-07-13 update.** `--exclude-dynamic-system-prompt-sections` was measured with the same
> harness (`floor.sh`, n=5, Sonnet 5, on top of `--effort low`): **TTFT median 6.39 s**
> (5.8710.54 s) vs **6.17 s** (5.876.44 s) for `--effort low` alone — i.e. **0.22 s worse, inside
> the noise band**, with one worse outlier; dropping that outlier does not change the verdict. n=5
> cannot prove "zero", only "no benefit detectable above noise" — but there is also a **mechanistic**
> reason not to expect one: `--help` says the flag *"Improves cross-user prompt-cache **reuse**"*, and
> **OCP is single-user** — there is no cross-user cache to share, so the flag has nothing to buy here.
> The banner stayed on `· Claude Max` (no billing-pool drop), but there is no win to bank. The ~6 s
> floor stands as stated below. Raw rows: [`prefill-spike-measurements.jsonl`](prefill-spike-measurements.jsonl).
After #1#3, the floor is **~6 s**, and it does not go lower. `claude` always injects the full
Claude Code system prompt + tool definitions (thousands to tens of thousands of prefill tokens)
regardless of what you ask it. `--exclude-dynamic-system-prompt-sections` exists and may shave
some of it — **unmeasured**; worth one spike, but do not expect to reach the direct API's
~1 s.
**Consequence to accept, and to state in the README**: even fully optimized, TUI mode has a
**~6 s TTFT floor**, so it cannot serve real-time / interactive-latency consumers. It remains
appropriate for batch, background, and cost-insensitive-latency use. The 知音 AI project
excluded it on this basis (their prompt-latency budget is 24 s) *independently* of the ToS
question already documented in the README.
---
## Reproduction
```bash
# harness never touches OCP's :3456 service or ocp-tui-* sessions, and never kill-server
bash docs/plans/2026-07-13-tui-latency/floor.sh 5 # baseline
TAG=effort-low EXTRA_ARGS="--effort low" bash .../floor.sh 5 # 40 %
TAG=bare EXTRA_ARGS="--bare" bash .../floor.sh 5 # the trap
# billing-pool check for ANY spawn-flag change — the banner is the only source of truth
tmux new-session -d -s probe -x 200 -y 50 -c "$HOME" \
"claude --model claude-sonnet-5 --session-id $(uuidgen) <your-flags-here>"
sleep 6; tmux capture-pane -p -t probe | grep -E "Claude Max|API Usage Billing"
tmux kill-session -t probe
```
## Interaction with OCP while the harness runs
- **Kill direction is safe both ways**: `reapStaleTuiSessions()` only `kill-session`s names
matching `ocp-tui-<port>-`, which `zhiyin-floor-*` never matches; and the harness only
`kill-session`s its own single session — it contains **no `kill-server`**.
- **One benign interaction** (only when TUI mode is enabled — the reap tick is itself gated on
`TUI_MODE`): OCP's periodic `kill-server` (zombie reaping) is gated on
`othersRemain`*any* foreign-prefixed tmux session suppresses it. So while the harness is
running, that sweep is skipped. This is the coexistence guard working as designed; it resumes
on the next tick.
## Harness caveats (stated so the numbers are not over-trusted)
- **n=5 per config**, single host, single model (Sonnet 5), single prompt size (~1850 tokens).
Enough to separate 6 s from 10 s from 30 s; **not** enough for a p95.
- TTFT is "marker visible in `capture-pane`", which includes tmux render latency (small, but
nonzero) — it is an upper bound on the true first-token time.
- **The harness's readiness marker is not OCP's.** `floor.sh` waits for `│ >||Try "`; OCP's
`tuiInputReady()` matches `/\? for shortcuts/`. These are different events, so the ~1.0 s
boot figure is **not** directly comparable to OCP's `BOOT_MS` gate (default cap 4000 ms). It
does not affect the conclusions (1 s ≪ 6 s TTFT), but it is not apples-to-apples.
- The first version of this harness reported TTFT **0.08 s** — a false positive: the prompt
literally contained the marker string it was grepping for, so the match fired the instant the
prompt was pasted. Fixed by describing the marker instead of spelling it. **The script exited 0
and "successfully" produced 5 samples both times** — exit status proves nothing here.
@@ -0,0 +1,3 @@
[] | ▝▜█████▛▘ Sonnet 5 with xhigh effort · Claude Max
[--effort low] | ▝▜█████▛▘ Sonnet 5 with low effort · Claude Max
[--bare] | ▝▜█████▛▘ Sonnet 5 with xhigh effort · API Usage Billing
+128
View File
@@ -0,0 +1,128 @@
#!/usr/bin/env bash
# OCP TUI-mode latency floor harness — see README.md in this directory.
#
# 目的:回答一个问题——如果把 OCP 现有的两个已知开销砍掉
# (a) 每请求 spawn + boot(可用预热进程池消除)
# (b) 假流式(等 turn_duration 才返回,可用增量读 pane 消除)
# 之后,订阅池路径的**真实 TTFT 地板**是多少?
#
# 判据:地板 ≤ 4s → OCP 作为"省钱选项"可行;> 8s → 死透,不再讨论。
#
# 红线:
# - 不经过生产 OCP 服务(:3456)—— 直接起 tmux+claudeOCP 进程零干扰
# - tmux session 前缀用 zhiyin-floor-**不是** ocp-tui-),避免被 OCP 的
# reaper 当成自己的会话杀掉,也避免我们杀到它的
# - 用 real HOME(凭据)—— scratch HOME + symlink 凭据会 fork OAuth 导致 401
# (见跨机记忆 tui_scratch_home_credential_fork
set -uo pipefail
N=${1:-5}
MODEL=${MODEL:-claude-sonnet-5}
EXTRA_ARGS=${EXTRA_ARGS:-} # 额外 CLI 参数(如 --effort low --bare
TAG=${TAG:-baseline}
OUT=${OUT:-$(dirname "$0")/measurements.jsonl}
PROMPT_FILE=$(mktemp)
PREFIX="zhiyin-floor"
mkdir -p "$(dirname "$OUT")"
# ── 构造提示:~2000 token 的假会议转写 + 明确的起始标记 ────────────────
# 单行(多行会在 tmux send-keys 时提前触发 Enter
build_prompt() {
local seg="Speaker A said the quarterly pipeline is tracking behind plan and the enterprise segment needs a different motion. Speaker B replied that the current onboarding flow loses roughly a third of trial accounts before the first integration is complete. They debated whether the fix belongs in product or in customer success. "
local body=""
for _ in $(seq 1 22); do body+="$seg"; done
printf '%s' "You are a real-time meeting copilot. Meeting transcript so far: $body --- Task: produce ONE prompt card as compact JSON with keys: points (array of 3 short Chinese bullet points), keyline (one English sentence the user can read aloud). IMPORTANT: your reply MUST begin with three hash characters immediately followed by the uppercase word CARD (no space between them), then the JSON. No preamble, no markdown fences." > "$PROMPT_FILE"
}
build_prompt
PROMPT_CHARS=$(wc -c < "$PROMPT_FILE" | tr -d ' ')
now_ms() { python3 -c 'import time;print(int(time.time()*1000))'; }
echo "配置: $TAG 参数: [$EXTRA_ARGS]"
echo "模型: $MODEL 样本: $N 提示长度: ${PROMPT_CHARS} chars (≈$((PROMPT_CHARS/4)) token)"
echo "输出: $OUT"
echo
for i in $(seq 1 "$N"); do
SESS="${PREFIX}-$$-$i"
SID=$(uuidgen)
# ── 冷启动:spawn + 等输入框就绪 ─────────────────────────────────
T_SPAWN=$(now_ms)
tmux new-session -d -s "$SESS" -x 200 -y 50 \
-e CLAUDE_CODE_DISABLE_CLAUDE_MDS=1 \
-e CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 \
-c "$HOME" \
"claude --model $MODEL --session-id $SID --strict-mcp-config --disallowedTools 'mcp__*' $EXTRA_ARGS" 2>/dev/null
if [ $? -ne 0 ]; then echo "[$i] tmux spawn 失败,跳过"; continue; fi
# 轮询输入框就绪(claude TUI 的输入提示符)
READY=0
for _ in $(seq 1 150); do # 上限 15s
PANE=$(tmux capture-pane -p -t "$SESS" 2>/dev/null || true)
if grep -qE '│ >||Try "' <<<"$PANE"; then READY=1; break; fi
sleep 0.1
done
T_READY=$(now_ms)
BOOT_MS=$((T_READY - T_SPAWN))
if [ "$READY" -ne 1 ]; then
echo "[$i] 启动超时(${BOOT_MS}ms),pane 末 3 行:"
tmux capture-pane -p -t "$SESS" 2>/dev/null | tail -3 | sed 's/^/ /'
tmux kill-session -t "$SESS" 2>/dev/null
continue
fi
# ── 热态:粘提示 → 回车 → 量首 token ─────────────────────────────
tmux send-keys -t "$SESS" -l "$(cat "$PROMPT_FILE")" 2>/dev/null
sleep 0.4 # 让粘贴落地(OCP 用 400ms 轮询粒度)
T0=$(now_ms)
tmux send-keys -t "$SESS" Enter 2>/dev/null
TTFT_MS=-1
for _ in $(seq 1 600); do # 上限 60s
if tmux capture-pane -p -t "$SESS" 2>/dev/null | grep -q '###CARD'; then
TTFT_MS=$(( $(now_ms) - T0 )); break
fi
sleep 0.1
done
# ── 完整回答:pane 连续 2s 不再变化 ──────────────────────────────
COMPLETE_MS=-1
if [ "$TTFT_MS" -ge 0 ]; then
LAST=""; STABLE=0
for _ in $(seq 1 900); do # 上限 90s
CUR=$(tmux capture-pane -p -t "$SESS" 2>/dev/null | cksum)
if [ "$CUR" = "$LAST" ]; then
STABLE=$((STABLE+1))
[ "$STABLE" -ge 20 ] && { COMPLETE_MS=$(( $(now_ms) - T0 - 2000 )); break; }
else
STABLE=0; LAST="$CUR"
fi
sleep 0.1
done
fi
printf '{"i":%d,"tag":"%s","model":"%s","extra_args":"%s","prompt_chars":%s,"boot_ms":%d,"ttft_ms":%d,"complete_ms":%d}\n' \
"$i" "$TAG" "$MODEL" "$EXTRA_ARGS" "$PROMPT_CHARS" "$BOOT_MS" "$TTFT_MS" "$COMPLETE_MS" | tee -a "$OUT"
tmux kill-session -t "$SESS" 2>/dev/null
sleep 1
done
rm -f "$PROMPT_FILE"
echo
echo "=== 汇总 ==="
python3 - "$OUT" <<'EOF'
import json,sys,statistics
rows=[json.loads(l) for l in open(sys.argv[1]) if l.strip()]
ok=[r for r in rows if r['ttft_ms']>=0]
if not ok: print("无有效样本"); sys.exit()
def s(k):
v=[r[k] for r in ok if r[k]>=0]
return f"n={len(v)} 中位={statistics.median(v)/1000:.2f}s 最小={min(v)/1000:.2f}s 最大={max(v)/1000:.2f}s" if v else "无"
print(f" 冷启动 boot : {s('boot_ms')} ← 预热进程池可完全消除")
print(f" TTFT(首 token : {s('ttft_ms')} ★ 这就是地板")
print(f" 完整回答 : {s('complete_ms')}")
print(f"\n 失败样本: {len(rows)-len(ok)}/{len(rows)}")
EOF
@@ -0,0 +1,15 @@
{"i": 1, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1077, "ttft_ms": 6172, "complete_ms": 9929}
{"i": 2, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1026, "ttft_ms": 6160, "complete_ms": 9996}
{"i": 3, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1010, "ttft_ms": 6437, "complete_ms": 9977}
{"i": 4, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1033, "ttft_ms": 5872, "complete_ms": 9944}
{"i": 5, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1154, "ttft_ms": 6387, "complete_ms": 9993}
{"i":1,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":1300,"ttft_ms":8321,"complete_ms":9939}
{"i":2,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":1070,"ttft_ms":10347,"complete_ms":11320}
{"i":3,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":911,"ttft_ms":13061,"complete_ms":15163}
{"i":4,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":1441,"ttft_ms":9981,"complete_ms":11066}
{"i":5,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":1036,"ttft_ms":17189,"complete_ms":17985}
{"i":1,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":429,"ttft_ms":-1,"complete_ms":-1}
{"i":2,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":437,"ttft_ms":-1,"complete_ms":-1}
{"i":3,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":444,"ttft_ms":-1,"complete_ms":-1}
{"i":4,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":446,"ttft_ms":-1,"complete_ms":-1}
{"i":5,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":441,"ttft_ms":-1,"complete_ms":-1}
@@ -0,0 +1,7 @@
{"hook_event_name": "MessageDisplay", "index": 0, "final": false, "delta": "## Mutex\n\n"}
{"hook_event_name": "MessageDisplay", "index": 1, "final": false, "delta": "A **mutual exclusion lock** prevents concurrent access to a shared resource, ensuring only one thread runs the critical section at a time.\n\n"}
{"hook_event_name": "MessageDisplay", "index": 2, "final": false, "delta": "- Acquiring a locked mutex blocks the caller until the current holder releases it.\n"}
{"hook_event_name": "MessageDisplay", "index": 3, "final": false, "delta": "- Failing to release a mutex causes a deadlock, freezing all waiting threads.\n\n```javascript\nconst { Mutex } = require('async-mutex');\n\nconst mutex = new Mutex();\n"}
{"hook_event_name": "MessageDisplay", "index": 4, "final": false, "delta": "let counter = 0;\n\nasync function increment() {\n const release = await mutex.acquire();\n try {\n"}
{"hook_event_name": "MessageDisplay", "index": 5, "final": false, "delta": " counter++; // only one caller here at a time\n } finally {\n release();\n }\n}\n"}
{"hook_event_name": "MessageDisplay", "index": 6, "final": true, "delta": "```"}
@@ -0,0 +1,5 @@
{"i":1,"tag":"effort-low-exclude-dynamic","model":"claude-sonnet-5","extra_args":"--effort low --exclude-dynamic-system-prompt-sections","prompt_chars":7451,"boot_ms":934,"ttft_ms":5867,"complete_ms":9953}
{"i":2,"tag":"effort-low-exclude-dynamic","model":"claude-sonnet-5","extra_args":"--effort low --exclude-dynamic-system-prompt-sections","prompt_chars":7451,"boot_ms":1275,"ttft_ms":6388,"complete_ms":9874}
{"i":3,"tag":"effort-low-exclude-dynamic","model":"claude-sonnet-5","extra_args":"--effort low --exclude-dynamic-system-prompt-sections","prompt_chars":7451,"boot_ms":874,"ttft_ms":10537,"complete_ms":11782}
{"i":4,"tag":"effort-low-exclude-dynamic","model":"claude-sonnet-5","extra_args":"--effort low --exclude-dynamic-system-prompt-sections","prompt_chars":7451,"boot_ms":1170,"ttft_ms":6379,"complete_ms":9947}
{"i":5,"tag":"effort-low-exclude-dynamic","model":"claude-sonnet-5","extra_args":"--effort low --exclude-dynamic-system-prompt-sections","prompt_chars":7451,"boot_ms":1329,"ttft_ms":6443,"complete_ms":9884}
@@ -0,0 +1,256 @@
# Backlog #2 (real streaming): **achievable** — via the `MessageDisplay` hook
**Date**: 2026-07-13
**Status**: prereq-spike result. **Streaming IS achievable on the TUI path**, byte-faithfully, on the
subscription pool. Three obvious sources are dead ends; a fourth one works.
**Scope**: answers the prereq spike that [`README.md`](README.md) § "Backlog #2" demanded *before* any
streaming design:
> **Prereq spike (do this before designing anything)**: does the transcript JSONL grow *during* a
> turn, or only at the end? If only at the end, (a) is dead and you are stuck with (b).
The answer: **(a) is dead, (b) is dead — and you are not stuck with either.** The CLI exposes its own
streaming interface as a **hook**, which the backlog did not consider.
**Measured on**: Mac mini / Claude Code **v2.1.207** / Sonnet 4.6 + Sonnet 5 / Claude Max /
real-home mode. Every claim below is reproducible from the commands given.
> **Honesty note on how this document was produced.** Its first version concluded the exact opposite —
> "streaming is not achievable; the CLI exposes no byte-faithful incremental source" — and was **wrong**.
> An adversarial reviewer, commissioned specifically to *refute* it, found `MessageDisplay` on a second
> pass; its own first pass had enumerated the hook registry with a truncated grep (it reported 21
> events — there are **30**). Both the wrong conclusion and its refutation are preserved here, because
> "we checked, it's impossible" is the most expensive kind of claim to get wrong: it closes a door and
> nobody re-opens it.
---
## ✅ The source that works: the `MessageDisplay` hook
`claude` fires a **`MessageDisplay`** hook as it renders each block of the assistant's reply. The
payload carries the **raw markdown source** of an incremental `delta`, plus a monotonic `index` and a
`final` flag:
```json
{ "hook_event_name": "MessageDisplay",
"turn_id": "6cb31d21-…", "message_id": "84ab9832-…",
"index": 0, "final": false, "delta": "## Mutex\n\n" }
```
*(payload also carries `session_id`, `transcript_path`, `prompt_id`, `cwd`)*
Registered as an ordinary command hook via `--settings` on a **plain interactive TUI spawn** (no `-p`,
no `--bare`), `claude-sonnet-4-6`, `--effort low`. Banner verified:
`▝▜█████▛▘ Sonnet 4.6 with low effort · Claude Max`**subscription pool, not metered billing**.
One live turn — 7 fires, spread across generation:
```
index=0 final=false len= 10 '## Mutex\n\n'
index=1 final=false len= 140 'A **mutual exclusion lock** prevents concurrent access to a shar…'
index=2 final=false len= 83 '- Acquiring a locked mutex blocks the caller until the current h…'
index=3 final=false len= 163 '- Failing to release a mutex causes a deadlock, freezing all wai…'
index=4 final=false len= 96 'let counter = 0;\n\nasync function increment() {\n const release =…'
index=5 final=false len= 84 ' counter++; // only one caller here at a time\n } finally {\n …'
index=6 final=true len= 3 '```'
```
**Every invariant a proxy needs — all hold:**
| requirement | result |
|---|---|
| **byte-faithful** — deltas are the model's *source*, not the rendered pane | ✅ `## `, `**`, ```` ```javascript ```` all present in the deltas |
| **exactness** — `concat(deltas) === T` (the transcript-authoritative text) | ✅ **true**, 579 == 579 bytes |
| **prefix-stable** — `T.startsWith(concat(deltas[0..n]))` at every n | ✅ **true at all 7 steps** |
| **incremental** — arrives during generation, not at the end | ✅ 7 fires spread across the turn |
| **no `-p`** — stays out of the metered `sdk-cli` pool | ✅ plain interactive TUI |
| **subscription pool** | ✅ banner `· Claude Max` |
This is exactly the contract a streaming design needs: deltas forward straight into SSE
`delta.content` chunks, and the transcript's final text `T` stays a cheap end-of-turn assertion
(`concat === T`) instead of a reconciliation problem.
### Caveats for the implementer
- **Block-level granularity, not token-level** — the hook fires **once per rendered block** (roughly one
per paragraph / list item / code block), so the chunk count **scales with answer length**: 7 fires for a
~600-byte answer, **18 for a ~2 KB one**. Plenty for SSE (`delta.content` has no minimum size), but do
not promise token-by-token output, and do not hard-code any assumption about chunk count.
- **🔴 The sink MUST be keyed by `session_id` — this is live TODAY, not a future concern.**
`OCP_TUI_MAX_CONCURRENT` defaults to **2**, so **two `claude` processes already run concurrently**. One
hook command writing to one shared sink would **interleave deltas from two different turns into one
stream** — request A's client receiving request B's text, the worst failure a proxy can have, and one a
single-request test will never surface. The payload carries `session_id` (and `turn_id` / `message_id`),
so demux is easy: derive the sink path from `session_id` (`<dir>/<session_id>.jsonl`) and read only your
own turn's file. This *also* keeps the design **warm-pool compatible**, because a pre-booted pane's
session-id is fixed at boot — one static hook script serves every pane. **Test it with ≥2 concurrent
streaming requests carrying distinguishable prompts and assert zero cross-contamination.**
- **⚠️ `forceSyncExecution: true` in the hook's source — `claude` BLOCKS on the hook.** A slow hook
adds latency to *every* delta. The hook must write and exit immediately (e.g. write to a FIFO / unix
socket that OCP reads; never work inline). **Measure the added per-delta latency.**
- **Thinking blocks appear to be excluded — but this is NOT yet stress-tested. Verify before shipping.**
The exclusion is inferred from `content.map(c => c.type === "text" ? c.text : "")` — but that snippet is
from the **`final:true`** call site, not the incremental one. Four live turns (incl. two at `--effort
high`) showed no thinking text in any delta and `concat === T` held — **but each transcript's thinking
block was empty (`thinking:""`, 0 chars)**, so the exclusion was never actually stressed. **The failure
mode is severe**: if thinking deltas *do* fire on some config (Opus, `xhigh`), `concat(deltas) !== T`
**and OCP streams the model's private reasoning to the caller**. The end-of-turn `concat === T` assertion
would *detect* that but **cannot prevent** it — SSE deltas cannot be un-sent. **Before shipping, run a
turn on a model+effort that produces substantive thinking** (a hard reasoning prompt on Opus / `xhigh`)
and confirm both (a) no thinking text in any delta and (b) `concat === T` still holds.
- OCP already owns the spawn (isolated HOME, its own flags), so injecting `--settings` with a
`MessageDisplay` hook sits inside the existing architecture.
- **`ALIGNMENT.md`**: this consumes `claude`'s **own** hook surface as emitted — forwarding, not
inventing. Not a new endpoint, not a fabricated protocol. (Class B / ADR 0007 — the TUI spawn is
OCP-owned; no `cli.js` citation applies.)
### Reproduce in 60 seconds
```bash
# hook script: append the payload (arrives on stdin) and exit immediately
printf '#!/bin/bash\ncat >> "$MD_LOG"; printf "\\n" >> "$MD_LOG"; exit 0\n' > /tmp/h.sh && chmod +x /tmp/h.sh
echo '{"hooks":{"MessageDisplay":[{"hooks":[{"type":"command","command":"MD_LOG=/tmp/deltas.jsonl /tmp/h.sh"}]}]}}' > /tmp/s.json
# plain interactive claude in tmux (prefix NOT ocp-tui-*, and never kill-server)
tmux new-session -d -s md-probe -x 220 -y 50 \
"claude --model claude-sonnet-4-6 --effort low --session-id $(uuidgen) --settings /tmp/s.json"
# …wait for '? for shortcuts', paste a markdown-producing prompt, press Enter…
jq -r '"\(.index) \(.final) \(.delta|@json)"' /tmp/deltas.jsonl # incremental raw-markdown deltas
# then assert: concat(deltas) == extractLatestAssistantText(<transcript>.jsonl)
```
---
## The three dead ends (still worth knowing — they say what NOT to build)
### (a) Incremental transcript reads — **dead: event granularity, not token granularity**
The transcript JSONL *does* grow during a turn, but one **whole event at a time**; the assistant's text
event is written as **one complete line**, appearing only ~0.3 s before the terminal `turn_duration`.
Observed (session `efd5b161`, `turn_duration: 7319 ms`):
```
#6 t+0.0s type=user (the prompt)
#15 t+4.7s type=assistant blocks=thinking
#16 t+7.0s type=assistant blocks=text ← the ENTIRE answer, in one line
#21 t+7.3s type=system subtype=turn_duration ← terminal
```
Cross-checked at **20 ms polling + `fs.watch`** (25× finer): a partial line **never touches disk** —
one write, `+1` line, carrying the complete answer. Also forced with the undocumented
`CLAUDE_CODE_INCLUDE_PARTIAL_MESSAGES=1`: still 1 assistant event, 0 partials (interactive mode has no
stream-json *sink* for it to write to).
**The transcript is still needed** — as the terminal-turn signal, as the authoritative `concat === T`
check, and as the input to the existing honesty gates (auth-banner detection, `truncated`). It is just
not the *streaming* source.
### (b) `tmux capture-pane` diffing — **dead: the pane is a RENDERED view, not the text**
The backlog expected to fall back to this, calling it "lossy … (wrapping, scrollback, spinner lines)".
The loss is far worse than formatting noise: **the pane does not contain the answer's source bytes at
all.** The TUI *renders* markdown, and `capture-pane -p` strips the ANSI that rendering produced.
Same turn, same lines:
```
TRANSCRIPT (authoritative T): PANE (capture-pane -p -J -S -500):
'## Semaphore' '⏺ Semaphore' ← heading marker gone
'' ''
'A **semaphore** is a synchro…' ' A semaphore is a synchro…' ← bold markers gone, indented
```
| token in the answer | in `T` | in the pane's answer region |
|---|---|---|
| `## ` (ATX heading) | yes | **no** — rendered as `` |
| `**` (bold markers) | yes | **no** — rendered to ANSI bold, then stripped by `-p` |
| ` ```javascript ` (fence + language) | yes | **no** — fence and language tag both gone |
| `- ` (list item) | yes | yes |
*(A literal `**` does appear elsewhere in the pane — in the **prompt echo**, because the prompt asked
for bold. Not in the answer.)*
**`capture-pane -e` (keeping the ANSI) does not rescue it — the inverse is provably non-unique.**
With `T` = ``"## Alpha\n\n**bravo**\n\n```javascript\nlet x=1;\n```"``:
```
⏺\e[39m \e[1mAlpha\n\n\e[0m \e[1mbravo\n\n\e[0m \e[34mlet\e[39m x=\e[32m1\e[39m;
```
`## Alpha` → **SGR 1 (bold)**. `**bravo**` → **SGR 1 (bold)**. *Identical ANSI* — an H2 and a bold span
are indistinguishable, never mind `**` vs `__`. The fence and its `javascript` tag are consumed by the
syntax highlighter into colours; recovering the tag would mean inverting a highlighter, and
`let x=1;` is valid in several languages.
So `T.startsWith(paneText)` is **false** — raw and indent-stripped, on essentially every markdown
answer. A proxy streaming pane text would be streaming **something the model did not say**. With
`MessageDisplay` available there is no reason to go near it.
### (c) `--debug-file` — **dead: it logs stream *timing*, never stream *content***
Worth stating precisely, because a casual check misleads in **both** directions here.
The default log level is `debug`, which **suppresses every `verbose` site**. Raise it and per-chunk
lines *do* appear, spread across generation:
```bash
CLAUDE_CODE_DEBUG_LOG_LEVEL=verbose claude --debug-file /tmp/d.log …
```
```
05:51:11.088 [VERBOSE] [shoji-engine] yield stream_event/- ← 16 of these, mid-turn,
05:51:11.537 [VERBOSE] [shoji-engine] yield stream_event/- over ~3.9 s of generation
05:51:15.192 [DEBUG] [shoji-engine] turn 1 end (usage in=575 out=255 api=6736ms stop=end_turn resultLen=857)
```
**But they carry no payload** — the format is `yield <type>/<subtype>`, a bare presence marker. Run with
no category filter (i.e. all categories) at verbose level: `content_block_delta` = **0**, `text_delta` =
**0**, `content_block_start` / `message_start` = **0**. The only byte-exact text in the log is the
end-of-turn `Stop` hook payload (`"last_assistant_message":"## Title\n\n**alpha bravo charlie**"`) —
transcript granularity. The log tells you **when** tokens arrive, never **what** they are. It is also
~2.7 MB per turn.
### Also checked, also not the answer
| candidate | outcome |
|---|---|
| `--output-format stream-json` (the one interface that emits `text_delta`) | **requires `--print`/`-p`** → `cc_entrypoint=sdk-cli` → the **metered** credit pool, which is exactly what TUI mode exists to avoid. Reproduced live. |
| `--input-format stream-json` | `Error: --input-format=stream-json requires output-format=stream-json` → same gate. |
| `CLAUDE_CODE_INCLUDE_PARTIAL_MESSAGES=1` (undocumented) | No stream-json sink in interactive mode → no partials. Banner stayed `· Claude Max`. |
| `sessionMirror` (undocumented) | Gated on `outputFormat === "stream-json"` → the `-p` family. |
| `--sdk-url` (hidden) | Forces stream-json + non-interactive → `sdk-cli`. *(inferred from the minified bundle; not banner-tested)* |
| `~/.claude/sessions/<pid>.json` | Registry metadata only (`{pid, sessionId, cwd, status, version, entrypoint:"cli", kind:"interactive"}`). No assistant text. *(Its `entrypoint:"cli"` incidentally confirms the TUI path stays on the subscription pool.)* |
| `~/.claude/history.jsonl` | User prompts only; the answer text is absent. |
| Asking the model to emit plain text (so the pane renders faithfully) | Would mean **mutating the caller's prompt** — a correctness violation for a proxy, and still not byte-faithful (wrapping + indent remain). Rejected. |
---
## Value: what streaming actually buys (read before building)
Streaming is *possible*. Whether it is *worth it* depends on the consumer, and the honest answer is
uncomfortable:
- **Streaming never makes the answer arrive sooner. It moves the *first* byte, not the *last*.** The
final token lands at the same wall-clock moment either way.
- So a consumer that must have the **complete** answer before it can act — e.g. one parsing a structured
JSON reply, **which is exactly the 知音 AI use case that motivated this entire investigation** — gains
**nothing at all**. Only a **progressively-rendering** consumer (a chat UI) gains.
And the number the backlog attached to this item was wrong:
- The backlog's "~20 s" was inferred from an external 3032 s report, **never measured through OCP**.
Measured through a real OCP instance (TUI mode, `claude-sonnet-4-6`, ~1850-token prompt, n=5):
**median 11.30 s** before [#156](https://github.com/dtzp555-max/ocp/pull/156), **9.55 s** after.
- **Same-turn decomposition** (baseline row `i=5`): **11.563 s** wall through OCP vs `turn_duration:
7.319 s` of CLI-internal time on that same turn → **OCP's own overhead ≈ 4.2 s** (n=1, baseline
`effort=high` config). *Caveats*: n=1; and `turn_duration` is the CLI's internal duration of an
**OCP-driven** turn, not a separate "native" baseline. Do **not** subtract this `effort=high` 7.3 s
from the `effort=low` 9.55 s median — a low-effort turn generates faster, so mixing them
*understates* the overhead.
- So OCP's own overhead is **single-digit seconds**, not ~20 s. The rest of any large number is the
model generating a long answer — which streaming hides but does not shorten.
**Recommendation**: build it — the contract is clean and the cost is small — but size the expectation
honestly. It is a *perceived-latency* feature for progressively-rendering consumers, not a throughput
win, and it does not move the **~6 s TTFT floor** ([`README.md`](README.md)) that rules TUI mode out for
interactive-latency consumers regardless.
+321
View File
@@ -0,0 +1,321 @@
import { rmSync } from "node:fs";
// TUI warm pane pool (docs/plans/2026-07-13-tui-latency backlog #3).
//
// WHAT IT IS: a small set of PRE-BOOTED `claude` panes, each already sitting at its
// input bar, so a request does not pay the cold boot. Opt-in: OCP_TUI_POOL_SIZE=0
// (default) disables it entirely and the request path is byte-for-byte today's.
//
// ── SINGLE-USE IS THE LOAD-BEARING RULE ─────────────────────────────────────
// A pooled pane serves EXACTLY ONE turn and is then killed and replaced in the
// background. Each pane carries its OWN fresh `--session-id`, fixed at boot, and the
// turn locates its transcript by that id. So OCP's one-session-per-request model is
// preserved: a session's transcript still holds exactly one logical exchange.
// That is what keeps lib/tui/transcript.mjs's extractLatestAssistantText (which returns
// the LAST text-bearing assistant entry in the whole file, not "text since the matching
// user line") correct — see the scoping note there. A pane MUST NEVER serve a second
// turn, and a session MUST NEVER be reset with /clear and reused: either would put two
// exchanges in one transcript and leak the earlier turn's text into the later turn's
// answer. Nothing here reuses a pane; keep it that way.
//
// ── WHY IT'S WORTH MORE THAN THE BOOT TIME ──────────────────────────────────
// Measured on this host (n=6 through OCP, Sonnet 4.6, --effort low): the cold path
// spends ~1.23 s reaching the input bar, but ALSO ~2.9 s inside the first turn beyond
// what claude itself reports as the turn duration — post-input-bar init that a pane
// which has been idle for a few seconds has already finished. A warm pane recovers both.
//
// ── COST (bounded, and paid whether or not a request arrives) ───────────────
// Each warm pane is a LIVE `claude` process (plus its tmux pane) sitting idle. Peak
// process count is (pool size) + (OCP_TUI_MAX_CONCURRENT in-flight turns) + (panes
// currently booting as replacements). Pool size is clamped to POOL_MAX_SIZE.
//
// Pure + injectable (bootPane / killPane / paneHealthy / now) so test-features.mjs can
// assert acquire / miss / refill / TTL / reaper-exemption with no tmux and no claude.
// Hard cap on OCP_TUI_POOL_SIZE. Each pane is an idle claude process; 4 is already a
// lot of resident memory on a small host (a Pi serving a family) for zero in-flight work.
export const POOL_MAX_SIZE = 4;
// A warm pane older than this is dropped on acquire rather than handed out. The periodic
// reap tick (server.mjs) drains the pool every 15 min anyway, so this only bites when
// that tick kept getting skipped because the TUI path was never idle. Guards against
// handing out a pane whose `claude` has been sitting so long it may have drifted
// (auto-compaction prompts, an idle-disconnect banner, an expired in-pane token).
export const POOL_MAX_AGE_MS = 10 * 60 * 1000;
// Clamp the operator-supplied size into [0, POOL_MAX_SIZE]. A garbage value disables the
// pool rather than guessing — an unparseable size must never silently boot 4 processes.
export function resolvePoolSize(raw) {
const n = parseInt(raw, 10);
if (!Number.isFinite(n) || n <= 0) return 0;
return Math.min(n, POOL_MAX_SIZE);
}
export class TuiPanePool {
// size: target number of warm panes (0 = disabled).
// maxAgeMs: per-pane TTL (see POOL_MAX_AGE_MS).
// mintPane: () => ({ sessionId, name }) — mints the identity of the NEXT pane. The POOL,
// not the boot function, owns this: the tmux session springs into existence the
// instant bootPane starts, so the pool must already know its NAME (see
// _bootingPane below). Deriving the name from the sessionId also makes `tmux ls`
// correlate to the transcript file.
// bootPane: async (model, {sessionId, name}) => { name, sessionId, model, bootedAt } —
// boots ONE pane under exactly that identity and resolves only once it is
// input-ready; throws if it never becomes ready.
// killPane: (name) => void — tmux kill-session. MUST be synchronous (see drain).
// paneHealthy:(name) => bool — pane still exists AND is still at its input bar.
constructor({ size, maxAgeMs = POOL_MAX_AGE_MS, mintPane, bootPane, killPane, paneHealthy, now = Date.now, log = () => {} }) {
this.size = Math.max(0, Math.min(parseInt(size, 10) || 0, POOL_MAX_SIZE));
// Fail fast at CONSTRUCTION, not at request time. refill() is called synchronously from
// the request path (runTuiTurn), so a missing collaborator would otherwise surface as a
// 500 on a live request instead of a loud error at boot.
if (this.size > 0) {
for (const [k, fn] of [["mintPane", mintPane], ["bootPane", bootPane], ["killPane", killPane], ["paneHealthy", paneHealthy]]) {
if (typeof fn !== "function") throw new TypeError(`TuiPanePool: ${k} must be a function`);
}
}
this.maxAgeMs = maxAgeMs;
this._mintPane = mintPane;
this._bootPane = bootPane;
this._killPane = killPane;
this._paneHealthy = paneHealthy;
this._now = now;
this._log = log;
this._panes = []; // warm, available panes: { name, sessionId, model, bootedAt }
// The pane currently BOOTING, BY NAME ({sessionId, name, model}) — or null.
//
// WHY A NAME AND NOT A COUNT (this is a fixed bug, don't regress it): bootTuiPane creates
// the tmux session SYNCHRONOUSLY and only THEN waits up to POOL_BOOT_MS (20 s) for the
// input bar. So for up to 20 s there is a LIVE pooled tmux session. When the pool tracked
// only a count, it could not NAME that session, so:
// - liveNames() could not spare it and the periodic reap sweep KILLED it (and
// kill-server'd on top), leaving the pool empty with nothing scheduled and firing the
// very tui_pool_boot_failed WARN operators are told to alert on; and
// - drain() could not kill it, so on shutdown it ORPHANED a live authenticated `claude`
// (the boot's .then that was supposed to clean up never runs — gracefulShutdown calls
// process.exit in the same tick).
// Both are fixed by holding the identity here, before the session exists.
this._bootingPane = null;
// Generation counter. Bumped whenever an in-flight boot is CANCELLED (drain / model
// switch). A boot compares the generation it started under against the current one:
// if they differ, its pane was already killed by us and its settle is inert — in
// particular a rejection is a CANCELLATION, not an operator-visible boot failure.
this._gen = 0;
this._paused = false; // true while drained; refill() is a no-op until resume()
this.warmModel = null; // the model the pool currently warms — learned from traffic (see acquire)
this.hits = 0; // requests served by a warm pane
this.misses = 0; // requests that fell back to the cold path
this.boots = 0; // panes successfully pre-booted
this.bootFailures = 0; // pre-boots that genuinely never reached the input bar
this.cancelled = 0; // in-flight boots WE killed (drain / model switch) — not failures
this.dropped = 0; // panes discarded unused (unhealthy / expired / wrong model / drained /
// cancelled — a cancelled in-flight boot also lands here via _drop)
}
get enabled() { return this.size > 0; }
get warm() { return this._panes.length; }
get booting() { return this._bootingPane ? 1 : 0; }
// The reaper's spare set: the EXACT names of every pane the pool currently owns and has NOT
// handed out — the warm ones AND the one currently booting (whose tmux session is already
// live; see _bootingPane). See the POOL/REAPER INVARIANT in lib/tui/session.mjs.
// Fail-safe by construction: a pane leaves this set the instant it is acquired, dropped, or
// cancelled, and if the pool is empty (or the process restarted) the set is empty — so an
// orphaned pooled pane looks exactly like any other stale session and IS reaped.
liveNames() {
const names = new Set(this._panes.map((p) => p.name));
if (this._bootingPane) names.add(this._bootingPane.name);
return names;
}
// Take a warm pane for `model`, or null (caller must fall back to the cold path — a MISS
// is always safe, never an error). Synchronous: paneHealthy is a cheap tmux capture.
//
// The pool warms the MOST RECENTLY REQUESTED model (`warmModel`). There is no boot-time
// pre-warm and no configured model: OCP cannot know which model the next caller wants, and
// pre-booting a process for a model nobody asks for is pure waste. Consequence, stated
// plainly: the FIRST request after start (and the first after a model switch) is always a
// MISS. The pool pays off for the steady repeat traffic it exists to serve.
acquire(model) {
if (!this.enabled) return null;
// Retarget on a model switch: --model is fixed at spawn, so panes for another model are
// useless. Drop them now (they are replaced by the next refill) rather than holding
// processes for a model that is no longer being asked for. This includes any pane
// currently BOOTING for the old model — its tmux session already exists, so leaving it to
// die on resolve would both hold a useless process and block the next refill (one boot at
// a time) for up to POOL_BOOT_MS.
if (model !== this.warmModel) {
for (const p of this._panes) { this._drop(p, "model_switch"); }
this._panes = [];
this._cancelBooting("model_switch");
this.warmModel = model;
}
while (this._panes.length) {
const p = this._panes.shift();
if (this._now() - p.bootedAt > this.maxAgeMs) { this._drop(p, "expired"); continue; }
if (!this._paneHealthy(p.name)) { this._drop(p, "unhealthy"); continue; }
this.hits++;
return p; // caller OWNS it now: it is out of the registry (so out of the spare set),
// and the caller's finally MUST kill it. Single-use — never returned here.
}
this.misses++;
return null;
}
// Bring the pool back up to `size` warm panes for `warmModel`. Fire-and-forget: never
// awaited on the request path and never throws into it.
//
// SLOT ACCOUNTING: a refill boot deliberately does NOT take a TuiSemaphore slot. Those
// slots bound concurrent *turns* (each up to the 120 s wallclock) and belong to real
// requests; charging a background pre-boot against them would let the pool starve the
// traffic it exists to speed up. It cannot leak a slot either, because it never holds one.
//
// SERIALIZED, ONE BOOT AT A TIME (and re-kicked on success until the pool is at target).
// An earlier version launched all `want` boots at once; live at size=2 that put two cold
// `claude` boots plus an in-flight turn on the CPU together, and a refill overran even the
// generous pool readiness cap (tui_pool_boot_failed). Booting sequentially keeps each boot
// near its uncontended ~1.2 s, bounds the CPU burst the pool can cause, and still has the
// replacement pane warm long before the next request arrives.
//
// A genuinely FAILED boot deliberately does NOT re-kick the chain — that is the backoff. A
// persistently failing boot (bad claude binary, no auth) would otherwise spin, respawning
// forever. The next natural trigger (the following request's refill, or the reap tick's
// resume) retries it. A CANCELLED boot is different: we killed it on purpose, nothing is
// wrong, and resume() is expected to start a fresh one immediately.
refill() {
if (!this.enabled || this._paused || !this.warmModel) return;
if (this._bootingPane) return; // one boot in flight at a time
if (this._panes.length >= this.size) return; // already at target
const model = this.warmModel;
const gen = this._gen;
// Mint the identity BEFORE booting: bootPane creates the tmux session synchronously, so
// the pool must be able to name (and therefore spare, and kill) it from this moment on.
const ident = this._mintPane();
this._bootingPane = { ...ident, model };
let enlisted = false;
Promise.resolve()
.then(() => this._bootPane(model, ident))
.then((pane) => {
// The world may have moved while we booted. If our generation was cancelled, kill the
// pane here rather than ASSUMING _cancelBooting already did.
//
// Why not just `return`: _cancelBooting kills by name, but the tmux session only EXISTS
// once _bootPane has actually run — and _bootPane is queued on a microtask (above). A
// caller that does refill() and then drain() in the SAME synchronous block would have
// _cancelBooting find nothing to kill (a no-op), bump the generation, and then this
// microtask would create the session, boot it fine, and — under a bare `return` — walk
// away from a LIVE authenticated `claude` that nothing owns. That is M1b in a new costume.
// No current call site does that, so this is defense-in-depth, not a live bug — but ADR
// 0008 and the reap-tick comment in server.mjs both explicitly contemplate a boot-time
// pre-warm, which is exactly the shape that would reach it.
//
// Killing an already-dead session is a harmless no-op (_drop swallows it), so this is
// idempotent whether or not _cancelBooting got there first.
if (gen !== this._gen) { this._drop(pane, "cancelled_late"); return; }
// Otherwise: still possible the pool filled or retargeted without a cancellation.
if (this._paused || model !== this.warmModel || this._panes.length >= this.size) {
this._drop(pane, "stale_boot");
return;
}
this._panes.push(pane);
this.boots++;
enlisted = true;
})
.catch((e) => {
// A rejection from a CANCELLED generation is not a fault: it is almost always
// "tui_pane_not_ready", thrown because WE killed the pane out from under the boot.
// Counting it as a bootFailure would fire the exact WARN operators are told to alert
// on, for a completely healthy drain. Stay silent — _cancelBooting already counted
// this as a cancellation, so do NOT count it again here.
if (gen !== this._gen) return;
this.bootFailures++;
this._log("warn", "tui_pool_boot_failed", { model, error: e && e.message });
})
.finally(() => {
// ONLY the current generation's boot owns the booting slot. A stale settle must not
// clear a slot that a newer boot (started by resume()) already holds.
if (gen === this._gen) this._bootingPane = null;
if (enlisted) this.refill(); // continue toward target, still one at a time
});
}
// Kill the in-flight boot's pane, SYNCHRONOUSLY, and invalidate its generation. Returns 1
// if there was one, else 0. The tmux session already exists (bootPane created it before it
// started waiting for readiness), so this is a real kill, not a cancellation flag.
_cancelBooting(reason) {
if (!this._bootingPane) return 0;
this._gen++; // the in-flight boot's settle is now inert
this._drop(this._bootingPane, reason); // synchronous kill-session
this._bootingPane = null;
this.cancelled++;
return 1;
}
// Kill every pane the pool owns — warm AND currently booting — and stop refilling. Returns
// how many were killed.
//
// Called (a) before the periodic reap sweep — reapStaleTuiSessions can only reap defunct
// `claude` zombies via kill-server, and kill-server is suppressed while any live pooled pane
// exists (including a booting one), so without this drain the pool would permanently disable
// zombie reaping; and (b) on graceful shutdown, so no pane outlives the process as an orphan.
//
// EVERY KILL HERE IS SYNCHRONOUS, and that is load-bearing. It is NOT safe to leave the
// booting pane to clean itself up on resolve: gracefulShutdown calls process.exit() in the
// same tick as this drain (TUI panes are children of the tmux SERVER, not of node, so
// node's activeProcesses set is empty on a TUI host and the "wait for children" path exits
// immediately). A .then()/.catch() scheduled here would never run, and the pane would
// survive as an orphaned, authenticated, idle `claude`.
drain() {
this._paused = true;
let n = this._panes.length;
for (const p of this._panes) this._drop(p, "drain");
this._panes = [];
n += this._cancelBooting("drain_booting");
return n;
}
// Undo drain() and start refilling again. Because drain() CANCELLED the in-flight boot
// (rather than leaving it pending), the booting slot is free and this really does start a
// fresh boot — the pool is never left empty with nothing scheduled.
resume() {
this._paused = false;
this.refill();
}
// /health surface (additive).
stats() {
return {
size: this.size,
warm: this._panes.length,
booting: this.booting,
model: this.warmModel,
hits: this.hits,
misses: this.misses,
boots: this.boots,
bootFailures: this.bootFailures,
cancelled: this.cancelled,
dropped: this.dropped,
};
}
_drop(pane, reason) {
this.dropped++;
try { this._killPane(pane.name); } catch { /* already gone */ }
// F5: every drop path (expired / unhealthy / model_switch / drain / cancelled_late /
// stale_boot) ends up here, and the reap tick drains the WHOLE pool on every tick — so
// without this, every warm pane's sink orphans in streamDir with no GC path (killPane only
// reaches the tmux session, never the pane's OWN files). Best-effort: pane.streamFile is
// undefined for a still-booting identity (the sink path is only known once bootPane
// resolves) and rmSync(force:true) is already a no-op on a missing file, so this never
// throws into the reaper regardless of which drop path got here.
if (pane.streamFile) {
try { rmSync(pane.streamFile, { force: true }); } catch { /* best-effort GC */ }
}
this._log("info", "tui_pool_pane_dropped", { name: pane.name, reason });
}
}
+41 -1
View File
@@ -139,7 +139,40 @@ export function recordTuiEntrypoint(tuiStats, observed, expectedMode = "cli") {
// Build the additive /health `tui` block (ADR 0007 PR-B amendment). Pure: given the
// config + live counters, returns the exact object embedded in /health. New fields only —
// behaviour-preserving for existing /health consumers (grandfathered B.2 under ADR 0006).
export function buildTuiHealthBlock({ enabled, entrypointMode, maxConcurrent }, tuiStats, semaphore) {
//
// `pool` (optional, warm pane pool — lib/tui/pool.mjs): a TuiPanePool, or null/undefined
// when the pool is off (the default). Reported as `pool: null` when off so the block's
// shape stays stable, and as the pool's stats (size / warm / hits / misses / …) when on —
// the operator's window onto both the hit rate and the standing idle-process cost.
//
// Streaming fields (backlog #2, OCP_TUI_STREAM) are ADDITIVE too:
// streamEnabled — is real (MessageDisplay-hook) SSE streaming on for TUI turns?
// streamTurns — streamed turns ATTEMPTED, counted before the truncation/auth-banner
// gates run (F6) — so a turn REFUSED by those gates still shows up
// here, which is exactly the turn an operator most wants visible.
// Counting only turns that survived the gates would silently exclude
// a turn's worst-case outcome from its own denominator.
// streamDeltas — MessageDisplay hook fires OBSERVED, including held-back ones (F6) —
// NOT only the ones forwarded to a client. This is what makes
// streamZeroDeltaTurns meaningful: a turn can have streamDeltas
// incrementing while still emitting nothing to the client (fully held
// back, e.g. a short answer), which is healthy, vs. a hook that fired
// zero times at all, which is not (see streamZeroDeltaTurns).
// streamTopUps — turns where the delta stream was a safe PREFIX of the transcript but
// not equal to it; OCP topped up from the transcript and served T.
// Benign but worth watching — a persistent rate means the hook is
// losing fires.
// streamDivergences — turns REFUSED because emitted bytes were not a prefix of the
// transcript. THE field to alert on for CORRECTNESS: it means the hook
// and the transcript disagreed and OCP chose to fail rather than serve
// unverifiable text.
// streamZeroDeltaTurns — streamed turns where the hook fired ZERO times (F7). THE field to
// alert on for AVAILABILITY: streamTopUps climbing is one fire dropped
// here and there (benign); this climbing means the hook is not firing
// AT ALL — e.g. `--settings` silently stopped registering it (a claude
// version bump), or F3's truncated-script failure mode — and every
// streamed turn is quietly degrading to fully-buffered with no error.
export function buildTuiHealthBlock({ enabled, entrypointMode, maxConcurrent, streamEnabled = false }, tuiStats, semaphore, pool = null) {
return {
enabled,
entrypointMode, // cli | auto | off
@@ -148,5 +181,12 @@ export function buildTuiHealthBlock({ enabled, entrypointMode, maxConcurrent },
inflight: semaphore.inflight, // current concurrent TUI turns
queued: semaphore.queued, // turns waiting for a slot
maxConcurrent,
pool: pool ? pool.stats() : null, // warm pane pool, or null when disabled
streamEnabled,
streamTurns: tuiStats.streamTurns ?? 0,
streamDeltas: tuiStats.streamDeltas ?? 0,
streamTopUps: tuiStats.streamTopUps ?? 0,
streamDivergences: tuiStats.streamDivergences ?? 0,
streamZeroDeltaTurns: tuiStats.streamZeroDeltaTurns ?? 0,
};
}
+340 -64
View File
@@ -14,6 +14,7 @@ import { mkdtempSync, writeFileSync, readFileSync, mkdirSync, existsSync, rmSync
import { tmpdir } from "node:os";
import { randomUUID } from "node:crypto";
import { readTuiTranscript } from "./transcript.mjs";
import { prepareStreamHook, streamFilePath, parseDeltaChunk } from "./stream.mjs";
// F7 fix (audit finding, LOW): the prefix used to be a bare, host-wide constant
// ("ocp-tui-"), so a SECOND OCP instance on the same host (e.g. a temporary
@@ -73,6 +74,36 @@ const defaultTmux = (args, opts = {}) =>
// `port` (required) is this instance's own listen port (server.mjs's PORT / lib/constants.mjs
// DEFAULT_PORT resolution) — the SPOT for "which sessions are ours."
//
// ── POOL/REAPER INVARIANT (warm pane pool — lib/tui/pool.mjs) ───────────────────────────
// A warm pooled pane is one of OUR OWN `ocp-tui-<port>-*` sessions that is ALIVE AND IDLE
// BY DESIGN — and the periodic sweep runs precisely when the instance is idle, i.e. exactly
// when the pool is full. Without an exemption the sweep would kill every warm pane on every
// tick (and kill-server on top). The exemption is `spare`: a set of EXACT session names the
// caller declares live. Three properties, all load-bearing:
//
// 1. A LIVE POOLED PANE IS NEVER REAPED — INCLUDING ONE THAT IS STILL BOOTING. It is in
// `spare` (the pool's live registry), so it is skipped by name. The booting case is not
// a footnote, it is the one that bit us: bootTuiPane creates the tmux session
// SYNCHRONOUSLY and only then waits up to POOL_BOOT_MS for the input bar, so a pooled
// session can be live for ~20 s before its boot resolves. The pool therefore mints the
// pane's NAME up front and holds it in `_bootingPane`, so liveNames() can name — and
// spare — a session whose boot has not finished. (An earlier version tracked only a
// COUNT of in-flight boots; the sweep could not name that session and killed it.)
// 2. A LEAKED/ORPHANED POOLED PANE IS STILL REAPED. Membership is by EXACT NAME from a
// live in-memory registry — NOT by "looks pooled" (name shape). A pane the pool no
// longer owns (handed out, dropped, cancelled, or left behind by a previous process
// generation — whose registry died with it) is absent from `spare` and is killed like
// any other stale session. Fail-safe: forgetting to pass `spare` reaps MORE, never less.
// 3. KILL-SERVER NEVER KILLS A LIVE POOL PANE. A spared session suppresses kill-server
// exactly as a foreign session does (it is a live child of the tmux server). The
// consequence — that a permanently-full pool would permanently disable the defunct-
// zombie reaping that ONLY kill-server can do — is resolved in server.mjs by DRAINING
// the pool immediately before the sweep, so `spare` is empty on the normal tick and
// kill-server still fires. `spare` is the belt-and-braces: a reap call site that
// forgets to drain still cannot kill a live pane.
//
// `spare` (default: none) — iterable of session names, or a Set. Ignored when the pool is off.
//
// `includeLegacy` (default false): when true, sessions matching the exact OLD bare-prefix
// shape (LEGACY_SESSION_NAME_RE) are ALSO treated as ours for kill-session purposes. This is
// the boot-time legacy migration: an operator upgrading past this fix could otherwise be left
@@ -88,14 +119,20 @@ const defaultTmux = (args, opts = {}) =>
// same class of residual risk the audit finding itself accepts ("no live instance of the new
// version creates them"); this PR does not regress that scenario, it only removes the far
// more common same-version collision (the actual F7 finding).
export function reapStaleTuiSessions({ tmux = defaultTmux, port, includeLegacy = false } = {}) {
export function reapStaleTuiSessions({ tmux = defaultTmux, port, includeLegacy = false, spare = null } = {}) {
const r = tmux(["list-sessions", "-F", "#{session_name}"]);
if (!r || r.status !== 0) return 0; // no tmux server / no sessions
const names = String(r.stdout || "").split("\n").map((s) => s.trim()).filter(Boolean);
const ownPrefix = sessionPrefixForPort(port);
const spared = spare instanceof Set ? spare : new Set(spare || []);
let killed = 0;
let othersRemain = false;
let sparedLive = 0;
for (const name of names) {
// Property 1+2: exemption is by EXACT NAME from the pool's live registry. A pooled-
// LOOKING name that is not in the registry is an orphan and falls through to the
// normal kill path below.
if (spared.has(name)) { sparedLive++; continue; }
const isOwn = name.startsWith(ownPrefix);
const isLegacyOwn = includeLegacy && LEGACY_SESSION_NAME_RE.test(name);
if (isOwn || isLegacyOwn) {
@@ -109,7 +146,11 @@ export function reapStaleTuiSessions({ tmux = defaultTmux, port, includeLegacy =
// Reap defunct `claude` zombies: safe ONLY when the server is now ours-only/empty.
// kill-server is what actually reaps (server exit reparents survivors to init); a
// per-session kill cannot, since node is not the zombies' parent.
if (!othersRemain) {
//
// Property 3: a SPARED session is a live child of this tmux server, so kill-server would
// kill it — it therefore suppresses kill-server exactly as a foreign session does. On the
// normal sweep the pool is drained first, so sparedLive is 0 and kill-server still fires.
if (!othersRemain && sparedLive === 0) {
tmux(["kill-server"]);
}
return killed;
@@ -119,8 +160,18 @@ export function reapStaleTuiSessions({ tmux = defaultTmux, port, includeLegacy =
// Boot + paste-settle timing. Conservative defaults validated on PI231; env-tunable.
const BOOT_MS = parseInt(process.env.OCP_TUI_BOOT_MS || "4000", 10); // max wait for input-ready
// Readiness cap for a POOL pre-boot. Deliberately far more generous than BOOT_MS: BOOT_MS is
// tight because a client is blocked on it, whereas a warm-pane boot happens in the background
// with nobody waiting. Observed live at size=2: a refill booting alongside an in-flight turn
// exceeded 4000 ms and was discarded (tui_pool_boot_failed), quietly costing hit rate for a
// pane that was merely slow, not broken. Scales with OCP_TUI_BOOT_MS if an operator raises it.
export const POOL_BOOT_MS = BOOT_MS * 5;
const READY_POLL_MS = parseInt(process.env.OCP_TUI_READY_POLL_MS || "400", 10); // readiness / paste-verify poll interval
const PASTE_VERIFY_MS = parseInt(process.env.OCP_TUI_PASTE_VERIFY_MS || "5000", 10); // max wait for pasted prompt to render
// Hook-sink drain interval when streaming. 100ms: the hook fires at BLOCK granularity
// (~5-7 fires per answer, seconds apart), so a finer poll buys nothing and a coarser one
// would add visible lag to the first delta. Cheap — one readFileSync of a small file.
const STREAM_POLL_MS = parseInt(process.env.OCP_TUI_STREAM_POLL_MS || "100", 10);
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
@@ -302,7 +353,25 @@ export function prepareTuiHome(realHome, tuiHome, cwd, { envTokenMode = false }
// A-PATH ONLY: built-in tools are left enabled (acceptable single-user). Deployment B
// (guest keys) MUST additionally pass --tools "" per spec §5.2(2) as the credential
// wall before this argv is reachable for owner_tier=guest — guard that in PR-3 wiring.
export function buildTuiCmd(claudeBin, model, sessionId, ehome, entrypointMode) {
//
// `stream` (optional, OCP_TUI_STREAM): { file, settings } — when present, the pane gets
// (a) OCP_TUI_STREAM_FILE in its env — read by the static MessageDisplay hook script to
// decide WHERE to append this pane's deltas. Delivered as env (not baked into the
// settings file) so the settings file stays STATIC and a pre-booted warm pane works.
// Verified live: a claude hook inherits the pane's environment.
// (b) --settings <file> — registers the MessageDisplay hook.
// VERIFIED LIVE (claude 2.1.207, this host) before shipping, because both were spawn-level
// risks:
// - the startup banner is UNCHANGED with --settings: "Sonnet 4.6 with low effort ·
// Claude Max" (subscription pool). --settings is NOT a --bare-class flag — it does not
// silently drop the subscription pool. Transcript entrypoint stayed "cli".
// - --settings MERGES into the settings hierarchy, it does NOT clobber <HOME>/.claude/
// settings.json: with --settings passed, the user-level settings.json's `env` block was
// still applied to the hook's environment. So the isolated-HOME settings story the TUI
// already relies on (permissions / additionalDirectories — see prepareTuiHome and the
// OCP_TUI_FULL_TOOLS note above) survives intact.
// When absent, the argv is byte-for-byte the pre-streaming argv.
export function buildTuiCmd(claudeBin, model, sessionId, ehome, entrypointMode, stream = null) {
// Deliver claude's env via an `env` prefix on the PANE COMMAND — tmux does NOT forward the
// spawning process's environment to the pane, and `new-session -e` needs tmux ≥3.2 (the cloud
// host runs 2.7), so this is the only portable, reliable mechanism (verified live 2026-06-01:
@@ -346,6 +415,8 @@ export function buildTuiCmd(claudeBin, model, sessionId, ehome, entrypointMode)
if (process.env.CLAUDE_CODE_OAUTH_TOKEN) {
sets.push(`CLAUDE_CODE_OAUTH_TOKEN=${shq(process.env.CLAUDE_CODE_OAUTH_TOKEN)}`);
}
// Streaming sink: the pane's own per-session delta file (see the `stream` note above).
if (stream && stream.file) sets.push(`OCP_TUI_STREAM_FILE=${shq(stream.file)}`);
const unset = ["CLAUDECODE", "ANTHROPIC_API_KEY", "ANTHROPIC_BASE_URL", "ANTHROPIC_AUTH_TOKEN"];
if (entrypointMode === "cli") sets.push("CLAUDE_CODE_ENTRYPOINT=cli");
else if (entrypointMode === "auto") unset.push("CLAUDE_CODE_ENTRYPOINT"); // let claude self-classify via TTY
@@ -377,47 +448,107 @@ export function buildTuiCmd(claudeBin, model, sessionId, ehome, entrypointMode)
} else {
toolArgs = ["--strict-mcp-config", "--disallowedTools", shq("mcp__*")];
}
// Effort: pass --effort EXPLICITLY. Without it, the pane's claude inherits a
// HOME-dependent effortLevel — real-home mode inherits the operator's
// ~/.claude/settings.json (whatever they set for their own interactive use),
// env-token scratch mode inherits claude's built-in default (prepareTuiHome never
// writes effortLevel) — so latency silently depends on which HOME mode
// resolveTuiHome() picked AND on an unrelated operator setting. Pinning it here
// removes both. Measured (docs/plans/2026-07-13-tui-latency): explicit low cuts
// direct-spawn TTFT p50 10.35s → 6.17s (40%) and collapses the spread ~15×;
// banner-verified to stay on the subscription pool (`· Claude Max`).
// OCP_TUI_EFFORT=inherit restores the pre-flag argv byte-for-byte (no --effort).
// An unknown value falls back to the default rather than reaching claude's argv:
// a typo'd --effort value must not risk a spawn-time usage error in the pane.
const EFFORT_LEVELS = ["low", "medium", "high", "xhigh", "max"]; // claude 2.1.207 --help
const effortRaw = (process.env.OCP_TUI_EFFORT || "low").trim().toLowerCase();
let effortArgs;
if (effortRaw === "inherit") {
effortArgs = [];
} else if (EFFORT_LEVELS.includes(effortRaw)) {
effortArgs = ["--effort", effortRaw];
} else {
console.error(`[tui] invalid OCP_TUI_EFFORT=${JSON.stringify(process.env.OCP_TUI_EFFORT)}; using "low" (valid: ${EFFORT_LEVELS.join("|")}, or "inherit" to omit the flag)`);
effortArgs = ["--effort", "low"];
}
// --settings registers the MessageDisplay hook. Omitted entirely when streaming is off,
// so the OFF argv is byte-for-byte the pre-streaming argv.
const settingsArgs = stream && stream.settings ? ["--settings", shq(stream.settings)] : [];
return [
envPrefix,
shq(claudeBin),
"--model", shq(model),
"--session-id", sessionId,
...toolArgs,
...effortArgs,
...settingsArgs,
].join(" ");
}
// Full per-request TUI lifecycle:
// 1. Pre-trust the scratch cwd (no trust dialog will appear).
// 2. Write prompt to a 0600 temp file (no shell injection from prompt content).
// 3. Boot an interactive `claude` in a fresh tmux session in the scratch cwd; poll
// capture-pane until the `? for shortcuts` input bar appears (readiness-poll
// replaces the old blind boot sleep). BOOT_MS is the max wait, not a fixed delay.
// 4. Paste the prompt via tmux load-buffer + paste-buffer -p (bracketed paste) —
// reliable for large multi-line prompts where send-keys -l is not (issue #130).
// Poll-verify the prompt landed in the input (placeholder gone / [Pasted text]);
// fast-fail with tui_paste_not_landed if it never lands (prevents the 120s
// wallclock "stuck typing" hang). Then submit with a SEPARATE Enter key event.
// 5. Block on the native JSONL transcript (located by session-id) until terminal
// marker or wall-clock cap.
// 6. Always teardown: kill session + rm temp dir (even on throw).
// Returns { text, entrypoint } from readTuiTranscript (entrypoint is the billing-pool
// classifier, e.g. "cli", or null if the transcript did not include a turn_duration).
export async function runTuiTurn({
prompt,
model,
claudeBin,
home,
realHome,
cwd,
port,
wallclockMs = 120000,
entrypointMode = "cli",
tmux = defaultTmux,
// Is a pane alive AND still sitting at its input bar? Used by the warm pool to decide,
// at hand-out time, whether a pre-booted pane is still usable (a dead/degraded pane must
// become a MISS → cold path, never a hung turn). capture-pane exits non-zero when the
// session no longer exists, so this covers "pane gone" and "pane not ready" in one call.
export function tuiPaneHealthy(tmux, tmuxName) {
const r = tmux(["capture-pane", "-p", "-t", tmuxName]);
if (!r || r.status !== 0 || typeof r.stdout !== "string") return false;
return tuiInputReady(r.stdout);
}
// Pool pane names carry a "p" marker after the port-scoped prefix:
// turn pane: ocp-tui-<port>-<8hex> (unchanged)
// pool pane: ocp-tui-<port>-p<8hex>
// Purely for operator legibility (`tmux ls` shows which panes are warm). It is NOT the
// reaper's exemption mechanism — that is the exact-name spare set (see the POOL/REAPER
// INVARIANT above), so a pooled-LOOKING orphan is still reaped. Both shapes start with
// sessionPrefixForPort(port), so both remain reapable as "ours", and neither can match
// LEGACY_SESSION_NAME_RE.
export function poolPaneName(port, sessionId) {
return sessionPrefixForPort(port) + "p" + sessionId.slice(0, 8);
}
// Boot ONE interactive `claude` pane and wait for its input bar. Shared by the cold
// request path (runTuiTurn) and the warm pool (lib/tui/pool.mjs) so a pooled pane is
// spawned with byte-for-byte the same argv, HOME, cwd and trust preparation as a
// cold-booted one — the pool must not become a second, drifting spawn path.
//
// Each pane gets its OWN fresh randomUUID() --session-id, fixed at boot. That is what
// keeps a pooled pane single-use-safe: its transcript holds exactly one exchange.
//
// requireReady: the cold path tolerates a readiness timeout (it falls through and lets
// the paste-verify decide — pre-existing behaviour, unchanged). The POOL sets it, because
// a pane that never reached its input bar is worthless as a warm pane and must not be
// enlisted: throw, let the pool count a bootFailure, and leave the request path to
// cold-boot as usual.
// bootMs: max wait for the input bar. Defaults to BOOT_MS (the REQUEST path's cap, which is
// deliberately tight — a client is blocked on it). The POOL passes POOL_BOOT_MS instead: a
// background pre-boot has nobody waiting on it, and capping it at the request-path's 4 s
// made real refills fail (observed live: a refill booting alongside an in-flight turn took
// >4 s and was discarded, silently lowering the hit rate). Slow != broken for a pre-boot.
// `sessionId` / `name` (both optional): the caller may supply the pane's identity instead of
// letting bootTuiPane mint it. The POOL does, because it must know the tmux session's NAME
// before this function runs — the session is created synchronously below, well before the
// readiness wait returns, so a pool that only learned the name on resolve could neither spare
// the session from the reaper nor kill it on shutdown. Supplying BOTH also keeps the name's
// hex suffix equal to the session-id's, so `tmux ls` correlates to the transcript file.
// `streamDir` (optional, OCP_TUI_STREAM): install claude's MessageDisplay hook on this pane.
// Done HERE, at boot — not at turn time — and that is the whole reason streaming survives the
// WARM POOL: the hook script + settings file are STATIC (one pair per streamDir), and the only
// per-turn thing, the sink path, is derived from the pane's own --session-id, which is fixed
// right here. So a pre-booted pane already carries its hook and its own sink and streams exactly
// like a cold-booted one; nothing request-specific is ever baked into the spawn.
export async function bootTuiPane({
model, claudeBin, home, realHome, cwd, port, entrypointMode = "cli",
tmux = defaultTmux, sessionId = null, name = null, requireReady = false, bootMs = BOOT_MS,
streamDir = null,
}) {
const sessionId = randomUUID();
const sid = sessionId || randomUUID();
// Port-scoped session name (F7 fix) — see sessionPrefixForPort / reapStaleTuiSessions
// for why this instance's own listen port is the namespace discriminator.
const tmuxName = sessionPrefixForPort(port) + sessionId.slice(0, 8);
const tmuxName = name || (sessionPrefixForPort(port) + sid.slice(0, 8));
const ehome = home || process.env.HOME; // HOME claude runs under (scratch or real)
const rhome = realHome || process.env.HOME; // real home (OAuth + onboarded config source)
@@ -434,10 +565,14 @@ export async function runTuiTurn({
if (!existsSync(cwd)) mkdirSync(cwd, { recursive: true });
prepareTuiHome(rhome, ehome, cwd, { envTokenMode });
// Write prompt to a temp file (mode 0600) so the content never touches argv.
const tmpDir = mkdtempSync(`${tmpdir()}/ocp-tui-`);
const promptFile = `${tmpDir}/prompt.txt`;
writeFileSync(promptFile, prompt, { mode: 0o600 });
// Streaming sink for THIS pane (see the streamDir note above). rmSync first so a
// re-used session-id can never replay a previous turn's deltas.
let streamFile = null, streamSettings = null;
if (streamDir) {
streamFile = streamFilePath(streamDir, sid);
streamSettings = prepareStreamHook(streamDir);
try { rmSync(streamFile, { force: true }); } catch { /* start from a fresh sink */ }
}
// Minimal env for spawnSync (tmux itself). The pane's claude env comes exclusively
// from the `env` prefix string built inside buildTuiCmd — tmux does NOT forward the
@@ -445,31 +580,147 @@ export async function runTuiTurn({
const env = { ...process.env };
env.HOME = ehome; // tmux needs HOME; all claude-specific vars go via buildTuiCmd prefix
// Boot the interactive session inside tmux, rooted at the scratch cwd.
// Capture the result: if tmux new-session fails (status !== 0) there is no PTY, no
// interactive spawn — abort BEFORE the boot wait rather than paste into a non-existent
// session or issue a billing request without a verified interactive context.
const spawnResult = tmux(
["new-session", "-d", "-s", tmuxName, "-x", "220", "-y", "50", "-c", cwd,
buildTuiCmd(claudeBin, model, sid, ehome, entrypointMode,
streamFile ? { file: streamFile, settings: streamSettings } : null)],
{ env },
);
if (!spawnResult || spawnResult.status !== 0) {
throw new Error("tui_spawn_failed: tmux session not created");
}
// Wait until claude's input bar is actually ready (not a blind sleep).
// bootMs is the MAX readiness wait, not a fixed delay.
const ready = await pollUntil(() => tuiInputReady(tuiCapturePane(tmux, tmuxName)),
{ timeoutMs: bootMs, intervalMs: READY_POLL_MS });
if (!ready) {
if (requireReady) {
try { tmux(["kill-session", "-t", tmuxName]); } catch { /* already gone */ }
throw new Error("tui_pane_not_ready: input bar did not appear within " + bootMs + "ms");
}
// Cold path (pre-existing behaviour): readiness timed out; rely on paste-verify.
console.error("[tui] input_not_ready", tmuxName);
}
return { name: tmuxName, sessionId: sid, model, ehome, streamFile, bootedAt: Date.now() };
}
// Full per-request TUI lifecycle:
// 1. Take a WARM pane from the pool if one is available for this model (opt-in;
// OCP_TUI_POOL_SIZE=0 => always null => steps 2-3 below are exactly today's path).
// A pooled pane is SINGLE-USE: it already carries its own fresh --session-id, it
// serves this one turn, and it is killed in the finally like any other pane.
// 2. On a MISS: pre-trust the scratch cwd, boot an interactive `claude` in a fresh tmux
// session in the scratch cwd, poll capture-pane until the `? for shortcuts` input bar
// appears (bootTuiPane). BOOT_MS is the max wait, not a fixed delay.
// 3. Write prompt to a 0600 temp file (no shell injection from prompt content).
// 4. Paste the prompt via tmux load-buffer + paste-buffer -p (bracketed paste) —
// reliable for large multi-line prompts where send-keys -l is not (issue #130).
// Poll-verify the prompt landed in the input (placeholder gone / [Pasted text]);
// fast-fail with tui_paste_not_landed if it never lands (prevents the 120s
// wallclock "stuck typing" hang). Then submit with a SEPARATE Enter key event.
// 5. Block on the native JSONL transcript (located by THIS pane's session-id) until
// terminal marker or wall-clock cap.
// 6. Always teardown: kill session + rm temp dir (even on throw), and kick a background
// pool refill so the next request finds a warm pane.
// Returns { text, entrypoint } from readTuiTranscript (entrypoint is the billing-pool
// classifier, e.g. "cli", or null if the transcript did not include a turn_duration).
//
// STREAMING (OCP_TUI_STREAM, default off). Pass `onDelta` and `streamDir`, and the pane's
// MessageDisplay hook (installed by bootTuiPane; see lib/tui/stream.mjs) appends each raw
// delta payload to the pane's own sink. This driver polls that sink and invokes onDelta(payload)
// per fire while the turn is still generating. A WARM pane already carries its sink from boot
// (pane.streamFile), so the pooled and cold paths stream identically.
//
// `streamDir` IS PASSED TO THE COLD BOOT UNCONDITIONALLY (not gated on `onDelta`) — F4 fix. The
// spawn argv is this project's billing-classification surface: a caller with OCP_TUI_STREAM on
// but THIS particular request non-streaming (stream:false) must still get the SAME argv whether
// it lands on a pool HIT or a cold-boot MISS, because a pre-booted pool pane cannot know in
// advance whether the request it will eventually serve wants streaming — it installs the hook
// unconditionally whenever the pool is warming at all (see server.mjs's bootPane closure). Gating
// the cold boot's hook install on `onDelta` made a stream:false request's argv depend on whether
// it happened to hit the pool or miss it — the exact drift this surface cannot tolerate. Whether
// the hook is actually POLLED is a separate, correctly-scoped decision: see `streaming` below,
// gated on onDelta && streamFile, so a non-streaming turn never reads its own sink even though
// the hook is running.
//
// The transcript stays AUTHORITATIVE regardless: it is still the terminal-turn signal, still the
// source of the returned `text`, and still the input to the caller's honesty gates. The delta
// stream is a low-latency MIRROR of it, never a replacement, and the caller asserts the two
// agree. With onDelta AND streamDir both omitted, nothing here changes: no poll, no hook.
//
// `abortSignal` (optional): aborts the transcript wait, so a client that disconnects mid-turn
// tears the pane down NOW (the finally below) instead of holding the pane — and therefore the
// caller's semaphore slot — until the turn or the wallclock cap ends.
export async function runTuiTurn({
prompt,
model,
claudeBin,
home,
realHome,
cwd,
port,
wallclockMs = 120000,
entrypointMode = "cli",
tmux = defaultTmux,
pool = null, // TuiPanePool | null — null (default) === today's cold-boot-only path
onPane = null, // optional observer: ({ warm }) => void, for logging/metrics
onDelta = null, // (payload) => void — invoked per MessageDisplay hook fire, mid-turn
streamDir = null, // hook sink dir, passed to the COLD boot UNCONDITIONALLY (F4 — see above);
// a warm pane brings its own, fixed at its own boot
abortSignal = null,
}) {
// 1. Warm pane, or cold boot. A MISS is never an error — it is exactly today's path.
let pane = pool ? pool.acquire(model) : null;
const warm = !!pane;
// Kick the refill IMMEDIATELY (not after the turn): the replacement pane then boots
// CONCURRENTLY with this turn and is warm by the time the next request arrives. Also
// runs on a MISS — acquire() has just retargeted the pool to this model, so the miss
// that cold-boots today warms the pool for the next caller. Fire-and-forget; it takes
// no TuiSemaphore slot (see pool.refill's SLOT ACCOUNTING note).
if (pool) pool.refill();
if (onPane) { try { onPane({ warm }); } catch { /* observer must never break a turn */ } }
if (!pane) {
// streamDir passed AS-IS (not gated on onDelta) — F4: see the STREAMING comment above.
pane = await bootTuiPane({ model, claudeBin, home, realHome, cwd, port, entrypointMode, tmux,
streamDir });
}
const tmuxName = pane.name;
const sessionId = pane.sessionId; // THIS pane's own session-id — one session, one turn
const ehome = pane.ehome || home || process.env.HOME;
// Streaming state is read off the PANE, not recomputed here — a warm pane fixed its sink at
// boot, and a cold one just did the same above. If the pool was booted WITHOUT a streamDir
// while onDelta is set, streamFile is null and the turn degrades to buffered: correct, just
// not fast. (server.mjs wires the same streamDir into both paths so that cannot happen.)
const streamFile = pane.streamFile || null;
const streaming = !!(onDelta && streamFile);
const streamCursor = { consumed: 0 };
let streamStopped = false;
let pollTimer = null;
// Drain every complete line appended since the last drain. Never throws into the turn: a
// malformed line is skipped by parseDeltaChunk, and an onDelta that throws is contained.
const drainDeltas = () => {
if (!streaming) return;
let text;
try { text = readFileSync(streamFile, "utf8"); } catch { return; } // absent until the first fire
const { deltas, consumed } = parseDeltaChunk(text, streamCursor.consumed);
streamCursor.consumed = consumed;
for (const d of deltas) {
try { onDelta(d); } catch { /* a sink error must never abort the turn */ }
}
};
// Write prompt to a temp file (mode 0600) so the content never touches argv.
const tmpDir = mkdtempSync(`${tmpdir()}/ocp-tui-`);
const promptFile = `${tmpDir}/prompt.txt`;
writeFileSync(promptFile, prompt, { mode: 0o600 });
try {
// 1. Boot the interactive session inside tmux, rooted at the scratch cwd.
// Capture the result: if tmux new-session fails (status !== 0) there is no
// PTY, no interactive spawn — abort BEFORE the boot sleep rather than paste
// into a non-existent session or issue a billing request without a verified
// interactive context. The finally teardown is still harmless (kill-session
// is a no-op when the session never existed).
const spawnResult = tmux(
["new-session", "-d", "-s", tmuxName, "-x", "220", "-y", "50", "-c", cwd,
buildTuiCmd(claudeBin, model, sessionId, ehome, entrypointMode)],
{ env },
);
if (!spawnResult || spawnResult.status !== 0) {
throw new Error("tui_spawn_failed: tmux session not created");
}
// 2. Wait until claude's input bar is actually ready (was: blind sleep(BOOT_MS)).
// BOOT_MS is now the MAX readiness wait, not a fixed delay.
const ready = await pollUntil(() => tuiInputReady(tuiCapturePane(tmux, tmuxName)),
{ timeoutMs: BOOT_MS, intervalMs: READY_POLL_MS });
if (!ready) {
// (readiness timed out; relying on paste-verify)
console.error("[tui] input_not_ready", tmuxName);
}
// 3. Paste the prompt via a tmux PASTE BUFFER with bracketed paste (-p), NOT
// `send-keys -l`. send-keys of a large multi-line prompt is unreliable: the
// embedded newlines arrive as separate key events (effectively repeated Enter),
@@ -495,12 +746,37 @@ export async function runTuiTurn({
// Submit (separate Enter key event).
tmux(["send-keys", "-t", tmuxName, "Enter"]);
// 4. Block on the native transcript (resolved by session-id) until terminal.
// Returns { text, entrypoint } from readTuiTranscript.
return await readTuiTranscript({ home: ehome, sessionId, wallclockMs });
// 5a. Streaming only: start polling the hook sink. Runs CONCURRENTLY with the
// transcript wait below — the deltas are what make the answer visible while the
// turn is still generating; the transcript is what makes it authoritative.
if (streaming) {
const loop = () => {
if (streamStopped) return;
drainDeltas();
pollTimer = setTimeout(loop, STREAM_POLL_MS);
};
pollTimer = setTimeout(loop, STREAM_POLL_MS);
}
// 5b. Block on the native transcript (resolved by THIS pane's session-id) until terminal.
// Returns { text, entrypoint, truncated } from readTuiTranscript.
const result = await readTuiTranscript({ home: ehome, sessionId, wallclockMs, abortSignal });
// 5c. FINAL drain. The terminal marker can land between two poll ticks, so the last
// delta(s) may still be unread — without this the tail would be missing from the
// stream and every turn would need a transcript top-up.
streamStopped = true;
if (pollTimer) clearTimeout(pollTimer);
drainDeltas();
return result;
} finally {
// 5. Teardown — always, even on throw.
// 6. Teardown — always, even on throw (including an abortSignal disconnect, which is
// exactly why the pane cannot outlive a client that walked away). A pooled pane is
// torn down here exactly like a cold-booted one: SINGLE-USE, never returned (pool.mjs).
streamStopped = true;
if (pollTimer) clearTimeout(pollTimer);
try { tmux(["kill-session", "-t", tmuxName]); } catch { /* already gone */ }
try { rmSync(tmpDir, { recursive: true, force: true }); } catch { /* best effort */ }
if (streamFile) { try { rmSync(streamFile, { force: true }); } catch { /* best effort */ } }
}
}
+273
View File
@@ -0,0 +1,273 @@
// TUI-mode real SSE streaming — the `MessageDisplay` hook sink.
//
// WHAT THIS IS. `claude` fires a **MessageDisplay** hook per rendered block of the
// assistant's reply, handing the hook the RAW MARKDOWN SOURCE of an incremental
// `delta` on stdin. Registered via `--settings` on the ordinary interactive TUI spawn
// (NO -p, NO --bare — the billing pool is untouched), it is the only byte-faithful
// incremental source the interactive CLI exposes. Everything here consumes that hook
// surface AS EMITTED — forwarding, not inventing.
//
// ALIGNMENT.md: **Class B**. We consume claude's own hook payload and re-emit it in the
// OpenAI chat/completions streaming shapes OCP already speaks (ADR 0006). There is no
// `cli.js` citation because no `cli.js` function is being mirrored: the TUI spawn is
// OCP-owned surface (ADR 0007), and the hook payload is claude's own published contract.
//
// THE VERIFIED CONTRACT (docs/plans/2026-07-13-tui-latency/streaming-spike.md, and
// independently reproduced on claude 2.1.207 / sonnet-4-6 / banner `· Claude Max`):
//
// payload (stdin, one JSON object per fire):
// { hook_event_name:"MessageDisplay", session_id, transcript_path, prompt_id, cwd,
// turn_id, message_id, index, final, delta }
//
// - deltas carry the raw markdown source (`## `, `**`, ```javascript all present)
// - concat(deltas of one message) === T, byte-exactly (T = extractLatestAssistantText)
// - T.startsWith(concat(deltas[0..n])) at EVERY n (prefix-stable)
// - block-level granularity (~5-7 fires per answer), NOT token-level
// - only `text` blocks fire it — thinking blocks are excluded (what OCP wants)
//
// ⚠️ THE HOOK IS SYNCHRONOUS. The hook's source sets `forceSyncExecution: true` —
// `claude` BLOCKS on every fire. The hook script must therefore write and exit, doing
// NO work inline. Measured cost of the script below: p50 7.2 ms / p90 14.7 ms per fire,
// i.e. ~50 ms added blocking across a whole ~7-delta turn against a 6-10 s turn. That is
// noise, so a plain append is the right sink — a FIFO would be faster on paper but a FIFO
// blocks its writer until a reader attaches, which would hand `claude` a way to hang.
//
// WARM-POOL COMPATIBILITY (load-bearing — a warm pane pool is a separate in-flight PR).
// The hook script and the settings file are BOTH STATIC: one copy per stream dir, written
// once, never per-request. The per-turn destination is carried in the PANE'S OWN ENV as
// `OCP_TUI_STREAM_FILE` (verified live: a hook inherits the pane's environment), and the
// path is derived from the session-id — which for a pre-booted pane is fixed at BOOT.
// Nothing about a request is baked into the settings file at spawn time, so a pane booted
// before its request arrives streams exactly the same way.
import { writeFileSync, mkdirSync, renameSync } from "node:fs";
import { detectTuiUpstreamError } from "./transcript.mjs";
// Default holdback before the first byte is released to the client. See TuiDeltaAssembler.
export const DEFAULT_HOLDBACK_CHARS = 100;
// The hook script. POSIX sh, no interpreter startup beyond /bin/sh, one fork (`cat`).
//
// - `printf` is a shell BUILTIN in sh/dash/bash, so the newline costs no fork.
// - the `{ cat; printf '\n'; } >>` group opens the file ONCE and appends both writes
// through the same O_APPEND fd, so a payload and its terminator can never be split
// by another writer. (They never race anyway: one file per pane, and MessageDisplay
// is synchronous within a pane.)
// - a payload JSON can never contain a literal newline — JSON.stringify escapes them —
// so "one line == one payload" holds, and a torn write is always a trailing partial
// line, which parseDeltaChunk() leaves unconsumed until it completes.
// - NO OCP_TUI_STREAM_FILE (e.g. a pane booted with streaming off, or any other claude
// session that happens to load this settings file) => swallow stdin and exit 0. The
// hook must NEVER fail or block: claude is waiting on it.
export const HOOK_SCRIPT = `#!/bin/sh
# OCP TUI streaming sink — claude fires this per MessageDisplay block and BLOCKS on it.
# Write and exit. Never do work here.
[ -n "\$OCP_TUI_STREAM_FILE" ] || exec cat >/dev/null
{ cat; printf '\\n'; } >> "\$OCP_TUI_STREAM_FILE"
`;
// The --settings payload registering the hook. Static: no per-request data.
export function buildStreamSettings(hookScriptPath) {
return { hooks: { MessageDisplay: [{ hooks: [{ type: "command", command: hookScriptPath }] }] } };
}
export const hookScriptPath = (streamDir) => `${streamDir}/md-hook.sh`;
export const streamSettingsPath = (streamDir) => `${streamDir}/settings.json`;
// One file per session-id. For a pre-booted (warm) pane the session-id is fixed at boot,
// so this path is knowable at boot — which is what keeps the pool compatible.
export const streamFilePath = (streamDir, sessionId) => `${streamDir}/${sessionId}.jsonl`;
// Atomic write: temp file + rename (same-directory, same-filesystem, so rename is atomic on
// POSIX). A process killed mid-`writeFileSync` leaves the TEMP file half-written, never the
// real path — `path` always names either the old complete content or the new complete
// content, never a torn one. That matters specifically for md-hook.sh: it is SYNCHRONOUS
// (claude blocks on every fire), so a truncated script would still pass `existsSync`, still
// get exec'd, and fail/hang on every single MessageDisplay fire with no operator-visible
// symptom short of streaming going silently dead (F7's streamZeroDeltaTurns is the backstop
// for exactly that). Mirrors ensureTuiCwdTrusted's tmp+renameSync pattern in session.mjs.
function writeFileAtomic(path, content, mode) {
const tmp = `${path}.${process.pid}.tmp`;
writeFileSync(tmp, content, { mode });
renameSync(tmp, path);
}
// Write the static hook script + settings file into `streamDir`. UNCONDITIONAL, not
// write-if-missing: these files persist across OCP restarts at `streamDir`, so a host that
// booted once under an older version and never had its stream dir cleared would otherwise be
// silently stuck on a stale HOOK_SCRIPT / buildStreamSettings() forever — no future OCP
// upgrade could ever reach it. Safe to call every boot: the content is static (no per-request
// data), so a same-content rewrite is the overwhelmingly common case and costs two tiny
// atomic writes, not a per-turn expense. Returns the settings path to hand to `claude
// --settings`.
export function prepareStreamHook(streamDir) {
mkdirSync(streamDir, { recursive: true });
const script = hookScriptPath(streamDir);
const settings = streamSettingsPath(streamDir);
writeFileAtomic(script, HOOK_SCRIPT, 0o700);
writeFileAtomic(settings, JSON.stringify(buildStreamSettings(script), null, 2), 0o600);
return settings;
}
// Parse newly-appended sink lines. `consumed` is the number of COMPLETE lines already
// taken; only lines terminated by "\n" are complete, so a payload caught mid-write stays
// unconsumed until its terminator lands. Returns the fresh MessageDisplay payloads plus
// the new consumed count. Pure — the caller owns the cursor.
export function parseDeltaChunk(text, consumed = 0) {
const lines = String(text ?? "").split("\n");
const complete = lines.slice(0, -1); // the tail after the last "\n" is a partial line
const deltas = [];
for (const line of complete.slice(consumed)) {
const t = line.trim();
if (!t) continue;
try {
const o = JSON.parse(t);
if (o && o.hook_event_name === "MessageDisplay" && typeof o.delta === "string") deltas.push(o);
} catch { /* not ours / not parseable — skip, never throw into the request path */ }
}
return { deltas, consumed: complete.length };
}
// ── The assembler: hook deltas → client bytes, with the honesty gates intact ──
//
// Two jobs, both load-bearing.
//
// 1. THE AUTH-BANNER HOLDBACK (C-1 / issue #133 must survive streaming).
// The interactive CLI renders an auth failure as ordinary assistant TEXT — so an
// expired-credential turn fires MessageDisplay with the BANNER as its delta, and a
// naive forwarder would stream "Please run /login · API Error: 401 …" to the client as
// a normal answer, exactly the silent-error case C-1 exists to prevent.
// detectTuiUpstreamError() classifies a WHOLE message, so it cannot be run per-delta.
// Instead we HOLD BACK the first `holdbackChars` characters. The default detector only
// ever fires on a message of <= 100 chars (TUI_ERR_MAX_LEN — real banners are 69 and 73),
// so once the TRIMMED accumulation EXCEEDS 100 chars the final text cannot be a banner by
// that detector's own length rule, and releasing is safe. An answer that never exceeds the
// holdback is simply delivered whole at terminal — i.e. exactly today's buffered
// behaviour, gates and all.
// THE GUARANTEE HAS TWO HALVES, both required — neither alone is sufficient:
// (i) Nothing is emitted for a message until its trimmed accumulation exceeds the
// detector's max banner length. This is what keeps the FIRST message of a turn
// safe: a banner-length message can never clear the holdback.
// (ii) Once a message boundary follows an emit (`restartedAfterEmit`), push() stops
// emitting ENTIRELY for the rest of the turn — a SECOND message (e.g. an
// auth-failure banner rendered mid-turn, after tool-using prose already streamed)
// gets zero bytes forwarded, not just a fresh holdback of its own. finalize() then
// refuses the whole turn (SSE error frame, no cache) precisely because the first
// message's bytes are unretractable and unverifiable against T. Without this half,
// (i) alone only protects the FIRST message per turn — see F1.
// ⚠️ Soundness is w.r.t. the DEFAULT detector. An operator who REPLACES it via
// CLAUDE_TUI_ERROR_PATTERNS with a pattern that can match a longer message must raise
// OCP_TUI_STREAM_HOLDBACK past their longest banner; server.mjs warns at boot. That is the
// one case (i) does not cover — (ii) still applies regardless. Even past both, the
// terminal gate still refuses to cache a banner and still ends the stream on an SSE error
// frame rather than finish_reason:"stop" — the holdback is the first of two layers, not
// the only one.
//
// 2. MESSAGE SCOPING (keeps `concat === T` the RIGHT assertion).
// The transcript's T is extractLatestAssistantText() — the LAST text-bearing assistant
// entry, not every assistant entry. A tool-using turn therefore has TWO messages
// (prose → tool_use → answer) and T is only the second. So the assembler scopes to the
// CURRENT message_id: when a new message_id appears and NOTHING has been emitted yet,
// the held text is DISCARDED — the transcript is about to discard it too, so this keeps
// us byte-identical to the buffered path instead of streaming prose the buffered path
// would have dropped. When a new message_id appears AFTER we have already emitted, the
// bytes are gone and cannot be retracted: finalize() then reports !ok and the caller
// fails the turn loudly (SSE error frame, no cache, counted on /health). Fail-loud is
// the correct posture — a proxy that silently serves text the transcript disagrees with
// is the exact class of bug ALIGNMENT.md exists to prevent.
// Sentinel for "no message seen yet". Deliberately not null/undefined — see the constructor.
const NO_MESSAGE_YET = Symbol("no-message-yet");
export class TuiDeltaAssembler {
constructor({ holdbackChars = DEFAULT_HOLDBACK_CHARS, detectError = detectTuiUpstreamError } = {}) {
this.holdbackChars = holdbackChars;
this.detectError = detectError;
this.emitted = ""; // bytes ALREADY written to the client — unretractable
this.pending = ""; // held back, not yet written
this.released = false;
// NOT null: a payload may legitimately carry message_id === null, and if the sentinel were
// also null the FIRST such payload would compare equal to it, register no boundary, and
// leave `messages` at 0 — which used to disarm the restartedAfterEmit guard below entirely.
// A unique object is === to nothing a JSON payload can produce, so the first fire ALWAYS
// registers as message 1, whatever its message_id is (or isn't).
this.messageId = NO_MESSAGE_YET;
this.deltas = 0; // hook fires seen
this.messages = 0; // distinct message_ids seen
this.restartedAfterEmit = false;
}
// All hook bytes for the CURRENT message (emitted + still held).
get full() { return this.emitted + this.pending; }
// Feed one MessageDisplay payload. Returns the text to emit NOW, or null (held back).
push(payload) {
const delta = payload && typeof payload.delta === "string" ? payload.delta : "";
const mid = payload ? payload.message_id : null;
if (mid !== this.messageId) {
this.messageId = mid;
this.messages++;
if (this.emitted === "") {
this.pending = ""; // safe: the transcript will drop this message too
} else {
// A boundary while bytes are ALREADY out is unrecoverable, full stop — the count of
// messages seen so far is irrelevant. The old `else if (this.messages > 1)` guard was
// the sole reason a null-message_id first payload could disarm F1: it left `messages`
// at 0, so the real boundary evaluated 1 > 1 === false and never armed. The invariant
// is "a boundary occurred while emitted !== ''", and that is exactly what this says.
this.restartedAfterEmit = true; // unrecoverable — finalize() will refuse the turn
}
}
this.deltas++;
// F1: once a message boundary has followed an emit, the turn is ALREADY unrecoverable —
// finalize() will refuse it (see restartedAfterEmit above). `this.released` stays true
// from the FIRST message's release and, uncorrected, lets every later message's deltas
// stream straight through unfiltered — exactly the auth-banner-mid-turn leak this class
// exists to prevent. Stop emitting HERE, permanently, for the rest of the turn: there is
// nothing left to gain from continuing to forward bytes for a turn that will be refused,
// and every byte forwarded now is one more the client cannot be told to un-see.
if (this.restartedAfterEmit) return null;
if (!delta) return null;
if (this.released) {
this.emitted += delta;
return delta;
}
this.pending += delta;
// Release only once the TRIMMED accumulation is past the banner detector's reach.
// detectTuiUpstreamError() trims before measuring length (TUI_ERR_MAX_LEN is a trimmed-
// length bound), so gating release on the UNTRIMMED pending.length let a run of >
// holdbackChars whitespace trim down to "" — detectError("") sees nothing to classify,
// returns null, and release fires with the holdback never having actually screened
// anything. Trimming here keeps both sides of the check talking about the same string.
if (this.pending.trim().length > this.holdbackChars && this.detectError(this.pending) == null) {
const out = this.pending;
this.pending = "";
this.released = true;
this.emitted += out;
return out;
}
return null;
}
// Reconcile against the AUTHORITATIVE transcript text T. Call only AFTER the truncation
// and auth-banner gates have passed. Returns:
// { ok:true, tail, exact } — tail is the remaining text to emit (may be ""). `exact`
// is concat(deltas) === T; when false we still serve exactly
// T, having topped up from the transcript, and the caller
// counts a topUp.
// { ok:false, ... } — what we already emitted is NOT a prefix of T. The client
// holds bytes the transcript disagrees with; the caller must
// NOT cache and must end the stream on an SSE error frame.
finalize(T) {
const text = typeof T === "string" ? T : "";
const full = this.full;
if (!text.startsWith(this.emitted)) {
return { ok: false, tail: null, exact: false, emitted: this.emitted.length, transcript: text.length };
}
return {
ok: true,
tail: text.slice(this.emitted.length),
exact: full === text,
emitted: this.emitted.length,
transcript: text.length,
};
}
}
+22 -1
View File
@@ -73,6 +73,17 @@ export function isTerminalLine(obj) {
// transcript holding one logical exchange). If a future warm-pool ever reuses a
// session WITHOUT a fresh session-id / clear, earlier-turn text could leak — that
// author must add user-line scoping here. See spec §7.2.
//
// STATUS (warm pool, lib/tui/pool.mjs — the "future warm-pool" this note anticipated):
// the pool does NOT reuse sessions, so the precondition above still holds and no
// user-line scoping was added. Each pooled pane is booted with its OWN fresh
// randomUUID() --session-id (bootTuiPane) and is SINGLE-USE: it serves exactly one turn
// and is then killed and replaced. One session still means one logical exchange, so the
// last assistant entry is still that request's answer.
// The warning therefore stands UNCHANGED for anyone who later wants a pane to serve a
// SECOND turn (or to reset one with /clear and reuse it): that is a leak, and it needs
// user-line scoping HERE before it can be safe. Do not relax pool.mjs's single-use rule
// without doing that work first.
export function extractLatestAssistantText(events) {
let text = "";
for (const ev of events) {
@@ -256,11 +267,21 @@ export function detectTuiUpstreamError(text, patternsRaw = process.env.CLAUDE_TU
// Resolution: pass an explicit `transcriptPath` (used by unit tests), OR pass
// `home` + `sessionId` to resolve by glob each poll (production) — the transcript
// file does not exist until the turn starts, so resolution happens inside the loop.
export async function readTuiTranscript({ transcriptPath: p, home, sessionId, wallclockMs = 120000, pollMs = 250 }) {
// `abortSignal` (optional): when it fires, stop waiting and throw TuiAbortError. The one
// caller that passes it is the STREAMING TUI path, which ties it to the client's socket:
// a client that disconnects mid-turn should not leave the pane running (and the caller's
// concurrency slot held) until the turn or the 120s cap ends. runTuiTurn's finally does the
// teardown. Omitted => the loop is byte-for-byte the pre-streaming loop.
export async function readTuiTranscript({ transcriptPath: p, home, sessionId, wallclockMs = 120000, pollMs = 250, abortSignal = null }) {
const deadline = Date.now() + wallclockMs;
let lastText = "";
let lastEntrypoint = null;
while (Date.now() < deadline) {
if (abortSignal && abortSignal.aborted) {
const err = new Error("tui_aborted: client disconnected before the turn completed");
err.name = "TuiAbortError";
throw err;
}
const resolved = p || findTranscriptPath(home, sessionId);
if (resolved && existsSync(resolved)) {
const events = parseTranscriptLines(readFileSync(resolved, "utf8"));
+371 -7
View File
@@ -22,6 +22,8 @@
* CLAUDE_MAX_CONCURRENT — max concurrent claude processes, -p/stream-json path (default: 8)
* CLAUDE_MAX_QUEUE — max requests waiting for a -p slot before HTTP 429 (default: 16)
* OCP_TUI_MAX_CONCURRENT — max concurrent interactive TUI turns, TUI-mode path (default: 2)
* OCP_TUI_POOL_SIZE — pre-booted warm `claude` panes held for TUI-mode (default: 0 = off;
* max 4). Each is a live idle process; cuts ~3-4s per request.
* OCP_SPAWN_REAL_HOME — "1" forces the -p spawn to use the real HOME (disables the
* latency spawn-home isolation; default: isolated when a token exists)
* CLAUDE_BREAKER_THRESHOLD — failures in window before circuit opens (default: 6)
@@ -32,7 +34,7 @@
* CLAUDE_HEARTBEAT_INTERVAL — SSE heartbeat interval in ms on streaming path (default: 0 = disabled)
*/
import { createServer } from "node:http";
import { spawn, execFileSync } from "node:child_process";
import { spawn, execFileSync, spawnSync } from "node:child_process";
import { randomUUID, timingSafeEqual } from "node:crypto";
import { readFileSync, readdirSync, accessSync, existsSync, constants, chmodSync, statSync, mkdirSync, writeFileSync, rmSync } from "node:fs";
import { fileURLToPath } from "node:url";
@@ -41,9 +43,11 @@ import { homedir } from "node:os";
import { validateKey, recordUsage, getUsageByKey, getUsageTimeline, getRecentUsage, createKey, listKeys, revokeKey, closeDb, checkQuota, updateKeyQuota, getKeyQuota, findKey, cacheHash, getCachedResponse, setCachedResponse, clearCache, getCacheStats, hasCacheControl, singleflight, getInflightStats } from "./keys.mjs";
import { DEFAULT_PORT } from "./lib/constants.mjs";
import { isLoopbackBind } from "./lib/net.mjs";
import { runTuiTurn, reapStaleTuiSessions, resolveTuiHome } from "./lib/tui/session.mjs";
import { runTuiTurn, reapStaleTuiSessions, resolveTuiHome, bootTuiPane, tuiPaneHealthy, poolPaneName, POOL_BOOT_MS } from "./lib/tui/session.mjs";
import { detectTuiUpstreamError } from "./lib/tui/transcript.mjs";
import { TuiSemaphore, SemaphoreAbortError, recordTuiEntrypoint, buildTuiHealthBlock } from "./lib/tui/semaphore.mjs";
import { TuiPanePool, resolvePoolSize, POOL_MAX_SIZE } from "./lib/tui/pool.mjs";
import { TuiDeltaAssembler, DEFAULT_HOLDBACK_CHARS } from "./lib/tui/stream.mjs";
import { createSerialMutex, createTtlCache, isTokenExpiring, orderLabelsLastGoodFirst } from "./lib/spawn-auth.mjs";
const __dirname = dirname(fileURLToPath(import.meta.url));
@@ -349,8 +353,98 @@ const tuiSemaphore = new TuiSemaphore(TUI_MAX_CONCURRENT);
const tuiStats = {
lastEntrypoint: null, // last observed cc_entrypoint from the transcript ("cli" | "sdk-cli" | null)
entrypointMismatches: 0, // count of cli-expected-but-got-other turns
streamTurns: 0, // streamed TUI turns ATTEMPTED (counted before the honesty gates — F6)
streamDeltas: 0, // MessageDisplay hook fires OBSERVED (forwarded + held-back — F6)
streamTopUps: 0, // turns where the delta stream != T but was a safe PREFIX of it
streamDivergences: 0, // turns REFUSED: emitted bytes were not a prefix of T
streamZeroDeltaTurns: 0, // streamed turns where the hook fired ZERO times (F7 — the hook is
// dead, not just one fire dropped; distinct from streamTopUps)
};
// ── TUI real streaming (backlog #2) — opt-in; default OFF ────────────────
// When ON *and* TUI_MODE is on *and* the client asked for stream:true, the turn is emitted
// as real SSE delta.content chunks as claude renders them, sourced from claude's own
// MessageDisplay hook (lib/tui/stream.mjs). When OFF, the buffered
// callClaudeTui → streamStringAsSSE path below is byte-for-byte unchanged — the spawn does
// not even get --settings. Opt-in is deliberate: the buffered path is stable production.
//
// Honest expectation (docs/plans/2026-07-13-tui-latency/streaming-spike.md): this moves the
// FIRST byte, not the last. A consumer that must parse a complete reply gains nothing; a
// progressively-rendering chat UI gains the ~4s between first delta and last. It does not
// move the ~6s TTFT floor of TUI mode.
const TUI_STREAM = process.env.OCP_TUI_STREAM === "1";
const TUI_STREAM_DIR = process.env.OCP_TUI_STREAM_DIR || `${process.env.HOME}/.ocp-tui/stream`;
// First-bytes holdback — the auth-banner gate's (C-1) survival mechanism under streaming.
// See TuiDeltaAssembler: nothing is emitted for a message until its TRIMMED accumulation
// exceeds this, which puts it out of the default banner detector's <=100-char reach — the
// FIRST of the two halves of the guarantee (see the assembler's class comment for the second:
// no further emission at all once a message boundary follows an emit). Only raise it.
const TUI_STREAM_HOLDBACK = parseInt(process.env.OCP_TUI_STREAM_HOLDBACK || String(DEFAULT_HOLDBACK_CHARS), 10);
if (TUI_MODE && TUI_STREAM && process.env.CLAUDE_TUI_ERROR_PATTERNS != null && TUI_STREAM_HOLDBACK <= DEFAULT_HOLDBACK_CHARS) {
// The holdback's FIRST-MESSAGE half (see TuiDeltaAssembler) is sound for the DEFAULT
// auth-banner detector (which cannot match a message longer than 100 chars). An
// operator-supplied pattern set has no such bound, so a banner longer than the holdback
// could reach the client before the terminal gate rejects the turn. (The second half — no
// further emission once a message boundary follows an emit — holds regardless of the
// detector; this warning is only about the first-message case.)
console.error(
`[tui] WARNING: OCP_TUI_STREAM=1 with a custom CLAUDE_TUI_ERROR_PATTERNS and holdback=${TUI_STREAM_HOLDBACK}.\n` +
" The streaming holdback's first-message coverage is sound only against the DEFAULT banner\n" +
" detector (<=100 chars). Raise OCP_TUI_STREAM_HOLDBACK above your longest custom banner, or\n" +
" the first chars of one could be streamed before the end-of-turn gate refuses the turn."
);
}
// ── Warm pane pool (docs/plans/2026-07-13-tui-latency #3) — opt-in; default OFF ─────────
// OCP_TUI_POOL_SIZE=0 (default) => tuiPool is null => runTuiTurn's cold-boot path is
// byte-for-byte unchanged. Set it to N (clamped to POOL_MAX_SIZE) to keep N pre-booted
// `claude` panes warm, each SINGLE-USE (see lib/tui/pool.mjs for why single-use is the
// load-bearing rule, and lib/tui/session.mjs for the POOL/REAPER INVARIANT).
//
// Default-off is deliberate on a stable production path: a warm pane is a LIVE idle
// `claude` process held whether or not a request ever arrives, so the operator must opt
// in to that standing cost. Measured saving when on (this host, Sonnet 4.6, --effort low):
// end-to-end p50 10.17 s (n=6, pool off) -> 6.00 s (n=12 warm hits), i.e. -41%.
// cli.js does NOT perform this operation (Class B, OCP-owned TUI spawn) — see ADR 0007.
const TUI_POOL_SIZE = TUI_MODE ? resolvePoolSize(process.env.OCP_TUI_POOL_SIZE) : 0;
const tuiPool = TUI_POOL_SIZE > 0
? new TuiPanePool({
size: TUI_POOL_SIZE,
// The POOL mints the pane's identity, not bootTuiPane: the tmux session exists the
// instant the boot starts, so the pool must be able to name (hence spare, hence kill)
// it before then. Name is derived from the session-id, so `tmux ls` correlates to the
// transcript file <HOME>/.claude/projects/*/<sessionId>.jsonl.
mintPane: () => {
const sessionId = randomUUID();
return { sessionId, name: poolPaneName(PORT, sessionId) };
},
bootPane: (model, ident) => bootTuiPane({
model,
claudeBin: CLAUDE,
home: TUI_HOME,
realHome: process.env.HOME,
cwd: TUI_CWD,
port: PORT,
entrypointMode: TUI_ENTRYPOINT,
sessionId: ident.sessionId,
name: ident.name,
requireReady: true, // a pane that never reached its input bar must not be enlisted
bootMs: POOL_BOOT_MS, // background pre-boot — no client is blocked, so be patient
// Warm panes must carry the MessageDisplay hook too, or every pool HIT would
// silently fall back to buffered while every MISS streamed — the two paths have to
// spawn identically (F4). Gated on TUI_STREAM, the deployment-wide switch — NOT on any
// particular request's stream:true/false, which does not exist yet at pre-boot time.
// The runTuiTurn cold-boot call site (callClaudeTui, below) mirrors this exact gate for
// the same reason. bootTuiPane derives the sink from the pane's own session-id, which is
// minted above, so nothing request-specific is baked in at pre-boot time.
streamDir: TUI_STREAM ? TUI_STREAM_DIR : null,
}),
killPane: (name) => { try { spawnSync(process.env.OCP_TUI_TMUX_BIN || "tmux", ["kill-session", "-t", name]); } catch { /* already gone */ } },
paneHealthy: (name) => tuiPaneHealthy((args) => spawnSync(process.env.OCP_TUI_TMUX_BIN || "tmux", args, { encoding: "utf8" }), name),
log: (level, event, data) => logEvent(level, event, data),
})
: null;
// ── FIX ③ (latency): default-path (-p / stream-json) spawn-home isolation ──────────────
// PROBLEM (measured, not theoretical): OCP's default spawn inherits the operator's real HOME
// (loading the global ~/.claude — plugins, skills, hooks) and runs with cwd=~/ocp (loading the
@@ -770,19 +864,45 @@ const cacheCleanupInterval = setInterval(() => {
// mechanism and the 15-min cadence makes the window negligible).
// Gated on TUI_MODE — zero effect (no kill-server, no list-sessions) when TUI is off.
// cli.js does NOT perform this operation (Class B, OCP-owned TUI spawn) — see ADR 0007.
//
// WARM POOL INTERACTION (the crux — see the POOL/REAPER INVARIANT in lib/tui/session.mjs).
// A warm pooled pane is one of OUR OWN ocp-tui-<port>-* sessions that is alive and idle BY
// DESIGN, and this sweep fires precisely when the instance is idle — i.e. exactly when the
// pool is full. Two things are therefore required, and both are done here:
// (a) DRAIN the pool BEFORE the sweep. Zombie reaping is possible ONLY via kill-server,
// and a live pooled pane suppresses kill-server (it is a live child of the tmux
// server). A permanently-full pool would otherwise permanently disable the very
// thing this tick exists to do. Draining costs one pane re-boot per tick (~1.2 s of
// background work every 15 min) and is invisible to callers: a request landing in the
// drain→refill gap simply MISSES the pool and takes today's cold path.
// (b) Pass the pool's live registry as `spare` anyway. After (a) it is empty, so this is
// belt-and-braces — it makes it impossible for THIS call site (or a future one) to
// kill a live pooled pane even if the drain were ever removed or reordered.
// RESIDUAL (unchanged in kind from the pre-pool code, and explicitly accepted there): a
// request arriving in the narrow window between the idle-check and kill-server has its pane
// torn down and fails cleanly via runTuiTurn's honesty gates. The drain widens that window
// by the cost of N kill-session calls (single-digit ms), not materially.
const TUI_REAP_INTERVAL_MS = 15 * 60 * 1000;
const tuiReapInterval = TUI_MODE ? setInterval(() => {
if (tuiSemaphore.inflight > 0 || tuiSemaphore.queued > 0) return; // a turn is live — defer
try {
const drained = tuiPool ? tuiPool.drain() : 0;
// F7 fix: scope to THIS instance's own port; a sibling ocp-tui-<otherPort>-* session
// (a second OCP instance on the same host) is treated as foreign, same as olp-tui-*.
// includeLegacy is NOT set here — see reapStaleTuiSessions' comment: the periodic sweep
// conservatively treats any lingering bare-prefix legacy session as foreign so it can
// never trigger kill-server on a steady-state tick; only the one-time boot reap below
// claims legacy-shaped zombies.
const n = reapStaleTuiSessions({ port: PORT });
if (n) logEvent("info", "tui_reaped_stale_sessions", { count: n, trigger: "periodic" });
const n = reapStaleTuiSessions({ port: PORT, spare: tuiPool ? tuiPool.liveNames() : null });
if (n || drained) {
logEvent("info", "tui_reaped_stale_sessions", { count: n, poolDrained: drained, trigger: "periodic" });
}
} catch (e) { logEvent("error", "tui_periodic_reap_failed", { error: e.message }); }
finally {
// Refill in the background regardless of how the sweep went — a throw mid-sweep must not
// leave the pool permanently paused (it would silently degrade to the cold path forever).
if (tuiPool) { try { tuiPool.resume(); } catch { /* best effort */ } }
}
}, TUI_REAP_INTERVAL_MS) : null;
if (tuiReapInterval && typeof tuiReapInterval.unref === "function") tuiReapInterval.unref();
@@ -1283,7 +1403,14 @@ async function callClaude(model, messages, conversationId, keyName, res) {
// Authority: claude CLI v2.1.158 interactive mode (cc_entrypoint=cli).
// SECURITY: A-path single-user ONLY — home is NOT isolation (see ADR 0007).
// `res` (optional, F2) is the client's http.ServerResponse — see closeSignalFor.
async function callClaudeTui(model, messages, _conversationId, _keyName, res) {
//
// `streamCtx` (optional, OCP_TUI_STREAM): { emit(text), signal } — when present the turn is
// ALSO streamed live via claude's MessageDisplay hook. The contract is unchanged: this still
// returns the TRANSCRIPT's text (T), the honesty gates still run on T before anything is
// committed, and the cache still stores T — never the concatenated deltas. streamCtx.emit is
// the SSE sink; streamCtx.signal is the client's disconnect signal, which tears the pane down
// mid-turn instead of holding the semaphore slot for a dead socket.
async function callClaudeTui(model, messages, _conversationId, _keyName, res, streamCtx = null) {
const cliModel = MODEL_MAP[model] || model;
const prompt = messagesToPrompt(messages); // includes system as [System] inline
recordModelRequest(cliModel, prompt.length);
@@ -1311,6 +1438,23 @@ async function callClaudeTui(model, messages, _conversationId, _keyName, res) {
// release() runs in a finally so any throw from runTuiTurn (tmux spawn failure,
// paste-not-landed) OR from the honesty gates below (truncation / error banner) can NEVER
// leak a slot. tuiSemaphore.inflight feeds /health.
// Streaming assembler (null when OCP_TUI_STREAM is off — then runTuiTurn gets no onDelta,
// spawns no hook, and behaves byte-for-byte as before). It owns the auth-banner holdback
// and the message scoping; see lib/tui/stream.mjs.
const assembler = streamCtx ? new TuiDeltaAssembler({ holdbackChars: TUI_STREAM_HOLDBACK }) : null;
// F6: counted here — the moment a streamed turn is ATTEMPTED — not after the honesty gates
// below. A turn refused by the truncation or auth-banner gate is exactly the turn an operator
// most wants visible in streamTurns; counting only turns that reached the gates made
// streamDivergences/streamTurns silently exclude its own worst cases from the denominator.
if (assembler) tuiStats.streamTurns++;
const onDelta = assembler
? (payload) => {
const out = assembler.push(payload);
tuiStats.streamDeltas++; // every hook fire OBSERVED, not just forwarded ones — see the
// /health field doc in lib/tui/semaphore.mjs (F6)
if (out) streamCtx.emit(out); // released past the holdback — safe to show the client
}
: null;
try {
const { text, entrypoint, truncated } = await runTuiTurn({
prompt,
@@ -1323,6 +1467,27 @@ async function callClaudeTui(model, messages, _conversationId, _keyName, res) {
// different port never collides with this instance's reap/kill-server logic.
wallclockMs: TUI_WALLCLOCK_MS,
entrypointMode: TUI_ENTRYPOINT,
// Warm pane pool (null unless OCP_TUI_POOL_SIZE > 0 → today's cold path exactly).
// A pooled pane is single-use: runTuiTurn kills it in its finally like any other.
pool: tuiPool,
// Only observe when the pool is ON — with it off (the default) no new log line is
// emitted, so the disabled path stays byte-for-byte today's, logs included.
onPane: tuiPool
? ({ warm }) => logEvent("info", warm ? "tui_pool_hit" : "tui_pool_miss",
{ model: cliModel, warmRemaining: tuiPool.warm })
: null,
onDelta,
// Gated on TUI_STREAM (the deployment-wide switch), NOT on `assembler` (this REQUEST's
// stream:true/false) — F4 fix. The pool's bootPane closure above installs the hook on
// every warm pane whenever TUI_STREAM is on, regardless of what any given future request
// asks for (a pre-booted pane cannot know that yet); the cold path must match, or a
// stream:false request gets --settings on a pool HIT and not on a pool MISS — two
// different spawn argvs for the identical request, which this project's alignment/billing
// posture cannot tolerate. Whether the hook's OUTPUT is actually consumed for THIS turn is
// decided downstream by `onDelta` (null when assembler is null), so a non-streaming
// request still never polls or emits — it just spawns identically either way.
streamDir: TUI_STREAM ? TUI_STREAM_DIR : null,
abortSignal: streamCtx ? streamCtx.signal : null,
});
// ── Honesty gates (issue #133) ─ run BEFORE recordModelSuccess / cache write-back.
// A throw here propagates to the catch below (recordModelError + reject), so the
@@ -1347,6 +1512,59 @@ async function callClaudeTui(model, messages, _conversationId, _keyName, res) {
throw new Error("tui_upstream_error: claude CLI returned an in-session error banner instead of an answer");
}
// ── Streaming safety net — the transcript is the authority, the deltas are the mirror.
// Runs AFTER the two gates above (so a truncated turn or an auth banner is never
// reconciled, let alone flushed) and BEFORE recordModelSuccess / the caller's cache
// write. Three outcomes:
// exact — concat(deltas) === T. The invariant held; emit whatever is still held
// back (a short answer never passes the holdback, so this is its whole text).
// top-up — what we emitted is a strict PREFIX of T but the deltas did not add up to
// it (a dropped/late fire). We serve exactly T by emitting the missing tail;
// the client still gets the right answer. Counted, and visible on /health.
// divergence— we already emitted bytes that are NOT a prefix of T. The client is holding
// text the transcript disagrees with and it cannot be retracted. REFUSE the
// turn: throw → SSE error frame, no cache, no success. Serving on would be
// exactly the "silently serve wrong text" failure this gate exists to stop.
// (Known trigger: a tool-using turn whose pre-tool prose exceeded the
// holdback — the transcript keeps only the LAST assistant message, so the
// prose we streamed is text T does not contain.)
if (assembler) {
// F7: a total hook failure (a claude version bump stops honoring --settings, or a
// truncated md-hook.sh per F3) produces zero fires for every turn, finalize() still
// reports ok:true/exact:false (the transcript alone carries the whole answer), and the
// turn succeeds NORMALLY — degrading to buffered with no error, no divergence, nothing
// but streamTopUps climbing (which the comment above calls "benign"). That is
// indistinguishable from one late fire dropped unless it is counted separately.
if (assembler.deltas === 0) {
tuiStats.streamZeroDeltaTurns++;
logEvent("warn", "tui_stream_zero_deltas", { model: cliModel });
}
const rec = assembler.finalize(text);
if (!rec.ok) {
tuiStats.streamDivergences++;
logEvent("error", "tui_stream_divergence", {
model: cliModel,
// The dominant cause in practice: a TOOL-USING turn whose pre-tool prose exceeded the
// holdback and was already streamed. The transcript keeps only the LAST assistant
// message, so that prose is text T does not contain. Remedy for such a deployment:
// raise OCP_TUI_STREAM_HOLDBACK above the model's typical narration length (later first
// chunk, but the prose stays held back and is then correctly discarded), or leave
// OCP_TUI_STREAM off. See README + ADR 0007 (2026-07-13 amendment).
reason: assembler.restartedAfterEmit ? "multi_message_after_emit (tool-use turn?)" : "delta_transcript_mismatch",
emittedChars: rec.emitted, transcriptChars: rec.transcript,
deltas: assembler.deltas, messages: assembler.messages,
});
throw new Error("tui_stream_divergence: streamed text is not a prefix of the transcript; refusing to serve it");
}
if (!rec.exact) {
tuiStats.streamTopUps++;
logEvent("warn", "tui_stream_topup", {
model: cliModel, emittedChars: rec.emitted, transcriptChars: rec.transcript, deltas: assembler.deltas,
});
}
if (rec.tail) streamCtx.emit(rec.tail);
}
recordModelSuccess(cliModel, 0); // elapsed not measurable here; wallclock at reader level
// Assert the subscription-pool classification. TUI exists to keep cc_entrypoint=cli
// (subscription pool); a silent degrade to sdk-cli (metered Agent SDK pool) would still
@@ -1361,6 +1579,14 @@ async function callClaudeTui(model, messages, _conversationId, _keyName, res) {
}
return text;
} catch (err) {
// A mid-turn client disconnect (streaming path only — abortSignal) is NOT an upstream
// failure: runTuiTurn's finally already tore the pane down, and this finally releases the
// slot. Mirror the queued-disconnect handling above (info, no recordModelError, no
// response) rather than booking a phantom model error against the socket going away.
if (err && err.name === "TuiAbortError") {
logEvent("info", "tui_turn_aborted", { reason: "client_disconnected", model: cliModel });
throw new RequestDisconnectedError("client disconnected mid-turn; TUI pane torn down");
}
recordModelError(cliModel, false);
throw err;
} finally {
@@ -1368,6 +1594,102 @@ async function callClaudeTui(model, messages, _conversationId, _keyName, res) {
}
}
// ── TUI-mode REAL streaming (OCP_TUI_STREAM=1) ──────────────────────────
// The stream:true + TUI_MODE + OCP_TUI_STREAM=1 path. Emits the turn as it is generated,
// from claude's own MessageDisplay hook, instead of buffering it and replaying it with
// streamStringAsSSE.
//
// WIRE SHAPES: every frame below is COPIED from callClaudeStreaming (the -p path) — the role
// chunk, the content-delta chunk, the stop chunk, `[DONE]`, and the post-header
// {error:{message,type}} frame. No new fields, no new shapes. (ALIGNMENT.md Rule 2 / Class B:
// the authority for the wire format is the OpenAI chat/completions streaming spec, adopted by
// ADR 0006; the authority for the TUI spawn is ADR 0007. No cli.js citation applies — see the
// commit body.)
//
// HEADERS ARE SENT EAGERLY, exactly as the -p path does, so the existing heartbeat
// (CLAUDE_HEARTBEAT_INTERVAL) covers the ~6s of silence before the first delta. The cost is
// the same one the -p path already pays: after the headers are out, an upstream failure can
// no longer be a JSON 500, so it is surfaced as the SSE error frame instead (issue #110).
async function callClaudeTuiStreaming(model, messages, conversationId, res, authInfo = {}) {
const id = `chatcmpl-${randomUUID()}`;
const created = Math.floor(Date.now() / 1000);
const t0 = Date.now();
const promptChars = messages.reduce((a, m) => a + contentToText(m.content).length, 0);
let headersSent = false;
function ensureHeaders() {
if (res.writableEnded || res.destroyed) return false;
if (headersSent) return true;
headersSent = true;
res.writeHead(200, {
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache",
"Connection": "keep-alive",
"X-Accel-Buffering": "no",
});
sendSSE(res, {
id, object: "chat.completion.chunk", created, model,
choices: [{ index: 0, delta: { role: "assistant" }, finish_reason: null }],
});
return true;
}
ensureHeaders();
const hb = startHeartbeat(res, HEARTBEAT_INTERVAL, conversationId);
// Held for the WHOLE turn (not just the queue wait): a disconnect must abort the transcript
// wait so runTuiTurn tears the pane down and callClaudeTui's finally frees the slot.
const { signal, detach } = closeSignalFor(res);
const streamCtx = {
signal,
emit(text) {
if (!text) return;
if (!ensureHeaders()) return; // client vanished — drop the write, the turn still unwinds
sendSSE(res, {
id, object: "chat.completion.chunk", created, model,
choices: [{ index: 0, delta: { content: text }, finish_reason: null }],
}, hb);
},
};
try {
// callClaudeTui returns the TRANSCRIPT text T after its honesty gates + the streaming
// reconciliation. Everything the client should see has been emitted by then.
const content = await callClaudeTui(model, messages, conversationId, authInfo.keyName, res, streamCtx);
// Cache T — never the concatenated deltas (mirrors the buffered TUI path).
if (CACHE_TTL > 0 && authInfo.cacheHash) {
try { setCachedResponse(authInfo.cacheHash, model, content); } catch (e) { logEvent("error", "cache_write_failed", { error: e.message }); }
}
if (!res.writableEnded && !res.destroyed) {
sendSSE(res, {
id, object: "chat.completion.chunk", created, model,
choices: [{ index: 0, delta: {}, finish_reason: "stop" }],
}, hb);
res.write("data: [DONE]\n\n");
res.end();
}
try { recordUsage({ keyId: authInfo.keyId, keyName: authInfo.keyName, model, promptChars, responseChars: content.length, elapsedMs: Date.now() - t0, success: true }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); }
} catch (err) {
// Client walked away (queued OR mid-turn): nothing to write to, nothing to record —
// same quiet outcome as every other disconnect path (L1 / F2).
if (err instanceof RequestDisconnectedError) { try { res.end(); } catch {} return; }
try { recordUsage({ keyId: authInfo.keyId, keyName: authInfo.keyName, model, promptChars, responseChars: 0, elapsedMs: Date.now() - t0, success: false }); } catch (e) { logEvent("error", "usage_record_failed", { error: e.message }); }
console.error(`[proxy] error: ${err.message}`);
// Headers are already out (eager, above), so — exactly like the -p path — the failure is
// surfaced as an SSE error frame, NOT a success-looking finish_reason:"stop". This is what
// keeps a truncated turn, an auth banner, or a stream divergence from being served as an
// answer: the client sees an error, and nothing was cached.
if (!res.writableEnded && !res.destroyed) {
sendSSE(res, { error: { message: sanitizeError(err.message), type: "provider_error" } }, hb);
res.write("data: [DONE]\n\n");
res.end();
}
} finally {
hb.stop();
detach();
}
}
// ── SSE heartbeat (opt-in idle watchdog) ────────────────────────────────
// Emits `: keepalive\n\n` SSE comment frames during silent windows on the
// streaming response. Design: docs/superpowers/specs/2026-04-25-47-sse-heartbeat-design.md
@@ -2260,6 +2582,11 @@ async function handleChatCompletions(req, res) {
}
if (stream) {
if (TUI_MODE && TUI_STREAM) {
// TUI-mode REAL streaming (opt-in): emit delta.content chunks as claude renders them,
// via its MessageDisplay hook. The transcript remains authoritative (gates + cache).
return callClaudeTuiStreaming(model, messages, conversationId, res, { keyId: req._authKeyId, keyName: req._authKeyName, cacheHash: req._cacheHash });
}
if (TUI_MODE) {
// TUI-mode: no real token stream — buffer the full turn via callClaudeTui,
// optionally write-back to cache, then replay as chunked SSE.
@@ -2558,9 +2885,17 @@ const server = createServer(async (req, res) => {
// still appears with enabled:false (cheap, harmless) so the shape is stable.
// entrypointMismatches/lastEntrypoint exist so an operator can poll /health to catch a
// silent metered-pool drift (the audit's top risk after the 6/15 billing flip).
// `pool` is a NEW nested field inside the (already additive) tui block: null when the
// warm pool is off (the default), so the disabled shape is unchanged apart from one
// explicit null. Lets the operator confirm hit rate + standing process cost.
//
// streamEnabled + the stream* counters are likewise ADDITIVE (new fields only, same
// grandfathered B.2 rationale — ADR 0006). streamDivergences is the one an operator
// must watch: a non-zero value means a streamed turn was REFUSED because the deltas
// disagreed with the transcript, which is the streaming path's only correctness risk.
tui: buildTuiHealthBlock(
{ enabled: TUI_MODE, entrypointMode: TUI_ENTRYPOINT, maxConcurrent: TUI_MAX_CONCURRENT },
tuiStats, tuiSemaphore,
{ enabled: TUI_MODE, entrypointMode: TUI_ENTRYPOINT, maxConcurrent: TUI_MAX_CONCURRENT, streamEnabled: TUI_MODE && TUI_STREAM },
tuiStats, tuiSemaphore, tuiPool,
),
});
}
@@ -2797,6 +3132,27 @@ function gracefulShutdown(signal) {
if (tuiReapInterval) clearInterval(tuiReapInterval);
closeDb();
// 2b. Drain the warm pane pool. A pooled `claude` is a child of the tmux SERVER, not of
// this node process, so it is NOT in activeProcesses and step 3 below cannot reach it —
// without this explicit drain every warm pane would outlive OCP as an orphan (and the
// pool's in-memory registry dies with the process, so nothing would remember it owned them).
//
// drain() kills the pane that is currently BOOTING too, and it does so SYNCHRONOUSLY. That
// is required, not incidental: step 4 below calls process.exit(0) in THIS SAME TICK whenever
// activeProcesses is empty — which on a TUI host it always is — so any cleanup a boot
// deferred to a .then()/.catch() would simply never run. (That was a real bug: the pool used
// to track in-flight boots as a count, could not name the booting session, and orphaned a
// live authenticated `claude` on every shutdown that landed mid-boot.)
//
// Orphans that survive anyway (SIGKILL, power loss) are still caught by the next instance's
// boot reap — this makes the graceful path clean, it is not the only safety net.
if (tuiPool) {
try {
const drained = tuiPool.drain();
if (drained) logEvent("info", "tui_pool_drained", { count: drained, trigger: "shutdown" });
} catch (e) { logEvent("error", "tui_pool_drain_failed", { error: e.message }); }
}
// 3. Kill all active child processes
for (const proc of activeProcesses) {
try { proc.kill("SIGTERM"); } catch {}
@@ -2866,11 +3222,19 @@ server.listen(PORT, BIND_ADDRESS, () => {
? (TUI_HOME === process.env.HOME ? "env-token (real home — unset OCP_TUI_HOME for credential isolation)" : "env-token (credential-isolated home — no credentials.json)")
: "credentials.json (no CLAUDE_CODE_OAUTH_TOKEN — see Troubleshooting #401)";
console.log(` TUI-mode: ON home=${TUI_HOME} cwd=${TUI_CWD} auth=${tuiAuth} wallclock=${TUI_WALLCLOCK_MS}ms maxConcurrent=${TUI_MAX_CONCURRENT}`);
console.log(TUI_POOL_SIZE > 0
? ` TUI warm pool: ON size=${TUI_POOL_SIZE}${TUI_POOL_SIZE} idle \`claude\` process(es) held warm; first request per model is still a cold MISS`
: ` TUI warm pool: OFF (set OCP_TUI_POOL_SIZE=1..${POOL_MAX_SIZE} to pre-boot panes and cut ~3-4s per request)`);
try {
// F7 fix: scope to THIS instance's own port (see reapStaleTuiSessions). includeLegacy:
// true ONLY here — the one-time boot reap is the designated point to claim orphaned
// bare-prefix ("ocp-tui-<uuid8>") zombie sessions left by a PRE-fix process generation
// of this same instance (no live post-fix instance ever creates that shape again).
// No `spare`: the warm pool is EMPTY at boot (there is no boot-time pre-warm — the pool
// learns its model from the first request), so this reap has no live pane to protect and
// it is exactly what SHOULD claim any ocp-tui-<port>-p* pool orphans left by a previous
// process generation of this instance (POOL/REAPER INVARIANT property 2). If a future
// change ever pre-warms at boot, this call MUST start passing tuiPool.liveNames().
const n = reapStaleTuiSessions({ port: PORT, includeLegacy: true });
if (n) logEvent("info", "tui_reaped_stale_sessions", { count: n });
} catch {}
+940 -4
View File
@@ -23,9 +23,30 @@ process.env.HOME = homedir(); // ensure consistent
let passed = 0;
let failed = 0;
// Pending promises from tests declared `async` but registered through the SYNC `test()` helper.
// 44 tests in this file are written that way. Before this, `test()` called fn(), got a promise back,
// and immediately printed ✓ and incremented `passed` — WITHOUT AWAITING IT. So for every async test:
// - ✓ meant "did not throw synchronously", NOT "passed";
// - a failed assertion escaped as an unhandled rejection, which crashes the process (CI still goes
// red on the non-zero exit) but is NOT counted, so the summary could print "N passed, 0 failed"
// and be wrong.
// The suite's own headline number was therefore not evidence for any async test — including the
// regression guards in this PR. Collected here and awaited before the summary prints.
const pendingAsync = [];
function test(name, fn) {
try {
fn();
const r = fn();
if (r && typeof r.then === "function") {
// Async body: settle it before counting. Do NOT print ✓ yet.
pendingAsync.push(
r.then(
() => { passed++; console.log(`${name}`); },
(e) => { failed++; console.log(`${name}: ${e.message}`); },
),
);
return;
}
passed++;
console.log(`${name}`);
} catch (e) {
@@ -1795,6 +1816,65 @@ test("buildTuiCmd shq-escapes a token containing shell metacharacters (no inject
}
});
// OCP_TUI_EFFORT (TUI latency, docs/plans/2026-07-13-tui-latency): the pane's claude
// must get an EXPLICIT --effort so its effort never depends on which HOME mode
// resolveTuiHome() picked (real-home inherits the operator's settings.json effortLevel;
// env-token scratch inherits claude's built-in default).
test("buildTuiCmd passes --effort low by default (OCP_TUI_EFFORT unset)", () => {
const save = process.env.OCP_TUI_EFFORT;
try {
delete process.env.OCP_TUI_EFFORT;
const cmd = buildTuiCmd("/usr/bin/claude", "m", "sid-eff1", "/home/u", "cli");
assert.ok(cmd.includes("--effort low"), "default must pin --effort low");
} finally {
if (save === undefined) delete process.env.OCP_TUI_EFFORT;
else process.env.OCP_TUI_EFFORT = save;
}
});
test("buildTuiCmd honors an explicit OCP_TUI_EFFORT level (case/space-normalized)", () => {
const save = process.env.OCP_TUI_EFFORT;
try {
process.env.OCP_TUI_EFFORT = " XHigh ";
const cmd = buildTuiCmd("/usr/bin/claude", "m", "sid-eff2", "/home/u", "cli");
assert.ok(cmd.includes("--effort xhigh"), "explicit level must be passed, normalized");
assert.ok(!cmd.includes("--effort low"), "default must not also appear");
} finally {
if (save === undefined) delete process.env.OCP_TUI_EFFORT;
else process.env.OCP_TUI_EFFORT = save;
}
});
test("buildTuiCmd OCP_TUI_EFFORT=inherit omits --effort entirely (pre-flag argv)", () => {
const save = process.env.OCP_TUI_EFFORT;
try {
process.env.OCP_TUI_EFFORT = "inherit";
const cmd = buildTuiCmd("/usr/bin/claude", "m", "sid-eff3", "/home/u", "cli");
assert.ok(!/--effort/.test(cmd), "inherit must not add --effort");
} finally {
if (save === undefined) delete process.env.OCP_TUI_EFFORT;
else process.env.OCP_TUI_EFFORT = save;
}
});
test("buildTuiCmd falls back to --effort low on an invalid OCP_TUI_EFFORT (never reaches argv)", () => {
const save = process.env.OCP_TUI_EFFORT;
const savedErr = console.error;
try {
process.env.OCP_TUI_EFFORT = "ludicrous'; rm -rf /;'";
let warned = "";
console.error = (...a) => { warned = a.join(" "); };
const cmd = buildTuiCmd("/usr/bin/claude", "m", "sid-eff4", "/home/u", "cli");
assert.ok(cmd.includes("--effort low"), "invalid value must fall back to low");
assert.ok(!cmd.includes("ludicrous"), "invalid raw value must NOT reach the shell string");
assert.ok(/invalid OCP_TUI_EFFORT/.test(warned), "must log a warning");
} finally {
console.error = savedErr;
if (save === undefined) delete process.env.OCP_TUI_EFFORT;
else process.env.OCP_TUI_EFFORT = save;
}
});
test("buildTuiCmd OCP_TUI_FULL_TOOLS=1 grants -p-equivalent tool surface (single-user opt-in)", () => {
const save = { ...process.env };
const restore = () => {
@@ -1968,6 +2048,489 @@ test("reaper with includeLegacy=true still spares a sibling instance's port-scop
assert.ok(!calls.includes("kill-server"), "sibling instance's live session still blocks kill-server");
});
// ── TUI warm pane pool (docs/plans/2026-07-13-tui-latency #3) ────────────
import { TuiPanePool, resolvePoolSize, POOL_MAX_SIZE, POOL_MAX_AGE_MS } from "./lib/tui/pool.mjs";
import { poolPaneName as poolName } from "./lib/tui/session.mjs";
// A pool wired to fakes: no tmux, no claude. bootPane resolves on the microtask queue; use
// `await settle()` after a refill() to let the SERIALIZED boot chain run to target.
// `live` models the real tmux server: bootTuiPane creates the session SYNCHRONOUSLY and only
// THEN waits (up to POOL_BOOT_MS) for the input bar, so the fake boot registers the session
// immediately and only afterwards resolves. `opts.hold` keeps a boot in that mid-flight window
// so tests can act on a pane that is live-but-not-yet-warm — the state that hid two bugs.
function makeFakePool(opts = {}) {
const killed = [];
const booted = [];
const live = new Set(); // "tmux sessions" that currently exist
let seq = 0;
let clock = 1_000_000;
const healthy = new Set();
// FIFO gate queue — one entry per in-flight held boot. A single `release` slot would be
// OVERWRITTEN by a later boot, so releasing "the first boot" would silently release the
// second instead (and mask the stale-settle bug this harness exists to test).
const gates = [];
const pool = new TuiPanePool({
size: opts.size ?? 2,
maxAgeMs: opts.maxAgeMs ?? POOL_MAX_AGE_MS,
now: () => clock,
mintPane: () => {
const n = ++seq;
return { sessionId: `sid-${n}`, name: `ocp-tui-3456-p${String(n).padStart(8, "0")}` };
},
bootPane: async (model, { sessionId, name }) => {
live.add(name); // session exists NOW
booted.push({ name, model });
if (opts.hold) {
await new Promise((r) => gates.push(r)); // ...stuck waiting for readiness
}
if (opts.bootThrows) { live.delete(name); throw new Error("boom"); }
// A pane whose session was killed while booting can never become ready — exactly what
// the real bootTuiPane does (it throws tui_pane_not_ready).
if (!live.has(name)) throw new Error("tui_pane_not_ready");
healthy.add(name);
return { name, sessionId, model, bootedAt: clock };
},
killPane: (name) => { killed.push(name); healthy.delete(name); live.delete(name); },
paneHealthy: (name) => healthy.has(name),
});
return {
pool, killed, booted, healthy, live,
releaseBoot: () => { const r = gates.shift(); if (r) r(); }, // release the OLDEST held boot
advance: (ms) => { clock += ms; }, at: () => clock,
};
}
const tick = () => new Promise((r) => setImmediate(r));
// Refills are SERIALIZED (one boot at a time, re-kicked on success), so settling the pool
// takes a chain of microtask turns, not one. 40 is far more than POOL_MAX_SIZE needs.
const settle = async () => { for (let i = 0; i < 40; i++) await tick(); };
console.log("\nTUI warm pane pool (acquire / miss / refill / TTL / reaper exemption):");
test("resolvePoolSize: default/garbage/negative disable the pool; size is clamped to POOL_MAX_SIZE", () => {
assert.equal(resolvePoolSize(undefined), 0, "unset => off (byte-for-byte today's cold path)");
assert.equal(resolvePoolSize("0"), 0);
assert.equal(resolvePoolSize("-3"), 0);
assert.equal(resolvePoolSize("banana"), 0, "garbage disables rather than guessing a size");
assert.equal(resolvePoolSize("2"), 2);
assert.equal(resolvePoolSize("99"), POOL_MAX_SIZE, "clamped — never boot an unbounded number of idle claudes");
});
test("pool size 0 is inert: acquire always misses and refill never boots", async () => {
const { pool, booted } = makeFakePool({ size: 0 });
assert.equal(pool.enabled, false);
assert.equal(pool.acquire("m1"), null, "disabled pool always MISSES → caller cold-boots");
pool.refill();
await settle();
assert.equal(booted.length, 0, "a disabled pool must never spawn a process");
});
test("acquire MISSES on an empty pool, and the miss refills for the requested model", async () => {
const { pool, booted } = makeFakePool({ size: 2 });
assert.equal(pool.acquire("sonnet"), null, "first request is always a MISS (no boot-time pre-warm)");
assert.equal(pool.misses, 1);
pool.refill();
await settle();
assert.equal(pool.warm, 2, "refilled to target");
assert.deepEqual(booted.map((b) => b.model), ["sonnet", "sonnet"], "warmed for the model that missed");
});
test("acquire HITS a warm pane, hands it out ONCE, and never returns it (single-use)", async () => {
const { pool } = makeFakePool({ size: 2 });
pool.acquire("sonnet"); pool.refill(); await settle();
assert.equal(pool.warm, 2);
const a = pool.acquire("sonnet");
assert.ok(a && a.name && a.sessionId, "warm pane handed out");
assert.equal(pool.hits, 1);
assert.equal(pool.warm, 1, "the pane LEAVES the registry when acquired");
const b = pool.acquire("sonnet");
assert.notEqual(b.name, a.name, "a pane is NEVER handed out twice — single-use");
assert.notEqual(b.sessionId, a.sessionId, "each pane carries its OWN fresh session-id (transcript.mjs scoping)");
assert.equal(pool.warm, 0);
assert.equal(pool.acquire("sonnet"), null, "exhausted pool MISSES rather than reusing a pane");
});
test("refill is bounded: never more than `size` panes, and concurrent refills do not overshoot", async () => {
const { pool, booted } = makeFakePool({ size: 2 });
pool.acquire("sonnet");
pool.refill(); pool.refill(); pool.refill(); // hammer it
await settle();
assert.equal(pool.warm, 2, "still exactly `size` warm panes");
assert.equal(booted.length, 2, "the _booting guard prevented duplicate boots");
});
// Live finding at size=2: two cold `claude` boots racing an in-flight turn made a refill
// overrun even the generous pool readiness cap. Boots are therefore SERIALIZED.
test("refill boots panes ONE AT A TIME (never two claude cold-boots racing each other)", async () => {
let concurrent = 0, peak = 0;
let seq = 0;
const pool = new TuiPanePool({
size: 3,
mintPane: () => { const n = ++seq; return { sessionId: `s${n}`, name: `p${n}` }; },
bootPane: async (model, { sessionId, name }) => {
concurrent++; peak = Math.max(peak, concurrent);
await new Promise((r) => setImmediate(r)); // simulate boot latency
concurrent--;
return { name, sessionId, model, bootedAt: Date.now() };
},
killPane: () => {},
paneHealthy: () => true,
});
pool.acquire("sonnet");
pool.refill();
await settle();
assert.equal(pool.warm, 3, "chain still reaches the target size");
assert.equal(peak, 1, "at most ONE boot in flight at any moment");
});
test("a FAILED boot does not re-kick the chain (backoff — a broken claude must not spin)", async () => {
const { pool, booted } = makeFakePool({ size: 3, bootThrows: true });
pool.acquire("sonnet");
pool.refill();
await settle();
assert.equal(pool.bootFailures, 1, "counted as a genuine failure (nobody cancelled it)");
assert.equal(booted.length, 1, "exactly ONE attempt — a failure stops the chain, it does not respawn forever");
assert.equal(pool.warm, 0);
assert.equal(pool.booting, 0, "and the booting slot is released, so the next trigger can retry");
});
test("acquire drops an UNHEALTHY warm pane (kills it) and falls through to a MISS", async () => {
const { pool, killed, healthy, booted } = makeFakePool({ size: 1 });
pool.acquire("sonnet"); pool.refill(); await settle();
const dead = booted[0].name;
healthy.delete(dead); // pane died / stopped being input-ready while idle
assert.equal(pool.acquire("sonnet"), null, "a dead pane must MISS, never hang a turn");
assert.ok(killed.includes(dead), "the dead pane is killed, not leaked");
assert.equal(pool.misses, 2);
});
test("acquire drops an EXPIRED warm pane (older than maxAgeMs)", async () => {
const { pool, killed, booted, advance } = makeFakePool({ size: 1, maxAgeMs: 60_000 });
pool.acquire("sonnet"); pool.refill(); await settle();
advance(60_001);
assert.equal(pool.acquire("sonnet"), null, "a pane past its TTL is not handed out");
assert.ok(killed.includes(booted[0].name), "expired pane is killed");
});
test("a model switch drops the wrong-model panes and retargets the pool (--model is fixed at spawn)", async () => {
const { pool, killed, booted } = makeFakePool({ size: 2 });
pool.acquire("sonnet"); pool.refill(); await settle();
const sonnetPanes = booted.map((b) => b.name);
assert.equal(pool.acquire("opus"), null, "different model => MISS (a sonnet pane cannot serve opus)");
assert.equal(pool.warm, 0, "sonnet panes dropped");
for (const p of sonnetPanes) assert.ok(killed.includes(p), "wrong-model pane killed, not leaked");
assert.equal(pool.warmModel, "opus", "pool retargeted to the model actually being asked for");
pool.refill(); await settle();
assert.deepEqual(booted.slice(2).map((b) => b.model), ["opus", "opus"], "refilled for the NEW model");
});
test("a boot that resolves AFTER a drain kills its own pane instead of enlisting it", async () => {
const { pool, live } = makeFakePool({ size: 1 });
pool.acquire("sonnet");
pool.refill(); // boot is in flight...
pool.drain(); // ...pool drained before it resolves (shutdown / reap sweep)
await settle();
assert.equal(pool.warm, 0, "the late pane must NOT be enlisted into a drained pool");
// Assert LIVENESS, not the kill-call COUNT. This assertion used to read
// `assert.equal(killed.length, 1)` — and it PASSED while the orphan it is named after was
// actually present: _cancelBooting kills BY NAME, and at drain time the tmux session does not
// exist yet (bootPane runs on a microtask), so that kill is a NO-OP which still increments the
// counter. "kill was called once" and "a live session is orphaned" were both true at the same
// time. The only honest question is whether the session is dead.
assert.equal(live.size, 0, "it kills itself — no orphan process left behind");
});
test("bootPane failure is counted, never thrown into the request path, and does not wedge refill", async () => {
const { pool } = makeFakePool({ size: 1, bootThrows: true });
pool.acquire("sonnet");
pool.refill();
await settle();
assert.equal(pool.warm, 0);
assert.equal(pool.bootFailures, 1);
assert.equal(pool.booting, 0, "the _booting counter is released on failure (else refill wedges forever)");
assert.equal(pool.acquire("sonnet"), null, "and the caller just MISSES → cold path");
});
test("drain kills every warm pane and pauses refills; resume restarts them", async () => {
const { pool, killed } = makeFakePool({ size: 2 });
pool.acquire("sonnet"); pool.refill(); await settle();
assert.equal(pool.warm, 2);
assert.equal(pool.drain(), 2, "drain reports how many it killed");
assert.equal(pool.warm, 0);
assert.equal(killed.length, 2, "both panes killed — none outlive the drain");
pool.refill(); await settle();
assert.equal(pool.warm, 0, "refill is a NO-OP while drained (paused)");
pool.resume(); await settle();
assert.equal(pool.warm, 2, "resume refills");
});
// ── The crux: pool ↔ reaper coexistence (POOL/REAPER INVARIANT, lib/tui/session.mjs) ──
console.log("\nTUI warm pool ↔ session reaper coexistence:");
test("INVARIANT 1: a LIVE pooled pane is NEVER reaped (it is in the spare set)", () => {
const killed = [];
const live = "ocp-tui-3456-pdeadbeef";
const fakeTmux = (args) => {
if (args[0] === "list-sessions") return { status: 0, stdout: `${live}\nocp-tui-3456-aaaa\n` };
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" };
};
const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456, spare: new Set([live]) });
assert.equal(n, 1, "only the stale turn session was reaped");
assert.ok(!killed.includes(live), "the live warm pane must survive the sweep");
assert.ok(killed.includes("ocp-tui-3456-aaaa"), "a genuinely stale own session is still reaped");
});
test("INVARIANT 2: an ORPHANED pooled pane (pool-shaped but NOT in the spare set) IS reaped", () => {
const killed = [];
// ocp-tui-3456-porphan01 LOOKS pooled but the live registry does not claim it — e.g. left
// behind by a previous process generation, whose in-memory registry died with it.
const fakeTmux = (args) => {
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-porphan01\nocp-tui-3456-plive0001\n" };
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" };
};
const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456, spare: new Set(["ocp-tui-3456-plive0001"]) });
assert.equal(n, 1);
assert.deepEqual(killed, ["ocp-tui-3456-porphan01"], "exemption is by EXACT NAME, never by name shape");
});
test("INVARIANT 2b: with NO spare set (the pre-pool call shape) pool-shaped panes are reaped — fail-safe", () => {
const killed = [];
const fakeTmux = (args) => {
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-pdeadbeef\n" };
if (args[0] === "kill-session") { killed.push(args[args.indexOf("-t") + 1]); return { status: 0 }; }
return { status: 0, stdout: "" };
};
const n = reapStaleTuiSessions({ tmux: fakeTmux, port: 3456 });
assert.equal(n, 1, "omitting `spare` reaps MORE, never less — forgetting it can't leak panes");
assert.deepEqual(killed, ["ocp-tui-3456-pdeadbeef"]);
});
test("INVARIANT 3: kill-server is SUPPRESSED while a live pooled pane is spared", () => {
const calls = [];
const live = "ocp-tui-3456-plive0001";
const fakeTmux = (args) => {
calls.push(args.join(" "));
if (args[0] === "list-sessions") return { status: 0, stdout: `${live}\nocp-tui-3456-aaaa\n` };
return { status: 0, stdout: "" };
};
reapStaleTuiSessions({ tmux: fakeTmux, port: 3456, spare: new Set([live]) });
assert.ok(!calls.includes("kill-server"), "kill-server would kill the live pane (a child of the tmux server)");
});
test("INVARIANT 3b: after a DRAIN the spare set is empty, so kill-server fires again (zombie reaping preserved)", async () => {
// This is the whole reason server.mjs drains BEFORE the periodic sweep: a permanently-full
// pool would otherwise permanently suppress the only mechanism that reaps defunct claudes.
const { pool } = makeFakePool({ size: 2 });
pool.acquire("sonnet"); pool.refill(); await settle();
assert.equal(pool.liveNames().size, 2, "pool is full → the sweep would be suppressed");
pool.drain();
assert.equal(pool.liveNames().size, 0, "drain empties the live registry");
const calls = [];
const fakeTmux = (args) => {
calls.push(args.join(" "));
if (args[0] === "list-sessions") return { status: 0, stdout: "ocp-tui-3456-aaaa\n" };
return { status: 0, stdout: "" };
};
reapStaleTuiSessions({ tmux: fakeTmux, port: 3456, spare: pool.liveNames() });
assert.ok(calls.includes("kill-server"), "kill-server fires post-drain — defunct zombies still get reaped");
});
// ── MID-BOOT: the state that hid M1a + M1b ────────────────────────────────────────────
// bootTuiPane creates the tmux session SYNCHRONOUSLY, then waits up to POOL_BOOT_MS (20s)
// for the input bar. So a pooled session can be LIVE for ~20s before its boot resolves.
// Every reaper test above uses a pool that is either full or drained — never mid-boot.
// That gap is exactly why both bugs shipped past the first round of tests.
test("M1a: a reap tick during an IN-FLIGHT BOOT must not orphan-kill the booting pane", async () => {
const { pool, live } = makeFakePool({ size: 1, hold: true });
pool.acquire("sonnet");
pool.refill();
await tick(); // boot started; session live; NOT yet warm
assert.equal(pool.warm, 0, "not warm yet");
assert.equal(pool.booting, 1, "a boot is in flight");
assert.equal(live.size, 1, "...and its tmux session ALREADY EXISTS");
const bootingName = [...live][0];
assert.ok(pool.liveNames().has(bootingName),
"REGRESSION GUARD: the booting pane MUST be nameable, or the sweep cannot spare it " +
"(the pool used to track in-flight boots as a COUNT and this was empty)");
});
test("M1a: the reap tick's drain kills the booting pane, and resume() starts a FRESH boot", async () => {
const warns = [];
const { pool, live, releaseBoot } = makeFakePool({ size: 1, hold: true });
pool._log = (lvl, ev) => { if (lvl === "warn") warns.push(ev); };
pool.acquire("sonnet");
pool.refill();
await tick();
const first = [...live][0];
// The reap tick, as server.mjs runs it: drain -> reap -> resume.
const drained = pool.drain();
assert.equal(drained, 1, "drain accounts for the booting pane");
assert.equal(live.size, 0, "its tmux session is killed — kill-server can now flush zombies");
assert.equal(pool.liveNames().size, 0, "nothing left to spare, so kill-server is not suppressed");
pool.resume();
await tick();
assert.equal(pool.booting, 1, "resume() started a FRESH boot — the pool is not left empty with nothing scheduled");
assert.notEqual([...live][0], first, "and it is a NEW pane, not the killed one");
// Now let the ORIGINAL (cancelled) boot settle. It rejects with tui_pane_not_ready because
// we killed its session — but that is OUR doing, not a fault.
releaseBoot();
await settle();
assert.equal(pool.bootFailures, 0,
"a cancelled boot must NOT be counted as a bootFailure — that is the WARN operators alert on");
assert.deepEqual(warns, [], "and it must not log tui_pool_boot_failed for a healthy drain");
assert.equal(pool.cancelled, 1, "it is counted as a cancellation instead (counted exactly once)");
});
test("M1b: shutdown drain kills the booting pane SYNCHRONOUSLY — no orphaned claude", async () => {
const { pool, live } = makeFakePool({ size: 1, hold: true });
pool.acquire("sonnet");
pool.refill();
await tick();
assert.equal(live.size, 1, "a live pooled session exists");
// gracefulShutdown: drain() then process.exit(0) IN THE SAME TICK (TUI panes are tmux
// children, not node children, so activeProcesses is empty and the exit is immediate).
// Nothing scheduled on the microtask queue can run. So we assert WITHOUT awaiting.
pool.drain();
assert.equal(live.size, 0,
"REGRESSION GUARD: the pane must be dead BEFORE any await. A .then()-based cleanup would " +
"never run before process.exit and would orphan a live authenticated `claude`.");
});
// M1b, second costume: drain() in the SAME synchronous block as refill(). The tmux session does
// not exist yet at drain time (bootPane runs on a microtask), so _cancelBooting's kill-by-name is
// a no-op — and a `.then` that merely `return`s on a stale generation would then let the boot
// CREATE the session and walk away from it. Not reachable from any current call site, but ADR 0008
// and the reap-tick comment both contemplate a boot-time pre-warm, which is exactly this shape.
// NOTE: deliberately NOT `hold: true`. A held boot never settles, so its `.then` never runs and
// the guard would vacuously pass — the test must let the boot actually SUCCEED, because the bug is
// precisely that a SUCCESSFUL boot on a cancelled generation walks away from its live session.
test("M1b': a boot cancelled BEFORE its session existed is still killed when it settles", async () => {
const { pool, live } = makeFakePool({ size: 1 });
pool.acquire("sonnet"); // miss → learns the model
pool.refill(); // mints the identity; bootPane is queued on a microtask — no session YET
pool.drain(); // SAME sync block: kill-by-name finds nothing to kill (no-op), bumps the gen
assert.equal(live.size, 0, "precondition: the session genuinely did not exist at cancel time");
await tick(); // NOW the boot microtask runs, CREATES the session, and settles on a stale gen
await tick();
assert.equal(live.size, 0,
"REGRESSION GUARD: a stale-generation boot must KILL its pane, not assume _cancelBooting " +
"already did. _cancelBooting kills BY NAME, and the tmux session does not exist until the " +
"boot microtask runs — so a cancellation landing first is a no-op, and a bare `return` here " +
"orphans a live authenticated `claude` that nothing owns.");
});
test("a stale boot settling after drain+resume must not clear the NEW boot's slot", async () => {
const { pool, releaseBoot } = makeFakePool({ size: 1, hold: true });
pool.acquire("sonnet");
pool.refill();
await tick();
pool.drain(); // cancels boot #1 (generation bumped)
pool.resume();
await tick();
assert.equal(pool.booting, 1, "boot #2 owns the slot");
releaseBoot(); // boot #1 finally settles (rejects)
await settle();
assert.equal(pool.booting, 1, "boot #2 STILL owns the slot — a stale settle must not free it");
});
test("a model switch cancels an in-flight boot for the OLD model (kills it, frees the slot)", async () => {
const { pool, live, killed } = makeFakePool({ size: 1, hold: true });
pool.acquire("sonnet");
pool.refill();
await tick();
const sonnetPane = [...live][0];
pool.acquire("opus"); // retarget mid-boot
assert.ok(killed.includes(sonnetPane), "the old model's booting pane is killed, not left to linger");
assert.equal(pool.booting, 0, "and its slot is freed immediately, so the new model can boot now");
assert.equal(pool.warmModel, "opus");
pool.refill();
await tick();
assert.equal(pool.booting, 1, "a boot for the NEW model starts without waiting out the old one");
});
test("a pane handed out for a turn leaves the spare set immediately (so its teardown is authoritative)", async () => {
const { pool } = makeFakePool({ size: 2 });
pool.acquire("sonnet"); pool.refill(); await settle();
const taken = pool.acquire("sonnet");
assert.ok(!pool.liveNames().has(taken.name),
"an acquired pane is the CALLER's — the pool must not also claim it live, or a crashed turn's pane would be spared forever");
});
test("N1: the pool mints ONE identity — the tmux name's hex is the transcript session-id's hex", async () => {
// Without this, `tmux ls` shows a pane whose name has no relation to any transcript file,
// so a live pane cannot be correlated to <HOME>/.claude/projects/*/<sessionId>.jsonl.
const seen = [];
const pool = new TuiPanePool({
size: 1,
mintPane: () => {
const sessionId = "deadbeef-1111-2222-3333-444444444444";
return { sessionId, name: poolName(3456, sessionId) };
},
bootPane: async (model, ident) => {
seen.push(ident);
return { ...ident, model, bootedAt: Date.now() };
},
killPane: () => {},
paneHealthy: () => true,
});
pool.acquire("sonnet");
pool.refill();
await settle();
assert.equal(seen.length, 1, "bootPane received the pool-minted identity");
assert.equal(seen[0].name, "ocp-tui-3456-pdeadbeef");
assert.ok(seen[0].name.endsWith(seen[0].sessionId.slice(0, 8)),
"the tmux session name carries the session-id's own hex — `tmux ls` correlates to the transcript");
const pane = pool.acquire("sonnet");
assert.equal(pane.sessionId, seen[0].sessionId,
"and the turn reads the transcript under THAT session-id — one identity end to end");
});
test("pool pane names are port-scoped (reapable as ours) and never match the legacy shape", () => {
const name = poolName(3456, "deadbeef-1111-2222-3333-444444444444");
assert.ok(name.startsWith(sessionPrefixForPort(3456)), "pool panes are OURS → reapable when orphaned");
assert.equal(name, "ocp-tui-3456-pdeadbeef");
assert.ok(!LEGACY_SESSION_NAME_RE.test(name), "must never be mistaken for a legacy bare-prefix session");
assert.ok(!poolName(9999, "aaaaaaaa-0000-0000-0000-000000000000").startsWith(sessionPrefixForPort(3456)),
"a sibling instance's pool pane is foreign to us");
});
test("buildTuiHealthBlock reports pool:null when off, and the pool's stats when on", async () => {
const st = { lastEntrypoint: "cli", entrypointMismatches: 0 };
const sem = { inflight: 0, queued: 0 };
const off = buildTuiHealthBlock({ enabled: true, entrypointMode: "cli", maxConcurrent: 2 }, st, sem, null);
assert.equal(off.pool, null, "pool disabled → explicit null (stable /health shape)");
const { pool } = makeFakePool({ size: 2 });
pool.acquire("sonnet"); pool.refill(); await settle();
const on = buildTuiHealthBlock({ enabled: true, entrypointMode: "cli", maxConcurrent: 2 }, st, sem, pool);
assert.equal(on.pool.size, 2);
assert.equal(on.pool.warm, 2);
assert.equal(on.pool.misses, 1);
assert.equal(on.pool.model, "sonnet");
});
// ── TUI home preparation (scratch vs real) ───────────────────────────────
import { prepareTuiHome, ensureTuiCwdTrusted } from "./lib/tui/session.mjs";
import { mkdtempSync as hMkdtemp, mkdirSync as hMkdir, writeFileSync as hWrite, readFileSync as hRead, existsSync as hExists, readlinkSync as hReadlink } from "node:fs";
@@ -2415,8 +2978,19 @@ test("buildTuiHealthBlock: shape + live counters (the additive /health tui block
const ts = { lastEntrypoint: "cli", entrypointMismatches: 3 };
const block = buildTuiHealthBlock(
{ enabled: true, entrypointMode: "cli", maxConcurrent: 2 }, ts, sem);
assert.deepEqual(Object.keys(block).sort(),
["enabled", "entrypointMismatches", "entrypointMode", "inflight", "lastEntrypoint", "maxConcurrent", "queued"]);
// Shape is ADDITIVE-only: the seven original keys must all still be present (existing
// /health consumers are grandfathered, ADR 0006), plus `pool` (warm pane pool) and the
// stream* fields (backlog #2). Asserting CONTAINMENT plus an exact added-set — rather than
// one flat deepEqual — is what makes "additive" itself the thing under test: a future field
// that silently REPLACED an original key would pass a flat equality check that was updated
// alongside it, but cannot pass this one.
const ORIGINAL_KEYS = ["enabled", "entrypointMismatches", "entrypointMode", "inflight", "lastEntrypoint", "maxConcurrent", "queued"];
const keys = Object.keys(block);
for (const k of ORIGINAL_KEYS) assert.ok(keys.includes(k), `original /health key must survive: ${k}`);
assert.deepEqual(keys.filter((k) => !ORIGINAL_KEYS.includes(k)).sort(),
["pool", "streamDeltas", "streamDivergences", "streamEnabled", "streamTopUps", "streamTurns", "streamZeroDeltaTurns"],
"only the documented pool + streaming fields may be added");
assert.equal(block.pool, null, "no pool passed → null (the default, pool disabled)");
assert.equal(block.enabled, true);
assert.equal(block.entrypointMode, "cli");
assert.equal(block.lastEntrypoint, "cli");
@@ -2874,8 +3448,370 @@ async function runAsyncTests() {
});
}
// ── TUI real streaming: MessageDisplay hook sink (backlog #2) ───────────────
// Pure-logic coverage for lib/tui/stream.mjs: sink parsing, the concat===T assertion,
// prefix-stability, the auth-banner holdback, message scoping, and the error paths.
import { TuiDeltaAssembler, parseDeltaChunk, buildStreamSettings, streamFilePath, HOOK_SCRIPT, prepareStreamHook } from "./lib/tui/stream.mjs";
test("stream: parseDeltaChunk consumes only COMPLETE lines (a torn write stays unread)", () => {
const p = (i, d, final = false) => JSON.stringify({ hook_event_name: "MessageDisplay", session_id: "s", message_id: "m", index: i, final, delta: d });
// second payload is mid-write — no trailing newline yet
const partial = `${p(0, "## A\n\n")}\n${p(1, "body").slice(0, 20)}`;
const r1 = parseDeltaChunk(partial, 0);
assert.equal(r1.deltas.length, 1, "only the terminated line is consumed");
assert.equal(r1.consumed, 1);
// now it lands complete
const whole = `${p(0, "## A\n\n")}\n${p(1, "body")}\n`;
const r2 = parseDeltaChunk(whole, r1.consumed);
assert.equal(r2.deltas.length, 1, "the once-partial line is picked up exactly once");
assert.equal(r2.deltas[0].delta, "body");
assert.equal(r2.consumed, 2);
// idempotent: nothing new
assert.equal(parseDeltaChunk(whole, r2.consumed).deltas.length, 0);
});
test("stream: parseDeltaChunk skips blank/garbage lines and foreign hook events", () => {
const md = JSON.stringify({ hook_event_name: "MessageDisplay", message_id: "m", index: 0, final: true, delta: "ok" });
const other = JSON.stringify({ hook_event_name: "Stop", message_id: "m", delta: "nope" });
const text = `\n{not json\n${other}\n${md}\n`;
const { deltas } = parseDeltaChunk(text, 0);
assert.equal(deltas.length, 1);
assert.equal(deltas[0].delta, "ok");
});
// The live-verified contract (claude 2.1.207): deltas are the raw markdown source and
// concat(deltas) === extractLatestAssistantText(transcript), byte-exactly.
const mdFire = (i, delta, { final = false, mid = "m1" } = {}) =>
({ hook_event_name: "MessageDisplay", session_id: "s1", message_id: mid, index: i, final, delta });
test("stream: concat(deltas) === T → exact, no top-up, prefix-stable at every n", () => {
const chunks = ["## Mutex\n\n", "A **mutual exclusion lock** prevents concurrent access.\n\n", "```javascript\nconst m = new Mutex();\n```"];
const T = chunks.join("");
const a = new TuiDeltaAssembler({ holdbackChars: 10 });
let acc = "";
chunks.forEach((c, i) => {
const out = a.push(mdFire(i, c, { final: i === chunks.length - 1 }));
if (out) acc += out;
assert.ok(T.startsWith(a.full), `prefix-stable at n=${i}`);
});
const rec = a.finalize(T);
assert.equal(rec.ok, true);
assert.equal(rec.exact, true, "concat(deltas) === T");
assert.equal(acc + rec.tail, T, "client's assembled stream === T");
assert.equal(a.deltas, 3);
});
test("stream: holdback withholds the first chars so the auth-banner gate can still fire", () => {
const banner = "Please run /login · API Error: 401 Invalid authentication credentials"; // 69 chars, a real one
const a = new TuiDeltaAssembler(); // default holdback 100
const out = a.push(mdFire(0, banner, { final: true }));
assert.equal(out, null, "a banner-length message must NEVER reach the client");
assert.equal(a.emitted, "", "nothing emitted");
// and the whole-message detector still classifies it — the gate runs on T, before any flush
assert.ok(detectTuiUpstreamError(a.full) !== null, "banner still detected at terminal");
});
test("stream: holdback releases once past the detector's reach, and only then", () => {
const a = new TuiDeltaAssembler({ holdbackChars: 100 });
assert.equal(a.push(mdFire(0, "x".repeat(80))), null, "80 chars: still held");
const out = a.push(mdFire(1, "y".repeat(40)));
assert.equal(out, "x".repeat(80) + "y".repeat(40), "released as one chunk once >100");
assert.equal(a.push(mdFire(2, "tail")), "tail", "subsequent deltas stream straight through");
});
test("stream: a short answer never passes the holdback and is delivered whole at terminal", () => {
const T = "The capital of France is Paris.";
const a = new TuiDeltaAssembler();
assert.equal(a.push(mdFire(0, T, { final: true })), null);
const rec = a.finalize(T);
assert.equal(rec.ok, true);
assert.equal(rec.exact, true);
assert.equal(rec.tail, T, "the whole short answer is flushed at terminal (buffered semantics)");
});
test("stream: a DROPPED delta is a safe prefix → top-up from the transcript, exact=false", () => {
const a = new TuiDeltaAssembler({ holdbackChars: 5 });
a.push(mdFire(0, "Hello world, this is the first block. "));
const T = "Hello world, this is the first block. And the tail the hook never delivered.";
const rec = a.finalize(T);
assert.equal(rec.ok, true, "prefix → recoverable");
assert.equal(rec.exact, false, "flagged: concat(deltas) !== T");
assert.equal(rec.tail, "And the tail the hook never delivered.");
assert.equal(a.emitted + rec.tail, T, "client still receives exactly T");
});
test("stream: emitted bytes NOT a prefix of T → divergence, refuse the turn", () => {
const a = new TuiDeltaAssembler({ holdbackChars: 5 });
a.push(mdFire(0, "Let me go and read that file for you first."));
const rec = a.finalize("A completely different final answer.");
assert.equal(rec.ok, false, "must NOT serve text the transcript disagrees with");
assert.equal(rec.tail, null);
});
// Message scoping: the transcript keeps only the LAST assistant message, so the assembler
// must too. Discarding is safe while nothing has been emitted; after that it is a divergence.
test("stream: new message_id BEFORE any emit → held text discarded, stays exact vs T", () => {
const a = new TuiDeltaAssembler({ holdbackChars: 100 });
a.push(mdFire(0, "I'll check the file.", { mid: "m1" })); // short pre-tool prose, held back
assert.equal(a.emitted, "");
const answer = "The file defines a Mutex class with acquire and release, and " + "z".repeat(90);
const out = a.push(mdFire(0, answer, { mid: "m2", final: true }));
assert.equal(out, answer, "only the FINAL message's text is emitted");
const rec = a.finalize(answer); // T = extractLatestAssistantText = the last message only
assert.equal(rec.ok, true);
assert.equal(rec.exact, true, "scoping to the last message_id keeps concat === T true");
assert.equal(a.messages, 2);
});
test("stream: new message_id AFTER an emit → unretractable, flagged and refused", () => {
const a = new TuiDeltaAssembler({ holdbackChars: 10 });
a.push(mdFire(0, "Long pre-tool prose that already went out to the client.", { mid: "m1" }));
assert.notEqual(a.emitted, "");
// Assert what push() RETURNS, not merely that finalize() refuses. This test used to check
// only restartedAfterEmit + finalize().ok, which left it passing while F1 was live: the
// second message's bytes were still being handed to the client. "The turn is refused" and
// "the client got the bytes anyway" were both true at once.
const out = a.push(mdFire(0, "The real answer.", { mid: "m2", final: true }));
assert.equal(out, null, "after a message boundary follows an emit, NOTHING more may be emitted");
assert.equal(a.restartedAfterEmit, true);
assert.equal(a.finalize("The real answer.").ok, false, "must refuse: prose already emitted is not in T");
});
test("F1: an auth banner rendered as a LATER message is never forwarded to the client", () => {
// The leak this class exists to prevent, in the shape production actually runs
// (OCP_TUI_FULL_TOOLS=1 → multi-message tool-using turns are the norm):
// 1. the model narrates past the holdback before a tool call → released, emitted != ""
// 2. credentials expire mid-turn → claude renders the 401 as ordinary assistant TEXT,
// as a NEW message
// 3. pre-fix: push() took the `if (this.released)` branch — `released` was never reset at
// a message boundary — and returned the BANNER verbatim, straight to the client.
// The holdback protected only the FIRST message of a turn. This asserts it protects the rest.
const a = new TuiDeltaAssembler({ holdbackChars: 100 });
const narration = "I'll check that file for you and then report back with what I find inside it.";
a.push(mdFire(0, narration + narration, { mid: "m1" })); // > holdback → released
assert.notEqual(a.emitted, "", "precondition: the narration really did reach the client");
const BANNER = "Please run /login · API Error: 401 Invalid authentication credentials";
const out = a.push(mdFire(1, BANNER, { mid: "m2", final: true }));
assert.equal(out, null, "the auth banner must NOT be forwarded once a later message begins");
assert.ok(!a.emitted.includes("401"), "no byte of the banner may have reached the client");
assert.equal(a.finalize(BANNER).ok, false, "and the turn is refused, not served");
});
test("F1: a first payload with message_id:null cannot disarm the guard", () => {
// The residual bypass the reviewer found by probing. `this.messageId` used to be initialized
// to null, so a first payload carrying message_id:null compared EQUAL to it → no boundary
// registered → `messages` stayed 0 → when the REAL boundary arrived, `messages > 1` evaluated
// 1 > 1 === false → restartedAfterEmit never armed → the released branch forwarded the banner.
// The whole F1 guard was disarmed by a single null field. parseDeltaChunk does not validate
// message_id, so such a payload does reach push().
const a = new TuiDeltaAssembler({ holdbackChars: 100 });
const narration = "I'll check that file for you and then report back with what I find inside it.";
a.push({ hook_event_name: "MessageDisplay", message_id: null, delta: narration + narration });
assert.notEqual(a.emitted, "", "precondition: the narration released to the client");
assert.equal(a.messages, 1, "a null message_id is still a MESSAGE — it must register as one");
const BANNER = "Please run /login · API Error: 401 Invalid authentication credentials";
const out = a.push({ hook_event_name: "MessageDisplay", message_id: "m2", delta: BANNER });
assert.equal(out, null, "the banner must not be forwarded — the guard must arm regardless");
assert.equal(a.restartedAfterEmit, true);
assert.ok(!a.emitted.includes("401"));
});
test("F1: whitespace cannot buy a release — the holdback screens TRIMMED length", () => {
// detectTuiUpstreamError() TRIMS before applying its <=100-char rule, so gating release on
// the UNTRIMMED pending.length let 101 spaces trim to "" → the detector has nothing to
// classify → returns null → release fires having screened nothing, and every subsequent
// delta of that message (a banner included) streams unfiltered.
const a = new TuiDeltaAssembler({ holdbackChars: 100 });
assert.equal(a.push(mdFire(0, " ".repeat(101), { mid: "m1" })), null,
"101 chars of whitespace must not clear a 100-char holdback");
assert.equal(a.released, false, "…and must not flip the assembler into released state");
const BANNER = "Please run /login · API Error: 401 Invalid authentication credentials";
assert.equal(a.push(mdFire(1, BANNER, { mid: "m1" })), null, "so the banner stays held back");
assert.ok(!a.emitted.includes("401"));
});
test("F3: a STALE or truncated hook script is overwritten, not trusted because it exists", () => {
// ~/.ocp-tui/stream/{md-hook.sh,settings.json} persist across OCP restarts. The old
// write-if-missing guard meant a host that booted once under an older version was stuck on
// that version's HOOK_SCRIPT forever — no upgrade could reach it. Worse, a non-atomic write
// interrupted mid-flight leaves a TRUNCATED md-hook.sh that existsSync() calls fine, and
// claude BLOCKS on that hook synchronously on every fire.
const dir = mkdtemp2(`${tmpdir2()}/ocp-hook-`);
mkdir2(dir, { recursive: true });
writeFile2(`${dir}/md-hook.sh`, "#!/bin/sh\n# stale, truncated leftov", { mode: 0o700 });
writeFile2(`${dir}/settings.json`, "{ TRUNCATED", { mode: 0o600 });
const settings = prepareStreamHook(dir);
assert.equal(readFile2(`${dir}/md-hook.sh`, "utf8"), HOOK_SCRIPT,
"the stale script must be replaced with the current one, not left because it existed");
assert.deepEqual(JSON.parse(readFile2(settings, "utf8")), buildStreamSettings(`${dir}/md-hook.sh`),
"…and so must the stale settings file");
});
test("stream: hook script is a write-and-exit sh script and tolerates a missing sink var", () => {
// forceSyncExecution: claude BLOCKS on this hook, so it must do no work inline.
assert.ok(HOOK_SCRIPT.startsWith("#!/bin/sh"));
assert.ok(HOOK_SCRIPT.includes('[ -n "$OCP_TUI_STREAM_FILE" ] || exec cat >/dev/null'),
"no sink configured => swallow stdin and exit 0; never fail, never block claude");
assert.ok(!/curl|node |python/.test(HOOK_SCRIPT), "no interpreter/network work in a blocking hook");
});
test("stream: settings registers exactly one MessageDisplay command hook (static, no per-request data)", () => {
const s = buildStreamSettings("/x/md-hook.sh");
assert.deepEqual(Object.keys(s.hooks), ["MessageDisplay"]);
assert.equal(s.hooks.MessageDisplay[0].hooks[0].type, "command");
assert.equal(s.hooks.MessageDisplay[0].hooks[0].command, "/x/md-hook.sh");
// Warm-pool compatibility: the settings file must NOT carry a session/request-specific path.
assert.ok(!JSON.stringify(s).includes(".jsonl"), "sink path comes from the pane env, not the settings file");
});
test("stream: sink path is keyed by session_id (concurrent panes cannot interleave)", () => {
// OCP_TUI_MAX_CONCURRENT defaults to 2 — two claude panes DO run at once. A shared sink
// would splice request A's deltas into request B's stream.
const A = streamFilePath("/d", "aaaa-1111");
const B = streamFilePath("/d", "bbbb-2222");
assert.notEqual(A, B, "one sink per session-id");
assert.ok(A.endsWith("/aaaa-1111.jsonl"));
});
test("stream: buildTuiCmd — OFF is byte-for-byte the pre-streaming argv; ON adds only env + --settings", () => {
const off = buildTuiCmd("/bin/claude", "m", "SID", "/h", "cli");
assert.ok(!off.includes("--settings"), "no --settings when streaming is off");
assert.ok(!off.includes("OCP_TUI_STREAM_FILE"), "no sink env when streaming is off");
const on = buildTuiCmd("/bin/claude", "m", "SID", "/h", "cli", { file: "/d/SID.jsonl", settings: "/d/s.json" });
assert.ok(on.includes("OCP_TUI_STREAM_FILE='/d/SID.jsonl'"), "sink delivered via the pane env");
assert.ok(on.includes("--settings '/d/s.json'"));
// must not regress the MCP wall or the pinned effort (#156)
assert.ok(on.includes("--strict-mcp-config") && on.includes("--disallowedTools 'mcp__*'"), "MCP wall intact");
assert.ok(on.includes("--effort low"), "OCP_TUI_EFFORT default intact");
assert.ok(!on.includes(" -p ") && !on.includes("--bare"), "still a plain interactive TUI spawn");
});
test("stream: /health block is additive and exposes the divergence counter", () => {
const stats = { lastEntrypoint: "cli", entrypointMismatches: 0, streamTurns: 3, streamDeltas: 21, streamTopUps: 1, streamDivergences: 0 };
const sem = { inflight: 0, queued: 0 };
const b = buildTuiHealthBlock({ enabled: true, entrypointMode: "cli", maxConcurrent: 2, streamEnabled: true }, stats, sem);
assert.equal(b.streamEnabled, true);
assert.equal(b.streamTurns, 3);
assert.equal(b.streamDivergences, 0);
// existing fields unchanged (grandfathered /health consumers)
assert.equal(b.enabled, true);
assert.equal(b.entrypointMode, "cli");
assert.equal(b.maxConcurrent, 2);
// a pre-streaming tuiStats (no stream* keys) must not produce undefined/NaN
const legacy = buildTuiHealthBlock({ enabled: false, entrypointMode: "cli", maxConcurrent: 2 }, { lastEntrypoint: null, entrypointMismatches: 0 }, sem);
assert.equal(legacy.streamEnabled, false);
assert.equal(legacy.streamDivergences, 0);
});
// ── Cleanup ──
runAsyncTests().then(() => {
// Settle the async-bodied tests registered through the sync `test()` helper BEFORE summarizing —
// otherwise their pass/fail is not reflected in the counts (see the `pendingAsync` comment above).
// ─── TUI streaming × warm pool: the INTEGRATION seam (backlog #2 rebased onto #158) ───
//
// The hook is installed by bootTuiPane at BOOT, and runTuiTurn reads the sink off the PANE
// (pane.streamFile). That indirection is the entire reason a POOLED pane streams: the pool
// pre-boots panes long before a request exists, so anything derived at turn time would leave
// every pool HIT silently buffered while every MISS streamed — a perf regression with no
// failing test and no error, visible only as "streaming mysteriously does nothing in prod".
// These three guard that seam.
console.log("\nTUI streaming × warm pane pool integration:");
import { bootTuiPane as bootPaneUnderTest, runTuiTurn as runTurnUnderTest } from "./lib/tui/session.mjs";
import { mkdtempSync as mkdtemp2, writeFileSync as writeFile2, mkdirSync as mkdir2, readFileSync as readFile2 } from "node:fs";
import { tmpdir as tmpdir2 } from "node:os";
// Fake tmux that records the spawned pane command and always looks ready + pasted.
function makeTmuxRecorder() {
const cmds = [];
const tmux = (args) => {
cmds.push(args);
if (args[0] === "capture-pane") {
// input bar present AND the prompt visibly landed → both polls pass immediately
return { status: 0, stdout: "[Pasted text #1 +2 lines]\n ? for shortcuts" };
}
return { status: 0, stdout: "" };
};
return { tmux, cmds, paneCmd: () => (cmds.find((a) => a[0] === "new-session") || []).slice(-1)[0] || "" };
}
// A HOME with one already-terminal transcript for `sid`, so readTuiTranscript returns at once.
function seedTranscript(home, sid, text) {
const dir = `${home}/.claude/projects/x`;
mkdir2(dir, { recursive: true });
writeFile2(`${dir}/${sid}.jsonl`, JSON.stringify({
type: "assistant",
message: { role: "assistant", content: [{ type: "text", text }], stop_reason: "end_turn" },
turn_duration: 1234, cc_entrypoint: "cli",
}) + "\n");
}
test("bootTuiPane with a streamDir installs the hook AT BOOT and hands the sink back on the pane", async () => {
const home = mkdtemp2(`${tmpdir2()}/ocp-t-`);
const streamDir = mkdtemp2(`${tmpdir2()}/ocp-s-`);
const rec = makeTmuxRecorder();
const pane = await bootPaneUnderTest({
model: "sonnet", claudeBin: "claude", home, realHome: home,
cwd: `${home}/wk`, port: 3456, tmux: rec.tmux, streamDir,
});
// The pane carries its OWN sink, named from its OWN session-id — which is what a pre-booted
// pool pane needs, since it is minted with no knowledge of the request it will eventually serve.
assert.ok(pane.streamFile, "a streamDir must yield a per-pane sink on the returned pane");
assert.ok(pane.streamFile.includes(pane.sessionId), "the sink is keyed by the pane's own session-id");
const cmd = rec.paneCmd();
assert.ok(cmd.includes("OCP_TUI_STREAM_FILE="), "the pane's env must carry its sink path");
assert.ok(cmd.includes(pane.streamFile), "…and it must be THIS pane's sink, not a shared one");
assert.ok(cmd.includes("--settings"), "the MessageDisplay hook must be registered at spawn");
});
test("bootTuiPane WITHOUT a streamDir spawns exactly today's pane — no hook, no --settings", async () => {
const home = mkdtemp2(`${tmpdir2()}/ocp-t-`);
const rec = makeTmuxRecorder();
const pane = await bootPaneUnderTest({
model: "sonnet", claudeBin: "claude", home, realHome: home,
cwd: `${home}/wk`, port: 3456, tmux: rec.tmux,
});
assert.equal(pane.streamFile, null, "no streamDir → no sink (streaming is opt-in, default OFF)");
const cmd = rec.paneCmd();
assert.ok(!cmd.includes("--settings"), "the default spawn must not gain --settings");
assert.ok(!cmd.includes("OCP_TUI_STREAM_FILE"), "the default spawn must not gain the hook env");
});
test("REGRESSION: a WARM (pooled) pane streams — the sink comes off the pane, not the turn", async () => {
const home = mkdtemp2(`${tmpdir2()}/ocp-t-`);
const streamDir = mkdtemp2(`${tmpdir2()}/ocp-s-`);
const rec = makeTmuxRecorder();
// Pre-boot a pane the way the POOL does (its own session-id + sink, fixed at boot).
const warm = await bootPaneUnderTest({
model: "sonnet", claudeBin: "claude", home, realHome: home,
cwd: `${home}/wk`, port: 3456, tmux: rec.tmux, streamDir,
});
// Its hook has already fired twice by the time the turn's transcript goes terminal.
writeFile2(warm.streamFile,
JSON.stringify({ hook_event_name: "MessageDisplay", delta: "Hello " }) + "\n" +
JSON.stringify({ hook_event_name: "MessageDisplay", delta: "world" }) + "\n");
seedTranscript(home, warm.sessionId, "Hello world");
const seen = [];
const pool = { acquire: () => warm, refill: () => {}, warm: 0 };
const out = await runTurnUnderTest({
prompt: "say hello", model: "sonnet", claudeBin: "claude", home, realHome: home,
cwd: `${home}/wk`, port: 3456, tmux: rec.tmux, pool,
onDelta: (d) => seen.push(d.delta),
// streamDir is deliberately NOT passed: on a pool HIT runTuiTurn never cold-boots, so if it
// recomputed the sink from a turn-time streamDir (the pre-rebase shape) this turn would emit
// ZERO deltas and silently serve buffered. Reading pane.streamFile is what makes it stream.
streamDir: null,
});
assert.deepEqual(seen, ["Hello ", "world"], "the pooled pane's deltas must reach the client");
assert.equal(out.text, "Hello world", "and the transcript stays authoritative for the final text");
});
runAsyncTests().then(() => Promise.all(pendingAsync)).then(() => {
closeDb();
console.log(`\n=== Results: ${passed} passed, ${failed} failed ===\n`);
process.exit(failed > 0 ? 1 : 0);