Compare commits

...
Author SHA1 Message Date
taodengandClaude Fable 5 4239744c86 docs(readme): review nit — 'pre-v3.22' → 'pre-flag' (next version not fixed yet)
Reviewer nit from the Iron Rule 10 independent review of PR #156.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR
2026-07-13 15:15:01 +10:00
taodengandClaude Fable 5 e95a7bb0f2 feat(tui): pin spawn effort via OCP_TUI_EFFORT (default low)
buildTuiCmd never passed --effort, so the pane's claude inherited a
HOME-dependent effortLevel: real-home mode inherits the operator's
~/.claude/settings.json (high/xhigh on typical operator hosts),
env-token scratch mode inherits claude's built-in default — proxied-turn
latency silently depended on which HOME mode resolveTuiHome() picked and
on an unrelated operator setting.

Pass --effort explicitly, from new env var OCP_TUI_EFFORT (default
"low"; allowlist low|medium|high|xhigh|max per `claude --help` 2.1.207;
"inherit" restores the pre-flag argv byte-for-byte; an invalid value
warns and falls back to "low" so a typo can never reach the pane argv).

Not endpoint-touching: no server.mjs change, no wire-level change — the
flag rides the existing interactive spawn (ADR 0007). Billing-pool
safety verified per the docs/plans/2026-07-13-tui-latency banner
protocol: startup banner stays "Claude Max" with "low effort".

Measured through a test OCP instance (:3979, TUI mode, real-home,
claude-sonnet-4-6, n=5+5, same ~1850-token prompt as floor.sh):

  before: median 11.30s  range 9.05-12.38s (spread 3.32s)  banner: high effort - Claude Max
  after:  median  9.55s  range 9.27-9.77s  (spread 0.50s)  banner: low effort - Claude Max
  OCP_TUI_EFFORT=inherit: banner back to "high effort" (pre-flag behavior restored)

README: new row in the Environment Variables table (release_kit
new_feature_doc_expectations: new env var -> README table). Tests: 4 new
buildTuiCmd cases (default, explicit level, inherit, invalid fallback);
suite 267 passed / 0 failed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VqgWJcjxrjjL9L9SkpZyXR
2026-07-13 15:08:50 +10:00
6854075c01 docs(plans): TUI-mode latency floor — measured decomposition + backlog (#155)
* docs(plans): TUI-mode latency floor — measured decomposition + backlog

An external consumer measured OCP's prompt path at TTFT p50 30-32s and excluded
OCP on that basis. This documents where those 30 seconds actually go, with a
reproducible harness (n=15) that bypasses OCP and measures the underlying
subscription path's true first-token time.

Findings:
- boot -> input-ready is only ~1.0s; it is NOT the bottleneck
- true TTFT is 6-10s; the remaining ~20s is runTuiTurn polling the transcript
  until turn_duration (ADR 0007 step 4) — i.e. waiting for the WHOLE turn.
  There is no streaming.
- buildTuiCmd never passes --effort, so the spawned claude inherits the
  operator's global effortLevel (xhigh on this host) — every request runs
  extended thinking. Passing --effort low: TTFT p50 9.70s -> 6.17s (-36%),
  spread 7.85-13.07s -> 5.87-6.44s. Stays on Claude Max.
- ⚠️ --bare SILENTLY drops off the subscription pool (banner flips
  'Claude Max' -> 'API Usage Billing'). It does cut boot to ~0.5s, but defeats
  the entire purpose of ADR 0007. Failure is silent: all 5 --bare samples
  produced no answer at all (no error, no crash, just never a token).
  Anyone optimizing boot MUST diff the banner line.
- Floor after all fixes is ~6s (claude always injects the full CC system prompt
  + tool definitions). TUI mode therefore cannot serve real-time consumers —
  a constraint worth stating in the README.

Backlog ranked by value/effort: (1) OCP_TUI_EFFORT env var, default low;
(2) real streaming instead of turn_duration polling (~20s, the big one);
(3) warm pane pool (~1s); (4) prefill trim (probably not worth it).

Docs-only; no version bump (matches repo convention — bump lands in the
chore(release) commit).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dx5Ncq6wWBrF27vJKHZ9Hr

* docs(plans): address review — restore --bare evidence, qualify effort claim, add banner captures

Reviewer (fresh-context, Iron Rule 10) returned REQUEST_CHANGES. All four technical
conclusions survived independent verification (source-read + live repro); the defects
were in the evidence file, and they were real:

- H-1: measurements.jsonl claimed n=15 but held 10 rows, and the --bare group — the
  basis of this PR's headline warning — had ZERO rows. The author had stripped them
  as 'invalid samples' (ttft_ms:-1) when they were in fact the evidence. Regenerated:
  n=15, three groups × 5, all with tag/extra_args. --bare reproduces exactly (5/5 no
  answer, boot 0.43-0.45s).
- M-1: the effort-inheritance claim was written unconditionally, but it depends on
  resolveTuiHome()'s mode. Real-home (current service config) inherits the operator's
  effortLevel: xhigh; env-token scratch home (~/.ocp-tui/home) has no effortLevel in
  its settings.json and prepareTuiHome() never writes one, so the pane gets claude's
  built-in default. Now documented as a table — and the mode split makes passing
  --effort explicitly MORE valuable, not less.
- M-2: baseline rows were produced by a pre-parameterized script and lacked
  tag/extra_args. Re-run with the committed script. Recomputed effect: -40% (was -36%).
- L-1: documented that the harness suppresses OCP's periodic kill-server tick via the
  othersRemain coexistence guard (by design, resumes next tick).
- L-2: documented that floor.sh's readiness marker differs from OCP's tuiInputReady(),
  so the ~1.0s boot figure is not apples-to-apples with BOOT_MS.
- Direct-API reference figure now explicitly labeled as external (not in this dataset).
- New: billing-banner.txt captures all three configs live, including confirmation that
  --effort low stays on Claude Max (reviewer noted this was asserted but unevidenced).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dx5Ncq6wWBrF27vJKHZ9Hr

* docs(plans): scope the effort claim to TUI mode (currently off), drop nonexistent bin/

Re-review (APPROVE_WITH_MINOR) caught two accuracy defects:
- MIN-1: 'every OCP request runs extended thinking' over-extrapolated. TUI mode is
  currently OFF on this host (CLAUDE_TUI_MODE=false; /health tui.enabled=false), so
  live traffic takes the -p path. The claim is about what happens WHEN TUI mode is
  enabled — now scoped, and the same qualifier applied to the kill-server interaction
  note (that reap tick is itself gated on TUI_MODE).
- NIT-2: the quoted grep included bin/, which does not exist in the repo (exit 2).
  Dropped; the zero-hit result over lib/ + server.mjs is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dx5Ncq6wWBrF27vJKHZ9Hr

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 14:22:17 +10:00
7 changed files with 422 additions and 0 deletions
+1
View File
@@ -958,6 +958,7 @@ See [Subscription-pool (TUI) mode](#subscription-pool-tui-mode) and ADR 0007 PR-
| `OCP_TUI_CWD` | `$HOME/.ocp-tui/work` | (TUI-mode) Scratch working directory where interactive claude sessions run. Transcripts land under `<HOME>/.claude/projects/<encoded-cwd>/`. Created automatically. |
| `OCP_TUI_HOME` | *(auto)* | (TUI-mode) `HOME` claude runs under. **When unset, OCP picks it for you:** if `CLAUDE_CODE_OAUTH_TOKEN` is set → a **credential-isolated** scratch home `$HOME/.ocp-tui/home` (no `credentials.json`, env-token auth — **recommended**); if no env token → the operator's real home (legacy shared `credentials.json`). Setting this to an **explicit** path overrides the auto-default. The credential handling at that path still follows the env token: **with** the env token it is credential-free (env-token auth, no `credentials.json` written); **without** the env token (and the path ≠ real home) it uses the legacy symlinked-credentials scratch mode, which carries the credential-fork caveat — see ADR 0007. |
| `OCP_TUI_ENTRYPOINT` | `cli` | (TUI-mode) Billing-classifier labeling: `cli` (default) pins `cc_entrypoint=cli` deterministically; `auto` lets claude self-classify via TTY detection; `off` leaves the inherited env untouched. Honest only when the spawn is a genuine interactive PTY — see ADR 0007. |
| `OCP_TUI_EFFORT` | `low` | (TUI-mode) Effort level passed to the interactive `claude` as an explicit `--effort` flag: `low` (default), `medium`, `high`, `xhigh`, `max`, or `inherit` to omit the flag (the pre-flag behaviour: the pane inherits a HOME-dependent effort — the operator's `~/.claude/settings.json` `effortLevel` in real-home mode, claude's built-in default in env-token scratch mode). Explicit `low` cuts measured TTFT p50 by ~40% and collapses run-to-run variance ~15× versus an inherited `xhigh` (see `docs/plans/2026-07-13-tui-latency/`); proxied requests rarely benefit from extended thinking. Banner-verified to stay on the subscription pool (`· Claude Max`). An invalid value logs a warning and falls back to `low`. |
| `OCP_TUI_MAX_CONCURRENT` | `2` | (TUI-mode) Max concurrent interactive TUI turns. **Independent** of `CLAUDE_MAX_CONCURRENT` (which bounds the `-p`/stream-json path; TUI never uses it). A TUI turn is heavy (per-request cold-boot of tmux+claude + up to `CLAUDE_TUI_WALLCLOCK_MS` wallclock), so the default is low to keep small hosts (e.g. a Pi 4) alive under a burst. Excess turns **queue** (bounded); a full queue yields a 503. See ADR 0007 PR-B amendment. |
| `OCP_SKIP_AUTH_TEST` | *(unset)* | When `=1`, skip the `claude -p` auth probe during `setup.mjs`. After 2026-06-15 this probe draws from the Agent SDK credit pool; set this to avoid burning a metered credit on re-installs or `ocp update` runs. Auth is validated at the first real request. |
| `OCP_TUI_FULL_TOOLS` | *(unset)* | (TUI-mode, **single-user only**) When `=1`, grant the interactive session the **same tool surface as the `-p` path** — `--allowedTools` (+ optional `--mcp-config`, read from `CLAUDE_ALLOWED_TOOLS` / `CLAUDE_MCP_CONFIG`) — instead of the default MCP-walled, built-in-tools-only set. Lets a trusted single-operator TUI deployment run a **tool-using / MCP agent** (e.g. an OpenClaw assistant) on the subscription pool. Safe because TUI **refuses to boot under `AUTH_MODE=multi`** (hard exit) — no guest key can ever reach the TUI path, so this gate cannot expose tools to an untrusted caller. (Under `AUTH_MODE=shared` + `OCP_TUI_ALLOW_LAN=1`, anyone holding the single shared key reaches it — that is the existing TUI trust model, unchanged.) Note: `--dangerously-skip-permissions` / `CLAUDE_SKIP_PERMISSIONS` is **not** supported for TUI — claude v2.1.x shows an interactive bypass-acceptance screen in headless tmux that cannot be answered, bricking the pane. Use scratch-home `settings.json` `additionalDirectories` instead. See [Subscription-pool (TUI) mode](#subscription-pool-tui-mode) and ADR 0007. |
+190
View File
@@ -0,0 +1,190 @@
# TUI-mode latency: measured floor, and the four things worth fixing
**Date**: 2026-07-13
**Status**: findings + backlog (no code changed yet)
**Measured on**: Mac mini / macOS 26.5.2 / Claude Code **v2.1.207** / Sonnet 5 / Claude Max subscription / **real-home mode** (no `CLAUDE_CODE_OAUTH_TOKEN`, no `OCP_TUI_HOME` in the service env)
**Evidence**: [`measurements.jsonl`](measurements.jsonl) — **n=15** (3 configs × 5) · banner captures [`billing-banner.txt`](billing-banner.txt) · harness [`floor.sh`](floor.sh)
## Why this exists
An external consumer (the 知音 AI project) benchmarked OCP's prompt path and measured
**TTFT p50 ≈ 3032 s**, and excluded OCP as a backend on that basis. That number is real,
but it is *not* the model being slow — this document decomposes where the 30 seconds
actually go, and what OCP can do about it.
**The harness deliberately does not go through OCP.** It spawns `tmux` + `claude` directly
(session prefix `zhiyin-floor-`, never `ocp-tui-*`) and polls `tmux capture-pane` for
incremental render, so it measures the **true first-token time** of the underlying
subscription path — the floor OCP could reach if it were perfect.
---
## Measurements
All rows in [`measurements.jsonl`](measurements.jsonl); every number below is recomputable from it.
| Config | n | boot→input-ready (median) | **TTFT (median)** | TTFT range | full answer (median) |
|---|---|---|---|---|---|
| baseline (inherits global `effortLevel: xhigh`) | 5 | 1.07 s | **10.35 s** | 8.32 17.19 s | 11.32 s |
| **`--effort low`** | 5 | 1.03 s | **6.17 s** | **5.87 6.44 s** | 9.98 s |
| `--bare` | 5 | 0.44 s | **no answer at all** (5/5 `ttft_ms: -1`) | — | — |
> **Not from this harness**: the direct Anthropic API reference figure (TTFT 0.841.64 s, n=2)
> comes from the 知音 AI project's own smoke test, not from `measurements.jsonl`. It is quoted
> only to size the gap; do not look for it in the evidence file.
### Where the 30 seconds go
```
~1.0 s spawn → claude's input bar is ready ← NOT the bottleneck
~6-10 s true TTFT (first token rendered in the pane)
~20 s ████ waiting for the whole turn to finish ████ ← this is the 30s
```
`runTuiTurn` blocks on the native transcript until a terminal event (`lib/tui/session.mjs`
"Block on the native transcript … until terminal"; `readTuiTranscript` in
`lib/tui/transcript.mjs`; ADR 0007 step 4) — i.e. it waits for the **entire turn** to complete
before returning anything. There is no streaming path. The ~20 s delta between this harness's
real TTFT and OCP's reported 3032 s is exactly that.
---
## ⚠️ Blocking constraint: `--bare` silently drops you off the subscription pool
Captured live ([`billing-banner.txt`](billing-banner.txt)) — the startup banner is the **only**
reliable indicator:
```
[] | Sonnet 5 with xhigh effort · Claude Max
[--effort low] | Sonnet 5 with low effort · Claude Max
[--bare] | Sonnet 5 with xhigh effort · API Usage Billing ← ❌
```
`--bare` ("skip hooks, LSP, plugin…") **also skips the subscription-credential resolution
path**. It really does cut boot to 0.430.45 s — but you are no longer on the subscription,
which defeats the entire purpose of TUI mode (ADR 0007 exists solely to reach the
subscription pool).
**The failure is silent.** All 5 `--bare` samples reached input-ready (boot 0.430.45 s), were
sent the prompt, and then produced **no answer at all** — 60 s timeout, no error, no crash, the
pane simply never rendered a token (the API-billing account had no credit balance). Nothing in
the transcript or the exit status reveals this.
**Anyone changing spawn flags must diff the banner line before and after.**
---
## Backlog — four items, ranked by value ÷ effort
### 1. Pass `--effort` explicitly on spawn — **do this first**
`buildTuiCmd` (`lib/tui/session.mjs`) does not pass `--effort``grep -rn -- "--effort\|effortLevel" lib/ server.mjs`
returns zero hits. What the pane's `claude` ends up using therefore depends on **which HOME mode
`resolveTuiHome()` picked**:
| mode | HOME | effort the pane gets |
|---|---|---|
| **real-home** (legacy default — *current* service config: no `CLAUDE_CODE_OAUTH_TOKEN`, no `OCP_TUI_HOME`) | `~` | **inherits the operator's `~/.claude/settings.json` → `effortLevel: xhigh` on this host** |
| env-token scratch (`CLAUDE_CODE_OAUTH_TOKEN` set — the direction #146/#150 pushed) | `~/.ocp-tui/home` | that settings.json contains only `permissions.additionalDirectories`; `prepareTuiHome()` never writes `effortLevel`**claude's built-in default** |
**Scope note**: TUI mode is currently *off* on this host (`CLAUDE_TUI_MODE=false`; `/health`
`"tui": {"enabled": false}`), so live traffic takes the `-p` path today. The statement below is
about what happens **when TUI mode is enabled**.
On the current HOME config, **every TUI request would run extended thinking** — pure waste
for the typical "generate this JSON" request, and it makes latency depend on an unrelated global
setting the operator may have changed for their own interactive use. And the mode split means
the effort level silently changes if the operator ever switches to env-token mode.
**Passing `--effort` explicitly fixes both problems at once.**
- **Effect (real-home, measured)**: TTFT p50 **10.35 s → 6.17 s (40 %)**, and the spread
collapses from 8.3217.19 s to **5.876.44 s**. For a proxy, the variance reduction matters
more than the median.
- **Cost**: one flag. Suggested: a new `OCP_TUI_EFFORT` env var (default `low`), documented in
README § "Environment Variables" per `release_kit.new_feature_doc_expectations`.
- **Risk**: none — banner confirms it stays on `Claude Max` (see `billing-banner.txt`).
- ⚠️ Do **not** reach for `--bare` to shave boot: see above.
### 2. Real streaming instead of blocking on turn-terminal — **the big one (~20 s)**
Today `runTuiTurn` blocks on the transcript until the turn is *finished*. The pane is already
rendering tokens incrementally the whole time — this harness proves you can observe first token
at ~6 s by polling `tmux capture-pane`.
- **Effect**: turns a 30 s wall into a ~6 s TTFT with progressive output; enables SSE streaming
on the OCP endpoint instead of a single blob at the end.
- **Cost**: real work. Pane capture is ANSI/redraw-based and lossy for exact text (wrapping,
scrollback, spinner lines). Two candidate sources: (a) incremental reads of the transcript
JSONL, (b) `capture-pane` diffing with a stable start marker. (a) is much cleaner **if it
holds**.
- **Prereq spike (do this before designing anything)**: does the transcript JSONL grow *during*
a turn, or only at the end? If only at the end, (a) is dead and you are stuck with (b).
### 3. Warm pane pool — ~1 s
Every request spawns a fresh tmux session + `claude` (`randomUUID()` + `new-session`, then
`kill-session` in `finally`; `grep -rn "pool\|warm\|reuse" lib/tui/*.mjs` → zero hits). Boot to
input-ready is ~1.0 s, paid on every request. A pool of pre-booted panes (single-use, replaced in
the background) amortizes it to zero for any workload below the pool refill rate.
- **Effect**: 1.0 s.
- **Cost**: moderate; interacts with the session reaper and the per-port prefix scoping added in
#148 — pooled panes must not look like zombies to the sweep.
- Lower priority than #1 and #2: it is the smallest slice.
### 4. Trim the prefill — **probably not worth it; know the floor**
After #1#3, the floor is **~6 s**, and it does not go lower. `claude` always injects the full
Claude Code system prompt + tool definitions (thousands to tens of thousands of prefill tokens)
regardless of what you ask it. `--exclude-dynamic-system-prompt-sections` exists and may shave
some of it — **unmeasured**; worth one spike, but do not expect to reach the direct API's
~1 s.
**Consequence to accept, and to state in the README**: even fully optimized, TUI mode has a
**~6 s TTFT floor**, so it cannot serve real-time / interactive-latency consumers. It remains
appropriate for batch, background, and cost-insensitive-latency use. The 知音 AI project
excluded it on this basis (their prompt-latency budget is 24 s) *independently* of the ToS
question already documented in the README.
---
## Reproduction
```bash
# harness never touches OCP's :3456 service or ocp-tui-* sessions, and never kill-server
bash docs/plans/2026-07-13-tui-latency/floor.sh 5 # baseline
TAG=effort-low EXTRA_ARGS="--effort low" bash .../floor.sh 5 # 40 %
TAG=bare EXTRA_ARGS="--bare" bash .../floor.sh 5 # the trap
# billing-pool check for ANY spawn-flag change — the banner is the only source of truth
tmux new-session -d -s probe -x 200 -y 50 -c "$HOME" \
"claude --model claude-sonnet-5 --session-id $(uuidgen) <your-flags-here>"
sleep 6; tmux capture-pane -p -t probe | grep -E "Claude Max|API Usage Billing"
tmux kill-session -t probe
```
## Interaction with OCP while the harness runs
- **Kill direction is safe both ways**: `reapStaleTuiSessions()` only `kill-session`s names
matching `ocp-tui-<port>-`, which `zhiyin-floor-*` never matches; and the harness only
`kill-session`s its own single session — it contains **no `kill-server`**.
- **One benign interaction** (only when TUI mode is enabled — the reap tick is itself gated on
`TUI_MODE`): OCP's periodic `kill-server` (zombie reaping) is gated on
`othersRemain`*any* foreign-prefixed tmux session suppresses it. So while the harness is
running, that sweep is skipped. This is the coexistence guard working as designed; it resumes
on the next tick.
## Harness caveats (stated so the numbers are not over-trusted)
- **n=5 per config**, single host, single model (Sonnet 5), single prompt size (~1850 tokens).
Enough to separate 6 s from 10 s from 30 s; **not** enough for a p95.
- TTFT is "marker visible in `capture-pane`", which includes tmux render latency (small, but
nonzero) — it is an upper bound on the true first-token time.
- **The harness's readiness marker is not OCP's.** `floor.sh` waits for `│ >||Try "`; OCP's
`tuiInputReady()` matches `/\? for shortcuts/`. These are different events, so the ~1.0 s
boot figure is **not** directly comparable to OCP's `BOOT_MS` gate (default cap 4000 ms). It
does not affect the conclusions (1 s ≪ 6 s TTFT), but it is not apples-to-apples.
- The first version of this harness reported TTFT **0.08 s** — a false positive: the prompt
literally contained the marker string it was grepping for, so the match fired the instant the
prompt was pasted. Fixed by describing the marker instead of spelling it. **The script exited 0
and "successfully" produced 5 samples both times** — exit status proves nothing here.
@@ -0,0 +1,3 @@
[] | ▝▜█████▛▘ Sonnet 5 with xhigh effort · Claude Max
[--effort low] | ▝▜█████▛▘ Sonnet 5 with low effort · Claude Max
[--bare] | ▝▜█████▛▘ Sonnet 5 with xhigh effort · API Usage Billing
+128
View File
@@ -0,0 +1,128 @@
#!/usr/bin/env bash
# OCP TUI-mode latency floor harness — see README.md in this directory.
#
# 目的:回答一个问题——如果把 OCP 现有的两个已知开销砍掉
# (a) 每请求 spawn + boot(可用预热进程池消除)
# (b) 假流式(等 turn_duration 才返回,可用增量读 pane 消除)
# 之后,订阅池路径的**真实 TTFT 地板**是多少?
#
# 判据:地板 ≤ 4s → OCP 作为"省钱选项"可行;> 8s → 死透,不再讨论。
#
# 红线:
# - 不经过生产 OCP 服务(:3456)—— 直接起 tmux+claudeOCP 进程零干扰
# - tmux session 前缀用 zhiyin-floor-**不是** ocp-tui-),避免被 OCP 的
# reaper 当成自己的会话杀掉,也避免我们杀到它的
# - 用 real HOME(凭据)—— scratch HOME + symlink 凭据会 fork OAuth 导致 401
# (见跨机记忆 tui_scratch_home_credential_fork
set -uo pipefail
N=${1:-5}
MODEL=${MODEL:-claude-sonnet-5}
EXTRA_ARGS=${EXTRA_ARGS:-} # 额外 CLI 参数(如 --effort low --bare
TAG=${TAG:-baseline}
OUT=${OUT:-$(dirname "$0")/measurements.jsonl}
PROMPT_FILE=$(mktemp)
PREFIX="zhiyin-floor"
mkdir -p "$(dirname "$OUT")"
# ── 构造提示:~2000 token 的假会议转写 + 明确的起始标记 ────────────────
# 单行(多行会在 tmux send-keys 时提前触发 Enter
build_prompt() {
local seg="Speaker A said the quarterly pipeline is tracking behind plan and the enterprise segment needs a different motion. Speaker B replied that the current onboarding flow loses roughly a third of trial accounts before the first integration is complete. They debated whether the fix belongs in product or in customer success. "
local body=""
for _ in $(seq 1 22); do body+="$seg"; done
printf '%s' "You are a real-time meeting copilot. Meeting transcript so far: $body --- Task: produce ONE prompt card as compact JSON with keys: points (array of 3 short Chinese bullet points), keyline (one English sentence the user can read aloud). IMPORTANT: your reply MUST begin with three hash characters immediately followed by the uppercase word CARD (no space between them), then the JSON. No preamble, no markdown fences." > "$PROMPT_FILE"
}
build_prompt
PROMPT_CHARS=$(wc -c < "$PROMPT_FILE" | tr -d ' ')
now_ms() { python3 -c 'import time;print(int(time.time()*1000))'; }
echo "配置: $TAG 参数: [$EXTRA_ARGS]"
echo "模型: $MODEL 样本: $N 提示长度: ${PROMPT_CHARS} chars (≈$((PROMPT_CHARS/4)) token)"
echo "输出: $OUT"
echo
for i in $(seq 1 "$N"); do
SESS="${PREFIX}-$$-$i"
SID=$(uuidgen)
# ── 冷启动:spawn + 等输入框就绪 ─────────────────────────────────
T_SPAWN=$(now_ms)
tmux new-session -d -s "$SESS" -x 200 -y 50 \
-e CLAUDE_CODE_DISABLE_CLAUDE_MDS=1 \
-e CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 \
-c "$HOME" \
"claude --model $MODEL --session-id $SID --strict-mcp-config --disallowedTools 'mcp__*' $EXTRA_ARGS" 2>/dev/null
if [ $? -ne 0 ]; then echo "[$i] tmux spawn 失败,跳过"; continue; fi
# 轮询输入框就绪(claude TUI 的输入提示符)
READY=0
for _ in $(seq 1 150); do # 上限 15s
PANE=$(tmux capture-pane -p -t "$SESS" 2>/dev/null || true)
if grep -qE '│ >||Try "' <<<"$PANE"; then READY=1; break; fi
sleep 0.1
done
T_READY=$(now_ms)
BOOT_MS=$((T_READY - T_SPAWN))
if [ "$READY" -ne 1 ]; then
echo "[$i] 启动超时(${BOOT_MS}ms),pane 末 3 行:"
tmux capture-pane -p -t "$SESS" 2>/dev/null | tail -3 | sed 's/^/ /'
tmux kill-session -t "$SESS" 2>/dev/null
continue
fi
# ── 热态:粘提示 → 回车 → 量首 token ─────────────────────────────
tmux send-keys -t "$SESS" -l "$(cat "$PROMPT_FILE")" 2>/dev/null
sleep 0.4 # 让粘贴落地(OCP 用 400ms 轮询粒度)
T0=$(now_ms)
tmux send-keys -t "$SESS" Enter 2>/dev/null
TTFT_MS=-1
for _ in $(seq 1 600); do # 上限 60s
if tmux capture-pane -p -t "$SESS" 2>/dev/null | grep -q '###CARD'; then
TTFT_MS=$(( $(now_ms) - T0 )); break
fi
sleep 0.1
done
# ── 完整回答:pane 连续 2s 不再变化 ──────────────────────────────
COMPLETE_MS=-1
if [ "$TTFT_MS" -ge 0 ]; then
LAST=""; STABLE=0
for _ in $(seq 1 900); do # 上限 90s
CUR=$(tmux capture-pane -p -t "$SESS" 2>/dev/null | cksum)
if [ "$CUR" = "$LAST" ]; then
STABLE=$((STABLE+1))
[ "$STABLE" -ge 20 ] && { COMPLETE_MS=$(( $(now_ms) - T0 - 2000 )); break; }
else
STABLE=0; LAST="$CUR"
fi
sleep 0.1
done
fi
printf '{"i":%d,"tag":"%s","model":"%s","extra_args":"%s","prompt_chars":%s,"boot_ms":%d,"ttft_ms":%d,"complete_ms":%d}\n' \
"$i" "$TAG" "$MODEL" "$EXTRA_ARGS" "$PROMPT_CHARS" "$BOOT_MS" "$TTFT_MS" "$COMPLETE_MS" | tee -a "$OUT"
tmux kill-session -t "$SESS" 2>/dev/null
sleep 1
done
rm -f "$PROMPT_FILE"
echo
echo "=== 汇总 ==="
python3 - "$OUT" <<'EOF'
import json,sys,statistics
rows=[json.loads(l) for l in open(sys.argv[1]) if l.strip()]
ok=[r for r in rows if r['ttft_ms']>=0]
if not ok: print("无有效样本"); sys.exit()
def s(k):
v=[r[k] for r in ok if r[k]>=0]
return f"n={len(v)} 中位={statistics.median(v)/1000:.2f}s 最小={min(v)/1000:.2f}s 最大={max(v)/1000:.2f}s" if v else "无"
print(f" 冷启动 boot : {s('boot_ms')} ← 预热进程池可完全消除")
print(f" TTFT(首 token : {s('ttft_ms')} ★ 这就是地板")
print(f" 完整回答 : {s('complete_ms')}")
print(f"\n 失败样本: {len(rows)-len(ok)}/{len(rows)}")
EOF
@@ -0,0 +1,15 @@
{"i": 1, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1077, "ttft_ms": 6172, "complete_ms": 9929}
{"i": 2, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1026, "ttft_ms": 6160, "complete_ms": 9996}
{"i": 3, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1010, "ttft_ms": 6437, "complete_ms": 9977}
{"i": 4, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1033, "ttft_ms": 5872, "complete_ms": 9944}
{"i": 5, "tag": "effort-low", "model": "claude-sonnet-5", "extra_args": "--effort low", "prompt_chars": 7451, "boot_ms": 1154, "ttft_ms": 6387, "complete_ms": 9993}
{"i":1,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":1300,"ttft_ms":8321,"complete_ms":9939}
{"i":2,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":1070,"ttft_ms":10347,"complete_ms":11320}
{"i":3,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":911,"ttft_ms":13061,"complete_ms":15163}
{"i":4,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":1441,"ttft_ms":9981,"complete_ms":11066}
{"i":5,"tag":"baseline","model":"claude-sonnet-5","extra_args":"","prompt_chars":7451,"boot_ms":1036,"ttft_ms":17189,"complete_ms":17985}
{"i":1,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":429,"ttft_ms":-1,"complete_ms":-1}
{"i":2,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":437,"ttft_ms":-1,"complete_ms":-1}
{"i":3,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":444,"ttft_ms":-1,"complete_ms":-1}
{"i":4,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":446,"ttft_ms":-1,"complete_ms":-1}
{"i":5,"tag":"bare","model":"claude-sonnet-5","extra_args":"--bare","prompt_chars":7451,"boot_ms":441,"ttft_ms":-1,"complete_ms":-1}
+26
View File
@@ -377,12 +377,38 @@ export function buildTuiCmd(claudeBin, model, sessionId, ehome, entrypointMode)
} else {
toolArgs = ["--strict-mcp-config", "--disallowedTools", shq("mcp__*")];
}
// Effort: pass --effort EXPLICITLY. Without it, the pane's claude inherits a
// HOME-dependent effortLevel — real-home mode inherits the operator's
// ~/.claude/settings.json (whatever they set for their own interactive use),
// env-token scratch mode inherits claude's built-in default (prepareTuiHome never
// writes effortLevel) — so latency silently depends on which HOME mode
// resolveTuiHome() picked AND on an unrelated operator setting. Pinning it here
// removes both. Measured (docs/plans/2026-07-13-tui-latency): explicit low cuts
// direct-spawn TTFT p50 10.35s → 6.17s (40%) and collapses the spread ~15×;
// banner-verified to stay on the subscription pool (`· Claude Max`).
// OCP_TUI_EFFORT=inherit restores the pre-flag argv byte-for-byte (no --effort).
// An unknown value falls back to the default rather than reaching claude's argv:
// a typo'd --effort value must not risk a spawn-time usage error in the pane.
const EFFORT_LEVELS = ["low", "medium", "high", "xhigh", "max"]; // claude 2.1.207 --help
const effortRaw = (process.env.OCP_TUI_EFFORT || "low").trim().toLowerCase();
let effortArgs;
if (effortRaw === "inherit") {
effortArgs = [];
} else if (EFFORT_LEVELS.includes(effortRaw)) {
effortArgs = ["--effort", effortRaw];
} else {
console.error(`[tui] invalid OCP_TUI_EFFORT=${JSON.stringify(process.env.OCP_TUI_EFFORT)}; using "low" (valid: ${EFFORT_LEVELS.join("|")}, or "inherit" to omit the flag)`);
effortArgs = ["--effort", "low"];
}
return [
envPrefix,
shq(claudeBin),
"--model", shq(model),
"--session-id", sessionId,
...toolArgs,
...effortArgs,
].join(" ");
}
+59
View File
@@ -1795,6 +1795,65 @@ test("buildTuiCmd shq-escapes a token containing shell metacharacters (no inject
}
});
// OCP_TUI_EFFORT (TUI latency, docs/plans/2026-07-13-tui-latency): the pane's claude
// must get an EXPLICIT --effort so its effort never depends on which HOME mode
// resolveTuiHome() picked (real-home inherits the operator's settings.json effortLevel;
// env-token scratch inherits claude's built-in default).
test("buildTuiCmd passes --effort low by default (OCP_TUI_EFFORT unset)", () => {
const save = process.env.OCP_TUI_EFFORT;
try {
delete process.env.OCP_TUI_EFFORT;
const cmd = buildTuiCmd("/usr/bin/claude", "m", "sid-eff1", "/home/u", "cli");
assert.ok(cmd.includes("--effort low"), "default must pin --effort low");
} finally {
if (save === undefined) delete process.env.OCP_TUI_EFFORT;
else process.env.OCP_TUI_EFFORT = save;
}
});
test("buildTuiCmd honors an explicit OCP_TUI_EFFORT level (case/space-normalized)", () => {
const save = process.env.OCP_TUI_EFFORT;
try {
process.env.OCP_TUI_EFFORT = " XHigh ";
const cmd = buildTuiCmd("/usr/bin/claude", "m", "sid-eff2", "/home/u", "cli");
assert.ok(cmd.includes("--effort xhigh"), "explicit level must be passed, normalized");
assert.ok(!cmd.includes("--effort low"), "default must not also appear");
} finally {
if (save === undefined) delete process.env.OCP_TUI_EFFORT;
else process.env.OCP_TUI_EFFORT = save;
}
});
test("buildTuiCmd OCP_TUI_EFFORT=inherit omits --effort entirely (pre-flag argv)", () => {
const save = process.env.OCP_TUI_EFFORT;
try {
process.env.OCP_TUI_EFFORT = "inherit";
const cmd = buildTuiCmd("/usr/bin/claude", "m", "sid-eff3", "/home/u", "cli");
assert.ok(!/--effort/.test(cmd), "inherit must not add --effort");
} finally {
if (save === undefined) delete process.env.OCP_TUI_EFFORT;
else process.env.OCP_TUI_EFFORT = save;
}
});
test("buildTuiCmd falls back to --effort low on an invalid OCP_TUI_EFFORT (never reaches argv)", () => {
const save = process.env.OCP_TUI_EFFORT;
const savedErr = console.error;
try {
process.env.OCP_TUI_EFFORT = "ludicrous'; rm -rf /;'";
let warned = "";
console.error = (...a) => { warned = a.join(" "); };
const cmd = buildTuiCmd("/usr/bin/claude", "m", "sid-eff4", "/home/u", "cli");
assert.ok(cmd.includes("--effort low"), "invalid value must fall back to low");
assert.ok(!cmd.includes("ludicrous"), "invalid raw value must NOT reach the shell string");
assert.ok(/invalid OCP_TUI_EFFORT/.test(warned), "must log a warning");
} finally {
console.error = savedErr;
if (save === undefined) delete process.env.OCP_TUI_EFFORT;
else process.env.OCP_TUI_EFFORT = save;
}
});
test("buildTuiCmd OCP_TUI_FULL_TOOLS=1 grants -p-equivalent tool surface (single-user opt-in)", () => {
const save = { ...process.env };
const restore = () => {