mirror of
https://github.com/dtzp555-max/ocp.git
synced 2026-07-22 13:35:08 +00:00
Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
2144e6769f |
+2
-17
@@ -1,25 +1,10 @@
|
|||||||
# Changelog
|
# Changelog
|
||||||
|
|
||||||
## v3.23.0 — 2026-07-17
|
## Unreleased
|
||||||
|
|
||||||
Minor release. Headline: **the default `sonnet` alias now resolves to Claude Sonnet 5** — a behavior change for every request that omits `model` (pin `claude-sonnet-4-6` by full ID to keep the previous default). Also: Windows-safe upgrade snapshots, two upgrade-system reliability fixes from a live fleet update, the `CLAUDE_SYSTEM_PROMPT` env var made functional, cache-key honesty for config changes, a billing-policy status correction (the 2026-06-15 `-p` split is PAUSED by Anthropic), and a major README restructure. No new endpoint; no new `cli.js` wire behavior. Every code PR carried a fresh-context reviewer (Iron Rule 10).
|
|
||||||
|
|
||||||
### Changed
|
### Changed
|
||||||
|
|
||||||
- **Default `sonnet` alias → `claude-sonnet-5` (#168, contributed by @vvlasy-openclaw).** The `sonnet` alias (the model used for every `/v1/chat/completions` request that omits `model`, and OpenClaw's OCP primary via `ocp-connect`) now resolves to `claude-sonnet-5` instead of `claude-sonnet-4-6`. `claude-sonnet-4-6` remains available by full ID for pinning. Mirrors the shipped Claude CLI's own `latest_per_family` mapping (`sonnet → claude-sonnet-5`, verified from binary 2.1.211). Split out from the additive model entry (#152) per Iron Rule 11.
|
- **Default `sonnet` alias → `claude-sonnet-5`.** The `sonnet` alias (the model used for every `/v1/chat/completions` request that omits `model`, and OpenClaw's OCP primary via `ocp-connect`) now resolves to `claude-sonnet-5` instead of `claude-sonnet-4-6`. `claude-sonnet-4-6` remains available by full ID for pinning. This is a behavior change for clients relying on the default — pin `claude-sonnet-4-6` explicitly to retain the previous model. Split out from the additive `claude-sonnet-5` model entry (#152) per Iron Rule 11.
|
||||||
- **`CLAUDE_SYSTEM_PROMPT` is now functional (#175).** The var was read, documented, and echoed on `/health.systemPrompt` but never reached a request (dead since the `APPEND_SYSTEM_PROMPT` retirement). It is now appended (last, trimmed) to the composed system prompt on the default `-p` path via the new pure `lib/prompt.mjs`; TUI-mode panes are unaffected. Unset ⇒ byte-identical composition to before. README § Environment Variables documents it, including the cache caveat below.
|
|
||||||
|
|
||||||
### Fixed
|
|
||||||
|
|
||||||
- **Windows-safe upgrade snapshot paths (#167, contributed by @nyxst4ck).** Snapshot directory timestamps now use `-` instead of `:` (Windows forbids `:` in names); legacy colon-named snapshots keep parsing, and `listSnapshots` now orders by **parsed timestamp** (with a deterministic name tie-breaker) so mixed legacy/new names sort chronologically — the initial revision's raw-string sort could delete the newest recovery snapshot at the format boundary and was caught in review; regression tests pin the same-hour mixed-format case.
|
|
||||||
- **`ocp update` reliability — two live-incident fixes (#174, closes #173).** (1) The doctor now runs `git fetch --tags` (offline-tolerant) before computing `latest_version` — previously it compared against the locally cached `origin/main`, so machines that hadn't pulled since a release reported "Already at latest" forever. (2) Post-flight now asserts `/health.version` equals the upgrade target (new `postFlightOk` predicate) instead of accepting any `auth.ok` — a stale orphan process holding the port used to pass post-flight while still serving the old version; the failure message now reports the last-seen version and points at `ss -ltnp`/`lsof -i`.
|
|
||||||
- **Response-cache key now carries a boot-config epoch (#177, closes #176).** The persistent cache keyed on model+key+params+messages but not on server config that shapes answers (`CLAUDE_SYSTEM_PROMPT`, wrapper text, `CLAUDE_ALLOWED_TOOLS`, `CLAUDE_NO_CONTEXT`) — changing any of these and restarting could serve stale-config answers until TTL expiry. A sha256 config-epoch is folded into every key; any config change is an instant whole-cache invalidation. One-time side effect: existing cache entries miss once after this upgrade.
|
|
||||||
|
|
||||||
### Docs
|
|
||||||
|
|
||||||
- **Billing-policy status corrected (#171).** Anthropic **paused** the announced 2026-06-15 `claude -p` billing split on its effective date (official help-article citation in README § How It Works): the default `-p` path currently bills the subscription, and TUI-mode is reframed as the ready-made **hedge** for if/when a reworked change lands. All in-force assertions of the split are now date-stamped and conditioned.
|
|
||||||
- **LAN mode scoped to chat-class workloads (#171).** New "workload fit" paragraph: multi-device OCP is for text-in/text-out workloads; client-machine coding agents are architecturally out of scope (tools execute on the OCP host).
|
|
||||||
- **README restructured, 1205 → ~500 lines (#172).** Operations-manual content moved to `docs/lan-mode.md`, `docs/tui-mode.md`, `docs/troubleshooting.md`, `docs/upgrading.md` (verbatim moves + two canonical dedups; zero content loss verified section-by-section). README keeps the quickstart, the release-kit-pinned reference tables, and summary stubs with links. Plus a staleness sweep (#170): 6-model examples, removal of the never-existed `ocp stop` command, `ocp-connect` claims corrected, current version examples.
|
|
||||||
|
|
||||||
## v3.22.1 — 2026-07-17
|
## v3.22.1 — 2026-07-17
|
||||||
|
|
||||||
|
|||||||
@@ -222,13 +222,13 @@ The canonical list lives in [`models.json`](./models.json) — the single source
|
|||||||
| `CLAUDE_MAX_CONCURRENT` | `8` | Max concurrent claude processes (`-p`/stream-json path) |
|
| `CLAUDE_MAX_CONCURRENT` | `8` | Max concurrent claude processes (`-p`/stream-json path) |
|
||||||
| `CLAUDE_MAX_QUEUE` | `16` | Max requests **waiting** for a `-p` concurrency slot. Beyond `CLAUDE_MAX_CONCURRENT`, requests queue (up to this cap) instead of being rejected; when the queue is **also** full, the request gets `HTTP 429` + `Retry-After` (not an opaque 500). Surfaced on `/health.concurrency` + `/health.stats.queueRejections`. |
|
| `CLAUDE_MAX_QUEUE` | `16` | Max requests **waiting** for a `-p` concurrency slot. Beyond `CLAUDE_MAX_CONCURRENT`, requests queue (up to this cap) instead of being rejected; when the queue is **also** full, the request gets `HTTP 429` + `Retry-After` (not an opaque 500). Surfaced on `/health.concurrency` + `/health.stats.queueRejections`. |
|
||||||
| `CLAUDE_QUEUE_RETRY_AFTER` | `5` | Seconds advertised in the `Retry-After` header on a `-p` concurrency-overflow `429`. |
|
| `CLAUDE_QUEUE_RETRY_AFTER` | `5` | Seconds advertised in the `Retry-After` header on a `-p` concurrency-overflow `429`. |
|
||||||
| `CLAUDE_MAX_PROMPT_CHARS` | *(derived)* | Prompt truncation limit in chars. Default derives from the models.json SPOT: `max(contextWindow) × 3` — currently **600,000** (≈150–200k tokens). Setting this env var (or the runtime settings API) overrides the derivation absolutely. See [ADR 0009](docs/adr/0009-spot-derived-prompt-budget.md). Note: very large prompts burn subscription-window quota quickly and slow TTFT; the TUI-mode paste path is untested beyond ~hundreds of KB. |
|
| `CLAUDE_MAX_PROMPT_CHARS` | `150000` | Prompt truncation limit (chars) |
|
||||||
| `CLAUDE_SESSION_TTL` | `3600000` | Session expiry (ms, default: 1 hour) |
|
| `CLAUDE_SESSION_TTL` | `3600000` | Session expiry (ms, default: 1 hour) |
|
||||||
| `CLAUDE_CACHE_TTL` | `0` | Response cache TTL (ms, 0 = disabled). Set to e.g. `300000` for 5-min cache. See [Response Cache](#response-cache). |
|
| `CLAUDE_CACHE_TTL` | `0` | Response cache TTL (ms, 0 = disabled). Set to e.g. `300000` for 5-min cache. See [Response Cache](#response-cache). |
|
||||||
| `CLAUDE_ALLOWED_TOOLS` | `Bash,Read,...,Agent` | Comma-separated tools to pre-approve |
|
| `CLAUDE_ALLOWED_TOOLS` | `Bash,Read,...,Agent` | Comma-separated tools to pre-approve |
|
||||||
| `CLAUDE_SKIP_PERMISSIONS` | `false` | Bypass all permission checks |
|
| `CLAUDE_SKIP_PERMISSIONS` | `false` | Bypass all permission checks |
|
||||||
| `CLAUDE_MCP_CONFIG` | *(unset)* | Path to an MCP server config JSON, passed to the spawned `claude` as `--mcp-config` (both the `-p` path and TUI `OCP_TUI_FULL_TOOLS` panes) |
|
| `CLAUDE_MCP_CONFIG` | *(unset)* | Path to an MCP server config JSON, passed to the spawned `claude` as `--mcp-config` (both the `-p` path and TUI `OCP_TUI_FULL_TOOLS` panes) |
|
||||||
| `CLAUDE_SYSTEM_PROMPT` | *(unset)* | Operator-wide system-prompt text appended (last) to every request's composed system prompt on the default `-p` path. TUI-mode panes are unaffected (they keep the interactive CLI's own system prompt). Echoed truncated on `/health.systemPrompt`. Note: changing this value and restarting auto-invalidates the response cache (the key carries a boot-config epoch, #177). |
|
| `CLAUDE_SYSTEM_PROMPT` | *(unset)* | Operator-wide system-prompt text appended (last) to every request's composed system prompt on the default `-p` path. TUI-mode panes are unaffected (they keep the interactive CLI's own system prompt). Echoed truncated on `/health.systemPrompt`. Note: the response cache key does not include server config — after changing this value, flush the cache (`ocp clear`) or let TTL expire. |
|
||||||
| `CLAUDE_NO_CONTEXT` | `false` | Suppress CLAUDE.md and auto-memory injection (pure API mode) |
|
| `CLAUDE_NO_CONTEXT` | `false` | Suppress CLAUDE.md and auto-memory injection (pure API mode) |
|
||||||
| `PROXY_API_KEY` | *(unset)* | Bearer token for shared-mode authentication |
|
| `PROXY_API_KEY` | *(unset)* | Bearer token for shared-mode authentication |
|
||||||
| `PROXY_ANONYMOUS_KEY` | *(unset)* | Well-known anonymous key (multi mode) — this exact string bypasses `validateKey()` and grants public access. Exposed via `/health.anonymousKey` only to localhost, or to all callers when `PROXY_ADVERTISE_ANON_KEY=1`. Full setup + security notes: [docs/lan-mode.md § Anonymous Access](docs/lan-mode.md#anonymous-access-optional). |
|
| `PROXY_ANONYMOUS_KEY` | *(unset)* | Well-known anonymous key (multi mode) — this exact string bypasses `validateKey()` and grants public access. Exposed via `/health.anonymousKey` only to localhost, or to all callers when `PROXY_ADVERTISE_ANON_KEY=1`. Full setup + security notes: [docs/lan-mode.md § Anonymous Access](docs/lan-mode.md#anonymous-access-optional). |
|
||||||
|
|||||||
@@ -1,54 +0,0 @@
|
|||||||
# ADR 0009 — Prompt-char budget derives from the models.json SPOT
|
|
||||||
|
|
||||||
Date: 2026-07-18
|
|
||||||
Status: Accepted (maintainer directive, 2026-07-18: "37.5k 截断未免太短了吧 … 这个在现在还适用吗")
|
|
||||||
|
|
||||||
## Context
|
|
||||||
|
|
||||||
`MAX_PROMPT_CHARS` (the tail-first truncation guard in `messagesToPrompt`) defaulted to a
|
|
||||||
hand-set constant of 150,000 chars ≈ 37.5k English tokens — set in the 200k-window era as a
|
|
||||||
runaway-context guard. Meanwhile `models.json` advertises `contextWindow: 200000` for every
|
|
||||||
model (and the underlying CLI registry carries 1M native windows for Opus 4.8 / Sonnet 5), and
|
|
||||||
`scripts/sync-openclaw.mjs` feeds that 200k into OpenClaw's compaction budget. The result was
|
|
||||||
a standing dishonesty identified in the PR #152 review: **no advertised contextWindow value was
|
|
||||||
true**, because the proxy silently guillotined every request at ~37.5k tokens — roughly 5×
|
|
||||||
below the advertised window — logging only a server-side warning the client never sees.
|
|
||||||
|
|
||||||
Raising the constant to another hand-set number would rot the same way. Following the model's
|
|
||||||
native 1M directly is also wrong: chars ≠ tokens (CJK runs ~1–1.5 chars/token vs ~4 for
|
|
||||||
English, so a 1M-token char cap would let CJK text sail past the model's real window into an
|
|
||||||
upstream rejection), single near-window requests can consume a large fraction of a 5-hour
|
|
||||||
subscription quota window, and the TUI paste path is untested at megabyte scale.
|
|
||||||
|
|
||||||
## Decision
|
|
||||||
|
|
||||||
The default budget **derives from the SPOT** instead of being a constant:
|
|
||||||
|
|
||||||
```
|
|
||||||
MAX_PROMPT_CHARS (default) = max(models.json models[].contextWindow) × 3 chars/token
|
|
||||||
= 200000 × 3 = 600,000 chars today
|
|
||||||
```
|
|
||||||
|
|
||||||
Implemented as the pure `derivePromptCharBudget(models, {charsPerToken = 3, floor = 150000})`
|
|
||||||
in `lib/prompt.mjs` (unit-tested; floor guards degenerate SPOT states). The multiplier ×3 is
|
|
||||||
deliberately conservative: full window for English, and CJK text reaches the model's real
|
|
||||||
window at roughly the same point the cap fires — so OCP truncates gracefully (tail-first)
|
|
||||||
instead of the upstream rejecting outright.
|
|
||||||
|
|
||||||
`CLAUDE_MAX_PROMPT_CHARS` (env) and the runtime settings API remain **absolute overrides**;
|
|
||||||
the derivation applies only when neither is set.
|
|
||||||
|
|
||||||
## Consequences
|
|
||||||
|
|
||||||
- The advertised `contextWindow: 200000` becomes honest: the proxy now actually accepts
|
|
||||||
prompts of that order (English ≈150–200k tokens) before truncating.
|
|
||||||
- If `models.json` ever advertises a larger window (e.g. 1M for the 1M-native models), the
|
|
||||||
budget scales automatically — no code change. Whether to advertise 1M is a **separate,
|
|
||||||
deliberate decision** (quota burn per request, OpenClaw compaction memory, TUI paste
|
|
||||||
limits) and is explicitly NOT made by this ADR; the current recommendation is to keep
|
|
||||||
200000 advertised until a real >200k use case appears.
|
|
||||||
- One-time behavior change: requests between 150k and 600k chars that were previously
|
|
||||||
truncated now pass through whole — longer TTFT and higher quota consumption for those
|
|
||||||
requests, by design.
|
|
||||||
- The truncation mechanism, logging, and the multimodal-path budget threading (PR #154's F2,
|
|
||||||
pending) are unchanged — only the default value's provenance changed.
|
|
||||||
@@ -25,7 +25,6 @@ New ADRs increment from the highest existing number. Filenames are
|
|||||||
| [0006](0006-openai-shim-scope.md) | OpenAI Shim Scope | The Class A / Class B taxonomy. Class A endpoints (`cli.js`-mirror) keep Rules 1–5 verbatim; Class B endpoints (OCP-owned compatibility surface — `/v1/chat/completions`, `/v1/models`, admin endpoints) are anchored to OpenAI's spec (B.1) or to an authorizing ADR (B.2). Triggered by PR #99 (external `response_format` honoring). Grandfathers the existing B.2 inventory at v3.16.4. |
|
| [0006](0006-openai-shim-scope.md) | OpenAI Shim Scope | The Class A / Class B taxonomy. Class A endpoints (`cli.js`-mirror) keep Rules 1–5 verbatim; Class B endpoints (OCP-owned compatibility surface — `/v1/chat/completions`, `/v1/models`, admin endpoints) are anchored to OpenAI's spec (B.1) or to an authorizing ADR (B.2). Triggered by PR #99 (external `response_format` honoring). Grandfathers the existing B.2 inventory at v3.16.4. |
|
||||||
| [0007](0007-tui-interactive-mode.md) | TUI Interactive Mode | Why TUI-mode spawns an interactive `claude` in a tmux pane (no `-p`) to reach the **subscription** billing pool (`cc_entrypoint=cli`) rather than the metered Agent SDK pool. Owns the TUI spawn machinery: entrypoint labeling, credential-isolated home, MCP hard-disable, session namespace + defunct-session reaping, the independent concurrency bound, and the `/health` `tui` block. **Single-user only** — hard FATAL on multi-user configs. |
|
| [0007](0007-tui-interactive-mode.md) | TUI Interactive Mode | Why TUI-mode spawns an interactive `claude` in a tmux pane (no `-p`) to reach the **subscription** billing pool (`cc_entrypoint=cli`) rather than the metered Agent SDK pool. Owns the TUI spawn machinery: entrypoint labeling, credential-isolated home, MCP hard-disable, session namespace + defunct-session reaping, the independent concurrency bound, and the `/health` `tui` block. **Single-user only** — hard FATAL on multi-user configs. |
|
||||||
| [0008](0008-tui-warm-pane-pool.md) | TUI Warm Pane Pool | Why `OCP_TUI_POOL_SIZE` pre-boots **single-use** `claude` panes (one turn each, own `--session-id`) — and why reuse is forbidden (`transcript.mjs` returns the last assistant entry in the file, so a reused session leaks the earlier turn's text). Measured −41% end-to-end. Defines the pool↔reaper invariant (exemption by exact name from a live registry; drain before every sweep so `kill-server` zombie reaping survives) and the standing idle-process cost. Extends ADR 0007. |
|
| [0008](0008-tui-warm-pane-pool.md) | TUI Warm Pane Pool | Why `OCP_TUI_POOL_SIZE` pre-boots **single-use** `claude` panes (one turn each, own `--session-id`) — and why reuse is forbidden (`transcript.mjs` returns the last assistant entry in the file, so a reused session leaks the earlier turn's text). Measured −41% end-to-end. Defines the pool↔reaper invariant (exemption by exact name from a live registry; drain before every sweep so `kill-server` zombie reaping survives) and the standing idle-process cost. Extends ADR 0007. |
|
||||||
| [0009](0009-spot-derived-prompt-budget.md) | SPOT-Derived Prompt Budget | Why `MAX_PROMPT_CHARS`'s default is `max(models.json contextWindow) × 3 chars/token` (600k chars today) instead of a hand-set constant — the old 150k silently under-delivered the advertised window ~5×. ×3 is the CJK-safe multiplier; env/settings stay absolute overrides; whether to advertise 1M windows is explicitly a separate decision. |
|
|
||||||
|
|
||||||
## When to write a new ADR
|
## When to write a new ADR
|
||||||
|
|
||||||
|
|||||||
@@ -15,40 +15,3 @@ export function appendOperatorPrompt(base, operatorAppend) {
|
|||||||
const op = typeof operatorAppend === "string" ? operatorAppend.trim() : "";
|
const op = typeof operatorAppend === "string" ? operatorAppend.trim() : "";
|
||||||
return op ? `${base}\n\n${op}` : base;
|
return op ? `${base}\n\n${op}` : base;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Derive the default prompt-char budget from the models.json SPOT (ADR 0009).
|
|
||||||
//
|
|
||||||
// The old default was a hand-set constant (150000 chars ≈ 37.5k English tokens) from the
|
|
||||||
// 200k-window era — silently far below what the advertised contextWindow promises. Instead
|
|
||||||
// of picking a new constant that will also rot, the default now FOLLOWS the SPOT:
|
|
||||||
//
|
|
||||||
// budget = max(models[].contextWindow) × charsPerToken
|
|
||||||
//
|
|
||||||
// charsPerToken = 3 is deliberately conservative: English runs ~4 chars/token, CJK ~1–1.5.
|
|
||||||
// At ×3, a 200k-token window yields 600,000 chars — full window for English, and CJK text
|
|
||||||
// hits the model's real window at roughly the same point the cap fires, so we truncate
|
|
||||||
// (graceful, tail-first) rather than let the upstream reject the request outright.
|
|
||||||
//
|
|
||||||
// The floor guards the degenerate cases (empty/missing models[], absent contextWindow):
|
|
||||||
// fall back to the historical constant rather than 0 — a zero budget would truncate every
|
|
||||||
// request to nothing, which is fail-OPEN in the "serve garbage" sense. CLAUDE_MAX_PROMPT_CHARS
|
|
||||||
// remains an absolute operator override at the call site (server.mjs); this function is only
|
|
||||||
// the unset-env default.
|
|
||||||
export function derivePromptCharBudget(models, { charsPerToken = 3, floor = 150000 } = {}) {
|
|
||||||
const windows = (Array.isArray(models) ? models : [])
|
|
||||||
.map(m => m?.contextWindow)
|
|
||||||
.filter(w => Number.isFinite(w) && w > 0);
|
|
||||||
if (windows.length === 0) return floor;
|
|
||||||
return Math.max(floor, Math.max(...windows) * charsPerToken);
|
|
||||||
}
|
|
||||||
|
|
||||||
// Resolve the effective budget from the env var + SPOT. TRUTHINESS (not != null) on the env
|
|
||||||
// value deliberately: an EMPTY value ("CLAUDE_MAX_PROMPT_CHARS=" in a systemd EnvironmentFile
|
|
||||||
// or .env) must mean "use the default" — exactly the old `parseInt(env || "150000")` contract.
|
|
||||||
// Treating "" as explicit gives parseInt("") = NaN, and a NaN cap silently DISABLES the
|
|
||||||
// runaway-context guard while injecting a false "[System] Note: 0 older messages were
|
|
||||||
// truncated" line into every prompt (caught in PR #179 review). Non-empty garbage still
|
|
||||||
// parses to NaN — the pre-existing class, slated for parseIntEnv routing in PR #154.
|
|
||||||
export function resolvePromptCharBudget(rawEnv, models, opts) {
|
|
||||||
return rawEnv ? parseInt(rawEnv, 10) : derivePromptCharBudget(models, opts);
|
|
||||||
}
|
|
||||||
|
|||||||
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
{
|
{
|
||||||
"name": "open-claude-proxy",
|
"name": "open-claude-proxy",
|
||||||
"version": "3.23.0",
|
"version": "3.22.1",
|
||||||
"description": "OCP (Open Claude Proxy) — use your Claude Pro/Max subscription as an OpenAI-compatible API for any IDE. Works with Cline, OpenCode, Aider, Continue.dev, OpenClaw, and more.",
|
"description": "OCP (Open Claude Proxy) — use your Claude Pro/Max subscription as an OpenAI-compatible API for any IDE. Works with Cline, OpenCode, Aider, Continue.dev, OpenClaw, and more.",
|
||||||
"type": "module",
|
"type": "module",
|
||||||
"bin": {
|
"bin": {
|
||||||
|
|||||||
+2
-8
@@ -49,7 +49,7 @@ import { TuiSemaphore, SemaphoreAbortError, recordTuiEntrypoint, buildTuiHealthB
|
|||||||
import { TuiPanePool, resolvePoolSize, POOL_MAX_SIZE } from "./lib/tui/pool.mjs";
|
import { TuiPanePool, resolvePoolSize, POOL_MAX_SIZE } from "./lib/tui/pool.mjs";
|
||||||
import { TuiDeltaAssembler, DEFAULT_HOLDBACK_CHARS, resolveStreamHoldback } from "./lib/tui/stream.mjs";
|
import { TuiDeltaAssembler, DEFAULT_HOLDBACK_CHARS, resolveStreamHoldback } from "./lib/tui/stream.mjs";
|
||||||
import { createSerialMutex, createTtlCache, isTokenExpiring, orderLabelsLastGoodFirst } from "./lib/spawn-auth.mjs";
|
import { createSerialMutex, createTtlCache, isTokenExpiring, orderLabelsLastGoodFirst } from "./lib/spawn-auth.mjs";
|
||||||
import { appendOperatorPrompt, resolvePromptCharBudget } from "./lib/prompt.mjs";
|
import { appendOperatorPrompt } from "./lib/prompt.mjs";
|
||||||
|
|
||||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||||
const _pkg = JSON.parse(readFileSync(join(__dirname, "package.json"), "utf8"));
|
const _pkg = JSON.parse(readFileSync(join(__dirname, "package.json"), "utf8"));
|
||||||
@@ -1148,13 +1148,7 @@ function buildCliArgs(cliModel, systemPrompt) {
|
|||||||
// Truncation guard: if total chars exceed MAX_PROMPT_CHARS, keep the system
|
// Truncation guard: if total chars exceed MAX_PROMPT_CHARS, keep the system
|
||||||
// message(s) + first user message + last N messages, dropping the middle.
|
// message(s) + first user message + last N messages, dropping the middle.
|
||||||
// This prevents runaway context from gateway-side conversation accumulation.
|
// This prevents runaway context from gateway-side conversation accumulation.
|
||||||
//
|
let MAX_PROMPT_CHARS = parseInt(process.env.CLAUDE_MAX_PROMPT_CHARS || "150000", 10);
|
||||||
// Default is SPOT-DERIVED (ADR 0009): max(models.json contextWindow) × 3 chars/token —
|
|
||||||
// currently 200000 × 3 = 600,000 chars — instead of the old hand-set 150000 (≈37.5k
|
|
||||||
// English tokens), which silently under-delivered the advertised window by ~5×. The env
|
|
||||||
// var (and the runtime settings API below) remain absolute operator overrides. If
|
|
||||||
// models.json ever advertises a bigger window, this budget follows automatically.
|
|
||||||
let MAX_PROMPT_CHARS = resolvePromptCharBudget(process.env.CLAUDE_MAX_PROMPT_CHARS, modelsConfig.models);
|
|
||||||
|
|
||||||
// Flatten OpenAI content (string | array of parts) to plain text for the prompt.
|
// Flatten OpenAI content (string | array of parts) to plain text for the prompt.
|
||||||
// Array content: concatenate text parts; replace non-text parts (e.g. image_url)
|
// Array content: concatenate text parts; replace non-text parts (e.g. image_url)
|
||||||
|
|||||||
+1
-43
@@ -844,49 +844,7 @@ test("doctor falls back to currentVersion when origin/main unreachable (no stale
|
|||||||
// contract lives in lib/prompt.mjs. Mutation-proof: make appendOperatorPrompt
|
// contract lives in lib/prompt.mjs. Mutation-proof: make appendOperatorPrompt
|
||||||
// return `base` unconditionally and the first test fails; make it stop trimming
|
// return `base` unconditionally and the first test fails; make it stop trimming
|
||||||
// and the whitespace test fails.
|
// and the whitespace test fails.
|
||||||
import { appendOperatorPrompt, derivePromptCharBudget, resolvePromptCharBudget } from "./lib/prompt.mjs";
|
import { appendOperatorPrompt } from "./lib/prompt.mjs";
|
||||||
|
|
||||||
console.log("\nPrompt-char budget (ADR 0009 — SPOT-derived):");
|
|
||||||
|
|
||||||
// Mutation-proof: drop the ×charsPerToken and the first test fails; drop the
|
|
||||||
// Math.max floor guard and the floor tests fail; use min() instead of max() over
|
|
||||||
// windows and the largest-window test fails.
|
|
||||||
test("derivePromptCharBudget: LARGEST contextWindow × 3 chars/token", () => {
|
|
||||||
const models = [{ contextWindow: 200000 }, { contextWindow: 100000 }];
|
|
||||||
assert.equal(derivePromptCharBudget(models), 600000);
|
|
||||||
});
|
|
||||||
|
|
||||||
test("derivePromptCharBudget: matches the live models.json SPOT (200k → 600k today)", () => {
|
|
||||||
const spot = JSON.parse(tuiReadFileSync(new URL("./models.json", import.meta.url), "utf8"));
|
|
||||||
assert.equal(derivePromptCharBudget(spot.models), 600000);
|
|
||||||
});
|
|
||||||
|
|
||||||
test("derivePromptCharBudget: floor wins over a tiny/absent window; empty input → floor", () => {
|
|
||||||
assert.equal(derivePromptCharBudget([{ contextWindow: 1000 }]), 150000, "3k chars would truncate everything — floor guards it");
|
|
||||||
assert.equal(derivePromptCharBudget([]), 150000);
|
|
||||||
assert.equal(derivePromptCharBudget(undefined), 150000);
|
|
||||||
assert.equal(derivePromptCharBudget([{ id: "x" }, { contextWindow: "junk" }, { contextWindow: -5 }]), 150000);
|
|
||||||
});
|
|
||||||
|
|
||||||
test("derivePromptCharBudget: charsPerToken and floor are tunable parameters", () => {
|
|
||||||
assert.equal(derivePromptCharBudget([{ contextWindow: 1000000 }], { charsPerToken: 3 }), 3000000);
|
|
||||||
assert.equal(derivePromptCharBudget([], { floor: 42 }), 42);
|
|
||||||
});
|
|
||||||
|
|
||||||
// PR #179 review regression: EMPTY env value must mean "use the default" (the old
|
|
||||||
// `parseInt(env || "150000")` contract). Mutation-proof: switch the resolver's
|
|
||||||
// truthiness check to `!= null` and the empty-string test fails (NaN ≠ 600000).
|
|
||||||
test("resolvePromptCharBudget: empty/unset env → SPOT-derived default, never NaN", () => {
|
|
||||||
const models = [{ contextWindow: 200000 }];
|
|
||||||
assert.equal(resolvePromptCharBudget("", models), 600000, "CLAUDE_MAX_PROMPT_CHARS= (empty) must fall back to derived");
|
|
||||||
assert.equal(resolvePromptCharBudget(undefined, models), 600000);
|
|
||||||
});
|
|
||||||
|
|
||||||
test("resolvePromptCharBudget: a set env value overrides the derivation absolutely", () => {
|
|
||||||
const models = [{ contextWindow: 200000 }];
|
|
||||||
assert.equal(resolvePromptCharBudget("300000", models), 300000);
|
|
||||||
assert.equal(resolvePromptCharBudget("150000", models), 150000, "explicit legacy value wins over the bigger derived default");
|
|
||||||
});
|
|
||||||
|
|
||||||
console.log("\nSystem-prompt operator append:");
|
console.log("\nSystem-prompt operator append:");
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user