mirror of
https://github.com/dtzp555-max/ocp.git
synced 2026-07-27 07:55:07 +00:00
* feat(models): align maxTokens with the CLI registry (#195)
Every Opus entry and claude-sonnet-5 go 16384 -> 64000; claude-sonnet-4-6 and
claude-haiku-4-5 go to 32000 (haiku was 8192). These are the
max_output_tokens.default each model declares in the compiled CLI 2.1.220
registry, extracted id-anchored:
claude-opus-5 max_output_tokens:{default:64000,upper:128000}
claude-opus-4-8 max_output_tokens:{default:64000,upper:128000}
claude-opus-4-7 max_output_tokens:{default:64000,upper:128000}
claude-opus-4-6 max_output_tokens:{default:64000,upper:128000}
claude-sonnet-5 max_output_tokens:{default:64000,upper:128000}
claude-sonnet-4-6 max_output_tokens:{default:32000,upper:128000}
claude-haiku-4-5 max_output_tokens:{default:32000,upper:64000}
This is not cosmetic metadata. OCP itself never enforces maxTokens — server.mjs
does not read it — but it is propagated to OpenClaw (setup.mjs,
scripts/sync-openclaw.mjs), and OpenClaw CAPS the outbound request with it:
maxTokens = Math.min(baseMaxTokens + thinkingBudget, modelMaxTokens)
with reasoning budgets of medium 8192 / high 16384, and a clamp
`if (maxTokens <= thinkingBudget) thinkingBudget = maxTokens - 1024`.
So at `high` reasoning against a model declared at 16384: maxTokens resolves to
16384, the clamp fires, thinking takes 15360, and roughly 1024 tokens are left
for the visible answer. That is what this repo shipped for every model until
now, and it is the actual harm — not the advertised-vs-real mismatch.
ocp-connect keeps its own family-prefix table (/v1/models does not expose
maxTokens, so it cannot derive them). Set to the family FLOOR — opus 64000,
sonnet 32000 because sonnet-4-6 is 32000 while sonnet-5 is 64000, haiku 32000 —
because prefixes cannot distinguish versions and under-advertising merely caps
a client lower than the model allows, whereas over-advertising would promise
capacity a family member does not have. Commented in place.
New test asserts the PRINCIPLE, not the numbers: every maxTokens must exceed
OpenClaw's `high` thinking budget so an answer still fits. A value-pinning test
would need editing on every model addition and would not explain itself.
Mutation-proven: reverting one entry to 16384 fails with
"claude-opus-5 declares maxTokens=16384 <= 16384: at OpenClaw's 'high'
reasoning level the thinking budget would consume all but ~1024 tokens".
Owner-approved decision (option: align with registry defaults). Tradeoff
recorded in CHANGELOG: longer answers, correspondingly higher per-request quota
burn on long generations.
Verified: bash -n on ocp-connect, and py_compile on its embedded python block
(lines 101-334) so the comment sits at a valid indentation.
Tests: 458 passed, 0 failed (457 + 1).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8
* fix(models): rewrite the maxTokens rationale — my justification was factually wrong
The numbers were right, the code was right, and the reason I gave for both was
wrong. Verified every part of the review's finding first-hand before rewriting.
WHAT I CLAIMED: OpenClaw caps outbound requests via
`maxTokens = Math.min(baseMaxTokens + thinkingBudget, modelMaxTokens)`, so a
16384 declaration at `high` reasoning left ~1024 tokens for the visible answer,
and raising it would produce longer answers and higher quota burn.
WHY THAT CANNOT HAPPEN FOR OCP — three independent breaks, each confirmed here:
scripts/sync-openclaw.mjs:75 api: "openai-completions"
The clamp I cited lives behind `case "anthropic-messages"`; the
openai-completions transport has no thinkingBudget concept at all, so that code
never runs for a claude-local provider.
buildCliArgs pushes: --input-format, --disallowedTools, --allowedTools,
--dangerously-skip-permissions, --mcp-config
No output-token flag of any kind reaches the CLI. OCP cannot cap output even if
it wanted to.
grep max_completion_tokens server.mjs lib/ -> 0 occurrences
That is the field OpenClaw would send to a local openai-completions endpoint,
and this repo never reads it.
So "expect longer answers and higher per-request quota burn" was a promise that
would not have materialised. That is worse than a wrong number: a release note
that tells operators to expect a behavioural change which cannot occur.
The change still stands, on the honest reason: models.json is the SPOT (ADR
0003) and its values should be true. maxTokens is ADVERTISED metadata consumed
only by clients that choose to honour it. CHANGELOG and the test comment now
say exactly that, and say plainly that OCP behavior is unchanged.
Test replaced, not just re-worded. The old one asserted `maxTokens > 16384`,
which review showed lets every entry sit at 16385 and still call itself
"registry-aligned" — the actual claim had no coverage. Now pinned per model
against a recorded table of CLI 2.1.220 id-anchored values, with an explicit
failure for any model missing a row. Mutation-proven on both holes:
all entries at 16385 -> CAUGHT (the old test passed this)
sonnet-4-6 wrongly at 64000 -> CAUGHT
Also fixed ocp-connect's unknown-id fallback (8192 -> 32000). Review flagged it
as contradicting the invariant: 32000 is the LOWEST max_output_tokens.default
in the registry, so it is the safe floor and no longer sits below every family
in its own table. Unreachable today, wrong to leave inconsistent.
Verified: bash -n on ocp-connect and py_compile on its embedded python block.
Tests: 458 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8
* fix(ocp-connect): revert the unknown-id fallback to 8192 — 32000 over-advertised
The previous commit raised ocp-connect's unknown-id fallback from 8192 to 32000 on
the stated grounds that 32000 was "the LOWEST max_output_tokens.default in the CLI
2.1.220 registry". That is false, and the change it justified made this PR violate
its own thesis.
Pairing-free enumeration of the compiled registry (17 occurrences, no grep-window
artifacts):
7 max_output_tokens:{default:64000,upper:128000}
5 max_output_tokens:{default:32000,upper:64000}
2 max_output_tokens:{default:8192,upper:8192}
2 max_output_tokens:{default:32000,upper:32000}
1 max_output_tokens:{default:32000,upper:128000}
distinct defaults: 8192 / 32000 / 64000 GLOBAL MINIMUM: 8192
The floor is 8192, held by claude-3-5-haiku and claude-3-5-sonnet — and those are
exactly the ids that reach the fallback, since "claude-3-5-*" matches none of the
family prefixes. So the raise over-advertised the only two real models the branch
serves, by 4x, in a PR whose entire claim is that advertised metadata should be true.
Verified by exercising the real classifier (verbatim source slice, not a replica):
claude-3-5-haiku 8192 FALLBACK registry 8192 exact
claude-3-5-sonnet 8192 FALLBACK registry 8192 exact
claude-opus-5 64000 family registry 64000 exact
claude-sonnet-4-6 32000 family registry 32000 exact
claude-haiku-4-5-20251001 32000 family registry 32000 exact
claude-sonnet-5 32000 family registry 64000 under (safe)
Also in this commit:
- The family table's comment said "FAMILY FLOOR of the CLI registry". It is the floor
over each family's members CURRENTLY IN models.json, not over the registry family:
claude-opus-4-0 / -4-5 are opus-family at 32000, so the opus row would over-advertise
them 2x. Unreachable today (ocp-connect derives ids from /v1/models, i.e. models.json),
but the comment now states the real scope and what to do if a legacy opus is added.
- The test table keyed haiku on "claude-haiku-4-5-20251001" while telling the reader to
extract values id-anchored. That id has ZERO id-anchored hits: the registry record is
id:"claude-haiku-4-5" and the dated string appears only as its provider_ids.first_party
(10 bare-string hits). Following the instruction on that row returns nothing, which is
precisely what pushes a reader back to the bare-string search this test exists to
prevent. The row now carries the registry id.
Reviewer credit: the reviewer that found this also retracted its own earlier finding,
which is what produced the 32000 fallback in the first place. Both directions verified
independently here before acting.
server.mjs: unchanged by this commit (ALIGNMENT.md requires no cli.js citation).
Suite: 462 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8
* docs(ocp-connect): my unreachability reason was wrong — state the invariant that holds
Review of cb87b82 approved the change and flagged the new comment's reasoning. Verified
both entry points it named; both are real.
I wrote that claude-opus-4-0/-4-5 are unreachable because "ocp-connect derives ids from
/v1/models, i.e. models.json". Neither half is reliable:
- base_url is user-supplied (ocp-connect:510 builds it from arguments), so /v1/models is
the REMOTE server's models.json — a different OCP version, or a fork, not this checkout.
- on a JSON parse failure the ids come from a hardcoded three-id list (:114), not from
/v1/models at all. Checked all three: opus-4-6 -> 64000, sonnet-4-6 -> 32000,
haiku-4 -> 32000. All exact against the registry, so safe today — but it is a genuine
non-SPOT entry point that my sentence denied existed.
The invariant that actually holds is narrower: no released OCP has ever served a legacy
opus id. Same conclusion, and now it names the thing that would break it.
This is the second reasoning defect in this PR where the number was right and the
justification was not. Recording it in the comment rather than quietly fixing it.
Out of scope, noted for a future touch: the bare `except:` at :114 catches BaseException,
so a KeyboardInterrupt mid-parse silently yields the fallback list.
server.mjs: unchanged (ALIGNMENT.md requires no cli.js citation).
Suite: 462 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8
---------
Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
73 lines
1.8 KiB
JSON
73 lines
1.8 KiB
JSON
{
|
|
"$schema": "./models.schema.json",
|
|
"version": 1,
|
|
"models": [
|
|
{
|
|
"id": "claude-opus-5",
|
|
"displayName": "Claude Opus 5",
|
|
"openclawName": "Claude Opus 5 (via CLI)",
|
|
"reasoning": true,
|
|
"contextWindow": 200000,
|
|
"maxTokens": 64000
|
|
},
|
|
{
|
|
"id": "claude-opus-4-8",
|
|
"displayName": "Claude Opus 4.8",
|
|
"openclawName": "Claude Opus 4.8 (via CLI)",
|
|
"reasoning": true,
|
|
"contextWindow": 200000,
|
|
"maxTokens": 64000
|
|
},
|
|
{
|
|
"id": "claude-opus-4-7",
|
|
"displayName": "Claude Opus 4.7",
|
|
"openclawName": "Claude Opus 4.7 (via CLI)",
|
|
"reasoning": true,
|
|
"contextWindow": 200000,
|
|
"maxTokens": 64000
|
|
},
|
|
{
|
|
"id": "claude-opus-4-6",
|
|
"displayName": "Claude Opus 4.6",
|
|
"openclawName": "Claude Opus 4.6 (via CLI)",
|
|
"reasoning": true,
|
|
"contextWindow": 200000,
|
|
"maxTokens": 64000
|
|
},
|
|
{
|
|
"id": "claude-sonnet-5",
|
|
"displayName": "Claude Sonnet 5",
|
|
"openclawName": "Claude Sonnet 5 (via CLI)",
|
|
"reasoning": true,
|
|
"contextWindow": 200000,
|
|
"maxTokens": 64000
|
|
},
|
|
{
|
|
"id": "claude-sonnet-4-6",
|
|
"displayName": "Claude Sonnet 4.6",
|
|
"openclawName": "Claude Sonnet 4.6 (via CLI)",
|
|
"reasoning": true,
|
|
"contextWindow": 200000,
|
|
"maxTokens": 32000
|
|
},
|
|
{
|
|
"id": "claude-haiku-4-5-20251001",
|
|
"displayName": "Claude Haiku 4.5",
|
|
"openclawName": "Claude Haiku 4.5 (via CLI)",
|
|
"reasoning": false,
|
|
"contextWindow": 200000,
|
|
"maxTokens": 32000
|
|
}
|
|
],
|
|
"aliases": {
|
|
"opus": "claude-opus-5",
|
|
"sonnet": "claude-sonnet-5",
|
|
"haiku": "claude-haiku-4-5-20251001"
|
|
},
|
|
"legacyAliases": {
|
|
"claude-opus-4": "claude-opus-4-7",
|
|
"claude-haiku-4": "claude-haiku-4-5-20251001",
|
|
"claude-haiku-4-5": "claude-haiku-4-5-20251001"
|
|
}
|
|
}
|