2 Commits
Author SHA1 Message Date
a17b6c0ad2 chore(release): v3.26.0 (#212)
* chore(release): v3.26.0 — maxTokens tells the truth, SPOT schema, release-job fix, ltBoot harness

Version 3.25.0 -> 3.26.0 (package.json; server.mjs reads VERSION from it, so /health and
/v1/models follow automatically). CHANGELOG's Unreleased section becomes v3.26.0.

Contents — five merged PRs, each with an independent fresh-context reviewer:

  #208  maxTokens aligned to the CLI registry (#195)      <- the only fleet-visible change
  #205  models.schema.json + CI validation (#196)
  #206  release.yml no-CHANGELOG path fix (#202)
  #204  ltBoot harness hardening + diagnostics (#199, #209)
  #207  cache-key guard comments (#200), AGENTS.md harness docs (#197)

The only user-visible change is advertised metadata: OpenClaw and other clients that read
maxTokens will see 64000/32000 instead of a uniform 16384. OCP's behavior is unchanged --
buildCliArgs passes no output-token flag, and max_completion_tokens appears nowhere in this
repo. No new endpoint, env var, CLI subcommand, or cli.js wire behavior. server.mjs is
untouched across the whole release, so ALIGNMENT.md requires no cli.js citation.

release_kit (Iron Rule 5.5) walked:
  - version_source package.json          bumped
  - changelog CHANGELOG.md               v3.26.0 section written
  - new file / SPOT / schema             models.schema.json documented in README
                                         (Available Models) and AGENTS.md (Key files)
  - Available Models table               no change needed; it carries no maxTokens column
  - Environment Variables / Endpoints    no additions this release
  - bootstrap_quirk_policy               no new one-time migration quirk

Remaining v3.25.0 mentions in README/docs are historical (describing what changed in that
release) and are correct as written; verified rather than bulk-replaced.

Suite: 462 passed, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8

* docs(release): correct a false claim in the v3.26.0 notes and disclose three gaps

Release review found four problems, all in the CHANGELOG text — which for a release PR
IS the deliverable, since release.yml copies it verbatim into the GitHub Release body.

1. "instead of a uniform 16384" was FALSE.

    git show v3.25.0:models.json  ->  six entries 16384, claude-haiku-4-5-20251001 8192
    distinct prior values: [8192, 16384]     uniform? False

Issue #195's own title says "for every Opus/Sonnet entry" — it never claimed haiku. The
framing also understated haiku, which moved from a lower base. Corrected, and stated: every
model except claude-sonnet-4-6 (1.95x) moves by the same 3.91x, haiku included.

(Review called haiku "the largest change in the release". Verified: it is a TIE at 3.91x,
not the largest alone — 16384->64000 is the same multiplier. The entry says so.)

2. models.schema.json contradicted this release's own notes.

It shipped in this release (#205) saying maxTokens "bounds request and compaction budgets".
The CHANGELOG correctly says maxTokens does not affect behavior. The schema's compaction
half is wrong: in OpenClaw 2026.7.1 both the compaction trigger and the summarisation chunk
size derive from contextWindow (minus reserveTokensFloor / softThreshold), never from
maxTokens. Fixed the schema rather than the CHANGELOG, since the CHANGELOG was right.

This makes the release PR touch one non-version file. Deliberate: shipping two canonical
documents that contradict each other is worse than a slightly wider release diff, and the
schema is published in this same release.

3. #203 was omitted, and the ltBoot entry read as "harness flakiness fixed".

42/200 -> 200/200 is true and #203 is still OPEN — with four sightings, one of them on #205
INSIDE this release. The entry made exactly this disclosure for ocp-connect (#210) and not
for the bug the harness was built to chase. Now states it is not fixed, and points at #211.

4. contextWindow is wrong by the same registry this release cites.

    claude-opus-5 / -4-8 / -4-7 / claude-sonnet-5   registry window:1e6   models.json 200000

Publishing "the SPOT tells the truth" with a 5x understatement in the adjacent field is the
kind of thing that ages badly. Unlike maxTokens it is NOT a safe metadata fix —
derivePromptCharBudget takes max(contextWindow) x 3 across ALL entries, so one 1M model
takes MAX_PROMPT_CHARS from 600k to 3M for every model including a genuinely-200k haiku.
Filed as #213 (needs per-model budgets, ADR-level) and disclosed in the notes.

Also: `claude-haiku-4-5` appeared where the models.json id is `claude-haiku-4-5-20251001` —
the exact legacyAlias-vs-id confusion #208's test file exists to prevent. All three
occurrences now use the full id.

Verified: models.schema.json still parses; suite 462 passed, 0 failed; release.yml's awk
extraction on the corrected file yields 6348 bytes, correct section, terminating at v3.25.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8

* docs(release): state the correction as fact, not as edit history

Review's style point, taken. The v3.26.0 section has never been published — no tag
exists — so release.yml's copy into the GitHub Release body is the FIRST time any
reader sees this text. "as an earlier draft of this entry claimed" corrects a claim
they never saw; it is internal editing history leaking into a public artifact. The
correction belongs in the commit record, where it already is in full.

Also split the 3.91x ratio into its own sentence — as an appositive it had to reach
back across an interrupting clause to find what it modified.

No factual change. Suite: 462 passed, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:28:18 +10:00
f2f9058db4 feat(models): commit models.schema.json and validate the SPOT against it in CI (#196) (#205)
* feat(models): commit models.schema.json and validate the SPOT against it in CI (#196)

models.json has declared `"$schema": "./models.schema.json"` since v3.11.0, but
that file was never committed — `git log --all -- models.schema.json` is empty.
So the SPOT that ADR 0003 makes canonical had NO structural validation: a
missing contextWindow, a typo'd openclawName, or a boolean written as a string
would pass every existing test and only surface downstream — in OpenClaw's
registry, or silently inside a truncation budget.

Validated with the repo's OWN validator (lib/structured-output.mjs, shipped for
#153) rather than adding ajv. AGENTS.md requires minimal dependencies and the
repo currently has ZERO; this keeps it that way and exercises that validator on
a second real input. Capability-probed first rather than assumed: it enforces
type / required / enum / const / additionalProperties (both the boolean and the
schema form) / items / minItems, which covers everything the SPOT needs.

Three tests:
  - the $schema reference resolves to a committed file (the #196 bug itself)
  - models.json validates against models.schema.json, strict
  - the schema actually REJECTS 8 corruption shapes — missing required field,
    wrong scalar type, non-integer window, unknown extra field, alias mapped to
    a non-string, empty models array, wrong document version, unknown
    top-level key

That third test is the guard-on-the-guard: without it a vacuous schema (`{}` or
a typo'd `properties`) would make the second test pass on anything.

Mutation-proven in three directions:
  schema replaced with {}              -> "schema failed to reject: missing required field"
  additionalProperties:false removed   -> guard fails (loosened constraint caught)
  models.json corrupted (drop a field) -> strict validation fails

What the schema deliberately does NOT encode: referential integrity — that
every aliases/legacyAliases target exists as a models[].id. That is not a JSON
Schema concept; it is already covered by two dedicated tests, and the schema
says so in its description so the next reader doesn't assume it is covered.

The field descriptions carry the non-obvious consequences that have bitten this
repo before: contextWindow is a GLOBAL lever (max across all entries x 3 sets
MAX_PROMPT_CHARS for every model, ADR 0009), maxTokens is advertising only and
is enforced by OpenClaw rather than OCP, and models[] order is load-bearing
because ocp-connect uses model_ids[0] as a fallback.

Documented per release_kit (new file/SPOT/schema -> contributor section):
AGENTS.md "Key files to know" and README's Available Models section.

Tests: 460 passed, 0 failed (457 + 3).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8

* fix(schema): warn that only a subset of draft 2020-12 is enforced; cover the gaps it cannot

Both review findings taken.

MED — the schema declared `$schema: draft/2020-12`, which promises full draft
semantics, while the validator that actually runs it enforces only a subset.
Someone adding `"minimum": 1` to contextWindow would get SILENT no-op with a
green suite — a trap made worse by the file looking authoritative. The
description now enumerates exactly what is enforced (type, required, const,
enum, additionalProperties in both forms, items, minItems/maxItems,
anyOf/allOf/oneOf, $ref, nullable) and names what is ignored (minimum,
maxLength, pattern, uniqueItems, propertyNames, not), with the instruction to
add such a constraint as a test instead.

LOW — three corruptions neither the schema nor any sibling test caught:
duplicate models[].id, empty-string id/displayName/openclawName, and
non-positive contextWindow/maxTokens. They are structurally NOT expressible
with the supported keyword set (no uniqueItems, no minLength, no minimum), so
they are asserted directly rather than by pretending the schema covers them.

Mutation-proven, all three previously slipped through:
  duplicate id       -> CAUGHT
  empty displayName  -> CAUGHT
  zero contextWindow -> CAUGHT

Tests: 461 passed, 0 failed (460 + 1).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8

* test(schema): close the fourth gap — uniqueness on all three name fields, not just id

Review approved the delta and named this exact follow-up as non-blocking:
"extend the same test to assert uniqueness of displayName/openclawName and use
m[f] === m[f].trim()". Implemented as specified.

Why it matters rather than being tidiness: scripts/sync-openclaw.mjs maps
`claude-local/<id>` -> { alias: displayName } and writes openclawName as the
registry label, so a duplicate in EITHER field collapses two models onto one
OpenClaw entry — the same defect class as a duplicate id, which the previous
version covered while leaving its two siblings open.

The trim check also changed shape. It was `m[f].trim().length > 0`, which a
PADDED id passes by construction — and that id is handed VERBATIM to
`claude --model`, so " claude-opus-5" would have failed upstream instead of
here. Now asserts `m[f] === m[f].trim()`.

Mutation-proven, all three previously slipped through:
  duplicate displayName  -> CAUGHT
  duplicate openclawName -> CAUGHT
  whitespace-padded id   -> CAUGHT

Tests: 461 passed, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017gbqUZ8HfBZpjjbzQ85oH8

---------

Co-authored-by: dtzp555 <dtzp555@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 12:04:53 +10:00