cold-audit catch from 2026-05-24 (round 3)
Three coordinated governance amendments per the major-decisions already
made (per "继续吧 unless major decision" + my F5/F11/F13+F14 option-(b)
calls earlier in the session). All P3-class, all docs-only. Single
PR per IDR cleanup-batch convention given the unified theme
("spec was aspirational; reality is X; amend spec to match").
Changes (3 docs files, +52 / -2):
1. docs/adr/0003-intermediate-representation.md — Amendment 1 (NEW
Amendments section, first ever for ADR 0003):
- F5 closure: removes the unimplemented `__irRoundTripTest()`
export-name claim from § Mitigations. The mental model
(symmetric irToNative+nativeToIR pair on each plugin) does NOT
fit the actual plugin shape (asymmetric spawn-wrappers:
request→CLI-args+stdin, response-stream→IR-chunks). Zero plugins
export this function.
- Documents the substitute test strategy that IS live at v0.1:
· Suite 3 covers entry-surface IR↔OpenAI round-trip
· Per-plugin spawn-with-mock test blocks cover IR→CLI→IR-chunk
· Lossy-translation edges documented inline in each plugin's
header comment (mistral.mjs DOCS-N markers being the most
thorough; anthropic.mjs + codex.mjs carry equivalent tables)
- Future plugins MUST satisfy both (a) spawn-with-mock test block
and (b) header lossy-field documentation — the enforcement
mechanism replacing the original export-name aspiration.
- Original § Mitigations bullet replaced with a pointer to
Amendment 1 (preserves history).
2. docs/adr/0005-cache-cross-provider.md — Amendment 5 (after D27
Amendment 4):
- F11 closure: acknowledges that the "Anthropic's own prompt cache
is consulted at the provider" half of § D2 is structurally
impossible at v0.1.
- Root cause: Anthropic plugin uses `claude -p --output-format text`
(verified at anthropic.mjs:196), a plain-text wire surface that
has no Messages-API field for cache_control markers. The markers
are also stripped at IR construction (openai-to-ir.mjs:45-89
whitelists fields, drops cache_control) per ADR 0003 IR design
reaffirmed in D27 Amendment 4 (F10).
- Clarifies v0.1 effective behavior: OLP-cache-bypass works
correctly (no double-caching when Anthropic + markers); the
delegate-to-Anthropic-prompt-cache half does not fire.
- Forward path (v1.x): switch Anthropic plugin to --output-format
json (NDJSON wire) or direct Messages API spawn — substantial
parser rewrite; deferred since no user has requested delegation.
- § D2 paragraph stays intact; an inline parenthetical reference
to Amendment 5 is added immediately after the affected sentence.
3. ALIGNMENT.md — new "Rule 4 exception class — Speculative-Candidate
Plugin" sub-section after Rule 5, before § Authorities:
- F13+F14 closure: formalizes the gap between strict ALIGNMENT
Rule 2 (no speculative shape assumptions) + Rule 4 (unalignable
plugins are deleted, not feature-flagged) and the actual practice
where Codex (D6) + Mistral (D8) plugins ship with explicit
UNPINNED assumptions documented in their headers, gated as
Candidate-not-Enabled.
- 5 precise conditions for a Speculative-Candidate plugin:
(1) STATIC_REGISTRY entry present, (2) candidate: true in
models-registry, (3) NOT in any user's providers.enabled config,
(4) header has explicit UNPINNED assumption section with labeled
IDs and planned-pin triggers, (5) defensive multi-shape parsers
tied to UNPINNED markers (greppable for future single-shape
replacement).
- Enablement criteria (Speculative-Candidate → Enabled): 3 explicit
requirements (remove UNPINNED labels from header; replace
multi-shape parsers with single-shape; file pin ADR amendment).
- Currently-in-class table:
· codex.mjs (D6) — A3 (auth token field name), A4 (NDJSON event
schema)
· mistral.mjs (D8) — A4 (JSON output event schema), A5 (model
flag), A6 (exact model IDs), A7 (streaming vs json output
mode), A8 (stdin prompt passing)
- Anthropic explicitly NOT in this class (CLI version pinned at
@anthropic-ai/claude-code v2.1.89 per Provider Authority Pins
table).
- CI enforcement noted as deferred (reviewer obligation today; a
future PR may add grep-based gate).
Pre-commit fold-in (per evidence-first checkpoint #4):
- **D31 reviewer flagged wording precision** ("Provider Inventory
table above" was structurally wrong — the table is BELOW the new
sub-section). Folded in: 1-word fix "above" → "below" on the
Speculative-Candidate condition #2.
Tests: 400/400 unchanged (docs-only).
Authority:
- ADR 0003 § Mitigations (modified) + new Amendment 1 (self-amendment)
- ADR 0005 § D2 (annotated) + new Amendment 5 (self-amendment)
- ALIGNMENT.md § Rules (extended with Rule 4 exception class)
- D27 F10 Amendment 4 of ADR 0005 (cited for cache_control IR-strip
precedent)
- CC 开发铁律 v1.6 § 10.x — Round-3 Cold Audit caught all 3 items
Reviewer (Iron Rule v1.6 § 10.x Mode A, fresh-context opus, independent
of drafter): APPROVE_WITH_MINOR. Verified:
- Suite 3 + 3 plugin spawn-with-mock blocks + 3 plugin lossy-field
header tables all present (F5 substitute test strategy claims
match implementation)
- anthropic.mjs:196 hardcodes --output-format text + openai-to-ir.mjs
drops cache_control from messages (F11 wire-limitation technically
accurate)
- All 8 UNPINNED assumption labels (codex A3/A4 + mistral A4/A5/A6/A7/A8)
match plugin header text exactly — no invention
- Anthropic-not-in-class claim verified against Provider Authority Pins
table
Follow-up items (reviewer's non-blocking notes, NOT in this PR):
1. **mistral.mjs A5 internal disconnect** — header lines 145-156 label
A5 as `UNPINNED-D-later-verifies`, but function body at 371-374
already cites DeepWiki enumeration as the pin and marks A5
`CONFIRMED-NOT-APPLICABLE` (no --model flag exists in vibe CLI).
The disconnect predates D31 and is a per-plugin header cleanup
(not a constitutional amendment), so it stays out of D31 scope per
IDR. File as separate follow-up issue — A5 should be moved out of
the Speculative-Candidate table once the header is updated.
2. **ADR 0005 Amendment ordering cosmetic** — strict reverse-chronological
would put Amendment 5 above Amendment 4 (current order is 4, 5, 3,
2 because both 4 and 5 share the 2026-05-24 date and 5 was inserted
after 4). Not blocking; future cleanup if maintainers want strict
numeric-descending.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
OLP — Open LLM Proxy
A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing, automatic fallback, content-addressed caching — so your IDEs and family clients keep working as long as any of your subscriptions has quota left.
Status: v0.1 — bootstrap. Most of this README is a skeleton; sections marked placeholder land alongside the relevant phase of work (see phase plan).
Why OLP
On 2026-05-14, Anthropic announced (effective 2026-06-15) that claude -p, the Agent SDK, and third-party agent traffic move out of the Pro/Max subscription pool into a separate fixed monthly Agent SDK Credit pool. OCP, OLP's predecessor, was a proxy around a single CLI — its core assumption was "subscription = unlimited within rate limits". That assumption breaks for Anthropic on the effective date.
The structural response is to stop relying on one provider's subscription terms remaining favourable. OLP spreads risk across multiple providers whose subscriptions still include CLI/programmatic use, routes intelligently between them, and caches aggressively so every request that does spawn a CLI counts.
OLP is not: a commercial multi-tenant SaaS; an enterprise gateway competing with LiteLLM / OpenCode / CLIProxyAPI on breadth; a model-capability router ("route to the smartest model" — you pick the model); a conversation-state store (your client handles that).
See ALIGNMENT.md for OLP's constitution and docs/adr/ for the founding ADRs.
Quick Start
placeholder — lands with Phase 1.
Anticipated shape:
# install
npm install -g @dtzp555-max/olp
# run setup (writes ~/.olp/config.json, asks which providers to enable)
olp setup
# start the proxy (default port 3456 — same as OCP if you migrate)
olp start
# point your IDE at http://localhost:3456/v1/chat/completions with the OLP API key from `olp keys list`.
Supported Providers
Source of truth: models-registry.json. This table is regenerated from the registry per the release_kit overlay; do not edit it out of sync.
OLP distinguishes Candidate Providers (declared as intended, not yet pinned) from Enabled Providers (authority pin filled + plugin landed + Phase audit passed). The v0.1 founding commit ships zero Enabled Providers — enablement is a Phase audit deliverable, not a bootstrap claim. See ALIGNMENT.md § Provider Inventory for the transition gate.
Candidate Providers
| Provider key | CLI | Subscription / auth | Anticipated Tier | Anticipated Phase |
|---|---|---|---|---|
anthropic |
claude -p |
Pro / Max OAuth (pre-2026-06-15); Agent SDK Credit pool after | D (re-eval post-2026-06-15) | Phase 1 |
openai |
codex exec --json |
ChatGPT Pro OAuth or API key | D | Phase 2 |
mistral |
vibe --prompt --output json |
Le Chat Pro API key | D | Phase 3 |
grok |
grok -p --output-format streaming-json |
xAI Build xai-... API key |
C | Phase 8+ |
kimi |
kimi -p --output-format stream-json |
Moonshot Kimi API key | C | Phase 8+ |
minimax |
TBD | MiniMax Token Plan (¥29+/mo) | B | Phase 8+ |
glm |
TBD | Zhipu Coding Plan ($10+/mo) | B | Phase 8+ |
qwen |
TBD | Alibaba Coding Plan ($50/mo) | B | Phase 8+ |
Risk tier guide. D = permissive / safe (eligible for default-enabled); C = tightening signal, no enforcement history (opt-in); B = service-level key revocation risk (opt-in + consent); A = excluded by default (cannot be opt-in enabled). Tier B providers prompt for explicit consent on first enable and record consent in ~/.olp/config.json. See ALIGNMENT.md § Risk Tier Framework.
Excluded by default (Tier A — evidence-backed, pending primary-source pin). Google Antigravity. See ADR 0006 for the named-prohibition + no-cost-advantage + reinstatement-friction rationale, and for the primary-source pinning follow-up that may force a Tier reconsideration if the Google FAQ language cannot be sourced within 90 days of 2026-05-23.
Configuration
placeholder — full configuration reference lands with Phase 4 (fallback engine).
OLP reads its config from ~/.olp/config.json. The minimum useful shape:
{
"routing": {
"chains": {
"<requested-model>": [
{ "provider": "<key>", "model": "<provider-model-id>" },
{ "provider": "<key>", "model": "<provider-model-id>" }
]
},
"soft_triggers": {
"<provider-key>": { "<trigger>": <threshold> }
}
}
}
Note:
routing.soft_triggersthresholds are parsed and stored but have no runtime effect at v0.1 — the quota polling path (quotaStatus()per hop) is deferred to v1.x per ADR 0004 Amendment 2. The evaluation logic exists and is tested; only the production data ingestion path is deferred.
Trigger types, fallback safety, idempotency rules, and the full example config land here when Phase 4 ships. See ADR 0004 (Fallback Engine Semantics & Safety) for the design.
API Endpoints
placeholder — full table lands as each endpoint lands.
| Endpoint | Method | Phase | Description |
|---|---|---|---|
/v1/chat/completions |
POST | 1 | OpenAI-compatible Chat Completions entry. Internally normalized to IR, dispatched to a provider plugin, response shape converted back. |
/v1/models |
GET | 1 | Lists models from models-registry.json. |
/health |
GET | 1 | Per-provider health snapshot (owner-only). |
/cache/stats |
GET | 5 | Cache hit rate, by-provider breakdown. |
/v0/management/quota |
GET | 6 | Per-provider quota / credit pool status (best-effort). |
/dashboard |
GET | 6 | Owner-only dashboard (localhost-bound by default). |
Environment Variables
placeholder — full table lands per-phase as variables are introduced.
| Variable | Default | Description |
|---|---|---|
OLP_PORT |
3456 |
HTTP listener port. |
OLP_CLAUDE_BIN |
claude (from PATH) |
Override path to the claude binary (Anthropic provider). Useful when multiple claude installs are present. |
OLP_CODEX_BIN |
codex (from PATH) |
Override path to the codex binary (OpenAI provider). |
OLP_VIBE_BIN |
vibe (from PATH) |
Override path to the vibe binary (Mistral provider). |
📋 Planned (Phase 2) — not yet read by the codebase:
OLP_HOME(~/.olp) — Config, providers, keys, cache, logs root. Currently hardcoded to~/.olp/config.jsoninloadFallbackConfigSync; the env override path is a Phase 2 config-layer deliverable.OLP_LOG_LEVEL(info) — Log level filter (error/warn/info/debug).logEventcurrently writes unconditionally; level filtering is a Phase 2 observability deliverable.
Further variables (per-provider auth path overrides, cache size limits, fallback-engine knobs) land with the relevant phase. See also the Implementation status table.
Response Headers
Every response served through OLP carries:
X-OLP-Provider-Used: <provider-key>— which provider's plugin served the request.X-OLP-Model-Used: <model-id>— which model the served provider used.X-OLP-Fallback-Hops: <n>— number of fallback hops (0if served by the primary chain entry).X-OLP-Cache: hit | miss | bypass— cache layer outcome.X-OLP-Latency-Ms: <ms>— end-to-end latency observed at the proxy.
If a fallback chain is exhausted, X-OLP-Fallback-Exhausted lists the tried providers in order.
Implementation status (as of 2026-05-24)
Phase 1 is in progress. This table reflects what is currently shipped vs. what is designed for later phases.
| File / artifact | Status | Notes |
|---|---|---|
server.mjs |
✅ Shipped | HTTP listener + dispatcher |
lib/ir/ |
✅ Shipped | IR definition + serializers (ADR 0003) |
lib/providers/anthropic.mjs |
✅ Shipped | claude -p spawn-binary plugin |
lib/providers/codex.mjs |
✅ Shipped | codex exec --json plugin |
lib/providers/mistral.mjs |
✅ Shipped | vibe --prompt plugin |
lib/cache/keys.mjs |
✅ Shipped | Content-addressed key computation |
lib/cache/store.mjs |
✅ Shipped | In-memory Map (file-backed layout: 📋 Phase 2 storage adapter) |
lib/fallback/engine.mjs |
✅ Shipped | Trigger evaluation + chain advancement (ADR 0004) |
Soft trigger data path (quotaStatus() polling) |
📋 Planned (v1.x) | Evaluation logic shipped + tested; data ingestion deferred per ADR 0004 Amendment 2 |
models-registry.json |
✅ Shipped | SPOT for (provider, model) metadata |
test-features.mjs |
✅ Shipped | Comprehensive test suite covering IR, cache, fallback, and integration paths (CI: test.yml) |
lib/keys.mjs |
📋 Planned (Phase 2) | Multi-key auth, per-key namespacing, audit log |
dashboard.html |
📋 Planned (Phase 6) | Owner-only multi-provider dashboard |
docs/provider-caveats.md |
📋 Planned (Phase 3+) | Lossy-translation reference; for now documented inline in each plugin header |
docs/openai-spec-pin.md |
✅ Shipped (D30) | OpenAI spec snapshot for annual audit; v0.1 baseline pinned 2026-05-24 |
docs/alignment-audits/ |
📋 Planned | Output directory for annual alignment audits (first audit: 2027-05-14) |
scripts/migrate-from-ocp.mjs |
📋 Planned (Phase 7) | OCP → OLP migration tool |
setup.mjs |
📋 Planned | Setup wizard / initial config |
Architecture
OLP is a Node.js (ESM, .mjs) HTTP proxy with no build step and minimal dependencies. The high-level shape:
- Entry surface —
server.mjshandles/v1/chat/completionsand the administrative endpoints. Governed by OpenAI's/v1/chat/completionsspecification as the wire authority. SeeALIGNMENT.md§ Authorities. - Intermediate Representation (IR) —
lib/ir/normalizes between the entry surface and provider-native shapes. The IR is the lingua franca; any extension is an ADR 0003 amendment. - Provider plugins —
lib/providers/<name>.mjs. Each plugin implements the contract in ADR 0002 (Plugin Architecture for Providers), spawns its CLI, and translates between IR and provider-native IO. - Cache layer —
lib/cache/is a content-addressed cache keyed on the 11-field tuple defined in ADR 0005 § "Cache key composition (v1.0)" (provider, model, messages, tools, temperature, response_format, cache_control, max_tokens, top_p, stop, tool_choice — as of Amendment 2/D15). Per-key isolation, prompt-caching bypass, chunked stream replay, and singleflight. See ADR 0005 (Cache Layer Cross-Provider Design). Current backing store is an in-memory Map; file-backed storage is planned for Phase 2. - Fallback engine —
lib/fallback/advances a configured chain one provider at a time on configured triggers, never retrying after the first response chunk has been emitted to the client. See ADR 0004. - Multi-key auth —
lib/keys.mjs(📋 planned, Phase 2) will carry OCP's per-OLP-key namespace isolation forward. Each OLP API key will have independent quota, cache namespace, and audit log; each key declares which providers it can access.
Read the ADRs in docs/adr/ in order before proposing structural changes.
Phase plan
OLP lands in phases. Each phase has its own PR series and Iron-Rule-10 reviewer; this README's placeholders are filled per-phase via the release_kit overlay.
- Phase 0 — Repo bootstrap,
ALIGNMENT.md, founding ADRs, CI workflows, PR template. (current) - Phase 1 —
server.mjsskeleton, IR, Anthropic plugin, cache D1+D4. Port from OCP. - Phase 2 — OpenAI Codex plugin.
- Phase 3 — Mistral Vibe plugin.
- Phase 4 — Fallback engine + routing chains config + quota poll worker.
- Phase 5 — Cache cross-provider hardening (D2+D3).
- Phase 6 — Dashboard + observability (
/v0/management/quota). - Phase 7 — Release v0.1, OCP enters maintenance.
- Phase 8+ — Optional Grok / Kimi / tier-2 plugins; provider-native protocol endpoints; deterministic triggers.
Full spec (decision rationale, open questions, risks): ~/.cc-rules/memory/projects/olp_v0_1_spec.md on the maintainer's workstations.
Migration from OCP
placeholder — scripts/migrate-from-ocp.mjs lands with Phase 7 (📋 planned, not yet authored).
Anticipated user-facing flow (target: <5 minutes):
- Stop OCP (
launchctl bootoutthe OCP service orocp stop). - Install OLP.
- Run
olp migrate-from-ocp— copies~/.ocp/keys/to~/.olp/keys/and points provider plugins at OCP's existing auth artifacts where applicable. - Start OLP. Clients pointing at port 3456 keep working; their existing OLP API keys remain valid.
OCP's cache directory is not migrated: OLP's cache key format includes provider+model and warms cold naturally. OCP enters maintenance mode (stability fixes only) when OLP v0.1 ships; new development happens in OLP.
License
MIT.
Acknowledgements
OLP evolved from OCP (Open Claude Proxy). OCP's per-key isolation model, cache-layer design (D1–D4), dashboard, and alignment-constitution discipline are all carried forward. The structural generalization from single-CLI to multi-provider is what makes this a new project rather than an OCP minor version — see ALIGNMENT.md § Reference: How OCP's cli.js discipline maps to OLP.
Authors: project maintainer (with AI drafting assistance).