taodengandClaude Opus 4.7 bafa6d1991 fix(fallback)+docs(adr-0004): D16 — honor "usable chunks streamed" qualifier on SPAWN_FAILED
cold-audit catch from 2026-05-23

Cold-audit Finding 17 (P2 fallback correctness). ADR 0004 § Trigger
taxonomy Hard triggers bullet 3 says "Provider CLI exit code ≠ 0
**with no usable response chunks streamed**" — the qualifier was
not honored. Pre-D16 code in `collectAllChunks` re-raised SPAWN_FAILED
unconditionally, discarding any partial chunks. Concrete failure:
provider yields 1000 chars of completion then exits non-zero (e.g.,
post-stream cleanup error) → chunks dropped, fallback to next provider,
user pays double spawn cost and loses the original provider's output.

Coordinated change across two layers, single commit per the
ADR-with-code pattern (D11 / D15 precedent):

1. docs/adr/0004-fallback-engine.md — Amendment 1 (top of doc, matching
   D11/D15 placement convention):
   - Documents the "usable chunks streamed" semantics precisely
   - Behavior split: chunks.length > 0 + SPAWN_FAILED → synthesize stop
     + return (Case B); chunks.length === 0 + SPAWN_FAILED → re-throw
     (Case A, hard trigger fires as before)
   - finish_reason='length' rationale (4 reasons documented)
   - Streaming-path note: ADR 0004 first-chunk rule already handles the
     analogous case for D10's real-streaming branch; D16 applies to
     buffered path only
   - Cache behavior: write-through `getOrCompute` (preserves D4 singleflight
     during truncation event) then evict via `set(ttlMs=0)` (so future
     fresh callers re-spawn). `__truncated` non-enumerable marker
     travels with the chunks array for follower visibility

2. server.mjs `collectAllChunks` salvage path:
   - try/catch around the for-await loop
   - On SPAWN_FAILED with chunks.length > 0: synthesize stop chunk
     `{type:'stop', finish_reason:'length'}`, log warn event
     `spawn_failed_after_usable_chunks`, mark chunks array with
     non-enumerable `__truncated`, return (no re-throw)
   - On SPAWN_FAILED with chunks.length === 0 OR any other error:
     re-throw (preserves existing hard-trigger semantics)

3. server.mjs `executeHopFn` cache eviction:
   - After `cacheStore.getOrCompute(...)` returns, check `result.__truncated`
   - If truncated: `cacheStore.set(keyId, hopCacheKey, result, 0)` —
     ttlMs=0 causes `_isAlive` to treat the entry as expired on next
     read (verified in lib/cache/store.mjs)

4. test-features.mjs Suite 13 — 3 new tests:
   - Case A regression: SPAWN_FAILED at iter 0 + 2-hop chain → openai
     serves, X-OLP-Fallback-Hops: 1 (no behavior change)
   - Case B 2-hop: 2 deltas + SPAWN_FAILED → anthropic serves with
     synthesized stop, hops=0, finish_reason='length', content
     concatenates, openai NOT called
   - Case B single-hop: 1 delta + SPAWN_FAILED → HTTP 200 (not 502),
     finish_reason='length', partial content visible

Tests: 297 → 300 (+3). All pass on Node 20.

Pre-commit fold-ins (per evidence-first checkpoint #4 — fold-ins
themselves need second-pass review):

- **Error-chunk-in-chunks fold-in (sonnet flagged)**: pre-D12 code
  pushed error chunks BEFORE throwing. Post-D16's `chunks.length > 0`
  check would incorrectly include an error chunk and trigger Case B
  for a path that's actually Case A. Moved the `type === 'error'`
  check BEFORE the push, restoring the invariant that the chunks
  array contains only delta/stop chunks. Verified all 3 scenarios:
  (1) error at iter 0 → throws before push → length=0 → Case A
  (2) delta×2 + error at iter 3 → throws before push → length=2 → Case B
      with delta×2 + synthesized stop (no error chunk leaks)
  (3) delta + non-zero exit from outside loop → length=1 → Case B

- **ADR doc-code drift fold-in (D16 reviewer flagged)**: original
  Amendment 1 text said the salvaged result "bypasses
  cacheStore.getOrCompute and is returned directly, exactly as the
  cache-bypass path does." This was factually wrong — the code
  write-throughs via getOrCompute then evicts via ttlMs=0. The drift
  was ironic: D16 was about removing doc-code drift in ADR 0004
  bullet 3 itself, and the amendment was about to ship fresh drift.
  Corrected to accurately describe the write-then-evict pattern and
  the rationale (preserving D4 singleflight during truncation events).

Authority:
- ADR 0004 § Trigger taxonomy Hard triggers bullet 3 (the qualifier
  this amendment makes load-bearing)
  https://github.com/dtzp555-max/olp/blob/main/docs/adr/0004-fallback-engine.md
- ADR 0005 § Cache write conditions item 1 — "response completed
  successfully (no truncation, no error mid-stream)"
- OpenAI Chat Completions finish_reason enum (stop|length|tool_calls|
  content_filter|function_call|null)
  https://platform.openai.com/docs/api-reference/chat/object
- ADR 0004 § Fallback safety — first-chunk rule (already governs the
  analogous case in the real-streaming path)
- ALIGNMENT.md Rule 2(c) spirit — ADR amendment + code change land
  in same merge (D11 / D15 precedent)
- CC 开发铁律 v1.6 § 10.x — Cold Audit caught this; diff-review
  passes focused on first-chunk rule for streaming missed the
  buffered path's truncation-vs-fallback decision point

Reviewer (Iron Rule v1.6 § 10.x Mode A, fresh-context opus, independent
of drafter): APPROVE_WITH_MINOR. Folded the ADR doc-code drift minor
before commit. Walked all 3 error-chunk scenarios against actual code
to verify the pre-commit fold-in is correct. Analyzed the eviction
race window (sub-ms post-inflight pre-eviction window where a fresh
caller could hit the truncated cache before eviction lands) and
concluded it's structurally bounded — one-shot leak per truncation
event; subsequent callers re-spawn. Acceptable as v0.1.

Follow-up items (reviewer's non-blocking suggestions, NOT in this PR):
- 4th test asserting second identical request triggers fresh spawn
  (defense-in-depth around the eviction; store.mjs ttlMs=0 semantics
  are independently established)
- `cacheStore.delete()` API (cleaner than set-with-ttlMs=0 — leaves
  no dead entry in the namespace map; future PR)
- `cache_evicted_truncated` log event for dashboard observability
- SPAWN_TIMEOUT salvage parity — same architectural argument as
  SPAWN_FAILED (user paid for partial content); deferred as a
  separate cold-audit finding for a future D-stage

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-24 11:40:52 +10:00

OLP — Open LLM Proxy

A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing, automatic fallback, content-addressed caching — so your IDEs and family clients keep working as long as any of your subscriptions has quota left.

Status: v0.1 — bootstrap. Most of this README is a skeleton; sections marked placeholder land alongside the relevant phase of work (see phase plan).


Why OLP

On 2026-05-14, Anthropic announced (effective 2026-06-15) that claude -p, the Agent SDK, and third-party agent traffic move out of the Pro/Max subscription pool into a separate fixed monthly Agent SDK Credit pool. OCP, OLP's predecessor, was a proxy around a single CLI — its core assumption was "subscription = unlimited within rate limits". That assumption breaks for Anthropic on the effective date.

The structural response is to stop relying on one provider's subscription terms remaining favourable. OLP spreads risk across multiple providers whose subscriptions still include CLI/programmatic use, routes intelligently between them, and caches aggressively so every request that does spawn a CLI counts.

OLP is not: a commercial multi-tenant SaaS; an enterprise gateway competing with LiteLLM / OpenCode / CLIProxyAPI on breadth; a model-capability router ("route to the smartest model" — you pick the model); a conversation-state store (your client handles that).

See ALIGNMENT.md for OLP's constitution and docs/adr/ for the founding ADRs.


Quick Start

placeholder — lands with Phase 1.

Anticipated shape:

# install
npm install -g @dtzp555-max/olp

# run setup (writes ~/.olp/config.json, asks which providers to enable)
olp setup

# start the proxy (default port 3456 — same as OCP if you migrate)
olp start

# point your IDE at http://localhost:3456/v1/chat/completions with the OLP API key from `olp keys list`.

Supported Providers

Source of truth: models-registry.json. This table is regenerated from the registry per the release_kit overlay; do not edit it out of sync.

OLP distinguishes Candidate Providers (declared as intended, not yet pinned) from Enabled Providers (authority pin filled + plugin landed + Phase audit passed). The v0.1 founding commit ships zero Enabled Providers — enablement is a Phase audit deliverable, not a bootstrap claim. See ALIGNMENT.md § Provider Inventory for the transition gate.

Candidate Providers

Provider key CLI Subscription / auth Anticipated Tier Anticipated Phase
anthropic claude -p Pro / Max OAuth (pre-2026-06-15); Agent SDK Credit pool after D (re-eval post-2026-06-15) Phase 1
openai codex exec --json ChatGPT Pro OAuth or API key D Phase 2
mistral vibe --prompt --output json Le Chat Pro API key D Phase 3
grok grok -p --output-format streaming-json xAI Build xai-... API key C Phase 8+
kimi kimi -p --output-format stream-json Moonshot Kimi API key C Phase 8+
minimax TBD MiniMax Token Plan (¥29+/mo) B Phase 8+
glm TBD Zhipu Coding Plan ($10+/mo) B Phase 8+
qwen TBD Alibaba Coding Plan ($50/mo) B Phase 8+

Risk tier guide. D = permissive / safe (eligible for default-enabled); C = tightening signal, no enforcement history (opt-in); B = service-level key revocation risk (opt-in + consent); A = excluded by default (cannot be opt-in enabled). Tier B providers prompt for explicit consent on first enable and record consent in ~/.olp/config.json. See ALIGNMENT.md § Risk Tier Framework.

Excluded by default (Tier A — evidence-backed, pending primary-source pin). Google Antigravity. See ADR 0006 for the named-prohibition + no-cost-advantage + reinstatement-friction rationale, and for the primary-source pinning follow-up that may force a Tier reconsideration if the Google FAQ language cannot be sourced within 90 days of 2026-05-23.


Configuration

placeholder — full configuration reference lands with Phase 4 (fallback engine).

OLP reads its config from ~/.olp/config.json. The minimum useful shape:

{
  "routing": {
    "chains": {
      "<requested-model>": [
        { "provider": "<key>", "model": "<provider-model-id>" },
        { "provider": "<key>", "model": "<provider-model-id>" }
      ]
    },
    "soft_triggers": {
      "<provider-key>": { "<trigger>": <threshold> }
    }
  }
}

Trigger types, fallback safety, idempotency rules, and the full example config land here when Phase 4 ships. See ADR 0004 (Fallback Engine Semantics & Safety) for the design.


API Endpoints

placeholder — full table lands as each endpoint lands.

Endpoint Method Phase Description
/v1/chat/completions POST 1 OpenAI-compatible Chat Completions entry. Internally normalized to IR, dispatched to a provider plugin, response shape converted back.
/v1/models GET 1 Lists models from models-registry.json.
/health GET 1 Per-provider health snapshot (owner-only).
/cache/stats GET 5 Cache hit rate, by-provider breakdown.
/v0/management/quota GET 6 Per-provider quota / credit pool status (best-effort).
/dashboard GET 6 Owner-only dashboard (localhost-bound by default).

Environment Variables

placeholder — full table lands per-phase as variables are introduced.

Variable Default Description
OLP_PORT 3456 HTTP listener port.
OLP_HOME ~/.olp Config, providers, keys, cache, logs root.
OLP_LOG_LEVEL info One of error, warn, info, debug.

Further variables (per-provider auth path overrides, cache size limits, fallback-engine knobs) land with the relevant phase.


Response Headers

Every response served through OLP carries:

  • X-OLP-Provider-Used: <provider-key> — which provider's plugin served the request.
  • X-OLP-Model-Used: <model-id> — which model the served provider used.
  • X-OLP-Fallback-Hops: <n> — number of fallback hops (0 if served by the primary chain entry).
  • X-OLP-Cache: hit | miss | bypass — cache layer outcome.
  • X-OLP-Latency-Ms: <ms> — end-to-end latency observed at the proxy.

If a fallback chain is exhausted, X-OLP-Fallback-Exhausted lists the tried providers in order.


Architecture

OLP is a Node.js (ESM, .mjs) HTTP proxy with no build step and minimal dependencies. The high-level shape:

  • Entry surfaceserver.mjs handles /v1/chat/completions and the administrative endpoints. Governed by OpenAI's /v1/chat/completions specification as the wire authority. See ALIGNMENT.md § Authorities.
  • Intermediate Representation (IR)lib/ir/ normalizes between the entry surface and provider-native shapes. The IR is the lingua franca; any extension is an ADR 0003 amendment.
  • Provider pluginslib/providers/<name>.mjs. Each plugin implements the contract in ADR 0002 (Plugin Architecture for Providers), spawns its CLI, and translates between IR and provider-native IO.
  • Cache layerlib/cache/ is a content-addressed cache keyed on (provider, model, messages, tools, temperature, response_format, cache_control). Per-key isolation, prompt-caching bypass, chunked stream replay, and singleflight. See ADR 0005 (Cache Layer Cross-Provider Design).
  • Fallback enginelib/fallback/ advances a configured chain one provider at a time on configured triggers, never retrying after the first response chunk has been emitted to the client. See ADR 0004.
  • Multi-key authlib/keys.mjs carries OCP's per-OLP-key namespace isolation forward. Each OLP API key has independent quota, cache namespace, and audit log; each key declares which providers it can access.

Read the ADRs in docs/adr/ in order before proposing structural changes.


Phase plan

OLP lands in phases. Each phase has its own PR series and Iron-Rule-10 reviewer; this README's placeholders are filled per-phase via the release_kit overlay.

  • Phase 0 — Repo bootstrap, ALIGNMENT.md, founding ADRs, CI workflows, PR template. (current)
  • Phase 1 — server.mjs skeleton, IR, Anthropic plugin, cache D1+D4. Port from OCP.
  • Phase 2 — OpenAI Codex plugin.
  • Phase 3 — Mistral Vibe plugin.
  • Phase 4 — Fallback engine + routing chains config + quota poll worker.
  • Phase 5 — Cache cross-provider hardening (D2+D3).
  • Phase 6 — Dashboard + observability (/v0/management/quota).
  • Phase 7 — Release v0.1, OCP enters maintenance.
  • Phase 8+ — Optional Grok / Kimi / tier-2 plugins; provider-native protocol endpoints; deterministic triggers.

Full spec (decision rationale, open questions, risks): ~/.cc-rules/memory/projects/olp_v0_1_spec.md on the maintainer's workstations.


Migration from OCP

placeholder — scripts/migrate-from-ocp.mjs lands with Phase 7.

Anticipated user-facing flow (target: <5 minutes):

  1. Stop OCP (launchctl bootout the OCP service or ocp stop).
  2. Install OLP.
  3. Run olp migrate-from-ocp — copies ~/.ocp/keys/ to ~/.olp/keys/ and points provider plugins at OCP's existing auth artifacts where applicable.
  4. Start OLP. Clients pointing at port 3456 keep working; their existing OLP API keys remain valid.

OCP's cache directory is not migrated: OLP's cache key format includes provider+model and warms cold naturally. OCP enters maintenance mode (stability fixes only) when OLP v0.1 ships; new development happens in OLP.


License

MIT.


Acknowledgements

OLP evolved from OCP (Open Claude Proxy). OCP's per-key isolation model, cache-layer design (D1D4), dashboard, and alignment-constitution discipline are all carried forward. The structural generalization from single-CLI to multi-provider is what makes this a new project rather than an OCP minor version — see ALIGNMENT.md § Reference: How OCP's cli.js discipline maps to OLP.

Authors: project maintainer (with AI drafting assistance).

S
Description
No description provided
Readme MIT
2.4 MiB
Languages
JavaScript 95.8%
HTML 2.3%
Shell 1.9%