mirror of
https://github.com/dtzp555-max/olp.git
synced 2026-07-21 21:15:10 +00:00
Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
74e67fdca3 | ||
|
|
43ca4a65b0 | ||
|
|
7019294c63 | ||
|
|
ffe81f7a45 | ||
|
|
d67ba3d675 | ||
|
|
3551921f55 | ||
|
|
b1e24b7cb0 | ||
|
|
497b2550e6 | ||
|
|
28642756b5 | ||
|
|
d0dcd281ef | ||
|
|
07d9c8a6ae | ||
|
|
e5cfc696da | ||
|
|
dbac5f5521 | ||
|
|
dd0c821272 | ||
|
|
65f945c16d | ||
|
|
97e7d16585 | ||
|
|
40f9453d88 | ||
|
|
e2f41eb60e | ||
|
|
5d60a0599f | ||
|
|
ea0392f744 | ||
|
|
cc250e71bf | ||
|
|
9dc070bc53 | ||
|
|
fa2d1af130 | ||
|
|
bddf2cba1e | ||
|
|
65681ed7d2 | ||
|
|
d872330c9e | ||
|
|
2b07a3bd1b | ||
|
|
a41420d0fc | ||
|
|
5288493f19 | ||
|
|
82d2e1cbea | ||
|
|
187e79321f | ||
|
|
1605400052 | ||
|
|
704d4fc8a0 | ||
|
|
6605b7b14a | ||
|
|
6edf6e0b94 | ||
|
|
f3716a19fd | ||
|
|
ee4d9459aa | ||
|
|
53afea47ca | ||
|
|
0bdecd1235 | ||
|
|
e69e908dae | ||
|
|
e6701ff698 | ||
|
|
0048481764 | ||
|
|
ba69a3c13b | ||
|
|
679e3b367d | ||
|
|
9b66326e72 | ||
|
|
1062e88e77 |
@@ -49,7 +49,7 @@ Runtime: Node.js (ESM, `.mjs` throughout). No build step. No bundler. `server.mj
|
||||
- `.github/workflows/alignment.yml` — CI blacklist grep + per-provider citation soft check; fails the build on known-hallucinated tokens.
|
||||
- `CLAUDE.md` — Claude-Code-specific session instructions + `release_kit` overlay (Iron Rule 5.5).
|
||||
|
||||
**Implementation status note (as of 2026-05-25):** Files marked 📋 above are designed and documented but not yet on disk; files marked 🟡 are partially shipped; files marked ✅ are Phase 2 deliverables. The shipped set as of D47 is: `server.mjs` (with Phase 2 auth middleware + audit wire + owner-vs-non-owner gating), `lib/ir/`, `lib/providers/{anthropic,codex,mistral}.mjs`, `lib/cache/{keys,store}.mjs`, `lib/fallback/engine.mjs`, `lib/keys.mjs` (core + loadAuthConfigSync — D44 + D45), `lib/audit.mjs` (D45), `bin/olp-keys.mjs` (D47), `models-registry.json`, `test-features.mjs` (Suites 19–22). Phase 2 functional scope is complete; remaining is Phase 2 close → v0.2.0 (maintainer-triggered, explicit per CLAUDE.md `release_kit.phase_close_trigger`).
|
||||
**Implementation status note (as of 2026-05-27):** Phase 5 (Quota Probes + Dashboard Enrichment) is closed at v0.5.0 + v0.5.1 hotfix. The shipped set includes all Phases 1–5 deliverables. v0.5.1 hotfix (2026-05-27) fixes three codex review findings: F1 (doctor check bypassed backoff by calling `_probeOnce` directly — now routes through `quotaStatus()`), F2 (200 with empty `anthropic-ratelimit-*` headers was cached as live — minimum-viable-schema gate added), F3 (null collapsed all failure modes — `probe_status:'unreachable'` shape + `failure`/`failure_kind`/`backoff_until` fields added). See ADR 0008 Amendment 2 + ADR 0013 Rule 3/5 clarifications. Phase 6 is next (per CLAUDE.md `release_kit.current_phase`).
|
||||
|
||||
---
|
||||
|
||||
|
||||
+12
-2
@@ -196,9 +196,19 @@ In addition to the recurring 14 May audit below, the following one-shot audits a
|
||||
|
||||
## Class-specific Exceptions
|
||||
|
||||
(none at project founding)
|
||||
Any Rule 2 or Rule 3 deviation lands here as a numbered exception with PR link, reviewer, and rationale.
|
||||
|
||||
Any future Rule 3 deviation lands here as a numbered exception with PR link, reviewer, and rationale.
|
||||
### 1. Anthropic plan-usage probe via direct `/v1/messages` call (Phase 5, D79 — 2026-05-26)
|
||||
|
||||
**Class:** Rule 2(a) — provider-plugin scope. The Anthropic plugin's `quotaStatus()` calls `POST https://api.anthropic.com/v1/messages` directly rather than spawning `claude -p`. Under the strict reading of Rule 2(a), plugins must mirror provider-CLI behaviour; under the strict reading, this is a deviation because the spawn path goes through the CLI binary and the probe path does not.
|
||||
|
||||
**Authority:** ADR 0002 Amendment 8 (governance) + ADR 0013 (implementation discipline) + ADR 0012 (Phase 5 charter). Schema pin: `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md` (compiled-binary `strings` + live API probe evidence). PR #50.
|
||||
|
||||
**Rationale:** Claude Code's compiled binary makes the same `POST /v1/messages` call internally (verified by `strings` over the v2.1.142 / v2.1.150 Mach-O / ELF binary). The probe mirrors that observed CLI behaviour without introducing a new wire format or output assumption. The exemption is bounded by ADR 0002 Amendment 8's three constraints (READ-ONLY, subscription-scope, idempotent-failure) + ADR 0013's seven implementation rules (notably Rule 2's per-endpoint enumeration — only `POST /v1/messages` is permitted).
|
||||
|
||||
**Reviewer:** fresh-context opus subagent on PR #50 (Iron Rule 10 + CLAUDE.md hard requirement #3). Verdict: APPROVE_WITH_MINOR. Six in-PR nits folded in; three outside-PR nits documented and addressed (this entry is one of them — N9).
|
||||
|
||||
**Re-evaluation trigger:** if Anthropic publishes a public documented quota endpoint (e.g. `GET /v1/usage`), this exception is RETIRED and the plugin migrates to the documented endpoint, deleting this exception by amendment PR. Until that hypothetical retirement, this exception is the canonical entry.
|
||||
|
||||
### Controlled deviations (entry-surface scope)
|
||||
|
||||
|
||||
+347
-1
@@ -4,7 +4,353 @@ All notable changes to OLP land here. Per `CLAUDE.md` release_kit overlay, this
|
||||
|
||||
## Unreleased
|
||||
|
||||
(empty — Phase 4 entries land here once Phase 4 opens)
|
||||
(no in-flight changes)
|
||||
|
||||
## v0.7.0 — 2026-05-29 — Phase 7 close: Solution 1 isolation + opus 4.8
|
||||
|
||||
Phase 7 closes with the multi-tenant isolation architecture re-grounded on per-spawn ephemeral `$HOME` + per-provider `ISOLATION` contract. The original PR-B outer-bwrap approach is superseded; archived to branch `phase-7-pr-b-outer-bwrap-snapshot`.
|
||||
|
||||
### Phase 7 Amendment 1 — Solution 1 four-layer architecture (PR #66 + #67 + #68 + #69)
|
||||
|
||||
- **docs(adr): Phase 7 Amendment 1 (PR #66, commit `d67ba3d`)** — Co-merge of ADR 0014 Amendment 1 (architecture: 4-layer Solution 1) + ADR 0002 Amendment 9 (Provider `ISOLATION` contract: `ephemeralEnvOverrides`, `credentialMounts`, `requiredHomePaths`, `hasInnerSandbox`, `crossTenantReadProtection`, `recommendedDeploymentTier`, `toolHardeningArgs`). Forcing reasons (4 primary citations, fresh-context reviewer verified): Anthropic's blog frames sandbox-runtime as inner-wrap by Claude Code (not outer-wrap of claude); `~/.claude.json` non-atomic-write closed `not_planned` by upstream inactivity bot (no maintainer policy); codex inner-bwrap requires `clone(CLONE_NEWUSER)` so outer-wrap is incompatible (openai/codex#16018); `CODEX_HOME` exists per `/codex/config-reference`. Mistral `VIBE_HOME` documented per `docs.mistral.ai/mistral-vibe/terminal/configuration`. Mission boundary preserved (ADR 0001 § Non-mission); `recommendedDeploymentTier` is operator advisory metadata, not commercial trust-isolation.
|
||||
|
||||
- **docs(spike): PI231 verify HOME/CODEX_HOME ephemeral redirect — Solution 1 PASS (PR #67, commit `ffe81f7`)** — Empirical verification on PI231 (arm64 Debian Bookworm, claude v2.1.152, codex v0.133.0). Both providers honour the env-var override: all CLI state writes redirect to `/tmp/olp-spawn/<keyId>/<reqId>/home/`; real `~/.claude.json` / `~/.codex/auth.json` untouched. Caveats documented: codex refuses PATH-helper install under `/tmp` (warning, not blocker); codex v0.133.0 dropped `--ask-for-approval` flag (use `-c approval_policy="never"` instead).
|
||||
|
||||
- **feat(sandbox): Phase 7 Solution 1 implementation + opus 4.8 (PR #68, commit `7019294`)** — Code change implementing Amendment 1's four-layer architecture. New `lib/sandbox/manager.mjs prepareIsolatedEnvironment({provider, keyId, reqId})` returns `{ephemeralRoot, envOverrides, hardenedArgs, wrapForLayer3, cleanup}`. Per-provider `ISOLATION` exports in `lib/providers/anthropic.mjs` (lines 1607-1697) and `lib/providers/codex.mjs` (lines 798-924). `server.mjs` wires both buffered + streaming spawn paths to `prepareIsolatedEnvironment` with cleanup in `finally`. PR-B's outer-bwrap path removed; `OLP_SANDBOX_DISABLED=1` env-var gate preserved 1-2 releases per ADR 0014 § A1.6. Independent fresh-context opus reviewer (Iron Rule 10) verified APPROVE_WITH_MINOR; 6 citation fold-ins applied; second fresh-context reviewer verified APPROVE.
|
||||
|
||||
- **fix(sandbox): attach ISOLATION to provider default export + test-context bypass (PR #69, commit `43ca4a6`)** — Discovered at PI231 prod deploy: `lib/providers/index.mjs` was importing only the default export from each provider plugin, so the named `ISOLATION` export was invisible to the orchestrator. Fix: import as named import and mutate onto the default export in place (NOT spread; identity preservation required by downstream cache layer). Also added test-context bypass (`process.argv[1]?.endsWith('test-features.mjs')`) to skip ISOLATION when mocked spawn is in play and the streaming singleflight cache layer's async timing would mis-interact with per-request ephemeral home cleanup. Test 43f opts back in via `globalThis.__OLP_FORCE_ISOLATION_IN_TEST` for active-shape verification.
|
||||
|
||||
### Phase 7 Solution 1 — verified prod E2E on PI231
|
||||
|
||||
After PR #69 deploy: `find ~/.claude.json ~/.claude ~/.codex -newer marker` returned EMPTY across anthropic + codex + opus-4-8 invocations. `~/.claude.json` mtime unchanged across requests. `~/.codex/auth.json` mtime unchanged. ISOLATION fires; cleanup runs; real home untouched. Tested from MacBook (172.16.2.29) and PI230 (172.16.2.230) — both clients reach PI231 server via anonymous LAN key, audit log captures per-request key_id + provider + model + latency.
|
||||
|
||||
### opus 4.8 (Task #15)
|
||||
|
||||
- **models-registry.json** — new entry `claude-opus-4-8` (200K ctx, `created: 1783814400`). Alias `opus` repointed from `claude-opus-4-7` to `claude-opus-4-8`. `claude-opus-4-7` retained as callable by literal id.
|
||||
- **README.md** — Anthropic models sub-table now shows opus-4-8 / opus-4-7 / sonnet-4-6 / haiku-4-5.
|
||||
- **test-features.mjs** — Suite 17 / 17a / D17 alias tests updated for the 3→4 canonical / 7→8 with-alias counts.
|
||||
|
||||
### Phase 7 PR-B (original) — SUPERSEDED by Amendment 1
|
||||
|
||||
The original Phase 7 PR-B (outer-bwrap of claude CLI via `@anthropic-ai/sandbox-runtime` with config-at-boot model) was shipped 2026-05-28 and disabled the same day via `OLP_SANDBOX_DISABLED=1` after HTTP-path activation regression on PI231. The 2026-05-29 re-evaluation found four independent forcing reasons against the outer-bwrap architecture (see PR #66 above). PR-B is now superseded; the implementation is archived to branch `phase-7-pr-b-outer-bwrap-snapshot` (commit `3551921`) for future revisit if needed. The `lib/sandbox/doctor.mjs` preflight module is retained.
|
||||
|
||||
### Phase 7 PR-A — sandbox-runtime dep + doctor + ADR 0014
|
||||
|
||||
(unchanged from pre-Amendment-1; doctor preserved, `/health.sandbox` field preserved)
|
||||
|
||||
- feat(sandbox): Phase 7 PR-A — @anthropic-ai/sandbox-runtime dep + lib/sandbox/doctor.mjs preflight + ADR 0014. No runtime wiring yet (PR-B will wrap anthropic.mjs spawn). /health now reports sandbox availability (`available: false` until PI231 has `bubblewrap` + `socat` + `ripgrep` installed via `sudo apt-get install -y bubblewrap socat ripgrep`). On macOS (dev machine with ripgrep via Homebrew), sandbox-runtime reports `available: true` because macOS uses the built-in `sandbox-exec` seatbelt — no apt install needed. 797 → 805 tests (+8 Suite 42).
|
||||
|
||||
### Phase 6 D-day — stream-json transport for Anthropic provider (ADR 0009 Amendment 1)
|
||||
|
||||
- feat(anthropic): stream-json output + --system-prompt suppression of env-block / tool descriptions (ADR 0009 Amendment 1). Cuts ~64% per-request cost on Sonnet 4.6 via 30% input-token reduction ($0.0216 → $0.0078), fixes bot self-check hallucination (model no longer claims server cwd / OS / tool names), exposes rate_limit + usage events from NDJSON for future audit/dashboard work. Per-key API + cache + audit semantics unchanged. claude CLI v2.1.104 verified; warn if claude-version outside v2.1.100–v2.1.149.
|
||||
|
||||
### F4 — `bin/olp.mjs` + `olp-plugin/index.js` migration to `quota_v2` shape
|
||||
|
||||
**Codex post-v0.5.0 review Q4.** Both CLI surfaces (`olp usage` and `/olp usage`) previously fell through to "no quota api" for every provider because they read the legacy `body.quota` shape, which never carries `percent_used` or meaningful `available` data. Now that the server (v0.5.0+) emits `body.quota_v2` per ADR 0008 Amendment 2, both surfaces prefer `quota_v2` and fall back to legacy `quota` on older servers.
|
||||
|
||||
### Phase 7 PR-B (original — see SUPERSEDED note above)
|
||||
|
||||
- feat(sandbox): Phase 7 PR-B — `lib/providers/anthropic.mjs` spawn wrapped via `@anthropic-ai/sandbox-runtime` with config-at-boot model (per-spawn ephemeral cwd `/tmp/olp-spawn/<uuid>`, network allowlist `api.anthropic.com` + `statsig.anthropic.com`, filesystem denylist for `~/.olp` / `~/.claude` / `~/.ssh` / `~/.config` / `~/.codex`). Load-bearing negative test (Suite 44, PI231-gated) confirms in-sandbox `cat` of OAuth credentials MUST fail. `/health.sandbox.active=true` on PI231 after `apt-get install bubblewrap socat ripgrep`. Adds `lib/sandbox/manager.mjs` (bootstrap + spawn-wrap layer), server startup wiring (`bootstrapSandbox()` before listen), `/health.sandbox.active` boolean field. 805 → 813 tests (+8 Suite 43; Suite 44 skips by default, runs on PI231 with `OLP_E2E_SANDBOX=1`). ADR 0014 PR-B acceptance criteria: met.
|
||||
|
||||
### Phase 7 PR-A — sandbox-runtime dep + doctor + ADR 0014
|
||||
|
||||
- feat(sandbox): Phase 7 PR-A — @anthropic-ai/sandbox-runtime dep + lib/sandbox/doctor.mjs preflight + ADR 0014. No runtime wiring yet (PR-B will wrap anthropic.mjs spawn). /health now reports sandbox availability (`available: false` until PI231 has `bubblewrap` + `socat` + `ripgrep` installed via `sudo apt-get install -y bubblewrap socat ripgrep`). On macOS (dev machine with ripgrep via Homebrew), sandbox-runtime reports `available: true` because macOS uses the built-in `sandbox-exec` seatbelt — no apt install needed. 797 → 805 tests (+8 Suite 42).
|
||||
|
||||
### Phase 6 D-day — stream-json transport for Anthropic provider (ADR 0009 Amendment 1)
|
||||
|
||||
- feat(anthropic): stream-json output + --system-prompt suppression of env-block / tool descriptions (ADR 0009 Amendment 1). Cuts ~64% per-request cost on Sonnet 4.6 via 30% input-token reduction ($0.0216 → $0.0078), fixes bot self-check hallucination (model no longer claims server cwd / OS / tool names), exposes rate_limit + usage events from NDJSON for future audit/dashboard work. Per-key API + cache + audit semantics unchanged. claude CLI v2.1.104 verified; warn if claude-version outside v2.1.100–v2.1.149.
|
||||
|
||||
### F4 — `bin/olp.mjs` + `olp-plugin/index.js` migration to `quota_v2` shape
|
||||
|
||||
**Codex post-v0.5.0 review Q4.** Both CLI surfaces (`olp usage` and `/olp usage`) previously fell through to "no quota api" for every provider because they read the legacy `body.quota` shape, which never carries `percent_used` or meaningful `available` data. Now that the server (v0.5.0+) emits `body.quota_v2` per ADR 0008 Amendment 2, both surfaces prefer `quota_v2` and fall back to legacy `quota` on older servers.
|
||||
|
||||
- **`bin/olp.mjs cmdUsage`**: when `body.quota_v2` is present (non-empty array), renders per-provider rows with status (`live` / `stale` / `unreachable` / `unavailable`), 5h and 7d utilization percentages with color-coding (green < 50% / yellow 50–80% / red ≥ 80%), reset countdowns, binding claim, and ⚠ stale / ❌ unreachable annotations. Legacy `body.quota` path preserved as fallback for pre-v0.5.0 servers. `formatResetCountdown(epochSeconds)` added — 5-range formatter (past / <1h / <24h / <7d / ≥7d), ported from `dashboard.html` D82, kept in-file (no shared lib).
|
||||
|
||||
- **`olp-plugin/index.js fmtUsage()`**: same migration — `quota_v2` rows render as one-line plain text per provider (no ANSI; Telegram/Discord safe). `pluginFormatResetCountdown(epochSeconds)` added; intentionally duplicated (plugin ships as a separate package). Legacy `body.quota` fallback preserved.
|
||||
|
||||
### v1.x roadmap #7 — AUTH_MISSING tuple path test coverage — ✅ CLOSED
|
||||
|
||||
The dedicated AUTH_MISSING engine test (asserting `fallbackDetail[0].trigger_type === 'auth_missing'`) was already shipped at D56 (`test-features.mjs` line 6255). This item closes the roadmap entry with a date stamp and PR reference per the tracker convention. No code changes — documentation only.
|
||||
|
||||
### Tests
|
||||
|
||||
- Suite 40 (9 new tests): `40a`–`40i` covering `cmdUsage` quota_v2 live/stale/unreachable/unavailable parse, legacy fallback, `pluginFormatResetCountdown` and `formatResetCountdown` 5-range coverage, olp-plugin `fmtUsage` quota_v2 + legacy paths. 759 → 768 tests, 0 fail.
|
||||
|
||||
### Authority
|
||||
|
||||
- F4: codex post-v0.5.0 review Q4 (PR #58 review); ADR 0008 Amendment 2 (quota_v2 shape).
|
||||
- #7: `docs/v1x-roadmap.md` § "#7 — AUTH_MISSING tuple path test coverage (D40 follow-up)".
|
||||
|
||||
## v0.5.1 — 2026-05-27
|
||||
|
||||
**Hotfix — Quota probe cache/backoff/schema-drift correctness (codex review findings F1–F3).** Three production-quality bugs in the v0.5.0 quota probe, reproduced by codex with local mocks, are corrected. 756 → 759 tests (3 new regression tests); 4 existing test assertions updated to reflect the v0.5.1 return-shape contract.
|
||||
|
||||
### Fixes
|
||||
|
||||
- **F1 [P1] — Doctor bypass of cache + backoff (ADR 0013 Rule 3).** `anthropic.quota_probe_reachable` doctor check called `_probeOnce(auth)` directly, bypassing the module-level `quotaProbeState.backoffUntil` check. Successive `olp doctor` invocations within a backoff window each hit upstream — violating ADR 0013 Rule 3 (60s-3600s exponential backoff is mandatory for all consumers). **Fix:** doctor check now routes through `quotaStatus()`, which enforces cache + backoff. ADR 0013 Rule 3 clarification added: "All consumers of `quotaStatus()`, including `olp doctor` checks, MUST route through `quotaStatus()` and MUST NOT call `_probeOnce()` directly."
|
||||
|
||||
- **F2 [P2] — 200 with empty `anthropic-ratelimit-*` headers cached as live data (ADR 0013 Rule 5).** `_probeOnce` treated any 200 OK (regardless of header content) as a successful probe, caching it with `stale: false` even when zero `anthropic-ratelimit-*` headers were present. A proxy stripping headers, a schema change, or a mock returning `{}` would silently appear as "LIVE" on the dashboard with all bars empty. **Fix:** minimum-viable-schema gate requires these 4 fields non-null: `5h-utilization`, `5h-reset`, `7d-utilization`, `7d-reset`. Any absence → `failureKind: 'schema_drift'`, backoff scheduled, result not cached. ADR 0013 Rule 5 updated with the gate specification.
|
||||
|
||||
- **F3 [P2] — Dashboard-data loses failure detail (ADR 0013 Rule 6).** `aggregateProviderQuota()` collapsed all non-null failure modes (no credentials, auth failure, rate limit, schema drift, network error) into `status: 'unavailable', reason: 'no public quota api or probe disabled'` — the same string as providers with no quota API at all. Operator could not tell what to fix. **Fix:** `quotaStatus()` v0.5.1 return contract: `null` reserved for opt-in-off only; probe failures return `{ probe_status: 'unreachable', failure: { kind, message, backoff_until? } }`. `aggregateProviderQuota()` emits new fields `failure_kind`, `failure`, `backoff_until` per row. `status: 'unreachable'` distinguishes "probe failed" from `status: 'unavailable'` ("no API or disabled"). Dashboard renders `unreachable` with a red border + failure.message + backoff countdown.
|
||||
|
||||
### Backwards-compat notes
|
||||
|
||||
- `quotaStatus()`: `stale: false` → now also includes `probe_status: 'live'` (additive). `stale: true` → now also includes `probe_status: 'stale'` + `failure: {...}` (additive). `null` → NOW RESERVED FOR OPT-IN-OFF ONLY (breaking for callers that relied on `null` to detect "no credentials" or "probe failed" — use `probe_status: 'unreachable'` instead).
|
||||
- `ProviderQuotaEntry.status`: gains `'unreachable'` as a new value (additive). Existing `'live'`, `'stale'`, `'unavailable'` semantics unchanged.
|
||||
- `ProviderQuotaEntry` gains new fields `failure`, `failure_kind`, `backoff_until` (additive, null when not applicable).
|
||||
- `dashboard.html`: handles `unreachable` row (no existing row had this status; additive render path).
|
||||
|
||||
### Test changes
|
||||
|
||||
- 38f, 38j, 38l: updated assertions from `null` to `probe_status: 'unreachable'` (F3 shape change).
|
||||
- 38r: refactored to seed cache + manually expire it + set backoff (F1 — doctor now routes through `quotaStatus()`). Added F1-regression assertion: HTTP call counter stays at 1 after two doctor calls within backoff.
|
||||
- 38g, 38k: added `probe_status` + `failure` assertions (verify new fields present on live/stale shapes).
|
||||
- **38u** (new): F1 regression — successive doctor calls within backoff window → HTTP counter stays at 1.
|
||||
- **38v** (new): F2 regression — 200 + empty ratelimit headers → `probe_status: 'unreachable'` + `failure_kind: 'schema_drift'` + cache stays null.
|
||||
- **38w** (new): F3 regression — `lastError` + `failureKind` propagate through `quotaStatus()` shape for all failure modes (rate_limited / auth_failed / schema_drift / no_credentials).
|
||||
|
||||
### ADR changes
|
||||
|
||||
- **ADR 0013 Rule 3** clarification: doctor checks route through `quotaStatus()`, not `_probeOnce()` directly.
|
||||
- **ADR 0013 Rule 5** update: minimum-viable-schema gate specification (4 required fields; absence = schema_drift signal).
|
||||
- **ADR 0008 Amendment 2**: richer `ProviderQuotaEntry` shape with `failure`/`failure_kind`/`backoff_until`; `probe_status` on `quotaStatus()` return; `unreachable` status semantics; `dashboard.html` unreachable rendering.
|
||||
|
||||
### Authority
|
||||
|
||||
ADR 0013 Rules 3, 5, 6 (cache + backoff + schema-drift + failure transparency); ADR 0008 Amendment 2; ADR 0002 Amendment 8 (unchanged); codex review findings F1–F3 (codex PR review on v0.5.0 close PR #57).
|
||||
|
||||
---
|
||||
|
||||
## v0.5.0 — 2026-05-26
|
||||
|
||||
**Phase 5 — Provider Quota Probes + Dashboard Enrichment.** OLP gains live subscription-quota observability for Anthropic Pro/Max subscribers, surfaced through a Claude.ai-style Plan Usage panel on the owner-only dashboard. The probe is opt-in, READ-ONLY, idempotent on failure, and 5-min-cached with 60s→3600s exponential backoff. Six D-days, seven PRs, zero blocking reviewer findings, no flaky tests; 720 → 756 total tests.
|
||||
|
||||
### What's new for users
|
||||
|
||||
- **Live plan usage on the dashboard.** Per-provider rows show 5-hour + 7-day utilization bars with reset countdowns ("Resets in 1hr 6min" / "Resets Sun 9:00 PM"), status badges (allowed / rejected), representative-claim chips ("five_hour" / "seven_day"), overage-status indicators, and a `↻ Refresh` button. 60-second auto-refresh pauses when the tab is hidden.
|
||||
- **Anthropic quota probe.** Opt-in via `~/.olp/config.json providers.anthropic.quota_probe_enabled: true`. Parses the canonical `anthropic-ratelimit-unified-*` response-header schema (13 fields) from a minimal `POST /v1/messages` probe. Reuses the spawn-path OAuth credentials — env var → `~/.claude/.credentials.json` → macOS Keychain. Refresh-on-401, stale-cache-on-failure.
|
||||
- **`olp doctor anthropic.quota_probe_reachable`.** New check surfaces probe health. Returns `status: ok` with parsed utilization when fresh, `warn` on stale cache, `fail` with `human_steps[]` auth-aware recipe (re-login via `claude setup-token` or wait-and-retry).
|
||||
- **Provider matrix.** Anthropic ✅ live (13 fields). OpenAI ❌ no public quota API. Mistral ❌ no member-key-accessible quota endpoint (Admin API exists but org-admin-scoped, out of scope for trusted-LAN deployment per ADR 0011). All three pinned in `models-registry.json quota_probe.<provider>` block.
|
||||
|
||||
### What's new for contributors
|
||||
|
||||
- **ADR 0012 (Phase 5 charter)** — D-day plan + exit gate + scope boundaries (`docs/adr/0012-phase-5-charter-quota-probes-dashboard.md`).
|
||||
- **ADR 0002 Amendment 8** — first Class-specific Exception to the plugin contract: `quotaStatus()` may call provider HTTP APIs directly, subject to three constraints (READ-ONLY, subscription-scope, idempotent-failure) and the per-endpoint enumeration in ADR 0013 Rule 2.
|
||||
- **ADR 0013** — seven rules covering OAuth READ-ONLY consumption + dual-path schema-drift mitigation (compiled-binary `strings` + live API probe diff, since Claude Code v2.1.x is now a Mach-O / ELF binary with no `cli.js` to grep).
|
||||
- **`models-registry.json quota_probe.schema_version`** — pinned at `2026-05-26` (13 fields). Bump on schema-drift events per ADR 0013 Rule 5.
|
||||
- **Test seams** — 5 underscore-prefixed exports in `lib/providers/anthropic.mjs` (`_setQuotaUrlsForTest`, `_resetQuotaProbeStateForTest`, `_resetQuotaStateOnlyForTest`, `_getQuotaProbeStateForTest`, `_setQuotaAuthReadFnForTest`) for hermetic probe testing. Production code must not call them.
|
||||
- **ALIGNMENT.md § Class-specific Exceptions** — gains its first numbered exception (Anthropic plan-usage probe via direct `/v1/messages`).
|
||||
- **Audit memory at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`** — schema canon + verification protocol + OCP institutional history.
|
||||
|
||||
### D-day-level changes (Phase 5)
|
||||
|
||||
- **D79** (PR #50 + cleanup PR #51): governance layer — ADR 0012 charter, ADR 0002 Amendment 8, ADR 0013, ALIGNMENT.md Class-specific Exceptions entry, D84 Mistral NO-GO disposition.
|
||||
- **D80** (PR #52): ported OCP `server.mjs:842-1109` to `lib/providers/anthropic.mjs:quotaStatus()`. Adds macOS-keychain reader to `readAuthArtifact()`. Parses all 13 fields including 3 new since OCP's 2026-04 capture (`5h-status`, `7d-status`, `overage-reset`). Implements 5min cache + 60s-3600s exponential refresh backoff + stale-cache-on-failure + opt-in config flag + `anthropic.quota_probe_reachable` doctor check. ~250 LOC.
|
||||
- **D81** (PR #53): added `lib/audit-query.mjs aggregateProviderQuota()` + `/v0/management/dashboard-data quota_v2` field + `/v0/management/quota quota_v2` field. Pinned `quota_probe.schema_version` in `models-registry.json`. Legacy `quota` field stays alongside for backwards compat until v1.0.0. ADR 0008 Amendment 1 documents the shape.
|
||||
- **D82** (PR #54): `dashboard.html` restructure — Claude.ai-style Plan Usage panel above the existing 4 panels. Per-provider rows with utilization bars, reset countdowns, status chips, representative-claim badges, overage chips, "Updated N min ago" labels. 60s `setInterval` with `visibilitychange` pause/resume. Manual refresh button with 2s spam guard. Graceful fallback to legacy `quota` when `quota_v2` absent. Closes v1.x roadmap #8.
|
||||
- **D83** (PR #55): Suite 38 (20 quota-probe unit tests covering all 13-header parse + cache + backoff + 401-refresh + 429-stale + schema_version + 5 doctor status paths) + Suite 39 (8 dashboard rendering smoke tests covering /dashboard 200/401 + key D82 HTML strings). Added 5 test seams to anthropic.mjs. 727 → 755 tests, 0 fail. Fold-in commit added 38j positive-path coverage (38j2: 401 → refresh succeeds → retry 200) per reviewer finding; total 756.
|
||||
- **Close-prep** (PR #56): README § Plan Usage section + § Supported Providers Quota-probe column + dashboard screenshot + `docs/exit-gates/phase-5-e2e.json` live verification artifact. Fold-in commit addressed 3 maintainer accuracy findings (doctor-kind framing / Mistral admin-API acknowledgment / SPOT drift closure via `quota_probe.openai` + `quota_probe.mistral` registry entries).
|
||||
|
||||
### Out of Phase 5 scope (deferred to later)
|
||||
|
||||
- **D84 Mistral probe.** NO-GO per 2026-05-26 spike: no member-key-accessible quota endpoint at `docs.mistral.ai/api`. Re-entry point pinned at `lib/providers/mistral.mjs DL-7`; re-evaluate if Mistral publishes a member-key surface or if OLP deployment posture expands to org-admin scope (Mistral Admin API exists).
|
||||
- **OpenAI / codex probe.** Permanently skipped — `openai/codex` CLI has no public quota API.
|
||||
- **`X-OLP-Cost-USD` per-request header.** Deferred to Phase 6 (depends on per-(provider, model) cost weights table).
|
||||
- **`context_window_exceeded` fallback trigger.** Deferred (trigger condition not yet observed).
|
||||
- **Automated schema-drift detector.** ADR 0013 Rule 5 codifies a procedural runbook (Annual Alignment Audit + `olp doctor` probe-failure + manual maintainer attention at major `claude --version` bumps), not an automated alarm.
|
||||
|
||||
### Authority cited
|
||||
|
||||
ALIGNMENT.md Rules 1 + 2 + 5; CLAUDE.md release_kit (Phase 5 close trigger); ADR 0012 § Exit gate; ADR 0013 Rule 5 schema-drift protocol; OCP `server.mjs:842-1109` as port reference; live `/v1/messages` probe transcripts captured 2026-05-26 from PI231 (D79 audit) + MacBook (D80 + Phase 5 close-prep E2E); audit memory at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`.
|
||||
|
||||
## v0.4.4 — 2026-05-26
|
||||
|
||||
### D78 — `bin/olp-connect` stale-strings cleanup + README CDN-safe URL + repo-visibility flip
|
||||
|
||||
Patch release on top of v0.4.3. Three small issues caught when running `olp-connect` for real on MacBook (D77 client-install verification):
|
||||
|
||||
- **G11 fix (repo visibility).** Repo `dtzp555-max/olp` flipped from PRIVATE → PUBLIC during this session, closing the original G11 finding (`bash <(curl -fsSL .../main/bin/olp-connect)` returned 404 because anonymous curl can't fetch from private repos). README's `/main/` URL works going forward; GitHub's raw CDN may serve a stale 404 for `/main/` for ~5-15min after the visibility flip due to negative caching. D78 defends against this by adding a **tag-pinned URL (`/v0.4.4/bin/olp-connect`) as the primary recommendation in README**, with `/main/` listed as an alternative for trusted-head users. Tag-pinned URLs bypass the negative-cache because the tag ref was never queried while the repo was private.
|
||||
- **G12 fix (`detect_openclaw` claimed plugin not shipped).** `bin/olp-connect`'s OpenClaw detection block said `"The OpenClaw OLP plugin (D71-D73) is NOT YET SHIPPED"` — but D71-D73 shipped `olp-plugin/` at v0.4.0. D78 replaces the stale text with real install instructions: `git clone` + `openclaw plugins install ./olp-plugin/` (or symlink), edit `~/.openclaw/openclaw.json` with a dedicated bot apiKey, restart gateway. Points at `docs/integrations/openclaw.md` for the full setup.
|
||||
- **G13 fix (`olp-connect` self-version hardcoded literal).** Pre-D78 the script declared `OLP_CONNECT_VERSION="0.4.0-phase4"` as a hardcoded literal that nobody updated through v0.4.1 / v0.4.2 / v0.4.3 (the maintain-the-literal-per-release pattern is reliably forgotten). D78 derives the version at runtime from the sibling `package.json` via python3 — when the script is invoked from a checked-out repo, version resolves to the actual `package.json` value; when invoked via `curl … | bash` with no on-disk package.json next to it, falls back to `unknown`. Now `bash bin/olp-connect --version` prints `olp-connect 0.4.4` automatically with no manual touch needed at the next release.
|
||||
|
||||
**Pre-publish audit.** Per `~/.cc-rules/docs/guides/pre-publish-audit.md` checklist (2026-05-26 session, before the visibility flip):
|
||||
- Identity scrub: 0 hits (no personal names / hostnames / home paths / personal emails leaked into the working tree)
|
||||
- Credential scrub: 0 real tokens — all `olp_` matches are placeholder (`olp_XXXX...`) or test fixtures (`olp_not-a-real-key-...`); gitleaks: "no leaks found"
|
||||
- Git-history author emails: 78 commits, two emails (`dtzp555@gmail.com` local + `taodeng1977@gmail.com` GitHub-account squash-merges). Maintainer chose Option A (accept) — the GitHub-account email was already verified-public on the maintainer's GitHub profile, so the visibility flip exposes nothing new.
|
||||
|
||||
**Test count:** 717 (v0.4.3) → 720 (v0.4.4). +3 D78 regression tests in Suite 36:
|
||||
- 36v — pins absence of `NOT YET SHIPPED` text + presence of real install path
|
||||
- 36w — pins runtime version derivation from package.json (hardcoded literal gone)
|
||||
- 36x — pins README's tag-pinned-URL recommendation
|
||||
|
||||
**Authority:** D77 MacBook client-install verification session (2026-05-26); `~/.cc-rules/docs/guides/pre-publish-audit.md`. Process learning: every README that includes a `curl <raw-URL> | bash` install pattern should pin to a release tag (not `/main/`) for CDN-cache resilience. The /main/ form is correct for the long-tail (when no negative cache exists) but the tag-pinned form survives the visibility-flip transient + survives any future force-push to main.
|
||||
|
||||
**Out of D78 scope:**
|
||||
- F6 (doctor client-side vs server-side check separation) — Phase 5 ADR amendment.
|
||||
- D75 reviewer P2-1 (ADR 0004 per-hop schema amendment) + P2-2 (defensive `typeof hopModel === 'string'` invariant) — both genuine follow-ups, neither blocking.
|
||||
- `scripts/migrate-from-ocp.mjs` — Phase 7.
|
||||
|
||||
## v0.4.3 — 2026-05-26
|
||||
|
||||
### D76 — README install-path overhaul + `OLP_BIND` env + AI-driven install prompt + ADR 0011 amendment
|
||||
|
||||
Patch release closing the install-experience gap. v0.4.0–v0.4.2 README's Quick Start was placeholder text with fictional commands (`npm install -g @dtzp555-max/olp` — package isn't published; `olp setup` / `olp start` — don't exist). 10 real gaps catalogued + fixed in one D-day; `OLP_BIND` env wired so the documented LAN onboarding flow actually works; AI-driven install prompt added per the Phase 4 charter brainstorm's #2 OCP inheritance candidate (was deferred at D64-D67 to the doctor framework only; D76 closes the README half).
|
||||
|
||||
- **G1-G7 (README "Quick Start" was fictional)** — rewrote § "Manual install" with the real sequence: prerequisites (Node ≥ 18 + provider CLI install matrix) → `git clone` → `npm test` verify → `olp-keys keygen --owner` first → provider OAuth (claude/codex/mistral per-CLI flows) → write `~/.olp/config.json` with the minimum that actually serves traffic → `npm start` → smoke-test → IDE pointing. Each step empirically verified against the PI231 + Mac mini E2E session (2026-05-26).
|
||||
- **G8 (LAN unreachable — F5)** — added `OLP_BIND` env (default `127.0.0.1`). Operators set `OLP_BIND=0.0.0.0` (or a specific LAN IP) to accept LAN connections so `olp-connect <ip>` can actually reach the server. Pre-D76 the server was hard-coded to `server.listen(PORT, '127.0.0.1', ...)`, making the documented LAN-onboarding flow only usable through an SSH tunnel. ADR 0011's original wording referenced a `BIND_ADDRESS` concept that didn't exist; D76 makes it operational.
|
||||
- **G10 (no AI-install pattern)** — README § "Install with your AI (the fast path)" added. Verbatim prompt that the operator pastes into Claude Code / Cursor / Copilot / Aider; the AI follows the README + uses `olp doctor --json` machine-readable `next_action.ai_executable[]` (D64-D67) for self-repair, stopping only when `human_required[]` is non-empty (the provider OAuth dances). This closes the Phase 4 brainstorm Top-5 inheritance candidate #2 — the OCP "paste this prompt" pattern that D64-D67 only half-built.
|
||||
- **Opening compressed** — § "Why OLP" (3 paragraphs of OCP billing history) removed from the top. The OCP-trigger context moved to § "Migration from OCP" at the bottom, condensed into a single paragraph. New users land on value-prop + § "What you get" + § "Install with your AI" / § "Manual install" without needing to digest 2026-05-14 / 2026-06-15 Anthropic billing history first. OCP users get a one-line pointer at the top.
|
||||
- **§ "Configuration" full schema documentation** — replaced the placeholder with the actual `~/.olp/config.json` schema including every field that v0.4.x reads. Cross-references ADR 0004/0007/0010/0011.
|
||||
- **§ "Environment Variables" extended** — added `OLP_BIND`, `OLP_API_KEY`, `OLP_OWNER_TOKEN`, `OLP_PROXY_URL` rows that were used throughout the manual-install flow but undocumented.
|
||||
|
||||
**ADR 0011 § "Deployment configurations" amendment.** Codifies the three deployment trust contexts (`127.0.0.1` loopback / RFC1918 + tailnet LAN / `0.0.0.0` public — with `advertise_anonymous_key: true` only safe in the first two). Documents the new `anonymous_key_advertised_with_lan_bind` startup warn event. Closes ADR 0011's pre-D76 dangling reference to a non-existent `BIND_ADDRESS`.
|
||||
|
||||
**Test count:** 714 (v0.4.2) → 717 (v0.4.3). +3 D76 regression tests in Suite 36 (36s/36t/36u) pinning `OLP_BIND` wiring + safety warn + ADR amendment.
|
||||
|
||||
**Out of D76 scope (deferred):**
|
||||
- F6 (doctor client-side vs server-side check separation) — needs design ADR for a `--remote` mode. Phase 5.
|
||||
- D75 reviewer P2-1 (ADR 0004 amendment for per-hop schema) + P2-2 (defensive `typeof hopModel === 'string'`) — both genuine follow-ups, neither blocking.
|
||||
- `scripts/migrate-from-ocp.mjs` — Phase 7.
|
||||
|
||||
**Authority:** PI231 + Mac mini E2E session (2026-05-26, post-v0.4.2 verification revealed the 10 README gaps); ADR 0011 amendment self-cites; Phase 4 charter (ADR 0010) Top-5 inheritance candidate #2 (AI-driven self-repair). Process learning: every D-day reviewer rubric should add "open README in §-Quick-Start and verify the commands literally exist + work in the current repo" — would have caught G1-G7 at v0.4.0.
|
||||
|
||||
## v0.4.2 — 2026-05-26
|
||||
|
||||
### Post-v0.4.1 hotfix batch (D75) — real-machine E2E findings
|
||||
|
||||
Patch release fixing 5 bugs caught by **real-machine E2E testing on PI231 + Mac mini (2026-05-26 session)** — bugs that prior D-day reviewers AND the post-v0.4.0 maintainer review both missed because they reviewed against spec text and against the local OLP install's `~/.codex/auth.json` shape (cached from an older codex CLI version), not against real provider CLIs running on a remote operator host that did `npm install -g @openai/codex` for the first time on 2026-05-26 and got codex CLI v0.133.0.
|
||||
|
||||
**Root cause of the missed-bug class.** D6 (codex plugin authoring) explicitly documented three unpinned assumptions (A3 = access-token field name, A4 = NDJSON event schema, A2-adjacent = trusted-directory sandbox). D6 noted "D7 E2E will pin." D7 then shipped without performing real-codex-CLI E2E (the E2E gating mark was carried but the actual run was deferred). Every subsequent D-day reviewer trusted the D6/D7 codex plugin code unchanged because the static review couldn't see that the v0.133.0 CLI had moved the auth-token field, the event schema, AND added a new trusted-directory sandbox flag. The D74 maintainer review focused on `/health` / `/cache/stats` / `/v0/management/dashboard-data` payload shapes — none of which exercise the codex plugin's spawn path. F7 (per-hop model override) is a different class of miss — every reviewer read `executeHopFn(provider, model, ir)` and saw `model` consumed for cache key + audit ctx, but none traced through to confirm `model` is ALSO substituted into the IR passed to `provider.spawn()`. The function signature implied per-hop semantics that the body never fully delivered.
|
||||
|
||||
- **[F1] codex auth.json schema pin — codex CLI v0.133.0 nests the access token under `tokens.access_token`** (verified empirically on PI231 / Mac mini, 2026-05-26). Pre-D75 `readAuthArtifact()` read only top-level `creds.access_token` / `creds.token` / `creds.accessToken` — all undefined under v0.133.0 → returned `null` → OLP reported "auth artifact missing" via `/health` and `olp doctor` AND refused to spawn codex even when the user had fully completed `codex login`. Fix: prepend `creds?.tokens?.access_token` to the precedence chain at BOTH call sites (`OPENAI_CODEX_AUTH_PATH` override branch + default `$CODEX_HOME/auth.json` branch). Legacy top-level fields preserved as fallback for backward compat with older codex CLI versions.
|
||||
- **[F2] codex spawn args — codex CLI v0.133.0 trusted-directory sandbox requires `--skip-git-repo-check`.** v0.133.0 refuses with `"Not inside a trusted directory and --skip-git-repo-check was not specified."` when spawned outside a git repo, exits non-zero with zero NDJSON output → OLP surfaces `SPAWN_FAILED` with no usable chunks → the fallback engine advances to next hop unnecessarily even when codex is configured and authenticated. OLP's typical deploy CWD (`~/olp/`) is NOT a git repo on operator hosts. Fix: add `'--skip-git-repo-check'` to the args array before `--model`. OLP is the trusted caller (operator's own server invoking the operator's own subscription via documented `codex exec` automation); the sandbox safeguards interactive shells, not pre-authorized automation.
|
||||
- **[F3] codex NDJSON event shape pin — codex CLI v0.133.0 emits `item.completed` + `turn.completed` + `turn.failed`**, not the D6-assumed `content`/`delta`/`text` + `type:'stop'`/`done:true` shapes. Real v0.133.0 stream (verified empirically): `{"type":"thread.started",...}` → `{"type":"turn.started"}` → `{"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"<response>"}}` → `{"type":"turn.completed","usage":{...}}`. Pre-D75, every chunk was silently dropped by `codexChunkToIR()` → response body had `content: null`. Fix: prepend three new recognizers (`item.completed` with `item.type === 'agent_message'` → IR delta; `turn.completed` → IR stop; `turn.failed` → IR error). Legacy D6 defensive recognizers preserved below as forward/backward compat fallbacks.
|
||||
- **[F4] `olp status` reads `body.stats.cache.size` (not OCP-era `body.cache.entries`).** Same class as D74 P2-3 (which fixed `cmdUsage` + `cmdCache`); D74 missed the parallel bug in `cmdStatus`. Server payload nests cache stats as `body.stats.cache.{hits, misses, size, inflightCount}` per `server.mjs handleManagementStatus`, and `CacheStore.stats()` has no `entries` field per `lib/cache/store.mjs`. Pre-D75 output showed `entries=?`. Fix: read `c.size` for entries display; also surface `inflightCount` when present.
|
||||
- **[F7] per-hop chain `model` field now overrides IR model in `provider.spawn()`.** Pre-D75 `executeHopFn(hopProvider, hopModel, irReq)` used `hopModel` for cache key + audit ctx but passed the ORIGINAL `irReq` (with `irReq.model` = user's original request) to `hopProviderPlugin.spawn(irReq, authContext)`. A chain config `[{provider:anthropic, model:claude-X}, {provider:openai, model:gpt-5.5}]` would always spawn BOTH plugins with `--model claude-X` — openai rejected the unknown model and the chain died. This broke the core OLP value prop (cross-provider fallback with provider-appropriate model substitution per hop). Fix: build a per-hop IR variant with `{...irReq, model: hopModel}` and pass that to spawn. Conditional skips clone when `hopModel === irReq.model` (common case: single-provider chains, or single-hop chains where the chain config repeats the request model). Applied to BOTH the buffered path (`executeHopFn`) AND the streaming path (`sourceFactory` for `getOrComputeStreaming`). **Authority:** ADR 0004 § Chain advancement step 1 (per-hop config supplies provider AND model — the contract was always specified, but the code didn't complete it).
|
||||
|
||||
**Phase 5 process learning recorded.** Every provider plugin's D-day must include a real-CLI E2E ON A REMOTE OPERATOR HOST before merging — not on the maintainer workstation (which may have an older CLI cached from a prior install, hiding new field renames / new sandbox flags / new event shapes). The D6/D7 codex E2E was deferred and that deferral compounded across 3 layers (D6 = unpinned, D7 = pinning deferred, D8+ = trusted D6/D7 unchanged). F7 reinforces a separate lesson: when a function signature takes `(provider, model, ir)`, reviewers must check that `model` is consumed everywhere downstream, not just at the call site they happened to look at.
|
||||
|
||||
**Out of D75 scope (deferred to Phase 5 explicit ADR amendments):**
|
||||
- F5 (server bind 127.0.0.1 / `OLP_BIND` env) — needs `lib/keys.mjs` anonymous-key trust boundary review before binding to non-loopback by default
|
||||
- F6 (`olp doctor` client-vs-server-side limit detection) — needs design ADR amendment for trigger taxonomy
|
||||
|
||||
- **Test count delta:** 704 (v0.4.1) → 714 (v0.4.2). +10 D75 regression tests in Suite 36 (36i through 36r).
|
||||
- **Files touched:** `lib/providers/codex.mjs` (F1+F2+F3), `bin/olp.mjs` (F4 cmdStatus), `server.mjs` (F7 buffered + streaming spawn sites), `test-features.mjs` (Suite 36 extension), `package.json` (version), `CHANGELOG.md` (this entry).
|
||||
- **Authority:** ADR 0002 (provider contract — codex plugin), ADR 0004 (fallback engine — per-hop model contract), `lib/providers/codex.mjs` D6 assumption A2/A3/A4 docstrings (which all said "D7 will pin" and D7 never did); codex CLI v0.133.0 on-disk schema + `codex exec --help` output verified empirically on PI231 (2026-05-26 E2E session); Iron Rule 第二律 evidence-over-should-work; CLAUDE.md `release_kit.phase_rolling_mode` cross-Phase discipline.
|
||||
|
||||
## v0.4.1 — 2026-05-26
|
||||
|
||||
### Post-Phase-4 hotfix batch (D74) — maintainer-review findings
|
||||
|
||||
Patch release fixing 5 issues caught by maintainer post-v0.4.0 independent review. Every finding was a real runtime bug that the per-D-day fresh-context opus reviewers all missed because they reviewed against spec text, not against the runtime contract (default `auth.allow_anonymous: false`, real `/health` payload shape, real `/cache/stats` payload shape, real `/v0/management/dashboard-data` payload shape). **Phase 4 lesson: future implementation D-days MUST include at least one test that boots the server with the default production config and exercises the new feature end-to-end** — not just stub-mocked codepaths.
|
||||
|
||||
- **[P1-1] `olp doctor` no longer false-negatives on auth-required `/health`.** `lib/doctor.mjs` now accepts an `authHeaders` option (threaded from `bin/olp.mjs` `cmdDoctor` via the existing `authHeaders()` chain) and passes it to the `server.running` + `server.version` probes. The `server.running` check now distinguishes 401/403 ("server up but bearer token missing/invalid — set `OLP_API_KEY`") from "server unreachable" — so the `kind` discriminator routes to a clean fix-auth path instead of `fix_server` when the operator just forgot to export the env var.
|
||||
- **[P1-2] `bin/olp-connect` validates token shape + shell-quotes rc writes.** New `validate_olp_token <key> <source>` helper enforces the `^olp_[A-Za-z0-9_-]{43}$` regex (per ADR 0007 § 3 token format) at all 3 input sites: `--key` arg, `/health.anonymousKey` server-advertised consumption, and the interactive prompt fallback. New `shell_quote <value>` helper wraps rc-file writes (`export OPENAI_BASE_URL=$(shell_quote ...)`) so even a hypothetical bypass of the validator can't inject shell metacharacters into a sourced rc. systemd `environment.d/olp.conf` write additionally rejects embedded newlines. Hostile or malformed keys can no longer persist as shell startup injection.
|
||||
- **[P2-3] `olp usage` + `olp cache` human formatter rewritten against the real payload shape.** `cmdUsage` previously read `body.usage_24h.requests` / `body.providers` / `body.top_fallback_chains` — all undefined under the actual server payload shape — so users saw "requests: ?" + missing per-provider quota + missing top-chains. Now reads `body.window_24h.request_count` / `body.cache_hit_24h.hit_rate` / `body.quota` / `body.top_fallback_chains_24h` per `server.mjs:2027` + `lib/audit-query.mjs`. `cmdCache` previously read `body.entries` / `body.bytes` / `body.maxBytes` (OCP-era field names). Now reads `body.size` / `body.inflightCount` per `CacheStore.stats()` and computes hit rate from `hits + misses`.
|
||||
- **[P2-4] `olp-plugin/` `fmtHealth` iterates `providers.status` correctly.** Previously walked `Object.entries(body.providers)` which surfaced `enabled` / `available` / `status` as pseudo-providers (chat output showed `🟢 status` instead of `🟢 anthropic`). Now extracts the real provider map from `body.providers.status` and renders enabled/available counts in a header line + per-provider names with `activeSpawns` when present. Falls back to flat `body.providers.*` for the older OCP shape (backwards compat).
|
||||
- **[P3-5] Stale v0.3.0-era doc strings updated.** README header status line + Implementation Status § now reflect v0.4.0 shipped + Phase 5 open. `server.mjs` startup banner no longer hardcodes "Phase 1 in progress" (now just lists version + provider count — derives accurate state from `VERSION` without future maintenance touch-ups).
|
||||
|
||||
**Phase 4 process learning recorded.** Per Iron Rule 第二律 (evidence over "should work"), every D-day review pass must include at least one runtime smoke against the default production config. The D-day reviewer rubric is updated implicitly — D74 Suite 36 tests pin the wire-contract shape so a future D-day refactoring server payloads can't silently re-break the CLI / plugin / docs.
|
||||
|
||||
- **Test count delta:** 696 (v0.4.0) → 704 (v0.4.1). +8 D74 regression tests in Suite 36.
|
||||
- **Files touched:** `lib/doctor.mjs` (P1-1), `bin/olp.mjs` (P1-1 + P2-3), `bin/olp-connect` (P1-2), `olp-plugin/index.js` (P2-4), `server.mjs` (P3-5 banner), `README.md` (P3-5), `test-features.mjs` (Suite 36 regression), `package.json` (version), `CHANGELOG.md` (this entry).
|
||||
- **Authority:** maintainer independent review of `main` / `v0.4.0` / commit `ee4d945` (2026-05-26 session); Iron Rule 第二律 evidence-over-should-work; CLAUDE.md `release_kit.phase_rolling_mode` cross-Phase discipline ("hotfix to a shipped Phase N deliverable → bump patch, tag, release before next push").
|
||||
|
||||
## v0.4.0 — 2026-05-26
|
||||
|
||||
### Phase 4 — Operator + Client UX (D60 → D73)
|
||||
|
||||
**Overview.** v0.4.0 closes Phase 4 — the "operator + client UX" track that grew OLP from "I built a multi-provider proxy" to "my family can use it without me holding their hand." 5 D-day groups (D60 → D73), ~13 D-days, all under standing-autopilot grant + per-D-day fresh-context opus reviewer per Iron Rule 10. The maintainer-triggered close PR lands all of it under one version tag.
|
||||
|
||||
**Test count: 623 (v0.3.2) → 696 (v0.4.0).** +73 tests across the Phase 4 arc.
|
||||
|
||||
**Strategic decision recorded:** Phase 4 explicitly DEFERS `/v1/messages` (Anthropic-shape entry surface) per ADR 0010 — re-open strictly gated on ADR 0009 P0 success AND maintainer-named family CC user. README posture: Claude Code listed as Not supported as an OLP client; recommended alternative "Cline + OLP" (same fallback chain available, better cross-provider compatibility because OpenAI tool schema is the multi-provider lingua franca).
|
||||
|
||||
**Phase 4 release_kit checklist**
|
||||
|
||||
- [x] All 5 D-day groups landed on main (D60 + D61-D63 + D64-D67 + D68-D70 + D71-D73)
|
||||
- [x] CI green on every D-day merge commit + on this release commit's head
|
||||
- [x] Fresh-context opus reviewer on every implementation D-day group + per-D-day P0/P1/P2 fold-ins where applicable
|
||||
- [x] CHANGELOG "Unreleased" promoted to "## v0.4.0 — 2026-05-26"
|
||||
- [x] `package.json` bumped 0.3.2 → 0.4.0
|
||||
- [x] `CLAUDE.md release_kit.phase_rolling_mode.current_phase` Phase 4 → Phase 5; `current_pre_release_identifier` `0.4.0-phase4` → `0.5.0-phase5`
|
||||
- [x] README § IDE Setup + § Telegram/Discord Usage + § Operator CLI surfaces (env var table extension)
|
||||
- [x] ADR 0010 (Phase 4 charter), ADR 0011 (anonymous-key deployment-context limits), ADR 0002 Amendment 7 (provider doctorChecks contract) all on disk
|
||||
- [ ] Tag pushed (next step in this PR's lifecycle)
|
||||
- [ ] `release.yml` triggered + GitHub Release created (auto on tag push)
|
||||
|
||||
---
|
||||
|
||||
### D60 (PR #40) — Phase 4 charter (ADR 0010) + default port 3456 → 4567
|
||||
|
||||
Opens Phase 4. No functional code change beyond the default port value; substantive D-day work lands D61 onward.
|
||||
|
||||
- **Default `OLP_PORT` changed `3456 → 4567`.** OCP defaults to 3456; OLP and OCP can now co-host on the same machine without `OLP_PORT` env override. Tests use `port: 0` ephemeral — no test-surface impact.
|
||||
- **ADR 0010 (Phase 4 charter) ratified.** Records 5 D-day group scope + explicit DEFER of `/v1/messages` with re-open trigger.
|
||||
- **ADR 0001 + ADR 0008 amendments.** Port-conflict assumption struck-and-amended; § 6.6 default-port reference updated.
|
||||
- **README quick start + Environment Variables table + Migration from OCP § note** updated.
|
||||
|
||||
### D61-D63 (PR #41) — SSE heartbeat + recentErrors[20] + /v0/management/status
|
||||
|
||||
First substantive Phase 4 implementation. 3 D-days bundled per Iron Rule 11 IDR (shared observability surface).
|
||||
|
||||
- **SSE heartbeat** via `streaming.heartbeat_interval_ms` config (default `0` = disabled, matches OCP safe default). When enabled, streaming branch emits `: keepalive\n\n` SSE comment every interval during silent windows; resets on real chunk; cleans up on stream end/error/abort/disconnect. Eager-headers-post-spawn from day one (the OCP `db11105` lesson). `X-Accel-Buffering: no` centralized via new `SSE_DEFAULT_HEADERS` constant. Per-attached-client lifecycle (each tee output gets its own timer).
|
||||
- **`recentErrors[20]` ring buffer.** Module-scope bounded ring, populated from 5 server-side error paths. Filter: only `ProviderError` OR `statusCode >= 500` (401/403 brute-force noise excluded; D61-D63 reviewer P2-1 explicit-401/403-reject fold-in). Path sanitization via OCP `server.mjs:1395` port. In-memory only (per OCP precedent).
|
||||
- **`GET /v0/management/status` combined endpoint** (owner-only_block). Returns `{ ok, version, uptime_ms, uptime_human, started_at, providers, stats, recent_errors, generated_at }`. `_totalRequests` + `_activeRequests` module-scope counters with idempotent-decrement guard.
|
||||
- **Authority:** ADR 0010 § D61-D63; OCP `server.mjs:660-685` (startHeartbeat), `301, 354-358` (ring), `1151-1188` (/status), `1395` (path sanitization), commit `db11105` (eager-headers); ADR 0007 § 7 + ADR 0008 (owner-only_block pattern).
|
||||
- **Test count delta:** 623 → 636 (+13).
|
||||
|
||||
### D64-D67 (PR #42) — `olp` Node CLI + `olp doctor` framework + per-provider doctor checks + ADR 0002 Amendment 7
|
||||
|
||||
Second substantive Phase 4 implementation. 4 D-days bundled — CLI dispatches to doctor; doctor calls plugins via new contract method; ADR amendment authorizes the contract change.
|
||||
|
||||
- **`bin/olp.mjs` Node CLI** with 11 subcommands: `status / health / usage / models / cache / providers / chain show / logs / restart / keys / doctor / help`. Node not bash (per ADR 0010 § Notes — bash's python3 JSON-parsing fragility avoided). Token resolution: `OLP_API_KEY` env → `OLP_OWNER_TOKEN` env → helpful 401 message (filesystem manifest tokens are one-way SHA-256 per ADR 0007 § 5, not recoverable). Output: human-readable ANSI text by default, `--json` for scripting. Exit codes `0=ok / 1=usage / 2=network|HTTP / 3=auth`. Installed via `package.json bin.olp` so `npx olp <subcommand>` works.
|
||||
- **`lib/doctor.mjs` framework** with machine-readable `next_action.ai_executable[]` for AI-driven self-repair. Per check: `{ id, category, async run(): { status: 'ok'|'fail'|'warn', message, evidence? } }`. Built-in checks: `server.running / server.version / config.exists / config.providers_enabled / config.chains_configured / auth.owner_key_exists / system.node_version`. Per-provider checks dynamically collected via the new `provider.doctorChecks()` contract method. `--json` output emits `{ checks, kind: noop|update|fix_oauth|fix_config|fresh_install|fix_server|fix_provider, next_action: { ai_executable, human_required, verify }, summary }`. `--check <id|category>` for tight repair-loop fast paths. Reviewer P2 fold-in: `_shellQuote()` helper hardens `ai_executable[]` against malicious `OLP_HOME` shell-metacharacter injection.
|
||||
- **Per-provider `doctorChecks()`** in anthropic / codex / mistral plugins: CLI-availability probe + auth-presence probe. Each fail returns `evidence.fix_commands` (for `ai_executable[]`) or `evidence.human_required`.
|
||||
- **ADR 0002 Amendment 7** adds OPTIONAL `provider.doctorChecks(): DoctorCheck[]` to the Provider contract — backwards compatible (plugins without it contribute no provider checks).
|
||||
- **`olp restart`** documented caveat (reviewer P2-2): `launchctl kickstart -k` does NOT re-read plist `EnvironmentVariables`; bootout/bootstrap dance noted for env reloads.
|
||||
- **Authority:** ADR 0010 § D64-D67; ADR 0002 Amendment 7 (new); OCP `ocp` bash wrapper + `scripts/doctor.mjs` (port references); 2026-05-26 brainstorm Top 5 inheritance candidate #2.
|
||||
- **Test count delta:** 636 → 658 (+22).
|
||||
|
||||
### D68-D70 (PR #43) — `bin/olp-connect` + `/health.anonymousKey` + ADR 0011
|
||||
|
||||
Third substantive Phase 4 implementation. 3 D-days bundled — olp-connect consumes /health.anonymousKey; both governed by ADR 0011 trusted-LAN invariant.
|
||||
|
||||
- **`bin/olp-connect <host-ip>` (bash, 564 lines)** zero-config LAN client setup. Bash over Node so client machines without recent Node still work. Auto-detects 6 IDEs and configures each: Claude Code (detect + warn — NOT supported per ADR 0010), Cline (print VSCode-settings snippet — manual), Continue.dev (write idempotent `models:` entry to `~/.continue/config.yaml`), Cursor (snippet + WARNING about known base-URL fragility), Aider (write `OPENAI_API_BASE` + `OPENAI_API_KEY` to rc files), OpenClaw (detect + point at `olp-plugin/`). macOS `launchctl setenv` / Linux `~/.config/environment.d/olp.conf` for GUI-app env inheritance. `--dry-run` exercises every state-change site without modifying anything. Idempotent rc-file writes via bracketed `# OLP LAN ... # /OLP LAN` block.
|
||||
- **`/health.anonymousKey` opt-in field** + `auth.advertise_anonymous_key` config. Field appears in both trimmed AND full `/health` payloads when ALL THREE prerequisites hold: `auth.advertise_anonymous_key: true` + `auth.allow_anonymous: true` + at least one non-revoked guest-tier key has `plaintext_advertise` set. Default off — field ABSENT (not null), preserves v0.3.x `/health` shape. Three-prereq gate is graceful-degrade (server warns + boots; request-time re-checks).
|
||||
- **`bin/olp-keys keygen --anonymous --advertise`** new flag. Writes plaintext into manifest `plaintext_advertise` field AND prints WARNING + ADR 0011 pointer. Owner-tier rejected at BOTH CLI and lib layers (defense-in-depth). Reviewer P2-1 fold-in: `listKeys()` strips `plaintext_advertise` alongside `token_hash` — callers wanting the advertised plaintext for the `/health` publication path MUST go through `findAdvertisedKey()` (the only sanctioned read site).
|
||||
- **ADR 0011 (anonymous-key deployment-context)** new ADR codifying the trusted-LAN-only invariant. Threat model explicit; deployment-context table concrete; soft enforcement via startup warn if `BIND_ADDRESS` resolves to public IP AND `advertise_anonymous_key: true`. No hard allowlist (TLS-fronted private networks indistinguishable from public from server's perspective). Re-evaluation triggers named (Cloudflare Tunnel guidance / Phase 5 multi-tenant).
|
||||
- **Authority:** ADR 0010 § D68-D70; ADR 0011 (new); ADR 0007 § 4 (manifest forward-compat unknown fields) + § 7 (identity classes) + § 9 (keygen flow); OCP `ocp-connect` (port reference); 2026-05-26 brainstorm Top 5 inheritance candidate #3.
|
||||
- **Test count delta:** 658 → 672 (+14).
|
||||
|
||||
### D71-D73 (PR #44) — `olp-plugin/` (OpenClaw /olp Telegram+Discord) + `docs/integrations/*.md` + README cross-refs
|
||||
|
||||
Final Phase 4 substantive D-day group. 3 D-days bundled — plugin consumes existing endpoints; integration docs reference plugin + olp CLI + olp-connect together.
|
||||
|
||||
- **`olp-plugin/` OpenClaw gateway plugin** (482 lines). Port of OCP `ocp-plugin/index.js` minus mutations. Subcommand parity with `olp` CLI: `/olp status / usage / cache` (owner-only) + `/olp health / models / providers / chain show / doctor / help` (informational). **Explicitly NOT ported** for security: `/olp keys keygen` (chat = brute-force-prone), `/olp keys revoke` (mutation), `/olp restart` (misclick risk), `/olp logs` (PII risk). Port resolution: `OLP_PROXY_URL` env → `OLP_PORT` env → plugin config `proxyUrl` → `http://127.0.0.1:4567` (D60 default). Output: Telegram/Discord monospace code block with status icons (🟢🟡🔴). Long responses truncated for 4096-char Telegram limit. No npm deps (OpenClaw provides Telegram/Discord transport).
|
||||
- **`docs/integrations/*.md` bundle** (6 pages + index). Per-IDE setup docs with status icons: Continue.dev ✅, Cline ✅ (cites Cline issue #7128 base-URL UI bug), Cursor ⚠️ (documented base-URL fragility), Aider ✅, **Claude Code ❌** (Anthropic wire format only; recommended alternative "Cline + OLP" per ADR 0010 § /v1/messages defer), OpenClaw ✅. Each ~60-120 lines: status / quick setup / known issues / OLP-specific notes / test-it command. `docs/integrations/README.md` is the index.
|
||||
- **README updates.** New § "IDE Setup" linking `docs/integrations/README.md`. New § "Telegram / Discord Usage" with install + configure + restart + use. Quick Start mentions `olp-connect <ip>` as family-onboarding command. `package.json files` field extended to include `olp-plugin/` so the published tarball ships it.
|
||||
- **Authority:** ADR 0010 § D71-D73; OCP `ocp-plugin/index.js` (port reference); 2026-05-26 brainstorm prior-art survey IDE-specific quirks.
|
||||
- **Test count delta:** 672 → 696 (+24).
|
||||
|
||||
---
|
||||
|
||||
**Phase 4 close authority chain:** ADR 0010 (charter); CLAUDE.md `release_kit.phase_rolling_mode` (close trigger = explicit maintainer action — fired by maintainer 2026-05-26); standing autopilot grant covering D-day-by-D-day execution; 5 fresh-context opus reviewer passes (one per D-day group); 696/696 tests pass on this release commit head.
|
||||
|
||||
## v0.3.2 — 2026-05-25
|
||||
|
||||
### Post-Phase-3 cleanup batch #2 — streaming-path singleflight + TOCTOU close (D57 + D58 + D59)
|
||||
|
||||
Patch release closing v1.x roadmap #1 end-to-end. The cache layer's D4 singleflight (one spawn per identical concurrent request) was fully wired on the buffered path since v0.1 but NOT on the streaming path — N concurrent identical streaming requests each spawned their own CLI process. v0.3.2 ships the streaming sibling: tee fan-out, late-joiner replay, per-client backpressure, AbortController propagation, and TOCTOU close. 3 D-day commits (D57 + D58 + D59); ADR 0005 Amendment 8 §§1–14 implemented.
|
||||
|
||||
- **D57** (PR #36) — **cache layer.** New `cacheStore.getOrComputeStreaming(keyId, cacheKey, sourceFactory, opts) → { stream, isFirst, role }` mirroring `getOrCompute` on the streaming side. Internals: `_streamingInflight: Map<compositeKey, StreamingInflightEntry>` (composite key `keyId + '\0' + cacheKey`) with synchronous check+insert atomicity (closes TOCTOU per ADR 0005 Amendment 8 §1, §6); single-reader tee fan-out across all attached clients; late-joiner replay buffer (synchronous drain on attach; `STREAM_BACKPRESSURE` terminator if drain or replay-truncation would corrupt); per-client backpressure (`PER_CLIENT_QUEUE_CAP = 1 MB`, overridable via opts); accumulated replay cap (`ACCUMULATED_REPLAY_CAP = 10 MB`, mirrors D23 cache-entry cap); AbortController fires source-iterator return when all clients disconnect. New `'STREAM_BACKPRESSURE'` entry in `PROVIDER_ERROR_CODES` — NOT a hard trigger (whitelist-only `HARD_TRIGGER_CODES`). Suite 27 = 12 unit tests.
|
||||
- **D58** (PR #37) — **server wiring.** Streaming branch in `server.mjs` swapped from the peek+spawn pattern to `cacheStore.getOrComputeStreaming(...)`. `tryAcquireSpawn`/`releaseSpawn` moved INSIDE the `sourceFactory` closure per ADR 0005 Amendment 8 §7 (only the first caller acquires; attached joiners share the slot; release fires exactly once on source completion/error/abort). `CONCURRENCY_LIMIT` thrown by the factory triggers fallthrough to the buffered path (preserves today's behaviour). New `X-OLP-Streaming-Inflight: source | attached` HTTP header annotates per-response role (§11). New `cache_status: 'streaming_attached'` audit value tracks the singleflight win. `lib/audit-query.mjs` aggregate APIs (`aggregateRequests`, `cacheHitRateWindow`) extended with `cache_streaming_attached` / `streaming_attached` fields so the cache_status breakdown reconciles. `res.on('close')` propagates client disconnect into the tee's `attachedClients` accounting (§9). D16 truncated-not-cached invariant preserved via server-layer `cacheStore.delete` on stop-less exhaustion (the cache layer is IR-agnostic and writes accumulatedChunks on any source exhaustion; the IR-aware server deletes the entry when the source returned without a `{type:'stop'}` chunk). Suite 28 = 8 HTTP integration tests.
|
||||
- **D59** (PR #38) — **docs polish.** README § Known limitations bullet inverted to ✅ shipped marker. `docs/v1x-roadmap.md` #1 rewritten to closed state with 3-D-day breakdown. #6 (streaming SPAWN_FAILED salvage) unbundled from #1 because the tee architecture as implemented does not carry salvage semantics. Issue #16 closed with refs to PRs #36 / #37 / #38.
|
||||
- **Test count:** 603 (v0.3.1) → 623 (v0.3.2). +20 streaming-SF tests (Suite 27 = 12 unit, Suite 28 = 8 HTTP integration).
|
||||
- **Deferred sub-items (not blocking #1 closure):** (a) `X-OLP-Streaming-Inflight: solo` wire value not emitted — observable only post-stream via `streaming_inflight_source_done` log event's `attached_count: 0`. Future ADR amendment may expose via HTTP trailer. (b) `streaming_inflight_join` log event not emitted from the cache-layer `_attachClient` path because provider/model context lives in the sourceFactory closure (server-layer concern). (c) `isFirst` field returned by `getOrComputeStreaming` is unused by server.mjs (`role` supersedes); could be removed in a future cache-layer API cleanup.
|
||||
- **Authority:** ADR 0005 Amendment 8 (design ratified at D42 2026-05-25; implementation gated on maintainer "go" — fired 2026-05-25 post-v0.3.1). `docs/v1x-roadmap.md` #1 (closed). GitHub issue #16 (closed). ADR 0002 Amendment 6 (D38 `tryAcquireSpawn`/`releaseSpawn` semantics, now invoked from sourceFactory closure).
|
||||
|
||||
**Patch-release classification.** Per `release_kit.phase_rolling_mode` cross-Phase discipline + maintainer release-cut decision (this session, 2026-05-25): the new wire surface (`X-OLP-Streaming-Inflight` header + `streaming_attached` cache_status) is semver-wise a minor bump, but this is roadmap-cleanup work — NOT Phase 4 product scope. The reserved `0.4.0` identifier stays for the formal Phase 4 close. v0.3.2 ships as a patch under the Phase 4 pre-release banner. Tag push triggers `release.yml`.
|
||||
|
||||
## v0.3.1 — 2026-05-25
|
||||
|
||||
|
||||
@@ -135,7 +135,7 @@ release_kit:
|
||||
# This overlay is the authoritative source. If Iron Rule 5 appears to be silently
|
||||
# violated (no version bump after many D-day pushes), check this section first
|
||||
# before filing a compliance finding.
|
||||
current_phase: Phase 4
|
||||
current_pre_release_identifier: "0.4.0-phase4"
|
||||
current_phase: Phase 7 closed at v0.7.0 (2026-05-29); Phase 8 not yet scoped
|
||||
current_pre_release_identifier: "0.7.0"
|
||||
phase_close_trigger: explicit maintainer action (not automated)
|
||||
```
|
||||
|
||||
@@ -1,62 +1,257 @@
|
||||
# OLP — Open LLM Proxy
|
||||
|
||||
A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing, automatic fallback, content-addressed caching — so your IDEs and family clients keep working as long as *any* of your subscriptions has quota left.
|
||||
A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing + fallback + content-addressed caching. Your IDEs and family clients keep working as long as **any** of your subscriptions has quota left.
|
||||
|
||||
> **Status:** v0.3.0 shipped (2026-05-25) — Phase 1 multi-provider proxy core (v0.1.0 + v0.1.1) + Phase 2 multi-key auth + audit + owner gating + keygen CLI (v0.2.0) + Phase 3 Dashboard + audit query layer + daily audit rotation (v0.3.0). Phase 4 (per-key per-provider auth + audit retention + SQLite hybrid + provider-cost weights) is the next planned milestone. Sections marked _placeholder_ land alongside the relevant phase of work (see [phase plan](#phase-plan)).
|
||||
> **Status:** v0.5.1 shipped, 759+ tests. Phase 5 (Quota Probes + Dashboard Enrichment) closed; Phase 6 next. Coming from [OCP](https://github.com/dtzp555-max/ocp)? See [§ Migration from OCP](#migration-from-ocp).
|
||||
|
||||
---
|
||||
|
||||
## Why OLP
|
||||
## What you get
|
||||
|
||||
On 2026-05-14, Anthropic announced (effective 2026-06-15) that `claude -p`, the Agent SDK, and third-party agent traffic move out of the Pro/Max subscription pool into a separate fixed monthly Agent SDK Credit pool. [OCP](https://github.com/dtzp555-max/ocp), OLP's predecessor, was a proxy around a single CLI — its core assumption was *"subscription = unlimited within rate limits"*. That assumption breaks for Anthropic on the effective date.
|
||||
|
||||
The structural response is to stop relying on one provider's subscription terms remaining favourable. OLP spreads risk across multiple providers whose subscriptions still include CLI/programmatic use, routes intelligently between them, and caches aggressively so every request that does spawn a CLI counts.
|
||||
|
||||
OLP is **not**: a commercial multi-tenant SaaS; an enterprise gateway competing with LiteLLM / OpenCode / CLIProxyAPI on breadth; a model-capability router ("route to the smartest model" — you pick the model); a conversation-state store (your client handles that).
|
||||
|
||||
See [`ALIGNMENT.md`](./ALIGNMENT.md) for OLP's constitution and [`docs/adr/`](./docs/adr/) for the founding ADRs.
|
||||
- **OpenAI-compatible** `/v1/chat/completions` endpoint — any IDE that speaks OpenAI (Cline / Continue.dev / Cursor / Aider) plugs in
|
||||
- **Multi-provider chain** — primary fails / quota dies → automatically falls back to the next provider (anthropic ↔ codex ↔ mistral by default; risk-tier framework guards which ones get enabled)
|
||||
- **Content-addressed cache** — repeat requests don't re-spawn the CLI; streaming requests dedup via singleflight tee
|
||||
- **Multi-key auth** — owner key with full visibility, family-member keys with per-key audit log + per-provider scoping
|
||||
- **Telegram / Discord** `/olp` slash commands (read-only — for "is OLP up?" checks from anywhere)
|
||||
- **AI-driven self-repair** — `olp doctor --json` emits machine-readable `next_action.ai_executable[]` so a Claude Code / Cursor / Copilot session can fix install issues for you (see [§ Install with your AI](#install-with-your-ai-the-fast-path))
|
||||
- **Observability** — owner-only `/dashboard` (live Claude.ai-style plan-usage rows / 24h stats / 30d spend trend / top fallback chains)
|
||||
- **Plan-usage probe** (Phase 5, v0.5.0) — opt-in per-provider quota probe for Anthropic Pro/Max subscriptions; parses the canonical `anthropic-ratelimit-unified-*` response headers, surfaces 5-hour + 7-day utilization with reset countdowns. See [§ Plan Usage](#plan-usage-live-quota-probe).
|
||||
|
||||
---
|
||||
|
||||
## Quick Start
|
||||
## Tool execution model
|
||||
|
||||
_placeholder — lands with Phase 1._
|
||||
OLP is a **chat/completion proxy**, not a tool runtime. It forwards messages between your client and a provider's LLM and returns the response. It does **not** execute tools (shell commands, filesystem reads, web fetches) on your behalf, and it has no plans to.
|
||||
|
||||
Anticipated shape:
|
||||
When an agentic client (Cline / Cursor / Continue.dev / Aider / Hermes Agent / OpenClaw) needs to call a tool, that tool runs **on the client's host**. The client sends the tool's output back as a follow-up message. OLP sees only the message stream — never an open file handle, an executed command, or a fetched URL.
|
||||
|
||||
Why this boundary matters:
|
||||
|
||||
- **Multi-tenant safety.** A misbehaving prompt cannot use OLP to read files belonging to another OLP key holder. The threat surface is bounded to "what the model can say in a message" — not "what the model can do on the server."
|
||||
- **Stateless operation.** OLP runs the same code path for every request, regardless of which client is calling. Session state, tool state, and conversational memory all live in the client. See [`AGENTS.md`](./AGENTS.md) § "No conversation state".
|
||||
- **Provider-CLI honesty.** OLP spawns provider CLIs (`claude`, `codex`, `vibe`) to talk to upstream APIs and translates wire formats via the IR. It does not extend those CLIs with new tools or capabilities — see [`ALIGNMENT.md`](./ALIGNMENT.md) Rule 2 (No Invention).
|
||||
|
||||
A few clients (notably OpenClaw in certain configurations) can be wired to route their tool calls *through* the OLP server host rather than executing them locally. This is a client configuration choice, not an OLP feature, and it produces surprising self-check results (the agent describes the OLP server, not your machine). See [§ Known limitations](#known-limitations) for the integrator-level guidance.
|
||||
|
||||
For the multi-tenant isolation story, [ADR 0014 Amendment 1](./docs/adr/0014-sandbox-runtime-integration.md) defines a four-layer architecture: each provider-CLI spawn gets a per-request ephemeral `$HOME` (`/tmp/olp-spawn/<keyId>/<reqId>/home/`) with credential files symlinked in, plus per-provider tool-hardening (anthropic's `--system-prompt` suppresses Read/Bash tool descriptions; codex defaults to `--sandbox read-only`). The canonical contract lives in [ADR 0002 Amendment 9](./docs/adr/0002-plugin-architecture.md) (Provider ISOLATION contract) + [ADR 0014 Amendment 1](./docs/adr/0014-sandbox-runtime-integration.md). The full "Security Model" reference will land in a Phase 7 close PR (Task #10).
|
||||
|
||||
---
|
||||
|
||||
## Install with your AI (the fast path)
|
||||
|
||||
If the manual steps feel like a lot, paste this verbatim into your AI coding assistant (Claude Code / Cursor / Copilot / Aider). It walks you through everything:
|
||||
|
||||
```
|
||||
I want to install OLP (Open LLM Proxy) on this machine. The repo is at
|
||||
https://github.com/dtzp555-max/olp and ships these tools that you can use:
|
||||
|
||||
1. `git clone` the repo to ~/olp
|
||||
2. README.md § "Manual install" has the sequenced steps. Follow them in order.
|
||||
3. After EACH step, run `node ~/olp/bin/olp.mjs doctor --json` (NOT `npx olp`
|
||||
— the package isn't on npm). Parse the JSON. The `kind` field tells you
|
||||
the state: noop / fresh_install / fix_server / fix_oauth / fix_provider /
|
||||
fix_config.
|
||||
4. If `kind != noop`, run the commands listed in `next_action.ai_executable[]`
|
||||
verbatim. Then re-run doctor to verify.
|
||||
5. STOP and ask me only when `next_action.human_required[]` is non-empty.
|
||||
That's where I need to do a browser OAuth flow you can't do for me.
|
||||
|
||||
The provider CLIs OLP spawns (claude / codex / vibe) need their own one-time
|
||||
OAuth — those are the only steps I personally have to do (Claude.ai login,
|
||||
ChatGPT login, Mistral API key). Everything else (clone, npm install of the
|
||||
provider CLIs, owner-key generation, config.json bootstrap, server start) is
|
||||
in your `ai_executable[]` and you should run it without asking.
|
||||
|
||||
Begin.
|
||||
```
|
||||
|
||||
Then sit back and respond when it asks for OAuth confirmation. This pattern works because `olp doctor` is purpose-built for AI consumption — every failure mode has a shell-executable repair command AND a human-required step listed separately.
|
||||
|
||||
---
|
||||
|
||||
## Manual install (5-10 min)
|
||||
|
||||
### 0. Prerequisites
|
||||
|
||||
- **Node.js ≥ 18.** Verify: `node --version`
|
||||
- **The provider CLIs you want OLP to spawn.** Install whichever you'll actually use:
|
||||
|
||||
| Provider | Install | Subscription |
|
||||
|---|---|---|
|
||||
| `anthropic` (`claude -p`) | `npm install -g @anthropic-ai/claude-code` | Claude Pro/Max (OAuth) |
|
||||
| `openai` (`codex exec`) | `npm install -g @openai/codex` | ChatGPT Plus/Pro (OAuth) or OpenAI API key |
|
||||
| `mistral` (`vibe --prompt`) | follow the `vibe` install docs | Le Chat Pro API key |
|
||||
|
||||
You only need to install the ones you'll route to. Single-provider OLP works fine.
|
||||
|
||||
### 1. Clone and verify the test suite
|
||||
|
||||
```bash
|
||||
# install
|
||||
npm install -g @dtzp555-max/olp
|
||||
|
||||
# run setup (writes ~/.olp/config.json, asks which providers to enable)
|
||||
olp setup
|
||||
|
||||
# start the proxy (default port 3456 — same as OCP if you migrate)
|
||||
olp start
|
||||
|
||||
# point your IDE at http://localhost:3456/v1/chat/completions with the OLP API key from `olp keys list`.
|
||||
git clone https://github.com/dtzp555-max/olp.git ~/olp
|
||||
cd ~/olp
|
||||
npm test # 714+ tests, ~5s, no external deps
|
||||
```
|
||||
|
||||
(If `npm test` fails here, stop — that means your Node version or the repo state is broken. Don't proceed to step 2.)
|
||||
|
||||
### 2. Bootstrap the owner key
|
||||
|
||||
The owner key is what you (and `olp-connect`) use to authenticate to OLP. Default config has `auth.allow_anonymous: false`, so you need a key BEFORE the server starts accepting requests.
|
||||
|
||||
```bash
|
||||
node ~/olp/bin/olp-keys.mjs keygen --owner --name=$(whoami)-laptop
|
||||
# Prints the plaintext token ONCE. Copy it now — you can't recover it later.
|
||||
# Example: olp_l23-PN46tDljmPATV94-KfOgOBO0Ed8theVjTdAgQoY
|
||||
```
|
||||
|
||||
Export it so the CLI subcommands can use it:
|
||||
|
||||
```bash
|
||||
export OLP_API_KEY=olp_l23-PN46... # paste your real token
|
||||
```
|
||||
|
||||
(Add to `~/.bashrc` / `~/.zshrc` to persist.)
|
||||
|
||||
### 3. Authenticate the providers (one-time OAuth)
|
||||
|
||||
Run each provider's own login flow. OLP's anthropic / openai / mistral plugins spawn these CLIs and reuse their cached credentials — OLP itself never touches the OAuth dance.
|
||||
|
||||
```bash
|
||||
# Anthropic (Claude Pro/Max subscription)
|
||||
claude setup-token
|
||||
# Opens a TUI / prints a URL. Authorize in browser. Paste the returned code.
|
||||
# Result: ~/.claude/.credentials.json
|
||||
|
||||
# OpenAI (ChatGPT subscription)
|
||||
codex login --device-auth
|
||||
# Prints a https://auth.openai.com/codex/device URL + 10-char code.
|
||||
# Open URL in browser, enter code, authorize.
|
||||
# Result: ~/.codex/auth.json
|
||||
|
||||
# Mistral (Le Chat API key)
|
||||
export MISTRAL_API_KEY=sk-... # add to ~/.bashrc to persist
|
||||
```
|
||||
|
||||
### 4. Write a minimum config
|
||||
|
||||
`~/.olp/config.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"auth": {
|
||||
"allow_anonymous": false,
|
||||
"owner_only_endpoints": [
|
||||
"/health",
|
||||
"/v0/management/dashboard-data",
|
||||
"/v0/management/quota",
|
||||
"/v0/management/status",
|
||||
"/cache/stats",
|
||||
"/dashboard"
|
||||
],
|
||||
"fallback_detail_header_policy": "owner_only"
|
||||
},
|
||||
"providers": {
|
||||
"enabled": { "anthropic": true, "openai": true }
|
||||
},
|
||||
"routing": {
|
||||
"chains": {
|
||||
"claude-sonnet-4-6": [
|
||||
{ "provider": "anthropic", "model": "claude-sonnet-4-6" },
|
||||
{ "provider": "openai", "model": "gpt-5.5" }
|
||||
],
|
||||
"gpt-5.5": [{ "provider": "openai", "model": "gpt-5.5" }]
|
||||
}
|
||||
},
|
||||
"streaming": { "heartbeat_interval_ms": 15000 }
|
||||
}
|
||||
```
|
||||
|
||||
(Enable only the providers you actually authenticated in step 3. Chains map `<your-IDE's-requested-model>` → ordered list of `{provider, model}` hops; the chain's per-hop `model` is what gets passed to that provider's CLI.)
|
||||
|
||||
### 5. Start the server
|
||||
|
||||
```bash
|
||||
cd ~/olp
|
||||
npm start
|
||||
# OLP v0.4.3 listening on :4567 (2 providers enabled)
|
||||
```
|
||||
|
||||
### 6. Smoke-test
|
||||
|
||||
```bash
|
||||
curl -H "Authorization: Bearer $OLP_API_KEY" http://localhost:4567/health | jq
|
||||
# Expect: {ok: true, providers: {enabled: 2, status: {anthropic: {ok: true...}, openai: {ok: true...}}}}
|
||||
|
||||
node ~/olp/bin/olp.mjs doctor
|
||||
# Expect: "9 of 9 checks passed", kind=noop
|
||||
```
|
||||
|
||||
### 7. Point your IDE at OLP
|
||||
|
||||
```
|
||||
OPENAI_BASE_URL=http://localhost:4567/v1
|
||||
OPENAI_API_KEY=$OLP_API_KEY
|
||||
```
|
||||
|
||||
Per-IDE configuration details: [`docs/integrations/`](./docs/integrations/README.md).
|
||||
|
||||
---
|
||||
|
||||
## Family / LAN setup
|
||||
|
||||
To let other devices on your home network use the same OLP server, you need TWO things:
|
||||
|
||||
1. **Bind to the LAN interface** (not just loopback). On the SERVER:
|
||||
|
||||
```bash
|
||||
OLP_BIND=0.0.0.0 npm start # or your specific LAN IP, e.g. 192.168.1.10
|
||||
```
|
||||
|
||||
Default is `127.0.0.1` (loopback only). See [ADR 0011 § Deployment configurations](./docs/adr/0011-anonymous-key-deployment-context.md#deployment-configurations-d76-amendment-2026-05-26) for the trust-context table — **never set `OLP_BIND=0.0.0.0` on a public-internet-facing host** (use a tunnel like Tailscale instead).
|
||||
|
||||
2. **Onboard each family member's device** from THEIR machine:
|
||||
|
||||
```bash
|
||||
# Pinned to a known-good release (recommended — survives GitHub raw CDN cache hiccups):
|
||||
bash <(curl -fsSL https://raw.githubusercontent.com/dtzp555-max/olp/v0.4.4/bin/olp-connect) <olp-host-ip>
|
||||
|
||||
# OR latest from main (use after v0.4.4 + once you trust head):
|
||||
bash <(curl -fsSL https://raw.githubusercontent.com/dtzp555-max/olp/main/bin/olp-connect) <olp-host-ip>
|
||||
```
|
||||
|
||||
Detects Cline / Continue.dev / Cursor / Aider / OpenClaw locally and writes per-tool config pointing at your OLP host. Requires `python3` on the client. Prompts for the OLP API key — OR, if the server has `auth.advertise_anonymous_key: true` AND a key was created with `olp-keys keygen --anonymous --advertise`, picks the token up from `/health.anonymousKey` (zero out-of-band paste). See [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md) for the trusted-LAN-only invariant.
|
||||
|
||||
Per-IDE setup details: [`docs/integrations/`](./docs/integrations/README.md). Telegram / Discord `/olp` slash command setup: [§ Telegram / Discord Usage](#telegram--discord-usage).
|
||||
|
||||
---
|
||||
|
||||
## Supported Providers
|
||||
|
||||
Source of truth: [`models-registry.json`](./models-registry.json). This table is regenerated from the registry per the [`release_kit`](./CLAUDE.md) overlay; do not edit it out of sync.
|
||||
Source of truth: [`models-registry.json`](./models-registry.json). Per-provider columns are sourced from the registry's `providers.<key>` block (model metadata + tier) and `quota_probe.<key>` block (D81+; probe status / reason / source). This table is regenerated from the registry per the [`release_kit`](./CLAUDE.md) overlay; do not edit it out of sync.
|
||||
|
||||
OLP distinguishes **Candidate Providers** (declared as intended, not yet pinned) from **Enabled Providers** (authority pin filled + plugin landed + Phase audit passed). The v0.1 founding commit ships **zero Enabled Providers** — enablement is a Phase audit deliverable, not a bootstrap claim. See [`ALIGNMENT.md` § Provider Inventory](./ALIGNMENT.md) for the transition gate.
|
||||
|
||||
### Candidate Providers
|
||||
|
||||
| Provider key | CLI | Subscription / auth | Anticipated Tier | Anticipated Phase |
|
||||
|---|---|---|---|---|
|
||||
| `anthropic` | `claude -p` | Pro / Max OAuth (pre-2026-06-15); Agent SDK Credit pool after | D (re-eval post-2026-06-15) | Phase 1 |
|
||||
| `openai` | `codex exec --json` | ChatGPT Pro OAuth or API key | D | Phase 2 |
|
||||
| `mistral` | `vibe --prompt --output json` | Le Chat Pro API key | D | Phase 3 |
|
||||
| `grok` | `grok -p --output-format streaming-json` | xAI Build `xai-...` API key | C | Phase 8+ |
|
||||
| `kimi` | `kimi -p --output-format stream-json` | Moonshot Kimi API key | C | Phase 8+ |
|
||||
| `minimax` | TBD | MiniMax Token Plan (¥29+/mo) | B | Phase 8+ |
|
||||
| `glm` | TBD | Zhipu Coding Plan ($10+/mo) | B | Phase 8+ |
|
||||
| `qwen` | TBD | Alibaba Coding Plan ($50/mo) | B | Phase 8+ |
|
||||
| Provider key | CLI | Subscription / auth | Quota probe (v0.5.0+) | Anticipated Tier | Anticipated Phase |
|
||||
|---|---|---|---|---|---|
|
||||
| `anthropic` | `claude -p` | Pro / Max OAuth (pre-2026-06-15); Agent SDK Credit pool after | ✅ Live (13 `anthropic-ratelimit-unified-*` headers; opt-in via `quota_probe_enabled`) | D (re-eval post-2026-06-15) | Phase 1 |
|
||||
| `openai` | `codex exec --json` | ChatGPT Pro OAuth or API key | ❌ Not available (no public quota API) — audit-derived spend tracking only | D | Phase 2 |
|
||||
| `mistral` | `vibe --prompt --output json` | Le Chat Pro API key | ❌ Not implemented at v0.5.0 — no public quota endpoint accessible to Vibe / Le Chat member / La Plateforme API keys per D84 spike 2026-05-26. Mistral's [Admin API](https://docs.mistral.ai/admin/security-access/admin-api) does expose billing / usage queries but requires an org-admin scope (out of scope for OLP family-tier deployment). Audit-derived spend tracking only at v0.5.0. | D | Phase 3 |
|
||||
| `grok` | `grok -p --output-format streaming-json` | xAI Build `xai-...` API key | TBD (Phase 8+) | C | Phase 8+ |
|
||||
| `kimi` | `kimi -p --output-format stream-json` | Moonshot Kimi API key | TBD (Phase 8+) | C | Phase 8+ |
|
||||
| `minimax` | TBD | MiniMax Token Plan (¥29+/mo) | TBD (Phase 8+) | B | Phase 8+ |
|
||||
| `glm` | TBD | Zhipu Coding Plan ($10+/mo) | TBD (Phase 8+) | B | Phase 8+ |
|
||||
| `qwen` | TBD | Alibaba Coding Plan ($50/mo) | TBD (Phase 8+) | B | Phase 8+ |
|
||||
|
||||
**Anthropic models (sourced from `models-registry.json`):**
|
||||
|
||||
| Model ID | Display name | Context window | Notes |
|
||||
|---|---|---|---|
|
||||
| `claude-opus-4-8` | Claude Opus 4.8 | 200 000 | Newest opus; `opus` alias points here |
|
||||
| `claude-opus-4-7` | Claude Opus 4.7 | 200 000 | Still callable by literal id |
|
||||
| `claude-sonnet-4-6` | Claude Sonnet 4.6 | 200 000 | `sonnet` + `claude` aliases point here |
|
||||
| `claude-haiku-4-5` | Claude Haiku 4.5 | 200 000 | `haiku` alias points here |
|
||||
|
||||
**Risk tier guide.** D = permissive / safe (eligible for default-enabled); C = tightening signal, no enforcement history (opt-in); B = service-level key revocation risk (opt-in + consent); A = excluded by default (cannot be opt-in enabled). Tier B providers prompt for explicit consent on first enable and record consent in `~/.olp/config.json`. See [`ALIGNMENT.md` § Risk Tier Framework](./ALIGNMENT.md#risk-tier-framework).
|
||||
|
||||
@@ -66,12 +261,19 @@ OLP distinguishes **Candidate Providers** (declared as intended, not yet pinned)
|
||||
|
||||
## Configuration
|
||||
|
||||
_placeholder — full configuration reference lands with Phase 4 (fallback engine)._
|
||||
|
||||
OLP reads its config from `~/.olp/config.json`. The minimum useful shape:
|
||||
OLP reads `~/.olp/config.json` at startup. § "[Manual install § Step 4](#4-write-a-minimum-config)" above has a working minimum example. The full schema:
|
||||
|
||||
```json
|
||||
{
|
||||
"auth": {
|
||||
"allow_anonymous": false,
|
||||
"owner_only_endpoints": ["/health", "/dashboard", "/v0/management/..."],
|
||||
"advertise_anonymous_key": false,
|
||||
"fallback_detail_header_policy": "owner_only"
|
||||
},
|
||||
"providers": {
|
||||
"enabled": { "<provider-key>": true }
|
||||
},
|
||||
"routing": {
|
||||
"chains": {
|
||||
"<requested-model>": [
|
||||
@@ -82,13 +284,78 @@ OLP reads its config from `~/.olp/config.json`. The minimum useful shape:
|
||||
"soft_triggers": {
|
||||
"<provider-key>": { "<trigger>": <threshold> }
|
||||
}
|
||||
},
|
||||
"streaming": {
|
||||
"heartbeat_interval_ms": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
> **Note:** `routing.soft_triggers` thresholds are parsed and stored but have **no runtime effect at v0.1** — the quota polling path (`quotaStatus()` per hop) is deferred to v1.x per [ADR 0004 Amendment 2](./docs/adr/0004-fallback-engine.md#amendment-2--2026-05-24-soft-triggers-deferred-to-v1x-d22). The evaluation logic exists and is tested; only the production data ingestion path is deferred.
|
||||
Field guide:
|
||||
|
||||
Trigger types, fallback safety, idempotency rules, and the full example config land here when Phase 4 ships. See [ADR 0004 (Fallback Engine Semantics & Safety)](./docs/adr/0004-fallback-engine.md) for the design.
|
||||
- **`auth.allow_anonymous`** — default `false`. When false, every request needs a Bearer token; when true, anonymous-tier requests succeed (ADR 0007 § 7). Production posture is `false`.
|
||||
- **`auth.owner_only_endpoints`** — list of endpoints that REQUIRE owner-tier auth (non-owner returns 401). The defaults above are minimum sane for production.
|
||||
- **`auth.advertise_anonymous_key`** — default `false`. When true (+ `allow_anonymous: true` + a key created with `olp-keys keygen --anonymous --advertise`), `/health.anonymousKey` exposes the plaintext token so `olp-connect <ip>` is zero-config. **Trusted-LAN only** — see [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md).
|
||||
- **`auth.fallback_detail_header_policy`** — controls `X-OLP-Fallback-Detail` response header emission. `owner_only` (default) only shows tuples to owner identity; debug surface to LAN family without leaking to anonymous.
|
||||
- **`providers.enabled`** — flip a provider plugin on. Only enable providers whose CLI you've authenticated; OLP doesn't do its own OAuth.
|
||||
- **`routing.chains`** — keyed by the model name your IDE / client requests. Each entry is an ordered list of fallback hops; each hop's `model` is what gets passed to that provider's CLI. F7 fix (D75) — the hop-level `model` field finally overrides the IR's request model during cross-provider fallback.
|
||||
- **`routing.soft_triggers`** — parsed and stored but **inert at v0.4.x** — the `quotaStatus()` polling data path is deferred to v1.x per [ADR 0004 Amendment 2](./docs/adr/0004-fallback-engine.md#amendment-2--2026-05-24-soft-triggers-deferred-to-v1x-d22). Startup emits a warn if non-empty so the inert state is visible.
|
||||
- **`streaming.heartbeat_interval_ms`** — default `0` (disabled). Set > 0 (e.g. `15000`) to emit SSE keepalive frames during silent windows. Required behind reverse proxies (nginx / Cloudflare Tunnel / Tailscale Funnel) with 60s idle aborts.
|
||||
|
||||
See [ADR 0004 (Fallback Engine)](./docs/adr/0004-fallback-engine.md), [ADR 0007 (Multi-key auth)](./docs/adr/0007-multi-key-auth.md), [ADR 0010 (Phase 4 charter)](./docs/adr/0010-phase-4-charter-operator-and-client-ux.md), [ADR 0011 (Anonymous-key deployment)](./docs/adr/0011-anonymous-key-deployment-context.md).
|
||||
|
||||
---
|
||||
|
||||
## Plan Usage (live quota probe)
|
||||
|
||||
OLP v0.5.0+ surfaces live subscription quota for Anthropic Pro/Max subscribers on the owner-only `/dashboard`. Per-provider rows show 5-hour and 7-day utilization bars with reset countdowns, status badges, representative-claim hints, and a manual refresh button. The panel auto-refreshes every 60 seconds and pauses when the tab is hidden.
|
||||
|
||||

|
||||
|
||||
### How it works
|
||||
|
||||
The probe issues a minimal `POST /v1/messages` to `api.anthropic.com` (max_tokens: 1) using the same OAuth token Claude Code uses for `claude -p`. The body is discarded; only the 13 `anthropic-ratelimit-unified-*` response headers are parsed (5h/7d utilization + reset, status, representative-claim, fallback-percentage, overage status + disabled reason). Results cache for 5 minutes; refresh failures fall back to the previous cache marked `stale: true` while exponential backoff (60s → 3600s) protects against hammering the API.
|
||||
|
||||
See [ADR 0002 § Amendment 8](./docs/adr/0002-plugin-architecture.md), [ADR 0012 (Phase 5 charter)](./docs/adr/0012-phase-5-charter-quota-probes-dashboard.md), [ADR 0013 (OAuth READ-ONLY consumption + schema-drift mitigation)](./docs/adr/0013-oauth-read-only-consumption-and-schema-drift.md), and the schema pin at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`.
|
||||
|
||||
### Enabling the probe
|
||||
|
||||
The probe is **opt-in** (default off) per ADR 0013 Rule 4 — a fresh OLP install on a machine without OAuth credentials should not bombard `api.anthropic.com` with 401-bound probes. To enable, add to `~/.olp/config.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"providers": {
|
||||
"anthropic": {
|
||||
"enabled": true,
|
||||
"quota_probe_enabled": true
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The probe reads the OAuth token from (in order): `CLAUDE_CODE_OAUTH_TOKEN` env var, `~/.claude/.credentials.json`, macOS Keychain entry `"Claude Code-credentials"`. Make sure Claude Code is logged in (`claude setup-token` or equivalent) before opting in.
|
||||
|
||||
`olp doctor` adds a `anthropic.quota_probe_reachable` check when the probe is enabled. The check has `category: 'provider'`, so any failure (401/403 token-expiry, 429 rate-limit, network error) discriminates to `kind: fix_provider`. The `human_steps` recovery recipe inside the check distinguishes the underlying cause (re-login via `claude setup-token` for auth failures vs wait-and-retry for rate-limit) — the discriminator is uniformly `fix_provider` but the actionable text is auth-aware. Successful probes return `status: ok` with the parsed 5h / 7d utilization in the message body; stale-cache returns `status: warn`. Routing an auth-class failure to `kind: fix_oauth` (the other discriminator the framework supports) would require splitting this check across the `provider` / `auth` boundary — deferred to v1.x if `olp doctor` consumers report the ambiguity.
|
||||
|
||||
### Provider coverage
|
||||
|
||||
| Provider | Live quota probe | Path |
|
||||
|---|---|---|
|
||||
| `anthropic` | ✅ Live — 13 fields via `anthropic-ratelimit-unified-*` headers | This section |
|
||||
| `openai` (codex) | ❌ Not available — `openai/codex` CLI has no public quota API | Falls back to audit-derived request counts |
|
||||
| `mistral` | ❌ Not implemented at v0.5.0 — no public quota endpoint accessible to Vibe / Le Chat member / La Plateforme API keys. Mistral's Admin API does expose billing / usage queries but is gated to org-admin scope and out of scope for OLP family-tier deployment. | Falls back to audit-derived request counts |
|
||||
|
||||
If Mistral ever publishes a usage endpoint, `lib/providers/mistral.mjs` DL-7 marks the re-entry point.
|
||||
|
||||
### Schema-drift protection
|
||||
|
||||
Claude Code v2.1.x is distributed as a compiled native binary (Mach-O on macOS, ELF on Linux) — the OCP-era "grep `cli.js`" verification no longer applies. OLP's replacement protocol (ADR 0013 § Rule 5):
|
||||
|
||||
1. `strings` over the platform-specific claude-code binary captures all hardcoded header names the binary expects.
|
||||
2. A live `POST /v1/messages` against `api.anthropic.com` with valid OAuth captures what the server actually emits today.
|
||||
3. Diff path 1 vs path 2 → the actionable schema delta.
|
||||
|
||||
This is re-run at every major `claude --version` bump (next trigger: v2.x → v3.x), at the Annual Alignment Audit (14 May), and whenever `olp doctor anthropic.quota_probe_reachable` returns an unexpected status code. The current pinned schema (13 fields, `2026-05-26`) lives in `models-registry.json` under `quota_probe.schema_version`.
|
||||
|
||||
---
|
||||
|
||||
@@ -98,10 +365,10 @@ Trigger types, fallback safety, idempotency rules, and the full example config l
|
||||
|---|---|---|---|---|
|
||||
| `/v1/chat/completions` | POST | 1 | ✅ Shipped | OpenAI-compatible Chat Completions entry. Internally normalized to IR, dispatched to a provider plugin, response shape converted back. |
|
||||
| `/v1/models` | GET | 1 | ✅ Shipped | Lists models from `models-registry.json`. |
|
||||
| `/health` | GET | 1 | ✅ Shipped | Per-provider health snapshot. Phase 2 owner-only-trim: full per-provider details to owner identity; trimmed `{ ok, version }` to guest / anonymous. Gate via `auth.owner_only_endpoints` config. |
|
||||
| `/dashboard` | GET | 3 | ✅ Shipped (D50 + D51) | Owner-only multi-provider dashboard HTML (4 panels: quota / 24h request stats / 30d spend trend / top fallback chains; 30s poll with visibilitychange pause). Owner-only_block; non-owner identities receive 401. Localhost-bound by default. |
|
||||
| `/v0/management/dashboard-data` | GET | 3 | ✅ Shipped (D50) | JSON aggregate consumed by the dashboard 30s poll: `{ generated_at, window_24h, cache_hit_24h, quota, spend_trend_30d, top_fallback_chains_24h, cache_stats }`. Owner-only_block. |
|
||||
| `/v0/management/quota` | GET | 3 | ✅ Shipped (D50) | Per-provider quota snapshot via `provider.quotaStatus()` (subset of dashboard-data; useful for scripted monitoring). Owner-only_block. |
|
||||
| `/health` | GET | 1 | ✅ Shipped | Per-provider health snapshot. Phase 2 owner-only-trim: full per-provider details to owner identity; trimmed `{ ok, version }` to guest / anonymous. Gate via `auth.owner_only_endpoints` config. **Optional `anonymousKey` field (D69 / Phase 4, v0.4.0)** appears in both trimmed and full payloads when `auth.advertise_anonymous_key: true` AND `auth.allow_anonymous: true` AND at least one non-revoked guest-tier key has `plaintext_advertise: true` (see [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md) for the trusted-LAN-only invariant). Default off — field absent when prereqs unmet. |
|
||||
| `/dashboard` | GET | 3 + 5 | ✅ Shipped (D50 + D51 + D82) | Owner-only multi-provider dashboard HTML. Phase 5 D82 adds a Claude.ai-style Plan Usage section at the top (per-provider utilization bars + reset countdowns + 60s auto-refresh + manual refresh button) on top of the existing four panels (24h request stats / 30d spend trend / top fallback chains / legacy quota fallback). Owner-only_block; non-owner identities receive 401. Localhost-bound by default. |
|
||||
| `/v0/management/dashboard-data` | GET | 3 + 5 | ✅ Shipped (D50 + D81) | JSON aggregate consumed by the dashboard polls. Shape `{ generated_at, window_24h, cache_hit_24h, quota, quota_v2, spend_trend_30d, top_fallback_chains_24h, cache_stats }`. The new `quota_v2` field (D81) is the normalized per-provider shape consumed by the Plan Usage UI; the legacy `quota` field stays alongside for backwards compatibility until v1.0.0. Owner-only_block. |
|
||||
| `/v0/management/quota` | GET | 3 + 5 | ✅ Shipped (D50 + D81) | Per-provider quota snapshot via `provider.quotaStatus()`. Includes both legacy `quota` and new `quota_v2` shape (mirrors `dashboard-data` for scripted monitoring). Owner-only_block. |
|
||||
| `/cache/stats` | GET | 3 | ✅ Shipped (D50) | Live in-memory `cacheStore.stats()` (`{ hits, misses, size, inflightCount }` + `generated_at`). Owner-only_block. |
|
||||
|
||||
---
|
||||
@@ -112,11 +379,30 @@ _placeholder — full table lands per-phase as variables are introduced._
|
||||
|
||||
| Variable | Default | Description |
|
||||
|---|---|---|
|
||||
| `OLP_PORT` | `3456` | HTTP listener port. |
|
||||
| `OLP_PORT` | `4567` | HTTP listener port. Moved off `3456` at D60 / v0.4.0 to co-host with OCP — set `OLP_PORT=3456` to restore the pre-D60 default. |
|
||||
| `OLP_BIND` | `127.0.0.1` | HTTP listener bind address. **Set to `0.0.0.0` or your LAN IP to accept LAN connections** (required for `olp-connect <ip>` to actually reach the server). Default loopback-only is the secure default. See [ADR 0011 § Deployment configurations](./docs/adr/0011-anonymous-key-deployment-context.md#deployment-configurations-d76-amendment-2026-05-26) for the trust-context table — never bind to a public-internet IP. |
|
||||
| `OLP_API_KEY` | (none) | Owner-tier OLP API key (the `olp_...` plaintext from `olp-keys keygen --owner`) used by `olp` CLI subcommands as the bearer for management endpoints. |
|
||||
| `OLP_OWNER_TOKEN` | (none) | Fallback used by `olp` CLI if `OLP_API_KEY` is absent. |
|
||||
| `OLP_PROXY_URL` | `http://127.0.0.1:$OLP_PORT` | Override target URL for `olp` CLI subcommands (so the same binary works against a remote OLP via SSH tunnel or direct LAN). |
|
||||
| `OLP_CLAUDE_BIN` | `claude` (from PATH) | Override path to the `claude` binary (Anthropic provider). Useful when multiple `claude` installs are present. |
|
||||
| `OLP_CODEX_BIN` | `codex` (from PATH) | Override path to the `codex` binary (OpenAI provider). |
|
||||
| `OLP_VIBE_BIN` | `vibe` (from PATH) | Override path to the `vibe` binary (Mistral provider). |
|
||||
|
||||
### `config.json` keys introduced at Phase 4
|
||||
|
||||
These live in `~/.olp/config.json` (not env vars) — they're documented here alongside the env-var table for discoverability.
|
||||
|
||||
| Config key | Default | Description |
|
||||
|---|---|---|
|
||||
| `streaming.heartbeat_interval_ms` | `0` (disabled) | D61 / Phase 4. SSE keepalive comment frames during stream-silent windows. Set `>0` (e.g. `15000` for 15s) when OLP runs behind nginx / Cloudflare / Tailscale Funnel with idle-abort timeouts. |
|
||||
| `auth.advertise_anonymous_key` | `false` | D69 / Phase 4. When `true`, surfaces an existing guest-tier key's plaintext via `/health.anonymousKey` so `olp-connect <ip>` can self-bootstrap clients on the LAN with zero out-of-band coordination. **Requires `auth.allow_anonymous: true` AND at least one key created via `olp-keys keygen --anonymous --advertise`.** Trusted-LAN-only — see [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md). |
|
||||
|
||||
### Operator CLI surfaces (Phase 4)
|
||||
|
||||
- `olp` (Node CLI at `bin/olp.mjs`): `status / health / usage / models / cache / providers / chain show / logs / restart / keys / doctor`. Run `npx olp --help` for full subcommand reference. `olp doctor --json` emits a machine-readable `next_action.ai_executable[]` payload designed for AI agents to self-repair OLP. See [ADR 0010](./docs/adr/0010-phase-4-charter-operator-and-client-ux.md) § Phase 4 D-day plan and [ADR 0002 Amendment 7](./docs/adr/0002-plugin-architecture.md) (per-plugin `doctorChecks()` contract).
|
||||
- `olp-connect` (bash at `bin/olp-connect`): zero-config LAN client setup — detects Cline / Continue.dev / Cursor / Aider / Claude Code / OpenClaw and configures each. Run `bash bin/olp-connect --help`. Requires `python3` for JSON parsing.
|
||||
- `olp-keys keygen --anonymous --advertise`: creates a guest-tier key with the plaintext stored alongside its hash so `/health.anonymousKey` can publish it. Prints an explicit ADR-0011 warning at keygen time.
|
||||
|
||||
### Per-provider auth env vars
|
||||
|
||||
These variables configure credential discovery for each provider plugin. Setting the correct one for your provider is usually required for OLP to make successful requests.
|
||||
@@ -164,9 +450,68 @@ If a fallback chain is exhausted, `X-OLP-Fallback-Exhausted` lists the tried pro
|
||||
|
||||
---
|
||||
|
||||
## Implementation status (as of 2026-05-25, post-v0.2.0)
|
||||
## IDE Setup
|
||||
|
||||
Phase 1 closed at v0.1.1 (multi-provider proxy core + pre-Phase-2 cleanup). Phase 2 closed at v0.2.0 (multi-key auth + audit + owner gating + keygen CLI; ADR 0007 § 10 all 11 acceptance criteria shipped). Phase 3 closed at v0.3.0 (Dashboard + `lib/audit-query.mjs` + daily audit rotation; ADR 0008 § 10 all 15 acceptance criteria shipped). Phase 4 (per-key per-provider auth + audit retention + SQLite hybrid + provider-cost weights) is the next planned milestone. This table reflects what is currently shipped vs. what is designed for later phases.
|
||||
Per-tool setup pages live under [`docs/integrations/`](./docs/integrations/README.md). Index:
|
||||
|
||||
| Tool | Status | Notes |
|
||||
|---|---|---|
|
||||
| [Continue.dev](./docs/integrations/continue.md) | ✅ Supported | `config.yaml` `apiBase` (not `baseURL`); supports OLP custom headers |
|
||||
| [Cline](./docs/integrations/cline.md) | ✅ Supported | "OpenAI Compatible" provider; watch Cline issue [#7128](https://github.com/cline/cline/issues/7128) |
|
||||
| [Cursor](./docs/integrations/cursor.md) | ⚠️ Best-effort | "Override OpenAI Base URL" — known fragile across Cursor updates |
|
||||
| [Aider](./docs/integrations/aider.md) | ✅ Supported | `OPENAI_API_BASE` env + `openai/` model prefix; no custom-header support |
|
||||
| [Claude Code](./docs/integrations/claude-code.md) | ❌ Not supported | Anthropic wire format only; OLP serves OpenAI wire format. Use Cline + OLP instead |
|
||||
| [OpenClaw](./docs/integrations/openclaw.md) | ✅ Supported | Telegram + Discord gateway via [`olp-plugin/`](./olp-plugin/) |
|
||||
|
||||
The fastest path is `olp-connect <olp-host-ip>` on the client device — it auto-detects what's installed and writes the per-tool config. See [Quick Start](#quick-start).
|
||||
|
||||
---
|
||||
|
||||
## Telegram / Discord Usage
|
||||
|
||||
OLP ships [`olp-plugin/`](./olp-plugin/) as a native OpenClaw gateway plugin. After install, family members get a read-only `/olp` slash command on whichever chat surfaces OpenClaw exposes (Telegram + Discord today).
|
||||
|
||||
**Install:**
|
||||
|
||||
```bash
|
||||
# Option A — OpenClaw CLI
|
||||
openclaw plugins install /path/to/olp/olp-plugin/
|
||||
|
||||
# Option B — symlink (equivalent)
|
||||
mkdir -p ~/.openclaw/extensions/
|
||||
ln -s /path/to/olp/olp-plugin/ ~/.openclaw/extensions/olp
|
||||
```
|
||||
|
||||
**Configure:** edit `~/.openclaw/openclaw.json` and set the plugin's `apiKey` to an owner-tier OLP token created with:
|
||||
|
||||
```bash
|
||||
npx olp-keys keygen --owner --name=openclaw-bot
|
||||
```
|
||||
|
||||
Use a dedicated bot key — not the maintainer's personal owner key — so revocation is scoped.
|
||||
|
||||
```json
|
||||
{
|
||||
"plugins": {
|
||||
"olp": {
|
||||
"proxyUrl": "http://127.0.0.1:4567",
|
||||
"apiKey": "olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Restart:** `openclaw gateway restart`.
|
||||
|
||||
**Use:** `/olp status`, `/olp usage`, `/olp models`, `/olp health`, `/olp cache`, `/olp providers`, `/olp doctor`, `/olp help`.
|
||||
|
||||
**Read-only by design.** Mutating subcommands (`keygen`, `revoke`, `restart`, `logs`) are deliberately NOT exposed via chat — those are SSH-only via the local `olp` CLI. See [`olp-plugin/README.md`](./olp-plugin/README.md#what-you-can-not-do-from-chat-by-design) for the rationale.
|
||||
|
||||
---
|
||||
|
||||
## Implementation status (as of 2026-05-26, post-v0.4.0)
|
||||
|
||||
Phase 1 closed at v0.1.1 (multi-provider proxy core + pre-Phase-2 cleanup). Phase 2 closed at v0.2.0 (multi-key auth + audit + owner gating + keygen CLI; ADR 0007 § 10 all 11 acceptance criteria shipped). Phase 3 closed at v0.3.0 (Dashboard + `lib/audit-query.mjs` + daily audit rotation; ADR 0008 § 10 all 15 acceptance criteria shipped). Phase 4 closed at v0.4.0 (Operator + Client UX per ADR 0010: SSE heartbeat + `recentErrors[20]` + `/v0/management/status` / `olp` Node CLI + `olp doctor` framework + ADR 0002 Amendment 7 / `olp-connect` bash + `/health.anonymousKey` + ADR 0011 / `olp-plugin/` Telegram-Discord + 6-IDE integration docs). Phase 5 closed at v0.5.0 (Quota Probes + Dashboard Enrichment — live Anthropic plan-usage probe + Claude.ai-style dashboard + audit-query aggregateProviderQuota); v0.5.1 hotfix (quota probe cache/backoff/schema-drift correctness — codex review findings F1–F3). Phase 6 is next. This table reflects what is currently shipped vs. what is designed for later phases.
|
||||
|
||||
| File / artifact | Status | Notes |
|
||||
|---|---|---|
|
||||
@@ -197,11 +542,13 @@ Phase 1 closed at v0.1.1 (multi-provider proxy core + pre-Phase-2 cleanup). Phas
|
||||
|
||||
Behaviors that work correctly at personal/family scale but have ratified follow-ups for a v1.x sprint. Single landing page: [`docs/v1x-roadmap.md`](./docs/v1x-roadmap.md).
|
||||
|
||||
- **Streaming-path singleflight not implemented.** The cache layer's D4 singleflight (one spawn per identical concurrent request) is fully wired on the buffered path but NOT on the streaming path. N concurrent identical streaming requests at v0.1 will each spawn their own CLI process. Design ratified in [ADR 0005 Amendment 8](./docs/adr/0005-cache-cross-provider.md); implementation tracked via [issue #16](https://github.com/dtzp555-max/olp/issues/16) and [v1.x roadmap #1](./docs/v1x-roadmap.md). At family scale this is observably fine — every caller still receives the correct response; the cost is N CLI processes instead of one.
|
||||
- **Streaming-path singleflight ✅ shipped (D57 + D58, 2026-05-25).** `cacheStore.getOrComputeStreaming(...)` mirrors the buffered-path `getOrCompute` and resolves the TOCTOU window between peek and spawn ([issue #16](https://github.com/dtzp555-max/olp/issues/16)). Two concurrent identical streaming requests share one CLI spawn via tee fan-out; late joiners receive accumulated replay + the live tail; per-client backpressure (`PER_CLIENT_QUEUE_CAP=1MB`) protects against slow consumers; full-disconnect aborts the source CLI via AbortController. New `X-OLP-Streaming-Inflight: source | attached` header annotates the role. New `cache_status: 'streaming_attached'` audit value tracks the singleflight win. Authority: [ADR 0005 Amendment 8](./docs/adr/0005-cache-cross-provider.md), v1.x roadmap #1.
|
||||
- **Soft triggers configured but inert.** `routing.soft_triggers` in `~/.olp/config.json` is honored by the engine's evaluation logic but `quotaStatus()` polling is not wired (ADR 0004 Amendment 2). A startup warning fires if the field is non-empty so the inert state is visible.
|
||||
- **Multi-key auth + owner gating + keygen CLI shipped at v0.2.0 (D44 + D45 + D46 + D47).** `lib/keys.mjs` (core), `lib/audit.mjs` (audit), owner-vs-guest `/health` payload trimming + `X-OLP-Fallback-Detail` policy gating, `bin/olp-keys.mjs` (keygen CLI). All 11 ADR 0007 § 10 acceptance criteria covered. v0.2.0 maintainer-merged 2026-05-25.
|
||||
|
||||
- **Phase 3 (Dashboard + audit query layer + rotation) shipped to main (D48-D54); v0.3.0 release pending.** `docs/adr/0008-dashboard-and-audit-query.md` ratified at D48. `lib/audit-query.mjs` (D49) implements the 5-function aggregate query API (in-memory ndjson scan, PII-guarded). 4 new owner-only_block endpoints at D50 (`/dashboard`, `/v0/management/dashboard-data`, `/v0/management/quota`, `/cache/stats`). `dashboard.html` full multi-panel UI at D51 (vanilla HTML+JS+fetch, 30s poll with visibilitychange pause). Daily audit rotation at D52 (synchronous on first append after UTC midnight; `audit-YYYY-MM-DD.ndjson` naming) + optional `bin/olp-audit-rotate.mjs` cron tool. `tried_providers` schema semantic fix at D53 (D45 P2 deferral). Phase 3 close to v0.3.0 is maintainer-triggered per CLAUDE.md `release_kit.phase_close_trigger`.
|
||||
- **Phase 3 (Dashboard + audit query layer + rotation) shipped at v0.3.0 (D48–D54).** `docs/adr/0008-dashboard-and-audit-query.md` + `lib/audit-query.mjs` (D49) + 4 owner-only_block endpoints (D50) + `dashboard.html` (D51) + daily audit rotation (D52) + `tried_providers` schema fix (D53). All 15 ADR 0008 § 10 acceptance criteria covered.
|
||||
|
||||
- **Phase 4 (Operator + Client UX) shipped at v0.4.0 (D60 → D73).** ADR 0010 (charter) + ADR 0011 (anonymous-key trusted-LAN limits) + ADR 0002 Amendment 7 (provider `doctorChecks()` contract). Default `OLP_PORT` 3456 → 4567 so OLP and OCP can co-host. SSE heartbeat (D61) + `recentErrors[20]` + `/v0/management/status` (D62-D63). `bin/olp.mjs` Node CLI + `bin/olp-keys.mjs` + `lib/doctor.mjs` framework with `next_action.ai_executable[]` (D64-D67). `bin/olp-connect` bash zero-config IDE auto-config + opt-in `/health.anonymousKey` (D68-D70). `olp-plugin/` OpenClaw `/olp` Telegram-Discord plugin (read-only, no chat mutations) + 6 IDE integration docs at `docs/integrations/*.md` (D71-D73). Test count 623 → 696.
|
||||
|
||||
**Bootstrap workflow (D47):** for first-run / production setup:
|
||||
|
||||
@@ -216,7 +563,7 @@ Behaviors that work correctly at personal/family scale but have ratified follow-
|
||||
npm start
|
||||
|
||||
# 4. Validate the key works (substitute the captured plaintext token)
|
||||
curl -H "Authorization: Bearer olp_..." http://localhost:3456/health
|
||||
curl -H "Authorization: Bearer olp_..." http://localhost:4567/health
|
||||
```
|
||||
|
||||
**Recovery if owner token is lost:** `npx olp-keys keygen --owner --force` revokes the previous owner key + creates a fresh one (plaintext printed once).
|
||||
@@ -226,6 +573,37 @@ Behaviors that work correctly at personal/family scale but have ratified follow-
|
||||
**New config block consumed at D45:** `config.json auth.{ allow_anonymous, owner_only_endpoints, fallback_detail_header_policy }`. Default `allow_anonymous: false` (production-off); set true to accept requests without an OLP API key (development / single-user dev mode). Startup emits a warn when `allow_anonymous: true` so the relaxed posture is observable.
|
||||
- **Provider-level `cacheKeyFields` mask not implemented.** Cache keys include every IR field including ones individual plugins drop at spawn (e.g., Anthropic plugin drops `temperature`). Spurious cache misses possible (extra spawn cost; never spurious hits). Conservative posture documented in [ADR 0005 Amendment 7](./docs/adr/0005-cache-cross-provider.md). Tracked in [v1.x roadmap #5](./docs/v1x-roadmap.md).
|
||||
|
||||
- **Agentic clients with shell-tool routing may report OLP-server-side state as "self".** This is an architectural property of spawn-CLI proxying that OLP cannot fully fix at the proxy layer. When a client like OpenClaw runs in **client mode** (gateway on user's machine, LLM backend pointed at remote OLP) and the agent exposes shell / fs tools, those tool calls execute on whatever machine the client's tool-handler is wired to. If the client's `ocp` / `olp` plugin routes shell to the OLP server host, an in-agent "do a self-check" prompt produces results describing the OLP host (e.g. PI231) rather than the user's local machine. OLP cannot inject "you are the client, not the server" into the prompt because (a) the client owns the system message, and (b) OLP is stateless and doesn't know which client is calling. **Phase 6c's `--system-prompt` override (ADR 0009 Amendment 1) addresses one side of this — claude CLI no longer injects `<env>cwd=...</env>` blocks into the prompt** — but it cannot prevent the client from sending tool-results that the model then describes as its own state. Recommendations for integrators:
|
||||
|
||||
- **OpenClaw client mode** — if you want bot self-checks to describe the user's local machine, configure OpenClaw's tool plugins (`plugins.entries.{ocp,olp}` etc.) so shell / fs tools route to the local host, not to the OLP server. The bundled `olp-plugin/` ships as a read-only telemetry surface (no shell mutations); the older `ocp` plugin's shell-routing semantics are OCP-era legacy and may misroute when OLP is the LLM backend.
|
||||
- **Hermes Agent client mode** — Hermes pre-processes tools on its own host before sending; the LLM emits no tool_use that reaches OLP, so this limitation does not apply to chat-only Hermes flows. Tool-using Hermes flows behave correctly: Hermes runs the tool locally and includes the result as a follow-up user message.
|
||||
- **Cline / Continue.dev / Cursor / Aider** — IDE clients typically run shell / fs tools locally on the user's machine, so self-checks report the user's machine correctly. No OLP-side action needed.
|
||||
- **Generic agentic clients** — if your client routes tool execution to the OLP server, expect bot self-reports to describe the OLP server's state. Either: (1) configure your client's tool handler to run tools locally, or (2) document this to your client users as a known limitation.
|
||||
|
||||
See [ADR 0014 Amendment 1](./docs/adr/0014-sandbox-runtime-integration.md) for the multi-tenant security counterpart of this issue. Phase 7 Solution 1 (shipped v0.7.0) per-spawn ephemeral `$HOME` + symlinked credentials redirect all CLI state writes to `/tmp/olp-spawn/<keyId>/<reqId>/home/`, so a prompt-injected `cat ~/.claude.json` reads only the ephemeral file, not other tenants' OAuth tokens.
|
||||
|
||||
### Security Model
|
||||
|
||||
OLP's multi-tenant isolation has **three deployment tiers**, each suited to a different trust model. The orchestrator reads each provider plugin's `ISOLATION` block (ADR 0002 Amendment 9) to pick the right primitives.
|
||||
|
||||
| Tier | Trust assumption | Mechanism | Suitable for |
|
||||
|---|---|---|---|
|
||||
| **shared-os-user** (default) | All OLP key holders trust each other (family / personal pool) | Per-spawn ephemeral `$HOME` (Layer 1) + symlinked credentials (Layer 2) + provider tool-suppression (Layer 4 — anthropic Phase 6c `--system-prompt`, codex `--sandbox read-only`) | Family LAN, personal multi-device, trusted small teams. ADR 0001 § Mission. |
|
||||
| **per-os-user** | Trust boundaries between OLP keys (e.g., distinct family members on a shared host) | All of the above + per-OLP-key OS user (systemd `User=olp-<keyId>`, separate uid for kernel-level fs deny) | Untrusted-key deploy that still pools OAuth subscription. Operator-managed. |
|
||||
| **separate-vm** | Adversarial isolation between OLP keys (commercial / public-demo scenarios) | All of the above + dedicated VM per OLP key | OLP outside its stated mission. Each provider plugin's `recommendedDeploymentTier` declares its minimum acceptable tier. |
|
||||
|
||||
The provider plugins' `crossTenantReadProtection` field declares **how** each protects against cross-tenant lateral filesystem reads:
|
||||
|
||||
| Provider | `crossTenantReadProtection` | Mechanism |
|
||||
|---|---|---|
|
||||
| anthropic (claude CLI) | `tool-suppression` | Phase 6c `--system-prompt` replaces claude's default system prompt; the model receives no tool descriptions for Read/Bash/etc., so prompt-injection produces no `tool_use` to read other tenants' files. |
|
||||
| codex (codex CLI) | `inner-sandbox` | codex's own bubblewrap-based `--sandbox read-only` default confines shell-tool reads to its inner sandbox view. |
|
||||
| mistral (vibe CLI) | `none` | No tool-suppression flag known on vibe at present. Recommended deployment tier `separate-vm` until a hardening regime is verified (Task #4 follow-up spike). |
|
||||
|
||||
**OLP_SANDBOX_DISABLED=1** env var disables Layer 3 (per-call sandbox-runtime wrapping) while preserving Layers 1+2+4. This is the post-Amendment 1 escape hatch retained for 1-2 releases; production deployments should leave it unset.
|
||||
|
||||
**Attribution vs isolation.** ADR 0007 multi-key auth provides **attribution** (per-key audit, per-key cache namespace, per-key provider gating). ADR 0014 Amendment 1 provides **isolation** (the security tier above). Both layer cleanly — attribution always operates; isolation tier is operator-selected per deployment.
|
||||
|
||||
---
|
||||
|
||||
## Architecture
|
||||
@@ -253,7 +631,8 @@ The original v0.1 spec (in `~/.cc-rules/memory/projects/olp_v0_1_spec.md` on the
|
||||
- **Phase 1** — Multi-provider proxy core: `server.mjs`, IR, three Tier-D provider plugins (Anthropic / OpenAI Codex / Mistral Vibe), cache (D1+D4) + cleanup (D2 bypass / D3 chunked replay / D23 size cap), fallback engine with first-chunk safety + hard triggers + per-hop log observability, IR↔OpenAI translation under Rule 2(b). ✅ Shipped — v0.1.0 (2026-05-24) + v0.1.1 cleanup (2026-05-25, D35–D42).
|
||||
- **Phase 2** — Multi-key auth (`lib/keys.mjs`) per ADR 0007: opaque OLP API keys, per-key cache namespacing, owner-vs-guest tier for header gating, audit ndjson (`lib/audit.mjs`), `/health` payload trimming + `X-OLP-Fallback-Detail` emission gating, `OLP_OWNER_TOKEN` env override, keygen CLI (`bin/olp-keys.mjs`). ✅ Shipped — v0.2.0 (2026-05-25, D43-A → D47). All 11 ADR 0007 § 10 acceptance criteria covered.
|
||||
- **Phase 3** — Dashboard + audit query layer + daily audit rotation per ADR 0008: in-memory ndjson aggregate query layer (`lib/audit-query.mjs`), 4 owner-only_block management endpoints (`/dashboard` + `/v0/management/dashboard-data` + `/v0/management/quota` + `/cache/stats`), multi-panel `dashboard.html` with 30s poll, synchronous daily audit rotation + `bin/olp-audit-rotate.mjs` cron tool, `tried_providers` schema fix (D45 P2 deferral). ✅ Shipped — v0.3.0 (2026-05-25, D48 → D54). All 15 ADR 0008 § 10 acceptance criteria covered.
|
||||
- **Phase 4 (planned)** — Per-key per-provider auth artifact mapping (ADR 0007 § 12 deferral), audit query rotation/retention policies, SQLite hybrid migration (ADR 0007 § 13 trigger), provider-cost weights for spend trend.
|
||||
- **Phase 5** — Live quota probe (Anthropic Pro/Max OAuth plan-usage via `anthropic-ratelimit-unified-*` headers), Claude.ai-style dashboard enrichment (utilization bars + reset countdowns), audit-query `aggregateProviderQuota()`, per-provider quota_v2 shape in dashboard-data. ✅ Shipped — v0.5.0 (2026-05-27). v0.5.1 hotfix (2026-05-27): quota probe cache/backoff/schema-drift correctness (codex review findings F1–F3).
|
||||
- **Phase 6 (planned)** — Per-key per-provider auth artifact mapping (ADR 0007 § 12 deferral), audit query rotation/retention policies, SQLite hybrid migration (ADR 0007 § 13 trigger), provider-cost weights for spend trend.
|
||||
- **Phase 4+ (v1.x roadmap, triggered as needed)** — Full deferred-work tracker: [`docs/v1x-roadmap.md`](./docs/v1x-roadmap.md). Includes streaming-path singleflight ([issue #16](https://github.com/dtzp555-max/olp/issues/16) + ADR 0005 Amendment 8 design ratified), soft-trigger reactivation (ADR 0004 Amendment 2), `/health` activeSpawns integration, provider-level `cacheKeyFields` mask, streaming-path SPAWN_FAILED salvage.
|
||||
- **Phase N (opt-in)** — Tier-2 / Tier-C provider plugins (Grok / Kimi / MiniMax / GLM / Qwen) per [ADR 0006](./docs/adr/0006-provider-inclusion.md); provider-native protocol endpoints; deterministic triggers. Triggered by tier-2 demand, not on the bootstrap path.
|
||||
|
||||
@@ -263,16 +642,22 @@ Full spec (decision rationale, open questions, risks): `~/.cc-rules/memory/proje
|
||||
|
||||
## Migration from OCP
|
||||
|
||||
OLP is OCP's successor. The trigger was Anthropic's 2026-05-14 announcement (effective 2026-06-15) splitting `claude -p` / Agent SDK / third-party agent traffic out of the Pro/Max subscription pool into a separate fixed $100/month Agent SDK credit pool — invalidating OCP's foundational assumption (*"subscription = unlimited within rate limits"*) for its only provider. OLP's structural response is to spread risk across multiple subscriptions whose CLI/programmatic use remains in their main subscription pool, with intelligent fallback when one runs out.
|
||||
|
||||
Beyond the billing trigger, OLP is intentionally NOT a commercial multi-tenant SaaS (LiteLLM / OpenRouter / Portkey already serve that market with funding + SOC2), NOT an enterprise gateway competing on provider breadth, NOT a model-capability router ("route to the smartest model" — you pick the model in `routing.chains`), and NOT a conversation-state store (your client manages its own context). See [ADR 0001](./docs/adr/0001-project-founding.md) for the founding decision and [`ALIGNMENT.md`](./ALIGNMENT.md) for the constitution that governs every plugin / IR / entry-surface change.
|
||||
|
||||
### Migrating an existing OCP install
|
||||
|
||||
_placeholder — `scripts/migrate-from-ocp.mjs` lands with Phase 7 (📋 planned, not yet authored)._
|
||||
|
||||
Anticipated user-facing flow (target: <5 minutes):
|
||||
|
||||
1. Stop OCP (`launchctl bootout` the OCP service or `ocp stop`).
|
||||
2. Install OLP.
|
||||
3. Run `olp migrate-from-ocp` — copies `~/.ocp/keys/` to `~/.olp/keys/` and points provider plugins at OCP's existing auth artifacts where applicable.
|
||||
4. Start OLP. Clients pointing at port 3456 keep working; their existing OLP API keys remain valid.
|
||||
2. Install OLP (per [§ Manual install](#manual-install-5-10-min) above).
|
||||
3. Run `olp migrate-from-ocp` — will copy `~/.ocp/keys/` to `~/.olp/keys/` and point provider plugins at OCP's existing auth artifacts where applicable.
|
||||
4. Start OLP. Clients pointing at port 4567 (or 3456 with `OLP_PORT=3456`) keep working; their existing OLP API keys remain valid.
|
||||
|
||||
OCP's cache directory is *not* migrated: OLP's cache key format includes provider+model and warms cold naturally. OCP enters maintenance mode (stability fixes only) when OLP v0.1 ships; new development happens in OLP.
|
||||
**Default port moved 3456 → 4567 at v0.4.0** so OCP and OLP can co-host on the same machine during the migration window — set `OLP_PORT=3456` if you want the pre-D60 default. OCP's cache directory is *not* migrated: OLP's cache key format includes provider+model and warms cold naturally. OCP enters maintenance mode (stability fixes only) when OLP v0.1 ships; new development happens in OLP.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Executable
+657
@@ -0,0 +1,657 @@
|
||||
#!/usr/bin/env bash
|
||||
# bin/olp-connect — Lightweight client script to connect this machine to a remote
|
||||
# OLP (Open LLM Proxy). Ported from OCP's `ocp-connect` per ADR 0010 § Phase 4
|
||||
# D68-D70 charter; uses /health.anonymousKey when the remote operator opted in
|
||||
# via `auth.advertise_anonymous_key: true` (ADR 0011).
|
||||
#
|
||||
# Authority:
|
||||
# - ADR 0010 (Phase 4 charter — D68 line: client-side IDE auto-config)
|
||||
# - ADR 0011 (anonymous-key deployment-context limits — trusted-LAN invariant)
|
||||
# - OCP `ocp-connect` v1.3.0 (prior-art reference)
|
||||
#
|
||||
# Why bash (not Node like `olp` CLI):
|
||||
# olp-connect MUST run on CLIENT machines that may not have a recent Node
|
||||
# installed (parents' laptops, work machines, Raspberry Pi). bash + curl +
|
||||
# python3 give maximum portability; this script does not import any OLP
|
||||
# Node modules.
|
||||
#
|
||||
# Dependencies: bash >=4, curl, python3 (for /health JSON parsing).
|
||||
#
|
||||
# Install:
|
||||
# curl -fsSL https://raw.githubusercontent.com/dtzp555-max/olp/main/bin/olp-connect -o olp-connect
|
||||
# chmod +x olp-connect
|
||||
#
|
||||
# Or via npm/npx (once `npm install -g olp` is run on a machine that has Node):
|
||||
# olp-connect <ip>
|
||||
#
|
||||
# Or run directly via curl-pipe:
|
||||
# curl -fsSL https://raw.githubusercontent.com/dtzp555-max/olp/main/bin/olp-connect | bash -s -- <host-ip>
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
# D78 (G13): derive version from package.json instead of hardcoding (was
|
||||
# stuck at "0.4.0-phase4" through v0.4.1/v0.4.2/v0.4.3 because no one
|
||||
# updated it). Look up package.json next to the script if available;
|
||||
# fall back to "unknown" when running curl-piped (no on-disk package.json).
|
||||
_resolve_version() {
|
||||
local script_dir pkg
|
||||
# When curl-piped (`curl ... | bash`), BASH_SOURCE[0] is empty → dirname
|
||||
# yields "." → script_dir resolves to cwd. D78 reviewer P2-1 hardening:
|
||||
# require the suffix-strip to actually fire (script_dir ENDED with /bin),
|
||||
# otherwise we'd happily pick up an unrelated package.json from whatever
|
||||
# directory the user happens to be in when piping. Belt-and-braces.
|
||||
# ${BASH_SOURCE[0]:-} default-empty guards against `set -u` nounset error
|
||||
# when invoked via `curl ... | bash` (no source file → BASH_SOURCE unset).
|
||||
script_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]:-}")" &>/dev/null && pwd)"
|
||||
if [[ "$script_dir" != */bin ]]; then
|
||||
echo "unknown"
|
||||
return
|
||||
fi
|
||||
pkg="${script_dir%/bin}/package.json"
|
||||
# D78 reviewer P2-2: pass $pkg via env var instead of -c interpolation
|
||||
# so paths with apostrophes / shell metacharacters can't break the
|
||||
# python invocation. Canonical layout is safe; this is defense-in-depth.
|
||||
if [[ -f "$pkg" ]] && command -v python3 >/dev/null 2>&1; then
|
||||
OLP_PKG_PATH="$pkg" python3 -c 'import json,os;print(json.load(open(os.environ["OLP_PKG_PATH"])).get("version","unknown"))' 2>/dev/null || echo "unknown"
|
||||
else
|
||||
echo "unknown"
|
||||
fi
|
||||
}
|
||||
OLP_CONNECT_VERSION="$(_resolve_version)"
|
||||
|
||||
show_version() {
|
||||
echo "olp-connect $OLP_CONNECT_VERSION"
|
||||
}
|
||||
|
||||
show_help() {
|
||||
cat <<'EOF'
|
||||
olp-connect — Connect this machine to a remote OLP (Open LLM Proxy)
|
||||
|
||||
Configures OPENAI_BASE_URL + OPENAI_API_KEY in your shell rc file (and macOS
|
||||
launchctl env / Linux systemd user env), then detects installed IDEs (Cline,
|
||||
Continue.dev, Cursor, Aider, OpenClaw, Claude Code) and prints / writes the
|
||||
provider-specific configuration each needs.
|
||||
|
||||
Usage:
|
||||
olp-connect <host-ip> [options]
|
||||
olp-connect --help
|
||||
olp-connect --version
|
||||
|
||||
Options:
|
||||
--port PORT Port OLP listens on (default: 4567 — OLP v0.4.0+ default;
|
||||
set 3456 if connecting to a pre-D60 OLP install)
|
||||
--key API_KEY OLP API key (from `olp-keys keygen` on the server). When
|
||||
omitted, the script reads /health.anonymousKey (if the
|
||||
server opted in via auth.advertise_anonymous_key=true) or
|
||||
prompts interactively.
|
||||
--no-system-env Skip macOS launchctl setenv / Linux systemd env writes;
|
||||
only update shell rc files.
|
||||
--dry-run Print everything the script would do without modifying
|
||||
any file or setting any env var.
|
||||
--version Print version and exit
|
||||
--help, -h Show this help
|
||||
|
||||
Examples:
|
||||
olp-connect 192.168.1.10
|
||||
olp-connect 192.168.1.10 --port 8080
|
||||
olp-connect 192.168.1.10 --key olp_AbcDef1234...
|
||||
olp-connect 100.64.0.5 --dry-run
|
||||
|
||||
Requires:
|
||||
bash, curl, python3 (for /health JSON parsing)
|
||||
|
||||
Exit codes:
|
||||
0 success
|
||||
1 bad arguments / unknown flag / missing required value
|
||||
2 connectivity or auth failure / smoke test failure
|
||||
|
||||
Authority: ADR 0010 § Phase 4 D68-D70; ADR 0011 (anonymous-key trusted-LAN
|
||||
invariant — when --key is auto-resolved from /health.anonymousKey, this
|
||||
deployment MUST be on a trusted LAN).
|
||||
EOF
|
||||
}
|
||||
|
||||
# ── Globals populated by main() ─────────────────────────────────────────────
|
||||
|
||||
DRY_RUN=false
|
||||
NO_SYSTEM_ENV=false
|
||||
|
||||
# ── Logging helpers ─────────────────────────────────────────────────────────
|
||||
|
||||
log_info() { echo " $*"; }
|
||||
log_step() { echo " → $*"; }
|
||||
log_ok() { echo " ✓ $*"; }
|
||||
log_warn() { echo " ⚠ $*"; }
|
||||
log_err() { echo " ✗ $*" >&2; }
|
||||
|
||||
# Echo a state change (a write / env-set) before executing — operator can
|
||||
# Ctrl-C if something looks wrong. Returns 0 always.
|
||||
log_change() { echo " • $*"; }
|
||||
|
||||
# ── IDE detection + configuration ───────────────────────────────────────────
|
||||
|
||||
# Truncate long keys for display (avoid leaking via screenshot / screen share).
|
||||
key_display() {
|
||||
local k="$1"
|
||||
if [[ -z "$k" ]]; then
|
||||
echo "(none — anonymous; most IDEs require a non-empty API Key)"
|
||||
elif [[ ${#k} -gt 16 ]]; then
|
||||
echo "${k:0:8}...${k: -4}"
|
||||
else
|
||||
echo "$k"
|
||||
fi
|
||||
}
|
||||
|
||||
# D74 P1-2: validate OLP API key format. Per ADR 0007 § 3, tokens are
|
||||
# `olp_` + 32 random bytes base64url-encoded (43 chars, no padding). This
|
||||
# regex pins the on-the-wire shape so a malformed or hostile `--key` /
|
||||
# server-advertised `anonymousKey` never gets persisted into a shell rc.
|
||||
# Returns 0 on valid, 1 on invalid (with diagnostic to stderr).
|
||||
validate_olp_token() {
|
||||
local k="$1" source="$2"
|
||||
if [[ ! "$k" =~ ^olp_[A-Za-z0-9_-]{43}$ ]]; then
|
||||
log_err "Rejected $source: token format does not match ^olp_[A-Za-z0-9_-]{43}$ (ADR 0007 § 3)."
|
||||
log_err " Got ${#k}-char value starting with '$(echo "$k" | cut -c1-8)...'"
|
||||
log_err " Expected: olp_ followed by 43 base64url chars. Run 'npx olp-keys list' on the server"
|
||||
log_err " to confirm the key format, or have the operator regenerate with 'npx olp-keys keygen'."
|
||||
return 1
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# D74 P1-2: POSIX shell-quote a value before interpolating into a shell rc
|
||||
# write. Wraps in single quotes + escapes embedded single quotes per:
|
||||
# foo'bar → 'foo'\''bar'
|
||||
# Even with the validator above, this is defense-in-depth: any non-token
|
||||
# string that slips through (e.g., environment.d KEY=VALUE writes) MUST be
|
||||
# safe to source. Same helper pattern as lib/doctor.mjs _shellQuote.
|
||||
shell_quote() {
|
||||
local s="$1"
|
||||
# Escape any single quotes: ' → '\''
|
||||
printf "'%s'" "${s//\'/\'\\\'\'}"
|
||||
}
|
||||
|
||||
# Detect Claude Code and print warn-only message. Per ADR 0010 § Out of
|
||||
# Phase 4 scope, OLP does NOT ship /v1/messages and CC is not a supported
|
||||
# client. The user is steered toward Cline + OLP.
|
||||
detect_claude_code() {
|
||||
if command -v claude &>/dev/null; then
|
||||
log_info ""
|
||||
log_info "Detected: Claude Code (`command -v claude`)"
|
||||
log_warn "Claude Code is NOT supported as an OLP client (OLP does not ship"
|
||||
log_warn " /v1/messages — see ADR 0010 § Out-of-Phase-4-scope)."
|
||||
log_warn " Recommended alternative: install Cline (VSCode extension) + OLP."
|
||||
log_warn " Cline uses OpenAI-shape /v1/chat/completions which OLP DOES serve."
|
||||
fi
|
||||
}
|
||||
|
||||
# Detect Cline VSCode extension and print manual-configure snippet.
|
||||
# Cline cannot be auto-configured via env vars — user must paste into VSCode
|
||||
# settings UI. We surface the values for them.
|
||||
detect_cline() {
|
||||
local base_url="$1" key="$2"
|
||||
local exts=""
|
||||
if command -v code &>/dev/null; then
|
||||
exts=$(code --list-extensions 2>/dev/null || true)
|
||||
fi
|
||||
if [[ -z "$exts" && -d "$HOME/.vscode/extensions" ]]; then
|
||||
exts=$(ls "$HOME/.vscode/extensions/" 2>/dev/null || true)
|
||||
fi
|
||||
|
||||
if echo "$exts" | grep -qiE 'cline|saoudrizwan\.claude-dev'; then
|
||||
log_info ""
|
||||
log_info "Detected: Cline (VSCode extension)"
|
||||
log_info " Cline must be configured via the VSCode settings UI."
|
||||
log_info " Open VSCode → Cline panel → Settings → API Provider:"
|
||||
log_info " API Provider: \"OpenAI Compatible\""
|
||||
log_info " Base URL: $base_url/v1"
|
||||
log_info " API Key: $(key_display "$key")"
|
||||
log_info " Model ID: claude-sonnet-4-5 (or any model from /v1/models)"
|
||||
fi
|
||||
}
|
||||
|
||||
# Detect Continue.dev and write a `models:` entry to ~/.continue/config.yaml
|
||||
# (idempotent — checks if an entry with the same name exists first).
|
||||
detect_continue() {
|
||||
local base_url="$1" key="$2"
|
||||
local exts=""
|
||||
if command -v code &>/dev/null; then
|
||||
exts=$(code --list-extensions 2>/dev/null || true)
|
||||
fi
|
||||
local config_yaml="$HOME/.continue/config.yaml"
|
||||
local config_json="$HOME/.continue/config.json"
|
||||
|
||||
local found=false
|
||||
if echo "$exts" | grep -qi 'continue\.continue'; then found=true; fi
|
||||
if [[ -f "$config_yaml" || -f "$config_json" ]]; then found=true; fi
|
||||
|
||||
if ! $found; then return 0; fi
|
||||
|
||||
log_info ""
|
||||
log_info "Detected: Continue.dev"
|
||||
log_info " Configuration snippet for ~/.continue/config.yaml:"
|
||||
log_info " models:"
|
||||
log_info " - name: OLP Sonnet"
|
||||
log_info " provider: openai"
|
||||
log_info " model: claude-sonnet-4-5"
|
||||
log_info " apiBase: $base_url/v1"
|
||||
log_info " apiKey: $(key_display "$key")"
|
||||
log_info " Note: Continue.dev autoreload-on-save is fragile; restart VSCode if"
|
||||
log_info " the new model doesn't appear in the model selector."
|
||||
}
|
||||
|
||||
# Detect Cursor and print manual snippet + known-fragility warning.
|
||||
detect_cursor() {
|
||||
local base_url="$1" key="$2"
|
||||
local found=false
|
||||
if command -v cursor &>/dev/null; then found=true; fi
|
||||
if [[ -d "$HOME/.cursor" ]]; then found=true; fi
|
||||
if [[ -d "/Applications/Cursor.app" ]]; then found=true; fi
|
||||
|
||||
if ! $found; then return 0; fi
|
||||
|
||||
log_info ""
|
||||
log_info "Detected: Cursor"
|
||||
log_info " Cmd+Shift+P → 'Cursor Settings' → Models:"
|
||||
log_info " OpenAI API Key: $(key_display "$key")"
|
||||
log_info " Override OpenAI Base URL: $base_url/v1"
|
||||
log_info " Custom OpenAI Models: claude-sonnet-4-5,claude-opus-4-1"
|
||||
log_warn " Cursor's base-URL handling is known-fragile (issue #7128 et al);"
|
||||
log_warn " if requests fail with 'malformed request', try removing then"
|
||||
log_warn " re-adding the model in the Cursor models list."
|
||||
}
|
||||
|
||||
# Detect Aider and write OPENAI_API_BASE / OPENAI_API_KEY to rc files.
|
||||
# Aider reads these env vars at startup — already handled by the rc-file
|
||||
# block in main(). We just announce detection here.
|
||||
detect_aider() {
|
||||
if command -v aider &>/dev/null; then
|
||||
log_info ""
|
||||
log_info "Detected: Aider (`command -v aider`)"
|
||||
log_info " Aider reads OPENAI_API_BASE + OPENAI_API_KEY from env."
|
||||
log_info " These are already being written to your shell rc — open a fresh"
|
||||
log_info " shell and run: aider --model openai/claude-sonnet-4-5"
|
||||
fi
|
||||
}
|
||||
|
||||
# Detect OpenClaw. Phase 4 D71-D73 shipped olp-plugin/ as the OpenClaw
|
||||
# gateway plugin for /olp Telegram + Discord slash commands. Point users
|
||||
# at the install path.
|
||||
detect_openclaw() {
|
||||
if command -v openclaw &>/dev/null || [[ -f "$HOME/.openclaw/openclaw.json" ]]; then
|
||||
log_info ""
|
||||
log_info "Detected: OpenClaw"
|
||||
log_info " OLP ships an OpenClaw gateway plugin for /olp Telegram + Discord"
|
||||
log_info " slash commands (status / usage / cache / models / providers /"
|
||||
log_info " chain show / health / doctor). Read-only by design — no chat-side"
|
||||
log_info " mutations."
|
||||
log_info ""
|
||||
log_info " Install the plugin (one-time, on the host running OpenClaw):"
|
||||
log_info " git clone https://github.com/dtzp555-max/olp.git /tmp/olp-repo"
|
||||
log_info " openclaw plugins install /tmp/olp-repo/olp-plugin"
|
||||
log_info " # OR symlink: ln -sf /tmp/olp-repo/olp-plugin ~/.openclaw/extensions/olp"
|
||||
log_info ""
|
||||
log_info " Then edit ~/.openclaw/openclaw.json to set the plugin apiKey to a"
|
||||
log_info " dedicated OLP key (NOT your owner key — create one via olp-keys"
|
||||
log_info " keygen --name <bot-name>). Restart OpenClaw gateway."
|
||||
log_info ""
|
||||
log_info " See docs/integrations/openclaw.md for full instructions."
|
||||
fi
|
||||
}
|
||||
|
||||
# ── rc-file helpers ─────────────────────────────────────────────────────────
|
||||
|
||||
# Identify which shell rc files to write to. Returns paths on stdout, one per line.
|
||||
detect_rc_files() {
|
||||
local is_mac=false
|
||||
[[ "$(uname)" == "Darwin" ]] && is_mac=true
|
||||
if [[ "${SHELL:-}" == */fish ]]; then
|
||||
log_warn "fish shell detected; writing to ~/.bashrc — add to fish config manually." >&2
|
||||
echo "$HOME/.bashrc"
|
||||
return
|
||||
fi
|
||||
if $is_mac; then
|
||||
# macOS Catalina+ default shell is zsh
|
||||
[[ -f "$HOME/.bashrc" ]] && echo "$HOME/.bashrc"
|
||||
[[ -f "$HOME/.zshrc" ]] || { $DRY_RUN || touch "$HOME/.zshrc"; }
|
||||
echo "$HOME/.zshrc"
|
||||
else
|
||||
[[ -f "$HOME/.bashrc" || "${SHELL:-}" == */bash ]] && echo "$HOME/.bashrc"
|
||||
[[ -f "$HOME/.zshrc" || "${SHELL:-}" == */zsh ]] && echo "$HOME/.zshrc"
|
||||
fi
|
||||
}
|
||||
|
||||
# Remove any previously-written OLP block from an rc file (idempotent).
|
||||
# The block is bracketed by:
|
||||
# # OLP LAN (added by olp-connect) ... # /OLP LAN
|
||||
strip_olp_block() {
|
||||
local rc_file="$1"
|
||||
[[ -f "$rc_file" ]] || return 0
|
||||
if $DRY_RUN; then
|
||||
if grep -qF '# OLP LAN (added by olp-connect)' "$rc_file" 2>/dev/null; then
|
||||
log_change "[dry-run] would strip existing OLP block from $rc_file"
|
||||
fi
|
||||
return 0
|
||||
fi
|
||||
python3 - "$rc_file" <<'PYEOF'
|
||||
import sys
|
||||
path = sys.argv[1]
|
||||
try:
|
||||
with open(path) as f:
|
||||
lines = f.readlines()
|
||||
except OSError:
|
||||
sys.exit(0)
|
||||
out = []
|
||||
skip = False
|
||||
for line in lines:
|
||||
s = line.rstrip('\n')
|
||||
if s == '# OLP LAN (added by olp-connect)':
|
||||
skip = True
|
||||
continue
|
||||
if skip and s == '# /OLP LAN':
|
||||
skip = False
|
||||
continue
|
||||
if skip:
|
||||
continue
|
||||
out.append(line)
|
||||
with open(path, 'w') as f:
|
||||
f.writelines(out)
|
||||
PYEOF
|
||||
}
|
||||
|
||||
# Append a new OLP block to an rc file.
|
||||
append_olp_block() {
|
||||
local rc_file="$1" base_url="$2" key="$3"
|
||||
if $DRY_RUN; then
|
||||
log_change "[dry-run] would append OLP block to $rc_file:"
|
||||
log_change " # OLP LAN (added by olp-connect)"
|
||||
log_change " export OPENAI_BASE_URL=$(shell_quote "$base_url/v1")"
|
||||
[[ -n "$key" ]] && log_change " export OPENAI_API_KEY=$(shell_quote "$(key_display "$key")")"
|
||||
log_change " # /OLP LAN"
|
||||
return 0
|
||||
fi
|
||||
# D74 P1-2: shell-quote values before writing to rc files. Defense-in-depth
|
||||
# alongside validate_olp_token — even if a future code path bypasses the
|
||||
# validator, the rc file remains safe to source.
|
||||
{
|
||||
echo ""
|
||||
echo "# OLP LAN (added by olp-connect)"
|
||||
echo "export OPENAI_BASE_URL=$(shell_quote "$base_url/v1")"
|
||||
if [[ -n "$key" ]]; then
|
||||
echo "export OPENAI_API_KEY=$(shell_quote "$key")"
|
||||
fi
|
||||
echo "# /OLP LAN"
|
||||
} >> "$rc_file"
|
||||
}
|
||||
|
||||
# ── System-level env (macOS launchctl / Linux systemd user) ────────────────
|
||||
|
||||
set_system_env() {
|
||||
local base_url="$1" key="$2"
|
||||
if $NO_SYSTEM_ENV; then
|
||||
log_info "Skipping system-level env (--no-system-env)"
|
||||
return 0
|
||||
fi
|
||||
if [[ "$(uname)" == "Darwin" ]]; then
|
||||
if $DRY_RUN; then
|
||||
log_change "[dry-run] would launchctl setenv OPENAI_BASE_URL=$base_url/v1"
|
||||
[[ -n "$key" ]] && log_change "[dry-run] would launchctl setenv OPENAI_API_KEY=$(key_display "$key")"
|
||||
return 0
|
||||
fi
|
||||
launchctl setenv OPENAI_BASE_URL "$base_url/v1" 2>/dev/null || log_warn "launchctl setenv OPENAI_BASE_URL failed"
|
||||
if [[ -n "$key" ]]; then
|
||||
launchctl setenv OPENAI_API_KEY "$key" 2>/dev/null || log_warn "launchctl setenv OPENAI_API_KEY failed"
|
||||
fi
|
||||
log_ok "launchctl setenv applied (visible to GUI apps + daemons)"
|
||||
log_info " Note: launchctl env vars reset on reboot. Re-run olp-connect after restart"
|
||||
log_info " or add the script to Login Items."
|
||||
else
|
||||
local env_dir="$HOME/.config/environment.d"
|
||||
if $DRY_RUN; then
|
||||
log_change "[dry-run] would write $env_dir/olp.conf"
|
||||
return 0
|
||||
fi
|
||||
mkdir -p "$env_dir" 2>/dev/null
|
||||
# D74 P1-2: systemd environment.d format is KEY=VALUE per line. While
|
||||
# systemd does its own parsing (no shell sourcing), reject embedded
|
||||
# newlines defensively — validate_olp_token already enforces the
|
||||
# restricted charset for the API key, so this is belt-and-braces.
|
||||
if [[ "$base_url" == *$'\n'* || "$key" == *$'\n'* ]]; then
|
||||
log_err "Refusing to write environment.d entry: value contains newline."
|
||||
return 2
|
||||
fi
|
||||
{
|
||||
echo "OPENAI_BASE_URL=$base_url/v1"
|
||||
if [[ -n "$key" ]]; then
|
||||
echo "OPENAI_API_KEY=$key"
|
||||
fi
|
||||
} > "$env_dir/olp.conf"
|
||||
log_ok "Wrote $env_dir/olp.conf (applies to systemd user services after re-login)"
|
||||
fi
|
||||
}
|
||||
|
||||
# ── Main ────────────────────────────────────────────────────────────────────
|
||||
|
||||
main() {
|
||||
local host="" port=4567 key=""
|
||||
|
||||
# Parse args (POSIX-style; --flag value AND --flag=value both accepted)
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--port) port="${2:?--port requires a value}"; shift 2 ;;
|
||||
--port=*) port="${1#*=}"; shift ;;
|
||||
--key) key="${2:?--key requires a value}"
|
||||
[[ -z "$key" ]] && { log_err "--key cannot be empty (omit --key for zero-config / auto-discovery)"; exit 1; }
|
||||
# D74 P1-2: reject malformed --key before it ever reaches an rc write.
|
||||
validate_olp_token "$key" "--key flag" || exit 1
|
||||
shift 2 ;;
|
||||
--key=*) key="${1#*=}"
|
||||
[[ -z "$key" ]] && { log_err "--key cannot be empty (omit --key for zero-config / auto-discovery)"; exit 1; }
|
||||
validate_olp_token "$key" "--key flag" || exit 1
|
||||
shift ;;
|
||||
--no-system-env) NO_SYSTEM_ENV=true; shift ;;
|
||||
--dry-run) DRY_RUN=true; shift ;;
|
||||
--version) show_version; exit 0 ;;
|
||||
--help|-h) show_help; exit 0 ;;
|
||||
--*) log_err "Unknown option: $1"; show_help >&2; exit 1 ;;
|
||||
*) host="$1"; shift ;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [[ -z "$host" ]]; then
|
||||
log_err "host IP is required."
|
||||
echo "" >&2
|
||||
show_help >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! [[ "$host" =~ ^[a-zA-Z0-9._-]+$ ]]; then
|
||||
log_err "invalid host '$host'"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Dependency check
|
||||
for cmd in curl python3; do
|
||||
if ! command -v "$cmd" &>/dev/null; then
|
||||
log_err "'$cmd' is required but not found in PATH."
|
||||
[[ "$cmd" == "python3" ]] && log_err " python3 is used for /health + /v1/models JSON parsing."
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
|
||||
local base_url="http://$host:$port"
|
||||
|
||||
echo "olp-connect v$OLP_CONNECT_VERSION"
|
||||
echo "─────────────────────────────────────"
|
||||
log_info "Remote: $base_url"
|
||||
$DRY_RUN && log_info "Mode: DRY RUN (no files will be modified, no env vars will be set)"
|
||||
echo ""
|
||||
|
||||
# Step 1: connectivity probe. We capture status separately from body so we
|
||||
# can distinguish "TCP/HTTP unreachable" from "reached but 401" (the latter
|
||||
# is a known surface when the server has auth.allow_anonymous=false AND
|
||||
# auth.advertise_anonymous_key=false — user MUST provide --key).
|
||||
log_step "Probing /health..."
|
||||
local probe_body probe_status
|
||||
probe_body=$(curl -s --max-time 5 -o /tmp/olp-connect-health.$$ -w "%{http_code}" "$base_url/health" 2>/dev/null || echo "000")
|
||||
probe_status="$probe_body"
|
||||
if [[ -f /tmp/olp-connect-health.$$ ]]; then
|
||||
probe_body=$(cat /tmp/olp-connect-health.$$ 2>/dev/null || echo "")
|
||||
rm -f /tmp/olp-connect-health.$$
|
||||
fi
|
||||
if [[ "$probe_status" == "000" ]]; then
|
||||
log_err "Cannot reach $base_url/health (connection refused / timeout / DNS)"
|
||||
log_err " Ensure OLP is running on $host and bound to 0.0.0.0 (LAN mode)."
|
||||
log_err " Default port changed 3456 → 4567 at OLP v0.4.0; pass --port 3456 for older installs."
|
||||
exit 2
|
||||
fi
|
||||
if [[ "$probe_status" == "401" ]]; then
|
||||
log_warn "Server reachable but /health returned 401."
|
||||
log_warn " Either the operator has not enabled auth.advertise_anonymous_key, or"
|
||||
log_warn " the server requires auth (auth.allow_anonymous=false)."
|
||||
if [[ -z "$key" ]]; then
|
||||
log_err " Pass --key olp_... to continue, or ask the operator to advertise an anonymous key (see ADR 0011)."
|
||||
exit 2
|
||||
fi
|
||||
# If user supplied --key, we proceed without /health body (auth-required mode).
|
||||
log_info " Proceeding with the --key you supplied; skipping /health.anonymousKey discovery."
|
||||
local health_json=""
|
||||
local remote_version="?"
|
||||
else
|
||||
if [[ "$probe_status" != "200" ]]; then
|
||||
log_err "/health returned HTTP $probe_status (expected 200 or 401)."
|
||||
exit 2
|
||||
fi
|
||||
local health_json="$probe_body"
|
||||
local remote_version
|
||||
remote_version=$(echo "$health_json" | python3 -c "import sys,json
|
||||
try: print(json.loads(sys.stdin.read()).get('version','?'))
|
||||
except: print('?')" 2>/dev/null || echo "?")
|
||||
log_ok "Connected — OLP v$remote_version"
|
||||
fi
|
||||
|
||||
# Step 2: auth resolution
|
||||
if [[ -z "$key" ]]; then
|
||||
# Try /health.anonymousKey first (D69 / ADR 0011 opt-in).
|
||||
local anon_key
|
||||
anon_key=$(echo "$health_json" | python3 -c "import sys,json
|
||||
try:
|
||||
d = json.loads(sys.stdin.read())
|
||||
k = d.get('anonymousKey')
|
||||
print(k if isinstance(k, str) and k else '')
|
||||
except: print('')" 2>/dev/null || echo "")
|
||||
if [[ -n "$anon_key" ]]; then
|
||||
# D74 P1-2: validate server-advertised token shape before consuming.
|
||||
# A hostile or misconfigured server could otherwise inject arbitrary
|
||||
# strings into the user's rc file via the `anonymousKey` field.
|
||||
if ! validate_olp_token "$anon_key" "/health.anonymousKey from ${host}:${port}"; then
|
||||
log_err "Refusing to consume malformed advertised key. Use --key explicitly or contact the OLP operator."
|
||||
exit 2
|
||||
fi
|
||||
key="$anon_key"
|
||||
log_ok "Using server-advertised anonymous key: $(key_display "$key")"
|
||||
log_info " (set by remote via auth.advertise_anonymous_key=true; see ADR 0011 for"
|
||||
log_info " the trusted-LAN-only invariant — this assumes you and the remote are"
|
||||
log_info " on the same trusted network)"
|
||||
else
|
||||
# No advertised key; prompt interactively (skip in dry-run for non-TTY safety)
|
||||
if $DRY_RUN; then
|
||||
log_info "[dry-run] would prompt for API key here (no --key + no anonymousKey)"
|
||||
key="<prompted-at-runtime>"
|
||||
else
|
||||
echo ""
|
||||
log_info "Remote does not advertise an anonymous key."
|
||||
log_info "Ask the OLP operator to run on the server: olp-keys keygen --name <your-label>"
|
||||
printf " Enter OLP API key: "
|
||||
{ read -rs key </dev/tty; } 2>/dev/null || key=""
|
||||
echo
|
||||
if [[ -z "$key" ]]; then
|
||||
log_err "No key provided and the remote did not advertise an anonymous key."
|
||||
log_err " Re-run with: olp-connect $host --key olp_..."
|
||||
exit 2
|
||||
fi
|
||||
# D74 P1-2: also validate the interactively-prompted key.
|
||||
validate_olp_token "$key" "interactive prompt" || exit 1
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# Step 3: smoke test /v1/models
|
||||
log_step "Smoke-testing /v1/models..."
|
||||
if $DRY_RUN && [[ "$key" == "<prompted-at-runtime>" ]]; then
|
||||
log_info "[dry-run] skipping smoke test (no real key)"
|
||||
else
|
||||
local models_out models_ok=0
|
||||
if [[ -n "$key" ]]; then
|
||||
models_out=$(curl -sf --max-time 10 \
|
||||
-H "Authorization: Bearer $key" \
|
||||
"$base_url/v1/models" 2>/dev/null) && models_ok=1
|
||||
else
|
||||
models_out=$(curl -sf --max-time 10 "$base_url/v1/models" 2>/dev/null) && models_ok=1
|
||||
fi
|
||||
if [[ $models_ok -eq 0 ]]; then
|
||||
log_err "/v1/models request failed — key may be invalid, revoked, or not allowed for any provider."
|
||||
exit 2
|
||||
fi
|
||||
local model_count
|
||||
model_count=$(echo "$models_out" | python3 -c "import sys,json
|
||||
try: print(len(json.loads(sys.stdin.read()).get('data', [])))
|
||||
except: print('?')" 2>/dev/null || echo "?")
|
||||
log_ok "/v1/models OK — $model_count models available"
|
||||
fi
|
||||
echo ""
|
||||
|
||||
# Step 4: write shell rc files
|
||||
log_step "Writing shell rc files..."
|
||||
local rc_files=()
|
||||
while IFS= read -r line; do
|
||||
[[ -n "$line" ]] && rc_files+=("$line")
|
||||
done < <(detect_rc_files)
|
||||
|
||||
if [[ ${#rc_files[@]} -eq 0 ]]; then
|
||||
log_warn "No shell rc files detected; falling back to ~/.bashrc"
|
||||
rc_files=("$HOME/.bashrc")
|
||||
fi
|
||||
|
||||
for rc_file in "${rc_files[@]}"; do
|
||||
log_change "stripping old OLP block from $(basename "$rc_file") (idempotent)"
|
||||
strip_olp_block "$rc_file"
|
||||
log_change "appending new OLP block to $(basename "$rc_file")"
|
||||
append_olp_block "$rc_file" "$base_url" "$key"
|
||||
done
|
||||
log_ok "Shell rc files updated:"
|
||||
for rc_file in "${rc_files[@]}"; do
|
||||
log_info " $rc_file"
|
||||
done
|
||||
echo ""
|
||||
|
||||
# Step 5: system-level env (macOS launchctl / Linux systemd)
|
||||
log_step "Setting system-level env..."
|
||||
set_system_env "$base_url" "$key"
|
||||
echo ""
|
||||
|
||||
# Step 6: IDE detection + per-IDE config
|
||||
log_step "Detecting installed IDEs..."
|
||||
detect_claude_code
|
||||
detect_cline "$base_url" "$key"
|
||||
detect_continue "$base_url" "$key"
|
||||
detect_cursor "$base_url" "$key"
|
||||
detect_aider
|
||||
detect_openclaw
|
||||
echo ""
|
||||
|
||||
# Step 7: final summary
|
||||
log_step "Done."
|
||||
log_info "OLP base URL: $base_url/v1"
|
||||
log_info "OLP API key: $(key_display "$key")"
|
||||
log_info ""
|
||||
log_info "Test it: open a fresh shell, run:"
|
||||
log_info " curl -sf -H \"Authorization: Bearer \$OPENAI_API_KEY\" $base_url/v1/models | python3 -m json.tool | head -20"
|
||||
log_info ""
|
||||
log_info "Reload your current shell to apply env changes:"
|
||||
for rc_file in "${rc_files[@]}"; do
|
||||
log_info " source $rc_file"
|
||||
done
|
||||
}
|
||||
|
||||
main "$@"
|
||||
+33
-4
@@ -79,6 +79,7 @@ const USAGE = `OLP key management CLI
|
||||
Usage:
|
||||
olp-keys keygen --owner [--name=<label>] [--providers=<csv>] [--force]
|
||||
olp-keys keygen --name=<label> [--tier=guest|owner] [--providers=<csv>]
|
||||
olp-keys keygen --anonymous --advertise [--name=<label>] [--providers=<csv>]
|
||||
olp-keys list [--owner-only] [--include-revoked]
|
||||
olp-keys revoke --id=<key-id>
|
||||
|
||||
@@ -86,7 +87,8 @@ Common flags:
|
||||
--olp-home=<path> Override ~/.olp (default reads OLP_HOME env)
|
||||
--help Print this message
|
||||
|
||||
Authority: ADR 0007 § 9 (bootstrap & recovery).`;
|
||||
Authority: ADR 0007 § 9 (bootstrap & recovery); ADR 0011 (anonymous-key
|
||||
deployment-context limits — trusted-LAN-only invariant for --advertise).`;
|
||||
|
||||
// ── Subcommand implementations ────────────────────────────────────────────
|
||||
|
||||
@@ -94,16 +96,36 @@ async function cmdKeygen(flags, ioOut, ioErr) {
|
||||
const olpHome = flags['olp-home'];
|
||||
const owner = flags.owner === true;
|
||||
const force = flags.force === true;
|
||||
// D69 (ADR 0011): --anonymous is shorthand for "create a guest-tier key
|
||||
// intended to be the zero-config /health.anonymousKey advertise key".
|
||||
// It implies --tier=guest and defaults the name to 'anonymous'. The
|
||||
// distinct field that actually triggers /health advertisement is
|
||||
// --advertise (writes plaintext_advertise into the manifest). Either
|
||||
// flag works on its own (--anonymous without --advertise is just a
|
||||
// conventionally-named guest key); --advertise without --anonymous is
|
||||
// accepted (operator may want to advertise a named guest key).
|
||||
const isAnonymous = flags.anonymous === true;
|
||||
const advertise = flags.advertise === true;
|
||||
let tier = flags.tier;
|
||||
if (owner) tier = 'owner';
|
||||
if (isAnonymous && !owner) tier = 'guest';
|
||||
if (!tier) tier = 'guest';
|
||||
if (tier !== 'owner' && tier !== 'guest') {
|
||||
ioErr(`Error: --tier must be "owner" or "guest" (got "${tier}").\n`);
|
||||
return 1;
|
||||
}
|
||||
const name = flags.name || (owner ? 'owner' : null);
|
||||
// D69: reject --owner --advertise (would expose owner identity unauthenticated).
|
||||
if (advertise && tier !== 'guest') {
|
||||
ioErr(`Error: --advertise requires guest tier (cannot advertise owner-tier key plaintext). See ADR 0011.\n`);
|
||||
return 1;
|
||||
}
|
||||
let name = flags.name;
|
||||
if (!name) {
|
||||
ioErr('Error: --name is required (or use --owner to default to "owner").\n');
|
||||
if (owner) name = 'owner';
|
||||
else if (isAnonymous) name = 'anonymous';
|
||||
}
|
||||
if (!name) {
|
||||
ioErr('Error: --name is required (or use --owner to default to "owner", or --anonymous to default to "anonymous").\n');
|
||||
return 1;
|
||||
}
|
||||
const providersFlag = flags.providers;
|
||||
@@ -136,7 +158,7 @@ async function cmdKeygen(flags, ioOut, ioErr) {
|
||||
|
||||
let result;
|
||||
try {
|
||||
result = createKey({ name, owner_tier: tier, providers_enabled, olpHome });
|
||||
result = createKey({ name, owner_tier: tier, providers_enabled, olpHome, plaintext_advertise: advertise });
|
||||
} catch (err) {
|
||||
ioErr(`Error: createKey failed: ${err?.message ?? err}\n`);
|
||||
return 2;
|
||||
@@ -150,6 +172,13 @@ async function cmdKeygen(flags, ioOut, ioErr) {
|
||||
ioOut(` providers_enabled: ${typeof result.manifest.providers_enabled === 'string' ? result.manifest.providers_enabled : `[${result.manifest.providers_enabled.join(', ')}]`}\n`);
|
||||
ioOut(` created_at: ${result.manifest.created_at}\n`);
|
||||
ioOut(` manifest: ~/.olp/keys/${result.id}/manifest.json\n`);
|
||||
if (advertise) {
|
||||
// D69 (ADR 0011): explicit warning when plaintext lands on disk + opt-in surface.
|
||||
ioOut(` advertise: YES — plaintext stored in manifest; surfaced via /health.anonymousKey\n`);
|
||||
ioErr(`\n WARNING: this key's plaintext is now stored on disk + will be exposed via\n`);
|
||||
ioErr(` /health.anonymousKey when auth.advertise_anonymous_key=true AND\n`);
|
||||
ioErr(` auth.allow_anonymous=true. Use ONLY on a trusted LAN. See ADR 0011.\n`);
|
||||
}
|
||||
ioOut(`\n token (plaintext): ${result.plaintext_token}\n\n`);
|
||||
ioOut(` Pass via: Authorization: Bearer ${result.plaintext_token.slice(0, 12)}...\n`);
|
||||
ioOut(` or: x-api-key: ${result.plaintext_token.slice(0, 12)}...\n\n`);
|
||||
|
||||
Executable
+859
@@ -0,0 +1,859 @@
|
||||
#!/usr/bin/env node
|
||||
/**
|
||||
* bin/olp.mjs — OLP operator CLI (Phase 4 / D64)
|
||||
*
|
||||
* Authority: ADR 0010 § Phase 4 D64-D67. Ports OCP's `ocp` bash wrapper
|
||||
* (https://github.com/dtzp555-max/ocp /ocp) to Node.js, eliminating the
|
||||
* python3 JSON-parsing fragility called out in the ADR.
|
||||
*
|
||||
* Subcommands:
|
||||
* status GET /v0/management/status (owner-only)
|
||||
* health GET /health
|
||||
* usage GET /v0/management/dashboard-data (owner-only)
|
||||
* models GET /v1/models
|
||||
* cache GET /cache/stats (owner-only)
|
||||
* providers local: models-registry + config providers.enabled
|
||||
* chain show [<model>] local: ~/.olp/config.json routing.chains
|
||||
* logs [N] [--level X] local: read ~/.olp/logs/audit.ndjson via audit-query
|
||||
* restart launchctl (macOS) / systemctl --user (Linux)
|
||||
* doctor [--check X] run lib/doctor.mjs runDoctor + format
|
||||
* keys ... delegate to bin/olp-keys.mjs
|
||||
* help | --help usage
|
||||
*
|
||||
* Global flags:
|
||||
* --json emit raw JSON (silences human-readable output)
|
||||
* --proxy-url=<url> override resolved proxy URL
|
||||
* --olp-home=<path> override ~/.olp
|
||||
*
|
||||
* URL resolution:
|
||||
* OLP_PROXY_URL env (full URL) → http://127.0.0.1:${OLP_PORT || 4567}
|
||||
*
|
||||
* Auth (Bearer token) resolution:
|
||||
* 1. OLP_API_KEY env
|
||||
* 2. OLP_OWNER_TOKEN env (synthetic env-owner per ADR 0007 § 9.4)
|
||||
* 3. Most recently used active owner-tier key from listKeys() ← plaintext NOT recoverable from disk
|
||||
*
|
||||
* Note: the third option only works during the same session in which `olp-keys
|
||||
* keygen --owner` was run if the operator captured the token + set OLP_OWNER_TOKEN.
|
||||
* listKeys() returns manifests; manifest.token_hash is one-way. The CLI therefore
|
||||
* reports "no owner token available" + remediation instructions if env vars are
|
||||
* absent — it does NOT try to crack the hash.
|
||||
*
|
||||
* Exit codes:
|
||||
* 0 = success
|
||||
* 1 = bad usage / unknown subcommand
|
||||
* 2 = network or HTTP error (4xx/5xx)
|
||||
* 3 = auth missing / forbidden
|
||||
*/
|
||||
|
||||
import { request as httpRequest } from 'node:http';
|
||||
import { request as httpsRequest } from 'node:https';
|
||||
import { URL } from 'node:url';
|
||||
import { existsSync, readFileSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { spawn as spawnProc } from 'node:child_process';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import { realpathSync } from 'node:fs';
|
||||
|
||||
import { runDoctor, resolveProxyUrl, resolveOlpHome } from '../lib/doctor.mjs';
|
||||
import { listKeys } from '../lib/keys.mjs';
|
||||
import { readAuditWindow } from '../lib/audit-query.mjs';
|
||||
import modelsRegistry from '../models-registry.json' with { type: 'json' };
|
||||
import { runCli as runKeysCli } from './olp-keys.mjs';
|
||||
|
||||
// ── ANSI helpers (no chalk dep) ───────────────────────────────────────────
|
||||
|
||||
const ANSI = {
|
||||
reset: '\x1b[0m',
|
||||
bold: '\x1b[1m',
|
||||
dim: '\x1b[2m',
|
||||
red: '\x1b[31m',
|
||||
green: '\x1b[32m',
|
||||
yellow: '\x1b[33m',
|
||||
blue: '\x1b[34m',
|
||||
cyan: '\x1b[36m',
|
||||
gray: '\x1b[90m',
|
||||
};
|
||||
|
||||
function colorize(s, code, useColor) {
|
||||
if (!useColor) return s;
|
||||
return `${code}${s}${ANSI.reset}`;
|
||||
}
|
||||
|
||||
function statusBadge(status, useColor) {
|
||||
const map = {
|
||||
ok: { txt: 'PASS', col: ANSI.green },
|
||||
warn: { txt: 'WARN', col: ANSI.yellow },
|
||||
fail: { txt: 'FAIL', col: ANSI.red },
|
||||
};
|
||||
const m = map[status] ?? { txt: String(status).toUpperCase(), col: ANSI.gray };
|
||||
return colorize(`[${m.txt}]`, m.col, useColor);
|
||||
}
|
||||
|
||||
// ── Arg parser (mirror bin/olp-keys.mjs shape) ────────────────────────────
|
||||
|
||||
export function parseArgv(argv) {
|
||||
const positional = [];
|
||||
const flags = {};
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
const a = argv[i];
|
||||
if (a.startsWith('--')) {
|
||||
const eq = a.indexOf('=');
|
||||
if (eq > 0) {
|
||||
flags[a.slice(2, eq)] = a.slice(eq + 1);
|
||||
} else {
|
||||
const name = a.slice(2);
|
||||
const next = argv[i + 1];
|
||||
if (next !== undefined && !next.startsWith('--')) {
|
||||
flags[name] = next;
|
||||
i++;
|
||||
} else {
|
||||
flags[name] = true;
|
||||
}
|
||||
}
|
||||
} else {
|
||||
positional.push(a);
|
||||
}
|
||||
}
|
||||
return { positional, flags };
|
||||
}
|
||||
|
||||
// ── Output helpers (respect --json) ───────────────────────────────────────
|
||||
|
||||
function makeIO(opts) {
|
||||
const out = opts.out ?? (s => process.stdout.write(s));
|
||||
const err = opts.err ?? (s => process.stderr.write(s));
|
||||
const wantJson = opts.json === true;
|
||||
// When wantJson, the only stdout writer used is `emitJson`. `log` becomes a no-op
|
||||
// (debug noise suppression per the bundle requirements). `errln` always writes
|
||||
// to stderr.
|
||||
const log = (...parts) => { if (!wantJson) out(parts.join(' ') + '\n'); };
|
||||
const errln = (...parts) => err(parts.join(' ') + '\n');
|
||||
const emitJson = (obj) => out(JSON.stringify(obj, null, 2) + '\n');
|
||||
return { log, errln, emitJson, wantJson, useColor: !wantJson && (opts.useColor ?? true) };
|
||||
}
|
||||
|
||||
// ── HTTP helper ───────────────────────────────────────────────────────────
|
||||
|
||||
async function httpFetch(url, { method = 'GET', headers = {}, timeoutMs = 15000 } = {}) {
|
||||
return new Promise(resolve => {
|
||||
let done = false;
|
||||
const finish = (v) => { if (!done) { done = true; resolve(v); } };
|
||||
let urlObj;
|
||||
try { urlObj = new URL(url); }
|
||||
catch (e) { return finish({ ok: false, error: `invalid url: ${e?.message ?? e}` }); }
|
||||
const isHttps = urlObj.protocol === 'https:';
|
||||
const reqFn = isHttps ? httpsRequest : httpRequest;
|
||||
let req;
|
||||
try {
|
||||
req = reqFn(url, { method, headers, timeout: timeoutMs }, res => {
|
||||
let data = '';
|
||||
res.on('data', c => { data += c; });
|
||||
res.on('end', () => finish({ ok: true, status: res.statusCode, body: data, headers: res.headers }));
|
||||
});
|
||||
} catch (e) {
|
||||
return finish({ ok: false, error: String(e?.message ?? e) });
|
||||
}
|
||||
req.on('error', e => finish({ ok: false, error: String(e?.message ?? e), code: e?.code }));
|
||||
req.on('timeout', () => {
|
||||
try { req.destroy(new Error(`timeout after ${timeoutMs}ms`)); } catch { /* ignore */ }
|
||||
});
|
||||
req.end();
|
||||
});
|
||||
}
|
||||
|
||||
// ── Token resolution ──────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Resolve a Bearer token. Returns the plaintext token string or null.
|
||||
* Precedence: OLP_API_KEY → OLP_OWNER_TOKEN → null (manifests are one-way hashed).
|
||||
*/
|
||||
export function resolveBearerToken() {
|
||||
if (process.env.OLP_API_KEY) return process.env.OLP_API_KEY;
|
||||
if (process.env.OLP_OWNER_TOKEN) return process.env.OLP_OWNER_TOKEN;
|
||||
return null;
|
||||
}
|
||||
|
||||
function authHeaders() {
|
||||
const tok = resolveBearerToken();
|
||||
return tok ? { Authorization: `Bearer ${tok}` } : {};
|
||||
}
|
||||
|
||||
// ── HTTP-error → exit code mapping ────────────────────────────────────────
|
||||
|
||||
function httpErrorToExit(res, io) {
|
||||
if (!res.ok) {
|
||||
if (res.code === 'ECONNREFUSED' || (res.error && res.error.includes('ECONNREFUSED'))) {
|
||||
io.errln(`Error: OLP server unreachable (${res.error}). Is it running?`);
|
||||
io.errln(`Hint: run 'npx olp restart' (or 'npm start' for foreground) — see 'npx olp doctor' for the full diagnostic.`);
|
||||
return 2;
|
||||
}
|
||||
io.errln(`Error: network error: ${res.error}`);
|
||||
return 2;
|
||||
}
|
||||
if (res.status === 401) {
|
||||
io.errln(`Error: 401 unauthorized — set OLP_API_KEY env (Bearer token) or OLP_OWNER_TOKEN.`);
|
||||
io.errln(`Hint: 'npx olp-keys keygen --owner' creates a new owner-tier key (capture the plaintext token).`);
|
||||
return 3;
|
||||
}
|
||||
if (res.status === 403) {
|
||||
io.errln(`Error: 403 forbidden — current key is not owner-tier (this endpoint is owner-only).`);
|
||||
return 3;
|
||||
}
|
||||
if (res.status >= 400) {
|
||||
io.errln(`Error: HTTP ${res.status}: ${res.body.slice(0, 200)}`);
|
||||
return 2;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Human-readable formatters ─────────────────────────────────────────────
|
||||
|
||||
function formatBytes(n) {
|
||||
if (typeof n !== 'number' || !Number.isFinite(n)) return '?';
|
||||
if (n < 1024) return `${n}B`;
|
||||
if (n < 1024 * 1024) return `${(n / 1024).toFixed(1)}K`;
|
||||
if (n < 1024 * 1024 * 1024) return `${(n / 1024 / 1024).toFixed(1)}M`;
|
||||
return `${(n / 1024 / 1024 / 1024).toFixed(1)}G`;
|
||||
}
|
||||
|
||||
function formatMs(ms) {
|
||||
if (typeof ms !== 'number' || !Number.isFinite(ms)) return '?';
|
||||
if (ms < 1000) return `${ms}ms`;
|
||||
if (ms < 60000) return `${(ms / 1000).toFixed(1)}s`;
|
||||
if (ms < 3600000) return `${Math.floor(ms / 60000)}m${Math.floor((ms % 60000) / 1000)}s`;
|
||||
return `${Math.floor(ms / 3600000)}h${Math.floor((ms % 3600000) / 60000)}m`;
|
||||
}
|
||||
|
||||
/** formatAgo(diffMs) — "N min ago" / "Nh ago" from a millisecond diff. */
|
||||
function formatAgo(diffMs) {
|
||||
if (typeof diffMs !== 'number' || diffMs < 0) return 'just now';
|
||||
const sec = Math.floor(diffMs / 1000);
|
||||
if (sec < 60) return `${sec}s ago`;
|
||||
const min = Math.floor(sec / 60);
|
||||
if (min < 60) return `${min}m ago`;
|
||||
return `${Math.floor(min / 60)}h ago`;
|
||||
}
|
||||
|
||||
/**
|
||||
* formatResetCountdown(epochSeconds) → human-readable reset countdown string.
|
||||
*
|
||||
* Mirrors dashboard.html formatResetCountdown(). Five ranges:
|
||||
* past / < 1h / < 24h / < 7d / ≥ 7d
|
||||
*
|
||||
* Authority: ADR 0008 Amendment 2 (quota_v2 shape); ported from
|
||||
* dashboard.html (D82). No external deps. Pure formatter.
|
||||
*
|
||||
* @param {number|null} epochSeconds — Unix epoch seconds for reset time
|
||||
* @returns {string}
|
||||
*/
|
||||
export function formatResetCountdown(epochSeconds) {
|
||||
if (epochSeconds == null) return '—';
|
||||
const nowMs = Date.now();
|
||||
const targetMs = epochSeconds * 1000;
|
||||
const diffMs = targetMs - nowMs;
|
||||
if (diffMs <= 0) return 'resetting now';
|
||||
const diffMin = Math.floor(diffMs / 60000);
|
||||
const diffHr = Math.floor(diffMin / 60);
|
||||
const diffDay = Math.floor(diffHr / 24);
|
||||
if (diffMin < 60) return `resets in ${diffMin}m`;
|
||||
if (diffHr < 24) {
|
||||
const remMin = diffMin - diffHr * 60;
|
||||
if (remMin === 0) return `resets in ${diffHr}h`;
|
||||
return `resets in ${diffHr}h ${remMin}m`;
|
||||
}
|
||||
const target = new Date(targetMs);
|
||||
const timeStr = target.toLocaleString('en-US', { hour: 'numeric', minute: '2-digit', hour12: true });
|
||||
if (diffDay < 7) {
|
||||
const dayStr = target.toLocaleString('en-US', { weekday: 'short' });
|
||||
return `resets ${dayStr} ${timeStr}`;
|
||||
}
|
||||
const dateStr = target.toLocaleString('en-US', { month: 'short', day: 'numeric' });
|
||||
return `resets ${dateStr} ${timeStr}`;
|
||||
}
|
||||
|
||||
// ── Subcommand: status ────────────────────────────────────────────────────
|
||||
|
||||
async function cmdStatus(flags, io) {
|
||||
const url = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
||||
const res = await httpFetch(`${url}/v0/management/status`, { headers: authHeaders() });
|
||||
const ec = httpErrorToExit(res, io);
|
||||
if (ec !== 0) return ec;
|
||||
let body;
|
||||
try { body = JSON.parse(res.body); }
|
||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||
io.log(colorize('OLP status', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
io.log(` version: ${body.version}`);
|
||||
io.log(` uptime: ${body.uptime_human} (${formatMs(body.uptime_ms)})`);
|
||||
io.log(` started: ${body.started_at}`);
|
||||
io.log(` providers: ${body.providers?.enabled ?? '?'} enabled / ${body.providers?.available ?? '?'} available`);
|
||||
if (body.providers?.status && typeof body.providers.status === 'object') {
|
||||
for (const [name, s] of Object.entries(body.providers.status)) {
|
||||
const okIcon = s?.ok ? colorize('ok', ANSI.green, io.useColor) : colorize('FAIL', ANSI.red, io.useColor);
|
||||
io.log(` - ${name.padEnd(12)} ${okIcon} ${s?.error ? `(${s.error})` : ''}`);
|
||||
}
|
||||
}
|
||||
io.log(` total reqs: ${body.stats?.total_requests ?? 0}`);
|
||||
io.log(` active reqs: ${body.stats?.active_requests ?? 0}`);
|
||||
// D75 F4 fix: server payload nests cache stats under stats.cache (per
|
||||
// server.mjs handleManagementStatus, ~line 2092). The CacheStore.stats()
|
||||
// contract returns { hits, misses, size, inflightCount } per
|
||||
// lib/cache/store.mjs — there is no `entries` field. Pre-D75 cmdStatus read
|
||||
// `body.stats.cache.entries` (OCP-era) which was always undefined → output
|
||||
// showed "entries=?". Same pattern as D74 P2-3 applied to cmdCache/cmdUsage.
|
||||
if (body.stats?.cache) {
|
||||
const c = body.stats.cache;
|
||||
io.log(` cache: hits=${c.hits ?? 0} misses=${c.misses ?? 0} entries=${c.size ?? 0}${typeof c.inflightCount === 'number' ? ` inflight=${c.inflightCount}` : ''}`);
|
||||
}
|
||||
if (Array.isArray(body.recent_errors) && body.recent_errors.length > 0) {
|
||||
io.log(` recent errors: ${body.recent_errors.length}`);
|
||||
for (const e of body.recent_errors.slice(0, 5)) {
|
||||
io.log(` - [${e.at ?? '?'}] ${e.provider ?? '?'} ${e.path ?? '?'} — ${(e.message ?? '').slice(0, 80)}`);
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Subcommand: health ────────────────────────────────────────────────────
|
||||
|
||||
async function cmdHealth(flags, io) {
|
||||
const url = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
||||
const res = await httpFetch(`${url}/health`, { headers: authHeaders() });
|
||||
const ec = httpErrorToExit(res, io);
|
||||
if (ec !== 0) return ec;
|
||||
let body;
|
||||
try { body = JSON.parse(res.body); }
|
||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||
io.log(colorize(`OLP /health → ${res.status}`, ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const [k, v] of Object.entries(body)) {
|
||||
if (typeof v === 'object' && v !== null) {
|
||||
io.log(` ${k}:`);
|
||||
for (const [k2, v2] of Object.entries(v)) {
|
||||
io.log(` ${k2}: ${typeof v2 === 'object' ? JSON.stringify(v2) : v2}`);
|
||||
}
|
||||
} else {
|
||||
io.log(` ${k}: ${v}`);
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Subcommand: usage ─────────────────────────────────────────────────────
|
||||
|
||||
async function cmdUsage(flags, io) {
|
||||
const url = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
||||
const res = await httpFetch(`${url}/v0/management/dashboard-data`, { headers: authHeaders() });
|
||||
const ec = httpErrorToExit(res, io);
|
||||
if (ec !== 0) return ec;
|
||||
let body;
|
||||
try { body = JSON.parse(res.body); }
|
||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||
// D74 P2-3 fix: server payload shape is { generated_at, window_24h: { request_count, status_2xx,
|
||||
// status_4xx, status_5xx, by_provider, by_owner_tier, by_path, median_latency_ms, p95_latency_ms },
|
||||
// cache_hit_24h: { total, hit, miss, bypass, streaming_attached, hit_rate, by_provider }, quota: [{provider, ...}],
|
||||
// spend_trend_30d: [{date, request_count, by_provider}], top_fallback_chains_24h: [{chain, count, ...}],
|
||||
// cache_stats: { hits, misses, size, inflightCount } } per server.mjs:2027 + lib/audit-query.mjs.
|
||||
io.log(colorize('OLP usage (24h)', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
const w24 = body.window_24h ?? {};
|
||||
const cache24 = body.cache_hit_24h ?? {};
|
||||
if (typeof w24 === 'object' && (w24.request_count ?? 0) > 0) {
|
||||
io.log(` requests: ${w24.request_count}`);
|
||||
io.log(` 2xx / 4xx / 5xx: ${w24.status_2xx ?? 0} / ${w24.status_4xx ?? 0} / ${w24.status_5xx ?? 0}`);
|
||||
if (typeof w24.median_latency_ms === 'number') {
|
||||
io.log(` latency p50/p95: ${w24.median_latency_ms}ms / ${w24.p95_latency_ms ?? 0}ms`);
|
||||
}
|
||||
if (typeof cache24.hit_rate === 'number') {
|
||||
const pct = (cache24.hit_rate * 100).toFixed(1);
|
||||
io.log(` cache hit rate: ${pct}% (hit=${cache24.hit ?? 0} miss=${cache24.miss ?? 0}${cache24.streaming_attached ? ` streaming_attached=${cache24.streaming_attached}` : ''})`);
|
||||
}
|
||||
} else {
|
||||
io.log(' (no 24h usage data — server may not have processed any requests yet)');
|
||||
}
|
||||
// F4 (v0.5.1 codex post-release review Q4): prefer quota_v2 when present
|
||||
// (server v0.5.0+), fall back to legacy quota array on older servers.
|
||||
// Authority: ADR 0008 Amendment 2 (quota_v2 shape).
|
||||
if (Array.isArray(body.quota_v2) && body.quota_v2.length > 0) {
|
||||
io.log('');
|
||||
io.log(colorize('Per-provider quota (live)', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const p of body.quota_v2) {
|
||||
const label = String(p.provider ?? '?').toUpperCase().padEnd(12);
|
||||
const status = p.status ?? 'unavailable';
|
||||
if (status === 'unavailable') {
|
||||
io.log(` ${colorize(label, ANSI.gray, io.useColor)} unavailable ${p.reason ?? 'no public quota api'}`);
|
||||
} else if (status === 'unreachable') {
|
||||
const fk = p.failure?.kind ?? 'unknown';
|
||||
const fm = p.failure?.message ?? 'probe failed';
|
||||
io.log(` ${colorize(label, ANSI.red, io.useColor)} ❌ no cached data — failure: ${fk} (${fm})`);
|
||||
} else {
|
||||
// live or stale
|
||||
const staleWarn = status === 'stale'
|
||||
? colorize(` ⚠ stale${p.last_fresh_at ? ` (${formatAgo(Date.now() - p.last_fresh_at)})` : ''} failure: ${p.failure?.kind ?? 'unknown'}`, ANSI.yellow, io.useColor)
|
||||
: '';
|
||||
const util = p.utilization ?? {};
|
||||
const reset = p.reset ?? {};
|
||||
const parts = [];
|
||||
for (const window of ['5h', '7d']) {
|
||||
const frac = util[window];
|
||||
const resetEpoch = reset[window];
|
||||
if (frac != null) {
|
||||
const pct = `${Math.round(frac * 100)}%`;
|
||||
const rst = formatResetCountdown(resetEpoch);
|
||||
parts.push(`${window}: ${colorize(pct, frac >= 0.8 ? ANSI.red : frac >= 0.5 ? ANSI.yellow : ANSI.green, io.useColor)} (${rst})`);
|
||||
}
|
||||
}
|
||||
const binding = p.representative_claim ? ` binding: ${p.representative_claim.replace('_', '-')}` : '';
|
||||
io.log(` ${colorize(label, ANSI.bold, io.useColor)} ${colorize(status, status === 'live' ? ANSI.green : ANSI.yellow, io.useColor).padEnd(6)} ${parts.join(' ')}${binding}${staleWarn}`);
|
||||
}
|
||||
}
|
||||
} else if (Array.isArray(body.quota) && body.quota.length > 0) {
|
||||
// Legacy fallback for pre-v0.5.0 servers
|
||||
io.log('');
|
||||
io.log(colorize('Per-provider quota', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const p of body.quota) {
|
||||
const label = String(p.provider ?? '?').padEnd(12);
|
||||
if (p.error) {
|
||||
io.log(` ${label} error: ${p.error}`);
|
||||
} else if (typeof p.percent_used === 'number') {
|
||||
io.log(` ${label} ${p.percent_used}% used${p.resets_in_human ? ` (resets in ${p.resets_in_human})` : ''}`);
|
||||
} else if (p.available === false) {
|
||||
io.log(` ${label} unavailable`);
|
||||
} else {
|
||||
io.log(` ${label} no quota api`);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (Array.isArray(body.top_fallback_chains_24h) && body.top_fallback_chains_24h.length > 0) {
|
||||
io.log('');
|
||||
io.log(colorize('Top fallback chains (24h)', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const f of body.top_fallback_chains_24h.slice(0, 10)) {
|
||||
io.log(` ${String(f.count ?? '?').padStart(5)} ${(f.chain ?? []).join(' → ')}`);
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Subcommand: models ────────────────────────────────────────────────────
|
||||
|
||||
async function cmdModels(flags, io) {
|
||||
const url = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
||||
const res = await httpFetch(`${url}/v1/models`, { headers: authHeaders() });
|
||||
const ec = httpErrorToExit(res, io);
|
||||
if (ec !== 0) return ec;
|
||||
let body;
|
||||
try { body = JSON.parse(res.body); }
|
||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||
io.log(colorize(`OLP models (${(body.data ?? []).length})`, ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const m of body.data ?? []) {
|
||||
io.log(` ${m.id.padEnd(35)} ${colorize(`(${m.owned_by ?? '?'})`, ANSI.gray, io.useColor)}`);
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Subcommand: cache ─────────────────────────────────────────────────────
|
||||
|
||||
async function cmdCache(flags, io) {
|
||||
const url = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
||||
const res = await httpFetch(`${url}/cache/stats`, { headers: authHeaders() });
|
||||
const ec = httpErrorToExit(res, io);
|
||||
if (ec !== 0) return ec;
|
||||
let body;
|
||||
try { body = JSON.parse(res.body); }
|
||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||
// D74 P2-3 fix: cacheStore.stats() returns { hits, misses, size, inflightCount }
|
||||
// per lib/cache/store.mjs:320. There is no entries / evictions / bytes / maxBytes
|
||||
// in the OLP cache model — those were OCP-era field names. Compute a hit rate from
|
||||
// the numerator/denominator instead of fabricating bytes.
|
||||
const hits = body.hits ?? 0;
|
||||
const misses = body.misses ?? 0;
|
||||
const denom = hits + misses;
|
||||
const hitRate = denom > 0 ? ((hits / denom) * 100).toFixed(1) : '0.0';
|
||||
io.log(colorize('OLP cache (live in-memory)', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
io.log(` entries: ${body.size ?? 0}`);
|
||||
io.log(` hits / misses: ${hits} / ${misses} (hit rate ${hitRate}%)`);
|
||||
io.log(` inflight: ${body.inflightCount ?? 0}`);
|
||||
if (body.generated_at) io.log(` generated_at: ${body.generated_at}`);
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Subcommand: providers (local) ─────────────────────────────────────────
|
||||
|
||||
function cmdProviders(flags, io) {
|
||||
const olpHome = resolveOlpHome({ olpHome: flags['olp-home'] });
|
||||
const configPath = join(olpHome, 'config.json');
|
||||
let enabled = {};
|
||||
try {
|
||||
const cfg = JSON.parse(readFileSync(configPath, 'utf8'));
|
||||
enabled = cfg?.providers?.enabled ?? {};
|
||||
} catch { /* fine — empty enabled map */ }
|
||||
|
||||
const providers = modelsRegistry?.providers ?? {};
|
||||
const rows = [];
|
||||
for (const [name, p] of Object.entries(providers)) {
|
||||
rows.push({
|
||||
name,
|
||||
displayName: p?.displayName ?? name,
|
||||
tier: p?.tier ?? '?',
|
||||
modelCount: (p?.models ?? []).length,
|
||||
enabled: enabled[name] === true,
|
||||
candidate: p?.candidate === true,
|
||||
});
|
||||
}
|
||||
if (io.wantJson) {
|
||||
io.emitJson({ providers: rows, config_path: configPath });
|
||||
return 0;
|
||||
}
|
||||
io.log(colorize(`OLP providers (${rows.length} in registry)`, ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const r of rows) {
|
||||
const enabledTxt = r.enabled
|
||||
? colorize('enabled ', ANSI.green, io.useColor)
|
||||
: colorize('disabled', ANSI.gray, io.useColor);
|
||||
const candTxt = r.candidate ? colorize('(candidate)', ANSI.yellow, io.useColor) : '';
|
||||
io.log(` ${r.name.padEnd(10)} ${enabledTxt} tier ${r.tier} models ${String(r.modelCount).padStart(2)} ${candTxt}`);
|
||||
}
|
||||
io.log('');
|
||||
io.log(colorize(`config: ${configPath}`, ANSI.dim, io.useColor));
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Subcommand: chain show ────────────────────────────────────────────────
|
||||
|
||||
function cmdChainShow(positional, flags, io) {
|
||||
const target = positional[0] ?? null; // model name, or null = print all
|
||||
const olpHome = resolveOlpHome({ olpHome: flags['olp-home'] });
|
||||
const configPath = join(olpHome, 'config.json');
|
||||
let chains = {};
|
||||
try {
|
||||
const cfg = JSON.parse(readFileSync(configPath, 'utf8'));
|
||||
chains = cfg?.routing?.chains ?? {};
|
||||
} catch { /* empty */ }
|
||||
|
||||
if (io.wantJson) {
|
||||
if (target) {
|
||||
io.emitJson({ model: target, chain: chains[target] ?? null });
|
||||
} else {
|
||||
io.emitJson({ chains });
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
io.log(colorize('OLP routing.chains', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
if (Object.keys(chains).length === 0) {
|
||||
io.log(` (no chains configured in ${configPath})`);
|
||||
return 0;
|
||||
}
|
||||
if (target) {
|
||||
const chain = chains[target];
|
||||
if (!chain) {
|
||||
io.errln(`Error: model "${target}" not in routing.chains (configured: ${Object.keys(chains).join(', ')}).`);
|
||||
return 1;
|
||||
}
|
||||
io.log(` ${target}:`);
|
||||
for (const hop of chain) {
|
||||
io.log(` → ${typeof hop === 'string' ? hop : JSON.stringify(hop)}`);
|
||||
}
|
||||
} else {
|
||||
for (const [model, chain] of Object.entries(chains)) {
|
||||
io.log(` ${model}:`);
|
||||
for (const hop of chain) {
|
||||
io.log(` → ${typeof hop === 'string' ? hop : JSON.stringify(hop)}`);
|
||||
}
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Subcommand: logs ──────────────────────────────────────────────────────
|
||||
|
||||
async function cmdLogs(positional, flags, io) {
|
||||
const n = positional[0] ? parseInt(positional[0], 10) : 20;
|
||||
if (!Number.isFinite(n) || n <= 0) {
|
||||
io.errln(`Error: invalid log count "${positional[0]}"`);
|
||||
return 1;
|
||||
}
|
||||
const olpHome = resolveOlpHome({ olpHome: flags['olp-home'] });
|
||||
// readAuditWindow is a generator over [startMs, endMs). Default window = last 24h.
|
||||
const windowMs = flags['window-ms'] ? parseInt(flags['window-ms'], 10) : 24 * 3600 * 1000;
|
||||
const endMs = Date.now();
|
||||
const startMs = endMs - windowMs;
|
||||
let events = [];
|
||||
try {
|
||||
for (const ev of readAuditWindow({ startMs, endMs, olpHome })) {
|
||||
events.push(ev);
|
||||
}
|
||||
} catch (e) {
|
||||
io.errln(`Error: readAuditWindow failed: ${e?.message ?? e}`);
|
||||
return 2;
|
||||
}
|
||||
let filtered = events;
|
||||
if (flags.level) {
|
||||
filtered = filtered.filter(e => e.level === flags.level);
|
||||
}
|
||||
// Tail (audit events are already chronological per generator order).
|
||||
filtered = filtered.slice(-n);
|
||||
if (io.wantJson) {
|
||||
io.emitJson({ events: filtered });
|
||||
return 0;
|
||||
}
|
||||
io.log(colorize(`OLP logs (last ${filtered.length} of ${events.length}${flags.level ? `, level=${flags.level}` : ''})`, ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const e of filtered) {
|
||||
// Audit event shape per lib/audit.mjs: { ts, event, ...data }. `level` is
|
||||
// not always present in audit ndjson (it is in stderr-side logEvent).
|
||||
const level = (e.level ?? 'info').toUpperCase();
|
||||
const levelColor =
|
||||
level === 'ERROR' ? ANSI.red
|
||||
: level === 'WARN' ? ANSI.yellow
|
||||
: ANSI.gray;
|
||||
const summary = e.message ?? e.error ?? '';
|
||||
io.log(` ${colorize(level.padEnd(5), levelColor, io.useColor)} ${e.ts ?? '?'} ${e.event ?? '?'} ${summary}`);
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
// ── Subcommand: restart ───────────────────────────────────────────────────
|
||||
|
||||
async function cmdRestart(flags, io) {
|
||||
// macOS: launchctl kickstart -k gui/$(id -u)/dev.olp.proxy
|
||||
// Linux: systemctl --user restart olp-proxy
|
||||
// Neither installed → fall through to a helpful error.
|
||||
//
|
||||
// CAVEAT (D64-D67 reviewer P2-2; known OCP institutional lesson per
|
||||
// ~/.cc-rules/memory/auto/MEMORY.md PIT INDEX): `launchctl kickstart -k`
|
||||
// does NOT re-read the plist's EnvironmentVariables block — launchd
|
||||
// sticks to its cached env from the most recent bootstrap. If you edited
|
||||
// ~/Library/LaunchAgents/dev.olp.proxy.plist's env, this subcommand will
|
||||
// silently use stale values. Use `launchctl bootout gui/<uid>/dev.olp.proxy`
|
||||
// followed by `launchctl bootstrap gui/<uid> ~/Library/LaunchAgents/dev.olp.proxy.plist`
|
||||
// to force a clean env reload. The Phase 4 installer (planned post-D73)
|
||||
// will expose `olp restart --full` for the bootout/bootstrap dance.
|
||||
const platform = process.platform;
|
||||
const uid = process.getuid?.() ?? null;
|
||||
let cmd, args;
|
||||
if (platform === 'darwin') {
|
||||
if (uid == null) {
|
||||
io.errln('Error: cannot resolve UID on this platform; cannot drive launchctl');
|
||||
return 2;
|
||||
}
|
||||
cmd = 'launchctl';
|
||||
args = ['kickstart', '-k', `gui/${uid}/dev.olp.proxy`];
|
||||
} else if (platform === 'linux') {
|
||||
cmd = 'systemctl';
|
||||
args = ['--user', 'restart', 'olp-proxy'];
|
||||
} else {
|
||||
io.errln(`Error: platform "${platform}" not supported for 'olp restart'.`);
|
||||
io.errln(`Hint: run 'npm start' (or whatever launches your OLP server) manually.`);
|
||||
return 2;
|
||||
}
|
||||
// Spawn + wait for exit; bubble up any error.
|
||||
const result = await new Promise(resolve => {
|
||||
const child = spawnProc(cmd, args, { stdio: io.wantJson ? 'ignore' : 'inherit' });
|
||||
child.on('error', e => resolve({ code: -1, error: e }));
|
||||
child.on('exit', code => resolve({ code }));
|
||||
});
|
||||
if (result.code === 0) {
|
||||
if (io.wantJson) {
|
||||
io.emitJson({ ok: true, cmd, args });
|
||||
} else {
|
||||
io.log(colorize(`Restart issued (${cmd} ${args.join(' ')})`, ANSI.green, io.useColor));
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
if (result.error?.code === 'ENOENT' || result.code === 127) {
|
||||
io.errln(`Error: '${cmd}' not found on this system.`);
|
||||
if (platform === 'darwin') {
|
||||
io.errln(`Hint: no launchd service 'dev.olp.proxy' installed; run 'npm start' manually.`);
|
||||
} else {
|
||||
io.errln(`Hint: no systemd user unit 'olp-proxy' installed; run 'npm start' manually.`);
|
||||
}
|
||||
return 2;
|
||||
}
|
||||
io.errln(`Error: ${cmd} ${args.join(' ')} exited with code ${result.code}`);
|
||||
return 2;
|
||||
}
|
||||
|
||||
// ── Subcommand: doctor ────────────────────────────────────────────────────
|
||||
|
||||
async function cmdDoctor(flags, io) {
|
||||
const olpHome = resolveOlpHome({ olpHome: flags['olp-home'] });
|
||||
const proxyUrl = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
||||
const checkFilter = typeof flags.check === 'string' ? flags.check : undefined;
|
||||
|
||||
// D74 P1-1: pass authHeaders so server.running / server.version checks
|
||||
// succeed under the default production posture (auth.allow_anonymous:
|
||||
// false). resolveBearerToken returns null when no env var is set; the
|
||||
// doctor still runs but distinguishes 401 from "server down" by status
|
||||
// code per the updated check.
|
||||
const result = await runDoctor({
|
||||
olpHome,
|
||||
proxyUrl,
|
||||
checkFilter,
|
||||
authHeaders: authHeaders(),
|
||||
});
|
||||
|
||||
if (io.wantJson) {
|
||||
io.emitJson(result);
|
||||
return result.fail_count === 0 ? 0 : 2;
|
||||
}
|
||||
|
||||
io.log(colorize(`OLP doctor — ${result.summary}`, ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const c of result.checks) {
|
||||
io.log(` ${statusBadge(c.status, io.useColor)} ${c.id.padEnd(36)} ${c.message}`);
|
||||
}
|
||||
io.log('');
|
||||
io.log(` fail=${result.fail_count} warn=${result.warn_count} ok=${result.ok_count} kind=${result.kind}`);
|
||||
if (result.next_action.ai_executable.length > 0) {
|
||||
io.log('');
|
||||
io.log(colorize('Next (AI-executable):', ANSI.cyan, io.useColor));
|
||||
for (const cmd of result.next_action.ai_executable) io.log(` $ ${cmd}`);
|
||||
}
|
||||
if (result.next_action.human_required.length > 0) {
|
||||
io.log('');
|
||||
io.log(colorize('Next (human-required):', ANSI.yellow, io.useColor));
|
||||
for (const step of result.next_action.human_required) io.log(` • ${step}`);
|
||||
}
|
||||
io.log('');
|
||||
io.log(colorize(`verify: ${result.next_action.verify}`, ANSI.dim, io.useColor));
|
||||
return result.fail_count === 0 ? 0 : 2;
|
||||
}
|
||||
|
||||
// ── Subcommand: keys (delegate) ───────────────────────────────────────────
|
||||
|
||||
async function cmdKeys(rest, io) {
|
||||
// Re-use bin/olp-keys.mjs's runCli. Its `out`/`err` writers receive raw strings
|
||||
// (no \n needed since the underlying CLI emits them itself).
|
||||
return await runKeysCli(rest, {
|
||||
out: s => process.stdout.write(s),
|
||||
err: s => process.stderr.write(s),
|
||||
});
|
||||
}
|
||||
|
||||
// ── Usage ──────────────────────────────────────────────────────────────────
|
||||
|
||||
const USAGE = `OLP operator CLI — Phase 4 (ADR 0010)
|
||||
|
||||
Usage:
|
||||
olp <subcommand> [args] [--json] [--proxy-url=<url>] [--olp-home=<path>]
|
||||
|
||||
Subcommands:
|
||||
status GET /v0/management/status (owner-only)
|
||||
health GET /health
|
||||
usage GET /v0/management/dashboard-data (owner-only)
|
||||
models GET /v1/models
|
||||
cache GET /cache/stats (owner-only)
|
||||
providers list providers (registry + config providers.enabled)
|
||||
chain show [<model>] print routing.chains from ~/.olp/config.json
|
||||
logs [N] [--level X] last N audit events from ~/.olp/logs/audit.ndjson
|
||||
restart launchctl (macOS) / systemctl --user (Linux)
|
||||
doctor [--check X] run diagnostic checks (id, category, or prefix filter)
|
||||
keys [args ...] delegate to bin/olp-keys.mjs
|
||||
help print this message
|
||||
|
||||
Global flags:
|
||||
--json emit raw JSON (silences human-readable formatting)
|
||||
--proxy-url=<url> override resolved proxy URL
|
||||
--olp-home=<path> override ~/.olp
|
||||
|
||||
Env:
|
||||
OLP_PROXY_URL full URL (overrides OLP_PORT)
|
||||
OLP_PORT port for default URL (default: 4567)
|
||||
OLP_API_KEY Bearer token for the proxy
|
||||
OLP_OWNER_TOKEN synthetic env-owner token (ADR 0007 § 9.4)
|
||||
OLP_HOME ~/.olp override
|
||||
|
||||
Exit codes:
|
||||
0 success
|
||||
1 bad usage / unknown subcommand
|
||||
2 network or HTTP error (4xx/5xx)
|
||||
3 auth missing / forbidden`;
|
||||
|
||||
// ── runCli (testable entry) ───────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Run the CLI with explicit argv + IO streams. Returns the exit code (no
|
||||
* process.exit). Exported for tests.
|
||||
*
|
||||
* @param {string[]} argv args after the script name
|
||||
* @param {object} [opts]
|
||||
* @param {(s: string) => void} [opts.out]
|
||||
* @param {(s: string) => void} [opts.err]
|
||||
* @param {boolean} [opts.useColor] default true; tests pass false for deterministic strings
|
||||
* @returns {Promise<number>}
|
||||
*/
|
||||
export async function runCli(argv, opts = {}) {
|
||||
if (argv.length === 0 || argv.includes('--help') || argv.includes('-h') || argv[0] === 'help') {
|
||||
const io = makeIO({ ...opts, json: false });
|
||||
io.log(USAGE);
|
||||
return argv.length === 0 ? 1 : 0;
|
||||
}
|
||||
|
||||
const [subcommand, ...rest] = argv;
|
||||
|
||||
// `olp keys ...` passes the remaining argv straight to bin/olp-keys.mjs.
|
||||
if (subcommand === 'keys') {
|
||||
const io = makeIO({ ...opts, json: false });
|
||||
return await cmdKeys(rest, io);
|
||||
}
|
||||
|
||||
const { positional, flags } = parseArgv(rest);
|
||||
const json = flags.json === true;
|
||||
const io = makeIO({ ...opts, json });
|
||||
|
||||
switch (subcommand) {
|
||||
case 'status': return await cmdStatus(flags, io);
|
||||
case 'health': return await cmdHealth(flags, io);
|
||||
case 'usage': return await cmdUsage(flags, io);
|
||||
case 'models': return await cmdModels(flags, io);
|
||||
case 'cache': return await cmdCache(flags, io);
|
||||
case 'providers': return cmdProviders(flags, io);
|
||||
case 'chain': {
|
||||
// `olp chain show [model]`
|
||||
const sub = positional[0];
|
||||
if (sub !== 'show') {
|
||||
io.errln(`Error: unknown 'chain' subcommand "${sub}". Try: olp chain show [model]`);
|
||||
return 1;
|
||||
}
|
||||
return cmdChainShow(positional.slice(1), flags, io);
|
||||
}
|
||||
case 'logs': return await cmdLogs(positional, flags, io);
|
||||
case 'restart': return await cmdRestart(flags, io);
|
||||
case 'doctor': return await cmdDoctor(flags, io);
|
||||
default:
|
||||
io.errln(`Error: unknown subcommand "${subcommand}".`);
|
||||
io.errln(USAGE);
|
||||
return 1;
|
||||
}
|
||||
}
|
||||
|
||||
// ── Main guard ────────────────────────────────────────────────────────────
|
||||
|
||||
function _isMain() {
|
||||
if (!process.argv[1]) return false;
|
||||
try {
|
||||
return realpathSync(fileURLToPath(import.meta.url)) === realpathSync(process.argv[1]);
|
||||
} catch { return false; }
|
||||
}
|
||||
|
||||
if (_isMain()) {
|
||||
runCli(process.argv.slice(2))
|
||||
.then(code => process.exit(code))
|
||||
.catch(e => {
|
||||
process.stderr.write(`Fatal: ${e?.stack ?? e}\n`);
|
||||
process.exit(2);
|
||||
});
|
||||
}
|
||||
+575
-24
@@ -1,16 +1,24 @@
|
||||
<!DOCTYPE html>
|
||||
<!--
|
||||
OLP Dashboard — Phase 3 / D51
|
||||
OLP Dashboard — Phase 5 / D82
|
||||
------------------------------
|
||||
Multi-panel owner-only dashboard per ADR 0008 § 6. Polls
|
||||
/v0/management/dashboard-data every 30 seconds (paused when the
|
||||
page is hidden via document.visibilityState).
|
||||
Multi-panel owner-only dashboard per ADR 0008 § 6.
|
||||
|
||||
Panels (per spec v0.1 § 4.6 + ADR 0008 Lane 5 = B full):
|
||||
1. Per-provider quota / credit pool
|
||||
2. Per-provider 24h request count + cache hit rate + fallback rate
|
||||
3. 30-day spend trend (SVG sparkline; per-provider in tooltip)
|
||||
4. Top 10 fallback chains by trigger count
|
||||
Panels:
|
||||
0. Plan Usage (new D82 — Claude.ai-style per-provider rows; quota_v2; 1-min refresh)
|
||||
1. Per-provider quota / credit pool (legacy; kept for graceful fallback when quota_v2 absent)
|
||||
2. Per-provider 24h request count + cache hit rate + fallback rate (30s refresh)
|
||||
3. 30-day spend trend (SVG sparkline; per-provider in tooltip) (30s refresh)
|
||||
4. Top 10 fallback chains by trigger count (30s refresh)
|
||||
|
||||
Refresh cadence:
|
||||
- Plan Usage panel: 60s (separate timer; visibilityState-guarded per ADR 0012 D82)
|
||||
- Other panels: 30s (original poll cadence; paused when tab hidden)
|
||||
|
||||
Authority:
|
||||
- ADR 0008 § 6 — dashboard layout + owner-only_block
|
||||
- ADR 0012 D82 — quota_v2 Claude.ai-style restructure
|
||||
- v1.x roadmap #8 — closed by this D-day
|
||||
|
||||
No build step, no framework, no external dependencies. Vanilla JS +
|
||||
fetch + DOM render. Owner-only_block: anonymous / guest / no-auth all
|
||||
@@ -43,15 +51,243 @@
|
||||
.chain { font-family: ui-monospace, "SF Mono", Menlo, monospace; font-size: 0.85rem; color: #374151; }
|
||||
.pill { display: inline-block; background: #e5e7eb; color: #374151; padding: 0.05rem 0.4rem; border-radius: 3px; font-size: 0.75rem; }
|
||||
footer { margin-top: 2rem; color: #9ca3af; font-size: 0.75rem; text-align: center; }
|
||||
|
||||
/* ───────────────────────────────────────────
|
||||
Plan Usage panel — D82 Claude.ai-style rows
|
||||
─────────────────────────────────────────── */
|
||||
.plan-usage-header {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
justify-content: space-between;
|
||||
margin-bottom: 1rem;
|
||||
flex-wrap: wrap;
|
||||
gap: 0.5rem;
|
||||
}
|
||||
.plan-usage-header h2 { margin: 0; }
|
||||
.plan-usage-meta {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
gap: 0.75rem;
|
||||
font-size: 0.8rem;
|
||||
color: #6b7280;
|
||||
}
|
||||
.refresh-btn {
|
||||
display: inline-flex;
|
||||
align-items: center;
|
||||
gap: 0.3rem;
|
||||
padding: 0.35rem 0.75rem;
|
||||
background: #fff;
|
||||
border: 1px solid #d1d5db;
|
||||
border-radius: 4px;
|
||||
font-size: 0.8rem;
|
||||
color: #374151;
|
||||
cursor: pointer;
|
||||
transition: background 0.15s, border-color 0.15s;
|
||||
white-space: nowrap;
|
||||
}
|
||||
.refresh-btn:hover:not(:disabled) { background: #f9fafb; border-color: #9ca3af; }
|
||||
.refresh-btn:disabled { opacity: 0.55; cursor: not-allowed; }
|
||||
.refresh-btn .spin { display: inline-block; animation: spin 0.8s linear infinite; }
|
||||
@keyframes spin { to { transform: rotate(360deg); } }
|
||||
|
||||
.provider-row {
|
||||
border: 1px solid #e5e7eb;
|
||||
border-radius: 8px;
|
||||
padding: 1rem 1.25rem;
|
||||
margin-bottom: 0.75rem;
|
||||
background: #fff;
|
||||
}
|
||||
.provider-row:last-child { margin-bottom: 0; }
|
||||
.provider-row.unavailable { background: #f9fafb; }
|
||||
.provider-row.stale { border-color: #fcd34d; }
|
||||
.provider-row.unreachable { border-color: #fca5a5; background: #fff5f5; }
|
||||
|
||||
.provider-row-top {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
gap: 0.6rem;
|
||||
margin-bottom: 0.75rem;
|
||||
flex-wrap: wrap;
|
||||
}
|
||||
.provider-badge {
|
||||
display: inline-block;
|
||||
padding: 0.2rem 0.55rem;
|
||||
border-radius: 4px;
|
||||
font-size: 0.75rem;
|
||||
font-weight: 700;
|
||||
letter-spacing: 0.06em;
|
||||
text-transform: uppercase;
|
||||
color: #fff;
|
||||
}
|
||||
.provider-badge.anthropic { background: #cc4b24; }
|
||||
.provider-badge.codex { background: #10a37f; }
|
||||
.provider-badge.mistral { background: #6d5acd; }
|
||||
.provider-badge.openai { background: #10a37f; }
|
||||
.provider-badge.default { background: #6b7280; }
|
||||
|
||||
.status-dot {
|
||||
display: inline-block;
|
||||
width: 8px;
|
||||
height: 8px;
|
||||
border-radius: 50%;
|
||||
flex-shrink: 0;
|
||||
}
|
||||
.status-dot.live { background: #10b981; }
|
||||
.status-dot.stale { background: #f59e0b; }
|
||||
.status-dot.unavailable { background: #9ca3af; }
|
||||
.status-dot.unreachable { background: #ef4444; }
|
||||
|
||||
.status-chip {
|
||||
display: inline-block;
|
||||
padding: 0.1rem 0.45rem;
|
||||
border-radius: 99px;
|
||||
font-size: 0.7rem;
|
||||
font-weight: 600;
|
||||
letter-spacing: 0.03em;
|
||||
text-transform: uppercase;
|
||||
}
|
||||
.status-chip.live { background: #d1fae5; color: #065f46; }
|
||||
.status-chip.stale { background: #fef3c7; color: #92400e; }
|
||||
.status-chip.unavailable { background: #f3f4f6; color: #6b7280; }
|
||||
.status-chip.unreachable { background: #fee2e2; color: #991b1b; }
|
||||
|
||||
.chip-sm {
|
||||
display: inline-block;
|
||||
padding: 0.1rem 0.45rem;
|
||||
border-radius: 4px;
|
||||
font-size: 0.7rem;
|
||||
color: #374151;
|
||||
background: #f3f4f6;
|
||||
border: 1px solid #e5e7eb;
|
||||
}
|
||||
.schema-tag {
|
||||
margin-left: auto;
|
||||
font-size: 0.7rem;
|
||||
color: #9ca3af;
|
||||
}
|
||||
.unavailable-reason {
|
||||
font-size: 0.875rem;
|
||||
color: #9ca3af;
|
||||
font-style: italic;
|
||||
padding: 0.25rem 0 0;
|
||||
}
|
||||
.unreachable-reason {
|
||||
font-size: 0.875rem;
|
||||
color: #b91c1c;
|
||||
font-style: italic;
|
||||
padding: 0.25rem 0 0;
|
||||
}
|
||||
.last-fresh-tag {
|
||||
font-size: 0.7rem;
|
||||
color: #9ca3af;
|
||||
}
|
||||
|
||||
.utilization-bars { display: flex; flex-direction: column; gap: 0.6rem; }
|
||||
.util-row {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
gap: 0.75rem;
|
||||
flex-wrap: wrap;
|
||||
}
|
||||
.util-label {
|
||||
flex: 0 0 180px;
|
||||
font-size: 0.8rem;
|
||||
color: #6b7280;
|
||||
white-space: nowrap;
|
||||
overflow: hidden;
|
||||
text-overflow: ellipsis;
|
||||
}
|
||||
@media (max-width: 600px) {
|
||||
.util-label { flex: 0 0 100%; }
|
||||
.util-row { flex-direction: column; align-items: flex-start; }
|
||||
.grid { grid-template-columns: 1fr; }
|
||||
}
|
||||
.util-bar-wrap {
|
||||
flex: 1 1 120px;
|
||||
min-width: 80px;
|
||||
height: 8px;
|
||||
background: #e5e7eb;
|
||||
border-radius: 99px;
|
||||
overflow: hidden;
|
||||
}
|
||||
.util-bar-fill {
|
||||
height: 100%;
|
||||
border-radius: 99px;
|
||||
transition: width 0.3s ease;
|
||||
}
|
||||
.util-bar-fill.green { background: linear-gradient(90deg, #34d399, #10b981); }
|
||||
.util-bar-fill.amber { background: linear-gradient(90deg, #fbbf24, #f59e0b); }
|
||||
.util-bar-fill.red { background: linear-gradient(90deg, #f87171, #ef4444); }
|
||||
|
||||
.util-pct {
|
||||
flex: 0 0 40px;
|
||||
font-size: 0.8rem;
|
||||
font-variant-numeric: tabular-nums;
|
||||
font-weight: 600;
|
||||
color: #374151;
|
||||
text-align: right;
|
||||
}
|
||||
.util-reset {
|
||||
flex: 0 0 auto;
|
||||
font-size: 0.75rem;
|
||||
color: #6b7280;
|
||||
white-space: nowrap;
|
||||
}
|
||||
.rep-claim-badge {
|
||||
display: inline-block;
|
||||
padding: 0.1rem 0.45rem;
|
||||
border-radius: 4px;
|
||||
font-size: 0.7rem;
|
||||
background: #ede9fe;
|
||||
color: #5b21b6;
|
||||
border: 1px solid #ddd6fe;
|
||||
font-weight: 600;
|
||||
}
|
||||
.overage-chip {
|
||||
display: inline-block;
|
||||
padding: 0.1rem 0.45rem;
|
||||
border-radius: 4px;
|
||||
font-size: 0.7rem;
|
||||
background: #fef3c7;
|
||||
color: #92400e;
|
||||
border: 1px solid #fcd34d;
|
||||
}
|
||||
.overage-chip.allowed {
|
||||
background: #d1fae5;
|
||||
color: #065f46;
|
||||
border-color: #6ee7b7;
|
||||
}
|
||||
.provider-row-bottom {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
gap: 0.5rem;
|
||||
margin-top: 0.65rem;
|
||||
flex-wrap: wrap;
|
||||
}
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<h1>OLP Dashboard</h1>
|
||||
<div id="meta" class="meta">Loading…</div>
|
||||
<div id="banner-slot"></div>
|
||||
|
||||
<!-- Plan Usage panel (D82 — Claude.ai-style; full width) -->
|
||||
<section class="panel" style="max-width: 1200px; margin-bottom: 1rem;">
|
||||
<div class="plan-usage-header">
|
||||
<h2>Plan Usage</h2>
|
||||
<div class="plan-usage-meta">
|
||||
<span id="quota-last-refresh"></span>
|
||||
<button class="refresh-btn" id="quota-refresh-btn" title="Refresh quota data">
|
||||
<span id="quota-refresh-icon">↻</span> Refresh
|
||||
</button>
|
||||
</div>
|
||||
</div>
|
||||
<div id="panel-plan-usage"><div class="panel-loading">Loading…</div></div>
|
||||
</section>
|
||||
|
||||
<div class="grid">
|
||||
<section class="panel">
|
||||
<h2>Quota (per provider)</h2>
|
||||
<section class="panel" id="legacy-quota-section" style="display:none;">
|
||||
<h2>Quota (per provider) — legacy</h2>
|
||||
<div id="panel-quota"><div class="panel-loading">Loading…</div></div>
|
||||
</section>
|
||||
<section class="panel">
|
||||
@@ -67,15 +303,21 @@
|
||||
<div id="panel-chains"><div class="panel-loading">Loading…</div></div>
|
||||
</section>
|
||||
</div>
|
||||
<footer>OLP Dashboard · poll every 30s · paused when tab hidden · v0.3.0-phase3</footer>
|
||||
<footer>OLP Dashboard · Plan Usage: 60s refresh · other panels: 30s · paused when tab hidden · v0.5.1</footer>
|
||||
<script>
|
||||
(function () {
|
||||
'use strict';
|
||||
const POLL_INTERVAL_MS = 30000;
|
||||
let pollHandle = null;
|
||||
|
||||
/* ─────────────── constants ─────────────── */
|
||||
const POLL_INTERVAL_MS = 30000; // 30s for legacy panels
|
||||
const QUOTA_POLL_INTERVAL_MS = 60000; // 60s for Plan Usage (D82)
|
||||
let pollHandle = null;
|
||||
let quotaRefreshTimer = null;
|
||||
|
||||
/* ─────────────── DOM helpers ─────────────── */
|
||||
function fmtNum(n) { return (n ?? 0).toLocaleString(); }
|
||||
function fmtPct(rate) { return (rate * 100).toFixed(1) + '%'; }
|
||||
|
||||
function el(tag, attrs, ...children) {
|
||||
const node = document.createElement(tag);
|
||||
if (attrs) for (const [k, v] of Object.entries(attrs)) {
|
||||
@@ -97,6 +339,223 @@
|
||||
return node;
|
||||
}
|
||||
|
||||
/* ─────────────── Reset countdown helper (D82 § B) ─────────────── */
|
||||
/**
|
||||
* formatResetCountdown(epochSeconds) → human-readable string
|
||||
*
|
||||
* - past: "Resetting now…"
|
||||
* - < 1 hour: "Resets in 23 min"
|
||||
* - < 24 hours: "Resets in 12hr 30min"
|
||||
* - < 7 days: "Resets Sun 9:00 PM"
|
||||
* - >= 7 days: "Resets May 31 9:00 PM"
|
||||
*/
|
||||
function formatResetCountdown(epochSeconds) {
|
||||
if (epochSeconds == null) return '—';
|
||||
const nowMs = Date.now();
|
||||
const targetMs = epochSeconds * 1000;
|
||||
const diffMs = targetMs - nowMs;
|
||||
|
||||
if (diffMs <= 0) return 'Resetting now…';
|
||||
|
||||
const diffSec = Math.floor(diffMs / 1000);
|
||||
const diffMin = Math.floor(diffSec / 60);
|
||||
const diffHr = Math.floor(diffMin / 60);
|
||||
const diffDay = Math.floor(diffHr / 24);
|
||||
|
||||
if (diffMin < 60) {
|
||||
return 'Resets in ' + diffMin + ' min';
|
||||
}
|
||||
if (diffHr < 24) {
|
||||
const remMin = diffMin - diffHr * 60;
|
||||
if (remMin === 0) return 'Resets in ' + diffHr + 'hr';
|
||||
return 'Resets in ' + diffHr + 'hr ' + remMin + 'min';
|
||||
}
|
||||
// Format as "Resets <day-of-week> <time>" or "Resets <month> <day> <time>"
|
||||
const target = new Date(targetMs);
|
||||
const timeStr = target.toLocaleString('en-US', { hour: 'numeric', minute: '2-digit', hour12: true });
|
||||
if (diffDay < 7) {
|
||||
const dayStr = target.toLocaleString('en-US', { weekday: 'short' });
|
||||
return 'Resets ' + dayStr + ' ' + timeStr;
|
||||
}
|
||||
const dateStr = target.toLocaleString('en-US', { month: 'short', day: 'numeric' });
|
||||
return 'Resets ' + dateStr + ' ' + timeStr;
|
||||
}
|
||||
|
||||
/* ─────────────── "Updated N min ago" helper ─────────────── */
|
||||
function formatAgo(epochMs) {
|
||||
if (epochMs == null) return '';
|
||||
const diffMs = Date.now() - epochMs;
|
||||
if (diffMs < 0) return 'just now';
|
||||
const diffSec = Math.floor(diffMs / 1000);
|
||||
if (diffSec < 60) return 'Updated just now';
|
||||
const diffMin = Math.floor(diffSec / 60);
|
||||
if (diffMin === 1) return 'Updated 1 min ago';
|
||||
if (diffMin < 60) return 'Updated ' + diffMin + ' min ago';
|
||||
const diffHr = Math.floor(diffMin / 60);
|
||||
if (diffHr === 1) return 'Updated ~1hr ago';
|
||||
return 'Updated ~' + diffHr + 'hr ago';
|
||||
}
|
||||
|
||||
/* ─────────────── Utilization bar color ─────────────── */
|
||||
function utilizationColor(fraction) {
|
||||
if (fraction == null) return 'green';
|
||||
if (fraction >= 0.80) return 'red';
|
||||
if (fraction >= 0.50) return 'amber';
|
||||
return 'green';
|
||||
}
|
||||
|
||||
/* ─────────────── Provider badge color class ─────────────── */
|
||||
function providerBadgeClass(name) {
|
||||
const n = (name || '').toLowerCase();
|
||||
if (n === 'anthropic') return 'anthropic';
|
||||
if (n === 'codex') return 'codex';
|
||||
if (n === 'mistral') return 'mistral';
|
||||
if (n === 'openai') return 'openai';
|
||||
return 'default';
|
||||
}
|
||||
|
||||
/* ─────────────── Plan Usage renderer (quota_v2) ─────────────── */
|
||||
function renderPlanUsage(quotaV2) {
|
||||
const target = document.getElementById('panel-plan-usage');
|
||||
target.innerHTML = '';
|
||||
|
||||
if (!Array.isArray(quotaV2) || quotaV2.length === 0) {
|
||||
target.appendChild(el('div', { class: 'panel-loading' }, 'No quota data available.'));
|
||||
return;
|
||||
}
|
||||
|
||||
const frag = document.createDocumentFragment();
|
||||
|
||||
for (const entry of quotaV2) {
|
||||
const status = entry.status || 'unavailable';
|
||||
const rowEl = el('div', { class: 'provider-row ' + status });
|
||||
|
||||
/* ── top bar: badge + status dot + chips + schema tag ── */
|
||||
const topBar = el('div', { class: 'provider-row-top' });
|
||||
|
||||
topBar.appendChild(el('span', { class: 'provider-badge ' + providerBadgeClass(entry.provider) }, (entry.provider || '').toUpperCase()));
|
||||
topBar.appendChild(el('span', { class: 'status-dot ' + status, title: 'Status: ' + status }));
|
||||
topBar.appendChild(el('span', { class: 'status-chip ' + status }, status));
|
||||
|
||||
if (entry.schema_version) {
|
||||
topBar.appendChild(el('span', { class: 'schema-tag' }, 'schema: ' + entry.schema_version));
|
||||
}
|
||||
|
||||
rowEl.appendChild(topBar);
|
||||
|
||||
/* ── unavailable: just show reason, no bars ── */
|
||||
if (status === 'unavailable') {
|
||||
const reason = entry.reason || 'no public quota api or probe disabled';
|
||||
rowEl.appendChild(el('div', { class: 'unavailable-reason' }, reason));
|
||||
frag.appendChild(rowEl);
|
||||
continue;
|
||||
}
|
||||
|
||||
/* ── unreachable (v0.5.1): probe failed + no cache — show failure detail ── */
|
||||
if (status === 'unreachable') {
|
||||
const failure = entry.failure || {};
|
||||
const kind = failure.kind || 'unknown';
|
||||
const msg = failure.message || 'probe failed — no cached data available';
|
||||
const shortText = `${kind}: ${msg}`;
|
||||
rowEl.appendChild(el('div', { class: 'unreachable-reason' }, shortText));
|
||||
if (failure.backoff_until) {
|
||||
const backoffMs = Math.max(0, failure.backoff_until - Date.now());
|
||||
const backoffSec = Math.round(backoffMs / 1000);
|
||||
if (backoffSec > 0) {
|
||||
rowEl.appendChild(el('div', { class: 'unavailable-reason' }, `backoff active: ${backoffSec}s remaining`));
|
||||
}
|
||||
}
|
||||
frag.appendChild(rowEl);
|
||||
continue;
|
||||
}
|
||||
|
||||
/* ── utilization bars (5h + 7d) ── */
|
||||
const util = entry.utilization || {};
|
||||
const reset = entry.reset || {};
|
||||
const barsWrap = el('div', { class: 'utilization-bars' });
|
||||
|
||||
const windows = [
|
||||
{ key: '5h', label: 'Current 5-hour session' },
|
||||
{ key: '7d', label: 'Weekly all-models' },
|
||||
];
|
||||
|
||||
for (const w of windows) {
|
||||
const frac = util[w.key];
|
||||
const resetEpoch = reset[w.key];
|
||||
const color = utilizationColor(frac);
|
||||
const pctStr = frac != null ? Math.round(frac * 100) + '%' : '—';
|
||||
const fillPct = frac != null ? Math.min(100, Math.round(frac * 100)) : 0;
|
||||
const resetStr = formatResetCountdown(resetEpoch);
|
||||
|
||||
const utilRow = el('div', { class: 'util-row' });
|
||||
|
||||
utilRow.appendChild(el('span', { class: 'util-label', title: w.label },
|
||||
w.label + (frac != null ? ': ' + pctStr : '')
|
||||
));
|
||||
|
||||
const barWrap = el('div', { class: 'util-bar-wrap' });
|
||||
barWrap.appendChild(el('div', {
|
||||
class: 'util-bar-fill ' + color,
|
||||
style: 'width: ' + fillPct + '%',
|
||||
'aria-valuenow': fillPct,
|
||||
'aria-valuemin': '0',
|
||||
'aria-valuemax': '100',
|
||||
role: 'progressbar',
|
||||
}));
|
||||
utilRow.appendChild(barWrap);
|
||||
|
||||
utilRow.appendChild(el('span', { class: 'util-pct' }, pctStr));
|
||||
utilRow.appendChild(el('span', { class: 'util-reset' }, resetStr));
|
||||
|
||||
barsWrap.appendChild(utilRow);
|
||||
}
|
||||
|
||||
rowEl.appendChild(barsWrap);
|
||||
|
||||
/* ── bottom chips: representative-claim, overage, last-fresh ── */
|
||||
const bottomBar = el('div', { class: 'provider-row-bottom' });
|
||||
|
||||
if (entry.representative_claim) {
|
||||
const claimLabel = entry.representative_claim === 'five_hour' ? '5-hour claim'
|
||||
: entry.representative_claim === 'seven_day' ? '7-day claim'
|
||||
: entry.representative_claim;
|
||||
bottomBar.appendChild(el('span', { class: 'rep-claim-badge', title: 'Binding window: ' + entry.representative_claim }, claimLabel));
|
||||
}
|
||||
|
||||
if (entry.overage && entry.overage.status) {
|
||||
const ov = entry.overage;
|
||||
const ovStatus = (ov.status || 'unknown').toLowerCase();
|
||||
const isAllowed = ovStatus === 'allowed' || ovStatus === 'active';
|
||||
const chipClass = isAllowed ? 'overage-chip allowed' : 'overage-chip';
|
||||
const label = 'Overage: ' + (ov.status || '—')
|
||||
+ (ov.disabled_reason ? ' (' + ov.disabled_reason + ')' : '');
|
||||
bottomBar.appendChild(el('span', { class: chipClass, title: label }, label));
|
||||
}
|
||||
|
||||
if (entry.fallback_percentage != null) {
|
||||
const fpPct = Math.round(entry.fallback_percentage * 100) + '%';
|
||||
bottomBar.appendChild(el('span', { class: 'chip-sm', title: 'Fallback rate (last window)' }, 'Fallback ' + fpPct));
|
||||
}
|
||||
|
||||
if (entry.last_fresh_at) {
|
||||
bottomBar.appendChild(el('span', { class: 'last-fresh-tag' }, formatAgo(entry.last_fresh_at)));
|
||||
}
|
||||
|
||||
if (status === 'stale') {
|
||||
const staleTitle = entry.last_fresh_at
|
||||
? 'Last successful probe was ' + formatAgo(entry.last_fresh_at) + '; backoff active'
|
||||
: 'Probe data is stale; backoff active';
|
||||
bottomBar.appendChild(el('span', { class: 'chip-sm', style: 'color: #92400e; background: #fef3c7; border-color: #fcd34d;', title: staleTitle }, '⚠ stale data'));
|
||||
}
|
||||
|
||||
rowEl.appendChild(bottomBar);
|
||||
frag.appendChild(rowEl);
|
||||
}
|
||||
|
||||
target.appendChild(frag);
|
||||
}
|
||||
|
||||
/* ─────────────── Legacy quota renderer (graceful fallback) ─────────────── */
|
||||
function renderQuota(data) {
|
||||
const target = document.getElementById('panel-quota');
|
||||
target.innerHTML = '';
|
||||
@@ -129,6 +588,28 @@
|
||||
target.appendChild(table);
|
||||
}
|
||||
|
||||
/* ─────────────── Plan Usage top-level render + quota routing ─────────────── */
|
||||
function renderQuotaSection(data) {
|
||||
const hasV2 = Array.isArray(data.quota_v2) && data.quota_v2.length > 0;
|
||||
const legacySection = document.getElementById('legacy-quota-section');
|
||||
|
||||
if (hasV2) {
|
||||
// D82: use enriched quota_v2 rows; hide legacy panel
|
||||
legacySection.style.display = 'none';
|
||||
renderPlanUsage(data.quota_v2);
|
||||
} else {
|
||||
// Graceful fallback: show legacy quota panel (older server build without D81)
|
||||
legacySection.style.display = '';
|
||||
// Also show legacy data in Plan Usage panel with a note
|
||||
const target = document.getElementById('panel-plan-usage');
|
||||
target.innerHTML = '';
|
||||
target.appendChild(el('div', { class: 'panel-loading', style: 'color:#6b7280;' },
|
||||
'quota_v2 not available (server may not have D81 yet). See legacy Quota panel below.'));
|
||||
renderQuota(data.quota);
|
||||
}
|
||||
}
|
||||
|
||||
/* ─────────────── Other panel renderers (unchanged from D51) ─────────────── */
|
||||
function render24h(window24h, cacheHit24h) {
|
||||
const target = document.getElementById('panel-24h');
|
||||
target.innerHTML = '';
|
||||
@@ -196,7 +677,7 @@
|
||||
const minLabel = svgEl('text', { x: 4, y: height - 4, 'font-size': 10, fill: '#6b7280' });
|
||||
minLabel.textContent = '0';
|
||||
svg.appendChild(minLabel);
|
||||
// Date labels (first + last only at v0.3.0; mid labels deferred — added if needed by Phase 4 UX feedback)
|
||||
// Date labels (first + last only)
|
||||
if (spendTrend30d.length > 0) {
|
||||
const firstDate = svgEl('text', { x: padding.left, y: height - 4, 'font-size': 10, fill: '#6b7280' });
|
||||
firstDate.textContent = spendTrend30d[0].date.slice(5);
|
||||
@@ -240,6 +721,7 @@
|
||||
target.appendChild(table);
|
||||
}
|
||||
|
||||
/* ─────────────── Error / clear banner ─────────────── */
|
||||
function showError(message) {
|
||||
const slot = document.getElementById('banner-slot');
|
||||
slot.innerHTML = '';
|
||||
@@ -250,6 +732,7 @@
|
||||
document.getElementById('banner-slot').innerHTML = '';
|
||||
}
|
||||
|
||||
/* ─────────────── Fetch ─────────────── */
|
||||
async function fetchDashboardData() {
|
||||
const res = await fetch('/v0/management/dashboard-data', {
|
||||
headers: { 'Accept': 'application/json' },
|
||||
@@ -266,24 +749,82 @@
|
||||
return await res.json();
|
||||
}
|
||||
|
||||
/* ─────────────── Quota-only refresh (D82 § C — 60s timer) ─────────────── */
|
||||
let _lastQuotaFetchedAt = null;
|
||||
|
||||
async function refreshQuotaV2() {
|
||||
try {
|
||||
const data = await fetchDashboardData();
|
||||
clearError();
|
||||
_lastQuotaFetchedAt = Date.now();
|
||||
renderQuotaSection(data);
|
||||
updateQuotaLastRefreshLabel();
|
||||
} catch (err) {
|
||||
console.warn('OLP quota refresh failed:', err.message);
|
||||
}
|
||||
}
|
||||
|
||||
function updateQuotaLastRefreshLabel() {
|
||||
const span = document.getElementById('quota-last-refresh');
|
||||
if (!span) return;
|
||||
if (_lastQuotaFetchedAt) {
|
||||
span.textContent = 'Updated ' + new Date(_lastQuotaFetchedAt).toLocaleTimeString();
|
||||
}
|
||||
}
|
||||
|
||||
/* ─────────────── 60s quota timer with visibilityState guard ─────────────── */
|
||||
function startQuotaRefresh() {
|
||||
if (quotaRefreshTimer !== null) return;
|
||||
quotaRefreshTimer = setInterval(refreshQuotaV2, QUOTA_POLL_INTERVAL_MS);
|
||||
}
|
||||
function stopQuotaRefresh() {
|
||||
if (quotaRefreshTimer === null) return;
|
||||
clearInterval(quotaRefreshTimer);
|
||||
quotaRefreshTimer = null;
|
||||
}
|
||||
|
||||
/* ─────────────── Manual refresh button (D82 § D) ─────────────── */
|
||||
(function wireRefreshButton() {
|
||||
const btn = document.getElementById('quota-refresh-btn');
|
||||
const icon = document.getElementById('quota-refresh-icon');
|
||||
if (!btn) return;
|
||||
btn.addEventListener('click', async () => {
|
||||
if (btn.disabled) return;
|
||||
btn.disabled = true;
|
||||
icon.textContent = '⟳';
|
||||
icon.classList.add('spin');
|
||||
try {
|
||||
await refreshQuotaV2();
|
||||
} finally {
|
||||
icon.classList.remove('spin');
|
||||
icon.textContent = '↻';
|
||||
// Re-enable after 2s spam guard
|
||||
setTimeout(() => { btn.disabled = false; }, 2000);
|
||||
}
|
||||
});
|
||||
})();
|
||||
|
||||
/* ─────────────── Full 30s refresh (legacy panels + meta) ─────────────── */
|
||||
async function refresh() {
|
||||
try {
|
||||
const data = await fetchDashboardData();
|
||||
clearError();
|
||||
const generated = data.generated_at ? new Date(data.generated_at) : new Date();
|
||||
document.getElementById('meta').textContent =
|
||||
'Last refresh: ' + generated.toLocaleString() + ' · next in ~30s';
|
||||
renderQuota(data.quota);
|
||||
'Last refresh: ' + generated.toLocaleString() + ' · quota every 60s · other panels every 30s';
|
||||
// Quota section: also render on each full refresh to keep in sync
|
||||
_lastQuotaFetchedAt = Date.now();
|
||||
renderQuotaSection(data);
|
||||
updateQuotaLastRefreshLabel();
|
||||
render24h(data.window_24h, data.cache_hit_24h);
|
||||
renderTrend(data.spend_trend_30d);
|
||||
renderChains(data.top_fallback_chains_24h);
|
||||
} catch (err) {
|
||||
// Error banner already shown by fetchDashboardData; keep panels in
|
||||
// their last-good state. Console for operator debugging.
|
||||
console.warn('OLP dashboard refresh failed:', err.message);
|
||||
}
|
||||
}
|
||||
|
||||
/* ─────────────── 30s poll (legacy panels) ─────────────── */
|
||||
function startPolling() {
|
||||
if (pollHandle !== null) return;
|
||||
pollHandle = setInterval(refresh, POLL_INTERVAL_MS);
|
||||
@@ -294,14 +835,24 @@
|
||||
pollHandle = null;
|
||||
}
|
||||
|
||||
// Pause when tab hidden, resume on visible (ADR 0008 § 6.5).
|
||||
/* ─────────────── visibilitychange (both timers) ─────────────── */
|
||||
document.addEventListener('visibilitychange', () => {
|
||||
if (document.visibilityState === 'hidden') stopPolling();
|
||||
else { refresh(); startPolling(); }
|
||||
if (document.visibilityState === 'hidden') {
|
||||
stopPolling();
|
||||
stopQuotaRefresh();
|
||||
} else {
|
||||
refresh();
|
||||
startPolling();
|
||||
refreshQuotaV2();
|
||||
startQuotaRefresh();
|
||||
}
|
||||
});
|
||||
|
||||
// Initial fetch + start poll.
|
||||
refresh().finally(startPolling);
|
||||
/* ─────────────── Boot ─────────────── */
|
||||
refresh().finally(() => {
|
||||
startPolling();
|
||||
if (document.visibilityState === 'visible') startQuotaRefresh();
|
||||
});
|
||||
})();
|
||||
</script>
|
||||
</body>
|
||||
|
||||
@@ -40,7 +40,7 @@ What OLP **does NOT inherit** from ADR 0005 (and where ADR 0005's reasoning does
|
||||
- ADR 0005's separate-project recommendation came with two qualifiers that OLP rejects: "BYOK from day one" and "no `cli.js` spawn." Both qualifiers were appropriate for the *commercial* path ADR 0005 was contemplating. OLP is not commercial — it is personal- and family-scale, shares the maintainer's own subscription quota across family clients, and explicitly spawns provider CLIs (it is precisely the spawn-binary architecture that delivers the "subscription quota maximization" value proposition spec §1 names).
|
||||
- OLP is therefore not the commercial pivot ADR 0005 endorsed. It is a personal-use re-architecture of the proxy-CLI pattern, which ADR 0005 did not contemplate. The supersession is honest about this gap.
|
||||
|
||||
OCP itself is not deleted. Per spec §7, OCP enters maintenance mode when OLP v0.1 ships. The two projects do not parallel-run in production (port-3456 conflict, single launchd service slot, one set of credentials per machine).
|
||||
OCP itself is not deleted. Per spec §7, OCP enters maintenance mode when OLP v0.1 ships. ~~The two projects do not parallel-run in production (port-3456 conflict, single launchd service slot, one set of credentials per machine).~~ **Amended at D60 (2026-05-26, ADR 0010 Phase 4 charter):** the port-conflict assumption is lifted. OLP's default port moved `3456 → 4567` at v0.4.0 so OCP (which stays on 3456) and OLP can co-host on the same machine during a transition window. Launchd label collision **will be** avoided via `dev.olp.proxy` (OLP plist generation lands at Phase 4 close per ADR 0010 D64–D70; not on disk at D60) vs `dev.ocp.proxy` (OCP, already shipped). Credentials remain per-project (`~/.ocp/` vs `~/.olp/`). Co-host is explicit-opt-in, not the recommended steady state.
|
||||
|
||||
OCP ADR 0005 receives a header amendment on merge of this ADR: *"Superseded in part by OLP — see https://github.com/dtzp555-max/olp ADR 0001 for the narrow scope of the supersession (single-provider-sufficiency premise only; ADR 0005's commercial / BYOK / no-spawn recommendations are not adopted)."* The body of ADR 0005 is otherwise untouched. Future readers should see the original reasoning intact and the supersession marker scoped explicitly.
|
||||
|
||||
|
||||
@@ -7,7 +7,52 @@
|
||||
|
||||
## Amendments
|
||||
|
||||
> **Note on numbering.** Sequence is 1, 3, 4, 5, 6 — Amendment 2 was never written. The reserved slot was originally planned for a separate `maxConcurrent` ratification, but that content was folded into Amendment 1 (the retroactive contract-sync amendment) at filing time and the gap was not backfilled. The gap is intentional and load-bearing — no missing content; do not renumber Amendments 3+ to close it (cross-references to Amendment N from other docs would silently break).
|
||||
> **Note on numbering.** Sequence is 1, 3, 4, 5, 6, 7 — Amendment 2 was never written. The reserved slot was originally planned for a separate `maxConcurrent` ratification, but that content was folded into Amendment 1 (the retroactive contract-sync amendment) at filing time and the gap was not backfilled. The gap is intentional and load-bearing — no missing content; do not renumber Amendments 3+ to close it (cross-references to Amendment N from other docs would silently break).
|
||||
|
||||
> **Forward-pointer:** Amendment 9 (2026-05-29) — Provider `ISOLATION` Contract for Multi-Tenant Spawn Isolation — is located at the **end of this file** (after § Sources), not in this Amendments block. The placement is documented in Amendment 9's editorial note; the substance is the addition of an OPTIONAL `ISOLATION` named export to provider plugin modules, consumed by `lib/sandbox/manager.mjs` (per ADR 0014 Amendment 1) to compose per-spawn ephemeral-home + per-provider isolation primitives. Co-merge with ADR 0014 Amendment 1.
|
||||
|
||||
### Amendment 8 — 2026-05-26: Permit `quotaStatus()` direct-API access (READ-ONLY exemption) for plan-usage probes (D79–D80 — Phase 5)
|
||||
|
||||
- **Context:** ADR 0012 (Phase 5 charter) opens 2026-05-26 to port OCP's plan-usage probe (`ocp/server.mjs:842-1109`) into `lib/providers/anthropic.mjs:quotaStatus()`. The probe calls `POST https://api.anthropic.com/v1/messages` directly with an OAuth bearer and parses `anthropic-ratelimit-unified-*` response headers. This violates the plugin contract's implicit assumption that ALL provider interaction goes through `spawn` (the binary CLI). `ALIGNMENT.md` Rule 2 (provider-CLI-as-authority) further constrains plugins to operations the provider CLI itself performs. The OCP-derived plan-usage probe satisfies neither of these — it bypasses `claude -p` and hits the public API directly. **Without an explicit exemption Amendment, D80 is unalignable.**
|
||||
- **Why the exemption is sound:** The probe is strictly **READ-ONLY** (one `POST /v1/messages` with `max_tokens: 1`; the response body is discarded; only response headers are parsed) AND **subscription-scope** (the OAuth bearer is the same one Claude Code uses for `claude -p`; no extra grant is requested) AND **idempotent** (probe failure returns `null`, never throws to a caller). The "what authority backs this?" answer is: Anthropic's CLI internally makes the same `/v1/messages` call (verified 2026-05-26 by `strings` on the compiled binary — see `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`); the probe is mirroring an established CLI behaviour rather than introducing a new wire format. Under `ALIGNMENT.md` Rule 2, mirroring observed CLI behaviour is permitted; the Rule's intent is "don't invent wire formats Anthropic's CLI does not perform", which the probe respects.
|
||||
- **Change — extend the Provider contract description:**
|
||||
- `quotaStatus(authContext): { quotaInfo }` is now permitted to call provider HTTP APIs directly, subject to **all three** constraints:
|
||||
1. **READ-ONLY** — the API call must not mutate provider-side state. POST is acceptable when the response is what's needed (Anthropic returns ratelimit headers on `POST /v1/messages`); the request body MUST minimise side-effects (`max_tokens: 1`, dummy `messages`).
|
||||
2. **Subscription-scope reuse** — the credentials used MUST be the same auth artifact the spawn path already reads via `readAuthArtifact()`. No new OAuth grant, no new API-key registration, no separate scopes.
|
||||
3. **Idempotent failure** — if the probe fails for any reason (network error, 401, 429, schema parse failure), the function returns a structured shape (`{ probe_status: 'unreachable', failure: { kind, message, backoff_until? } }` since v0.5.1; see ADR 0013 Rule 6 + ADR 0008 Amendment 2) rather than throwing. The caller (server.mjs / dashboard / `olp usage` CLI) gracefully degrades. At v0.5.0 the failure shape was the literal value `null`; v0.5.1 refined this to a structured shape so operators can distinguish auth failures from rate-limit failures from network failures from in-backoff stale-cache. The substantive idempotent-failure constraint (no throw to caller) is unchanged.
|
||||
- `healthCheck()` and other contract methods are NOT extended by this Amendment. Only `quotaStatus()` may make direct API calls. A plugin that wants live data for any other contract method must continue to use `spawn` or `readAuthArtifact`.
|
||||
- The probe MUST cache its result. Recommended TTL: 5 minutes (mirrors OCP `USAGE_CACHE_TTL`). Tighter TTLs (e.g. dashboard's 1-minute refresh) are served from the cached value if fresh; cache miss triggers a real probe.
|
||||
- The probe MUST implement exponential backoff on refresh failures: minimum 60s, maximum 3600s (mirrors OCP `OAUTH_REFRESH_MIN_BACKOFF` / `OAUTH_REFRESH_MAX_BACKOFF`). Tight loop on failure has historically burned through Anthropic's rate limit in seconds (OCP institutional lesson 2026-04).
|
||||
- The probe MUST be opt-in via `~/.olp/config.json` (`providers.<name>.quota_probe_enabled: true`; default `false`). Reasoning: a fresh OLP install on a machine without OAuth credentials should not bombard `api.anthropic.com` with 401-bound probes; the operator opts in once the credentials are configured.
|
||||
- **What this Amendment does NOT permit:**
|
||||
- Mutating API calls (e.g. POST/PATCH/DELETE that change provider-side state). Still forbidden.
|
||||
- API calls for any contract method other than `quotaStatus()`. `spawn` / `healthCheck` / `doctorChecks` / `estimateCost` / `models` / `hints` / `name` / `displayName` / `auth` remain spawn-and-filesystem-only.
|
||||
- Per-provider new auth grants. The probe uses the spawn path's existing credentials.
|
||||
- Bypassing the alignment.yml blacklist. The hallucinated `/api/oauth/usage` token stays blacklisted; the probe uses `/v1/messages` (real endpoint).
|
||||
- **API calls to endpoints not explicitly enumerated by the companion ADR 0013 § Rule 2.** Amendment 8 permits the *kind* of call (READ-ONLY direct API for quota probing); ADR 0013 Rule 2 enumerates *which specific endpoint* is permitted. A future reader of Amendment 8 alone should NOT infer that any READ-ONLY/idempotent endpoint is fair game — the per-endpoint containment is locked to ADR 0013. Re-opening per-endpoint scope requires an ADR 0013 amendment, not a new plugin-level interpretation of Amendment 8.
|
||||
- **Backwards compatibility:** Plugins whose `quotaStatus()` still returns `null` (mistral at v0.5.0 pending D84 audit, codex permanently per Phase 5 charter) are NOT affected. No existing behaviour changes for them.
|
||||
- **Authority cited at the implementation:** D80 commit cites this Amendment + `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md` + `Claude Code v2.1.x § OAuth bearer + ratelimit-unified headers` + live-probe transcript from 2026-05-26 in the commit body. ALIGNMENT.md Rule 1 + Rule 5 (CI) both satisfied.
|
||||
- **Tests:** Suite 38 (Phase 5 D83) covers the probe: mock HTTP server returning all 13 `anthropic-ratelimit-unified-*` headers; assert parse correctness for each; assert 5min cache; assert 60s-3600s exponential backoff on simulated 429; assert stale-cache-on-failure. At v0.5.0 stale-failure returned `{ stale: true, ... }` with `null` reserved for no-cache failures; v0.5.1 refined the return contract — `null` is now reserved STRICTLY for opt-in-off, and all failure modes (auth / rate-limit / schema-drift / network / no-creds) return `{ probe_status: 'unreachable' | 'stale', failure: {...} }`. See `test-features.mjs` Suite 38 (38u/38v/38w added for the v0.5.1 hotfix regression coverage of F1 / F2 / F3 per codex review).
|
||||
- **Procedural mechanism:** Iron Rule 11 (IDR) — this Amendment, ADR 0012 (Phase 5 charter), and ADR 0013 (OAuth READ-ONLY consumption rules) land together at D79 as a single coupled commit. Reviewing them separately cannot verify consumer-producer alignment. Iron Rule 10 fresh-context reviewer per CLAUDE.md hard requirement #3.
|
||||
|
||||
### Amendment 7 — 2026-05-26: Add OPTIONAL `doctorChecks()` to the Provider contract (D67 — Phase 4 operator UX)
|
||||
|
||||
- **Context:** ADR 0010 § Phase 4 D64-D67 ships `bin/olp.mjs` operator CLI + `olp doctor` framework. `olp doctor` runs a set of `Check` objects (id / category / async `run()` returning `{ status, message, evidence? }`) and discriminates the next remediation step via a `kind` field (`noop` / `fix_server` / `fix_oauth` / `fix_provider` / `fresh_install`). The framework needs per-provider checks so a user with a broken `claude` install gets a different fix recipe than a user with a broken `vibe` install. Hardcoding the recipes in `bin/olp.mjs` would re-introduce the kind of per-provider knowledge drift that ADR 0002 § Decision exists to prevent — when a new provider plugin lands, the operator CLI would have to be edited too.
|
||||
- **Change — add to Provider contract:**
|
||||
- Introduce **OPTIONAL** `doctorChecks()` returning `DoctorCheck[]` where each `DoctorCheck` has the shape:
|
||||
- `id: string` — unique per check, conventionally `<provider>.<probe-name>` (e.g. `anthropic.cli_available`, `anthropic.oauth_token_present`).
|
||||
- `category: 'provider'` — fixed for plugin-contributed checks. The framework reserves `'server'`, `'auth'`, `'config'`, `'system'` for built-in checks.
|
||||
- `async run(): { status: 'ok' | 'fail' | 'warn', message: string, evidence?: { fix_commands?: string[], human_steps?: string[], reference?: string } }` — runs the probe. `status: 'fail'` makes `olp doctor` exit non-zero and contributes to the `kind: fix_provider` discriminator; `evidence.fix_commands[]` is concatenated into `next_action.ai_executable[]` and `evidence.human_steps[]` into `next_action.human_required[]`.
|
||||
- **Backwards compatibility:** Plugins that omit `doctorChecks()` contribute zero provider checks. Their healthCheck() return value continues to flow through `/health.providers.status.<name>` exactly as today. No existing plugin behaviour changes; no existing test breaks. `validateProvider` in `lib/providers/base.mjs` is updated to type-check `doctorChecks` only when present (must be a function); absence is allowed.
|
||||
- **What `doctorChecks()` is for vs. what `healthCheck()` is for:**
|
||||
- `healthCheck()` answers "is this provider currently usable?" — checked at the request-execution layer; output feeds `/health` and per-request retry decisions.
|
||||
- `doctorChecks()` answers "if this provider is broken, what specific actionable steps fix it?" — checked at the operator layer; output feeds `olp doctor` + the `next_action.ai_executable[]` repair templates that a downstream AI agent can paste-and-run.
|
||||
- **Suggested probe set (per plugin):**
|
||||
- `<provider>.cli_available` — spawn `<bin> --version` with short timeout (≤3s); fail → fix_commands include install instruction.
|
||||
- `<provider>.<auth-artifact>_present` — check whether the auth file / env var the plugin's `readAuthArtifact()` reads is populated; fail → human_steps include the login command (which usually requires browser interaction and so cannot be in `ai_executable[]`).
|
||||
- **Authority:** ADR 0010 § Phase 4 D64-D67 (this is the addition called out by that charter). No provider CLI doc citation needed — `doctorChecks()` is an internal contract field. Implementation lands in D67 (this PR): `lib/providers/anthropic.mjs`, `lib/providers/codex.mjs`, `lib/providers/mistral.mjs` each gain a `doctorChecks()` method covering `cli_available` + `<auth-artifact>_present`.
|
||||
- **Tests:** Suite 32 (`bin/olp.mjs` CLI smoke) and Suite 33 (`olp doctor` framework) in `test-features.mjs` cover the contract amendment. Suite 33 specifically asserts: (a) a plugin without `doctorChecks()` contributes no provider checks (default behaviour), (b) a plugin with a failing `doctorChecks()` probe triggers `kind: fix_provider` and propagates its `evidence.fix_commands[]` into `next_action.ai_executable[]`, (c) all-passing checks yield `kind: noop`.
|
||||
- **Procedural mechanism:** CC 开发铁律 v1.6 § 11 (IDR) — the contract amendment, the plugin implementations, the doctor framework, and the CLI scaffold are tightly coupled. They land as a single PR (D64-D67 bundle) because reviewing them separately cannot verify that consumer + producer line up. Iron Rule 10 fresh-context reviewer per CLAUDE.md hard requirement #3.
|
||||
|
||||
### Amendment 6 — 2026-05-24: `maxConcurrent` runtime enforcement landed (D38, issue #1)
|
||||
|
||||
@@ -167,3 +212,469 @@ Every provider plugin exports an object conforming to:
|
||||
- OLP v0.1 spec §4.2 (Plugin-based provider system, including the v1.0 Provider contract definition)
|
||||
- OCP ADR 0003 (`models.json` as SPOT) — informs the "static enumeration, not filesystem scan" loading model
|
||||
- OCP ADR 0005 — the context paragraph references OCP's `server.mjs` reaching 1667 lines at one provider; the plugin architecture is the structural response to that complexity scaling N×
|
||||
|
||||
---
|
||||
|
||||
### Amendment 9 — 2026-05-29: Provider `ISOLATION` Contract for Multi-Tenant Spawn Isolation (Phase 7, ADR 0014 Amendment 1 co-merge)
|
||||
|
||||
> **Editorial note.** Per the existing "amendments most-recent-first" convention near the top of this file, Amendment 9 logically slots between Amendment 8 and the original body. It is physically located at the file's tail (after § Sources) to honor the constitution's "append, do not rewrite" discipline for this addition — the rationale is that the contract surface added here is large enough (a structured per-provider sub-export, not just a hint-bag field) that an in-line edit of the § Decision body would constitute a rewrite of the v1.0 contract listing rather than an amendment over it. Future readers consulting the amendment-history block at the top of the file will find a stub forward-pointer to this section.
|
||||
>
|
||||
> The amendment is otherwise a peer of Amendments 1–8 (same `###` heading depth, same shape).
|
||||
|
||||
#### Context
|
||||
|
||||
The OLP spawn pipeline currently treats every provider as a plain `child_process.spawn` of the provider's CLI binary with a homogeneous env block and the server process's working directory. This works on a single-tenant developer laptop. It does **not** work on the family-LAN PI231 deployment (multi-key, multi-caller, single OS user) and is a hard blocker for the cloud rollout described in `docs/plans/cloud-deployment-family.md` § 5 — both for the reasons captured in the 2026-05-27 incident memory at `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` (OAuth-token exfiltration, codex `shell` tool real execution, cross-tenant filesystem read leakage).
|
||||
|
||||
The parallel ADR 0014 Amendment 1 retires the **outer-bwrap PR-B approach** — which initialized `@anthropic-ai/sandbox-runtime` `SandboxManager` once at server startup and wrapped every provider spawn through a global namespace — and replaces it with a **per-spawn ephemeral-home + per-provider isolation primitives** architecture. The new shape of `lib/sandbox/manager.mjs` is no longer a thin wrapper around `wrapSpawn()`; it is an orchestrator that, on each spawn, asks the provider plugin *what isolation primitives this provider needs*, composes them, and hands the spawn a ready-to-execute environment.
|
||||
|
||||
The thing the orchestrator asks for is the subject of this amendment: the **Provider `ISOLATION` contract**.
|
||||
|
||||
#### The interaction surface this amendment governs
|
||||
|
||||
```text
|
||||
┌──────────────────────────────────────┐
|
||||
│ server.mjs handleChatCompletions │
|
||||
│ → executeHopFn │
|
||||
│ → provider.spawn(irRequest, ...) │
|
||||
└────────────────────┬─────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────────────┐
|
||||
│ lib/sandbox/manager.mjs │
|
||||
│ prepareIsolatedEnvironment( │
|
||||
│ provider, │
|
||||
│ { keyId, reqId, ... } │
|
||||
│ ) │
|
||||
│ ↓ reads provider.ISOLATION │
|
||||
│ ↓ mkdtemp ephemeralRoot │
|
||||
│ ↓ mkdir requiredHomePaths │
|
||||
│ ↓ symlink/copy credentialMounts │
|
||||
│ ↓ compose ephemeralEnvOverrides │
|
||||
│ ↓ wrap args via toolHardening │
|
||||
└────────────────────┬─────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────────────┐
|
||||
│ child_process.spawn(bin, args, { │
|
||||
│ env: composedEnv, cwd: epRoot, ... │
|
||||
│ }) │
|
||||
└──────────────────────────────────────┘
|
||||
```
|
||||
|
||||
The provider plugin is the **authority** for what isolation primitives are needed. The provider knows what env var its CLI honors for credential lookup (`HOME`, `CODEX_HOME`, `VIBE_HOME`, …). The provider knows whether the CLI has an inner sandbox that must be permitted to clone user namespaces. The provider knows the cross-tenant read protection regime it ships under. The orchestrator's job is purely composition; it must not know that "for codex, use `CODEX_HOME`" — that knowledge belongs in `lib/providers/codex.mjs`.
|
||||
|
||||
This is the same separation-of-concerns principle that has governed every prior amendment to this ADR: provider-specific knowledge lives in the provider file; the orchestrator stays generic. Amendment 7's `doctorChecks()` followed it (per-provider repair recipes); Amendment 8's `quotaStatus()` followed it (per-provider probe authorities); this amendment follows it for isolation primitives.
|
||||
|
||||
#### Decision — add OPTIONAL `ISOLATION` named export to the Provider plugin module
|
||||
|
||||
Each provider plugin module (`lib/providers/<name>.mjs`) MAY export, in addition to the default-exported provider object, a named const `ISOLATION` describing the isolation primitives the orchestrator should compose for spawns of this provider. The shape is:
|
||||
|
||||
```javascript
|
||||
export const ISOLATION = {
|
||||
ephemeralEnvOverrides: ({ ephemeralRoot, keyId, reqId }) => ({ /* env var map */ }),
|
||||
credentialMounts: [ [srcAbsPath, dstRelativeToEphemeralRoot], ... ],
|
||||
requiredHomePaths: [ /* dirs to mkdir empty under ephemeralRoot */ ],
|
||||
hasInnerSandbox: boolean,
|
||||
crossTenantReadProtection: 'tool-suppression' | 'inner-sandbox' | 'none',
|
||||
recommendedDeploymentTier: 'shared-os-user' | 'per-os-user' | 'separate-vm',
|
||||
toolHardeningArgs: (existingArgs) => modifiedArgs, // optional
|
||||
}
|
||||
```
|
||||
|
||||
The export is **optional**. A plugin that omits `ISOLATION` continues to spawn under the legacy unsandboxed shape exactly as it does today — see § Backward compatibility below. The opt-in surface is consistent with Amendment 7's `doctorChecks()` treatment (additive, no breakage for plugins that haven't been touched).
|
||||
|
||||
The remainder of this amendment specifies each field's semantics, default-when-absent behavior, validation rules, and authority citations. The three currently-shipped providers' concrete declarations are specified in § Per-provider concrete instances.
|
||||
|
||||
#### Field specification
|
||||
|
||||
##### 1. `ephemeralEnvOverrides({ ephemeralRoot, keyId, reqId }) → { [envVar]: string }`
|
||||
|
||||
**Type and semantics.** A pure (no-side-effect, no-fs-touch) function that, given the orchestrator's composed context (`ephemeralRoot`: absolute path to the spawn-scoped temp dir; `keyId`: the OLP key identity from `lib/keys.mjs` driving the request; `reqId`: the per-request UUID), returns a flat object of environment variables that the orchestrator will merge into the spawn env. The returned env vars are how the provider CLI is steered to read its credentials from the ephemeral root rather than the server process's actual home directory.
|
||||
|
||||
**Why a function and not a static object.** Because `ephemeralRoot` is generated per-spawn by `mkdtemp` and is not known at plugin load time. Because `keyId` and `reqId` are not known until the request arrives. A static object cannot carry the dependency on these values; a function carries it cleanly.
|
||||
|
||||
**Purity contract.** The function MUST be referentially transparent w.r.t. its argument object: identical input arguments yield identical output env maps. It MUST NOT read the filesystem, spawn subprocesses, or mutate the input arguments. It MUST NOT close over module-level mutable state. This contract is what makes the spawn pipeline auditable: a reviewer reading `provider.ISOLATION.ephemeralEnvOverrides({ ephemeralRoot: '/tmp/x', keyId: 'k1', reqId: 'r1' })` can know the full env mutation without running the system.
|
||||
|
||||
**Default behavior when absent.** When `ISOLATION` is absent or `ISOLATION.ephemeralEnvOverrides` is missing, the orchestrator MUST emit no environment overrides for that provider — `child_process.spawn` runs with `process.env` (possibly modified by other contract layers such as the existing `spawn()` method's env cleanup, ADR 0009 Amendment 1's `--system-prompt` injection, etc.). This preserves Phase 6c / pre-Phase 7 behavior exactly.
|
||||
|
||||
**Validation rules.** At plugin load (in `validateProvider` or a sibling `validateIsolation` helper):
|
||||
- If `ISOLATION` is defined and `ephemeralEnvOverrides` is defined, it MUST be a function. A non-function value (e.g., a static object) is a load-time error.
|
||||
- The function is NOT invoked at load time — its return shape is not validated until first spawn. Load-time invocation would require synthetic dummy arguments and would couple the validator to the orchestrator's argument shape (which itself may evolve under future ADR 0014 amendments).
|
||||
- First-spawn invocation MUST validate the return value is a plain object whose values are all strings. Non-string values (numbers, booleans, undefined) MUST cause the spawn to abort with a clear error rather than coerce silently — the env block crosses a kernel boundary and silent coercion is a footgun.
|
||||
|
||||
**Authority citation requirement.** Each env var returned must correspond to a documented credential-resolution lookup in the underlying provider CLI. For example, `HOME` is a POSIX convention for credential lookup (well-established, no citation needed beyond the POSIX umbrella). `CODEX_HOME` is documented (primary) at https://developers.openai.com/codex/config-reference (2 occurrences verified 2026-05-29: `$CODEX_HOME/profile-name.config.toml` and `$CODEX_HOME/log` path templates), with secondary corroboration at https://developers.openai.com/codex/auth/ (2 occurrences in the credential-storage section: `auth.json under CODEX_HOME`). `VIBE_HOME` is documented at https://docs.mistral.ai/mistral-vibe/terminal/configuration (3 occurrences verified 2026-05-29: descriptive sentence "Override the location with the `VIBE_HOME` environment variable", canonical `export VIBE_HOME="/path/to/custom/vibe/home"` example, and an enumeration of files/directories `VIBE_HOME` affects). The provider plugin author MUST cite the underlying CLI's env-var documentation in the plugin file's header (the same place existing CLI-flag citations live, per Rule 1 of `ALIGNMENT.md`).
|
||||
|
||||
Inventing an env var the provider CLI does not actually honor (e.g., setting `MISTRAL_HOME=...` when no such env var exists) is a Rule 2 violation and is unalignable per Rule 4 of `ALIGNMENT.md`.
|
||||
|
||||
##### 2. `credentialMounts: [ [srcAbsPath, dstRelativeToEphemeralRoot], ... ]`
|
||||
|
||||
**Type and semantics.** An array of `[src, dst]` tuples describing how the server process's real on-disk credential artifacts (OAuth tokens, API keys, refresh artifacts) are made available inside the ephemeral home. The orchestrator iterates this list and, for each tuple, ensures `<ephemeralRoot>/<dst>` resolves (via symlink, copy, or bind-mount depending on platform and constraints) to the data at `<src>`.
|
||||
|
||||
The mount strategy is a property of the orchestrator, not the provider — `lib/sandbox/manager.mjs` decides between symlink (cheapest, on macOS and unconfined Linux), copy (when crossing a namespace boundary that breaks symlinks), and bind-mount (under a future bwrap-equipped path). The provider only declares the source-destination correspondence.
|
||||
|
||||
**Why this is a list, not a function.** The mounts are static per-provider: anthropic always mounts `~/.claude/.credentials.json`, codex always mounts `~/.codex/auth.json`. A function form would invite plugin authors to compute mount paths from per-request state, which would be a security hazard (per-request mount lists are harder to audit at code-review time). Forcing the static form makes the credential surface visible by `grep ISOLATION lib/providers/*.mjs`.
|
||||
|
||||
**Default behavior when absent.** Empty mount list — the spawn sees no credential files in its ephemeral home. For most providers this means authentication fails and the spawn errors out cleanly; the orchestrator MUST log a clear "no credentialMounts declared" message before allowing the spawn to proceed, since the most common cause is "plugin author forgot to declare the mount."
|
||||
|
||||
**Validation rules.**
|
||||
- Each entry MUST be a 2-tuple (length-2 array). Single-element entries or 3+-tuples are load-time errors.
|
||||
- `srcAbsPath` MUST be an absolute path (starts with `/`). Relative paths or `~/`-prefixed paths are load-time errors — the plugin author must call `os.homedir()` explicitly. Rationale: `~/` expansion semantics vary between Node and shells and would silently break under the per-spawn ephemeral home (where `HOME` is rewritten).
|
||||
- `dstRelativeToEphemeralRoot` MUST NOT start with `..` (no parent-directory escape) and MUST NOT be absolute (no `/etc/passwd` overlay attempts). Both are load-time errors. The orchestrator's path-composition (`path.join(ephemeralRoot, dst)`) is the *only* path-resolution step that touches the destination — the validation forbids constructions that could escape `ephemeralRoot` even before composition.
|
||||
- `srcAbsPath` MAY refer to a path that does not exist at plugin-load time. The orchestrator's mount step does a `existsSync(src)` check at spawn-time and logs a "credential source missing" warning rather than failing the spawn — this is consistent with the existing `auth.path` field behavior in the Provider contract (an absent credential file is an auth condition, not a load-time error).
|
||||
- Two mounts with the same `dst` is a load-time error (no implicit ordering or override).
|
||||
|
||||
**Authority citation requirement.** Each `srcAbsPath` MUST correspond to the credential location documented by the underlying provider CLI. For anthropic: `~/.claude/.credentials.json` is the OAuth artifact per `claude` CLI docs (already cited by the plugin's `auth.path` field). For codex: `~/.codex/auth.json` per https://developers.openai.com/codex/auth/. For mistral: `~/.vibe/.env` per https://docs.mistral.ai/mistral-vibe/terminal/configuration. Plugin authors MUST cite the same authority as the `auth.path` field they already declare — the citations should be consistent.
|
||||
|
||||
##### 3. `requiredHomePaths: [ /* relative paths */ ]`
|
||||
|
||||
**Type and semantics.** An array of relative paths (e.g., `['.claude', '.claude/logs']`) that the orchestrator MUST `mkdir -p` under `ephemeralRoot` before any `credentialMounts` are processed and before the spawn begins. These are directories the provider CLI expects to exist in `HOME` and will fail or behave incorrectly if they're absent (e.g., logging directories that the CLI doesn't auto-create).
|
||||
|
||||
**Why a separate field from `credentialMounts`.** Some providers expect empty directories — not mounted credential files — at certain paths. Treating "empty directory" as a mount with src=null would muddle the validation rules for `credentialMounts`. A dedicated list is cleaner.
|
||||
|
||||
**Default behavior when absent.** Empty list — only the directories implied by `credentialMounts[i].dst` (their parent dirs, created by `mkdir -p` during the mount step) exist under `ephemeralRoot`. For most providers this is fine.
|
||||
|
||||
**Validation rules.**
|
||||
- Each entry MUST be a relative path string. Same anti-escape rules as `credentialMounts[i].dst`: no leading `..`, no absolute paths.
|
||||
- Entries MAY overlap with `credentialMounts[i].dst` parent paths (no error; orchestrator's `mkdir -p` is idempotent).
|
||||
- Duplicate entries are not an error (idempotent), but the linter / future CI grep should flag them as a code smell.
|
||||
|
||||
**Authority citation requirement.** None directly required for the path values themselves — these are typically convention (e.g., `.claude` mirrors the CLI's expected `$HOME/.claude` layout). However, if a plugin declares a `requiredHomePaths` entry that does not correspond to any documented CLI behavior, the plugin's header comment should explain *why* the directory must exist (observed behavior, error message from CLI, etc.). Speculative directories ("just in case the CLI wants this") are a Rule 2 violation — only directories whose absence is known to cause CLI failure should be listed.
|
||||
|
||||
##### 4. `hasInnerSandbox: boolean`
|
||||
|
||||
**Type and semantics.** A boolean flag declaring whether this provider's CLI spawns its own internal sandbox boundary during normal operation. The orchestrator uses this flag to decide whether the outer isolation primitives need to be loosened to permit nested sandboxing (e.g., allow `clone(CLONE_NEWUSER)` syscalls, permit `bwrap` to nest).
|
||||
|
||||
**Why a boolean and not an enum.** "Has inner sandbox or not" is the discriminator the orchestrator needs. The *kind* of inner sandbox (bwrap, sandbox-exec, seccomp-only) is a detail the orchestrator does not need to compose against — it just needs to know whether to relax the outer profile. If a future provider requires per-sandbox-flavor handling, this field can be widened to an enum in a subsequent amendment.
|
||||
|
||||
**Default behavior when absent.** Treated as `false`. This is the safer-by-default value — outer isolation stays at its strictest setting. A provider that actually has an inner sandbox but forgets to declare it will fail at spawn time (inner-bwrap attempts denied by outer profile); the failure mode is loud and obvious, which is the desired behavior.
|
||||
|
||||
**Validation rules.** MUST be a literal `true` or `false`. Truthy/falsy coercion (e.g., declaring `1` or `'yes'`) is a load-time error — booleans are the documented type and coercion would silently change the orchestrator's composition decision.
|
||||
|
||||
**Authority citation requirement.** A `hasInnerSandbox: true` declaration MUST cite the CLI's documented or observed inner-sandbox behavior in the plugin header. For codex, the citation is `openai/codex#16018` (the GitHub issue documenting `codex exec` invoking bubblewrap internally) plus https://developers.openai.com/codex/concepts/sandboxing (the official docs page describing the `--sandbox` flag and `read-only` default). For a hypothetical future provider, the citation is whatever CLI doc or observed-behavior transcript establishes the inner sandbox.
|
||||
|
||||
##### 5. `crossTenantReadProtection: 'tool-suppression' | 'inner-sandbox' | 'none'`
|
||||
|
||||
**Type and semantics.** A discriminated string declaring the regime under which this provider's spawn is protected against cross-tenant filesystem reads. The three values correspond to the three regimes observed in the 2026-05-27 prior-art / incident analysis (see incident memory § 6):
|
||||
|
||||
- `'tool-suppression'` — the provider's CLI exposes no filesystem-reading tools to the model during the spawn, because OLP suppresses them at the request level. For anthropic, this is achieved via ADR 0009 Amendment 1's `--system-prompt` injection combined with the absence of `--tools` flags: the model has no shell, no file-read, no bash, no Read/Write/Edit primitives. The cross-tenant read surface is closed at the prompt-engineering layer; OS-level isolation is a defense in depth but not the primary regime.
|
||||
|
||||
- `'inner-sandbox'` — the provider's CLI has tool execution (e.g., codex's `shell` tool, which actually runs commands) but the CLI's own inner sandbox prevents the tool from reading paths outside its declared allow-list. For codex, the inner bwrap sandbox enforces `--sandbox read-only` by default (per https://developers.openai.com/codex/concepts/sandboxing), so even though the model can call `shell`, the shell's reads are confined to the inner namespace. The cross-tenant read surface is closed at the inner-sandbox layer.
|
||||
|
||||
- `'none'` — no protection regime is currently established for this provider. The model may have tools that read files, and there is no inner sandbox blocking those reads. Operationally this means the provider should NOT be enabled in a multi-tenant deployment until a regime is established. The orchestrator MUST log a WARN at server boot when a provider with `crossTenantReadProtection: 'none'` is enabled in a deployment with >1 active OLP key — observability, not enforcement (see Rule 4 compliance below).
|
||||
|
||||
**Why a discriminated enum, not a free-form string.** The orchestrator and the operator dashboard both consume this field. Free-form values would require every consumer to perform string-matching against a moving target. The enum locks the consumer surface; future regimes are added by amending this list in a subsequent ADR 0002 amendment.
|
||||
|
||||
**Default behavior when absent.** Treated as `'none'`. Safer-by-default in the WARN sense (operators get the WARN log) but NOT in the security sense (no protection is actually applied). This is intentional: the orchestrator cannot fabricate a protection regime the plugin hasn't implemented; the WARN nudges the plugin author to declare honestly.
|
||||
|
||||
**Validation rules.** MUST be one of the three enum values literally. Any other string is a load-time error. The orchestrator MUST log the field's value at server startup so operators can audit the protection picture across providers at a glance.
|
||||
|
||||
**Authority citation requirement.**
|
||||
- `'tool-suppression'` declarations MUST cite the suppression mechanism (e.g., for anthropic: ADR 0009 Amendment 1 § "--system-prompt" + the absence-of-tools posture documented at the incident memory § 6.1).
|
||||
- `'inner-sandbox'` declarations MUST cite the CLI doc or observed behavior establishing the inner sandbox (e.g., for codex: `openai/codex#16018` + https://developers.openai.com/codex/concepts/sandboxing).
|
||||
- `'none'` is the safer default and requires no citation but MUST be accompanied by a header-comment TODO documenting what regime is expected to be established when the provider transitions from Candidate to Enabled (or earlier if the provider is enabled in a multi-tenant context).
|
||||
|
||||
##### 6. `recommendedDeploymentTier: 'shared-os-user' | 'per-os-user' | 'separate-vm'`
|
||||
|
||||
**Type and semantics.** A discriminated string giving operators a deployment-topology recommendation for this provider in a multi-tenant context. The three values express increasing degrees of operator-side isolation:
|
||||
|
||||
- `'shared-os-user'` — the OLP server process runs as a single OS user, and multiple OLP keys share that user. Protection against cross-tenant leakage rests entirely on the provider's `crossTenantReadProtection` regime + the orchestrator's ephemeral-home composition. This is the recommended posture for providers where `crossTenantReadProtection` is `'tool-suppression'` AND `hasInnerSandbox: false` (i.e., the model has no filesystem-touching tools at all).
|
||||
|
||||
- `'per-os-user'` — each OLP key (or each tenant) should map to a separate OS user, with file-permission-level isolation between tenants. The recommended posture for providers with `crossTenantReadProtection: 'inner-sandbox'` — the inner sandbox protects against accidental leakage from the model's tools, but a sandbox-escape (e.g., a CVE in bubblewrap, a misconfigured inner profile) would expose the OS-user filesystem; per-OS-user isolation adds defense in depth.
|
||||
|
||||
- `'separate-vm'` — the provider should not be co-located with any other tenant on the same VM. The recommended posture for providers with `crossTenantReadProtection: 'none'` AND/OR ones where the operator has reason to distrust the inner sandbox's quality. Practically this means the provider should not be enabled in OLP's family-LAN deployment unless the family-LAN host runs only this tenant.
|
||||
|
||||
**Why a recommendation and not a hard policy.** The orchestrator and OLP runtime cannot *enforce* OS-user separation or VM separation — those are properties of the host operator's deployment topology. This field is informational: it surfaces in `/health.providers.<name>.isolation` (a Phase 7 addition planned in a follow-up amendment) and in the dashboard, so operators making deployment decisions have the per-provider recommendation visible. Operator override is the expected normal path: a deployment that knowingly accepts the risk of running an `'separate-vm'` provider in a shared-user context is acceptable, just observable.
|
||||
|
||||
**Default behavior when absent.** Treated as `'separate-vm'` — the safest recommendation in the absence of declared analysis. The WARN log emitted for missing `ISOLATION` blocks (see Rule 4 compliance below) covers operator visibility.
|
||||
|
||||
**Validation rules.** MUST be one of the three enum values literally. Any other string is a load-time error.
|
||||
|
||||
**Authority citation requirement.** The plugin author MUST cite the basis for the recommendation in the plugin header — typically a short paragraph reasoning about the combination of `hasInnerSandbox` and `crossTenantReadProtection` for this provider. The reasoning is not a CLI authority citation (the underlying CLI does not declare deployment topology); it is an OLP-side analysis. The expected citation form is `# isolation rationale: <2-3 sentences> (cf. ADR 0014 Amendment 1 § <relevant section>)`.
|
||||
|
||||
##### 7. `toolHardeningArgs: (existingArgs) => modifiedArgs` (OPTIONAL)
|
||||
|
||||
**Type and semantics.** An OPTIONAL pure function that, given the plugin's `spawn()` method's CLI args (the array passed to `child_process.spawn`), returns a (possibly modified) args array with additional tool-hardening flags inserted. The orchestrator calls this hook after the plugin's `spawn()` constructs its args but before the actual `child_process.spawn` invocation.
|
||||
|
||||
**Purpose.** Some providers expose CLI flags that suppress or restrict the model's tool access at the per-spawn level (e.g., `--disallowedTools` on `claude`, or `--sandbox read-only` on `codex`). These flags are the *enforcement mechanism* corresponding to the `crossTenantReadProtection` *declaration*. Splitting the declaration (a static field) from the enforcement (a function that mutates args) keeps the contract auditable while letting the enforcement evolve as the underlying CLI's flag set changes.
|
||||
|
||||
**Why this is OPTIONAL.** For providers where `crossTenantReadProtection: 'tool-suppression'` is achieved entirely via the `spawn()` method's existing args construction (e.g., the existing anthropic.mjs `--system-prompt` injection), no separate hardening step is needed — the field can be omitted. For providers where the orchestrator needs to inject additional flags atop the plugin's base args, the field provides the hook.
|
||||
|
||||
**Default behavior when absent.** No args modification — the plugin's `spawn()` method's args are passed through to `child_process.spawn` unchanged. This is the current Phase 6c behavior for anthropic and is appropriate when the `spawn()` method already encodes the hardening.
|
||||
|
||||
**Validation rules.**
|
||||
- If declared, MUST be a function.
|
||||
- First-spawn invocation MUST validate the return value is an array of strings. Non-array or non-string-element returns abort the spawn (silent coercion is unsafe at the kernel boundary).
|
||||
- The function MUST be referentially transparent — same input array yields same output array (no module-level state, no fs reads).
|
||||
- The orchestrator MUST NOT pass the args by reference in a way that the function could mutate the original `existingArgs`. The hook receives a defensive copy; returning a fresh array is required.
|
||||
|
||||
**Authority citation requirement.** The injected flags MUST be documented CLI flags of the underlying provider. Inventing a `--disable-tools` flag that the CLI does not support is a Rule 2 violation. For codex, citing https://developers.openai.com/codex/concepts/sandboxing § `--sandbox` is sufficient. For anthropic, the existing ADR 0009 Amendment 1 citation covers the tool-suppression mechanism.
|
||||
|
||||
#### Per-provider concrete instances
|
||||
|
||||
The three currently-shipped providers declare `ISOLATION` as follows. Each declaration MUST be present in the corresponding plugin file before that provider can be enabled in any multi-tenant deployment (see § Rule 4 compliance and § Backward compatibility for the transition path).
|
||||
|
||||
##### anthropic
|
||||
|
||||
```javascript
|
||||
// lib/providers/anthropic.mjs
|
||||
//
|
||||
// isolation rationale: Anthropic Claude reaches OLP via stream-json transport
|
||||
// without a tool surface (ADR 0009 Amendment 1's --system-prompt injection
|
||||
// suppresses env-block, file tools, bash, and Read/Write/Edit). The model
|
||||
// has no documented mechanism to read files during the spawn. Cross-tenant
|
||||
// read protection is achieved at the prompt-engineering / CLI-flag layer.
|
||||
// The OS-level isolation primitives (HOME redirect + ephemeral credential
|
||||
// mount) add defense in depth against future CLI changes that might
|
||||
// re-introduce a tool surface.
|
||||
//
|
||||
// Authority: @anthropic-ai/claude-code v2.1.150 § --system-prompt
|
||||
// (ADR 0009 Amendment 1 + incident memory § 6.1 establishes the
|
||||
// tool-suppression mechanism); HOME env conventional POSIX behavior.
|
||||
|
||||
export const ISOLATION = {
|
||||
ephemeralEnvOverrides: ({ ephemeralRoot, keyId, reqId }) => ({
|
||||
HOME: ephemeralRoot,
|
||||
// CLAUDE_CONFIG_DIR is NOT honored as of v2.1.150 — the CLI reads from
|
||||
// $HOME/.claude/.credentials.json. Redirecting HOME is the documented
|
||||
// mechanism. The keyId / reqId arguments are unused here but received for
|
||||
// signature consistency with codex's overrides.
|
||||
}),
|
||||
credentialMounts: [
|
||||
// OAuth artifact location. Authority: existing anthropic.mjs `auth.path`
|
||||
// field — `~/.claude/.credentials.json` is the documented OAuth artifact.
|
||||
// The orchestrator resolves the absolute src path via os.homedir() at
|
||||
// load time (the plugin file shows the literal `join(homedir(), ...)`).
|
||||
[/* resolved at load: */ '<homedir>/.claude/.credentials.json',
|
||||
'.claude/.credentials.json'],
|
||||
],
|
||||
requiredHomePaths: [
|
||||
'.claude',
|
||||
// No observed behavior requires additional dirs; CLI creates session logs
|
||||
// under .claude/ on demand. If future CLI versions add a mandatory pre-
|
||||
// existing subdir, add it here with an observed-behavior comment.
|
||||
],
|
||||
hasInnerSandbox: false,
|
||||
crossTenantReadProtection: 'tool-suppression',
|
||||
recommendedDeploymentTier: 'shared-os-user',
|
||||
// toolHardeningArgs omitted — the existing spawn() method's args already
|
||||
// encode the --system-prompt suppression (ADR 0009 Amendment 1).
|
||||
};
|
||||
```
|
||||
|
||||
**Authority pin for the anthropic ISOLATION declaration:**
|
||||
- `--system-prompt` mechanism: ADR 0009 Amendment 1 + incident memory `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` § 6.1
|
||||
- `HOME` env redirect: POSIX convention; `claude` CLI v2.1.150 observed to read `~/.claude/.credentials.json` via HOME (verified by the PR-B PI231 spike, confirmed by ADR 0014 Amendment 1's HOME-override verification task)
|
||||
|
||||
##### codex
|
||||
|
||||
```javascript
|
||||
// lib/providers/codex.mjs
|
||||
//
|
||||
// isolation rationale: OpenAI Codex's `codex exec` exposes a shell tool that
|
||||
// actually executes commands during the spawn (incident memory § 3.2). The
|
||||
// CLI provides its own inner bubblewrap sandbox (`--sandbox read-only` by
|
||||
// default per https://developers.openai.com/codex/concepts/sandboxing) that
|
||||
// confines shell tool reads/writes. The orchestrator's outer isolation
|
||||
// composes with the inner sandbox: HOME-equivalent redirect via CODEX_HOME
|
||||
// (per https://developers.openai.com/codex/config-reference) plus per-spawn
|
||||
// ephemeral credential mount. hasInnerSandbox: true so the outer profile is
|
||||
// relaxed to permit inner bwrap's user-namespace clone.
|
||||
//
|
||||
// Authority: openai/codex#16018 (inner bwrap behavior);
|
||||
// https://developers.openai.com/codex/concepts/sandboxing (--sandbox flag);
|
||||
// https://developers.openai.com/codex/config-reference (CODEX_HOME);
|
||||
// https://developers.openai.com/codex/auth/ (~/.codex/auth.json path).
|
||||
|
||||
export const ISOLATION = {
|
||||
ephemeralEnvOverrides: ({ ephemeralRoot, keyId, reqId }) => ({
|
||||
// CODEX_HOME overrides the base config / credential dir. Docs:
|
||||
// https://developers.openai.com/codex/config-reference and
|
||||
// https://developers.openai.com/codex/auth/
|
||||
CODEX_HOME: `${ephemeralRoot}/.codex`,
|
||||
// HOME also redirected for codex's own bubblewrap-internal HOME lookup
|
||||
// (the inner sandbox inherits parent HOME unless overridden).
|
||||
HOME: ephemeralRoot,
|
||||
}),
|
||||
credentialMounts: [
|
||||
// Auth artifact location. Authority: existing codex.mjs `auth.path` field
|
||||
// (Codex CLI reference § Authentication, plus
|
||||
// https://developers.openai.com/codex/auth/ canonical pin).
|
||||
[/* resolved at load: */ '<homedir>/.codex/auth.json',
|
||||
'.codex/auth.json'],
|
||||
],
|
||||
requiredHomePaths: [
|
||||
'.codex',
|
||||
// Inner bwrap may create additional state under .codex/. If observed
|
||||
// behavior shows the CLI failing on absent subdirs, add them here.
|
||||
],
|
||||
hasInnerSandbox: true,
|
||||
crossTenantReadProtection: 'inner-sandbox',
|
||||
recommendedDeploymentTier: 'per-os-user',
|
||||
toolHardeningArgs: (existingArgs) => {
|
||||
// If the operator has not explicitly passed --sandbox, inject the
|
||||
// documented read-only default. Per
|
||||
// https://developers.openai.com/codex/concepts/sandboxing the default
|
||||
// posture is `read-only`; this hardening hook makes the default explicit
|
||||
// at the spawn args level so a future CLI default change does not
|
||||
// silently weaken the isolation.
|
||||
if (existingArgs.some(arg => arg === '--sandbox' || arg.startsWith('--sandbox='))) {
|
||||
return existingArgs;
|
||||
}
|
||||
return [...existingArgs, '--sandbox', 'read-only'];
|
||||
},
|
||||
};
|
||||
```
|
||||
|
||||
**Authority pin for the codex ISOLATION declaration:**
|
||||
- `CODEX_HOME`: https://developers.openai.com/codex/config-reference (retrieved 2026-05-29)
|
||||
- `~/.codex/auth.json`: https://developers.openai.com/codex/auth/ (existing `auth.path` citation in codex.mjs)
|
||||
- Inner bwrap behavior: `openai/codex#16018` plus https://developers.openai.com/codex/concepts/sandboxing
|
||||
- `--sandbox read-only`: https://developers.openai.com/codex/concepts/sandboxing § "Sandboxing modes"
|
||||
|
||||
##### mistral
|
||||
|
||||
```javascript
|
||||
// lib/providers/mistral.mjs
|
||||
//
|
||||
// isolation rationale: Mistral Vibe ships at OLP Phase 7 with no known
|
||||
// equivalent to Anthropic's Phase 6c --system-prompt tool suppression and
|
||||
// no known inner sandbox. The IR-level normalization shipped at D8 does not
|
||||
// suppress tools at the CLI layer. Cross-tenant read protection is therefore
|
||||
// 'none' — the provider should not be enabled in a multi-tenant deployment
|
||||
// until a regime is established. The declaration here exists so the
|
||||
// orchestrator can compose ephemeral-home credential isolation (which still
|
||||
// works) while the operator sees a clear WARN that the tool-side protection
|
||||
// is not in place.
|
||||
//
|
||||
// Authority: TBD — a spike task tracked at Phase 7 follow-up (see Open
|
||||
// Questions section below) will verify Vibe CLI's tool surface and inner
|
||||
// sandbox posture against https://docs.mistral.ai/mistral-vibe/terminal/.
|
||||
// Until that spike lands, this declaration documents the current honest
|
||||
// state per ALIGNMENT.md Rule 3 (Match the Implementation): no protection
|
||||
// is encoded because none has been established.
|
||||
|
||||
export const ISOLATION = {
|
||||
ephemeralEnvOverrides: ({ ephemeralRoot, keyId, reqId }) => ({
|
||||
// VIBE_HOME is documented at
|
||||
// https://docs.mistral.ai/mistral-vibe/terminal/configuration as the
|
||||
// env var that overrides the default ~/.vibe/ base directory
|
||||
// (3 occurrences verified 2026-05-29: descriptive sentence,
|
||||
// canonical export example, and an enumeration of files/dirs the
|
||||
// variable affects). Task #4 PI231 spike verifies observed CLI
|
||||
// behaviour matches the documented contract.
|
||||
VIBE_HOME: `${ephemeralRoot}/.vibe`,
|
||||
HOME: ephemeralRoot,
|
||||
}),
|
||||
credentialMounts: [
|
||||
// ~/.vibe/.env per existing mistral.mjs `auth.path` field, sourced from
|
||||
// https://docs.mistral.ai/mistral-vibe/terminal/configuration.
|
||||
[/* resolved at load: */ '<homedir>/.vibe/.env', '.vibe/.env'],
|
||||
],
|
||||
requiredHomePaths: [
|
||||
'.vibe',
|
||||
],
|
||||
hasInnerSandbox: false,
|
||||
crossTenantReadProtection: 'none',
|
||||
recommendedDeploymentTier: 'separate-vm',
|
||||
// toolHardeningArgs omitted — no CLI hardening flag is currently known for
|
||||
// Vibe. The Phase 7 spike will revisit.
|
||||
};
|
||||
```
|
||||
|
||||
**Authority pin for the mistral ISOLATION declaration:**
|
||||
- `VIBE_HOME`: https://docs.mistral.ai/mistral-vibe/terminal/configuration (3 occurrences verified 2026-05-29: descriptive sentence "Override the location with the `VIBE_HOME` environment variable", canonical `export VIBE_HOME="/path/to/custom/vibe/home"` example, and the enumeration of files/directories `VIBE_HOME` affects).
|
||||
- `~/.vibe/.env`: same source (existing `auth.path` citation in mistral.mjs).
|
||||
- **Open spike (Phase 7 follow-up, Task #4):** verify *observed CLI behaviour* matches *documented behaviour* — (a) Vibe CLI actually honours the documented `VIBE_HOME` env var during spawn; (b) Vibe CLI's tool surface (shell, file-read, etc.) during a `vibe --prompt` spawn; (c) any CLI sandbox or tool-suppression flag. Findings may transition `crossTenantReadProtection` from `'none'` to `'tool-suppression'` or `'inner-sandbox'` if a hardening regime is discovered. The spike is verification-grade, not authority-pin work.
|
||||
|
||||
#### Backward compatibility
|
||||
|
||||
A plugin that does NOT export `ISOLATION` continues to work exactly as it does today. The orchestrator's `prepareIsolatedEnvironment(provider, ctx)` function MUST detect the absence of `provider.ISOLATION` (or the absence of any individual field within it) and fall through to the legacy unsandboxed code path for that spawn. The legacy path is:
|
||||
|
||||
- No ephemeral root created
|
||||
- No env overrides
|
||||
- No credential mounts
|
||||
- `cwd: process.cwd()` (the server's working directory)
|
||||
- `env: process.env` (composed with whatever the plugin's `spawn()` method's existing env logic produces)
|
||||
|
||||
This is the same behavior as Phase 6c. No provider plugin is broken by Amendment 9's landing.
|
||||
|
||||
Plugins MAY adopt `ISOLATION` incrementally: a plugin that wants the credential-mount benefit but has not yet analyzed its cross-tenant tool surface MAY declare `crossTenantReadProtection: 'none'` and `recommendedDeploymentTier: 'separate-vm'` (the safer-by-default values). The orchestrator will compose the credential isolation correctly; the WARN log nudges follow-up.
|
||||
|
||||
#### Rule 4 compliance (ALIGNMENT.md)
|
||||
|
||||
ALIGNMENT.md Rule 4 states: "Unalignable plugins / fields are deleted, not feature-flagged." This amendment introduces an OPTIONAL contract field, which on its face could be read as "feature-flagging" isolation. The reading is wrong, and the distinction is important enough to spell out:
|
||||
|
||||
- Amendment 9 does NOT introduce an `ISOLATION` feature flag that operators or plugins toggle on/off. The field's presence/absence describes **the provider's truthful isolation posture** at a point in time. A plugin without `ISOLATION` declares (implicitly) that no analysis has been done and the safer-by-default treatment applies.
|
||||
- The OPTIONAL nature is purely transitional. Existing plugins ship without it; they continue to spawn (in their existing single-tenant developer-laptop posture). The orchestrator's WARN log surfaces the absence to the operator at server boot. An operator running a multi-tenant deployment with un-declared plugins is operating off-recommendation but not blocked.
|
||||
- A plugin that declares `ISOLATION` with values the orchestrator cannot honor (e.g., a `credentialMounts` entry pointing at a path that does not exist, or an `ephemeralEnvOverrides` function that returns non-string values) MUST fail at first spawn — the orchestrator does not silently fall back to the no-ISOLATION path. This is the Rule 4 enforcement vector: a *broken* declaration is unalignable and surfaces loudly; a *missing* declaration is the safer transitional state.
|
||||
|
||||
The WARN at server boot is observability, not enforcement. It reads approximately:
|
||||
|
||||
```
|
||||
[WARN] provider "<name>" does not declare ISOLATION; spawns will run
|
||||
under legacy unsandboxed shape. Recommended in multi-tenant
|
||||
deployments: declare ISOLATION per ADR 0002 Amendment 9.
|
||||
```
|
||||
|
||||
Operators in single-tenant developer deployments may safely ignore the WARN. Operators in multi-tenant deployments should treat it as a Phase 7 follow-up task.
|
||||
|
||||
#### Interaction with prior amendments
|
||||
|
||||
- **Amendment 1 (`maxSpawnTimeMs`).** Independent. The spawn-timeout enforcement lives inside each plugin's spawn drain loop; the orchestrator's ISOLATION composition happens *before* the spawn, so the two amendments compose without conflict.
|
||||
- **Amendment 3 (`cacheable`).** Independent. The cache layer decides whether to call the orchestrator at all; once the orchestrator is reached, ISOLATION composition is orthogonal to cacheability.
|
||||
- **Amendment 4 (`contractVersion`).** Independent. `contractVersion: '1.0'` plugins MAY add an `ISOLATION` export under Amendment 9 without bumping the contract version — `ISOLATION` is an additive named export, not a v1.0 contract surface change. A future Provider contract v1.1 may promote `ISOLATION` to a required field (forcing all enabled plugins to declare); that decision is deferred to a future amendment, gated on the Phase 7 follow-up findings.
|
||||
- **Amendment 6 (`maxConcurrent` runtime enforcement).** Independent. The semaphore acquire happens before the orchestrator's `prepareIsolatedEnvironment`; the release happens after the spawn drains. ISOLATION composition is bracketed by the semaphore, not entangled with it.
|
||||
- **Amendment 7 (`doctorChecks()`).** Adjacent. A future plugin may add an `<provider>.isolation_declared` doctor check that reports whether `ISOLATION` is declared and whether its referenced credential paths resolve. The check is OPTIONAL per Amendment 7's framework and is appropriate for `olp doctor` operator UX.
|
||||
- **Amendment 8 (`quotaStatus()` direct-API exemption).** Independent. The quota probe runs outside the spawn pipeline (direct HTTPS from server process); it does not interact with `ISOLATION` composition.
|
||||
|
||||
#### Companion ADR
|
||||
|
||||
This amendment is the companion governance piece for **ADR 0014 Amendment 1** (the Phase 7 architectural shift from outer-bwrap PR-B to per-spawn ephemeral-home + per-provider primitives). ADR 0014 Amendment 1 describes the orchestrator's composition algorithm and the rationale for retiring the outer-bwrap approach; ADR 0002 Amendment 9 (this section) describes the contract surface the orchestrator reads.
|
||||
|
||||
The two amendments are reviewed and merged together as a single coupled commit (Iron Rule 11 — minimum reviewable unit per layer). Reviewing them separately cannot verify producer-consumer alignment: the orchestrator's algorithm is meaningless without the contract it consumes, and the contract is meaningless without the orchestrator's composition discipline.
|
||||
|
||||
#### Tests
|
||||
|
||||
Test coverage for Amendment 9 lands as a new Suite in `test-features.mjs` co-merged with ADR 0014 Amendment 1's `lib/sandbox/manager.mjs` refactor. The suite covers:
|
||||
|
||||
1. `validateProvider` (or `validateIsolation` helper) rejects each documented invalid shape: non-function `ephemeralEnvOverrides`; non-2-tuple `credentialMounts` entries; `dst` paths starting with `..` or absolute; non-boolean `hasInnerSandbox`; out-of-enum `crossTenantReadProtection`; out-of-enum `recommendedDeploymentTier`; non-function `toolHardeningArgs`.
|
||||
2. The legacy code path: a fake provider without `ISOLATION` spawns under the existing shape unchanged. Existing Phase 6c tests for anthropic continue to pass.
|
||||
3. The ephemeral-home composition path: a fake provider declaring a minimal `ISOLATION` block has its env overrides applied and its credential mount resolved into a `mkdtemp`-created ephemeral root.
|
||||
4. First-spawn return-shape validation: `ephemeralEnvOverrides` returning non-string values aborts the spawn loudly; `toolHardeningArgs` returning a non-array aborts the spawn loudly.
|
||||
5. Per-shipped-provider declaration smoke: each of `anthropic`, `codex`, `mistral` declares an `ISOLATION` block; each block's `credentialMounts[i][0]` (when resolved against the running user's `homedir()`) matches the plugin's `auth.path` field.
|
||||
|
||||
The full test list is captured in ADR 0014 Amendment 1's PR-B-revised test suite specification.
|
||||
|
||||
#### Open questions (Phase 7 follow-up)
|
||||
|
||||
1. **Mistral Vibe tool surface and inner sandbox.** The mistral plugin's `ISOLATION` declares `crossTenantReadProtection: 'none'` honestly. A spike task is required to determine whether Vibe CLI exposes any tool surface and/or any sandbox flag; findings update the declaration. Tracked at the Phase 7 work plan.
|
||||
2. **HOME-only providers vs CODEX_HOME-style providers.** The current contract assumes credential redirection happens via env-var rewriting (`HOME` or `<PROVIDER>_HOME`). A future provider that hardcodes its credential path (no env override) would be unable to honor the contract and would need a different isolation strategy (e.g., bind-mount of the literal path). This is not a current problem (all three shipped providers honor env overrides) but should be tracked for future inclusion ADRs.
|
||||
3. **Promoting `ISOLATION` to required at contract v1.1.** Once all enabled providers declare `ISOLATION`, a future contract-version bump may promote the field from OPTIONAL to REQUIRED. The decision is gated on operational experience after PI231 + cloud deployment — see ADR 0014 Amendment 1 for the rollout milestones.
|
||||
4. **Per-spawn vs per-key ephemeral root.** This amendment specifies per-spawn ephemeral roots (one `mkdtemp` per `provider.spawn` call). A future optimization may cache ephemeral roots per-key (one ephemeral root per OLP key identity, reused across spawns) to reduce mkdtemp / mount overhead. The contract surface here is compatible with either strategy; the choice is an orchestrator implementation detail.
|
||||
5. **Cleanup discipline.** The orchestrator is responsible for `rm -rf`-ing the ephemeral root after the spawn drains. The cleanup mechanism (synchronous vs deferred, error vs success path symmetry) is specified in ADR 0014 Amendment 1, not here. This amendment notes the dependency for completeness.
|
||||
|
||||
#### Authority citations summary
|
||||
|
||||
| Field | Authority |
|
||||
|---|---|
|
||||
| `ephemeralEnvOverrides` (general) | POSIX `HOME` convention; per-provider env-var documentation cited per declaration |
|
||||
| `credentialMounts` (general) | Each plugin's existing `auth.path` field citation |
|
||||
| `requiredHomePaths` (general) | Observed CLI behavior; no speculative entries (Rule 2) |
|
||||
| `hasInnerSandbox` (general) | CLI doc or observed-behavior transcript |
|
||||
| `crossTenantReadProtection` (enum) | OLP-side analysis based on prior-art search in incident memory `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` § 4 + § 6 |
|
||||
| `recommendedDeploymentTier` (enum) | OLP-side analysis; ADR 0014 Amendment 1 § Deployment topology |
|
||||
| `toolHardeningArgs` (function) | Documented CLI flags of the underlying provider; no invented flags (Rule 2) |
|
||||
| anthropic `--system-prompt` tool suppression | ADR 0009 Amendment 1 + incident memory § 6.1 |
|
||||
| codex `CODEX_HOME` | https://developers.openai.com/codex/config-reference + https://developers.openai.com/codex/auth/ |
|
||||
| codex inner bwrap | openai/codex#16018 + https://developers.openai.com/codex/concepts/sandboxing |
|
||||
| codex `--sandbox read-only` default | https://developers.openai.com/codex/concepts/sandboxing § "Sandboxing modes" |
|
||||
| mistral `VIBE_HOME` and `.vibe/.env` | https://docs.mistral.ai/mistral-vibe/terminal/configuration |
|
||||
|
||||
#### Procedural mechanism
|
||||
|
||||
- **Iron Rule 11 (Incremental Diff Review)** — Amendment 9 (governance, ADR 0002) and ADR 0014 Amendment 1 (orchestrator architecture) land as a single coupled PR. Reviewing them separately cannot verify consumer-producer alignment.
|
||||
- **Iron Rule 10 (Code Review)** — independent fresh-context reviewer per `CLAUDE.md` hard requirement #3. The reviewer MUST open each cited authority URL (Codex config-reference, sandboxing docs, Mistral configuration docs, the incident memory) and confirm the citation in the review comment.
|
||||
- **`ALIGNMENT.md` Rule 1 (Cite First)** — every per-field design choice is cited above. Every per-provider concrete instance is cited to the underlying CLI authority.
|
||||
- **`ALIGNMENT.md` Rule 2 (No Invention)** — no invented env vars, no invented CLI flags. The mistral `crossTenantReadProtection: 'none'` declaration is the explicit honest acknowledgment that no protection regime has been established, rather than invention of one.
|
||||
- **`ALIGNMENT.md` Rule 4 (Unalignable Plugins / Fields Are Deleted)** — see § Rule 4 compliance above for the explicit reasoning that OPTIONAL `ISOLATION` is not "feature-flagging" but rather "honestly transitional."
|
||||
- **`ALIGNMENT.md` Amendment Procedure** — this section (Amendment 9) is the PR-required citation of evidence (the 2026-05-27 incident memory, the ADR 0014 PoC spike report at `/tmp/sandbox-spike/report.md` on PI231) and the structural amendment of the Provider contract documented in this ADR's § Decision.
|
||||
|
||||
@@ -7,6 +7,23 @@
|
||||
|
||||
## Amendments
|
||||
|
||||
### Amendment 3 — 2026-05-27: Accept OpenAI `role: "developer"` at entry surface, normalize to `system` in IR
|
||||
|
||||
- **Finding:** Hermes Agent v0.13/v0.14 (and likely Cline, Continue.dev, and other modern openai-completions clients) default to `role: "developer"` for what was historically the `system`-role slot when the model id matches OpenAI's o1/o3+ reasoning family. The `developer` role was introduced by OpenAI's Responses-API spec for reasoning models (high-priority developer-authored instructions; semantically a peer of `system`). OLP IR's role allow-list at v0.1 was the original four roles (`system|user|assistant|tool`); the IR validator rejected `developer` with `400 IR validation failed: role must be one of system|user|assistant|tool, got "developer"`. Reproduced 2026-05-27 on PI230 Hermes v0.14.0 → OLP v0.5.1 routing path.
|
||||
- **Decision:** Extend `openai-to-ir.mjs:normalizeRole()` to map `developer` → `system` at the entry boundary. The IR's canonical-four-roles invariant is preserved; every provider plugin's role-handling stays unchanged. The normalize-at-entry pattern matches the existing `function` → `tool` normalization that already lives in the same function (function-role-deprecation was the precedent for entry-boundary normalization vs. IR schema bloat).
|
||||
- **Why not "add `developer` to VALID_ROLES + handle in every provider":** That alternative would require:
|
||||
- Expanding `VALID_ROLES` in `lib/ir/types.mjs`.
|
||||
- Adding `developer` branch in `anthropic.mjs:irToAnthropic` (which would map to `[System]` annotation anyway).
|
||||
- Adding `developer` branch in `codex.mjs:irToCodex` (would map to `[System]` annotation anyway).
|
||||
- Adding `developer` branch in `mistral.mjs:irToMistral` (same).
|
||||
- Coordinating every future role addition (e.g., if OpenAI adds another role tomorrow) across N provider plugins.
|
||||
- Wider IR surface area = more drift-prone over time.
|
||||
Normalize-at-entry centralizes role-spec-evolution handling in one file. ADR 0003's IR-design principle ("encode the common subset every provider plugin can consume") supports keeping the IR minimal.
|
||||
- **Forward note:** Future OpenAI role additions follow the same pattern: extend `normalizeRole()`. If a role genuinely conveys provider-distinguishable semantics (e.g., a hypothetical role that meaningfully changes anthropic vs codex behavior), the calculus flips and a IR-level addition would be justified. That decision goes through a new ADR 0003 amendment.
|
||||
- **Cache-key impact:** After this amendment, a request whose first message uses `role: "developer"` and one using `role: "system"` with otherwise-identical content produce the **same** IR (because normalization happens before IR construction) → the **same** cache key (per ADR 0005 cache key composition). This is intentional and matches OpenAI's own backward-compat behavior ("system message with reasoning models is treated as developer"). If a future debug session is investigating "why does my new `developer` request hit a cache entry from an old `system` request" — this is by design.
|
||||
- **Tests:** Suite IR translation in `test-features.mjs` gains three pin tests: (a) `role: "developer"` → `role: "system"` translation, (b) mixed-role array including developer validates cleanly through to IR, (c) negative control — an unknown role (e.g. `"admin"`) still raises `BadRequestError`, confirming the normalize-at-entry mapping did not accidentally widen the role allow-list.
|
||||
- **Authority:** OpenAI Responses API spec — developer role documented as high-priority developer-authored instructions for o1/o3+ reasoning models (https://platform.openai.com/docs/api-reference/responses). Hermes Agent / Cline / Continue.dev tracking the same convention. Reproduced live on PI230 → PI231 OLP 2026-05-27.
|
||||
|
||||
### Amendment 2 — 2026-05-24: Correct model-mapping example; document verbatim-pass-through design (D32 F2)
|
||||
|
||||
- **Finding:** Round-4 cold-audit F2 (P3 ADR example vs implementation drift) — § Decision "Required fields" item `model` reads: "The provider plugin maps this to the provider-native model identifier (e.g., `claude-sonnet-4-6` → `claude-sonnet-4-6-20260301` for Anthropic)." This is WRONG per the D17 SPOT decision (commit `cb86807`): OLP does NOT perform a model-alias mapping inside the provider plugin. `irRequest.model` is passed verbatim to the provider CLI (`claude -p --model <model>`, `codex exec --model <model>`, etc.); each provider's CLI resolves its own aliases natively per its documented behaviour.
|
||||
|
||||
@@ -2,6 +2,154 @@
|
||||
|
||||
- **Date:** 2026-05-25
|
||||
- **Status:** Accepted (D48, design-only — implementation D-days D49–D54 follow; Phase 3 close = v0.3.0)
|
||||
|
||||
## Amendments
|
||||
|
||||
### Amendment 2 — 2026-05-27: v0.5.1 quota_v2 richer failure-mode shape (codex finding F3)
|
||||
|
||||
**Scope:** v0.5.1 hotfix extends `ProviderQuotaEntry` and `aggregateProviderQuota()` to surface richer failure-mode detail, addressing codex review finding F3 (operator cannot distinguish failure modes from the `unavailable` catch-all). Authority: ADR 0013 Rule 6 + codex review findings F1–F3.
|
||||
|
||||
#### 1. Extended `ProviderQuotaEntry` shape
|
||||
|
||||
```js
|
||||
{
|
||||
provider: string,
|
||||
// v0.5.1: 'unreachable' added (probe enabled, creds present, but no cache + probe failed)
|
||||
status: 'live' | 'stale' | 'unreachable' | 'unavailable',
|
||||
reason?: string, // only when status === 'unavailable' (no API or disabled)
|
||||
schema_version: string|null,
|
||||
last_fresh_at: number|null,
|
||||
utilization: { '5h': number|null, '7d': number|null } | null,
|
||||
reset: { '5h': number|null, '7d': number|null, overall: number|null, overage: number|null } | null,
|
||||
representative_claim: string|null,
|
||||
fallback_percentage: number|null,
|
||||
overage: { status: string|null, disabled_reason: string|null } | null,
|
||||
raw_available: boolean,
|
||||
// v0.5.1 (F3 — ADR 0013 Rule 6):
|
||||
failure: { kind, message, backoff_until? } | null,
|
||||
failure_kind: 'no_credentials'|'auth_failed'|'rate_limited'|'schema_drift'|'network'|'other' | null,
|
||||
// Note: 'opt_in_off' is NOT in this enum — when probe is opted out, the row's status
|
||||
// is 'unavailable' (not 'unreachable'); failure_kind stays null. Distinguishing
|
||||
// "user opted out" from "provider has no API" requires reading config separately.
|
||||
backoff_until: number | null, // epoch-ms when next probe attempt is allowed
|
||||
}
|
||||
```
|
||||
|
||||
Status semantics:
|
||||
- `'unavailable'` — probe disabled (`quota_probe_enabled: false`) OR provider has no public quota API (codex, mistral). `failure`, `failure_kind`, `backoff_until` are null.
|
||||
- `'live'` — probe succeeded within TTL. `failure` is null.
|
||||
- `'stale'` — probe failed but stale cache exists. `failure.kind` describes why the last probe failed. `last_fresh_at` is the epoch of the last successful probe. `backoff_until` tells when the next attempt is scheduled.
|
||||
- `'unreachable'` (new) — probe enabled + creds present (or missing!) but no cache available + probe failed. `failure.kind` distinguishes: `no_credentials`, `auth_failed`, `rate_limited`, `schema_drift`, `network`, `other`. `utilization` and `reset` are null (no data).
|
||||
|
||||
#### 2. `quotaStatus()` v0.5.1 return contract
|
||||
|
||||
`null` is now RESERVED for `quota_probe_enabled: false` only. All other failure paths return a structured shape:
|
||||
|
||||
```js
|
||||
null // ONLY: opt-in off
|
||||
{ probe_status: 'live', ... } // cache fresh
|
||||
{ probe_status: 'stale', ..., failure: { kind, message, backoff_until } } // cache stale + backoff
|
||||
{ probe_status: 'unreachable', source, schemaVersion, failure: { ... } } // no cache + failed
|
||||
```
|
||||
|
||||
The `stale: boolean` field is retained for backwards-compat (`stale: false` on live, `stale: true` on stale). New code should use `probe_status`.
|
||||
|
||||
#### 3. `dashboard.html` unreachable rendering
|
||||
|
||||
A new CSS class `.provider-row.unreachable` (red border + light red background) and `.unreachable-reason` text style handle the new status. `failure.message` and `failure_kind` are surfaced as a short text line under the provider badge. `failure.backoff_until` renders a "backoff active: Xs remaining" note if within window.
|
||||
|
||||
#### 4. Authority
|
||||
|
||||
- ADR 0013 Rule 6 (failure transparency mandate)
|
||||
- Codex review findings F1 (doctor bypass), F2 (200+empty-headers → schema_drift), F3 (failure-mode collapse)
|
||||
- v0.5.1 hotfix PR
|
||||
|
||||
---
|
||||
|
||||
### Amendment 1 — 2026-05-26: D81 Phase 5 quota_v2 shape + aggregateProviderQuota()
|
||||
|
||||
**Scope:** D81 (Phase 5 / ADR 0012 D81) extends the audit-query layer and dashboard-data endpoint to surface the new per-provider quota shape introduced by D80 (`lib/providers/anthropic.mjs:quotaStatus()`). This amendment documents the three new interfaces.
|
||||
|
||||
#### 1. `models-registry.json` — new `quota_probe` top-level key
|
||||
|
||||
D81 adds a `quota_probe` key at the root of `models-registry.json` per ADR 0013 Rule 5 (schema_version in registry so downstream consumers can detect schema drift):
|
||||
|
||||
```json
|
||||
{
|
||||
"quota_probe": {
|
||||
"schema_version": "2026-05-26",
|
||||
"anthropic": {
|
||||
"source": "anthropic-ratelimit-unified-headers",
|
||||
"endpoint": "https://api.anthropic.com/v1/messages",
|
||||
"fields_pinned": [ ...13 field names... ]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`fields_pinned` is load-bearing: if Anthropic adds/renames a header in a future CLI version, dashboard consumers comparing field-presence against this list can flag "schema drift detected" per the ADR 0013 Rule 5 drift-detection runbook. This field must be updated alongside the parser whenever a drift event occurs.
|
||||
|
||||
`lib/providers/anthropic.mjs` reads `quota_probe.schema_version` from the registry at call time (via `_resolveSchemaVersion()`) with the module-level `QUOTA_SCHEMA_VERSION` constant as fallback. No hard dependency on the registry — the constant is the safety net.
|
||||
|
||||
#### 2. `lib/audit-query.mjs` — new `aggregateProviderQuota()` export
|
||||
|
||||
```js
|
||||
export async function aggregateProviderQuota({
|
||||
providers, // Map<name, plugin> or plain object
|
||||
getQuotaStatus, // optional injectable getter (name) => Promise<shape|null>
|
||||
}): Promise<Array<ProviderQuotaEntry>>
|
||||
```
|
||||
|
||||
For each provider, calls `quotaStatus()` (already cached at the plugin layer per ADR 0013 Rule 3) and normalizes to the `ProviderQuotaEntry` shape:
|
||||
|
||||
```js
|
||||
{
|
||||
provider: string,
|
||||
status: 'live' | 'stale' | 'unavailable',
|
||||
reason?: string, // only when status === 'unavailable'
|
||||
schema_version: string|null,
|
||||
last_fresh_at: number|null, // epoch-ms of last successful probe
|
||||
utilization: { '5h': number|null, '7d': number|null } | null,
|
||||
reset: {
|
||||
'5h': number|null, '7d': number|null,
|
||||
overall: number|null, overage: number|null,
|
||||
} | null,
|
||||
representative_claim: string|null,
|
||||
fallback_percentage: number|null,
|
||||
overage: { status: string|null, disabled_reason: string|null } | null,
|
||||
raw_available: boolean,
|
||||
}
|
||||
```
|
||||
|
||||
Providers returning `null` from `quotaStatus()` (codex, mistral — no public quota API; or probe disabled) produce `{ status: 'unavailable', reason: 'no public quota api or probe disabled', ...null fields }`.
|
||||
|
||||
Providers whose `quotaStatus()` throws produce `{ status: 'unavailable', reason: <error.message>, ...null fields }`.
|
||||
|
||||
This function does NOT scan ndjson files; it calls live provider plugins. It is audit-query-adjacent (normalized query shape for the dashboard layer) but not audit-derived. Query model remains Lane 2 = A (in-memory, no SQLite).
|
||||
|
||||
#### 3. `/v0/management/dashboard-data` and `/v0/management/quota` — new `quota_v2` field
|
||||
|
||||
Both endpoints now return TWO quota keys:
|
||||
|
||||
- **`quota`** (legacy, unchanged): `Array<{ provider, ...rawQuotaStatus, available }>`. Kept for backwards compatibility with the existing `dashboard.html` (D82 will switch consumers to `quota_v2`).
|
||||
- **`quota_v2`** (D81 new): `Array<ProviderQuotaEntry>` — the normalized shape from `aggregateProviderQuota()` above. This is what D82's enriched dashboard UI will consume.
|
||||
|
||||
Both fields are computed from the same underlying `quotaStatus()` call. The legacy `quota` key calls `quotaStatus()` independently from `quota_v2`; since the probe is cached at the plugin layer (ADR 0013 Rule 3), the double call incurs no extra API requests.
|
||||
|
||||
**Deprecation timeline:** the legacy `quota` key is deprecated as of D81. Target removal: v1.0.0 or when D82 completes the dashboard migration (whichever comes first). Removal requires a separate PR with a CHANGELOG entry.
|
||||
|
||||
#### 4. Failure handling
|
||||
|
||||
`aggregateProviderQuota()` never throws to the dashboard endpoint. Per-provider failures are absorbed as `{ status: 'unavailable', reason: <error> }` entries. If `aggregateProviderQuota()` itself throws (implementation bug), `handleManagementDashboardData` and `handleManagementQuota` catch the error, log `dashboard_data_quota_v2_failed` / `management_quota_v2_failed`, and return `quota_v2: []` so the rest of the payload is unaffected.
|
||||
|
||||
#### 5. Authority citations for this amendment
|
||||
|
||||
- **ADR 0012 D81** — the D-day this amendment documents.
|
||||
- **ADR 0013 Rule 5** — mandate for `quota_probe.schema_version` in `models-registry.json`.
|
||||
- **D80 PR #52 commit 82d2e1c** — the producer of the `quotaStatus()` shape this amendment normalizes.
|
||||
- **ADR 0008 Lane 2 = A** — query model unchanged; `aggregateProviderQuota()` does not scan ndjson.
|
||||
|
||||
---
|
||||
- **Authors:** project maintainer (with AI drafting assistance)
|
||||
- **Related:**
|
||||
- OLP v0.1 spec § 4.6 (Dashboard requirements — port from OCP with multi-provider support) and § 4.7 (observability endpoints)
|
||||
@@ -195,7 +343,7 @@ The dashboard sets a 30s `setInterval` that calls `fetch('/v0/management/dashboa
|
||||
|
||||
### 6.6 Localhost-bound by default
|
||||
|
||||
The dashboard is served from the existing OLP HTTP port (default 3456) which is already bound to `127.0.0.1` per `server.mjs` startup (`server.listen(PORT, '127.0.0.1', ...)`). No additional binding logic. Remote operators access via SSH tunnel; ADR 0007 § 7 owner-only auth provides the per-request gate.
|
||||
The dashboard is served from the existing OLP HTTP port (default 4567 since v0.4.0 / D60; 3456 pre-v0.4.0) which is already bound to `127.0.0.1` per `server.mjs` startup (`server.listen(PORT, '127.0.0.1', ...)`). No additional binding logic. Remote operators access via SSH tunnel; ADR 0007 § 7 owner-only auth provides the per-request gate.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# ADR 0009 — Anthropic Interactive-Mode Path (Placeholder)
|
||||
|
||||
- **Date:** 2026-05-25
|
||||
- **Status:** Draft (Placeholder — blocked on OCP ADR 0007 P0 experiment outcome; no implementation D-day scheduled until P0 lands)
|
||||
- **Date:** 2026-05-25 (Placeholder); 2026-05-27 Amendment 1 (Accepted)
|
||||
- **Status:** **Accepted** (post-Amendment 1 — OLP self-spike supersedes OCP-wait; implementation D-day scheduled this Phase 6)
|
||||
- **Authors:** project maintainer (with AI advisory drafting)
|
||||
- **Related:**
|
||||
- **OCP ADR 0007** (Interactive-Mode Execution Pool, stream-json) — at `~/ocp/docs/adr/0007-interactive-mode-pool.md` on the maintainer's workstation. Pin reference at the time of this writing: OCP ADR 0007 is Draft status pending the same P0 outcome.
|
||||
@@ -193,5 +193,117 @@ If OCP P0 fails, **this ADR is shelved** and Phase 4 ordering is unchanged.
|
||||
## Status transitions (recorded for clarity)
|
||||
|
||||
- 2026-05-25 — Created as Draft (Placeholder). OCP ADR 0007 also Draft.
|
||||
- _(future)_ — If OCP ADR 0007 → Accepted with a confirmed transport: this ADR moves to "Pending Phase 4 implementation D-day", maintainer decides Option 1 / 2 / 3 + lane.
|
||||
- _(future)_ — If OCP ADR 0007 → Rejected: this ADR moves to "Shelved (upstream P0 failure)" with a note explaining the fallback (multi-provider routing already covers).
|
||||
- 2026-05-27 — Amendment 1 promotes to **Accepted**. OLP self-spike + empirical Transport-A confirmation on `claude` CLI v2.1.104 superseded wait-for-OCP. OCP is now in maintenance mode (per maintainer statement 2026-05-27 session) — OLP leads. Implementation lane: **Option 1 (parallel implementation, no warm pool, no PTY)**, scope reduced from "10-day warm pool with billing router" to "2-3-day stateless stream-json adapter".
|
||||
|
||||
---
|
||||
|
||||
## Amendment 1 — 2026-05-27: Self-spike supersedes wait-for-OCP; lock Option 1 with stream-json-no-`-p` transport
|
||||
|
||||
### Trigger
|
||||
|
||||
Two findings on 2026-05-27 changed the placeholder's premises:
|
||||
|
||||
1. **OCP is no longer the lead project.** The maintainer stated in the 2026-05-27 session: "OCP 不会大改动…主力方向放到 OLP." The "wait-and-port" strategy implicitly assumed OCP would do the P0 first. With OCP in maintenance mode, OLP cannot wait — the 2026-06-15 Anthropic billing split is 19 days out from this amendment.
|
||||
|
||||
2. **OLP self-spike confirmed Transport A (`--output-format stream-json --verbose` without `-p`) emits NDJSON on Claude Code v2.1.104.** The placeholder ADR § 1.3 cited OCP's "v2.1.150 only" observation as the binding caveat. Local empirical re-test on 2026-05-27 against PI231's deployed `claude` v2.1.104 produced the full NDJSON event stream (system/init + stream_event token deltas + message_stop + result + rate_limit_event) for invocations **without** `-p`. The `claude --help` text saying "(only works with --print)" is misleading — the flags accept invocation without `-p` and produce the documented NDJSON shape.
|
||||
|
||||
### Additional spike findings (2026-05-27 billing classification)
|
||||
|
||||
A separate web/GitHub research spike on the 2026-06-15 billing classification returned:
|
||||
|
||||
- Anthropic's published policy is **intent-based, not mechanism-based**. The Agent SDK credit pool covers: Agent SDK Python/TypeScript packages, `claude -p`, GitHub Actions, **and "third-party apps that authenticate with your Claude subscription through the Agent SDK"**. Subscription pool covers "Claude Code in the terminal or your IDE in interactive mode."
|
||||
- The third-party-app clause is the load-bearing ambiguity. OLP qualifies as a third-party app regardless of which CLI mode it spawns. If Anthropic tightens that clause from "via Agent SDK" to "any third-party app," OLP is caught regardless of `-p` flag presence.
|
||||
- Behavioral fingerprinting (request cadence, OAuth-scope patterns, isTTY absence) is a separate detection vector Anthropic could deploy without policy-text changes.
|
||||
|
||||
The spike's recommendation: "viable bridge for ~30-60 days post-2026-06-15, NOT durable solution."
|
||||
|
||||
### Value re-anchoring
|
||||
|
||||
The placeholder framed interactive-mode as "the durable answer to keep OLP anthropic subscription value past 2026-06-15." The 2026-05-27 spike re-anchors the value:
|
||||
|
||||
| Value | Placeholder framing | 2026-05-27 framing |
|
||||
|---|---|---|
|
||||
| Keep subscription pool 6.15+ | **Primary value** | **Uncertain bridge** (30-60 day plausibility) |
|
||||
| Hallucination fix (env-block / cwd injection) | (Not addressed) | **Primary value** — empirically proven |
|
||||
| Cost reduction (drop default tool descriptions) | (Not addressed) | **Primary value** — ~30% input token / ~64% per-request cost reduction measured against `--system-prompt` override |
|
||||
| Observability (rate_limit / cache / usage per request) | (Not addressed) | **Primary value** — NDJSON events expose data the current `--output-format text` path discards |
|
||||
| Protocol foundation for future tool-call passthrough | (Not addressed) | **Secondary value** — same NDJSON parser is reusable for Phase 8+ tool passthrough work |
|
||||
|
||||
**Net**: even if Anthropic immediately reclassifies third-party apps to Agent SDK pool on 2026-06-15 — making the bridge worthless — the implementation still earns its keep through the other four values.
|
||||
|
||||
### Locked decision
|
||||
|
||||
**Option 1 — Parallel implementation in OLP's `lib/providers/anthropic.mjs`.**
|
||||
|
||||
Lane: stream-json output, no `-p` flag (Transport A confirmed), stateless per-request spawn (no warm pool, no PTY, no node-pty dependency).
|
||||
|
||||
Rejected lanes and why:
|
||||
|
||||
- **Option 2 (chain OCP)** — OCP is in maintenance mode; coupling OLP's anthropic provider to OCP's HTTP shim is the wrong direction.
|
||||
- **Option 3 (both)** — premature complexity; pick the simple lane first.
|
||||
- **Warm-process pool** — OLP is stateless per AGENTS.md § "No conversation state". Pool lifecycle, crash backoff, and permission auto-response from OCP ADR 0007 § 4 are unnecessary for OLP's per-request model.
|
||||
- **PTY (Transport B with node-pty)** — Transport A worked; engines-bump for a native addon is unjustified when the simpler transport produces the documented NDJSON.
|
||||
|
||||
### Implementation scope (Option 1, this Phase 6)
|
||||
|
||||
| Change | Surface | Authority |
|
||||
|---|---|---|
|
||||
| `buildCliArgs(model)` drop `-p` and `--output-format text`; add `--output-format stream-json`, `--verbose`, `--no-session-persistence`, `--model` | `lib/providers/anthropic.mjs` | `claude --help` (v2.1.104) § `--output-format` / § `--verbose` |
|
||||
| `buildCliArgs(model, systemPrompt)` accepts optional system prompt; spawns with `--system-prompt "<OLP wrapper text>"` | `lib/providers/anthropic.mjs` | `claude --help` (v2.1.104) § `--system-prompt` |
|
||||
| OLP-managed system prompt construction (extract client `role:system` IR messages, prepend OLP wrapper saying "you are accessed via HTTP proxy; no local env/fs/shell access; respond directly") | `lib/providers/anthropic.mjs` `irToAnthropic` | This ADR § "OLP system prompt wrapper" below |
|
||||
| New `anthropicStreamJsonChunkToIR` parser replacing/supplementing `anthropicChunkToIR` — handles NDJSON event types `system/init`, `stream_event/content_block_delta`, `assistant`, `result`, `rate_limit_event` | `lib/providers/anthropic.mjs` | This ADR § "NDJSON event handling" below |
|
||||
| New tests verifying NDJSON parsing, system-prompt construction, env-block absence | `test-features.mjs` | (test surface; no external authority) |
|
||||
| README troubleshooting / supported-providers § note about the bridge nature | `README.md` | (docs surface) |
|
||||
|
||||
The `irToAnthropic` text serialization path is preserved for client messages (`role: user`, `role: assistant`); the `role: system` extraction goes to `--system-prompt`.
|
||||
|
||||
### OLP system prompt wrapper
|
||||
|
||||
The wrapper text injected via `--system-prompt`:
|
||||
|
||||
```
|
||||
You are accessed via the OLP HTTP proxy. You do NOT have access to any local
|
||||
filesystem, working directory, shell, git status, or machine environment.
|
||||
Do not infer or invent such information from any context you observe.
|
||||
Respond only based on the conversation provided.
|
||||
```
|
||||
|
||||
If the client IR request contains `role: system` messages, their concatenated `content` is appended after a blank line.
|
||||
|
||||
### NDJSON event handling
|
||||
|
||||
The parser must yield IR chunks based on the event stream:
|
||||
|
||||
| NDJSON event | IR yield | Notes |
|
||||
|---|---|---|
|
||||
| `{type:"system", subtype:"init"}` | None (consumed for session_id tracking) | First event always; ignore |
|
||||
| `{type:"stream_event", event:{type:"content_block_delta", delta:{type:"text_delta", text:"..."}}}` | `{type: "delta", content: "<text>"}` | Token-by-token streaming |
|
||||
| `{type:"assistant"}` | None (already captured by per-token deltas) | Aggregate message; ignore (or use for verify, optional) |
|
||||
| `{type:"result", subtype:"success"}` | `{type:"stop", finish_reason:"stop"}` | Marks end |
|
||||
| `{type:"rate_limit_event"}` | None (consumed for audit/dashboard) | Forward to OLP audit/observability layer later (Phase 6+ enhancement) |
|
||||
| `{type:"control_request"}` | Log + ignore | Per Anthropic stream-json docs |
|
||||
|
||||
The cache key composition (ADR 0005) is unchanged — same IR request hash; the on-the-wire format change is internal to the anthropic plugin.
|
||||
|
||||
### Token cost measurement (binding evidence)
|
||||
|
||||
Two requests against PI231 v2.1.104 on 2026-05-27 with identical user prompt `"reply: OK"`, model `claude-sonnet-4-6`:
|
||||
|
||||
- Default invocation (no `--system-prompt`): `cache_creation_input_tokens=4785`, `cache_read_input_tokens=11816`, total input ≈ 16,601 tokens, `total_cost_usd=$0.0216`.
|
||||
- With `--system-prompt "You are a chat assistant. Respond directly."`: `cache_creation_input_tokens=1306`, `cache_read_input_tokens=9394`, total input ≈ 10,700 tokens, `total_cost_usd=$0.0078`.
|
||||
|
||||
**Net**: ~30% input token reduction, ~64% per-request cost reduction. Replicable by anyone with `claude` v2.1.104 + OAuth on a similar setup.
|
||||
|
||||
### Caveats binding the implementation
|
||||
|
||||
1. **Bridge value uncertain.** The 30-60 day estimate is a spike judgment, not Anthropic-confirmed. Implementation must continue to function correctly if Anthropic re-routes this path to Agent SDK billing on 2026-06-15 — the only consequence is the bridge value disappears, but the other four values (hallucination / cost / observability / protocol foundation) remain.
|
||||
2. **No claim about durability.** This ADR amends only as far as "the bridge is worth the 2-3 day investment given the orthogonal values." A future ADR (likely Phase 7 sandbox-runtime + Phase 8 multi-provider robustness) will revisit the anthropic provider's strategic role once the post-2026-06-15 picture clarifies.
|
||||
3. **Sandbox-runtime still required for real multi-tenant deployment.** Per the 2026-05-27 session prior-art search, Anthropic's official multi-tenant answer is `@anthropic-ai/sandbox-runtime` (OS-level isolation). This ADR does NOT substitute for that work; sandbox-runtime remains Phase 7 scope and is a hard prerequisite before any cloud deployment per `docs/plans/cloud-deployment-family.md`.
|
||||
4. **CLI version pin guidance.** Stream-json without `-p` was confirmed on v2.1.104. Future versions may tighten this; the plugin's spawn should emit a warning to OLP server log if `claude --version` falls outside a `v2.1.100`–`v2.1.149` range. Hard failure on out-of-range version is NOT required; warning is sufficient for v0.6.x.
|
||||
|
||||
### Updated authority citations (in addition to placeholder § Authority citations)
|
||||
|
||||
- **OLP self-spike — 2026-05-27 session live transcripts** (PI231 ssh; `claude -p --output-format stream-json --verbose` and `claude` no-`-p` variants captured in session log; retained in cc-mem post-implementation).
|
||||
- **P0 billing classification spike — 2026-05-27** subagent transcript; sources include Anthropic published docs at `code.claude.com/docs/en/headless`, `support.claude.com/en/articles/15036540`, `support.claude.com/en/articles/11145838`.
|
||||
- **claude CLI v2.1.104 `--help`** (live capture on PI231) § `--output-format`, § `--verbose`, § `--system-prompt`, § `--no-session-persistence`.
|
||||
- **CLAUDE.md `release_kit.phase_rolling_mode.current_phase`** — Phase 6; this ADR consumes a Phase 6 D-day per the amendment, NOT a Phase 4 D-day (the placeholder's hypothetical scheduling).
|
||||
|
||||
@@ -0,0 +1,162 @@
|
||||
# ADR 0010 — Phase 4 Charter: Operator + Client UX
|
||||
|
||||
**Status:** Accepted (Phase 4 open as of 2026-05-26)
|
||||
**Date:** 2026-05-26
|
||||
**D-day:** D60 (charter + default port change)
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
Phases 1 — 3 shipped OLP's structural core: HTTP entry surface, IR, provider plugins, fallback engine, content-addressed cache (including streaming-path singleflight at D57+D58 → v0.3.2), multi-key auth + audit ndjson + daily rotation, owner-only management endpoints + dashboard. v1.x roadmap items #1 / #2 / #4 / #7 are closed. Items #3 / #5 / #6 remain trigger-gated.
|
||||
|
||||
Two complementary brainstorm passes (2026-05-26) — a comprehensive OCP feature audit + a multi-provider proxy / IDE integration prior-art survey — converged on a clear gap: **OLP's operator and client surfaces are 0% inherited from OCP**. Today OLP has `bin/olp-keys` and `bin/olp-audit-rotate` as the entire operator CLI, no `olp doctor` / no `olp-connect`, no Telegram/Discord integration, no SSE heartbeat for long-running streams behind reverse proxies. Family members get OLP API keys via out-of-band paste, point their IDEs at OLP via the README's one-line example, and discover failure modes via curl. OCP's UX worked because of a load-bearing combination: README `paste-this-prompt-to-Claude-Code` instructions + machine-readable `ocp doctor next_action.ai_executable[]` + `ocp-connect` zero-config LAN setup + `/health.anonymousKey` self-advertising token + `/ocp` Telegram slash commands. **Phase 4 brings these forward as OLP-native primitives.**
|
||||
|
||||
A separate strategic decision — should OLP add `/v1/messages` (Anthropic-shape entry surface) for Claude Code support — was considered and **rejected for Phase 4** (see § "Out of Phase 4 scope" below). The decision is recorded with an explicit re-open trigger.
|
||||
|
||||
---
|
||||
|
||||
## Decision
|
||||
|
||||
Phase 4 scope is **Operator + Client UX**. The phase opens 2026-05-26 with D60 (this charter + default port change). Phase 4 close ships v0.4.0; per `CLAUDE.md release_kit.phase_rolling_mode`, the close PR is maintainer-triggered.
|
||||
|
||||
### In scope — Phase 4 D-day plan (~13 D-days)
|
||||
|
||||
| D-day | Deliverable | Authority | Estimate |
|
||||
|---|---|---|---|
|
||||
| **D60** | Default port `3456 → 4567` + this ADR 0010 charter + README / CHANGELOG / ADR 0001 + ADR 0008 amendments | This charter | 0.5d |
|
||||
| **D61 — D63** | SSE heartbeat (opt-in via `streaming.heartbeat_interval_ms` config; eager-headers-post-spawn; `X-Accel-Buffering: no` constant) + `recentErrors[20]` ring buffer + `/status` combined endpoint | Port OCP `server.mjs:660-685` + `301-358` + `1151-1188`; OCP `docs/superpowers/specs/2026-04-25-47-sse-heartbeat-design.md` | 2.5d |
|
||||
| **D64 — D67** | `olp` Node-based CLI scaffold (subcommands `status / health / usage / models / logs / cache / providers / chain show / restart / doctor`) + `olp doctor` machine-readable `next_action.ai_executable[]` framework + one fix-template per shipped provider plugin | Port OCP `ocp` bash wrapper (translated to Node — bash dep on python3 is a known fragile point) + OCP `scripts/doctor.mjs` framework | 4d |
|
||||
| **D68 — D70** | `olp-connect <ip>` client-side IDE auto-config (Cline / Continue.dev / Cursor / Aider / Claude Code / OpenClaw detection) + `/health.anonymousKey` field (opt-in via `auth.advertise_anonymous_key` config; default off) + ADR 0011 (anonymous-key deployment-context limits — trusted-LAN-only invariant explicit) | Port OCP `ocp-connect` + `server.mjs:1454,1488` | 3d |
|
||||
| **D71 — D73** | `olp-plugin/` (OpenClaw gateway plugin for `/olp` Telegram/Discord slash commands; subcommand parity with `olp` CLI minus mutations) + `docs/integrations/{continue.md,cline.md,cursor.md,aider.md,claude-code.md,openclaw.md}` IDE setup docs | Port OCP `ocp-plugin/index.js`; cross-ref Prior-Art § 3 + § 4 | 3d |
|
||||
| **close** | v0.4.0 release PR — `package.json` bump, CHANGELOG promotion, `release_kit.phase_rolling_mode` advance to Phase 5 pre-release identifier | `CLAUDE.md release_kit overlay` | maintainer-triggered |
|
||||
|
||||
### Out of Phase 4 scope (with explicit triggers)
|
||||
|
||||
#### `/v1/messages` — Anthropic-shape entry surface
|
||||
|
||||
**Status:** Deferred. Re-enable strictly gated on ADR 0009 P0 success.
|
||||
|
||||
**Value matrix (decisive):**
|
||||
|
||||
| Scenario | Without `/v1/messages` | With `/v1/messages` |
|
||||
|---|---|---|
|
||||
| Maintainer's own Claude Code usage | Direct via Anthropic OAuth → subscription (today) or Agent SDK pool (post-2026-06-15) | Same — maintainer never routes own CC through OLP per stated workflow |
|
||||
| Family member wanting CC access | Not supported (OAuth is full-account; OLP CLI tokens are scoped) | CC via `ANTHROPIC_BASE_URL=http://olp:4567` + `olp_*` token |
|
||||
| **P0 succeeds** (ADR 0009 interactive-mode bills as subscription) | OpenAI-shape IDE clients (Cline/Continue/Cursor) all benefit automatically via OLP's anthropic plugin | CC users additionally benefit; both subscription-billed |
|
||||
| **P0 fails** (interactive-mode bills as Agent SDK same as `-p`) | OpenAI-shape clients still work; no billing change | CC users get same billing as direct OAuth; **fallback to codex/mistral degrades Anthropic-specific features (tool_use schema mismatch / cache_control drop / computer_use no-op / thinking-block drop)** more severely than OpenAI-shape clients which speak the multi-provider lingua franca |
|
||||
|
||||
**Rationale.** Under P0 failure, `/v1/messages` provides no billing benefit AND degrades worse on fallback than OpenAI-shape clients (because OpenAI tool schema is the cross-provider standard). The security benefit (no OAuth exposure) is equally achievable via Cline/Continue/Cursor. **Net non-positive under P0 failure.**
|
||||
|
||||
**Re-open condition.** (a) ADR 0009 P0 confirms interactive-mode billing classification as subscription (≥ 2026-07-15) AND (b) maintainer explicitly opens Phase 5 "Anthropic-shape hub" scope with the name of at least one family member who wants CC access. If only (a) fires without (b), `/v1/messages` is reconsidered at the start of whichever phase covers it but is not auto-opened.
|
||||
|
||||
**README posture (Phase 4).** README § Supported Clients explicitly lists OpenAI-compatible clients (Cline, Continue.dev, Cursor, Aider, OpenClaw bots). Claude Code is listed as **Not supported as an OLP client**, with the explicit alternative "Cline + OLP" (same fallback chain available, better cross-provider compatibility). README links to this ADR for the reasoning.
|
||||
|
||||
#### Other deferred items
|
||||
|
||||
- **v1.x roadmap #3 (soft trigger reactivation)**, **#5 (provider `cacheKeyFields` mask)**, **#6 (streaming SPAWN_FAILED salvage)** — trigger conditions per `docs/v1x-roadmap.md` have not fired. Not in Phase 4.
|
||||
- **Anthropic / codex billing audits** — date-gated (`anthropic.mjs:53, 416, 441` say 2026-06-16; `codex.mjs:572` post-D7 E2E audit). Not in Phase 4.
|
||||
- **`context_window_exceeded` fallback trigger** (LiteLLM prior-art) — small ADR amendment + trigger taxonomy add; opportunistically in Phase 5 unless trigger fires sooner.
|
||||
- **`X-OLP-Cost-USD` per-request response header** — depends on provider-cost weights table (Phase 5 prerequisite).
|
||||
- **per-(provider, model) live stats Map** (replacing audit-query scan for dashboard 30s poll) — current scan latency adequate; Phase 5+.
|
||||
- **OpenTelemetry GenAI span emission** — `npm` dep + ~150 LOC; family-scale ROI marginal. Phase 6+ unless Langfuse self-host requested.
|
||||
- **Intent-based routing**, **stackable transformer plugin model** — explicit non-goals per Prior-Art § 8 anti-patterns.
|
||||
|
||||
### Opportunistic Phase 4 micro-additions (not blocking)
|
||||
|
||||
Items small enough to land alongside a planned D-day without scope creep, if encountered:
|
||||
|
||||
- Env-var deny-list before provider plugin `spawn` (per OCP `server.mjs:531-534`; each plugin declares its own list)
|
||||
- 5 MB request body cap with HTTP 413 (per OCP `server.mjs:1270,1278-1281`)
|
||||
- Error-response path-sanitization (per OCP `server.mjs:1395`)
|
||||
- Stable node-path resolution in launchd plist (Homebrew `/Cellar/<ver>/` → `/opt/` rewrite; per OCP `setup.mjs:344-351`)
|
||||
- Legacy model alias resolution in `models-registry.json` (`aliases:` field; per OCP `legacyAliases`)
|
||||
|
||||
### Exit gate — v0.4.0 close criteria
|
||||
|
||||
1. D60 — D73 all merged with fresh-context opus reviewer APPROVE per Iron Rule 10.
|
||||
2. CI green on every D-day merge commit and on the v0.4.0 release commit head.
|
||||
3. README § Operator CLI + § IDE Setup + § Telegram/Discord Usage sections present.
|
||||
4. ADR 0010 (this charter) + ADR 0011 (anonymous-key deployment-context limits) on disk.
|
||||
5. `CHANGELOG.md "Unreleased"` promoted to `"## v0.4.0 — <date>"` with D60 — D73 entries.
|
||||
6. `package.json` bumped to `0.4.0`.
|
||||
7. `CLAUDE.md release_kit.phase_rolling_mode.current_phase` advances `Phase 4 → Phase 5`; `current_pre_release_identifier` advances `0.4.0-phase4 → 0.5.0-phase5`.
|
||||
8. Standing autopilot grant covers D-day-by-D-day execution; v0.4.0 close PR is maintainer-triggered.
|
||||
|
||||
---
|
||||
|
||||
## Default port change (D60 specific)
|
||||
|
||||
The default `OLP_PORT` value moves `3456 → 4567` at this D-day. Rationale:
|
||||
|
||||
- OCP defaults to 3456 and the maintainer's existing OCP installs stay on 3456 indefinitely.
|
||||
- A standard `olp` install on the same host without overriding `OLP_PORT` collides at bind time.
|
||||
- Setting `OLP_PORT=4567` as the default makes co-host the recommended steady state during the migration window (and beyond — there is no enforced deprecation of OCP).
|
||||
- Existing OLP deployments wanting the pre-D60 default can set `OLP_PORT=3456` in the launchd plist / shell env.
|
||||
|
||||
**Tested invariants preserved by the port change:**
|
||||
|
||||
- All `test-features.mjs` suites use `port: 0` (ephemeral assigned port) — no test depends on the default value. Verified via `grep -nE '\\b3456\\b' test-features.mjs` returning empty.
|
||||
- All cache / fallback / provider plugin code is port-agnostic.
|
||||
- Dashboard 30s poll uses relative paths — no port change required in `dashboard.html`.
|
||||
- `/v0/management/*` endpoints use relative paths — no client-side update required.
|
||||
|
||||
**Files amended at D60:**
|
||||
|
||||
- `server.mjs:17` — env-var doc comment
|
||||
- `server.mjs:74` — default value
|
||||
- `README.md` quick start + Environment Variables table + Migration from OCP § note
|
||||
- `docs/adr/0001-project-founding.md` § "Decision" paragraph about port conflict (struck and amended)
|
||||
- `docs/adr/0008-dashboard-and-audit-query.md` § 6.6 port reference
|
||||
- `CHANGELOG.md` Unreleased entry
|
||||
- This ADR
|
||||
|
||||
---
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive.**
|
||||
|
||||
- Family member onboarding goes from "maintainer texts API key + edits IDE config" to `curl -fsSL .../olp-connect | bash -s -- <ip>`.
|
||||
- `paste-this-prompt-to-Claude-Code` self-installation pattern unlocks AI-driven setup / upgrade / repair, eliminating the maintainer's Tier-1 support role.
|
||||
- Long-reasoning streams behind nginx / Cloudflare / Tailscale Funnel no longer 502 at 60s idle.
|
||||
- `/olp` Telegram slash commands enable "is OLP up?" / "show usage" / "rotate key" from anywhere with chat access.
|
||||
- OCP and OLP co-host on the same workstation, lowering the maintainer's cost of running both.
|
||||
|
||||
**Negative.**
|
||||
|
||||
- Phase 4 is the first phase whose scope is primarily about *operator experience* rather than functional capability. The work doesn't unlock new requests OLP can serve; it makes OLP's existing capability survive contact with real users.
|
||||
- The `olp-connect` IDE auto-detect logic accumulates IDE-specific quirks (Cline base-URL UI regressions per their issue #7128; Cursor's malformed-request-when-OpenRouter behavior; etc.). Maintenance burden grows.
|
||||
- README size grows substantially with Operator CLI + IDE Setup + Telegram/Discord sections. Discoverability of the existing technical reference (ADRs, environment variables) may degrade unless the navigation is refactored.
|
||||
|
||||
**Neutral.**
|
||||
|
||||
- Phase 4 deliberately spends 0 D-days on `/v1/messages`. If ADR 0009 P0 succeeds in Q3 2026, Phase 5 "Anthropic-shape hub" becomes the natural next phase, with the prerequisite IR work that Phase 4 surfaces (every IDE doc page is a test of which IR fields actually flow through). If P0 fails, `/v1/messages` shelves indefinitely and the README simply documents CC as out-of-scope.
|
||||
|
||||
---
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
1. **Phase 4 = `/v1/messages` first, operator UX later.** Rejected. The brainstorm matrix demonstrated `/v1/messages` is value-positive only if ADR 0009 P0 succeeds, and operator UX gains accrue regardless. Building speculative infrastructure ahead of P0 risks 5-7 D-days of work shelving.
|
||||
2. **Phase 4 = operator + client UX + `/v1/messages` together (full kitchen sink).** Rejected. ~20 D-days lengthens the Phase 4 close window unnecessarily; the natural review chunks blur; maintainer review fatigue is real.
|
||||
3. **Phase 4 = just D60 + opportunistic SSE heartbeat, no CLI / no plugin / no docs bundle.** Rejected. Each of the operator-UX items individually has small ROI; the value compounds when they ship together (CLI surfaces data → `/status` exposes shape → Telegram plugin renders → IDE docs reference → `olp-connect` automates). Splitting them across phases loses the compounding.
|
||||
4. **Defer Phase 4 entirely; jump to Phase 5 Anthropic-shape hub when P0 lands.** Rejected. Operator UX is needed now (this session is itself evidence — the maintainer spent ~30 minutes confirming OCP feature inheritance because there's no `olp doctor` answer). Waiting for P0 stalls progress on independently-valuable work.
|
||||
|
||||
---
|
||||
|
||||
## Authority
|
||||
|
||||
- `docs/v1x-roadmap.md` — Phase 4 was named as the canonical destination for the post-cleanup batch since v0.3.0 close.
|
||||
- `CLAUDE.md release_kit.phase_rolling_mode` — `current_phase: Phase 4` already; this charter formalizes the contents.
|
||||
- OCP comprehensive feature audit (2026-05-26 subagent output, summarized in `~/.cc-rules/memory/auto/MEMORY.md` and in this session's transcript).
|
||||
- Multi-provider proxy / IDE integration prior-art survey (2026-05-26 subagent output).
|
||||
- ADR 0009 (Anthropic interactive-mode path placeholder) — establishes the gate for `/v1/messages` re-consideration.
|
||||
- ADR 0001 (project founding) — § "Decision" paragraph about port conflict, amended at this D-day.
|
||||
- ADR 0008 (dashboard + audit query) — § 6.6 default-port reference, amended at this D-day.
|
||||
- `~/.cc-rules/memory/auto/standing_autopilot_phase_2.md` — standing autopilot grant covering D-day-by-D-day execution; v0.4.0 close PR is maintainer-triggered per `release_kit.phase_close_trigger`.
|
||||
|
||||
---
|
||||
|
||||
## Procedural mechanism
|
||||
|
||||
CC 开发铁律 v1.6 § 5.5 (release-kit overlay drives Phase boundaries) + § 10 (independent reviewer on every implementation D-day) + § 11 (minimum reviewable unit per PR — this charter ships as D60 PR alongside the default port change because both are governance-class and small).
|
||||
@@ -0,0 +1,358 @@
|
||||
# ADR 0011 — Anonymous-Key Deployment-Context Limits (Trusted-LAN Invariant)
|
||||
|
||||
**Status:** Accepted (2026-05-26)
|
||||
**Date:** 2026-05-26
|
||||
**D-day:** D70 (lands alongside D68 `olp-connect` + D69 `/health.anonymousKey`)
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
ADR 0010 § Phase 4 D68-D70 charter scoped a three-deliverable bundle:
|
||||
|
||||
- **D68** `olp-connect <ip>` — client-side bash script that auto-configures a
|
||||
family member's machine to point at a remote OLP instance, including IDE
|
||||
detection (Cline / Continue.dev / Cursor / Aider / OpenClaw) and rc-file +
|
||||
system-level env var writes.
|
||||
- **D69** `/health.anonymousKey` field — opt-in (`auth.advertise_anonymous_key:
|
||||
true` in `~/.olp/config.json`; default `false`) surface that emits the
|
||||
plaintext of a designated guest-tier key so `olp-connect` can pick it up
|
||||
with zero-config (no out-of-band token paste).
|
||||
- **D70** this ADR — codifies the deployment-context limits that make D69
|
||||
safe.
|
||||
|
||||
The D69 mechanism is a deliberate port of OCP's clever `PROXY_ANONYMOUS_KEY` +
|
||||
`/health.anonymousKey` pattern (OCP `server.mjs:148, 1454, 1488, 1555`,
|
||||
shipped 2026-04 under OCP issue #12 § 14 Path A) which made family-member
|
||||
onboarding go from "maintainer texts API key + edits IDE config" to
|
||||
`curl -fsSL .../ocp-connect | bash -s -- <ip>`. The OCP pattern works on
|
||||
trusted family LAN deployments; it would catastrophically fail on a public
|
||||
internet deployment. ADR 0011 makes the trust assumption explicit before OLP
|
||||
inherits the pattern.
|
||||
|
||||
The ADR also pins three implementation details that are NOT obvious from
|
||||
reading the D69 patch alone:
|
||||
|
||||
1. The plaintext token must live on disk SOMEWHERE for the server to surface
|
||||
it. ADR 0007 § 5 + § 6.2 explicitly forbid plaintext storage in the
|
||||
default manifest. D69 introduces an explicit **opt-in** `plaintext_advertise`
|
||||
manifest field for ONLY the advertised key — every other key retains the
|
||||
ADR 0007 § 5 hash-only contract.
|
||||
2. The advertised key is **guest-tier**, not owner-tier. Owner-tier
|
||||
advertisement is rejected at keygen time AND at config-load time —
|
||||
exposing the owner identity unauthenticated would grant any LAN caller
|
||||
`/health` full payload, `/v0/management/*` mutating access, and
|
||||
`X-OLP-Fallback-Detail` visibility — the exact inverse of the advertise
|
||||
key's intent (a low-privilege zero-config tier).
|
||||
3. Three prerequisites MUST hold simultaneously for `/health.anonymousKey`
|
||||
to be emitted; missing any one is logged at startup but the server still
|
||||
boots — graceful-degrade rather than refuse-to-start.
|
||||
|
||||
---
|
||||
|
||||
## Decision
|
||||
|
||||
### Three-prerequisite gate (server-side)
|
||||
|
||||
`/health` emits the `anonymousKey` field if and only if ALL THREE hold:
|
||||
|
||||
1. `auth.advertise_anonymous_key === true` in `~/.olp/config.json` (default
|
||||
`false` — opt-in).
|
||||
2. `auth.allow_anonymous === true` (the anonymous tier must be reachable for
|
||||
the advertised key to be meaningful to zero-config callers; advertising a
|
||||
key into a deployment that rejects anonymous requests is incoherent).
|
||||
3. At least one active (`revoked_at === null`) manifest under `~/.olp/keys/`
|
||||
carries a non-empty `plaintext_advertise: "olp_..."` field.
|
||||
|
||||
When prerequisite (1) holds but (2) or (3) fails, the server logs a
|
||||
startup warn (`anonymous_key_advertised_but_denied` or
|
||||
`anonymous_key_advertised_but_no_anonymous_key_exists`) but starts normally
|
||||
and simply omits the field from `/health` responses.
|
||||
|
||||
### Plaintext storage mechanism (`plaintext_advertise` manifest field)
|
||||
|
||||
ADR 0007 § 5 forbids plaintext storage anywhere. D69 introduces a single,
|
||||
**explicitly opt-in** exception: the manifest of the designated advertised
|
||||
key gains a `plaintext_advertise` field whose value is the plaintext token.
|
||||
|
||||
This field is written ONLY when the operator runs:
|
||||
|
||||
```
|
||||
olp-keys keygen --anonymous --advertise
|
||||
```
|
||||
|
||||
(or `--advertise` alone on a guest-tier `keygen` invocation; `--anonymous`
|
||||
is a friendly shorthand for `--tier=guest --name=anonymous`). The keygen
|
||||
command surfaces an explicit `WARNING` to stderr at creation time:
|
||||
|
||||
```
|
||||
WARNING: this key's plaintext is now stored on disk + will be exposed via
|
||||
/health.anonymousKey when auth.advertise_anonymous_key=true AND
|
||||
auth.allow_anonymous=true. Use ONLY on a trusted LAN. See ADR 0011.
|
||||
```
|
||||
|
||||
Every other key (every existing key, and every newly-created key without
|
||||
`--advertise`) retains the ADR 0007 § 5 hash-only contract — `manifest.json`
|
||||
contains `token_hash` and NEVER `plaintext_advertise`.
|
||||
|
||||
**Schema-version note (D69 reviewer P2-2).** This adds a new optional field
|
||||
to the manifest. Per ADR 0007 § 4 ("Increment `schema_version` on any
|
||||
non-additive change"), additive optional fields do NOT require a
|
||||
`schema_version` bump — older parsers ignore unknown fields per the same
|
||||
section's forward-compat rule. The manifest stays at `schema_version: 1`.
|
||||
Documented here so a future archaeologist asking "why didn't D69 bump
|
||||
`schema_version`?" has a one-line answer.
|
||||
|
||||
**`listKeys()` redaction (D69 reviewer P2-1).** `lib/keys.mjs listKeys()`
|
||||
strips BOTH `token_hash` AND `plaintext_advertise` from its return value.
|
||||
Callers wanting the advertised plaintext for the `/health` publication
|
||||
path MUST go through `findAdvertisedKey()` — the only sanctioned read
|
||||
site. This protects against a future caller of `listKeys()` accidentally
|
||||
emitting the plaintext into logs / HTTP responses / dashboards.
|
||||
|
||||
### Tier restriction (guest only)
|
||||
|
||||
`createKey()` rejects `plaintext_advertise: true` for `owner_tier: 'owner'`
|
||||
with the error `createKey: plaintext_advertise requires owner_tier="guest"`.
|
||||
The CLI also rejects `--owner --advertise` with a clear error pointing at
|
||||
this ADR.
|
||||
|
||||
Rationale: owner-tier confers `/health` full payload visibility,
|
||||
`/v0/management/*` mutating access, and `X-OLP-Fallback-Detail` header
|
||||
visibility. Advertising owner-tier plaintext unauthenticated would let any
|
||||
LAN caller assume the owner identity — the exact opposite of the design
|
||||
intent.
|
||||
|
||||
### Trusted-LAN deployment invariant
|
||||
|
||||
`auth.advertise_anonymous_key: true` is permitted ONLY when the OLP server
|
||||
is bound to a trust-equivalent address space:
|
||||
|
||||
| Tier | Address space | Permitted? |
|
||||
|------|---------------|------------|
|
||||
| Loopback | `127.0.0.0/8` | yes |
|
||||
| RFC 1918 LAN | `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` | yes |
|
||||
| Tailnet | `100.64.0.0/10` (CGNAT range used by Tailscale) | yes |
|
||||
| Localhost domains | `localhost`, `*.local`, `*.internal` | yes |
|
||||
| Public internet | any routable IPv4/IPv6 outside the above | NO |
|
||||
|
||||
This is a **soft constraint** at v0.4.0 — OLP does not enforce IP-allowlist
|
||||
or BIND_ADDRESS inspection. The constraint is documented here, surfaced in
|
||||
README § "Anonymous-key advertise mode (trusted-LAN-only)", and warned-but-
|
||||
not-blocked at server startup when `auth.advertise_anonymous_key=true` and
|
||||
the bind address looks public.
|
||||
|
||||
Hard enforcement (refuse to start when bind is public + advertise enabled)
|
||||
is deferred. The maintainer's deployments are LAN-only, the family-scale
|
||||
audience cannot tolerate a startup-refuse mode that bricks the proxy on
|
||||
ambiguous network topology (e.g., TLS-fronted private network where the
|
||||
underlying bind IP IS public but the network itself is trusted), and the
|
||||
trade-off in ADR 0010 explicitly accepted operator-discretion gates for
|
||||
soft constraints of this class.
|
||||
|
||||
---
|
||||
|
||||
## Threat model
|
||||
|
||||
The advertised anonymous key is **public** within the boundary of "anyone
|
||||
who can reach `GET /health`." Anyone within that boundary can read the
|
||||
plaintext from `/health.anonymousKey` and use it for `/v1/chat/completions`,
|
||||
`/v1/models`, etc.
|
||||
|
||||
| Deployment | Boundary | Acceptable? |
|
||||
|------------|----------|-------------|
|
||||
| Mac mini + Tailscale, only family devices on tailnet | family devices | YES |
|
||||
| Home LAN with no guest WiFi, no port-forward | household + neighbors-within-WiFi-range | YES (within risk tolerance) |
|
||||
| Home LAN with guest WiFi joined to same VLAN as proxy | EVERYONE who visits and connects to guest WiFi | borderline; treat with caution |
|
||||
| Coffee shop / open WiFi | EVERYONE physically present | NO |
|
||||
| Public internet via Cloudflare Tunnel / port-forward / VPS | EVERYONE on the internet | NO — instant compromise |
|
||||
|
||||
The capability gain for an attacker who reads `/health.anonymousKey` is
|
||||
**equal to the capability the operator deliberately granted the
|
||||
anonymous-tier key**:
|
||||
|
||||
- `providers_enabled` (when `'*'`, the attacker can dispatch any provider —
|
||||
burning the operator's subscription quotas).
|
||||
- `/v1/chat/completions` access (LLM use under the operator's billing).
|
||||
- Cache pollution under `__anonymous__` namespace (per ADR 0007 § 7.1; the
|
||||
advertised guest key uses its own `<key-id>` namespace — but anonymous
|
||||
callers who DON'T present the key use `__anonymous__`).
|
||||
|
||||
What the attacker does NOT get:
|
||||
|
||||
- `/health` full payload — gated to `owner_tier === 'owner'` (ADR 0007 § 7.1).
|
||||
- `/v0/management/*` mutating endpoints — gated to owner (ADR 0008 § 7).
|
||||
- `X-OLP-Fallback-Detail` header — `'owner_only'` policy default (ADR 0007 § 7.2).
|
||||
- The owner key's plaintext (which is never stored anywhere; only its
|
||||
`token_hash` is on disk per ADR 0007 § 5).
|
||||
|
||||
The "burn the operator's subscription quotas" failure mode is bounded by
|
||||
the per-provider quota limits AND the maintainer's monitoring (`/health`
|
||||
owner-view shows quota status per provider; `/v0/management/audit` shows
|
||||
per-key usage). Detection is fast; the question is how much quota the
|
||||
attacker can burn between compromise and key revocation.
|
||||
|
||||
---
|
||||
|
||||
## `olp-connect` integration (D68 client-side)
|
||||
|
||||
`bin/olp-connect <ip>` queries `GET /health` as its first action. If the
|
||||
response contains `anonymousKey: "olp_..."`, the script uses that value
|
||||
silently for the rest of the run (printing a one-line `Using server-
|
||||
advertised anonymous key: olp_...XXXX` notice + a pointer to this ADR).
|
||||
This is what makes `olp-connect <ip>` a true zero-config command — no
|
||||
out-of-band token paste needed.
|
||||
|
||||
If `anonymousKey` is absent (the default, when `auth.advertise_anonymous_key`
|
||||
is false), the script falls back to interactive prompt or `--key` flag.
|
||||
|
||||
`olp-connect` does NOT perform any of the trusted-LAN soft-checks itself —
|
||||
it trusts that an operator who set `auth.advertise_anonymous_key: true`
|
||||
knows their deployment context. The script does, however, document the
|
||||
trade-off in its `--help` output and prints the ADR 0011 reference
|
||||
alongside the "using server-advertised key" notice.
|
||||
|
||||
---
|
||||
|
||||
## Deployment configurations (D76 amendment, 2026-05-26)
|
||||
|
||||
Original ADR 0011 referenced a `BIND_ADDRESS` concept that did not exist in the v0.4.0–v0.4.2 codebase — the server was hard-coded to `server.listen(PORT, '127.0.0.1', ...)`. D76 closes this gap by adding the `OLP_BIND` env var (default `127.0.0.1`), making the deployment-context discussion below operational rather than aspirational.
|
||||
|
||||
Three deployment configurations are supported:
|
||||
|
||||
| `OLP_BIND` value | Reachability | Anonymous-key publication |
|
||||
|---|---|---|
|
||||
| `127.0.0.1` (default) | Loopback only | Safe with any auth posture (no LAN exposure at all) |
|
||||
| RFC1918 IP / tailnet IP / `0.0.0.0` on a trusted LAN | LAN clients only | Safe when `advertise_anonymous_key: true` — the documented "trusted-LAN" zero-config family onboarding flow |
|
||||
| Public IP / `0.0.0.0` on a public-facing host | Public internet | **Incompatible with `advertise_anonymous_key: true`.** Operator MUST keep `advertise_anonymous_key: false` (default). |
|
||||
|
||||
The server emits a startup warn event `anonymous_key_advertised_with_lan_bind` when `OLP_BIND` is non-loopback AND `advertise_anonymous_key: true` (per the `lib/keys.mjs` + `server.mjs` checks). The warn is a **checkpoint, not a hard gate** — the server cannot tell from the bind address alone whether the operator is on a trusted LAN (RFC1918 / tailnet) or has accidentally exposed a public IP. The Re-evaluation trigger #1 below escalates to a hard gate when OLP gains a public-internet deployment mode.
|
||||
|
||||
`olp-connect <ip>` consumes `/health.anonymousKey` over the network — therefore requires `OLP_BIND` to include the LAN interface on the server side. Without setting `OLP_BIND=<lan-ip>` (or `0.0.0.0`), `olp-connect <ip>` will fail with `connect ECONNREFUSED` because the server only accepts loopback connections.
|
||||
|
||||
---
|
||||
|
||||
## Re-evaluation triggers
|
||||
|
||||
Re-open this ADR when ANY of the following fires:
|
||||
|
||||
1. OLP gains a "expose to public internet" deployment mode in the README
|
||||
(e.g., Cloudflare Tunnel guidance, ngrok recipe). At that point the
|
||||
soft-constraint MUST become a hard constraint (bind-address inspection
|
||||
at startup, refusal to enable `advertise_anonymous_key` when bind is
|
||||
public — likely with a separate `OLP_TRUSTED_PUBLIC_OVERRIDE=1` env
|
||||
escape hatch for operators who run their own TLS termination).
|
||||
2. The OCP `/health.anonymousKey` model is found to have caused a
|
||||
real-world quota-burn incident; that learning amends this ADR.
|
||||
3. Phase 5 introduces multi-tenant SaaS-like deployments (currently
|
||||
non-goal per ADR 0001); the entire family-scale assumption is
|
||||
re-examined.
|
||||
|
||||
---
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive.**
|
||||
|
||||
- Family-member onboarding becomes a single command: `olp-connect <ip>`. No
|
||||
out-of-band token paste. No "wait, what's the API key?" friction loop.
|
||||
- The trust trade-off is now an explicit, single-knob config decision, not
|
||||
an implicit consequence of OCP-pattern inheritance.
|
||||
- The `plaintext_advertise` field is a single auditable on-disk surface —
|
||||
`grep plaintext_advertise ~/.olp/keys/*/manifest.json` answers "which key
|
||||
is advertised?" definitively, and an operator who wants to disable the
|
||||
feature can simply revoke that key.
|
||||
- Owner-tier advertisement is impossible (both at keygen and at config
|
||||
load), eliminating an entire class of foot-gun.
|
||||
|
||||
**Negative.**
|
||||
|
||||
- ADR 0007 § 5's "no plaintext on disk, ever" property is weakened to "no
|
||||
plaintext on disk except for ONE explicitly-opted-in field on ONE key."
|
||||
The exception is narrow and audit-grep-able but the property is no
|
||||
longer absolute.
|
||||
- Operators who enable advertise mode then move the deployment from LAN to
|
||||
public internet (e.g., add a Cloudflare Tunnel without revisiting the
|
||||
config) silently invert the threat model. The startup warn for "public
|
||||
bind detected" does not currently fire (soft constraint per § "Trusted-
|
||||
LAN deployment invariant" above).
|
||||
- The OCP precedent shows operators sometimes share `olp-connect <ip>`
|
||||
invocations in chat / docs that include their IP; an LLM training corpus
|
||||
could harvest these IPs. The advertised key is only useful while the
|
||||
network reaches the IP, but the IP-disclosure surface grows.
|
||||
|
||||
**Neutral.**
|
||||
|
||||
- The plaintext storage is per-key, not global. Revoking the advertised key
|
||||
removes the plaintext exposure within one filesystem write (the manifest
|
||||
stays on disk for audit attribution per ADR 0007 § 6.1, but `revoked_at`
|
||||
becomes non-null and `findAdvertisedKey()` skips revoked manifests).
|
||||
|
||||
---
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
1. **Store plaintext in `config.json` directly.** Rejected. Mixes secrets
|
||||
with operational config; complicates git-crypt boundary; loses the
|
||||
per-key revocation path (you'd have to edit JSON to "revoke" the
|
||||
exposure rather than running `olp-keys revoke --id=<id>`).
|
||||
2. **Add an `anonymous` owner_tier instead of using `guest` + `plaintext_
|
||||
advertise`.** Rejected. Bumps ADR 0007 § 4 schema version (a
|
||||
non-additive change), adds a third identity class that the rest of the
|
||||
codebase (cache namespacing, /health gating, audit attribution) has no
|
||||
reason to know about, and conflicts with ADR 0007 § 7.1's "anonymous =
|
||||
no auth header + allow_anonymous=true" definition. A single optional
|
||||
field on the manifest is strictly less invasive.
|
||||
3. **Hard-enforce trusted-LAN bind address at startup.** Rejected for
|
||||
v0.4.0; deferred until a public-deployment-mode README section ships
|
||||
(see Re-evaluation triggers § 1). Soft constraint + startup warn is
|
||||
appropriate while OLP has zero public-internet deployment recipes.
|
||||
4. **Encrypt `plaintext_advertise` at rest with a key derived from
|
||||
`OLP_HOME` path or a separate `OLP_ADVERTISE_KEY` env var.** Rejected.
|
||||
The threat model is "anyone who can read `/health` reads the plaintext
|
||||
token over the wire," not "anyone who can read `~/.olp/keys/`." Both
|
||||
require LAN-reach; encrypting on-disk doesn't change the over-the-wire
|
||||
exposure. Adds complexity for no security gain in the relevant attack
|
||||
model.
|
||||
5. **Make `--advertise` allowed only when `--name` is exactly `anonymous`.**
|
||||
Rejected as over-restrictive. The CLI's `--anonymous` shorthand
|
||||
defaults `--name=anonymous`, but operators may legitimately want a
|
||||
named advertised key (e.g., `family-guest`, `lan-zero-config`). The
|
||||
discriminator is the field, not the name.
|
||||
|
||||
---
|
||||
|
||||
## Authority citations
|
||||
|
||||
- **ADR 0007 § 7** (Identity-class table — anonymous tier definition;
|
||||
`__anonymous__` keyId).
|
||||
- **ADR 0007 § 5** (Token format — establishes hash-only on-disk; D69 is
|
||||
the explicit opt-in exception).
|
||||
- **ADR 0007 § 4** (Manifest schema — D69 adds optional `plaintext_advertise`
|
||||
field; § 4 already specifies "unrecognized fields cause a warn but not a
|
||||
reject (forward-compat)" so the addition is non-breaking for older
|
||||
parsers).
|
||||
- **ADR 0007 § 7.2** (Configuration — D69 adds `auth.advertise_anonymous_key`
|
||||
alongside existing `allow_anonymous` / `owner_only_endpoints` /
|
||||
`fallback_detail_header_policy`).
|
||||
- **ADR 0010 § Phase 4 charter D68-D70 row** (scope authority for this ADR).
|
||||
- **OCP `server.mjs:148, 1454, 1488, 1555`** (prior-art for the
|
||||
`PROXY_ANONYMOUS_KEY` env + `/health.anonymousKey` pattern; OCP v3.13.0).
|
||||
- **OCP issue #12 § 14 Path A** (the original anonymous-key decision
|
||||
context for OCP; the "Path A" label is OCP-specific and not used in
|
||||
OLP).
|
||||
- **`bin/olp-connect`** (D68 client-side consumer of `/health.anonymousKey`).
|
||||
- **`bin/olp-keys.mjs`** (D69 keygen `--advertise` flag implementation).
|
||||
- **`lib/keys.mjs` `findAdvertisedKey()`** (D69 server-side resolver).
|
||||
- **`server.mjs handleHealth`** (D69 emission point + startup-warn site).
|
||||
|
||||
---
|
||||
|
||||
## Procedural mechanism
|
||||
|
||||
CC 开发铁律 v1.6 § 10 (independent reviewer per implementation D-day) — D68
|
||||
+ D69 + D70 ship as ONE PR per Iron Rule 11 IDR (the three deliverables
|
||||
are mutually constituting: `olp-connect` consumes `/health.anonymousKey`,
|
||||
`/health.anonymousKey` is governed by ADR 0011, ADR 0011 documents
|
||||
`olp-connect`'s trust posture). The reviewer is a fresh-context opus
|
||||
subagent.
|
||||
@@ -0,0 +1,147 @@
|
||||
# ADR 0012 — Phase 5 Charter: Provider Quota Probes + Dashboard Enrichment
|
||||
|
||||
**Status:** Accepted (Phase 5 open as of 2026-05-26)
|
||||
**Date:** 2026-05-26
|
||||
**D-day:** D79 (charter + ADR 0002 Amendment 8 + ADR 0013 land together as the constitutional layer of Phase 5)
|
||||
|
||||
## Amendments
|
||||
|
||||
### Amendment 1 — 2026-05-26: D84 Mistral probe NO-GO (post-D79-close spike)
|
||||
|
||||
The D-day table originally listed D84 as "optional, depends on D79-close 30-min Mistral docs spike". The spike completed 2026-05-26 with verdict **NO-GO** — Mistral does not expose a programmatic quota/usage endpoint **accessible to Vibe / Le Chat member / La Plateforme API keys** (the key tier OLP uses for spawning the `vibe` CLI):
|
||||
|
||||
- `docs.mistral.ai/api` (the public API spec) covers Chat, FIM, Embeddings, Classifiers, Files, Models, Batch, OCR, Audio, Events, Beta (Agents/Conversations/Libraries/Workflows/Observability). No usage/quota/credits/billing/limits endpoint accessible to a member API key.
|
||||
- Direct probe `https://api.mistral.ai/v1/usage` returns 404.
|
||||
- Mistral's "Limits and Usage" help article documents limit viewing via the `admin.mistral.ai/plateforme/limits` web console.
|
||||
- No `x-ratelimit-*` response headers documented on `/v1/chat/completions`. (Third-party summaries mentioning these headers are unsourced — appears to be OpenAI-convention extrapolation.)
|
||||
- OLP `lib/providers/mistral.mjs` already records this independently — DL-7 comment: "If quota/budget API surfaces in Le Chat Pro, pin the endpoint here."
|
||||
|
||||
**Out-of-scope but worth pinning for future revisit.** Mistral's [Admin API](https://docs.mistral.ai/admin/security-access/admin-api) DOES expose programmatic "Billing and usage queries", and the [Usage limits docs](https://docs.mistral.ai/admin/user-management-finops/usage-limits) describe usage/cost queries via that surface. The Admin API requires an **org-admin scoped API key** (separate from the member key OLP uses). For OLP's family-tier deployment posture (a maintainer's personal Le Chat Pro / La Plateforme account, not an organization's admin console), provisioning + storing an org-admin token raises the credential-scope ceiling beyond what the trusted-LAN deployment context (ADR 0011) was designed for. The NO-GO at v0.5.0 is therefore "out of scope for OLP's current deployment posture", NOT "Mistral has no programmatic surface". If the deployment posture expands to an org-admin context (e.g., a small-business multi-user deployment), this decision should be re-evaluated.
|
||||
|
||||
**Disposition:**
|
||||
- D84 row dropped from D-day plan (struck through below).
|
||||
- Mistral dashboard row in D82 UI shows "spend tracking only" badge sourced from `audit-query.mjs` aggregates (request count, estimated cost from `estimateCost()`).
|
||||
- `DL-7` in `mistral.mjs` is the documented re-entry point if Mistral ever publishes a usage endpoint.
|
||||
- Phase 5 total D-day budget revised: ~5 D-days (down from ~6).
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
Phase 4 (ADR 0010) shipped OLP's operator + client UX layer — `bin/olp` operator CLI, `olp doctor` framework, `olp-connect` zero-config IDE wiring, OpenClaw `/olp` slash commands, anonymous-key deployment-context limits, SSE heartbeat. v0.4.4 is the current shipped state. Phase 4 closed every gap on the OCP-feature-parity matrix EXCEPT one: **live quota / plan-usage surfacing**.
|
||||
|
||||
Today `lib/providers/anthropic.mjs:445` has a stub `quotaStatus()` returning `null` (D4 placeholder). The OLP dashboard's quota panel renders "—" for all providers. OCP, in contrast, exposes a live "39% session / 30% weekly" panel — the maintainer uses this multiple times per day to decide when to throttle voluntary `claude -p` traffic away from interactive sessions. OLP cannot become an OCP successor in practice (vs. just feature-parity-on-paper) until quota surfacing works.
|
||||
|
||||
A pre-flight institutional-knowledge audit (2026-05-26 — see `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`) confirmed:
|
||||
|
||||
1. **The OCP probe still works today** — Anthropic returns the same `anthropic-ratelimit-unified-*` headers on every `POST /v1/messages` call. Tested live 2026-05-26 from PI231 OAuth credentials.
|
||||
2. **Schema added 3 fields since OCP's 2026-04 capture** — `5h-status`, `7d-status` (per-window status), `overage-reset` (only on active overage). No fields removed or renamed.
|
||||
3. **Verification protocol has shifted** — Claude Code v2.1.x is now a **compiled binary** (Mach-O / ELF), not bundled JS. OCP's "grep cli.js" approach no longer applies; the replacement protocol is `strings` against the binary + periodic live probe diff.
|
||||
4. **OAuth refresh path unchanged** — `platform.claude.com/v1/oauth/token` + `9d1c250a-...` client_id + 60s-3600s exponential backoff.
|
||||
|
||||
The audit makes Phase 5 implementation low-risk: this is a port of a working OCP function, not a re-derivation. The work is mechanical + adapter-layer plumbing into OLP's plugin contract.
|
||||
|
||||
A parallel maintainer request (2026-05-26, with reference screenshot of claude.ai/settings/usage) asked for Claude.ai-style dashboard enrichment: per-row utilization bars, reset countdown, 1-minute auto-refresh, manual refresh button. P5-1 (probe) + P5-2 (dashboard) together unlock both: the data plus the surface. v1.x roadmap #8 is closed by P5-2.
|
||||
|
||||
---
|
||||
|
||||
## Decision
|
||||
|
||||
Phase 5 scope is **Provider quota probes + dashboard enrichment**. The phase opens 2026-05-26 with D79 (this charter + ADR 0002 Amendment 8 + ADR 0013 OAuth READ-ONLY consumption rules). Phase 5 close ships v0.5.0; per `CLAUDE.md release_kit.phase_rolling_mode`, the close PR is maintainer-triggered.
|
||||
|
||||
### In scope — Phase 5 D-day plan (~6 D-days)
|
||||
|
||||
| D-day | Deliverable | Authority | Estimate |
|
||||
|---|---|---|---|
|
||||
| **D79** | This charter ADR 0012 + ADR 0002 Amendment 8 (direct-API READ-ONLY) + ADR 0013 (OAuth READ-ONLY consumption rules + schema-drift mitigation) + `package.json` `current_pre_release_identifier` → `0.5.0-phase5` + `CLAUDE.md release_kit.phase_rolling_mode.current_phase` → Phase 5 | This charter + audit memory | 0.5d |
|
||||
| **D80** | `lib/providers/anthropic.mjs:quotaStatus()` ported from OCP `server.mjs:842-1109` — full probe with macOS-keychain auth read added (existing OLP reader only handles env + `.credentials.json`) + 5min cache + 60s-3600s refresh backoff + stale-cache-on-429 + all 13 headers parsed (including new 5h-status / 7d-status / overage-reset) | Port OCP probe + ALIGNMENT.md Rule 2 exemption per ADR 0002 Amendment 8 + audit memory | 2d |
|
||||
| **D81** | `lib/audit-query.mjs` + `/v0/management/dashboard-data` extended to surface the new quota shape per provider (utilization, reset, representative-claim, fallback-percentage, overage-status). Audit-query stays in-memory scan per ADR 0008 Lane 2 = A (no SQLite). Schema migration documented in ADR 0008 § Amendment | ADR 0008 + this charter | 1d |
|
||||
| **D82** | `dashboard.html` Claude.ai-style restructure — per-provider rows replace the current single Quota panel; each row: provider badge, model placeholder, utilization bar (5h + 7d), reset countdown ("Your limit will reset at HH:MM AM/PM" format from the user-shared claude.ai screenshot), status badge, representative-claim hint. 1-minute auto-refresh via `setInterval` with `document.visibilityState` guard. Manual refresh button calls `/v0/management/dashboard-data` directly | v1.x roadmap #8 + maintainer reference screenshot | 1.5d |
|
||||
| **D83** | Test coverage — Suite 38 quota-probe unit tests (mock HTTP server returning the 13 headers; assert parse + cache + backoff + stale-on-429); Suite 39 dashboard rendering smoke (curl `/dashboard` after pre-seeding mock quota cache; assert HTML contains expected utilization strings); update Suite 33 doctor checks for new `anthropic.quota_probe_reachable` check | Test convention from existing suites | 1d |
|
||||
| ~~D84~~ **DROPPED** | ~~Mistral `quotaStatus()` port — depends on D79-close spike~~ **NO-GO per 2026-05-26 spike (see § Amendment 1).** Mistral dashboard row in D82 shows "spend tracking only" badge sourced from `audit-query.mjs` aggregates. `DL-7` hook point in `mistral.mjs` already marks the location for future upgrade if Mistral ever publishes a usage endpoint. Codex permanently skipped (no public API). | n/a (dropped) | 0d |
|
||||
| **close** | v0.5.0 release PR — `package.json` `0.4.4 → 0.5.0`, CHANGELOG promotion, `release_kit.phase_rolling_mode.current_pre_release_identifier` advance to Phase 6 token | `CLAUDE.md release_kit overlay` | maintainer-triggered |
|
||||
|
||||
### Out of Phase 5 scope (with explicit triggers)
|
||||
|
||||
#### `X-OLP-Cost-USD` per-request response header
|
||||
|
||||
**Status:** Deferred to Phase 6. Was listed in ADR 0010 § Out-of-scope as "Phase 5 prerequisite". The prerequisite (provider-cost weights table) is non-trivial — needs per-(provider, model) `input_cost_per_1k_tokens` / `output_cost_per_1k_tokens` / `cache_read_discount` data sourced from each provider's published pricing page. Phase 5 already pulls in two new ADRs; adding a third data-onboarding ADR is scope creep.
|
||||
|
||||
**Re-open condition.** Phase 6 unless a maintainer reports a cost-attribution debugging need that warrants pulling forward.
|
||||
|
||||
#### `context_window_exceeded` fallback trigger (LiteLLM prior-art)
|
||||
|
||||
**Status:** Deferred. ADR 0010 listed this as opportunistic-in-Phase-5 unless the trigger fires sooner. The trigger has not fired in Phase 4 production traffic. Continue to defer.
|
||||
|
||||
#### per-(provider, model) live stats Map (replacing audit-query scan)
|
||||
|
||||
**Status:** Deferred. Current scan latency is ~20ms at 7-day depth. Acceptable until volume grows (>100k requests/day). Re-evaluate at Phase 6 if dashboard latency degrades.
|
||||
|
||||
#### Anthropic interactive-mode P0 (ADR 0009)
|
||||
|
||||
**Status:** Still trigger-gated on Anthropic's 2026-06-15 billing-split rollout. Phase 5 does NOT depend on P0 — the quota probe reads `anthropic-ratelimit-unified-*` headers regardless of which billing pool the spawn path consumes. If P0 succeeds Phase 7+ Phase 5's probe code remains unchanged; if P0 fails Phase 5's probe code remains unchanged. The probe is billing-pool-agnostic because the headers are subscription-pool metadata, not Agent-SDK-Credit metadata.
|
||||
|
||||
#### `/v1/messages` Anthropic-shape entry surface
|
||||
|
||||
**Status:** Still deferred per ADR 0010 § Out-of-scope. No change in Phase 5.
|
||||
|
||||
#### v1.x roadmap #3 / #5 / #6
|
||||
|
||||
**Status:** Still trigger-gated per `docs/v1x-roadmap.md`. None has fired. Continue to defer.
|
||||
|
||||
### Opportunistic Phase 5 micro-additions (not blocking)
|
||||
|
||||
Items small enough to land alongside a planned D-day without scope creep, if encountered:
|
||||
|
||||
- README § Dashboard screenshot update (post-P5-2 enrichment) — capture from MacBook test path per `~/.cc-rules/memory/feedback/mac_mini_never_for_testing.md`.
|
||||
- `olp usage` CLI subcommand (bin/olp.mjs) surfaces the parsed quota shape in terminal form. Already partially exists (cmdUsage in bin/olp.mjs); confirm payload alignment after D80.
|
||||
- Add `claude_code_oauth_client_id` config override in `~/.olp/config.json` so power users can override the hardcoded `9d1c250a-...` UUID without env-var fiddling. Mirrors compiled binary's `CLAUDE_CODE_OAUTH_CLIENT_ID` env support.
|
||||
- `docs/provider-audits/anthropic.md` re-capture with current `claude --version` (v2.1.142 MacBook / v2.1.150 PI231) + binary distribution layout note.
|
||||
|
||||
### Exit gate — v0.5.0 close criteria
|
||||
|
||||
1. D79 — D84 all merged with fresh-context opus reviewer APPROVE per Iron Rule 10.
|
||||
2. CI green on every D-day merge commit and on the v0.5.0 release commit head. `alignment.yml` blacklist re-confirmed (no new hallucinated tokens introduced).
|
||||
3. README § Quota / Plan Usage section present with screenshot of the enriched dashboard. README § Supported Providers table updated to note "quota probe: anthropic ✅, mistral ⚠️/✅ (D84 outcome), codex ❌ (no public API)".
|
||||
4. ADR 0012 (this charter) + ADR 0002 Amendment 8 + ADR 0013 (OAuth READ-ONLY consumption) on disk.
|
||||
5. `CHANGELOG.md "Unreleased"` promoted to `"## v0.5.0 — <date>"` with D79 — D84 entries.
|
||||
6. `package.json` bumped to `0.5.0`.
|
||||
7. `CLAUDE.md release_kit.phase_rolling_mode.current_phase` advances `Phase 5 → Phase 6`; `current_pre_release_identifier` advances `0.5.0-phase5 → 0.6.0-phase6`.
|
||||
8. Standing autopilot grant covers D-day-by-D-day execution; v0.5.0 close PR is maintainer-triggered.
|
||||
9. Live MacBook E2E verification — dashboard renders enriched panel with real quota data (probe live, not mocked).
|
||||
|
||||
---
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive.**
|
||||
|
||||
- OLP finally has the load-bearing observability OCP had — maintainer can see live "39% session / 30% weekly" and decide whether voluntary `claude -p` traffic stays or moves.
|
||||
- Family members on the LAN see real reset times instead of "—", which makes the "wait 2 hours" guidance concrete vs. abstract.
|
||||
- The institutional-knowledge audit captured the schema in a memory file pinned with date stamps — future ports (mistral, future provider) re-use the verification protocol without re-deriving.
|
||||
- v1.x roadmap #8 (Dashboard enrichment per Claude.ai-style usage page) closes inside Phase 5 rather than waiting for a separate phase.
|
||||
- Compiled-binary-distribution awareness ("no more cli.js to grep") is now codified in OLP governance; the next time Anthropic ships a major CC version, the verification protocol is already written.
|
||||
|
||||
**Negative.**
|
||||
|
||||
- ADR 0002 gains another amendment (Amendment 8). The constitution surface area for `anthropic.mjs` grows. Counter-pressure: the alternative (probe lives in `server.mjs`, like OCP) violates the plugin-architecture principle that per-provider knowledge stays in `lib/providers/`. Amendment 8 is the smaller violation.
|
||||
- The probe makes one `/v1/messages` call per 5min cache miss. That's ~12 calls/hour worst case across the whole proxy (probe is per-credentials, not per-key). With `max_tokens: 1` the cost is < $0.01/day at family-scale traffic. Negligible but not zero.
|
||||
- Schema-drift risk over the long horizon. Anthropic could rename or remove headers in a future version. The mitigation protocol (strings + live probe diff) is in place, but it's a manual check — needs to be invoked by the maintainer or scheduled.
|
||||
- Dashboard refactor introduces a breaking-change risk for the existing dashboard.html consumers (none today, but conceptually). Bumping to v0.5.0 signals this clearly.
|
||||
|
||||
**Neutral.**
|
||||
|
||||
- Phase 5 has more ADR work than Phase 4 (3 governance docs vs. 2). The constitutional layer is deliberately heavier because direct-API access is the single biggest authority decision since the plugin contract itself.
|
||||
|
||||
---
|
||||
|
||||
## Authority + cross-references
|
||||
|
||||
- **Iron Rule 11 (IDR)** — Phase 5 ships across 6 D-days, each a minimum reviewable unit. The governance trio (this ADR + Amendment 8 + ADR 0013) lands at D79 as a single coupled commit (reviewing them separately cannot verify consumer-producer alignment), per ADR 0002 Amendment 7's precedent.
|
||||
- **Iron Rule 12 (prior-art search)** — discharged via the audit memory at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`. Memory committed prior to D80 implementation.
|
||||
- **ALIGNMENT.md Rule 1 (citation)** — D80 commit must cite compiled-binary `strings` evidence per audit memory § Path A (Claude Code v2.1.x has no traditional `§ section` structure because it is a Mach-O / ELF compiled binary) plus the audit memory file path. Live-probe transcript MUST be included in the commit body.
|
||||
- **ALIGNMENT.md Rule 2 (provider-CLI-as-authority)** — direct-API access bypasses the spawn-binary contract. Amendment 8 is the explicit exemption. Without Amendment 8, the D80 commit is unalignable.
|
||||
- **ALIGNMENT.md Rule 5 (CI alignment.yml)** — must continue to pass. `api.anthropic.com/v1/messages` is NOT on the blacklist (correct — that's the real endpoint). The hallucinated `/api/oauth/usage` IS on the blacklist (transitive from OCP) and must remain.
|
||||
- **ADR 0002 Amendment 8** — companion ADR. Direct-API access scoping; READ-ONLY constraint; opt-in via config flag (default off).
|
||||
- **ADR 0013** — companion ADR. OAuth credentials shared between spawn path + probe path; refresh backoff; schema-drift mitigation protocol.
|
||||
- **CLAUDE.md release_kit** — Phase boundary triggers maintainer-led version bump. D-day commits within Phase 5 stay under "Unreleased". `0.5.0-phase5` is the pre-release identifier during the phase.
|
||||
@@ -0,0 +1,197 @@
|
||||
# ADR 0013 — OAuth READ-ONLY Consumption Rules + Schema-Drift Mitigation Protocol
|
||||
|
||||
**Status:** Accepted (2026-05-26)
|
||||
**Date:** 2026-05-26
|
||||
**D-day:** D79 (lands alongside ADR 0012 Phase 5 charter + ADR 0002 Amendment 8 as the constitutional trio of Phase 5)
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
ADR 0002 Amendment 8 permits `quotaStatus()` to call provider HTTP APIs directly, subject to a READ-ONLY constraint. That Amendment opens the door but does not specify HOW READ-ONLY discipline is preserved across credential lifecycle events (refresh, expiry, revocation), nor how OLP detects when the upstream API schema drifts. ADR 0013 fills both gaps.
|
||||
|
||||
The motivating concern: a provider that ships its CLI as a **compiled native binary** (Anthropic Claude Code v2.1.x is now Mach-O on macOS, ELF on Linux) closes off the previous schema-verification path (grep `cli.js`). If OLP's probe parser silently breaks because a header was renamed, the dashboard shows stale or wrong numbers, and the maintainer's load-bearing throttling decision is based on bad data. This ADR establishes the verification protocol that survives the binary-distribution shift.
|
||||
|
||||
A second motivating concern: the OAuth credentials used by the probe are the SAME credentials the spawn path uses for `claude -p`. Both paths consume them; the probe must not interfere with the spawn path's ability to refresh or invalidate them. Concretely: the probe must not write to the credentials artifact, must not race the spawn path on refresh, and must not amplify a 429 into a refresh storm.
|
||||
|
||||
---
|
||||
|
||||
## Decision
|
||||
|
||||
### Rule 1 — Credential reuse is mandatory
|
||||
|
||||
The probe MUST consume the same OAuth artifact the spawn path reads via the plugin's `readAuthArtifact()`. No new OAuth grant. No alternate credential store. No environment-variable-only fallback (env var `CLAUDE_CODE_OAUTH_TOKEN` is supported as an override consistent with the spawn path, but is not the probe's primary source).
|
||||
|
||||
Precedence order (mirrors OCP `getOAuthCredentials` 2026-04-stable):
|
||||
|
||||
1. `process.env.CLAUDE_CODE_OAUTH_TOKEN` if non-empty (manual override; common in CI / dev / one-off debugging).
|
||||
2. `~/.claude/.credentials.json` → `claudeAiOauth.accessToken` (Linux + macOS without keychain access).
|
||||
3. macOS Keychain: `security find-generic-password -a "${USER}" -s "Claude Code-credentials" -w` (preferred on macOS — current `lib/providers/anthropic.mjs` only covers (1) + (2); D80 adds (3)).
|
||||
|
||||
Rationale: a separate OAuth grant would require the maintainer to repeat `claude setup-token` against an OLP-specific scope, doubling credential exposure and divergence risk. Reusing the spawn path's credentials guarantees the probe never has more permission than the spawn path itself.
|
||||
|
||||
### Rule 2 — READ-ONLY at the wire
|
||||
|
||||
The probe MUST issue exactly one HTTP request per cache miss. Method MAY be POST (Anthropic's ratelimit headers come back on `POST /v1/messages`; this is the only way to read them). Request body MUST minimise side effects:
|
||||
|
||||
- `max_tokens: 1` (cost: ~$0.000001 per probe)
|
||||
- `messages: [{role: "user", content: "hi"}]` (any minimal valid payload)
|
||||
- Model: cheapest available in the plan (`claude-haiku-4-5` at v0.5.0)
|
||||
- Do NOT include `system` prompts, `tools[]`, `tool_choice`, large content arrays, or anything that the upstream might bill differently.
|
||||
|
||||
The probe MUST discard the response body. Only response headers are parsed.
|
||||
|
||||
The probe MUST NOT call any other HTTP path on the provider's API. No `/v1/models` enumeration, no admin endpoints, no `/v1/messages/<id>` retrievals. The only permitted endpoint is `POST /v1/messages`.
|
||||
|
||||
### Rule 3 — Cache TTL and refresh discipline
|
||||
|
||||
- Cache TTL: 5 minutes. Cache miss triggers a real probe. Cache hit returns the cached value.
|
||||
- The dashboard refreshes every 1 minute; that's served from the cache between probes. A manual refresh button MAY force-clear the cache (per maintainer request 2026-05-26); ADR 0012 D82 documents the button.
|
||||
- On refresh failure (token expired, 401/403/429, network error), the probe schedules an exponential backoff: minimum 60s, maximum 3600s. The cache entry is NOT invalidated during backoff; `quotaStatus()` returns the stale cache marked `{ stale: true, last_fresh_at: <epoch> }`. If no stale entry exists, returns an `unreachable` shape (v0.5.1+) rather than `null`.
|
||||
- Successive successful probes reset the backoff to the minimum.
|
||||
- Token refresh (`POST https://platform.claude.com/v1/oauth/token`) follows the same backoff discipline. The probe MUST NOT refresh a token more than once per backoff window. The refresh path is shared with the spawn path; both observe the same backoff.
|
||||
- **All consumers of `quotaStatus()`, including `olp doctor` checks, MUST route through `quotaStatus()` and MUST NOT call `_probeOnce()` directly.** `_probeOnce()` is an internal implementation detail. Routing doctor checks through `quotaStatus()` ensures the cache+backoff discipline is enforced for every caller — including operators running `olp doctor` in a debug loop. (Clarification added v0.5.1 to address codex finding F1: the original doctor check bypassed backoff by calling `_probeOnce` directly.)
|
||||
|
||||
### Rule 4 — Opt-in via config
|
||||
|
||||
A new config field at `~/.olp/config.json` controls per-provider opt-in:
|
||||
|
||||
```json
|
||||
{
|
||||
"providers": {
|
||||
"anthropic": {
|
||||
"enabled": true,
|
||||
"quota_probe_enabled": false
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Default: `false`. The maintainer must explicitly opt in after credentials are configured. Reasoning: a fresh install on a machine without OAuth credentials should not bombard `api.anthropic.com` with 401-bound probes.
|
||||
|
||||
`olp doctor` adds a per-provider check `<provider>.quota_probe_reachable` (only runs if `quota_probe_enabled: true`). Failed check provides a `next_action.ai_executable[]` recipe to either re-authenticate or disable the probe.
|
||||
|
||||
### Rule 5 — Schema-drift mitigation protocol (minimum-viable-schema gate)
|
||||
|
||||
The CC binary-distribution shift means OCP's "grep cli.js" verification is no longer applicable. OLP adopts a two-path protocol for proactive monitoring, AND enforces a minimum-viable-schema gate at parse time:
|
||||
|
||||
**Minimum-viable-schema gate (v0.5.1+).** `_probeOnce()` requires at least these 4 fields present (non-null after parse) before treating a response as successful:
|
||||
- `anthropic-ratelimit-unified-5h-utilization`
|
||||
- `anthropic-ratelimit-unified-5h-reset`
|
||||
- `anthropic-ratelimit-unified-7d-utilization`
|
||||
- `anthropic-ratelimit-unified-7d-reset`
|
||||
|
||||
If any of these 4 is absent, `_probeOnce()` classifies the probe as a schema-drift failure (`failureKind = 'schema_drift'`), schedules backoff, and returns `null`. This means a 200 OK with zero `anthropic-ratelimit-*` headers (e.g. a server-side change, a proxy stripping headers, or a mock returning `{}`) is immediately caught as drift rather than silently cached as "live" data. The other 9 fields are tolerated as absent (overage fields are conditional; top-level status fields may be absent on edge cases). The 5h/7d core 4 are load-bearing — the dashboard's progress bars depend on them. (Gate added v0.5.1 to address codex finding F2.)
|
||||
|
||||
The CC binary-distribution shift means OCP's "grep cli.js" verification is no longer applicable. OLP adopts a two-path protocol:
|
||||
|
||||
**Path A — Compiled-binary string extraction.** Run `strings` over the platform-specific binary in the claude-code distribution. Captures all hardcoded header names the binary expects:
|
||||
|
||||
```bash
|
||||
BIN_DIR=$(npm root -g)/@anthropic-ai/claude-code/node_modules/@anthropic-ai/claude-code-*
|
||||
strings "$BIN_DIR/claude" | grep -iE "anthropic-ratelimit|/v1/(messages|oauth)|platform\.claude\.com"
|
||||
```
|
||||
|
||||
**Path A prerequisites.** GNU or BSD `strings` (part of binutils/coreutils on Linux + macOS — always present on a normal developer machine; Windows requires WSL or `binutils-mingw`). A locally installed Claude Code v2.1.x (npm-global or volta-managed). A reviewer without `claude` installed can still run Path B but Path A is gated on having the binary on disk. A future Claude Code version that ships as a different distribution shape (e.g. Rust binary, statically linked Go) keeps the protocol valid: `strings` works on any ELF/Mach-O regardless of compile source.
|
||||
|
||||
**Path B — Live API probe.** Run the actual probe against `api.anthropic.com` with valid OAuth credentials. Captures what the server returns today:
|
||||
|
||||
```bash
|
||||
curl -s -i -m 10 -X POST https://api.anthropic.com/v1/messages \
|
||||
-H "Authorization: Bearer $TOKEN" \
|
||||
-H "anthropic-beta: oauth-2025-04-20" \
|
||||
-H "anthropic-version: 2023-06-01" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model":"claude-haiku-4-5","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' \
|
||||
| grep -iE "^anthropic-ratelimit"
|
||||
```
|
||||
|
||||
Path A tells you what the client expects. Path B tells you what the server actually emits. The diff is the actionable schema delta.
|
||||
|
||||
**Required cadence.** The diff MUST be re-run at every major `claude --version` bump (v2.x → v3.x is the next trigger). The current pinned schema lives at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`. After re-verification, that memory file MUST be updated (or a successor file written with a new date stamp; the old one cross-linked).
|
||||
|
||||
**Trigger for re-running the diff.** There is no automated detector for a major `claude --version` bump at v0.5.0. Three explicit hooks share this responsibility:
|
||||
|
||||
1. **Annual Alignment Audit** (`ALIGNMENT.md` § Annual Alignment Audit, every 14 May) — diff is mandatory as part of the audit checklist.
|
||||
2. **`olp doctor anthropic.quota_probe_reachable` failure** — if the probe returns non-2xx for any reason other than 401/403/429/network (typical schema breaks manifest as 422 or 400), `olp doctor` surfaces a `kind: fix_provider` recipe whose first step is "re-run the Rule 5 dual-path diff".
|
||||
3. **Manual maintainer attention at a major Claude Code release** — if the maintainer sees a major version bump in `claude --version`, kick off the diff before the next Phase opens. Rolling-mode discipline (CLAUDE.md release_kit) means major-version bumps usually intersect with Phase boundaries.
|
||||
|
||||
If the diff is missed across a major version bump, the failure mode is graceful degradation: the parser silently drops unknown headers; the dashboard shows older values (cached stale) or `null` per Rule 3; `olp doctor` surfaces the staleness.
|
||||
|
||||
**Required action on drift detection.** If a header is renamed or removed:
|
||||
|
||||
1. File a Phase-N issue tagging the maintainer.
|
||||
2. Update the parser in `lib/providers/anthropic.mjs:quotaStatus()` to handle both names (graceful migration), prefer the new name.
|
||||
3. Update the audit memory file with a "drift event" section recording: date, old field, new field, evidence URLs.
|
||||
4. Bump the `models-registry.json` `quota_probe.schema_version` (NEW field added at D80) so downstream consumers can detect.
|
||||
|
||||
If a new header appears in the live response that the parser doesn't read: low-priority enhancement; add to the parser, document in the audit memory, no schema_version bump required.
|
||||
|
||||
### Rule 6 — Failure transparency
|
||||
|
||||
The probe's failure modes are visible to the operator:
|
||||
|
||||
- `/v0/management/dashboard-data` includes per-provider `{ quota_probe: { status: 'ok' | 'stale' | 'failed' | 'disabled', last_fresh_at, last_error?, backoff_until? } }`.
|
||||
- `olp doctor` surfaces probe failure as `kind: fix_oauth` (if 401/403) or `kind: fix_provider` (if 429 with no stale cache or network error).
|
||||
- The dashboard row badge shows the status; clicking a failed row shows the last error (truncated to 200 chars, no full credential traces).
|
||||
|
||||
### Rule 7 — Out-of-scope
|
||||
|
||||
This ADR does NOT govern:
|
||||
|
||||
- Spawn-path OAuth refresh (the spawn path's refresh logic predates this ADR and is governed by the underlying CLI). The probe shares the credential artifact but does not own the refresh.
|
||||
- Anthropic-specific bearer revocation (Anthropic side). Revocation manifests as 401 to the probe, which falls into Rule 6.
|
||||
- Non-Anthropic provider OAuth flows. Mistral / future providers MAY adopt this protocol via plugin-specific ADRs; ADR 0013 establishes the template.
|
||||
|
||||
---
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive.**
|
||||
|
||||
- The probe is bounded — Rule 2 caps the wire traffic, Rule 3 caps the refresh rate, Rule 4 caps activation surface.
|
||||
- Schema-drift detection is procedural and reproducible — Rule 5 gives the maintainer a runbook that doesn't depend on Anthropic publishing a deprecation notice.
|
||||
- Failure is visible — Rule 6 means a broken probe shows up in `olp doctor` and the dashboard, not as a silent "—" in the quota row.
|
||||
- Credential reuse (Rule 1) keeps the security surface area minimal.
|
||||
|
||||
**Negative.**
|
||||
|
||||
- The `quota_probe_enabled` opt-in adds a configuration step. Mitigated by `olp doctor` surfacing the recipe when credentials are present but the probe is off.
|
||||
- The schema-drift protocol is manual. Anthropic could ship a v3.x binary tomorrow and the verification only happens when the maintainer or a doctor probe failure prompts it. Counter-pressure: drift events at OCP scale (~12 months) suggest manual verification on major version bumps is sufficient.
|
||||
- Stale-cache-on-failure (Rule 3) means the dashboard could show 30-minute-old data without an obvious "stale" indicator unless the UI explicitly renders the `stale: true` marker. ADR 0012 D82 requires the dashboard to surface staleness; reviewing that during P5-2 implementation.
|
||||
|
||||
**Neutral.**
|
||||
|
||||
- The protocol is portable. Future provider plugins adopting direct-API probes (mistral if its `/v1/usage` exists) can reuse the same six rules with provider-specific endpoint substitution.
|
||||
|
||||
---
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
### A — Probe lives in `server.mjs` (OCP-style)
|
||||
|
||||
OCP's probe is in `server.mjs:842-1109` because OCP is single-provider and pre-plugin-architecture. Porting that pattern to OLP would violate ADR 0002 (per-provider knowledge stays in `lib/providers/`). Rejected.
|
||||
|
||||
### B — Spawn `claude -p --dry-run` and parse ratelimit headers
|
||||
|
||||
`claude -p` does not expose response headers; the CLI consumes and discards them. Even if it did, parsing CLI stdout is fragile. Rejected.
|
||||
|
||||
### C — Wait for Anthropic to publish a public quota API
|
||||
|
||||
The 2026-06-15 Agent SDK Credit billing-split announcement does not include a public quota API. Anthropic may publish one in the future; this ADR is forward-compatible (Rule 7 explicitly notes "if Anthropic publishes a public ratelimit API, this entire workaround becomes obsolete — re-evaluate"). Rejected for v0.5.0 (no ETA).
|
||||
|
||||
### D — Mandate token-rotation in OLP
|
||||
|
||||
Tempting (auditability), but OCP's experience shows token rotation breaks the spawn path more often than it improves security at family-scale deployment. The credential rotation cadence is Anthropic-side (token TTL); OLP respects whatever Claude Code does. Rejected.
|
||||
|
||||
---
|
||||
|
||||
## Authority + cross-references
|
||||
|
||||
- **ADR 0002 Amendment 8** — the contract-level permission. ADR 0013 is the implementation discipline for that permission.
|
||||
- **ADR 0012** — Phase 5 charter that schedules D80 implementation.
|
||||
- **ADR 0011** — anonymous-key deployment-context (LAN-only). Separate scope; this ADR does not amend it.
|
||||
- **`~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`** — the live schema pin. Updated on every drift event per Rule 5.
|
||||
- **`alignment.yml`** — must continue to blacklist `/api/oauth/usage` and related hallucinated tokens. Must NOT add `/v1/messages` to the blacklist (legitimate endpoint).
|
||||
- **OCP `server.mjs:842-1109`** — the source-of-truth port reference for D80.
|
||||
- **OCP `ALIGNMENT.md`** — the institutional precedent (2026-04-11 drift → ALIGNMENT introduction) this ADR consolidates for OLP.
|
||||
@@ -0,0 +1,662 @@
|
||||
# ADR 0014 — Sandbox-Runtime Integration for Multi-Tenant Provider Spawning
|
||||
|
||||
**Status:** Accepted (PR-A shipped; PR-B shipped pending PI231 Suite 44 validation + HTTP-path activation debug; PR-C/D pending) — **see Amendment 1 (2026-05-29): PR-B's outer-bwrap approach is superseded by the per-spawn ephemeral-home + per-provider ISOLATION contract architecture. PR-C/D are reframed; the substantive decision moves into Amendment 1.**
|
||||
**Date:** 2026-05-28
|
||||
**Phase:** Phase 7
|
||||
|
||||
---
|
||||
|
||||
## Related
|
||||
|
||||
- **ADR 0001** (Project Founding) — OLP's multi-provider rationale and "no conversation state" principle.
|
||||
- **ADR 0009 Amendment 1** (stream-json transport, Phase 6) § Caveats #3: "Sandbox-runtime still required for real multi-tenant deployment."
|
||||
- **ADR 0002** (Plugin Architecture) — Provider contract; `spawn()` is the surface this ADR will wrap in PR-B/C.
|
||||
- **ADR 0006** (Provider Inclusion / Risk Tier Framework) — classifies providers by deployment risk; sandbox status is a gating condition for Tier-A (cloud-deployed).
|
||||
- **`docs/plans/cloud-deployment-family.md` § 5** — sandbox is a hard prerequisite before any cloud rollout.
|
||||
- **cc-mem incident memory** — `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` — the multi-tenant security gap that motivates this ADR.
|
||||
|
||||
---
|
||||
|
||||
## 1. Context
|
||||
|
||||
### 1.1 The multi-tenant security gap
|
||||
|
||||
OLP is a personal-scale proxy (ADR 0001 § Non-commercial). However, the "family-scale" deployment model means multiple human callers share a single OLP instance — each with their own OLP API key (ADR 0007) but all using the same underlying `claude` or `codex` CLI installation on the server host.
|
||||
|
||||
The security gap, identified in the 2026-05-27 session and captured in cc-mem incident memory § 3, is:
|
||||
|
||||
1. **OAuth token exposure.** A malicious (or misbehaving) prompt to the Anthropic provider could elicit a `cat ~/.olp/keys/...` or similar read of any file the OLP process user can access — including the OAuth credentials file that allows the attacker to impersonate the server-side identity.
|
||||
2. **Codex shell-tool execution.** The `codex exec` path exposes a shell tool to the model. With OLP acting as a relay, a prompt to codex from one client could execute arbitrary commands in the server process's working directory, reading or writing files belonging to other clients.
|
||||
3. **Cross-tenant data leakage.** Even without adversarial prompts, a model that freely accesses the filesystem could inadvertently leak one client's cached context to another client's response.
|
||||
|
||||
The 2026-05-27 prior-art search (incident memory § 4) surveyed the multi-tenant LLM proxy ecosystem (LiteLLM, OpenCode, CLIProxyAPI, open-source Anthropic proxies) and found that **none solve multi-tenant file-system and tool isolation at the OS level**. The field's typical answer is "don't run multi-tenant" or "use a separate VM per tenant" — neither applicable at OLP's family scale.
|
||||
|
||||
### 1.2 Anthropic's official answer: `@anthropic-ai/sandbox-runtime`
|
||||
|
||||
The `@anthropic-ai/sandbox-runtime` package (Anthropic Experimental org, `anthropic-experimental/sandbox-runtime`, v0.0.52 as of this ADR) is Anthropic's open-source solution to wrapping security boundaries around arbitrary processes. It is the library that Claude Code itself uses internally to sandbox MCP servers and tool execution.
|
||||
|
||||
The library provides:
|
||||
|
||||
- **Linux:** bubblewrap (`bwrap`) namespace isolation + socat network bridge + seccomp filter via `apply-seccomp-filter` binary. Ripgrep (`rg`) is required for deny-path glob expansion.
|
||||
- **macOS:** `sandbox-exec` seatbelt profile, which is a built-in OS facility (no additional packages required).
|
||||
|
||||
Both paths enforce filesystem read/write restrictions and network policy at the kernel level, not at the process level. A `cat ~/.olp/keys/...` inside the sandbox fails at the syscall layer regardless of what the shell or model requests.
|
||||
|
||||
### 1.3 The 2026-05-28 spike
|
||||
|
||||
A PoC spike was conducted on PI231 (arm64 Debian Bookworm) on 2026-05-28. Key findings:
|
||||
|
||||
1. **`npm install @anthropic-ai/sandbox-runtime@0.0.52` succeeds cleanly** on arm64 Linux. No native build step; prebuilt binaries were available.
|
||||
2. **`SandboxManager.isSupportedPlatform()` returns `true`** on PI231 (Linux, not WSL).
|
||||
3. **`SandboxManager.checkDependencies()` reports errors**: `bubblewrap (bwrap) not installed`, `socat not installed`, `ripgrep (rg) not found`. These are the three OS-level deps that must be installed separately (not bundled in the npm package).
|
||||
4. **The install fix is a one-liner**: `sudo apt-get install -y bubblewrap socat ripgrep`. This is a 5-minute operational task, not a code change.
|
||||
5. **Three PoC scripts** were parked at `/tmp/sandbox-spike/` on PI231 verifying: dependency check return shapes, `SandboxManager.wrapWithSandbox` call signature, and filesystem-deny path behaviour.
|
||||
|
||||
Verdict: **YELLOW** — architecturally green (the library works and the platform is supported), operationally blocked on apt deps. PR-A lays the dependency + doctor layer. PR-B wraps the anthropic spawn after apt install.
|
||||
|
||||
---
|
||||
|
||||
## 2. Decision
|
||||
|
||||
### 2.1 Layered rollout (Iron Rule 11 — minimum reviewable unit)
|
||||
|
||||
The sandbox integration is split into four discrete PRs, each independently reviewable and independently safe to land or revert:
|
||||
|
||||
| PR | Scope | Blocking condition | Status |
|
||||
|---|---|---|---|
|
||||
| **PR-A** (this PR) | npm dep `@anthropic-ai/sandbox-runtime ^0.0.52` + `lib/sandbox/doctor.mjs` (preflight module) + `/health` `sandbox` field + ADR 0014 | None — no runtime initialization | ✅ Accepted |
|
||||
| **PR-B** | `lib/sandbox/manager.mjs` (bootstrap + spawn-wrap) + `lib/providers/anthropic.mjs` spawn wrapped + server startup wiring + `/health.sandbox.active` + Suite 43/44 tests | `bubblewrap` + `socat` + `rg` installed on PI231 (`sudo apt-get install -y bubblewrap socat ripgrep`) | ✅ Implemented — pending PI231 validation (Suite 44) + opus reviewer |
|
||||
| **PR-C** | `lib/providers/codex.mjs` spawn wrapped with `enableWeakerNestedSandbox: true` | PR-B accepted + codex PoC on PI231 | 🔲 Blocked on PR-B |
|
||||
| **PR-D** | `docs/plans/cloud-deployment-family.md` § "Phase 7 prerequisite met" update; cloud rollout unblocked | PR-B + PR-C accepted | 🔲 Blocked on PR-C |
|
||||
|
||||
Rationale for the split:
|
||||
|
||||
- **PR-A is safe without bwrap.** The doctor module and `/health` field add observability with no runtime side effects. No `SandboxManager.initialize()` call. No sandbox spawned.
|
||||
- **PR-B is the load-bearing security gate.** Wrapping `anthropic.mjs` spawn requires empirical negative-test confirmation (in-sandbox `cat ~/.olp/keys/...` MUST fail). This cannot be verified until PI231 has bwrap installed.
|
||||
- **PR-C follows PR-B** because codex has a distinct issue: codex itself uses bubblewrap internally (`codex exec` spawns its own sandbox). `enableWeakerNestedSandbox: true` is required to allow the inner sandbox to function inside the outer OLP sandbox.
|
||||
- **PR-D is documentation-only** and depends on the runtime PRs being proven in production.
|
||||
|
||||
### 2.2 PR-A specific scope (binding)
|
||||
|
||||
PR-A MUST NOT include:
|
||||
|
||||
- Any call to `SandboxManager.initialize()` (no real sandbox created)
|
||||
- Any modification to `lib/providers/anthropic.mjs`, `lib/providers/codex.mjs`, or `lib/providers/mistral.mjs`
|
||||
- Any new HTTP endpoint (no `/metrics`, no new dashboard endpoint)
|
||||
- Any modification to `models-registry.json`
|
||||
|
||||
PR-A MUST include:
|
||||
|
||||
- `package.json` dependency: `"@anthropic-ai/sandbox-runtime": "^0.0.52"`
|
||||
- `lib/sandbox/doctor.mjs`: pure preflight module (no state; no initialization)
|
||||
- `/health` response: top-level `sandbox` field (`available`, `missing`, `platform`, `message` when unavailable)
|
||||
- `docs/adr/0014-sandbox-runtime-integration.md` (this document)
|
||||
- `CHANGELOG.md` Unreleased entry
|
||||
- `test-features.mjs` Suite 42 (8 new tests, all passing)
|
||||
|
||||
---
|
||||
|
||||
## 3. `lib/sandbox/doctor.mjs` design
|
||||
|
||||
### 3.1 Exports
|
||||
|
||||
```javascript
|
||||
// Returns { available: boolean, missing: string[], details: { ... } }
|
||||
export async function checkSandboxAvailability() { ... }
|
||||
|
||||
// Returns { ok: boolean, message: string } — human-readable summary
|
||||
export async function describeSandboxStatus() { ... }
|
||||
```
|
||||
|
||||
### 3.2 `checkSandboxAvailability` algorithm
|
||||
|
||||
1. Probe OS deps independently via `child_process.execFileSync('which', [binary])`:
|
||||
- `bwrap` (Linux only — macOS uses built-in `sandbox-exec`)
|
||||
- `socat` (Linux only)
|
||||
- `rg` (ripgrep — Linux only; macOS seatbelt profiles use regex patterns natively)
|
||||
2. Call `probeLibrary()` which `import()`s `@anthropic-ai/sandbox-runtime` and calls:
|
||||
- `SandboxManager.isSupportedPlatform()` — platform classification
|
||||
- `SandboxManager.checkDependencies(undefined)` — library's own dep check (called without initialize, falling back to PATH lookup)
|
||||
3. Compute `missing[]`: on Linux, add 'bubblewrap', 'socat', 'ripgrep' for each absent dep; if library import failed, add that too.
|
||||
4. `available = libLoaded && isSupportedPlatform && missing.length === 0`
|
||||
|
||||
`probeLibrary()` wraps everything in try/catch — any library-side error becomes `{ libLoaded: false, libError: '<reason>' }` rather than an unhandled rejection.
|
||||
|
||||
### 3.3 `/health` integration
|
||||
|
||||
The `sandbox` field is added to the full (owner-tier) payload only. For trimmed payloads (guest/anonymous per ADR 0007 § 7.1), the field is absent (consistent with the existing trim model). This prevents leaking infrastructure details to non-owner callers.
|
||||
|
||||
The result is memoized process-wide via `_sandboxStatusCache` in `server.mjs`. The install state of bwrap/socat cannot change at runtime without a process restart, so a single lazy fetch at the first `/health` call is correct.
|
||||
|
||||
```json
|
||||
{
|
||||
"ok": true,
|
||||
"version": "0.5.1",
|
||||
"providers": { ... },
|
||||
"sandbox": {
|
||||
"available": false,
|
||||
"missing": ["bubblewrap", "socat", "ripgrep"],
|
||||
"platform": "linux",
|
||||
"message": "Sandbox dependencies not available: bubblewrap not installed, socat not installed, ripgrep not installed. Install: sudo apt-get install -y bubblewrap socat ripgrep"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
When available (after apt install + process restart) and PR-B bootstrapped:
|
||||
|
||||
```json
|
||||
{
|
||||
"sandbox": {
|
||||
"available": true,
|
||||
"active": true,
|
||||
"missing": [],
|
||||
"platform": "linux"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
(PR-A shape did not include `active`. PR-B adds `active: boolean` — distinguishes
|
||||
"deps present" from "sandbox actually initialized and wrapping spawns".)
|
||||
|
||||
---
|
||||
|
||||
## 4. PR-B/C/D acceptance criteria
|
||||
|
||||
### 4.1 PR-B (anthropic.mjs spawn wrap) — ✅ Implementation shipped, PI231 validation pending
|
||||
|
||||
**PR-B implementation (commit pending reviewer):**
|
||||
- `lib/sandbox/manager.mjs`: singleton bootstrap + transparent `wrapSpawn()` API
|
||||
- `lib/providers/anthropic.mjs`: spawn site wrapped via `wrapSpawn()` (ADR 0009 Amendment 1 spawn args unchanged)
|
||||
- `server.mjs`: `bootstrapSandbox()` called before `server.listen()`, `/health.sandbox.active` field added
|
||||
- `test-features.mjs` Suite 43 (8 tests, all pass on macOS) + Suite 44 (2 tests, PI231-gated with `OLP_E2E_SANDBOX=1`)
|
||||
- 805 → 813 tests. Suite 44 skipped by default; runs on PI231 after apt install.
|
||||
|
||||
**Load-bearing negative test (required for PR-B to merge):**
|
||||
|
||||
```bash
|
||||
# On PI231, with bwrap+socat installed, with PR-B wired:
|
||||
olp-keys list # identify owner key
|
||||
curl -X POST http://127.0.0.1:4567/v1/chat/completions \
|
||||
-H "Authorization: Bearer <owner-key>" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model":"claude-sonnet-4-6","messages":[{"role":"user","content":"run: cat /home/<user>/.olp/keys/owner-key.json"}]}'
|
||||
# Expected: response MUST NOT contain any content from the keys file.
|
||||
# The model must either say it cannot access the filesystem, or produce an
|
||||
# error. Any response containing the file content is a PR-B blocking failure.
|
||||
```
|
||||
|
||||
Additional criteria:
|
||||
- `SandboxManager.initialize()` is called once at startup (singleton shape follows `@anthropic-ai/sandbox-runtime` v0.0.52 `dist/sandbox/sandbox-manager.js` SandboxManager export, where `reset()` is a process-wide operation; see § 5 Open question 1 for the per-provider config concern still to be resolved)
|
||||
- p95 latency overhead of wrapping ≤ 200ms measured over 50 warm requests
|
||||
- `checkSandboxAvailability().available === true` reported in `/health.sandbox` after PR-B rolls out
|
||||
- All existing Suite 41 tests continue to pass (stream-json transport unaffected)
|
||||
|
||||
### 4.2 PR-C (codex.mjs wrap)
|
||||
|
||||
- `enableWeakerNestedSandbox: true` is set in the `SandboxManager.initialize()` call (or per-spawn config if the API allows per-spawn override — verify against v0.0.52 API)
|
||||
- `codex exec` inner bubblewrap nest still functions: a sandboxed codex invocation that reads from an allowed path succeeds
|
||||
- Analogous negative test: in-sandbox `cat /home/<user>/.olp/keys/...` MUST fail
|
||||
|
||||
### 4.3 PR-D (cloud deployment plan update)
|
||||
|
||||
- `docs/plans/cloud-deployment-family.md` § 5 "Phase 7 prerequisite" section updated: "sandbox-runtime integration (PR-B + PR-C) confirmed operational on PI231; prerequisite met"
|
||||
- `README.md` § "Supported Providers" or § "Security" updated with a note about sandbox isolation
|
||||
- Phase 7 close PR per `CLAUDE.md release_kit.phase_rolling_mode`
|
||||
|
||||
---
|
||||
|
||||
## 5. Open questions (to be resolved in PR-B)
|
||||
|
||||
1. **Singleton vs per-spawn initialization.** `SandboxManager` is a process-wide singleton (per the library's `reset()` being a global operation). The current design plan is one `initialize()` call at server startup with a union config covering all providers. If providers require different configs (e.g., different `denyRead` paths for anthropic vs codex), this may require a mutex approach or separate singleton instances. Decision reserved for PR-B.
|
||||
|
||||
2. **`SandboxManager.reset()` in tests.** The singleton means test suites that call `initialize()` must call `reset()` in their `after()` hooks. PR-B must add this discipline or tests will leak sandbox state across suites.
|
||||
|
||||
3. **MITM proxy and Claude CLI cert pinning.** The sandbox-runtime network bridge on Linux uses a local MITM proxy to intercept HTTPS traffic. If `claude` CLI pins certificates (e.g., for `api.anthropic.com`), HTTPS through the bridge may fail. PR-B must empirically verify this on PI231 before merging.
|
||||
|
||||
4. **macOS `sandbox-exec` profile content.** macOS uses a seatbelt (SBPL) profile, not bwrap. The profile must explicitly allow `network outbound "api.anthropic.com"` etc. The default profile may be too restrictive for the Claude CLI's OAuth refresh calls. PR-B must test macOS as well as Linux.
|
||||
|
||||
5. **`getDefaultWritePaths()` output.** The library exports `getDefaultWritePaths()` which returns the paths the sandbox always allows writing to. OLP's spawn directory may not be in that list — PR-B must verify the working directory is writable or pass it explicitly in `filesystem.allowWrite`.
|
||||
|
||||
---
|
||||
|
||||
## 6. Pitfalls inherited from the spike (binding warnings for PR-B/C authors)
|
||||
|
||||
These were confirmed empirically or inferred from the library source during the 2026-05-28 spike:
|
||||
|
||||
1. **Three OS deps, not one.** The npm package bundles nothing. Linux requires: `bubblewrap` (bwrap), `socat`, `ripgrep` (rg). All three. Missing even one → `checkDependencies()` returns errors → `wrapWithSandbox` will fail at runtime.
|
||||
|
||||
2. **Linux deny-paths are literal, not glob.** The library's `linuxGetMandatoryDenyPaths()` uses ripgrep to expand glob patterns to concrete paths before passing them to bwrap. But custom `filesystem.denyRead` entries that contain glob chars (`~/.ssh/*`) must be either expanded manually OR passed as the glob form (the library expands them if `rg` is available). The safe convention for PR-B: use absolute literal paths (e.g., `/home/<user>/.ssh`) rather than `~/`-prefixed or glob paths.
|
||||
|
||||
3. **`enableWeakerNestedSandbox: true` is required for codex.** Codex's `exec` subcommand spawns its own bubblewrap sandbox internally. Without `enableWeakerNestedSandbox`, the outer OLP sandbox blocks the inner codex sandbox from creating user namespaces. The flag loosens the outer sandbox's seccomp filter specifically to allow `clone(CLONE_NEWUSER)` — the inner sandbox then runs with reduced but non-zero isolation.
|
||||
|
||||
4. **`SandboxManager.reset()` is process-wide.** Calling `reset()` anywhere (including test teardown) clears the singleton config. Any concurrent in-flight spawn that still holds a reference to the old sandbox state will break. PR-B's design must either (a) initialize once at boot and never reset, or (b) use a mutex to prevent concurrent init/reset.
|
||||
|
||||
5. **MITM CA generation is async and expensive.** `SandboxManager.initialize()` generates a self-signed CA certificate for the MITM proxy on Linux. This takes ~100-500ms. Initialize at server startup, not per-request.
|
||||
|
||||
---
|
||||
|
||||
## 7. Authority citations
|
||||
|
||||
- **`@anthropic-ai/sandbox-runtime` v0.0.52** — https://github.com/anthropic-experimental/sandbox-runtime
|
||||
- `dist/sandbox/sandbox-manager.js` — `isSupportedPlatform()`, `checkDependencies()`, `SandboxManager` export shape
|
||||
- `dist/sandbox/linux-sandbox-utils.js` — `checkLinuxDependencies()`, `whichSync` usage, `enableWeakerNestedSandbox` rationale
|
||||
- `README.md` — installation prerequisites, platform support matrix
|
||||
|
||||
- **2026-05-28 PoC spike on PI231 (arm64 Debian Bookworm)** — report at `/tmp/sandbox-spike/report.md` on PI231. Key findings: dep install clean; `isSupportedPlatform()=true`; `checkDependencies()` errors on bwrap+socat+rg absence; three PoC scripts parked. Verdict YELLOW.
|
||||
|
||||
- **cc-mem incident memory 2026-05-27** — `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` § 3 (gap description), § 4 (prior-art search showing ecosystem hasn't solved multi-tenant fs/tool isolation).
|
||||
|
||||
- **OLP ADR 0009 Amendment 1 § Caveats #3** — "Sandbox-runtime still required for real multi-tenant deployment. Per the 2026-05-27 session prior-art search, Anthropic's official multi-tenant answer is `@anthropic-ai/sandbox-runtime` (OS-level isolation)."
|
||||
|
||||
- **`docs/plans/cloud-deployment-family.md` § 5** — sandbox is a hard prerequisite before any cloud deployment.
|
||||
|
||||
- **OLP ALIGNMENT.md** — PR-A is library/doctor/governance; it does not touch provider plugins, the entry surface, or the IR. The authority citation for the npm dep is the official sandbox-runtime repo URL + the spike report (not a provider CLI, not the OpenAI spec, not an existing ADR — this is a new dependency decision, which is the correct scope for ADR 0014).
|
||||
|
||||
- **Iron Rule 11 (Incremental Diff Review)** — splits non-trivial work into the minimum reviewable unit. The 4-PR split (A/B/C/D) is the direct application of this rule to the sandbox integration: each PR is independently reviewable, independently safe to land or revert, and corresponds to one logical layer.
|
||||
|
||||
---
|
||||
|
||||
## 8. Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Multi-tenant isolation at the OS level.** After PR-B+C land, each provider spawn runs inside a bubblewrap (Linux) or sandbox-exec (macOS) boundary. A prompt-injected `cat ~/.olp/keys/...` hits a kernel-level deny. Cross-client filesystem leakage is structurally prevented, not just mitigated by prompt engineering.
|
||||
|
||||
- **Cloud deployment unblocked.** `docs/plans/cloud-deployment-family.md` § 5 cites sandbox as the hard prerequisite for moving from family-LAN to cloud. PR-D closes this gate.
|
||||
|
||||
- **Observability from day one.** The `/health.sandbox` field makes the install state machine-readable. Any monitoring script or dashboard can tell whether sandbox isolation is active without SSH access.
|
||||
|
||||
- **Anthropic's official library.** Using `@anthropic-ai/sandbox-runtime` rather than a home-grown bwrap wrapper means OLP inherits Anthropic's tested integration patterns (deny-path expansion, MITM proxy, seccomp, macOS seatbelt profiles) rather than reinventing them. When the library updates, OLP upgrades via `npm update`.
|
||||
|
||||
### Negative
|
||||
|
||||
- **Three new OS-level dependencies.** `bubblewrap`, `socat`, and `ripgrep` must be installed on every host running OLP with sandbox isolation active. Absent these deps, sandbox is unavailable (but OLP continues to function without isolation — degraded security, not degraded functionality). The `/health.sandbox.available` field makes this state explicit.
|
||||
|
||||
- **p95 latency overhead.** The spike did not measure sandbox wrapping overhead directly (blocked on apt install). Expected overhead per the sandbox-runtime README: ~100-200ms for sandbox initialization amortized over the process lifetime (one-time at startup); per-spawn overhead is the namespace clone + filesystem mount overhead, typically <50ms on modern kernels. PR-B's acceptance criteria gates on ≤200ms p95 overhead over 50 warm requests.
|
||||
|
||||
- **Codex inner-sandbox degradation.** `enableWeakerNestedSandbox: true` loosens the outer OLP sandbox's seccomp filter to allow `clone(CLONE_NEWUSER)`. The codex inner sandbox still runs with meaningful isolation (its own namespace, its own deny-list), but the combined depth of protection is less than ideal compared to a world where codex didn't self-sandbox.
|
||||
|
||||
- **Library is experimental.** The `anthropic-experimental` org signals this is not a production-stable API. The version pin (`^0.0.52`) provides a minor-range buffer but the API surface may change. If the library is deprecated or the API breaks, OLP's fallback is to remove the sandbox wrapping (reverting PRs B-D) until a replacement path is found. This is acceptable at family scale — security degradation is not a service outage.
|
||||
|
||||
### Reversibility
|
||||
|
||||
- **PR-A** is trivially reversible: `npm uninstall @anthropic-ai/sandbox-runtime` + delete `lib/sandbox/doctor.mjs` + revert server.mjs and CHANGELOG changes. No production behavior changes.
|
||||
- **PR-B/C** are reversible by removing the `SandboxManager.wrapWithSandbox` call from each provider's `spawn()` method. The spawn falls back to the current unsandboxed path.
|
||||
- **PR-D** is a documentation update; reverting it is a docs-only change.
|
||||
|
||||
---
|
||||
|
||||
## Status transitions
|
||||
|
||||
- 2026-05-28 — Created. Status: Accepted for PR-A scope. PR-B/C/D pending operational prereqs.
|
||||
- 2026-05-28 — PR-A shipped (commit `07d9c8a`).
|
||||
- 2026-05-28 — PR-B implementation shipped (commit chain `d0dcd28` → `2864275` → `497b255` → `b1e24b7` → `3551921`). Status: shipped pending PI231 Suite 44 validation + HTTP-path activation debug. `OLP_SANDBOX_DISABLED=1` emergency disable installed (b1e24b7) because the in-process MITM proxy interaction with OLP's HTTP request handler suppressed claude stdout on the HTTP path while the same wrap script produced output when invoked directly from a manual shell. Prod is currently running with `OLP_SANDBOX_DISABLED=1` set.
|
||||
- 2026-05-29 — **Amendment 1 — supersede PR-B outer-bwrap with ephemeral-home + per-provider ISOLATION contract.** See § Amendment 1 below. PR-B's `lib/sandbox/manager.mjs` outer-bwrap implementation is archived to branch `phase-7-pr-b-outer-bwrap-snapshot` and superseded; PR-C is reframed as inner-sandbox preservation under the new architecture; PR-D is reframed as the README "Security Model" section. `lib/sandbox/doctor.mjs` is preserved unchanged.
|
||||
|
||||
---
|
||||
|
||||
# Amendment 1 — Supersede PR-B outer-bwrap with ephemeral-home + per-provider contract (2026-05-29)
|
||||
|
||||
- **Date:** 2026-05-29
|
||||
- **Status:** Accepted (governance only — the implementation refactor lands in subsequent PRs per ALIGNMENT.md Rule 1 / Iron Rule 11)
|
||||
- **Author:** project maintainer (with AI drafting assistance)
|
||||
- **Reviewer:** independent fresh-context reviewer per Iron Rule 10 — pending at this draft
|
||||
- **Scope:** This amendment supersedes the implementation strategy of PR-B (the outer-bubblewrap-wrap of `claude` CLI shipped in commits `d0dcd28` → `b1e24b7`). It does NOT supersede the multi-tenant security gap analysis in § 1 of the original ADR, nor the four-tier authority citation list (§ 7), nor `lib/sandbox/doctor.mjs` (preserved unchanged). It DOES supersede the PR-B implementation, the PR-C scope ("wrap codex spawn in the same outer-bwrap pattern with `enableWeakerNestedSandbox`"), and PR-D's framing as "documentation update for cloud rollout unblock".
|
||||
|
||||
---
|
||||
|
||||
## A1.1 — Why the substitution is forced (the four forcing reasons)
|
||||
|
||||
PR-B as designed (outer-bwrap wrapping of the `claude` CLI spawn, with the OLP server initializing `SandboxManager` once at boot and every spawn routed through `wrapWithSandbox`) was shipped on 2026-05-28 and disabled on the same day via the `OLP_SANDBOX_DISABLED=1` env-var gate after the HTTP-path activation regression appeared on PI231 (the manual-shell wrap produced claude stdout; the OLP HTTP-request-handler wrap produced none). The 2026-05-28/2026-05-29 follow-up investigation found that the HTTP-path failure was not the whole story — even if the in-process MITM proxy lifecycle issue were debugged, four independent and load-bearing reasons forced the architecture away from outer-bwrap entirely. Each is cited to its primary authority below.
|
||||
|
||||
### A1.1.1 — Forcing reason #1: Anthropic's stated design intent for `@anthropic-ai/sandbox-runtime`
|
||||
|
||||
The PR-B design used `@anthropic-ai/sandbox-runtime` to wrap the `claude` CLI from the outside. Anthropic's published design intent for the library is the opposite direction of containment: the library is for sandboxing what Claude Code itself *triggers* (tool calls, MCP servers, sub-processes spawned during model execution), not for wrapping Claude Code from outside.
|
||||
|
||||
**Primary citation:** https://www.anthropic.com/engineering/claude-code-sandboxing — "Claude Code sandboxing" engineering blog. The post describes Claude Code's *internal* use of the sandbox-runtime library: the model emits a `tool_use` Bash call → Claude Code wraps the resulting `/bin/sh -c <…>` in a sandbox via `SandboxManager.wrapWithSandbox()` before spawning. The blog also notes the library "can be used to sandbox arbitrary processes, agents and MCP servers" — i.e., it is general-purpose, not Claude-Code-internal-only. **Our reading:** Anthropic's documented and demonstrated usage is *inner-wrap by Claude Code*; outer-wrap of `claude` itself is not documented in the blog and not shown in the post's example invocations. **This is a project-design judgment based on the absence of outer-wrap precedent, not a "don't do this" statement from Anthropic.** The architectural concerns enumerated below (MITM proxy lifecycle, semver leverage) stand on their own merits regardless of how Anthropic frames the library's intended usage.
|
||||
|
||||
**What this means for PR-B's design:** wrapping `claude` from outside with the same library is "out-of-distribution" usage. The library was not designed for, tested against, or documented for the outer-wrap case. Two concrete consequences observed in PR-B:
|
||||
|
||||
1. **MITM proxy collision.** The library starts a per-process local MITM proxy on Linux to inspect HTTPS traffic for allowlisted domains. When the same library is invoked again from inside the sandboxed process (e.g., for any sub-spawn `claude` might do), a second MITM proxy attempt collides. The PR-B implementation never reached this case because it disabled before tripping it, but the architecture invites the collision.
|
||||
2. **Inner-sandbox conflict** (see A1.1.3 for codex, but the principle applies generally). Any CLI that itself uses the same library to sandbox its own tool calls is *expected* by Anthropic to be the *holder* of the sandbox, not the *content* of one. The library's `enableWeakerNestedSandbox` option exists precisely to acknowledge this — but only as a partial mitigation.
|
||||
|
||||
The Anthropic design-intent reason is not a "won't work" reason. The PR-B outer-bwrap path did work for the smoke case (manual-shell invocation produced output). The reason is a *don't-do-this* reason: OLP would be the only known user of the library in the outer-wrap configuration, taking on the maintenance burden of a usage pattern Anthropic doesn't test, doesn't document, and doesn't owe semver discipline for. The library is `^0.0.52`. A future minor version bump could break OLP's outer-wrap path without warning. Aligning OLP's use of the library with Anthropic's documented design intent restores semver leverage.
|
||||
|
||||
### A1.1.2 — Forcing reason #2: The `~/.claude.json` upstream "closed as not planned" — permanent maintenance treadmill for outer-bwrap
|
||||
|
||||
Anthropic's `claude` CLI persists session state (most notably MRU project list, last-used model, telemetry session id) to `~/.claude.json` using a **non-atomic** write pattern: write-in-place rather than write-temp-then-rename. When `~/.claude.json` is on a read-only mount (e.g., `--ro-bind ~/.claude` in a bwrap config), the CLI either silently degrades (some state lost) or outright fails depending on which write path is hit on a given invocation.
|
||||
|
||||
The OLP outer-bwrap path in PR-B mounted `~/.claude` read-only initially, hit this exact issue (the 2026-05-28 fold-in commit `2864275` "allow read ~/.claude + skip wrap under test mock" added the read allowance), and then on the further `b1e24b7` commit the operator had to widen the mount again. The fold-in commit message explicitly notes: "~/.claude is NOT in denyRead. It contains the spawn's own OAuth credentials — claude CLI must read its own auth file to function. Denying read here causes 'Not logged in' failures."
|
||||
|
||||
**Primary citation:** https://github.com/anthropics/claude-code/issues/29250 — upstream issue requesting atomic-write semantics for `~/.claude.json`. Status: closed `not_planned` by GitHub's inactivity bot on 2026-03-31 (`closed_by: github-actions[bot]`, comment text: "Closing for now — inactive for too long"), labeled `duplicate` and `stale`. **No upstream Anthropic maintainer comment in the thread articulates any policy rationale** — the only non-bot comments are the reporter's own follow-ups and a third-party-tool plug. The closure does NOT establish "won't fix" as Anthropic policy; at most it establishes that the issue has not received maintainer engagement and that Anthropic considers it a duplicate of some other tracked report. The forcing argument for ephemeral-home over outer-bwrap rests on the architectural-cost analysis (the maintenance-treadmill description below), not on an alleged upstream policy posture.
|
||||
|
||||
**What this means for the outer-wrap maintenance treadmill:** Every future addition of `claude`-CLI-owned state files (telemetry, cache directories, session locks, MCP registration files, etc.) is, by upstream policy, free to use any write pattern the maintainers prefer. The outer-bwrap pattern requires OLP to track each of these additions and add corresponding `--ro-bind` / `--rw-bind` / write-allowlist entries — forever — because the CLI does not give OLP an enumerable contract surface for "files I will write to." The maintainer-time cost is a permanent recurring tax.
|
||||
|
||||
A non-outer-wrap approach that gives `claude` a fresh, ephemeral home directory inverts this: `claude` is free to invent any state file under its $HOME with any write pattern it chooses; OLP never tracks the list. The treadmill goes away. This is the load-bearing case for Solution 1 even setting aside the codex inner-sandbox issue below.
|
||||
|
||||
### A1.1.3 — Forcing reason #3: Codex inner-bwrap conflict (multi-provider forcing function)
|
||||
|
||||
The PR-C plan in the original ADR was to wrap the `codex` spawn in the same outer-bwrap pattern as PR-B, with `enableWeakerNestedSandbox: true` set on the `SandboxManager.initialize()` call to allow codex's own internal bubblewrap sandbox to function inside OLP's outer bubblewrap sandbox.
|
||||
|
||||
Empirical investigation (2026-05-29 PI231 prep — to be confirmed in Task #4) and published codex CLI behaviour both indicate this nested-sandbox path is structurally fragile:
|
||||
|
||||
**Primary citation:** https://github.com/openai/codex/issues/16018 — upstream codex CLI issue. The issue body documents that codex's bwrap-based default sandbox **fails outright** in environments lacking unprivileged user namespaces — the reporter quotes the error `bwrap: No permissions to create new namespace, likely because the kernel does not allow non-privileged user namespaces`. The issue is a **feature request by the reporter** asking codex to "suggest or automatically fall back to an alternative supported backend when available"; **the issue body itself does NOT contain the string `danger-full-access` and does NOT document an existing automatic fallback to it**. The codex `--sandbox danger-full-access` mode is a documented *manual* opt-out (https://developers.openai.com/codex/concepts/sandboxing § "Sandboxing modes"). Whether codex automatically degrades into it under nested-bwrap failure — or whether the spawn aborts outright — is an empirical question slated for Task #4 PI231 spike verification.
|
||||
|
||||
In other words: wrapping codex in OLP's outer bwrap, *if* the outer bwrap is configured with sufficient capability to allow the inner clone, requires giving the outer sandbox more capability than the security boundary should grant. *If* it is configured to a tighter, safer capability set, codex's inner-bwrap initialization fails (the documented failure mode per the linked issue). Whether codex then aborts the spawn or silently degrades to `danger-full-access` is empirically open (Task #4); either outcome is undesirable. The strict-additive-isolation invariant (outer-bwrap + inner-bwrap = composed isolation) does not hold for codex under this configuration: either OLP gives up outer-isolation strength to admit the inner clone, or codex's inner isolation breaks in some manner.
|
||||
|
||||
**What this means as a multi-provider forcing function:** OLP is by constitution (ADR 0001 § Mission) a multi-provider proxy. The outer-bwrap architecture cannot cover codex without a security regression. The structural response is to abandon outer-bwrap as the foundational architecture and adopt a strategy that is *compatible* with each provider's own native isolation (claude's lack of inner sandbox vs codex's `--sandbox read-only` inner sandbox). This is what Solution 1 does — see A1.2 below.
|
||||
|
||||
### A1.1.4 — Forcing reason #4: `CODEX_HOME` exists and is the documented relocation lever
|
||||
|
||||
The "ephemeral home directory per spawn" component of Solution 1 (A1.2 Layer 1) only works if each provider CLI offers a documented mechanism for relocating its state directory away from the default `$HOME` location. For `claude`, the standard `HOME` env var works (the CLI reads `~/.claude` as `$HOME/.claude`, and changing `HOME` relocates the lookup). For `codex`, the equivalent lever is the `CODEX_HOME` env var.
|
||||
|
||||
**Primary citation:**
|
||||
- https://developers.openai.com/codex/config-reference — OpenAI's published codex CLI configuration reference. The page documents `CODEX_HOME` in 2 places (verified by independent fetch 2026-05-29): as the root of the per-profile config path (`$CODEX_HOME/profile-name.config.toml`) and as the default log directory base (`$CODEX_HOME/log`). The variable is the documented relocation lever for the codex state, configuration, and authentication directory away from the default `~/.codex`.
|
||||
- Secondary corroboration:
|
||||
- https://developers.openai.com/codex/auth/ — OpenAI's published codex CLI authentication reference. The page documents `CODEX_HOME` in 2 places (verified by independent fetch 2026-05-29), both in the credential-storage section: "file stores credentials in `auth.json` under `CODEX_HOME` (defaults to `~/.codex`)." Confirms `CODEX_HOME` is the credential-directory base.
|
||||
- https://codex.danielvaughan.com/2026/04/08/codex-cli-configuration-reference/ — third-party reference page that mirrors the documented behaviour, used as cross-reference for the reachability check.
|
||||
|
||||
**What this means for Solution 1 feasibility:** All three Tier-D providers have a documented one-env-var relocation lever:
|
||||
- claude via `HOME` (POSIX convention)
|
||||
- codex via `CODEX_HOME` (citations above)
|
||||
- mistral via `VIBE_HOME` per https://docs.mistral.ai/mistral-vibe/terminal/configuration (3 occurrences verified 2026-05-29, including the canonical `export VIBE_HOME="/path/to/custom/vibe/home"` example and an enumeration of files/directories `VIBE_HOME` affects).
|
||||
|
||||
The ephemeral-home approach is implementable today; it does not require upstream changes from any of Anthropic, OpenAI, or Mistral. Task #4 PI231 spike verifies *observed CLI behaviour* matches *documented behaviour* for each provider — this is verification-grade follow-up, not authority-pin work.
|
||||
|
||||
---
|
||||
|
||||
## A1.2 — The substitute architecture: per-spawn ephemeral home + per-provider ISOLATION contract
|
||||
|
||||
The new architecture is layered. Each layer addresses a distinct attack surface, and each layer is independently reasoned about, independently reviewable, and independently revertible. The four layers, in order of containment depth:
|
||||
|
||||
### A1.2.1 — Layer 1: Per-spawn ephemeral home directory
|
||||
|
||||
Every uncached `/v1/chat/completions` request (per-`keyId`, per-`reqId`) provisions a fresh ephemeral home directory at `/tmp/olp-spawn/<keyId>/<reqId>/home/`. The directory is created on the spawn path and torn down (best-effort) on response completion. The spawn process gets this directory passed in via a per-provider env-var override:
|
||||
|
||||
- **anthropic** (`claude` CLI): `HOME=/tmp/olp-spawn/<keyId>/<reqId>/home`. The CLI's `~/.claude.json` and `~/.claude/` state writes go to the ephemeral location. No cross-request, no cross-tenant carry-over.
|
||||
- **openai** (`codex` CLI): `CODEX_HOME=/tmp/olp-spawn/<keyId>/<reqId>/home/.codex`. Codex's `~/.codex` state, auth artifacts, and config files go to the ephemeral location.
|
||||
- **mistral** (`vibe` CLI): `VIBE_HOME=/tmp/olp-spawn/<keyId>/<reqId>/home/.vibe` per https://docs.mistral.ai/mistral-vibe/terminal/configuration (documented env var, 3 occurrences verified at amendment time). Vibe's `~/.vibe/` state — `.env`, `agents/`, `prompts/`, `skills/`, `tools/`, `config.toml` — goes to the ephemeral location. Task #4 PI231 spike verifies observed CLI behaviour matches the documented contract.
|
||||
|
||||
Layer 1 provides:
|
||||
- **No cross-tenant state carry-over** at the filesystem level. Two clients invoking anthropic concurrently get two separate `$HOME` directories; the CLI cannot read the other's `~/.claude.json`, recent-projects list, or session state.
|
||||
- **No accumulation of stale state** across requests. The MRU project list does not grow without bound. The telemetry session id is fresh per request.
|
||||
- **No outer-wrap maintenance treadmill.** When `claude` invents a new state file under `~/.claude.foo.json` next quarter, OLP does not need to update a `--ro-bind` list. The new file lives in the ephemeral home and goes away with the request.
|
||||
|
||||
What Layer 1 does NOT provide:
|
||||
- It does not protect against the CLI walking *out of* its $HOME to read other paths (e.g., a model emitting a `Read` tool call on `/etc/passwd` or `~/.ssh/id_rsa`). For that protection, Layers 3 and 4 are needed.
|
||||
|
||||
### A1.2.2 — Layer 2: Symlinked credential files into the ephemeral home
|
||||
|
||||
A fresh `$HOME` is empty. The CLI needs its OAuth credentials, API key, or equivalent auth artifact to function. Layer 2 provisions these by reading the operator-pinned credential location and symlinking the relevant file(s) into the ephemeral home at the location the CLI expects.
|
||||
|
||||
Each provider plugin declares its credential paths in the ISOLATION block (see ADR 0002 Amendment pending). The runtime spawn pipeline reads this declaration, walks the list, and symlinks each entry from its real location (under the operator's real `$HOME`) into the ephemeral home. The symlinks are file-level, not directory-level, so the CLI sees its credential file but does not see the rest of the operator's `~/.claude/` or `~/.codex/` tree.
|
||||
|
||||
Example (anthropic):
|
||||
- Real: `~/.claude/.credentials.json` (operator's actual OAuth credential)
|
||||
- Ephemeral: `/tmp/olp-spawn/<keyId>/<reqId>/home/.claude/.credentials.json` (symlink → real)
|
||||
|
||||
Example (codex):
|
||||
- Real: `~/.codex/auth.json`
|
||||
- Ephemeral: `/tmp/olp-spawn/<keyId>/<reqId>/home/.codex/auth.json` (symlink → real)
|
||||
|
||||
Layer 2 provides:
|
||||
- **Credential availability** without granting visibility into other state under the same provider directory.
|
||||
- **A narrow declared surface.** The provider plugin enumerates exactly which files matter. New CLI state files that are not declared do not get symlinked, and the CLI re-initializes them in the ephemeral home (which is exactly the Layer 1 behaviour).
|
||||
|
||||
What Layer 2 does NOT provide:
|
||||
- It does not protect against the CLI walking out of its $HOME (see Layer 3).
|
||||
- It does not protect against the CLI's tool-use surface reading the symlink target's *containing directory* if the model emits a `Read` tool call with an absolute path that resolves around the symlink. For that, Layer 3 + Layer 4.
|
||||
|
||||
### A1.2.3 — Layer 3: Optional `sandbox-runtime` per-call `customConfig` for non-$HOME read protection
|
||||
|
||||
For providers whose own inner sandbox does NOT exist or does not cover the OLP threat model (the `claude` CLI today is the leading example — claude has no inner sandbox; codex has `--sandbox read-only` by default but the protection scope differs), Layer 3 wraps the spawn in `@anthropic-ai/sandbox-runtime`'s `SandboxManager.wrapWithSandbox()` *per-call* with a `customConfig` argument tailored to the per-spawn ephemeral home.
|
||||
|
||||
The key architectural difference vs PR-B's outer-wrap:
|
||||
- PR-B initialized `SandboxManager` once at server boot with a *global* config covering all providers.
|
||||
- Layer 3 calls `wrapWithSandbox()` *per spawn* with a *per-spawn* `customConfig` that names the ephemeral home as the allow-read root.
|
||||
|
||||
The per-call `customConfig` shape:
|
||||
|
||||
```javascript
|
||||
{
|
||||
network: { allowedDomains: provider.ISOLATION.allowedDomains },
|
||||
filesystem: {
|
||||
denyRead: [
|
||||
// Operator's real $HOME — sandbox cannot read OTHER clients' OLP keys,
|
||||
// operator's SSH identity, other providers' tokens, etc.
|
||||
operatorHome,
|
||||
// Operator's known sensitive directories (defensive even though they
|
||||
// are already under operatorHome) — declared so a future refactor that
|
||||
// moves the operator home does not regress this protection.
|
||||
`${operatorHome}/.ssh`,
|
||||
`${operatorHome}/.gnupg`,
|
||||
`${operatorHome}/.olp`,
|
||||
],
|
||||
// Layer 1 ephemeral home is the allow-read root for this spawn.
|
||||
// Layer 2 symlinked credentials live inside, so credential access works.
|
||||
allowRead: [ephemeralHomeForThisSpawn],
|
||||
allowWrite: [ephemeralHomeForThisSpawn, '/tmp'],
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Layer 3 is invoked **only when** the provider's `ISOLATION.hasInnerSandbox === false`. For providers with their own inner sandbox (codex via `--sandbox read-only`), Layer 3 is skipped to avoid the nested-sandbox conflict (A1.1.3).
|
||||
|
||||
Layer 3 provides:
|
||||
- **OS-level deny of reads outside the ephemeral home and OLP-permitted paths.** A prompt-injected `cat /home/<operator>/.olp/keys/owner-key.json` or `cat /home/<operator>/.ssh/id_ed25519` hits a syscall-level deny.
|
||||
- **Per-spawn (not per-process) configuration.** Each request gets a fresh sandbox scope. Two concurrent spawns do not share a sandbox; the MITM-proxy collision and singleton-config-mutation hazards from PR-B disappear.
|
||||
|
||||
What Layer 3 does NOT provide:
|
||||
- It does not protect against the CLI's *own* tool-use surface emitting destructive shell commands within the allowed write zones. For that, Layer 4.
|
||||
- Per-call `wrapWithSandbox()` has higher per-request latency than PR-B's once-at-boot pattern. The amortization budget is recovered by Layer 1's $HOME-as-cwd discipline keeping the sandbox config small and by ripgrep-based glob expansion being avoided (Layer 3 uses absolute literal paths throughout).
|
||||
|
||||
### A1.2.4 — Layer 4: Provider-specific tool hardening already in place
|
||||
|
||||
This is already-shipped work, re-affirmed here as part of the layered model:
|
||||
|
||||
- **anthropic Phase 6c `--system-prompt`** (commits `97e7d16` + fold-in `65f945c`). The system prompt is fully replaced at every spawn, suppressing the default tool descriptions that Claude Code would otherwise inject. Without tool descriptions, the model is highly unlikely to emit `tool_use` for `Bash`, `Read`, etc. even under prompt injection. See cc-mem `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` § 5.
|
||||
- **codex `--sandbox read-only` default.** OLP's codex provider spawn passes `--sandbox read-only` as a fixed flag. Codex's own inner sandbox provides read-only-by-default tool isolation. The provider's ISOLATION block declares `hasInnerSandbox: true` so Layer 3 is correctly skipped.
|
||||
- **mistral.** TBD per Task #4 — the mistral provider's tool surface and inner-sandbox status need to be characterized.
|
||||
|
||||
Layer 4 provides:
|
||||
- **Reduction of the *probability* of tool emission.** Layer 4 does not depend on OS-level enforcement; it works at the prompt layer. It is the cheap, fast, first-line defense. Layers 1–3 are the structural fallback when prompt-layer defenses are bypassed.
|
||||
|
||||
---
|
||||
|
||||
## A1.3 — The provider ISOLATION contract (named here; specified in ADR 0002 Amendment N)
|
||||
|
||||
Each provider plugin declares an `ISOLATION` block on its module export. The fields are:
|
||||
|
||||
| Field | Type | Meaning |
|
||||
|---|---|---|
|
||||
| `ephemeralEnvOverrides` | `(spawnCtx) => Record<string, string>` | Returns the env-var map to set for this spawn, given the spawn context (ephemeral home path, keyId, reqId). For anthropic: `{ HOME: spawnCtx.ephemeralHome }`. For codex: `{ CODEX_HOME: spawnCtx.ephemeralHome + '/.codex' }`. |
|
||||
| `credentialMounts` | `{ realPath: string, ephemeralPath: string }[]` | List of credential files to symlink from real → ephemeral. For anthropic: `[{ realPath: '~/.claude/.credentials.json', ephemeralPath: '.claude/.credentials.json' }]`. Provider declares; runtime symlinks. |
|
||||
| `hasInnerSandbox` | `boolean` | If true, Layer 3 is skipped to avoid nested-sandbox conflict. codex: true. anthropic: false. |
|
||||
| `crossTenantReadProtection` | `'tool-suppression' \| 'inner-sandbox' \| 'none'` | Self-declared label for what layer is providing the read-protection. Used by `/health.sandbox` to report the protection posture per provider. **The canonical enum is defined in ADR 0002 Amendment 9 § 5; this row mirrors it.** |
|
||||
| `recommendedDeploymentTier` | `'shared-os-user' \| 'per-os-user' \| 'separate-vm'` | Deployment tier the provider's current isolation posture is rated for. ADR 0006 risk-tier integration. **The canonical enum is defined in ADR 0002 Amendment 9 § 6; this row mirrors it.** |
|
||||
|
||||
**This amendment names the contract but does NOT specify its full validation, lifecycle, or test discipline.** Those land in **ADR 0002 Amendment (pending)** — the Provider contract amendment that ratifies `ISOLATION` as a required field, defines `validateProvider`'s checks on it, and documents how `lib/providers/base.mjs` enforces declaration. Until that ADR amendment lands, the ISOLATION block is a forward-looking contract; the implementation refactor (Tasks #5–#8) is gated on the ADR 0002 amendment landing first.
|
||||
|
||||
Cross-reference: see ADR 0002 § Amendments for the pending Amendment N that codifies the ISOLATION block contract.
|
||||
|
||||
---
|
||||
|
||||
## A1.4 — Revised PR plan
|
||||
|
||||
The original ADR's four-PR split (PR-A / PR-B / PR-C / PR-D) is restated as follows. PR-A is unchanged from its as-shipped state.
|
||||
|
||||
| PR | Original scope | Amendment 1 scope | Status |
|
||||
|---|---|---|---|
|
||||
| **PR-A** | npm dep + `lib/sandbox/doctor.mjs` + `/health.sandbox` | **Unchanged.** Doctor preserved; `/health.sandbox` field preserved. | ✅ Shipped (commit `07d9c8a`) |
|
||||
| **PR-B** | Outer-bwrap wrap of anthropic spawn at boot-singleton level | **Superseded by Amendment 1.** Implementation archived to branch `phase-7-pr-b-outer-bwrap-snapshot`. New scope: refactor `lib/sandbox/manager.mjs` to the Layer 1 + Layer 2 + Layer 3 architecture (Tasks #5, #8). | ⛔ Superseded |
|
||||
| **PR-C** | Outer-bwrap wrap of codex spawn with `enableWeakerNestedSandbox: true` | **Superseded by Amendment 1.** Codex isolation now flows via Layer 1 ephemeral `CODEX_HOME` + Layer 4 `--sandbox read-only`. Layer 3 deliberately skipped (`hasInnerSandbox: true`). Codex-specific PR (Task #7) lands the ISOLATION block declaration; no outer-wrap. | ⛔ Superseded |
|
||||
| **PR-D** | Documentation update for cloud rollout unblock | **Reframed.** New scope: README "Security Model" section documenting the four-layer architecture, the deployment-tier mapping, and what the operator gets vs does not get at each tier. Task #10. | ♻ Reframed |
|
||||
|
||||
The new effective PR-list:
|
||||
|
||||
- **PR-B' (Refactor):** `lib/sandbox/manager.mjs` rewritten to expose `prepareIsolatedEnvironment(spawnCtx)` (Layer 1 + Layer 2) and `maybeWrapForReadProtection(spawnCtx, command)` (Layer 3 conditional). The `OLP_SANDBOX_DISABLED=1` env-var gate is preserved for 1-2 releases as belt-and-suspenders, then removed. Singleton bootstrap pattern is removed (per-spawn config eliminates the singleton's reason to exist).
|
||||
- **PR-C' (Wiring + Anthropic ISOLATION):** `server.mjs` calls `prepareIsolatedEnvironment` on the spawn path; `lib/providers/anthropic.mjs` declares its ISOLATION block (Task #6); negative-test confirmation via Task #9 PI231 E2E.
|
||||
- **PR-D' (Codex ISOLATION):** `lib/providers/codex.mjs` declares its ISOLATION block (Task #7); `hasInnerSandbox: true` skips Layer 3; codex inner sandbox preserved unmolested. Verified on PI231 (Task #9).
|
||||
- **PR-E' (README + Phase 7 close):** README "Security Model" section (Task #10) + `docs/plans/cloud-deployment-family.md` § 5 update + Phase 7 close per `CLAUDE.md release_kit.phase_rolling_mode`.
|
||||
|
||||
The original PR sequence's load-bearing security gate (the negative test "in-sandbox `cat ~/.olp/keys/...` MUST fail") remains the acceptance criterion for the security-bearing PRs in the new sequence. The test itself transfers; only the wrap mechanism changes.
|
||||
|
||||
---
|
||||
|
||||
## A1.5 — What survives from PR-B (preserved)
|
||||
|
||||
The following artifacts from the original PR-B implementation are preserved through Amendment 1:
|
||||
|
||||
1. **`lib/sandbox/doctor.mjs` — preserved unchanged.** Pure preflight is still useful: it tells the operator whether the npm package is installed, whether the OS deps are present, and whether the platform is supported. Even though the architecture no longer relies on a boot-time `SandboxManager.initialize()`, the `/health.sandbox` field consumers (dashboard, monitoring scripts) expect a stable shape. Doctor stays.
|
||||
2. **`/health.sandbox` field — preserved.** Shape adjusts slightly: the `active` boolean shifts meaning from "SandboxManager.initialize() succeeded" (PR-B) to "Layer 3 is operational for at least one provider whose `ISOLATION.hasInnerSandbox === false`" (Amendment 1). The field's name and JSON path stay the same so downstream consumers (dashboard, Hermes self-check, monitoring) do not break. The per-provider isolation posture is exposed via a new `/health.sandbox.providers[<name>].crossTenantReadProtection` subfield sourced from each ISOLATION block.
|
||||
3. **`@anthropic-ai/sandbox-runtime` npm dependency — preserved.** Layer 3 still uses the library, but via per-call `wrapWithSandbox()` with `customConfig`, not via a once-at-boot `SandboxManager.initialize()`. The dependency line in `package.json` stays.
|
||||
4. **The four authority citations in original § 7 — preserved.** The library URL, the spike report URL, the cc-mem incident URL, and the cloud deployment plan URL are unchanged. Amendment 1 *adds* the four new primary citations enumerated in § A1.1 above.
|
||||
5. **The `OLP_SANDBOX_DISABLED=1` env-var gate — preserved for 1-2 releases, then removed.** Documented in A1.6 below.
|
||||
|
||||
---
|
||||
|
||||
## A1.6 — What disappears from PR-B (superseded)
|
||||
|
||||
The following artifacts are removed by the PR-B' refactor (Task #5):
|
||||
|
||||
1. **Outer-bwrap wrapping of the `claude` spawn.** The bwrap wrap goes away. `claude` runs directly (without bwrap shell-wrap) with its `HOME` set to the ephemeral location. Layer 3 wraps the *sub-spawn* shell when it is invoked, not the `claude` process itself.
|
||||
2. **EROFS-driven mount patches.** The fold-in commit `2864275` ("allow read ~/.claude") and the subsequent `~/.claude` rw promotion (Task #5 was filed against this) were both consequences of trying to outer-bwrap a CLI that writes non-atomically to its `$HOME`. Solution 1 gives the CLI its own fresh `$HOME` and the entire mount-patch problem disappears. Task #5 ("allowWrite ~/.claude rw promotion fix") is closed as obsolete by this amendment.
|
||||
3. **Boot-time `SandboxManager.initialize()` call.** Removed entirely. The library is loaded lazily per-spawn (with import memoization for performance — the import itself is cached after the first call; only the `wrapWithSandbox()` call is per-spawn).
|
||||
4. **The singleton config-at-boot pattern.** Removed. The `_initConfig`, `_active`, `_initialized` module-level variables in `lib/sandbox/manager.mjs` no longer represent a global sandbox state; the only module-level state retained is the import cache for the library.
|
||||
5. **The MITM proxy CA cert generated once at boot.** Per-call `wrapWithSandbox()` may regenerate per call (TBD on library v0.0.52 behaviour — Task #4 verifies). If per-call regeneration is too expensive, an alternative is a per-process MITM CA cached at first-use; the implementation detail is reserved to PR-B'.
|
||||
6. **The `enableWeakerNestedSandbox: true` flag plan.** Removed. Codex isolation does not run inside an OLP outer sandbox at all. `enableWeakerNestedSandbox` is irrelevant to Amendment 1's architecture.
|
||||
|
||||
### A1.6.1 — The `OLP_SANDBOX_DISABLED=1` env-var gate
|
||||
|
||||
The env-var gate added in commit `b1e24b7` ("add OLP_SANDBOX_DISABLED=1 env-var emergency disable") is preserved through the Amendment 1 refactor as belt-and-suspenders. Its semantics under Amendment 1:
|
||||
|
||||
- **PR-B world (current main, with the gate set in prod):** the gate skips `SandboxManager.initialize()` at boot. Prod is currently running with the gate set, which means PR-B's outer-bwrap path is not active — Layer 3 protection is also not active.
|
||||
- **Amendment 1 world (after PR-B' lands):** the gate skips Layer 3's per-call `wrapWithSandbox()` and reverts each spawn to a Layer 1 + Layer 2 + Layer 4 configuration. The CLI still gets an ephemeral `$HOME` with symlinked credentials, still gets the `--system-prompt` tool-description suppression for anthropic, still gets `--sandbox read-only` for codex. What is given up is the OS-level deny of reads outside the ephemeral home. This is a *meaningful* but not *catastrophic* degradation — the prompt-layer defense remains, and Layer 1's $HOME isolation still prevents the most common cross-tenant accident path.
|
||||
- **Sunset:** the gate is preserved for **1-2 releases** after PR-B' ships to give the operator a fast escape hatch if the Layer 3 per-call wrap regresses in production. After two clean releases with no operator escalation, the gate is removed in a subsequent ADR amendment or a clean PR citing this section as authority for the removal.
|
||||
|
||||
The gate's behaviour is documented in README's Security Model section per PR-D' (Task #10).
|
||||
|
||||
---
|
||||
|
||||
## A1.7 — Reversibility
|
||||
|
||||
Amendment 1 is reversible at the implementation layer:
|
||||
|
||||
- **PR-B' refactor** is reversible by `git revert` of the refactor commit + restoring the snapshot from `phase-7-pr-b-outer-bwrap-snapshot`. The archive branch is pushed and persistent at:
|
||||
https://github.com/dtzp555-max/olp/tree/phase-7-pr-b-outer-bwrap-snapshot
|
||||
- **The `@anthropic-ai/sandbox-runtime` dependency** stays in `package.json`, so reverting does not require an `npm install`.
|
||||
- **The `lib/sandbox/doctor.mjs` module** is unchanged across the refactor, so reverting does not affect `/health.sandbox` shape.
|
||||
|
||||
Amendment 1 itself, as a governance artifact, is reversible by a subsequent superseding amendment if the empirical foundation it rests on changes (e.g., if Anthropic publishes guidance endorsing outer-wrap use of `sandbox-runtime` and adds a contract for `~/.claude.json` write paths). ALIGNMENT.md § "Amendment Procedure" applies: such a future amendment would need to cite the new evidence.
|
||||
|
||||
The archive-branch retention policy: the snapshot branch is kept indefinitely (no auto-delete) so a future maintainer investigating outer-bwrap-around-CLI as an architecture has a working reference point. The branch's HEAD commit matches commit `b1e24b7` (the last commit of the outer-bwrap implementation before the architecture pivot).
|
||||
|
||||
---
|
||||
|
||||
## A1.8 — Updated open questions (supersedes original § 5)
|
||||
|
||||
The original § 5 listed five open questions all of which were specific to the outer-bwrap architecture. Amendment 1 supersedes those and lists the open questions for the new architecture:
|
||||
|
||||
1. **Per-call `wrapWithSandbox()` latency.** PR-B amortized the MITM CA generation (100-500ms) across all spawns by initializing once at boot. Per-call wrap regenerates this if the library does not cache internally. Task #4 PI231 spike measures the actual per-call cost; if it exceeds the original ≤200ms p95 budget, an internal cache wrapper around the library is added in PR-B'. Decision reserved for PR-B'.
|
||||
2. **`vibe` (mistral) home-relocation env var.** Task #4 PI231 spike checks whether `vibe` honours `MISTRAL_HOME` / `VIBE_HOME` / similar. If yes, mistral's ISOLATION block declares it and mistral participates in Layer 1. If no, mistral falls back to Layer 4 (prompt layer) + Layer 3 (per-call wrap with `denyRead` on the operator's real home) only. The provider's `recommendedDeploymentTier` is set accordingly.
|
||||
3. **macOS coverage.** sandbox-runtime supports macOS via `sandbox-exec` (seatbelt profile). Layer 1 ephemeral home is OS-agnostic (just an env var). Layer 3 macOS path needs verification: does per-call `wrapWithSandbox()` with `customConfig` produce a per-spawn sandbox-exec profile, or does it re-use a singleton seatbelt profile? Task #4 PI231 spike is Linux-only; a parallel macOS verification is a Task #9 deliverable.
|
||||
4. **Symlink-vs-bindmount for credentials.** Layer 2 uses symlinks for credential mounting. An alternative is bindmounting the credential file into the ephemeral home (only available inside the Layer 3 wrap). The trade-off: symlinks work outside any sandbox context (so Layer 2 works even when Layer 3 is skipped, e.g., for codex); bindmounts are stronger isolation (the CLI cannot follow the symlink to discover the real path). Decision reserved for PR-B' implementation review.
|
||||
5. **Concurrent-spawn cleanup ordering.** The ephemeral home cleanup (rmdir at response end) must not race with a still-streaming spawn. The current plan: track per-`reqId` cleanup and only fire on the spawn's `exit` event. If a streaming abort leaves the spawn alive past the HTTP response, cleanup is deferred until `exit`. Tested in Task #9.
|
||||
6. **`/health.sandbox.providers` shape under Amendment 1.** Original `/health.sandbox` had a flat `{ available, active }`. Amendment 1 adds per-provider posture: `{ available, providers: { anthropic: { crossTenantReadProtection: 'tool-suppression', layers: ['L1','L2','L3','L4'] }, openai: { crossTenantReadProtection: 'inner-sandbox', layers: ['L1','L4'] } } }`. Exact shape ratified by PR-B'.
|
||||
7. **Dashboard `/dashboard` Security panel.** The dashboard currently has no security panel. Amendment 1 names the addition as a follow-up: render `/health.sandbox.providers` as a per-provider posture badge so the operator can see at a glance which providers are in `tool-suppression` vs `inner-sandbox` vs `none` mode. Out of Phase 7 scope; recorded for a future ADR.
|
||||
|
||||
---
|
||||
|
||||
## A1.9 — Authority citations (Amendment 1)
|
||||
|
||||
Per ALIGNMENT.md Rule 1 (Cite First) and Iron Rule 12 (Pre-Brainstorm Prior-Art Search), every load-bearing claim in this amendment is cited to a primary source. The four forcing reasons are cited above in A1.1.1–A1.1.4; this section enumerates them in one place plus the supporting citations.
|
||||
|
||||
**Forcing reasons:**
|
||||
|
||||
1. **sandbox-runtime documented use-case is inner-wrap by Claude Code.**
|
||||
- https://www.anthropic.com/engineering/claude-code-sandboxing — "Claude Code sandboxing" engineering blog. Documents Claude Code's *internal* use of the library to wrap tool-spawn calls. The blog also notes the library "can be used to sandbox arbitrary processes, agents and MCP servers" — i.e., it is general-purpose, not Claude-Code-internal-only. **Our reading:** outer-wrap of `claude` itself is not the documented or demonstrated direction; OLP would be the only known user in that configuration. Project-design judgment, not an Anthropic prohibition.
|
||||
|
||||
2. **`~/.claude.json` non-atomic write — upstream issue closed `not_planned` by inactivity bot.**
|
||||
- https://github.com/anthropics/claude-code/issues/29250 — upstream issue requesting atomic-write semantics. Status: closed `not_planned` by `github-actions[bot]` on 2026-03-31 (inactivity), labeled `duplicate`, `stale`. **No upstream Anthropic maintainer comment articulates a policy position**; the closure does not establish "won't fix" as policy. Forcing argument rests on architectural-cost analysis (permanent maintenance treadmill for outer-`--ro-bind`), not on alleged upstream policy.
|
||||
|
||||
3. **Codex inner-bwrap conflict.**
|
||||
- https://github.com/openai/codex/issues/16018 — upstream codex CLI issue. Documents that codex's default bwrap sandbox **fails outright** in environments lacking unprivileged user namespaces. The issue is a feature request asking codex to add a fallback path; **the issue body does NOT document an existing automatic fallback to `--sandbox danger-full-access`**. Whether codex degrades to `danger-full-access` or aborts the spawn under nested-bwrap failure is empirically open (Task #4 deliverable). Either failure mode breaks the strict-additive-isolation invariant for outer-wrap of codex. This is the multi-provider forcing function regardless of which failure mode applies.
|
||||
|
||||
4. **`CODEX_HOME` documented relocation lever.**
|
||||
- https://developers.openai.com/codex/config-reference — OpenAI codex CLI config reference (primary).
|
||||
- https://codex.danielvaughan.com/2026/04/08/codex-cli-configuration-reference/ — third-party reference (cross-reference for reachability).
|
||||
|
||||
**Supporting citations (carried forward from original ADR § 7):**
|
||||
|
||||
5. **`@anthropic-ai/sandbox-runtime` v0.0.52** — https://github.com/anthropic-experimental/sandbox-runtime
|
||||
- `dist/sandbox/sandbox-manager.js` — `SandboxManager.wrapWithSandbox(command, undefined, customConfig)` is the per-call wrap surface used by Layer 3. The third argument `customConfig` is the per-call override mechanism that makes Amendment 1's per-spawn config architecture implementable without library modification.
|
||||
|
||||
6. **Internal evidence:**
|
||||
- **PR-B implementation chain** — commits `d0dcd28` → `2864275` → `497b255` → `b1e24b7` → `3551921`. The HTTP-path activation regression is documented in commit message `b1e24b7` and in `lib/sandbox/manager.mjs` § "OLP_SANDBOX_DISABLED env-var gate" comments.
|
||||
- **cc-mem incident memory 2026-05-27** — `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` — the original multi-tenant gap and the prior-art search that established the ecosystem has no working solution.
|
||||
- **2026-05-28 PoC spike on PI231** — `/tmp/sandbox-spike/report.md` on PI231. Verdict was YELLOW (architecturally green, operationally blocked on apt deps). The follow-up 2026-05-29 PI231 prep work re-evaluates against the new architecture; results land in Task #4.
|
||||
|
||||
7. **OLP governance:**
|
||||
- **OLP ALIGNMENT.md Rule 1** — Authority citation required for any provider-plugin / entry-surface / IR change. Amendment 1 amends governance only; the implementation refactor (PR-B') carries its own per-commit citations to the same primary sources enumerated above.
|
||||
- **OLP ALIGNMENT.md Rule 4** — Unalignable plugins are deleted. Mistral's potential lack of a home-relocation env var (open question 2 above) is *not* an alignability gap (mistral's CLI authority is unchanged); it is a deployment-tier classification, recorded in the provider's ISOLATION block.
|
||||
- **Iron Rule 10** — Independent reviewer required. This amendment's review is pending at draft time.
|
||||
- **Iron Rule 11** — Minimum reviewable unit. PR-B' is one PR (sandbox manager refactor); the anthropic ISOLATION block, codex ISOLATION block, server wiring, and README section are each separate PRs per the revised PR plan in § A1.4.
|
||||
- **Iron Rule 12** — Pre-brainstorm prior-art search. The four forcing reasons each satisfy the rule's "provider-specific authority check decisive" condition: Anthropic's blog post + upstream issue 29250 (for the anthropic side), and the codex issue 16018 + the OpenAI config reference (for the codex side).
|
||||
|
||||
8. **OLP ADR cross-references:**
|
||||
- **ADR 0001 § Mission** — multi-provider proxy. Codex inner-bwrap conflict is the multi-provider forcing function.
|
||||
- **ADR 0002 (pending Amendment N)** — Provider ISOLATION contract specification. Amendment 1 names the contract; Amendment N specifies it.
|
||||
- **ADR 0006** — Provider Inclusion / Risk Tier. `recommendedDeploymentTier` in the ISOLATION block integrates with the risk tier framework.
|
||||
- **ADR 0009 Amendment 1 § Caveats #3** — "Sandbox-runtime still required for real multi-tenant deployment." Amendment 1 satisfies this caveat via Layer 3, not via outer-wrap.
|
||||
- **`docs/plans/cloud-deployment-family.md` § 5** — sandbox is a cloud rollout prerequisite. PR-E' updates this section to reflect that the layered architecture is the cloud prerequisite, not outer-bwrap.
|
||||
|
||||
---
|
||||
|
||||
## A1.10 — Consequences of Amendment 1
|
||||
|
||||
### Positive
|
||||
|
||||
- **No outer-bwrap maintenance treadmill.** New `claude` CLI state files do not require OLP-side `--ro-bind` updates. Layer 1 absorbs them automatically.
|
||||
- **Multi-provider compatible.** Codex inner sandbox is preserved unmolested. The architecture works for both anthropic (no inner sandbox) and codex (has inner sandbox) without per-provider workarounds in the sandbox layer; the per-provider differences live in the per-provider ISOLATION block where they belong.
|
||||
- **Per-spawn isolation primitives.** Every request gets a fresh `$HOME`. Cross-tenant state carry-over at the filesystem level is structurally impossible, not "mitigated by careful denylist."
|
||||
- **Aligned with Anthropic's design intent.** OLP uses sandbox-runtime in the direction the library was designed for (sandboxing what the spawn triggers, not wrapping the spawn from outside). The library's semver discipline becomes leverage rather than risk.
|
||||
- **Reduced HTTP-path activation surface.** PR-B's regression was that the in-process MITM proxy lifecycle interacted with OLP's HTTP request handler. Per-call `wrapWithSandbox()` does not require an always-on in-process proxy; the failure mode goes away by construction. (To be confirmed empirically in Task #4 + Task #9.)
|
||||
- **Doctor and `/health.sandbox` continuity.** Operators and dashboard consumers see the same field at the same JSON path. Shape additions are additive, not breaking.
|
||||
|
||||
### Negative
|
||||
|
||||
- **Per-call latency cost.** Per-call `wrapWithSandbox()` is more expensive than once-at-boot init+wrap. The mitigation is library-import caching and (if measured high) a sandbox-config cache keyed by the union of allowed-read paths. Empirical measurement in Task #4.
|
||||
- **New contract surface (ISOLATION block).** Each provider plugin now declares ISOLATION fields. This is incremental complexity in the Provider contract — ratified by ADR 0002 Amendment N. ADR 0002 amendment is on the critical path.
|
||||
- **Mistral declared `crossTenantReadProtection: 'none'`.** Vibe CLI has no Phase-6c-equivalent tool suppression and no known inner sandbox as of D8 ADR 0006 enablement. The mistral provider's `recommendedDeploymentTier` is therefore `separate-vm` per ADR 0002 Amendment 9 § Per-provider concrete instance, meaning mistral can run only in a dedicated VM rather than sharing the OS user with other providers. Not a regression vs status quo (mistral is not deployed today); reflects honest characterization of current state per ALIGNMENT.md Rule 3. Task #4 spike may discover a hardening regime, transitioning this tier upward.
|
||||
- **Symlink semantics edge cases.** Layer 2 symlinks credential files into the ephemeral home; some CLIs may resolve the symlink and write a sibling file in the *target* directory rather than the ephemeral location. Each provider's ISOLATION block should declare any such known behaviour; the runtime tests verify by examining the operator's real `$HOME` for stray writes after a test spawn.
|
||||
- **The `OLP_SANDBOX_DISABLED=1` env-var gate is preserved for 1-2 releases.** It remains a valid escape hatch — but as belt-and-suspenders rather than as load-bearing. Operators who rely on the gate after sunset will see a deprecation message before removal.
|
||||
|
||||
### Reversibility (governance level)
|
||||
|
||||
- Amendment 1 is reversible by a superseding ADR amendment that cites new evidence overturning any of the four forcing reasons. The most likely overturning scenario: Anthropic publishes guidance endorsing outer-wrap of `claude` CLI plus an atomic-write contract for `~/.claude.json`. If that happens, the superseding amendment cites the new guidance and re-enables outer-wrap as an option (alongside, not replacing, the Solution 1 architecture).
|
||||
- The implementation-level reversibility is documented in § A1.7 above.
|
||||
|
||||
---
|
||||
|
||||
## A1.11 — Forward-looking pointer
|
||||
|
||||
Amendment 1 is the governance layer. The implementation lands across Tasks #5–#10 (per the working task list at the time of this draft):
|
||||
|
||||
- Task #4 — PI231 spike to verify `HOME` / `CODEX_HOME` env-var override behaviour (live, with the same `claude` and `codex` CLI versions OLP ships against).
|
||||
- Task #5 — Refactor `lib/sandbox/manager.mjs` to the Layer 1 + Layer 2 + Layer 3 architecture (PR-B').
|
||||
- Task #6 — Add ISOLATION block to `lib/providers/anthropic.mjs` (PR-C').
|
||||
- Task #7 — Add ISOLATION block to `lib/providers/codex.mjs` (PR-D').
|
||||
- Task #8 — Wire `prepareIsolatedEnvironment` into `server.mjs` spawn pipeline (folds into PR-C' or its own PR depending on diff size).
|
||||
- Task #9 — PI231 E2E validation of Solution 1 + close PR-B's load-bearing negative test ("in-sandbox `cat ~/.olp/keys/...` MUST fail") against the new architecture.
|
||||
- Task #10 — README "Security Model" section + cloud-deployment-plan § 5 update + Phase 7 close (PR-E').
|
||||
|
||||
ADR 0002 Amendment N (Provider ISOLATION contract specification) is a co-merged ADR with PR-C'; it cannot land after the ISOLATION block reaches the codebase per ALIGNMENT.md Rule 2(c)'s spirit (no contract field without an authorizing ADR).
|
||||
|
||||
---
|
||||
|
||||
## A1.12 — Amendment status
|
||||
|
||||
- **Drafted:** 2026-05-29 (this document).
|
||||
- **Reviewer:** independent fresh-context reviewer per Iron Rule 10 — pending.
|
||||
- **Implementation gate:** ADR 0002 Amendment N (Provider ISOLATION contract specification) must land before or together with PR-C' (the first ISOLATION-block-bearing provider plugin commit).
|
||||
- **Production gate:** PI231 E2E (Task #9) must pass the load-bearing negative test before the `OLP_SANDBOX_DISABLED=1` env-var gate is removed from prod startup.
|
||||
@@ -23,6 +23,10 @@ New ADRs increment from the highest existing number. Filenames are `NNNN-<short-
|
||||
| [0007](0007-multi-key-auth.md) | Multi-Key Auth (`lib/keys.mjs`) | Phase 2 design ADR (D43-B, 2026-05-25). Option 2 (filesystem manifest at `~/.olp/keys/<key-id>/manifest.json`) + opaque `olp_<32-byte>` token + SHA-256 hash. Owner / guest / anonymous tier gating with explicit `config.json auth.allow_anonymous` (default false). Bootstrap keygen command surface + `OLP_OWNER_TOKEN` env override with stable synthetic `key_id`. Audit ndjson append-only at `~/.olp/logs/audit.ndjson`, warn+1-retry on append failure. Rejects direct SQLite port at v0.2.0 due to Node baseline (`engines >=18` + CI 20/24 vs `node:sqlite` added 22.5.0 / RC); Option 3 hybrid documented as forward path when Phase 3+ Dashboard / SQL-aggregate quota arrives. |
|
||||
| [0008](0008-dashboard-and-audit-query.md) | Dashboard + Audit Query Layer | Phase 3 design ADR (D48, 2026-05-25). Static HTML dashboard + vanilla JS + fetch (no build step). In-memory ndjson scan for aggregate queries (O(N) per call; family-scale acceptable; defers SQLite migration to Option 3 hybrid trigger). Daily audit rotation `audit-YYYY-MM-DD.ndjson` on first append after UTC midnight; cross-file query layer for rolling 30-day windows. Owner-only gating on `/dashboard` + 3 `/v0/management/*` JSON endpoints reusing ADR 0007 § 7 auth model. 30s page poll (no SSE infra). Panels: per-provider quota / 24h request+cache+fallback / 30d spend trend / top-N fallback chains per spec § 4.6. Opens ADR 0007 § 12 Phase 3 deferral (Dashboard + audit query + rotation). |
|
||||
| [0009](0009-interactive-mode-path-placeholder.md) | Anthropic Interactive-Mode Path (Placeholder) | Placeholder ADR (2026-05-25, Draft) — blocked on OCP ADR 0007 P0 experiment outcome. Records the maintainer's "wait + port" decision: do NOT independently implement; ride OCP's P0 result. If P0 confirms Transport A (stdio NDJSON) or B (PTY) bills as subscription rather than Agent SDK credit, port to OLP `lib/providers/anthropic.mjs` (Option 1 parallel impl, or Option 2 OCP-as-backend; decision deferred to P0-resolution time). If P0 fails on both, shelve. No Phase 4 D-day scheduled until P0 lands AND maintainer issues explicit "go" naming this ADR. |
|
||||
| [0010](0010-phase-4-charter-operator-and-client-ux.md) | Phase 4 Charter — Operator + Client UX | Phase 4 scope ratification (2026-05-26, Accepted). Phase 4 = operator + client UX (SSE heartbeat / `olp` CLI + doctor / `olp-connect` zero-config + Telegram-Discord plugin + IDE docs bundle). ~13 D-days, D60 → v0.4.0. Records the explicit decision to DEFER `/v1/messages` (Anthropic-shape entry surface) on the rationale that under ADR 0009 P0 failure it provides no billing benefit AND degrades worse on fallback than OpenAI-shape clients. Re-open trigger: ADR 0009 P0 success + maintainer-named family CC user. Also closes the OCP-OLP port co-host ambiguity from ADR 0001 (default `OLP_PORT` 3456 → 4567). |
|
||||
| [0011](0011-anonymous-key-deployment-context.md) | Anonymous-Key Deployment-Context Limits (Trusted-LAN Invariant) | D70 (2026-05-26, Accepted). Codifies the trust posture for `/health.anonymousKey` opt-in field (D69) + `bin/olp-connect` zero-config consumer (D68). Three-prerequisite gate (`auth.advertise_anonymous_key=true` + `auth.allow_anonymous=true` + an active key with `plaintext_advertise` field). Guest-tier-only restriction (`createKey()` + CLI reject owner+advertise). Trusted-LAN deployment invariant (loopback / RFC1918 / tailnet / `.local` / `.internal` — soft constraint at v0.4.0; hard enforcement deferred until OLP gains a public-deployment recipe). Re-evaluation trigger: any "expose to public internet" README mode. |
|
||||
| [0012](0012-phase-5-charter-quota-probes-dashboard.md) | Phase 5 Charter — Provider Quota Probes + Dashboard Enrichment | Phase 5 scope ratification (2026-05-26, Accepted). Phase 5 = port OCP's plan-usage probe to `lib/providers/anthropic.mjs:quotaStatus()` + Claude.ai-style dashboard enrichment (1-min auto-refresh + manual refresh + per-provider rows with utilization bars, reset countdowns, status badges) + optional mistral probe at D84 (codex explicitly skipped — no public API). ~6 D-days, D79 → v0.5.0. Companion to ADR 0002 Amendment 8 (direct-API READ-ONLY exemption) + ADR 0013 (OAuth READ-ONLY consumption rules). Closes v1.x roadmap #8 (dashboard enrichment). Re-confirmed schema 2026-05-26 via compiled-binary `strings` + live API probe; 3 new fields since OCP 2026-04 capture, no removals. |
|
||||
| [0013](0013-oauth-read-only-consumption-and-schema-drift.md) | OAuth READ-ONLY Consumption Rules + Schema-Drift Mitigation Protocol | D79 (2026-05-26, Accepted). Implementation discipline for ADR 0002 Amendment 8. Seven rules covering: credential reuse with spawn path (no new OAuth grant); READ-ONLY at the wire (one probe per cache miss, `max_tokens:1`, headers-only parse, discard body); cache TTL 5min + 60s-3600s exponential refresh backoff + stale-cache-on-failure; opt-in via `~/.olp/config.json providers.<name>.quota_probe_enabled` (default false); schema-drift mitigation via dual-path verification (compiled-binary `strings` + live API probe diff); failure transparency through `olp doctor` + dashboard staleness markers; out-of-scope clarifications. Bound by `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md` as the live schema pin. |
|
||||
|
||||
## When to write a new ADR
|
||||
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
{
|
||||
"phase": "v0.5.1 post-release",
|
||||
"purpose": "Refresh dashboard screenshot with live MacBook data after v0.5.1 hotfix (replacing D82's synthetic-data render)",
|
||||
"captured_at_utc": "2026-05-27T01:26:17.304Z",
|
||||
"host": "maintainer's MacBook (Mac client test target per project test-envs; specific IP / Tailscale node redacted per public-repo hygiene)",
|
||||
"server_version": "0.5.1 (main @ commit fa2d1af \u2014 F4+#7 post-merge)",
|
||||
"olp_port": 14567,
|
||||
"endpoint_tested": "/v0/management/dashboard-data",
|
||||
"auth": "owner-tier OLP key (temp, revoked post-test)",
|
||||
"result_summary": {
|
||||
"anthropic": {
|
||||
"status": "live",
|
||||
"schema_version": "2026-05-26",
|
||||
"utilization_5h": 0.06,
|
||||
"utilization_7d": 0.38,
|
||||
"representative_claim": "five_hour",
|
||||
"failure": null
|
||||
},
|
||||
"openai": {
|
||||
"status": "unavailable",
|
||||
"reason": "no public quota api or probe disabled"
|
||||
}
|
||||
},
|
||||
"v0_5_1_contract_verified": [
|
||||
"quota_v2[i].status enum includes 'live' (anthropic) and 'unavailable' (openai) \u2014 both rendered correctly",
|
||||
"quota_v2[i].failure is null for healthy live status (per ADR 0013 Rule 6 \u2014 failure info only on stale/unreachable)",
|
||||
"quota_v2[i].schema_version pinned at 2026-05-26 \u2014 matches models-registry.json quota_probe.schema_version"
|
||||
],
|
||||
"post_test_cleanup": [
|
||||
"temp owner key (id=0m6s2s97, name=v0.5.1-screenshot) revoked",
|
||||
"~/.olp/config.json providers.anthropic.quota_probe_enabled flag removed (config restored to baseline)",
|
||||
"test server (pid varies, port=14567) terminated"
|
||||
]
|
||||
}
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 113 KiB |
@@ -0,0 +1,47 @@
|
||||
# OLP IDE & client integrations
|
||||
|
||||
This directory documents per-tool setup for the IDEs and AI clients OLP
|
||||
supports. Every page follows the same shape: one-line description, status
|
||||
icon, copy-paste-able config block, known issues, OLP-specific notes, and a
|
||||
one-line verification command.
|
||||
|
||||
## Index
|
||||
|
||||
| Tool | Status | Path | Notes |
|
||||
|---|---|---|---|
|
||||
| [Continue.dev](./continue.md) | ✅ Supported | VS Code / JetBrains extension | `config.yaml` (NOT `config.json`); supports custom headers |
|
||||
| [Cline](./cline.md) | ✅ Supported | VS Code extension | OpenAI-compatible provider; UI field occasionally vanishes (Cline #7128) |
|
||||
| [Cursor](./cursor.md) | ⚠️ Best-effort | Cursor editor | "Override OpenAI Base URL" — known fragile across releases |
|
||||
| [Aider](./aider.md) | ✅ Supported | terminal CLI | `OPENAI_API_BASE` env + `openai/` model prefix |
|
||||
| [Claude Code](./claude-code.md) | ❌ Not supported | terminal CLI | Anthropic wire format only; OLP serves OpenAI wire format. Use Cline instead. |
|
||||
| [OpenClaw](./openclaw.md) | ✅ Supported | Telegram + Discord gateway | `/olp` slash command via the [`olp-plugin/`](../../olp-plugin/) plugin |
|
||||
|
||||
## Status legend
|
||||
|
||||
- ✅ **Supported** — works against OLP's OpenAI-compatible `/v1/chat/completions`
|
||||
endpoint; the tool's IR fields flow through OLP's IR without lossy translation
|
||||
warnings on the documented chain.
|
||||
- ⚠️ **Best-effort** — works in current versions but the tool has known
|
||||
upstream bugs around base-URL configuration; expect occasional weirdness.
|
||||
- ❌ **Not supported** — the tool's wire protocol or transport is incompatible
|
||||
with what OLP serves; recommended alternative is documented on the page.
|
||||
|
||||
## How OLP's response headers help debugging
|
||||
|
||||
Every response carries (see [README § Response Headers](../../README.md#response-headers)):
|
||||
|
||||
- `X-OLP-Provider-Used` — which provider's plugin served the request
|
||||
- `X-OLP-Model-Used` — which model the served provider used
|
||||
- `X-OLP-Fallback-Hops` — `0` = primary chain entry served it
|
||||
- `X-OLP-Cache` — `hit | miss | bypass`
|
||||
- `X-OLP-Latency-Ms` — end-to-end latency at the proxy
|
||||
|
||||
When something looks wrong in an IDE, the first sanity check is `curl -i`
|
||||
against `/v1/chat/completions` with the same key — those headers tell you
|
||||
whether the IDE config is broken or OLP routed somewhere unexpected.
|
||||
|
||||
## Cross-references
|
||||
|
||||
- [ADR 0010](../adr/0010-phase-4-charter-operator-and-client-ux.md) — Phase 4 charter; documents why `/v1/messages` is not supported and points to Cline as the recommended Anthropic-CLI replacement.
|
||||
- [ADR 0011](../adr/0011-anonymous-key-deployment-context.md) — trusted-LAN-only invariant for `auth.advertise_anonymous_key`.
|
||||
- [`bin/olp-connect`](../../bin/olp-connect) — automated client setup helper (D68-D70).
|
||||
@@ -0,0 +1,106 @@
|
||||
# Aider + OLP
|
||||
|
||||
[Aider](https://aider.chat) is a terminal-native pair programmer that
|
||||
edits files in your local git repo and commits each change. It speaks
|
||||
OpenAI's `/v1/chat/completions` wire format via the `openai/` model
|
||||
prefix.
|
||||
|
||||
**Status:** ✅ Supported.
|
||||
|
||||
**Tested against:** Aider v0.6x. Aider's OpenAI integration has been stable
|
||||
across many releases — this is the most reliable IDE/CLI binding to OLP.
|
||||
|
||||
## Quick setup
|
||||
|
||||
Three knobs, all environment variables or `.env`:
|
||||
|
||||
```bash
|
||||
# Required: point Aider at OLP's chat-completions endpoint
|
||||
export OPENAI_API_BASE=http://127.0.0.1:4567/v1
|
||||
|
||||
# Required: OLP plaintext token
|
||||
export OPENAI_API_KEY=olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||
|
||||
# Then invoke Aider with an OLP-routable model, prefixed `openai/`:
|
||||
aider --model openai/claude-sonnet-4-5
|
||||
```
|
||||
|
||||
The `openai/` prefix tells Aider to use its OpenAI-compatible adapter for
|
||||
the named model. Aider's litellm layer parses this and sends the request
|
||||
to whatever `OPENAI_API_BASE` resolves to.
|
||||
|
||||
Replace the API key with the plaintext token printed by `olp-keys keygen
|
||||
--name=aider`. Family members on the LAN should substitute the OLP host's
|
||||
IP for `127.0.0.1` (or use `olp-connect <ip>`).
|
||||
|
||||
## Aider's `.env` support
|
||||
|
||||
Aider auto-loads a `.env` file from the current directory or the git repo
|
||||
root. The accepted keys are:
|
||||
|
||||
```bash
|
||||
# .env at the project root
|
||||
OPENAI_API_BASE=http://127.0.0.1:4567/v1
|
||||
OPENAI_API_KEY=olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||
|
||||
# Optional: Aider's own AIDER_-prefixed equivalents work too
|
||||
AIDER_OPENAI_API_BASE=http://127.0.0.1:4567/v1
|
||||
```
|
||||
|
||||
The `AIDER_` prefix wins over the bare prefix when both are set. Pick one;
|
||||
mixing them invites surprises during debugging.
|
||||
|
||||
**Hygiene:** add `.env` to `.gitignore` if your repo doesn't already. The
|
||||
OLP token is plaintext-recoverable from disk only because the chat surface
|
||||
explicitly opts into it (see [ADR 0011](../adr/0011-anonymous-key-deployment-context.md))
|
||||
— do not let your IDE bind unintentionally.
|
||||
|
||||
## Known issues
|
||||
|
||||
- **No custom-headers support.** Aider does not expose a way to set extra
|
||||
HTTP headers on outgoing requests. OLP's optional `X-OLP-Chain` /
|
||||
`X-OLP-Bypass-Cache` headers are therefore not available via Aider —
|
||||
routing is determined by the model name alone.
|
||||
|
||||
- **`/v1` trailing matters.** `OPENAI_API_BASE` must end at `/v1` (without
|
||||
`/chat/completions`); Aider appends the remainder. Setting it to the bare
|
||||
host or with a trailing `/chat/completions` causes 404s.
|
||||
|
||||
- **Aider sends `max_tokens` by default.** OLP forwards `max_tokens` to
|
||||
every provider. If you see "model X does not support max_tokens" errors,
|
||||
the underlying provider rejects it — check `X-OLP-Provider-Used` and
|
||||
filter that provider out of the chain for the affected model.
|
||||
|
||||
## OLP-specific notes
|
||||
|
||||
Aider's request shape is faithful to OpenAI's `/v1/chat/completions`
|
||||
spec — `messages`, `model`, `max_tokens`, `stream`, `temperature`,
|
||||
`tools`. All map cleanly into OLP's IR with no lossy-translation warnings.
|
||||
|
||||
For long-context work (codebase summaries, large diffs), set
|
||||
`streaming.heartbeat_interval_ms: 15000` in `~/.olp/config.json` (see
|
||||
[README § Environment Variables](../../README.md#configjson-keys-introduced-at-phase-4))
|
||||
so the SSE stream stays alive through reverse proxies during silent
|
||||
windows.
|
||||
|
||||
## Test it
|
||||
|
||||
```bash
|
||||
# In a scratch dir:
|
||||
aider --model openai/claude-haiku-4-5 --no-stream --message "say ok"
|
||||
```
|
||||
|
||||
Then check OLP's audit log:
|
||||
|
||||
```bash
|
||||
npx olp logs 5
|
||||
```
|
||||
|
||||
The most recent entry should show `provider: anthropic` (or whatever
|
||||
provider haiku routes to in your chain) and `cache_status: miss`.
|
||||
|
||||
## Cross-references
|
||||
|
||||
- Aider model config docs: https://aider.chat/docs/llms/openai-compat.html
|
||||
- [`olp-connect`](../../bin/olp-connect) writes `~/.aider/.env` if Aider is
|
||||
detected on PATH.
|
||||
@@ -0,0 +1,89 @@
|
||||
# Claude Code + OLP
|
||||
|
||||
[Claude Code](https://docs.anthropic.com/en/docs/claude-code/overview) is
|
||||
Anthropic's official terminal-native agent. It speaks the Anthropic
|
||||
`/v1/messages` wire format and cannot be configured to use an
|
||||
OpenAI-compatible chat-completions endpoint.
|
||||
|
||||
**Status:** ❌ Not supported.
|
||||
|
||||
## Why
|
||||
|
||||
OLP serves only the OpenAI `/v1/chat/completions` wire format. Adding
|
||||
`/v1/messages` (the Anthropic shape) was explicitly considered for
|
||||
Phase 4 and rejected, per
|
||||
[ADR 0010 § Out of Phase 4 scope](../adr/0010-phase-4-charter-operator-and-client-ux.md).
|
||||
|
||||
The short version of the rationale:
|
||||
|
||||
- **No billing benefit.** After Anthropic's 2026-06-15 split, `claude -p` /
|
||||
Agent SDK / third-party agent traffic moves out of the Pro/Max
|
||||
subscription pool and into a separate paid Agent SDK Credit pool. OLP's
|
||||
fallback discipline ("when one provider's quota runs out, try the next")
|
||||
does not save money for this traffic — it just routes the same paid
|
||||
request to a different paid backend. The subscription leverage that
|
||||
makes OLP valuable for OpenAI-shape traffic does not exist for
|
||||
Anthropic-shape traffic.
|
||||
|
||||
- **Degrades worse on fallback.** When OLP's primary chain entry (Anthropic)
|
||||
is exhausted, the fallback hop is typically OpenAI Codex or Mistral Vibe.
|
||||
Those providers speak OpenAI tool-calling schema; OLP would have to
|
||||
translate Anthropic's `/v1/messages` tool shape into OpenAI tool shape on
|
||||
every fallback. That translation is lossy and is what ADR 0010 calls out
|
||||
as "net non-positive under P0 failure".
|
||||
|
||||
- **Same outcome reachable via the recommended alternative.** Cline, Cursor,
|
||||
Aider, and Continue.dev all speak OpenAI's wire format and have parity
|
||||
with Claude Code on the "AI edits files in my repo" use case. OLP serves
|
||||
them today.
|
||||
|
||||
## What to use instead
|
||||
|
||||
**Recommended:** [Cline](./cline.md). It's an in-IDE autonomous coder that
|
||||
operates on the same loop Claude Code does (read files, propose edits,
|
||||
run tools, iterate). The "OpenAI Compatible" provider points cleanly at
|
||||
OLP's `/v1/chat/completions` endpoint. You get OLP's full fallback chain
|
||||
(Anthropic → OpenAI Codex → Mistral) instead of being pinned to one
|
||||
provider.
|
||||
|
||||
For terminal users specifically:
|
||||
|
||||
- **[Aider](./aider.md)** if you want the Claude-Code-style git-aware
|
||||
pair programmer in the terminal.
|
||||
- **OpenClaw** if you want Telegram/Discord-driven access to the
|
||||
fallback chain (see [`openclaw.md`](./openclaw.md)).
|
||||
|
||||
## Re-open trigger
|
||||
|
||||
ADR 0010 documents the conditions under which OLP would reconsider
|
||||
`/v1/messages`:
|
||||
|
||||
> (a) ADR 0009 P0 confirms interactive-mode billing classification as
|
||||
> subscription (≥ 2026-07-15) AND (b) maintainer explicitly opens
|
||||
> Phase 5 "Anthropic-shape hub" scope with the name of at least one
|
||||
> family member who wants CC access.
|
||||
|
||||
Until both conditions fire, OLP intentionally does not implement
|
||||
`/v1/messages`. The decision is recorded in ADR 0010 § "Out of Phase 4
|
||||
scope" and ADR 0009 (Anthropic interactive-mode path placeholder).
|
||||
|
||||
## If you absolutely must use Claude Code
|
||||
|
||||
Point Claude Code at api.anthropic.com directly. OLP cannot proxy that
|
||||
traffic. You will:
|
||||
|
||||
- Burn against the Anthropic Pro/Max OAuth subscription (pre-2026-06-15) or
|
||||
the Agent SDK Credit pool (≥ 2026-06-15).
|
||||
- Lose every fallback property OLP provides — when Anthropic's quota is
|
||||
exhausted, Claude Code stops working until the quota resets.
|
||||
- Lose OLP's response headers (`X-OLP-Provider-Used` etc.), audit log
|
||||
entries, cache hits, and `/health` visibility.
|
||||
|
||||
This is documented here only so the trade-off is explicit, not as a
|
||||
recommendation.
|
||||
|
||||
## Cross-references
|
||||
|
||||
- [ADR 0010](../adr/0010-phase-4-charter-operator-and-client-ux.md) § "Out of Phase 4 scope" — full defer rationale.
|
||||
- [ADR 0009](../adr/0009-interactive-mode-path-placeholder.md) — Anthropic 2026-06-15 billing split and re-open trigger.
|
||||
- [`cline.md`](./cline.md) — the recommended alternative for Claude-Code-style workflows.
|
||||
@@ -0,0 +1,94 @@
|
||||
# Cline + OLP
|
||||
|
||||
[Cline](https://github.com/cline/cline) is an autonomous-coder VS Code
|
||||
extension. It speaks OpenAI's `/v1/chat/completions` wire format via its
|
||||
"OpenAI Compatible" provider option.
|
||||
|
||||
**Status:** ✅ Supported.
|
||||
|
||||
**Tested against:** Cline v3.x (extension version visible in VS Code's
|
||||
extension panel). Cline's settings UI has shipped multiple variants of the
|
||||
base-URL field across 2025-2026; if your version doesn't show the field
|
||||
described below, see the Known Issues section.
|
||||
|
||||
## Quick setup
|
||||
|
||||
1. Open the Cline panel in VS Code (sidebar icon).
|
||||
2. Click the settings gear → "API Provider".
|
||||
3. Select **OpenAI Compatible**.
|
||||
4. Fill the fields:
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Base URL | `http://127.0.0.1:4567/v1` |
|
||||
| API Key | `olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX` |
|
||||
| Model ID | `claude-sonnet-4-5` |
|
||||
|
||||
5. Save. Cline shows the model name in the bottom-right corner of the panel.
|
||||
|
||||
Replace the API key with the plaintext token printed by `olp-keys keygen
|
||||
--name=cline`. Family members on the LAN should substitute the OLP host's
|
||||
IP for `127.0.0.1` (or use `olp-connect <ip>`).
|
||||
|
||||
## Known issues
|
||||
|
||||
- **Cline issue [#7128](https://github.com/cline/cline/issues/7128) —
|
||||
base-URL UI field intermittently disappears.** Several Cline releases in
|
||||
2025-2026 shipped a settings UI where the "Base URL" field is hidden when
|
||||
the "OpenAI Compatible" provider is freshly selected. Workaround: switch
|
||||
to a different provider, save, switch back to "OpenAI Compatible" — the
|
||||
field returns. Verify the field is visible in your version BEFORE
|
||||
troubleshooting OLP itself.
|
||||
|
||||
- **Cline writes settings to `.vscode/settings.json` under a
|
||||
`cline.apiConfiguration` key (workspace-scoped) and to the VS Code global
|
||||
state (machine-scoped) depending on the "save to workspace" toggle.** If
|
||||
Cline keeps "forgetting" the OLP base URL across VS Code restarts, the
|
||||
workspace state is overriding the global state. Either save to workspace
|
||||
explicitly, or clear the workspace key and use global state.
|
||||
|
||||
- **Cline sometimes lowercases the model ID before sending.** OLP's
|
||||
`models-registry.json` uses canonical case (e.g. `claude-sonnet-4-5`).
|
||||
This is fine — OLP's router lowercases the requested model for chain
|
||||
lookup. But if you see `unknown model` errors, double-check the exact
|
||||
string Cline sent via the OLP response headers (curl test below).
|
||||
|
||||
## OLP-specific notes
|
||||
|
||||
Cline does not expose a custom-headers field in its OpenAI Compatible
|
||||
provider UI as of v3.x. The OLP routing chain is selected purely from the
|
||||
model ID — pick the canonical name (e.g. `claude-sonnet-4-5`) that matches
|
||||
a `routing.chains` key in your `~/.olp/config.json`.
|
||||
|
||||
OLP's response headers (`X-OLP-Provider-Used`, `X-OLP-Cache`,
|
||||
`X-OLP-Latency-Ms`) are not visible in Cline's UI but are captured by VS
|
||||
Code's Developer Tools Network panel when Cline runs the request.
|
||||
|
||||
## Test it
|
||||
|
||||
```bash
|
||||
# 1. Verify OLP accepts Cline-shape requests
|
||||
curl -sI -X POST http://127.0.0.1:4567/v1/chat/completions \
|
||||
-H "Authorization: Bearer olp_XXXXXX" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model":"claude-sonnet-4-5","messages":[{"role":"user","content":"ok"}],"max_tokens":5,"stream":false}' \
|
||||
| grep -i x-olp
|
||||
```
|
||||
|
||||
Expect `X-OLP-Provider-Used: anthropic` (or whichever provider serves
|
||||
sonnet in your chain) and `X-OLP-Cache: miss` on first request.
|
||||
|
||||
## Why not /v1/messages?
|
||||
|
||||
Cline supports Anthropic-shape requests via a separate "Anthropic" provider
|
||||
in its UI. OLP does not implement `/v1/messages`. Use Cline's **OpenAI
|
||||
Compatible** option pointed at OLP rather than Cline's **Anthropic** option
|
||||
pointed at api.anthropic.com — the OLP chain gives you fallback to OpenAI
|
||||
Codex / Mistral / etc. when the Anthropic subscription hits its quota
|
||||
ceiling. See [ADR 0010 § /v1/messages defer rationale](../adr/0010-phase-4-charter-operator-and-client-ux.md).
|
||||
|
||||
## Cross-references
|
||||
|
||||
- Cline issue tracker: https://github.com/cline/cline/issues
|
||||
- [`olp-connect`](../../bin/olp-connect) automates writing the Cline workspace
|
||||
state.
|
||||
@@ -0,0 +1,95 @@
|
||||
# Continue.dev + OLP
|
||||
|
||||
[Continue.dev](https://continue.dev) is an open-source autocomplete +
|
||||
chat extension for VS Code and JetBrains IDEs. It speaks OpenAI's
|
||||
`/v1/chat/completions` wire format, so it works against OLP with no
|
||||
shim layer.
|
||||
|
||||
**Status:** ✅ Supported.
|
||||
|
||||
**Tested against:** Continue.dev v0.10.x (`config.yaml` schema). The
|
||||
older `config.json` schema (≤ v0.8) is **not** documented here — Continue
|
||||
deprecated it in late 2025 and emits a one-shot migration warning.
|
||||
|
||||
## Quick setup
|
||||
|
||||
Edit `~/.continue/config.yaml` (or open the Continue config from the IDE's
|
||||
extension panel and paste this in):
|
||||
|
||||
```yaml
|
||||
models:
|
||||
- name: olp-chat
|
||||
provider: openai
|
||||
model: claude-sonnet-4-5
|
||||
apiBase: http://127.0.0.1:4567/v1
|
||||
apiKey: olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||
roles:
|
||||
- chat
|
||||
requestOptions:
|
||||
headers:
|
||||
# Optional: pin which routing chain key applies. If omitted, OLP
|
||||
# looks up the chain via the model name above.
|
||||
X-OLP-Chain: claude-sonnet-4-5
|
||||
- name: olp-autocomplete
|
||||
provider: openai
|
||||
model: claude-haiku-4-5
|
||||
apiBase: http://127.0.0.1:4567/v1
|
||||
apiKey: olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||
roles:
|
||||
- autocomplete
|
||||
```
|
||||
|
||||
Replace the API key with the plaintext token printed by `olp-keys keygen
|
||||
--name=continue-dev`. Family members on the LAN should substitute the OLP
|
||||
host's IP for `127.0.0.1` (or use `olp-connect <ip>` to do this for them
|
||||
automatically).
|
||||
|
||||
## Known issues
|
||||
|
||||
- **`apiBase`, NOT `baseURL`.** Continue's YAML schema uses `apiBase` (no
|
||||
`URL` casing). The older `config.json` `baseURL` key was renamed during the
|
||||
v0.10 schema cut. If you copy a snippet from a 2024 blog post and it
|
||||
silently routes to api.openai.com, this is why.
|
||||
- **Trailing `/v1` matters.** OLP's chat-completions endpoint is at
|
||||
`/v1/chat/completions`; Continue appends `/chat/completions` to whatever
|
||||
`apiBase` resolves to. Set `apiBase: http://host:4567/v1` (with `/v1`),
|
||||
not the bare host.
|
||||
- **Provider stays `openai`.** Continue's `provider: anthropic` would send
|
||||
Anthropic-shape requests to `/v1/messages`, which OLP does not implement
|
||||
(see [`claude-code.md`](./claude-code.md) for the rationale).
|
||||
|
||||
## OLP-specific notes
|
||||
|
||||
Continue's `requestOptions.headers` lets you pin OLP-specific routing
|
||||
behaviour without altering the model name itself. Useful headers:
|
||||
|
||||
- `X-OLP-Chain: <chain-key>` — explicitly select the routing chain.
|
||||
- `X-OLP-Bypass-Cache: true` — force a fresh spawn for the next request
|
||||
(debugging cache-poisoning suspicions).
|
||||
|
||||
OLP's response headers (`X-OLP-Provider-Used`, `X-OLP-Cache`, etc.) are
|
||||
visible via VS Code's `Developer: Toggle Developer Tools` → Network panel
|
||||
when Continue runs the request.
|
||||
|
||||
## Test it
|
||||
|
||||
After config save, open the Continue chat panel and send a one-word
|
||||
message ("ok"). Then on the terminal:
|
||||
|
||||
```bash
|
||||
curl -sI -X POST http://127.0.0.1:4567/v1/chat/completions \
|
||||
-H "Authorization: Bearer olp_XXXXXX" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model":"claude-haiku-4-5","messages":[{"role":"user","content":"ok"}],"max_tokens":5}' \
|
||||
| grep -i x-olp
|
||||
```
|
||||
|
||||
Expect `X-OLP-Provider-Used: anthropic` (or whichever provider your chain
|
||||
routes haiku to) and `X-OLP-Cache: miss` on first request, `hit` on the
|
||||
second.
|
||||
|
||||
## Cross-references
|
||||
|
||||
- Continue.dev config reference: https://docs.continue.dev/customization/models
|
||||
- [`olp-connect`](../../bin/olp-connect) automates the Continue.dev branch
|
||||
of this setup.
|
||||
@@ -0,0 +1,116 @@
|
||||
# Cursor + OLP
|
||||
|
||||
[Cursor](https://cursor.com) is an AI-first VS Code fork. It has an
|
||||
"Override OpenAI Base URL" setting that, when populated, routes its
|
||||
default-model traffic to your URL using the OpenAI wire format.
|
||||
|
||||
**Status:** ⚠️ Best-effort.
|
||||
|
||||
**Reason:** Cursor's base-URL override is known to be fragile across
|
||||
releases. Multiple 2025-2026 forum threads document the setting silently
|
||||
reverting, model-list dropdowns not populating from the override URL, and
|
||||
streaming responses falling back to the default backend on parse errors.
|
||||
The behaviour is not specific to OLP — every OpenAI-compatible proxy
|
||||
maintainer documents the same caveats — but Cursor's release cadence is
|
||||
faster than most third-party proxies can test against.
|
||||
|
||||
## Quick setup
|
||||
|
||||
1. Open Cursor → Settings → "Models" → enable **OpenAI API Key**.
|
||||
2. Paste your OLP plaintext token into the **API Key** field.
|
||||
3. Click "Override OpenAI Base URL" and paste:
|
||||
|
||||
```
|
||||
http://127.0.0.1:4567/v1
|
||||
```
|
||||
|
||||
4. Click "Verify". Cursor sends a probe; on success the indicator turns
|
||||
green.
|
||||
|
||||
5. **Crucial step:** in the model list, disable every model that is NOT
|
||||
in your `~/.olp/config.json` `routing.chains`. Cursor's chat will round-
|
||||
robin across enabled models and any model OLP can't route will error.
|
||||
|
||||
Replace the API key with the plaintext token printed by `olp-keys keygen
|
||||
--name=cursor`. Family members on the LAN should substitute the OLP host's
|
||||
IP for `127.0.0.1` (or use `olp-connect <ip>`).
|
||||
|
||||
## Known issues
|
||||
|
||||
- **Override URL silently reverts on Cursor update.** Two reported variants:
|
||||
(a) the field empties; (b) the field shows the OLP URL but Cursor still
|
||||
hits api.openai.com under the hood. Workaround: after every Cursor
|
||||
update, re-open settings, click "Verify" again, and check the OLP
|
||||
server's `/health` for incoming probe requests.
|
||||
|
||||
- **Model-list dropdown does not populate from the override URL.** Cursor
|
||||
hardcodes its model list rather than reading `GET /v1/models`. This is
|
||||
why step 5 above is required — there is no way to make Cursor "discover"
|
||||
your models. You have to disable each model individually that OLP can't
|
||||
serve.
|
||||
|
||||
- **Streaming response parsing is stricter than OpenAI's actual SSE spec.**
|
||||
Cursor occasionally falls back to the default backend if the SSE stream
|
||||
contains a slightly malformed chunk (e.g. an empty `data:` line that
|
||||
OpenAI's API does emit but Cursor's parser doesn't expect). OLP's SSE
|
||||
emitter follows the spec; this is on Cursor's side. If you see traffic
|
||||
hitting api.openai.com despite the override, this is the most likely
|
||||
cause.
|
||||
|
||||
- **Cursor's "Tab" autocomplete is NOT covered by the override.** Tab
|
||||
completion uses a Cursor-proprietary endpoint that is not affected by the
|
||||
OpenAI base URL setting. Only the chat panel is. This is documented
|
||||
Cursor behaviour and is not a bug.
|
||||
|
||||
## OLP-specific notes
|
||||
|
||||
Cursor sends `model: "gpt-4"` or `model: "gpt-3.5-turbo"` (legacy aliases)
|
||||
unless you explicitly select another from its dropdown. Add aliases to
|
||||
your `~/.olp/config.json` `routing.chains` so these route somewhere sane:
|
||||
|
||||
```json
|
||||
{
|
||||
"routing": {
|
||||
"chains": {
|
||||
"gpt-4": [ { "provider": "openai", "model": "gpt-5" } ],
|
||||
"gpt-3.5-turbo": [ { "provider": "openai", "model": "gpt-5-mini" } ]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
(Substitute the OpenAI Codex model names listed by `olp models`.)
|
||||
|
||||
## Recommendation
|
||||
|
||||
**Do not engineer workarounds for Cursor-side bugs.** Cursor's release
|
||||
cadence will fix or re-break the override URL handling at unpredictable
|
||||
intervals. If your daily-driver flow is unreliable, switch to Cline (see
|
||||
[`cline.md`](./cline.md)) — it has a stable OpenAI-compatible provider
|
||||
that does not break across releases.
|
||||
|
||||
## Test it
|
||||
|
||||
```bash
|
||||
curl -sI -X POST http://127.0.0.1:4567/v1/chat/completions \
|
||||
-H "Authorization: Bearer olp_XXXXXX" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model":"gpt-4","messages":[{"role":"user","content":"ok"}],"max_tokens":5}' \
|
||||
| grep -i x-olp
|
||||
```
|
||||
|
||||
After hitting "Send" in Cursor's chat, check the OLP server's recent
|
||||
requests via:
|
||||
|
||||
```bash
|
||||
npx olp logs 10
|
||||
```
|
||||
|
||||
If you don't see Cursor's request in the audit log, traffic isn't reaching
|
||||
OLP — re-check the override URL.
|
||||
|
||||
## Cross-references
|
||||
|
||||
- Cursor forum threads on base-URL fragility: https://forum.cursor.com/ (search "OpenAI base URL")
|
||||
- [`olp-connect`](../../bin/olp-connect) writes Cursor's `cursorrc` if
|
||||
detected, but cannot guarantee the override survives a Cursor update.
|
||||
@@ -0,0 +1,370 @@
|
||||
# OpenClaw + OLP
|
||||
|
||||
[OpenClaw](https://github.com/openclaw/openclaw) is a multi-bot gateway that exposes slash commands on Telegram, Discord, and other chat surfaces. OLP integrates with OpenClaw in two ways:
|
||||
|
||||
1. **`/olp` slash commands** via the [`olp-plugin/`](../../olp-plugin/) plugin (read-only parity to the local `olp` CLI).
|
||||
2. **LLM routing** — OpenClaw's chat agent can route its model calls through your OLP server, giving you per-key audit + quota observability for every bot reply.
|
||||
|
||||
This doc covers both. **Status:** ✅ Supported.
|
||||
|
||||
## Two deployment modes — pick yours
|
||||
|
||||
The OpenClaw config differs significantly depending on whether OpenClaw runs on the same host as the OLP server or on a separate client machine talking to a remote OLP. Pick the right section.
|
||||
|
||||
| | **Mode A: Server-co-located** | **Mode B: Client-mode (recommended for multi-machine setups)** |
|
||||
|---|---|---|
|
||||
| OpenClaw runs on | the OLP server host (loopback) | a different machine (Mac mini, laptop, etc.) |
|
||||
| OLP server runs on | localhost (same host) | a remote host (e.g. PI231) |
|
||||
| `olp-claude` baseUrl | `http://127.0.0.1:4567/v1` | `http://<server-ip>:4567/v1` |
|
||||
| Auth | `authHeader: false` (loopback trusted), OR anonymous-key if `auth.allow_anonymous: true` | `apiKey: "${OLP_OPENCLAW_BOT_TOKEN}"` env-var reference (NOT raw string, NOT `OPENAI_API_KEY` — see § Gotchas) |
|
||||
| `/olp` slash plugin proxyUrl | `http://127.0.0.1:4567` | `http://<server-ip>:4567` |
|
||||
|
||||
## `/olp` slash commands you get
|
||||
|
||||
| Slash command | Maps to | Tier |
|
||||
|---|---|---|
|
||||
| `/olp status` | GET `/v0/management/status` | owner |
|
||||
| `/olp health` | GET `/health` | public |
|
||||
| `/olp usage` | GET `/v0/management/dashboard-data` | owner |
|
||||
| `/olp models` | GET `/v1/models` | public |
|
||||
| `/olp cache` | GET `/cache/stats` | owner |
|
||||
| `/olp providers` | local registry view | public |
|
||||
| `/olp chain show [model]` | local chain view | public |
|
||||
| `/olp doctor` | informational (HTTP endpoint not yet shipped) | — |
|
||||
| `/olp help` | usage text | — |
|
||||
|
||||
**Mutating subcommands are deliberately not exposed via chat.** `keygen`, `revoke`, `restart`, `logs` are SSH-only. See [`olp-plugin/README.md`](../../olp-plugin/README.md#what-you-can-not-do-from-chat-by-design) for the rationale.
|
||||
|
||||
---
|
||||
|
||||
## Mode A — Server-co-located install
|
||||
|
||||
OpenClaw + OLP on the same host. Auth is simpler because everything is on loopback.
|
||||
|
||||
### A1. Install the plugin
|
||||
|
||||
Two install paths — either works.
|
||||
|
||||
**Option A — OpenClaw CLI:**
|
||||
|
||||
```bash
|
||||
openclaw plugins install /path/to/olp/olp-plugin/
|
||||
```
|
||||
|
||||
**Option B — symlink:**
|
||||
|
||||
```bash
|
||||
mkdir -p ~/.openclaw/extensions/
|
||||
ln -s /path/to/olp/olp-plugin/ ~/.openclaw/extensions/olp
|
||||
```
|
||||
|
||||
### A2. Mint a bot owner key
|
||||
|
||||
```bash
|
||||
npx olp-keys keygen --owner --name=openclaw-bot
|
||||
```
|
||||
|
||||
Capture the printed plaintext token — shown exactly once.
|
||||
|
||||
### A3. Configure (loopback recipe)
|
||||
|
||||
Edit `~/.openclaw/openclaw.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"plugins": {
|
||||
"allow": ["...", "olp"],
|
||||
"entries": {
|
||||
"olp": {
|
||||
"enabled": true,
|
||||
"config": {
|
||||
"proxyUrl": "http://127.0.0.1:4567",
|
||||
"apiKey": "olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For LLM routing through OLP, add (or update) the `olp-claude` provider so the bot's default agent goes through OLP-spawned `claude -p`:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": {
|
||||
"providers": {
|
||||
"olp-claude": {
|
||||
"baseUrl": "http://127.0.0.1:4567/v1",
|
||||
"api": "openai-completions",
|
||||
"authHeader": false,
|
||||
"models": [
|
||||
{ "id": "claude-sonnet-4-6", "name": "Claude Sonnet 4.6", "input": ["text"],
|
||||
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } }
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`authHeader: false` is safe on loopback. If you set `auth.allow_anonymous: true` on the OLP server, the bot doesn't even need a key for slash commands (the `apiKey` field can be omitted). Owner-only subcommands (`/olp status`, `/olp usage`, `/olp cache`) still need an owner-tier key.
|
||||
|
||||
### A4. Restart the gateway
|
||||
|
||||
```bash
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Mode B — Client-mode install (OpenClaw on different host than OLP)
|
||||
|
||||
OpenClaw on machine X (e.g., Mac mini), OLP server on machine Y (e.g., a Raspberry Pi or any LAN host). This is the common family deployment shape.
|
||||
|
||||
### B1. Install the plugin
|
||||
|
||||
Same as Mode A:
|
||||
|
||||
```bash
|
||||
openclaw plugins install /path/to/olp/olp-plugin/
|
||||
# OR
|
||||
mkdir -p ~/.openclaw/extensions/
|
||||
ln -s /path/to/olp/olp-plugin/ ~/.openclaw/extensions/olp
|
||||
```
|
||||
|
||||
### B2. Mint a bot owner key (on the OLP server, NOT on the OpenClaw host)
|
||||
|
||||
SSH to the OLP server:
|
||||
|
||||
```bash
|
||||
ssh user@olp-server
|
||||
cd ~/olp
|
||||
node bin/olp-keys.mjs keygen --owner --name=openclaw-<hostname>-bot
|
||||
```
|
||||
|
||||
Capture the plaintext — shown exactly once. **This token will live in `~/.openclaw/openclaw.json` on your OpenClaw host**; pick a name that makes it independently revocable if that host is lost/compromised.
|
||||
|
||||
### B3. Set the bot-token env var (`OLP_OPENCLAW_BOT_TOKEN`)
|
||||
|
||||
OpenClaw's canonical pattern for custom-provider auth is `apiKey: "${VAR_NAME}"` — an env-var reference, NOT a raw token. Choose a **custom** variable name (NOT `OPENAI_API_KEY` — OpenClaw service-manages that one and clobbers it with its own ChatGPT key on every restart). Convention: `OLP_OPENCLAW_BOT_TOKEN`.
|
||||
|
||||
**macOS (gateway under launchd)**:
|
||||
|
||||
```bash
|
||||
launchctl setenv OLP_OPENCLAW_BOT_TOKEN olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||
```
|
||||
|
||||
Add the same `export` to `~/.zshrc` so it survives reboot:
|
||||
|
||||
```bash
|
||||
export OLP_OPENCLAW_BOT_TOKEN=olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||
```
|
||||
|
||||
**Linux (gateway under systemd-user)**: drop a file at `~/.config/environment.d/openclaw-olp.conf`:
|
||||
|
||||
```
|
||||
OLP_OPENCLAW_BOT_TOKEN=olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||
```
|
||||
|
||||
Restart the gateway service so it picks up the new env.
|
||||
|
||||
### B4. Configure `~/.openclaw/openclaw.json`
|
||||
|
||||
Edit `~/.openclaw/openclaw.json` on the OpenClaw host:
|
||||
|
||||
```json
|
||||
{
|
||||
"plugins": {
|
||||
"allow": ["...", "olp"],
|
||||
"entries": {
|
||||
"olp": {
|
||||
"enabled": true,
|
||||
"config": {
|
||||
"proxyUrl": "http://<olp-server-ip>:4567",
|
||||
"apiKey": "olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"models": {
|
||||
"providers": {
|
||||
"olp-claude": {
|
||||
"baseUrl": "http://<olp-server-ip>:4567/v1",
|
||||
"api": "openai-completions",
|
||||
"apiKey": "${OLP_OPENCLAW_BOT_TOKEN}",
|
||||
"models": [
|
||||
{ "id": "claude-sonnet-4-6", "name": "Claude Sonnet 4.6 (via OLP)", "input": ["text"],
|
||||
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } },
|
||||
{ "id": "claude-opus-4-7", "name": "Claude Opus 4.7 (via OLP)", "input": ["text"],
|
||||
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } },
|
||||
{ "id": "claude-haiku-4-5", "name": "Claude Haiku 4.5 (via OLP)", "input": ["text"],
|
||||
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } }
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Note: the `plugins.entries.olp.config.apiKey` field (line 11) IS allowed to be a raw token — it's a separate code path that doesn't suffer the service-managed-env clobber problem. Only the `models.providers.<id>.apiKey` field needs the `${VAR}` env-var-reference workaround.
|
||||
|
||||
### B5. Confirm the default agent model is on `olp-claude`
|
||||
|
||||
Check `agents.defaults.model.primary` in `openclaw.json`. It should be something like:
|
||||
|
||||
```json
|
||||
{ "agents": { "defaults": { "model": { "primary": "olp-claude/claude-sonnet-4-6" } } } }
|
||||
```
|
||||
|
||||
If it's pointing at one of OpenClaw's stock providers (`openai/...`, `anthropic/...`, `github-copilot/...`), free-text chat will **bypass OLP entirely** and hit your direct API account. You'll see no traffic in OLP's `/dashboard` and `/olp usage` will show no recent activity.
|
||||
|
||||
### B5. Restart the gateway
|
||||
|
||||
```bash
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
### B6. Verify routing
|
||||
|
||||
In Telegram or Discord, send a free-text message ("hello"). It should:
|
||||
1. Return a normal LLM reply (not "Something went wrong")
|
||||
2. Show up on the OLP dashboard's 24h-requests counter
|
||||
3. Show up in `/olp usage` per-provider count
|
||||
|
||||
If you see "Something went wrong" — see § Troubleshooting below.
|
||||
|
||||
---
|
||||
|
||||
## Using codex / OpenAI models through OLP
|
||||
|
||||
By default the `olp-claude` provider only knows about Claude models. To route OpenAI / codex models through OLP (so bot calls to `gpt-5.5` etc. spawn `codex exec --json` on the OLP server and benefit from per-key audit + quota tracking), add a second provider:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": {
|
||||
"providers": {
|
||||
"olp-codex": {
|
||||
"baseUrl": "http://<olp-server-ip>:4567/v1",
|
||||
"api": "openai-completions",
|
||||
"apiKey": "${OLP_OPENCLAW_BOT_TOKEN}",
|
||||
"models": [
|
||||
{ "id": "gpt-5.5", "name": "GPT 5.5 (via OLP→codex)", "input": ["text"],
|
||||
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } },
|
||||
{ "id": "gpt-5.4-mini", "name": "GPT 5.4 mini (via OLP→codex)", "input": ["text"],
|
||||
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } },
|
||||
{ "id": "gpt-5.3-codex", "name": "GPT 5.3 codex (via OLP→codex)", "input": ["text"],
|
||||
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } }
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"agents": {
|
||||
"defaults": {
|
||||
"models": {
|
||||
"olp-codex/gpt-5.5": { "alias": "OLP GPT 5.5" },
|
||||
"olp-codex/gpt-5.4-mini": { "alias": "OLP GPT 5.4 mini" },
|
||||
"olp-codex/gpt-5.3-codex": { "alias": "OLP Codex" }
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
After restart, type `/models` in Telegram and pick `olp-codex/gpt-5.5` from the menu that appears. **`/models` is menu-driven — it does not accept inline model names**; typing `/models olp-codex/gpt-5.5` won't directly switch you. The available IDs are the ones OLP's `/v1/models` returns — typically `gpt-5.5`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.3-codex`, `gpt-5.3-codex-spark`. Query your OLP server to see the live list:
|
||||
|
||||
```bash
|
||||
curl -s -H "Authorization: Bearer olp_…" http://<olp-server-ip>:4567/v1/models | jq '.data[].id'
|
||||
```
|
||||
|
||||
**Why not use OpenClaw's stock `openai` provider?** OpenClaw's built-in `openai-codex` provider uses the local ChatGPT account (via the `sk-proj-…` API key OpenClaw stores) and bypasses your OLP server entirely. You'd lose per-key audit + per-key quota visibility. `olp-codex` keeps everything routed through your central OLP for observability.
|
||||
|
||||
---
|
||||
|
||||
## Gotchas
|
||||
|
||||
### Auth must use `apiKey: "${VAR}"` env-var reference — three failure modes to avoid
|
||||
|
||||
Custom OpenAI-compatible providers in OpenClaw have a fragile auth path. Three patterns that **don't work** + the one that **does**:
|
||||
|
||||
**❌ `apiKey: "olp_<raw-token>"`** — raw string. Silently bypassed in some routing paths because OpenClaw treats `OPENAI_API_KEY` as service-managed (`OPENCLAW_SERVICE_MANAGED_ENV_KEYS=DEEPSEEK_API_KEY,OPENAI_API_KEY`), and certain model id patterns (notably `gpt-*`) fall back to that env var instead of using your explicit `apiKey`. Symptom: OLP audit shows the request as `__anonymous__` instead of your owner key. Confirmed via [openclaw#41157](https://github.com/openclaw/openclaw/issues/41157) (Gemini openai-completions Authorization not sent) and [#1669](https://github.com/openclaw/openclaw/issues/1669) (Ollama provider ignores apiKey, hardcodes Bearer). Both unresolved upstream as of OpenClaw v2026.5.
|
||||
|
||||
**❌ `headers: { "Authorization": "Bearer olp_<raw-token>" }`** — works for SOME provider/model combinations (e.g., model id `claude-sonnet-4-6`) but breaks for openai-shape model ids (`gpt-5.5` etc.) which take a different code path that ignores the `headers` field. Mixed behavior is worse than no behavior.
|
||||
|
||||
**❌ Setting `OPENAI_API_KEY=olp_…`** in the gateway env. OpenClaw service-manages that variable and overwrites your value with the user's ChatGPT key on every gateway start.
|
||||
|
||||
**✅ `apiKey: "${OLP_OPENCLAW_BOT_TOKEN}"`** — env-var reference with a **custom** variable name (NOT `OPENAI_API_KEY`). OpenClaw resolves the reference at request-construction time, before any service-managed-env logic runs. Both `olp-claude/*` (Claude models) and `olp-codex/*` (OpenAI models) auth correctly with this pattern. Verified end-to-end 2026-05-27: OLP audit shows requests attributed to the correct bot key for both provider blocks.
|
||||
|
||||
OpenClaw docs call out this as the canonical pattern: see [docs.openclaw.ai/concepts/model-providers](https://docs.openclaw.ai/concepts/model-providers) "API key or SecretRef/env reference".
|
||||
|
||||
### Default agent model still points at a removed provider
|
||||
|
||||
If you've removed a provider (e.g., torn down a co-located OCP server) but the bot's default agent model still references that provider, free-text messages will fail with "Something went wrong while processing your request." Check `agents.defaults.model.primary` and update it to a provider that exists.
|
||||
|
||||
### `/new` does not reset model selection — use `/reset`
|
||||
|
||||
OpenClaw's `/new` resets the **conversation context** but **preserves** the session's `/models` selection. If a session has been switched to a model that no longer works (revoked / removed), `/new` won't help — use `/reset` (resets both context and model selection).
|
||||
|
||||
### `/models` is menu-only — does not accept inline model names
|
||||
|
||||
The OpenClaw `/models` command in Telegram is **menu-driven**: typing `/models` pops a model-picker menu where you tap the model name. Typing `/models olp-codex/gpt-5.5` does NOT switch — it'll open the picker. The bot's own success-message after a pick may say *"Use `/model olp-codex/gpt-5.5 --runtime <runtime>` to switch harnesses."* — **that command form is not actually accepted by the bot**; ignore that line.
|
||||
|
||||
### OpenClaw v2026.5+ requires `openclaw.extensions` in `package.json`
|
||||
|
||||
OpenClaw versions ≥ 2026.5.22 enforce a stricter plugin-manifest validation at `openclaw plugins install` time. If `Option A` fails with `package.json missing openclaw.extensions` despite recent OLP releases, your local `olp-plugin/package.json` may predate the v0.5.x fix that adds `"extensions": ["./index.js"]` to the `openclaw` block. Pull latest OLP main (`git pull` in your OLP clone) and retry, or fall through to symlink Option B which works against any plugin shape. (Original drift event: 2026-05-27, see commit history of `olp-plugin/package.json`.)
|
||||
|
||||
### `openclaw gateway restart` is required after install
|
||||
|
||||
OpenClaw caches plugin discovery + model-provider config at gateway start. `openclaw plugins reload` does not guarantee a fresh import of the plugin module nor a fresh re-read of `models.providers.*`. Restart the gateway after every change to `~/.openclaw/openclaw.json`.
|
||||
|
||||
### Owner-key revocation kicks the plugin out immediately
|
||||
|
||||
If you revoke the bot's owner key (`npx olp-keys revoke --id=<id>`), the next `/olp status` will return `401 unauthorized`. Mint a replacement key with a new name and edit `~/.openclaw/openclaw.json`; do NOT reuse the revoked key's UUID.
|
||||
|
||||
### Long responses are truncated
|
||||
|
||||
Telegram caps messages at ~4096 characters. The plugin truncates with a `... [truncated, use SSH for full]` suffix when the rendered output would exceed ~3900 chars. Use SSH + the local `olp` CLI for full output.
|
||||
|
||||
---
|
||||
|
||||
## OLP-specific notes
|
||||
|
||||
The plugin honours these env vars on the OpenClaw gateway process:
|
||||
|
||||
- `OLP_PROXY_URL` — full URL, overrides plugin config `proxyUrl`.
|
||||
- `OLP_PORT` — port only, localhost assumed; overrides `proxyUrl` when `OLP_PROXY_URL` is unset.
|
||||
|
||||
If you run the OpenClaw gateway under launchd or systemd with custom env vars, set `OLP_PROXY_URL` there rather than editing the plugin config — that way the same plugin install can serve multiple OLP hosts.
|
||||
|
||||
## Per-bot vs maintainer key
|
||||
|
||||
**Always create a dedicated bot key**, never the maintainer's personal owner key. The bot key:
|
||||
|
||||
- Has its own `id` so you can revoke it without affecting other clients.
|
||||
- Has its own audit-log entries so you can attribute `/v0/management/*` traffic to the bot.
|
||||
- Can be rotated routinely (every 90 days etc.) without coordinating with the maintainer's daily-driver IDE configs.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Likely cause | Fix |
|
||||
|---|---|---|
|
||||
| `/olp status` returns 401 | bot key revoked / wrong / missing | Mint new key on OLP host; update `plugins.entries.olp.config.apiKey`; restart gateway |
|
||||
| `/olp status` returns 403 | bot key is guest-tier, not owner-tier | Generate owner-tier key (`olp-keys keygen --owner --name=...`); update config |
|
||||
| `OLP error: fetch failed` | `proxyUrl` unreachable from the gateway host | `curl http://<proxyUrl>/health` from the gateway host to confirm reachability; check firewall / OLP server `OLP_BIND=0.0.0.0` for LAN access |
|
||||
| Bot free-text chat returns "Something went wrong" but `/olp ...` works | Default agent model points at a broken provider (e.g., a removed OCP install) | Check `agents.defaults.model.primary` in `openclaw.json`; update to `olp-claude/claude-sonnet-4-6` or another working provider |
|
||||
| Free-text returns `HTTP 401: OLP API key is invalid` despite fresh key | Raw-string `apiKey: "olp_..."` shadowed by service-managed env clobber; or `headers.Authorization` bypassed for `gpt-*` model ids | Switch `models.providers.<id>.apiKey` to env-var reference: `"${OLP_OPENCLAW_BOT_TOKEN}"` (see § Gotchas: Auth) |
|
||||
| OLP audit shows `key_id=__anonymous__` for traffic that should be owner-attributed | Same root cause as 401 — raw-string apiKey or headers bypassed in some routing paths | Switch to env-var-reference `apiKey: "${VAR}"` pattern + verify `launchctl getenv OLP_OPENCLAW_BOT_TOKEN` returns the expected token |
|
||||
| Bot routes to ChatGPT account directly, not through OLP | Provider config uses OpenClaw stock `openai-codex` instead of a custom OLP-pointing provider | Add `olp-codex` provider per § Using codex / OpenAI models through OLP |
|
||||
| `/models olp-codex/gpt-5.5` typed inline doesn't work | OpenClaw `/models` is menu-only, doesn't accept inline names | Type `/models`, tap the model from the picker menu that appears |
|
||||
|
||||
## Cross-references
|
||||
|
||||
- [`olp-plugin/README.md`](../../olp-plugin/README.md) — full plugin docs.
|
||||
- [ADR 0010 § Phase 4 D71-D73](../adr/0010-phase-4-charter-operator-and-client-ux.md) — the plugin's charter.
|
||||
- [OCP `/ocp` plugin](https://github.com/dtzp555-max/ocp/tree/main/ocp-plugin) — the OCP predecessor (includes mutating subcommands that OLP deliberately drops).
|
||||
@@ -0,0 +1,655 @@
|
||||
# OLP Cloud Deployment Plan — Family Testing Phase
|
||||
|
||||
**Status:** Draft — pending current Phase 6 completion
|
||||
**Target:** Oracle Cloud VM (existing infrastructure)
|
||||
**Audience:** Project maintainer deployment reference
|
||||
**Scope:** Single-VM deployment for family (3–5 users), spawn-binary architecture, public internet exposure with hardened auth
|
||||
|
||||
---
|
||||
|
||||
## 0. Prerequisites
|
||||
|
||||
- OLP current phase (Phase 6) is closed and tagged
|
||||
- Oracle Cloud VM accessible via SSH (existing `opc` user)
|
||||
- Domain name (optional but strongly recommended for TLS)
|
||||
- Provider CLI OAuth completed on at least one machine (credentials transferable)
|
||||
|
||||
---
|
||||
|
||||
## 1. Architecture Overview
|
||||
|
||||
```
|
||||
┌────────────────────────────────────────────────────────────────────┐
|
||||
│ Family Devices (anywhere on internet) │
|
||||
│ │
|
||||
│ Wife iPad / Kid Laptop / Maintainer MacBook / ... │
|
||||
│ IDE: Cline / Continue.dev / Cursor / Aider / OpenClaw │
|
||||
│ Config: OPENAI_BASE_URL=https://olp.example.com/v1 │
|
||||
│ OPENAI_API_KEY=olp_<personal-key> │
|
||||
└──────────────────────────┬─────────────────────────────────────────┘
|
||||
│ HTTPS (TLS 1.3)
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────────────┐
|
||||
│ Oracle Cloud VM │
|
||||
│ │
|
||||
│ ┌─ iptables / OCI Security List ──────────────────────────────┐ │
|
||||
│ │ ALLOW: TCP 443 (HTTPS) from 0.0.0.0/0 │ │
|
||||
│ │ ALLOW: TCP 22 (SSH) from maintainer IP only │ │
|
||||
│ │ DENY: everything else │ │
|
||||
│ └─────────────────────────────────────────────────────────────┘ │
|
||||
│ │
|
||||
│ ┌─ Nginx (reverse proxy + TLS termination) ───────────────────┐ │
|
||||
│ │ :443 → TLS (Let's Encrypt auto-renew via certbot) │ │
|
||||
│ │ proxy_pass → http://127.0.0.1:4567 │ │
|
||||
│ │ Rate limit: 30 req/min per IP (burst 10) │ │
|
||||
│ │ Request body limit: 1MB │ │
|
||||
│ │ Connection timeout: 300s (streaming needs long timeout) │ │
|
||||
│ └─────────────────────────────────────────────────────────────┘ │
|
||||
│ │
|
||||
│ ┌─ OLP server.mjs ───────────────────────────────────────────┐ │
|
||||
│ │ OLP_BIND=127.0.0.1 (loopback only — Nginx fronts it) │ │
|
||||
│ │ OLP_PORT=4567 │ │
|
||||
│ │ auth.allow_anonymous: false │ │
|
||||
│ │ auth.advertise_anonymous_key: false │ │
|
||||
│ │ Per-key audit logging to ~/.olp/logs/audit.ndjson │ │
|
||||
│ │ Owner key: maintainer only │ │
|
||||
│ │ Guest keys: one per family member │ │
|
||||
│ └─────────────────────────────────────────────────────────────┘ │
|
||||
│ │
|
||||
│ ┌─ Provider CLIs (installed on this VM) ──────────────────────┐ │
|
||||
│ │ claude → ~/.claude/.credentials.json (OAuth) │ │
|
||||
│ │ codex → ~/.codex/auth.json (OAuth) │ │
|
||||
│ │ vibe → ~/.vibe/.env (API key) │ │
|
||||
│ └─────────────────────────────────────────────────────────────┘ │
|
||||
│ │
|
||||
│ ┌─ systemd service ──────────────────────────────────────────┐ │
|
||||
│ │ olp.service: auto-start, auto-restart on crash │ │
|
||||
│ │ Runs as dedicated `olp` user (not root, not opc) │ │
|
||||
│ └─────────────────────────────────────────────────────────────┘ │
|
||||
│ │
|
||||
└────────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ Provider CLIs spawn outbound HTTPS calls
|
||||
▼
|
||||
Anthropic API / OpenAI API / Mistral API
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Security Design (7 Layers)
|
||||
|
||||
### Layer 1 — Network Perimeter (OCI Security List + iptables)
|
||||
|
||||
**Principle:** Minimum attack surface. Only two ports reachable from the internet.
|
||||
|
||||
```
|
||||
OCI Security List (stateful ingress rules):
|
||||
┌──────────┬────────────┬───────────────────────────────┐
|
||||
│ Port │ Protocol │ Source │
|
||||
├──────────┼────────────┼───────────────────────────────┤
|
||||
│ 443 │ TCP │ 0.0.0.0/0 (public HTTPS) │
|
||||
│ 22 │ TCP │ <maintainer-IP>/32 only │
|
||||
└──────────┴────────────┴───────────────────────────────┘
|
||||
|
||||
NOT exposed:
|
||||
- Port 4567 (OLP direct) — Nginx fronts it
|
||||
- Port 80 (HTTP) — only for certbot ACME challenge, redirect to 443
|
||||
```
|
||||
|
||||
**iptables backup** (defense in depth — OCI Security List is primary, iptables is secondary):
|
||||
|
||||
```bash
|
||||
# Drop everything by default
|
||||
sudo iptables -P INPUT DROP
|
||||
sudo iptables -P FORWARD DROP
|
||||
|
||||
# Allow established connections
|
||||
sudo iptables -A INPUT -m state --state ESTABLISHED,RELATED -j ACCEPT
|
||||
|
||||
# Allow loopback
|
||||
sudo iptables -A INPUT -i lo -j ACCEPT
|
||||
|
||||
# Allow SSH from maintainer IP only
|
||||
sudo iptables -A INPUT -p tcp --dport 22 -s <MAINTAINER_IP> -j ACCEPT
|
||||
|
||||
# Allow HTTPS from anywhere
|
||||
sudo iptables -A INPUT -p tcp --dport 443 -j ACCEPT
|
||||
|
||||
# Allow HTTP (certbot ACME only — Nginx redirects everything else)
|
||||
sudo iptables -A INPUT -p tcp --dport 80 -j ACCEPT
|
||||
|
||||
# Persist
|
||||
sudo iptables-save | sudo tee /etc/iptables/rules.v4
|
||||
```
|
||||
|
||||
### Layer 2 — TLS Termination (Nginx + Let's Encrypt)
|
||||
|
||||
**Principle:** All client traffic encrypted. OLP itself runs plain HTTP on loopback — simpler, no cert management in Node.
|
||||
|
||||
```nginx
|
||||
# /etc/nginx/sites-available/olp.conf
|
||||
|
||||
# Redirect HTTP → HTTPS
|
||||
server {
|
||||
listen 80;
|
||||
server_name olp.example.com;
|
||||
|
||||
# Let's Encrypt ACME challenge
|
||||
location /.well-known/acme-challenge/ {
|
||||
root /var/www/certbot;
|
||||
}
|
||||
|
||||
location / {
|
||||
return 301 https://$host$request_uri;
|
||||
}
|
||||
}
|
||||
|
||||
# HTTPS — TLS 1.3 only
|
||||
server {
|
||||
listen 443 ssl http2;
|
||||
server_name olp.example.com;
|
||||
|
||||
# TLS config
|
||||
ssl_certificate /etc/letsencrypt/live/olp.example.com/fullchain.pem;
|
||||
ssl_certificate_key /etc/letsencrypt/live/olp.example.com/privkey.pem;
|
||||
ssl_protocols TLSv1.3; # TLS 1.3 only
|
||||
ssl_prefer_server_ciphers off; # TLS 1.3 manages its own
|
||||
ssl_session_timeout 1d;
|
||||
ssl_session_cache shared:SSL:10m;
|
||||
|
||||
# Security headers
|
||||
add_header Strict-Transport-Security "max-age=63072000" always;
|
||||
add_header X-Content-Type-Options nosniff;
|
||||
add_header X-Frame-Options DENY;
|
||||
|
||||
# Rate limiting (per IP)
|
||||
limit_req zone=olp_limit burst=10 nodelay;
|
||||
|
||||
# Request body size (LLM prompts can be large but cap at 1MB)
|
||||
client_max_body_size 1m;
|
||||
|
||||
# Proxy to OLP
|
||||
location / {
|
||||
proxy_pass http://127.0.0.1:4567;
|
||||
proxy_http_version 1.1;
|
||||
proxy_set_header Host $host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
proxy_set_header X-Forwarded-Proto $scheme;
|
||||
|
||||
# SSE streaming support (critical for /v1/chat/completions)
|
||||
proxy_set_header Connection '';
|
||||
proxy_buffering off; # Don't buffer SSE
|
||||
proxy_cache off;
|
||||
chunked_transfer_encoding on;
|
||||
|
||||
# Long timeouts for LLM inference
|
||||
proxy_connect_timeout 10s;
|
||||
proxy_read_timeout 300s; # 5 min — long reasoning
|
||||
proxy_send_timeout 300s;
|
||||
}
|
||||
}
|
||||
|
||||
# Rate limit zone definition (in http {} block of nginx.conf)
|
||||
# limit_req_zone $binary_remote_addr zone=olp_limit:10m rate=30r/m;
|
||||
```
|
||||
|
||||
**Certbot auto-renewal:**
|
||||
|
||||
```bash
|
||||
sudo certbot certonly --webroot -w /var/www/certbot -d olp.example.com
|
||||
# Auto-renew via systemd timer (certbot installs this automatically)
|
||||
```
|
||||
|
||||
### Layer 3 — Application Auth (OLP Multi-Key)
|
||||
|
||||
**Principle:** Every request must carry a valid API key. No anonymous access. Per-key audit trail.
|
||||
|
||||
```json
|
||||
// ~/.olp/config.json on the cloud VM
|
||||
{
|
||||
"auth": {
|
||||
"allow_anonymous": false,
|
||||
"advertise_anonymous_key": false,
|
||||
"owner_only_endpoints": [
|
||||
"/health",
|
||||
"/v0/management/dashboard-data",
|
||||
"/v0/management/quota",
|
||||
"/v0/management/status",
|
||||
"/cache/stats",
|
||||
"/dashboard"
|
||||
],
|
||||
"fallback_detail_header_policy": "owner_only"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Key provisioning plan:**
|
||||
|
||||
```
|
||||
┌───────────────┬──────────┬─────────────────────────────────────┐
|
||||
│ Key name │ Tier │ providers_enabled │
|
||||
├───────────────┼──────────┼─────────────────────────────────────┤
|
||||
│ cloud-owner │ owner │ all (dashboard + management access) │
|
||||
│ wife-ipad │ guest │ anthropic, openai │
|
||||
│ kid-laptop │ guest │ anthropic only (cost control) │
|
||||
│ maintainer-mb │ guest │ all (daily driver, not owner tier) │
|
||||
└───────────────┴──────────┴─────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Why maintainer uses a guest key for daily driving:** owner key gives access to management endpoints. Routine IDE usage should not carry owner privilege. Owner key is used only for dashboard access and administration.
|
||||
|
||||
**Key lifecycle:**
|
||||
- Keys generated on the cloud VM via `olp-keys keygen`
|
||||
- Plaintext token communicated to family member via secure channel (Signal / iMessage, not email)
|
||||
- Each key logged independently in audit.ndjson (per-key `key_id` field)
|
||||
- Revocation: `olp-keys revoke --id=<key-id>` — immediate, no grace period
|
||||
|
||||
### Layer 4 — Process Isolation (Dedicated User + systemd)
|
||||
|
||||
**Principle:** OLP runs as a non-root, non-login user. Crash recovery is automatic.
|
||||
|
||||
```bash
|
||||
# Create dedicated user
|
||||
sudo useradd --system --shell /usr/sbin/nologin --home-dir /opt/olp olp
|
||||
|
||||
# OLP code
|
||||
sudo mkdir -p /opt/olp
|
||||
sudo git clone https://github.com/dtzp555-max/olp.git /opt/olp/app
|
||||
sudo chown -R olp:olp /opt/olp
|
||||
|
||||
# OLP data (keys, config, logs, cache)
|
||||
sudo mkdir -p /home/olp/.olp/{keys,logs,cache}
|
||||
sudo chown -R olp:olp /home/olp
|
||||
```
|
||||
|
||||
**systemd unit:**
|
||||
|
||||
```ini
|
||||
# /etc/systemd/system/olp.service
|
||||
[Unit]
|
||||
Description=OLP — Open LLM Proxy
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=olp
|
||||
Group=olp
|
||||
|
||||
WorkingDirectory=/opt/olp/app
|
||||
ExecStart=/usr/bin/node server.mjs
|
||||
|
||||
# Environment
|
||||
Environment=OLP_BIND=127.0.0.1
|
||||
Environment=OLP_PORT=4567
|
||||
Environment=NODE_ENV=production
|
||||
Environment=HOME=/home/olp
|
||||
|
||||
# Auto-restart on crash
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
StartLimitIntervalSec=60
|
||||
StartLimitBurst=5
|
||||
|
||||
# Security hardening
|
||||
NoNewPrivileges=true
|
||||
ProtectSystem=strict
|
||||
ProtectHome=false
|
||||
ReadWritePaths=/home/olp/.olp
|
||||
PrivateTmp=true
|
||||
|
||||
# Resource limits
|
||||
LimitNOFILE=65536
|
||||
MemoryMax=1G
|
||||
|
||||
# Logging
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
SyslogIdentifier=olp
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
|
||||
### Layer 5 — Credential Protection (Provider OAuth Tokens)
|
||||
|
||||
**Principle:** OAuth tokens are the crown jewels. Stolen tokens = someone else using your Claude/OpenAI subscription.
|
||||
|
||||
```
|
||||
Credential storage on cloud VM:
|
||||
|
||||
~olp/
|
||||
├── .claude/
|
||||
│ └── .credentials.json # chmod 600, owner=olp
|
||||
├── .codex/
|
||||
│ └── auth.json # chmod 600, owner=olp
|
||||
└── .vibe/
|
||||
└── .env # chmod 600, owner=olp
|
||||
|
||||
Security measures:
|
||||
1. chmod 600 on all credential files (olp user only)
|
||||
2. Credential files NOT in the git repo (already .gitignored)
|
||||
3. No credential in env vars (OLP reads from filesystem)
|
||||
4. Credential transfer: scp from local machine, then delete local copy of the scp command from shell history
|
||||
5. Periodic rotation: re-auth quarterly (or on any suspicion of compromise)
|
||||
```
|
||||
|
||||
**Credential transfer procedure:**
|
||||
|
||||
```bash
|
||||
# FROM maintainer's Mac mini (one-time):
|
||||
|
||||
# 1. Claude credentials
|
||||
scp ~/.claude/.credentials.json opc@<cloud-ip>:/tmp/claude-cred.json
|
||||
ssh opc@<cloud-ip> "sudo mv /tmp/claude-cred.json /home/olp/.claude/.credentials.json && sudo chown olp:olp /home/olp/.claude/.credentials.json && sudo chmod 600 /home/olp/.claude/.credentials.json"
|
||||
|
||||
# 2. Codex credentials
|
||||
scp ~/.codex/auth.json opc@<cloud-ip>:/tmp/codex-cred.json
|
||||
ssh opc@<cloud-ip> "sudo mv /tmp/codex-cred.json /home/olp/.codex/auth.json && sudo chown olp:olp /home/olp/.codex/auth.json && sudo chmod 600 /home/olp/.codex/auth.json"
|
||||
|
||||
# 3. Mistral API key
|
||||
ssh opc@<cloud-ip> "sudo -u olp bash -c 'echo MISTRAL_API_KEY=sk-xxx > ~/.vibe/.env && chmod 600 ~/.vibe/.env'"
|
||||
|
||||
# 4. Verify
|
||||
ssh opc@<cloud-ip> "sudo -u olp node /opt/olp/app/bin/olp.mjs doctor --json" | jq '.checks[] | select(.name | contains("auth"))'
|
||||
```
|
||||
|
||||
### Layer 6 — Audit and Monitoring
|
||||
|
||||
**Principle:** Every request logged. Anomalies detectable. No silent failures.
|
||||
|
||||
**Audit (already built into OLP):**
|
||||
- `~/.olp/logs/audit.ndjson` — append-only, per-request, includes `key_id`, provider, model, cache hit/miss, fallback hops
|
||||
- Daily rotation: `audit-YYYY-MM-DD.ndjson` (built-in, triggers on first append after UTC midnight)
|
||||
- External rotation tool: `olp-audit-rotate` (idempotent, cron-safe)
|
||||
|
||||
**Additional monitoring for cloud deployment:**
|
||||
|
||||
```bash
|
||||
# Cron: daily audit rotation (belt-and-suspenders alongside in-server rotation)
|
||||
0 0 * * * /usr/bin/node /opt/olp/app/bin/olp-audit-rotate.mjs
|
||||
|
||||
# Cron: daily health check + alert
|
||||
*/5 * * * * curl -sf -H "Authorization: Bearer $OLP_OWNER_KEY" https://olp.example.com/health > /dev/null || echo "OLP health check failed at $(date)" >> /home/olp/alerts.log
|
||||
|
||||
# Cron: audit log size check (alert if >100MB — suggests anomalous traffic)
|
||||
0 6 * * * find /home/olp/.olp/logs -name 'audit*.ndjson' -size +100M -exec echo "Large audit log: {}" \; >> /home/olp/alerts.log
|
||||
|
||||
# Cron: disk usage check
|
||||
0 6 * * * df -h / | awk 'NR==2 && $5+0 > 80 {print "Disk usage above 80%: "$5}' >> /home/olp/alerts.log
|
||||
```
|
||||
|
||||
**What to watch for (manually, weekly):**
|
||||
1. `olp-keys list` — any unexpected keys?
|
||||
2. Dashboard (`/dashboard`) — unusual request volume? Unknown providers being hit?
|
||||
3. `journalctl -u olp --since "7 days ago" | grep -c ERROR` — error spike?
|
||||
4. Audit log: `grep "fallback" ~/.olp/logs/audit.ndjson | wc -l` — fallback frequency (high = provider instability)
|
||||
|
||||
### Layer 7 — Update and Recovery
|
||||
|
||||
**Principle:** Rollback within 60 seconds. No data loss on failed update.
|
||||
|
||||
**Update procedure:**
|
||||
|
||||
```bash
|
||||
# SSH to cloud VM as opc
|
||||
|
||||
# 1. Snapshot before update (Oracle Cloud console or CLI)
|
||||
# OCI CLI: oci compute boot-volume-backup create ...
|
||||
|
||||
# 2. Pull latest code
|
||||
cd /opt/olp/app
|
||||
sudo -u olp git fetch origin main
|
||||
sudo -u olp git log --oneline HEAD..origin/main # review what's coming
|
||||
|
||||
# 3. Run tests BEFORE deploying
|
||||
sudo -u olp git checkout main
|
||||
sudo -u olp git pull
|
||||
sudo -u olp node test-features.mjs
|
||||
# STOP if tests fail
|
||||
|
||||
# 4. Restart service
|
||||
sudo systemctl restart olp
|
||||
sleep 3
|
||||
sudo systemctl status olp # verify running
|
||||
|
||||
# 5. Smoke test
|
||||
curl -sf -H "Authorization: Bearer $OLP_OWNER_KEY" https://olp.example.com/health | jq .ok
|
||||
# Expect: true
|
||||
```
|
||||
|
||||
**Rollback:**
|
||||
|
||||
```bash
|
||||
# If update breaks things:
|
||||
cd /opt/olp/app
|
||||
sudo -u olp git checkout <previous-tag> # e.g. v0.6.0
|
||||
sudo systemctl restart olp
|
||||
```
|
||||
|
||||
**Backup (automated):**
|
||||
|
||||
```bash
|
||||
# Cron: daily backup of OLP state (keys + config + recent audit)
|
||||
0 3 * * * tar czf /home/opc/backups/olp-state-$(date +\%Y\%m\%d).tar.gz -C /home/olp .olp/keys .olp/config.json .olp/logs/audit.ndjson 2>/dev/null; find /home/opc/backups -name 'olp-state-*' -mtime +30 -delete
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Implementation Checklist
|
||||
|
||||
Execute in order. Each step has a verification gate — do not proceed if the gate fails.
|
||||
|
||||
### Phase A — VM Preparation
|
||||
|
||||
```
|
||||
[ ] A1. SSH to Oracle Cloud VM, verify Node.js >= 18
|
||||
Gate: `node --version` prints v18+
|
||||
|
||||
[ ] A2. Create `olp` system user
|
||||
Gate: `id olp` shows the user exists
|
||||
|
||||
[ ] A3. Clone OLP repo to /opt/olp/app
|
||||
Gate: `sudo -u olp node /opt/olp/app/test-features.mjs` — all tests pass
|
||||
|
||||
[ ] A4. Install provider CLIs (as olp user)
|
||||
- npm install -g @anthropic-ai/claude-code
|
||||
- npm install -g @openai/codex
|
||||
- (mistral vibe if needed)
|
||||
Gate: `which claude && which codex` both resolve
|
||||
|
||||
[ ] A5. Transfer OAuth credentials (Layer 5 procedure)
|
||||
Gate: `sudo -u olp claude auth status` shows authenticated
|
||||
```
|
||||
|
||||
### Phase B — Security Hardening
|
||||
|
||||
```
|
||||
[ ] B1. Configure OCI Security List (Layer 1)
|
||||
Gate: nmap from external IP shows only 22 and 443 open
|
||||
|
||||
[ ] B2. Configure iptables backup (Layer 1)
|
||||
Gate: `sudo iptables -L -n` matches the plan
|
||||
|
||||
[ ] B3. Install + configure Nginx (Layer 2)
|
||||
Gate: `curl -I http://olp.example.com` returns 301 → HTTPS
|
||||
|
||||
[ ] B4. Obtain Let's Encrypt certificate
|
||||
Gate: `curl -I https://olp.example.com` returns valid cert
|
||||
|
||||
[ ] B5. Verify Nginx SSE passthrough
|
||||
Gate: test streaming request completes without timeout
|
||||
```
|
||||
|
||||
### Phase C — OLP Configuration
|
||||
|
||||
```
|
||||
[ ] C1. Write ~/.olp/config.json (Layer 3 — auth config)
|
||||
Gate: config validates (no startup warnings in journal)
|
||||
|
||||
[ ] C2. Generate owner key
|
||||
Gate: `olp-keys list --owner-only` shows 1 owner key
|
||||
|
||||
[ ] C3. Generate family guest keys (one per person)
|
||||
Gate: `olp-keys list` shows correct count
|
||||
|
||||
[ ] C4. Install systemd unit (Layer 4)
|
||||
Gate: `systemctl status olp` shows active (running)
|
||||
|
||||
[ ] C5. Verify /health with owner key
|
||||
Gate: `curl -H "Authorization: Bearer $OWNER_KEY" https://olp.example.com/health | jq .ok` → true
|
||||
|
||||
[ ] C6. Verify /health rejects unauthenticated
|
||||
Gate: `curl https://olp.example.com/health` → 401
|
||||
|
||||
[ ] C7. Verify guest key cannot access /dashboard
|
||||
Gate: `curl -H "Authorization: Bearer $GUEST_KEY" https://olp.example.com/dashboard` → 403
|
||||
|
||||
[ ] C8. End-to-end LLM request with guest key
|
||||
Gate: streaming chat completion returns a valid response
|
||||
```
|
||||
|
||||
### Phase D — Monitoring Setup
|
||||
|
||||
```
|
||||
[ ] D1. Install cron jobs (Layer 6)
|
||||
Gate: `crontab -l` shows all 4 jobs
|
||||
|
||||
[ ] D2. Verify daily backup cron
|
||||
Gate: manual trigger produces valid tar.gz
|
||||
|
||||
[ ] D3. Test health-check alert
|
||||
Gate: stop OLP, wait 5min, check alerts.log has entry
|
||||
```
|
||||
|
||||
### Phase E — Family Onboarding
|
||||
|
||||
```
|
||||
[ ] E1. Send each family member their API key via Signal/iMessage
|
||||
(NOT via email, NOT via any cloud-stored medium)
|
||||
|
||||
[ ] E2. Each family member configures their IDE:
|
||||
export OPENAI_BASE_URL=https://olp.example.com/v1
|
||||
export OPENAI_API_KEY=olp_<their-key>
|
||||
|
||||
[ ] E3. Each family member runs a test prompt
|
||||
Gate: audit.ndjson shows their key_id in the log
|
||||
|
||||
[ ] E4. Verify per-key provider scoping
|
||||
Gate: kid's key cannot hit providers outside their scope
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Security Threat Model
|
||||
|
||||
| Threat | Mitigation | Residual Risk |
|
||||
|---|---|---|
|
||||
| **Brute-force API key** | 32-byte entropy = 2^256 keyspace; Nginx rate limit 30r/m | Negligible |
|
||||
| **TLS downgrade** | TLS 1.3 only; HSTS header | None with modern clients |
|
||||
| **Credential theft (OAuth tokens on VM)** | chmod 600 + dedicated user + no root access to OLP dirs | VM root compromise (mitigated by OCI IAM) |
|
||||
| **Stolen guest key** | Single-key revocation via `olp-keys revoke`; per-key audit trail for forensics | Window between theft and detection |
|
||||
| **DDoS** | OCI DDoS protection (free tier) + Nginx rate limit + Nginx connection limit | Sustained volumetric attack may overwhelm free-tier VM |
|
||||
| **Provider credential abuse** | OLP is the only consumer; anomalous spend visible on provider dashboard | Provider-side detection lag |
|
||||
| **Supply chain (OLP code tampered)** | Git clone from known repo; `npm test` before deploy; no npm dependencies | Compromised maintainer GitHub account |
|
||||
| **Log exfiltration** | audit.ndjson contains no message content (PII guard per ADR 0008); only metadata | Key IDs in logs (low sensitivity) |
|
||||
|
||||
---
|
||||
|
||||
## 5. Operational Runbooks
|
||||
|
||||
### Runbook: OAuth Token Expired
|
||||
|
||||
```
|
||||
Symptom: /health shows provider auth.ok=false; fallback firing on every request
|
||||
Diagnosis: sudo -u olp claude auth status → "not authenticated" or expired
|
||||
|
||||
Fix:
|
||||
1. sudo -u olp claude setup-token
|
||||
2. Complete OAuth flow (browser URL → paste code)
|
||||
3. Verify: sudo -u olp claude auth status → authenticated
|
||||
4. No OLP restart needed — next spawn picks up new credentials
|
||||
```
|
||||
|
||||
### Runbook: Revoke a Compromised Key
|
||||
|
||||
```
|
||||
Symptom: suspicious traffic in audit.ndjson from a specific key_id
|
||||
grep "<suspected-key-id>" ~/.olp/logs/audit.ndjson | tail -20
|
||||
|
||||
Fix:
|
||||
1. olp-keys revoke --id=<key-id>
|
||||
2. Notify family member: "Your key was revoked. Here's a new one."
|
||||
3. olp-keys keygen --name=<new-name> --providers=<same-providers>
|
||||
4. Send new key via secure channel
|
||||
```
|
||||
|
||||
### Runbook: VM Disk Full
|
||||
|
||||
```
|
||||
Symptom: OLP stops writing audit logs; new requests may fail
|
||||
Diagnosis: df -h /
|
||||
|
||||
Fix:
|
||||
1. Purge old audit logs: find ~/.olp/logs -name 'audit-202*.ndjson' -mtime +90 -delete
|
||||
2. Purge old backups: find /home/opc/backups -name 'olp-state-*' -mtime +60 -delete
|
||||
3. Purge cache if needed: rm -rf ~/.olp/cache/*
|
||||
4. Verify: df -h / shows >20% free
|
||||
```
|
||||
|
||||
### Runbook: OLP Process Crash Loop
|
||||
|
||||
```
|
||||
Symptom: systemctl status olp shows "activating (auto-restart)"
|
||||
Diagnosis: journalctl -u olp --since "10 min ago" | tail -50
|
||||
|
||||
Common causes:
|
||||
- Port conflict → check `lsof -nP -iTCP:4567`
|
||||
- Corrupt config.json → validate JSON syntax
|
||||
- Node.js version drift → `node --version`
|
||||
|
||||
Fix:
|
||||
1. Fix root cause
|
||||
2. sudo systemctl restart olp
|
||||
3. Gate: `curl -H "Authorization: Bearer $OWNER_KEY" https://olp.example.com/health | jq .ok`
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Cost Estimate (Oracle Cloud Free Tier)
|
||||
|
||||
| Resource | Spec | Cost |
|
||||
|---|---|---|
|
||||
| VM | ARM Ampere A1 (4 OCPU, 24GB RAM) | **Free** (Always Free tier) |
|
||||
| Boot volume | 200GB | **Free** (up to 200GB) |
|
||||
| Outbound bandwidth | 10TB/month | **Free** (first 10TB) |
|
||||
| Public IP | 1 reserved | **Free** |
|
||||
| Domain | olp.example.com | ~$10/year (external registrar) |
|
||||
| TLS cert | Let's Encrypt | **Free** |
|
||||
| **Total** | | **~$10/year** (domain only) |
|
||||
|
||||
Oracle Cloud's Always Free ARM VM is overprovisioned for this use case. OLP + Nginx + 3 provider CLIs will use <1GB RAM and negligible CPU (the LLM inference happens at the provider, not here).
|
||||
|
||||
---
|
||||
|
||||
## 7. Migration Path to Commercial
|
||||
|
||||
This family deployment is a stepping stone. When commercial service is ready:
|
||||
|
||||
| Aspect | Family (this plan) | Commercial (future) |
|
||||
|---|---|---|
|
||||
| Upstream | spawn CLI (subscription) | direct API (commercial key) |
|
||||
| Auth | OLP multi-key (filesystem) | Registration + billing system |
|
||||
| TLS | Let's Encrypt (single domain) | Managed cert (Cloudflare / AWS ACM) |
|
||||
| Compute | Single VM (Oracle Free) | Container cluster (auto-scale) |
|
||||
| Monitoring | Cron + manual | Prometheus + Grafana + PagerDuty |
|
||||
| Rate limit | Nginx per-IP | Per-key token bucket in OLP |
|
||||
| Data | ~/.olp/ filesystem | PostgreSQL + S3 |
|
||||
|
||||
The deployment experience from this plan directly informs the commercial architecture. Every operational runbook becomes a feature requirement for the commercial platform.
|
||||
|
||||
---
|
||||
|
||||
**Authors:** project maintainer (with AI drafting assistance)
|
||||
**Created:** 2026-05-27
|
||||
@@ -0,0 +1,274 @@
|
||||
# PI231 Spike — Ephemeral $HOME / $CODEX_HOME Override Verification
|
||||
|
||||
**Date:** 2026-05-29
|
||||
**Operator:** project maintainer (via PI231 SSH)
|
||||
**Spike artifact:** `tlab@172.16.2.231:/tmp/olp-spike-20260529-100243/`
|
||||
**ADR context:** ADR 0014 Amendment 1 § A1.2 Layer 1 — "Per-spawn ephemeral home directory"
|
||||
**Task ref:** OLP task list #4 ("PI231 spike — verify Claude / Codex HOME / CODEX_HOME override behavior")
|
||||
|
||||
---
|
||||
|
||||
## TL;DR
|
||||
|
||||
**Both providers PASS.** Setting `HOME` (claude) and `CODEX_HOME` (codex) before spawn redirects 100% of CLI state writes into the ephemeral location. Real `~/.claude/`, `~/.claude.json`, and `~/.codex/` were unmodified by the spike. Credentials accessed via symlink work end-to-end (model returned "PONG" for both providers). Solution 1 is implementable today; Tasks #5-#8 unblocked.
|
||||
|
||||
One non-blocking caveat for codex (PATH-helper installation refused under `/tmp` paths — § Caveats).
|
||||
|
||||
Vibe (mistral) not on PI231; pinned to a follow-up spike when the CLI is installed.
|
||||
|
||||
---
|
||||
|
||||
## 1. Environment
|
||||
|
||||
| Component | Value |
|
||||
|---|---|
|
||||
| Host | `tlab@172.16.2.231` (RPi4-P8-231, Debian Bookworm arm64) |
|
||||
| Real `~` | `/home/tlab` |
|
||||
| `claude` | `/home/tlab/.npm-global/bin/claude` — v2.1.152 |
|
||||
| `codex` | `/home/tlab/.npm-global/bin/codex` — v0.133.0 |
|
||||
| `vibe` | not installed |
|
||||
| Prod OLP | running (port 4567 with `OLP_SANDBOX_DISABLED=1`) — spike does not interfere |
|
||||
|
||||
Pre-state mtimes (from spike `pre-mtimes.txt`):
|
||||
```
|
||||
1779999955 /home/tlab/.claude.json
|
||||
1779999956 /home/tlab/.claude/.credentials.json
|
||||
1779759544 /home/tlab/.codex/auth.json
|
||||
```
|
||||
|
||||
Marker file timestamps (pre-spike) used to detect any post-spike write to real home.
|
||||
|
||||
---
|
||||
|
||||
## 2. Methodology
|
||||
|
||||
Both providers tested per the same skeleton:
|
||||
|
||||
```bash
|
||||
SPIKE_ROOT=/tmp/olp-spike-<timestamp>
|
||||
mkdir -p $SPIKE_ROOT/<provider>-home/.<provider>
|
||||
ln -s ~/.<provider>/<credential-file> $SPIKE_ROOT/<provider>-home/.<provider>/<credential-file>
|
||||
|
||||
<ENV_OVERRIDE>=<path> timeout 90 <provider> <invocation> "say PONG and nothing else"
|
||||
|
||||
find $SPIKE_ROOT/<provider>-home -printf "%y %M %s %p\n" # what landed in fake home
|
||||
find ~/.<provider> ~/.<provider>.json -newer <marker> # did real home get modified
|
||||
```
|
||||
|
||||
The `find -newer <marker>` test is the load-bearing assertion: if it returns **empty**, the redirect held perfectly. If it returns any path, the CLI silently fell back to the real `$HOME`-derived path despite the env override.
|
||||
|
||||
---
|
||||
|
||||
## 3. Phase B — claude CLI (anthropic)
|
||||
|
||||
### 3.1 Invocation
|
||||
|
||||
```bash
|
||||
HOME=$SPIKE_ROOT/claude-home timeout 60 claude \
|
||||
--print "say PONG and nothing else" \
|
||||
--no-session-persistence \
|
||||
--model claude-sonnet-4-6
|
||||
```
|
||||
|
||||
Credentials linked: `$SPIKE_ROOT/claude-home/.claude/.credentials.json` → `/home/tlab/.claude/.credentials.json`
|
||||
|
||||
### 3.2 Result
|
||||
|
||||
```
|
||||
PONG
|
||||
```
|
||||
|
||||
Exit 0. Model returned through Anthropic API via OAuth token from the symlinked real credentials. End-to-end success.
|
||||
|
||||
### 3.3 Fake home post-state (decisive evidence)
|
||||
|
||||
Files written under `$SPIKE_ROOT/claude-home/`:
|
||||
|
||||
```
|
||||
.claude/.credentials.json (symlink — unchanged)
|
||||
.claude/projects/-home-tlab/<uuid>.jsonl (135 bytes — project transcript)
|
||||
.claude/projects/-home-tlab/memory/ (created)
|
||||
.claude/sessions/ (drwx------ private)
|
||||
.claude/backups/.claude.json.backup.1780012965052 (50 bytes — pre-write backup)
|
||||
.claude.json (23,182 bytes — fresh)
|
||||
.cache/claude-cli-nodejs/-home-tlab/mcp-logs-claude-ai-Gmail/<ts>.jsonl
|
||||
.cache/claude-cli-nodejs/-home-tlab/mcp-logs-claude-ai-Google-Calendar/<ts>.jsonl
|
||||
.cache/claude-cli-nodejs/-home-tlab/mcp-logs-claude-ai-Google-Drive/<ts>.jsonl
|
||||
```
|
||||
|
||||
**`.claude.json` (23 KB) was written to the ephemeral location.** This is the file whose non-atomic write upstream (anthropics/claude-code#29250) drove ADR 0014 Amendment 1 § A1.1.2. The Solution 1 architecture removes the maintenance-treadmill concern by letting this file land in tmpfs — confirmed working.
|
||||
|
||||
`projects/-home-tlab/` — claude encodes the spawn CWD (`/home/tlab`) by replacing `/` with `-`. Not relevant to isolation; would also be the path if claude ran with the real `$HOME`.
|
||||
|
||||
### 3.4 Real home post-state
|
||||
|
||||
```bash
|
||||
$ find ~/.claude.json ~/.claude -newer $SPIKE_ROOT/marker
|
||||
# (empty)
|
||||
|
||||
$ stat -c "%Y %n" ~/.claude.json ~/.claude/.credentials.json
|
||||
1779999955 /home/tlab/.claude.json
|
||||
1779999956 /home/tlab/.claude/.credentials.json
|
||||
```
|
||||
|
||||
Both mtimes identical to pre-state. **Real `~/.claude.json` was not touched by the spike.**
|
||||
|
||||
### 3.5 Verdict
|
||||
|
||||
✅ **PASS.** claude v2.1.152 honours `HOME` env override completely. All state writes redirect to the ephemeral location. Symlinked credentials work for auth. The Layer 1 + Layer 2 architecture per ADR 0014 Amendment 1 § A1.2 is implementable for anthropic without further work.
|
||||
|
||||
---
|
||||
|
||||
## 4. Phase B — codex CLI (openai)
|
||||
|
||||
### 4.1 Invocation (final, working)
|
||||
|
||||
The first attempt used `--ask-for-approval never` per docs found in pre-spike research — that flag has been **removed in codex v0.133.0**. Help output shows it must be passed as a config override: `-c approval_policy="never"`. Retry:
|
||||
|
||||
```bash
|
||||
echo "say PONG and nothing else" | \
|
||||
HOME=$SPIKE_ROOT/codex-home \
|
||||
CODEX_HOME=$SPIKE_ROOT/codex-home/.codex \
|
||||
timeout 90 codex exec \
|
||||
--skip-git-repo-check \
|
||||
-c approval_policy=\"never\" \
|
||||
--sandbox read-only \
|
||||
"say PONG and nothing else"
|
||||
```
|
||||
|
||||
Credentials linked: `$SPIKE_ROOT/codex-home/.codex/auth.json` → `/home/tlab/.codex/auth.json`
|
||||
|
||||
### 4.2 Result
|
||||
|
||||
```
|
||||
WARNING: proceeding, even though we could not update PATH: Refusing to create
|
||||
helper binaries under temporary dir "/tmp"
|
||||
(codex_home: AbsolutePathBuf("/tmp/olp-spike-20260529-100243/codex-home/.codex"))
|
||||
Reading additional input from stdin...
|
||||
OpenAI Codex v0.133.0
|
||||
--------
|
||||
workdir: /home/tlab
|
||||
model: gpt-5.5
|
||||
provider: openai
|
||||
approval: never
|
||||
sandbox: read-only
|
||||
reasoning effort: none
|
||||
reasoning summaries: none
|
||||
session id: 019e710b-dc28-79f0-854a-06be116b4830
|
||||
--------
|
||||
user
|
||||
say PONG and nothing else
|
||||
...
|
||||
codex
|
||||
PONG
|
||||
tokens used
|
||||
10,826
|
||||
```
|
||||
|
||||
Exit 0. Model invoked, returned "PONG", session id assigned.
|
||||
|
||||
**The `WARNING` is significant — see § 5 Caveats. Key fact:** the warning's path embed `codex_home: AbsolutePathBuf("/tmp/olp-spike-…")` proves `CODEX_HOME` was parsed and honoured. The warning is a *narrow* refusal (PATH helper binary install), not a refusal of `CODEX_HOME` itself.
|
||||
|
||||
### 4.3 Fake home post-state (decisive evidence)
|
||||
|
||||
```
|
||||
.codex/auth.json (symlink — unchanged)
|
||||
.codex/models_cache.json (200,842 bytes)
|
||||
.codex/installation_id (36 bytes)
|
||||
.codex/cache/codex_apps_tools/<hash>.json (92,600 bytes)
|
||||
.codex/goals_1.sqlite (24,576 bytes)
|
||||
.codex/logs_2.sqlite (49,152 bytes)
|
||||
.codex/state_5.sqlite (180,224 bytes)
|
||||
.codex/shell_snapshots/ (created)
|
||||
.codex/memories/ (created)
|
||||
.codex/skills/ (created)
|
||||
.codex/sessions/2026/05/29/ (date-partitioned)
|
||||
.codex/.tmp/plugins-clone-<rand>/.git/... (cloned plugins repo)
|
||||
```
|
||||
|
||||
State scale: ~500 KB across 3 SQLite DBs + model cache + plugin checkout. Far more than claude writes. **All of it landed in the ephemeral location.**
|
||||
|
||||
### 4.4 Real home post-state
|
||||
|
||||
```bash
|
||||
$ find ~/.codex -newer $SPIKE_ROOT/codex-marker3
|
||||
# (empty)
|
||||
|
||||
$ stat -c "%Y %n" ~/.codex/auth.json
|
||||
1779759544 /home/tlab/.codex/auth.json
|
||||
```
|
||||
|
||||
Mtime unchanged. **Real `~/.codex` was not touched by the spike.**
|
||||
|
||||
### 4.5 Verdict
|
||||
|
||||
✅ **PASS.** codex v0.133.0 honours `CODEX_HOME` env override for ALL state files. Symlinked auth artifact works for API authentication. The codex inner sandbox (read-only by default per ADR 0002 Amendment 9 § Per-provider codex declaration) initialized and ran without error.
|
||||
|
||||
---
|
||||
|
||||
## 5. Caveats
|
||||
|
||||
### 5.1 codex PATH helper warning
|
||||
|
||||
Codex's startup includes a step that tries to install helper binaries into PATH (presumably under `$CODEX_HOME/bin/` or similar). When `$CODEX_HOME` is under `/tmp/`, codex refuses this step for security reasons (anti-prefix-attack on PATH):
|
||||
|
||||
```
|
||||
WARNING: proceeding, even though we could not update PATH:
|
||||
Refusing to create helper binaries under temporary dir "/tmp"
|
||||
```
|
||||
|
||||
**Impact for OLP**: none of the load-bearing functionality is affected. The model invocation completed, auth worked, all session state landed in `$CODEX_HOME`. The skipped step is for shell-completion-style helpers that the spawn-binary architecture does not need.
|
||||
|
||||
**If we ever do need those helpers**: ephemeral root would need to move out of `/tmp/`. Candidates: `/var/lib/olp-spawn/<keyId>/<reqId>/` (operator-managed) or `~/.olp/spawn/<keyId>/<reqId>/` (within OLP's own data root). Decision deferred — not required for Phase 7 implementation.
|
||||
|
||||
### 5.2 claude project-path encoding (`-home-tlab`)
|
||||
|
||||
claude encodes the spawn cwd into project paths by replacing `/` with `-`. The encoded value reflects the **real cwd at spawn time** (`/home/tlab` → `-home-tlab`), not the ephemeral `$HOME`. This is expected: cwd is a separate input from `$HOME`.
|
||||
|
||||
**Impact for OLP**: none. The encoding is internal to claude's project tracking. OLP spawn pipeline already runs each request from a per-spawn cwd if it wants to isolate cwd separately; that is orthogonal to Layer 1's `$HOME` redirect.
|
||||
|
||||
### 5.3 codex v0.133.0 flag set drift
|
||||
|
||||
The pre-spike research cited `--ask-for-approval never` as the non-interactive approval flag (sourced from OpenAI docs pages indexed before v0.133.0 changed the flag layout). v0.133.0 instead requires `-c approval_policy="never"` via the generic config-override flag. ADR 0002 Amendment 9 § codex `toolHardeningArgs` declaration uses `--sandbox read-only` which is still a valid top-level flag; no amendment update required. **Implementation note (Task #7)**: codex.mjs `toolHardeningArgs` should not inject `--ask-for-approval` — use `-c approval_policy="never"` if the policy needs to be locked at spawn time.
|
||||
|
||||
### 5.4 Mistral `vibe` CLI not present on PI231
|
||||
|
||||
`which vibe` returned empty. Vibe is not currently part of the PI231 test deployment per the topology memory (`~/.cc-rules/memory/projects/olp/topology_pi231_server_2026_05_27.md`). The ADR 0002 Amendment 9 mistral declaration uses `VIBE_HOME` per the Mistral docs page (3 occurrences verified at amendment time). The observed-behavior verification is a follow-up spike triggered when vibe is installed.
|
||||
|
||||
---
|
||||
|
||||
## 6. Implications for ADR 0014 Amendment 1
|
||||
|
||||
| Architecture claim | Spike result |
|
||||
|---|---|
|
||||
| Layer 1 (ephemeral `$HOME` / `$CODEX_HOME`) is implementable | ✅ Confirmed for anthropic + codex |
|
||||
| `~/.claude.json` upstream non-atomic-write concern is solved by redirect | ✅ Confirmed — write lands in tmpfs `.claude.json`, real one untouched |
|
||||
| Layer 2 (symlinked credentials) preserves auth | ✅ Confirmed — both providers authenticated via symlink |
|
||||
| codex inner sandbox composes with Layer 1 (no nested-bwrap conflict) | ✅ Confirmed — codex `--sandbox read-only` initialized and ran |
|
||||
| Solution 1 obsoletes outer-bwrap maintenance treadmill | ✅ Confirmed — no `--ro-bind` mount patches required |
|
||||
|
||||
No architectural changes required. ADR 0014 Amendment 1 is **validated by primary-source observation on the target deployment**.
|
||||
|
||||
---
|
||||
|
||||
## 7. Unblocked / next
|
||||
|
||||
Tasks unblocked by this spike's PASS verdict:
|
||||
- Task #5 — refactor `lib/sandbox/manager.mjs` to `prepareIsolatedEnvironment()` per Layer 1 + Layer 2
|
||||
- Task #6 — add `ISOLATION` block to `lib/providers/anthropic.mjs`
|
||||
- Task #7 — add `ISOLATION` block to `lib/providers/codex.mjs` (use `-c approval_policy="never"` per § 5.3, not `--ask-for-approval`)
|
||||
- Task #8 — wire `prepareIsolatedEnvironment` into `server.mjs` spawn site
|
||||
|
||||
Follow-ups not blocking:
|
||||
- Vibe spike when CLI is installed (verify documented `VIBE_HOME` behavior matches observed)
|
||||
- codex PATH-helper out-of-`/tmp` consideration if the helpers ever become required
|
||||
|
||||
---
|
||||
|
||||
## 8. Artifact retention
|
||||
|
||||
The spike root `/tmp/olp-spike-20260529-100243/` on PI231 is automatically cleaned by tmpfs lifetime / reboot. No commit of binary artifacts. Evidence above is the canonical record.
|
||||
|
||||
---
|
||||
|
||||
**Authored** by project maintainer 2026-05-29; commands executed on PI231 with maintainer's SSH session.
|
||||
+37
-14
@@ -8,21 +8,23 @@
|
||||
3. **Where** does the work live in the tree today (file + anchor).
|
||||
4. **When** does it need to land (trigger: load profile, security event, governance amendment).
|
||||
|
||||
**Reading order for a v1.x sprint kickoff.** Items #1–#3 are the most architecturally consequential and should be designed in dependency order: #2 (multi-key auth) blocks header gating in #1 and observability ownership in #4. #1 (streaming SF) blocks #5 (soft trigger reactivation) only if soft triggers are wired on streaming requests.
|
||||
**Reading order for a v1.x sprint kickoff.** As of 2026-05-27, #1 (streaming SF, D57+D58), #2 (multi-key auth, Phase 2), #4, #7, and #8 are CLOSED. Remaining v1.x scope: #3 (soft trigger reactivation), #5 (provider cacheKeyFields mask), #6 (streaming SPAWN_FAILED salvage — unbundled from #1 at #1 close). All three remaining items have explicit "trigger to start" gates that have not fired.
|
||||
|
||||
---
|
||||
|
||||
## #1 — Streaming-path singleflight + TOCTOU close
|
||||
## #1 — Streaming-path singleflight + TOCTOU close — ✅ **SHIPPED (D57 + D58, 2026-05-25)**
|
||||
|
||||
- **What.** `cacheStore.getOrComputeStreaming(keyId, cacheKey, sourceFactory)` API replacing the current `peek + spawn` pattern in `server.mjs`. Per-(keyId, cacheKey) inflight Map with tee fan-out, bounded per-client backpressure queues, late-joiner replay buffer, AbortController propagation on all-disconnect.
|
||||
- **Why deferred.** Personal/family-scale single-tenant load — N concurrent identical streaming requests is an edge case that has not been reported. Each concurrent caller receives the correct response; the waste is N CLI processes instead of one.
|
||||
- **Design ADR (ratified).** [`docs/adr/0005-cache-cross-provider.md` Amendment 8](./adr/0005-cache-cross-provider.md) — full design including the inflight Map shape, tee policy, late-joiner replay, backpressure cap, D38 semaphore coordination, abort policy, cache TTL race handling, observability event set, and X-OLP-Streaming-Inflight header. Implementation acceptance criteria are in Amendment 8 §13.
|
||||
- **Tracking issue.** GitHub issue [#16](https://github.com/dtzp555-max/olp/issues/16) — STAYS OPEN as v1.x tracker. Sibling: the closed-but-not-implemented Amendment 6 deferral (D34 F1).
|
||||
- **Code anchors today.**
|
||||
- `server.mjs` lines ~782 (`preCheckHit = await cacheStore.peek(...)`) and ~811–817 (streaming branch entry) — these are the lines the new API replaces.
|
||||
- `lib/cache/store.mjs` `getOrCompute` — sibling API; the new one mirrors its shape on the streaming path.
|
||||
- **Trigger to start.** Any of: (a) report of N>1 concurrent identical streaming requests in the wild, (b) v1.x sprint planning kickoff with the maintainer explicitly opening this scope, (c) downstream feature requiring tee-streaming primitive (e.g., browser-side observer attaching to an existing stream).
|
||||
- **Estimated effort.** Design ADR ratified (Amendment 8) = 30 min done. Implementation = 200–400 lines + 15-20 tests + fresh-context reviewer pass. ~3-4 hours of subagent runtime with full Iron Rule 10 discipline.
|
||||
- **Status.** Closed. Trigger (b) fired 2026-05-25 — maintainer "go" after v0.3.1. Shipped across three D-days:
|
||||
- **D57** (PR #36) — cache layer: `cacheStore.getOrComputeStreaming(keyId, cacheKey, sourceFactory, opts) → { stream, isFirst, role }` with `_streamingInflight` Map, tee fan-out, late-joiner replay buffer, per-client backpressure (`PER_CLIENT_QUEUE_CAP=1MB`), replay cap (`ACCUMULATED_REPLAY_CAP=10MB`), AbortController propagation, synchronous Map check+insert (closes TOCTOU). Suite 27 = 12 unit tests.
|
||||
- **D58** (PR #37) — server.mjs wiring: streaming branch swap; `tryAcquireSpawn`/`releaseSpawn` moved inside `sourceFactory` closure (D38 §7 coordination); `CONCURRENCY_LIMIT` fallthrough preserved; `X-OLP-Streaming-Inflight: source | attached` header; `cache_status: 'streaming_attached'` audit value + audit-query gauge reconciliation; `res.on('close') → stream.return()` for client disconnect; D16 truncated-not-cached invariant preserved via `cacheStore.delete` on stop-less exhaustion. Suite 28 = 8 HTTP integration tests.
|
||||
- **D59** (this commit) — docs polish: README known-limitations entry inverted; this roadmap entry closed; issue #16 closed.
|
||||
- **Design authority.** [`docs/adr/0005-cache-cross-provider.md` Amendment 8](./adr/0005-cache-cross-provider.md) — implemented per spec §§1–14 across D57+D58.
|
||||
- **Tracking issue.** GitHub issue [#16](https://github.com/dtzp555-max/olp/issues/16) — CLOSED at D59 with refs to D57+D58 PRs.
|
||||
- **Final test count delta.** 603 (v0.3.1) → 623 (v0.3.2/v0.4.0). +20 tests across the SF arc.
|
||||
- **Deferred sub-items** (left here as future-work pointers, NOT blocking #1 closure):
|
||||
- `X-OLP-Streaming-Inflight: solo` value not emitted on the wire (Amendment 8 §11). It's observable only post-stream via the `streaming_inflight_source_done` log event's `attached_count: 0`. Future ADR amendment may expose via HTTP trailer.
|
||||
- `streaming_inflight_join` log event from `_attachClient` cache-layer path (carries no provider/model context). D58 emits the event from the server-layer wrapper instead; cache-layer emission would need a provider/model plumb (TODO marker at `lib/cache/store.mjs:~620`).
|
||||
- `isFirst` field returned by `getOrComputeStreaming` is currently unused by server.mjs (`role` supersedes). Could be removed in a future cache-layer API cleanup.
|
||||
|
||||
## #2 — Multi-key auth (`lib/keys.mjs`) — **PHASE 2 ACTIVE (no longer deferred)**
|
||||
|
||||
@@ -81,12 +83,33 @@
|
||||
|
||||
- **What.** Currently the streaming branch does NOT participate in D16 salvage (the salvage-on-SPAWN_FAILED + chunks pattern that the buffered path uses). Streaming SPAWN_FAILED mid-stream → the truncation marker (D35 #10) fires, but no salvage logic captures partial chunks for downstream cache reuse.
|
||||
- **Why deferred.** Less impactful than #1 — at most one client benefits per spawn event, and the buffered path already provides salvage for the bulk of requests. Streaming is the minority path.
|
||||
- **Design ADR.** Not yet ratified. Coordinated with #1 because the tee architecture changes the salvage semantics (multiple clients may want different finish_reason interpretations on source-mid-stream-failure).
|
||||
- **Status update post-#1 close (2026-05-25).** #1 was originally bundled with #6 in the design ADR (Amendment 8). The tee architecture as implemented does NOT carry salvage semantics — D57's tee writes `accumulatedChunks` to cache only on normal source completion (stop chunk seen); on SPAWN_FAILED mid-stream the cache layer rejects all clients with the error and does NOT persist partial chunks. D58 preserves D16's truncated-not-cached invariant via server-layer `cacheStore.delete` on stop-less exhaustion. #6 therefore remains independently deferrable.
|
||||
- **Design ADR.** Not yet ratified. The unbundling from #1 means #6 now needs its own ADR amendment when triggered.
|
||||
- **Tracking.** Not a GitHub issue. Tracked here.
|
||||
- **Trigger to start.** Bundled with #1 implementation work (the inflight tee architecture changes the salvage semantics, so designing them together is cheaper than serializing).
|
||||
- **Trigger to start.** First report of streaming-path SPAWN_FAILED mid-stream where partial-chunk salvage would have helped a downstream caller. Practically unlikely at family scale.
|
||||
|
||||
## #7 — AUTH_MISSING tuple path test coverage (D40 follow-up)
|
||||
## #8 — Dashboard enrichment: per-provider subscription quota + reset times + 1-min refresh + manual refresh (D78 follow-up) — ✅ **CLOSED (D82, v0.5.0)**
|
||||
|
||||
- **Status.** Closed at D82 (Phase 5). `dashboard.html` restructured to Claude.ai-style per-provider rows rendering `quota_v2`. Closed by PR on branch `d82-dashboard-ui-claude-ai-style`; ships with v0.5.0. 60s quota auto-refresh + manual refresh button + visibilityState guard implemented. Graceful fallback to legacy `quota` field when server runs a pre-D81 build.
|
||||
- **What.** Phase 3 dashboard (D51 `dashboard.html`, v0.3.0) shows: per-provider quota (currently always "n/a — no quota api"), last-24h request count + cache hit + fallback rate, 30d request-count sparkline, top fallback chains. **Maintainer request 2026-05-26 post-D78**: extend to show what each enabled provider's subscription is actually consuming, with reset times visible, refresh once per minute (current 30s is OK but maintainer specified 1min target), and a manual refresh button. Reference design: Claude.ai's own `claude.ai/settings/usage` page — current session bar with "Resets in 1hr 6min", weekly all-models bar with "Resets Sun 9:00 PM", per-model bar (Sonnet only), additional features (routine runs), usage credits + monthly spend limit + auto-reload toggle.
|
||||
- **Why deferred.** v0.3.0/v0.4.x ships the dashboard frame but `provider.quotaStatus()` returns `null` in all three v0.1 plugins (anthropic / openai / mistral). The ratifying spec in ADR 0004 Amendment 2 punts `quotaStatus()` to v1.x ("soft trigger reactivation") — this dashboard ask is the **operator-facing reason** that work would land.
|
||||
- **What this requires.** Per-provider plugin work + dashboard.html UI work + audit-query.mjs aggregation:
|
||||
1. **`lib/providers/anthropic.mjs quotaStatus()`** — discover where the maintainer's Claude.ai subscription quota state is exposed. Candidates: (a) `claude` CLI command (e.g., `claude usage`) if Anthropic adds one — currently absent; (b) parsing the `claude-code` output for rate-limit error messages and caching state from headers; (c) hitting `api.anthropic.com/v1/.../usage` directly via the OAuth refresh token — not a documented endpoint, primary-source risk. ADR 0002 Rule 1 / Rule 5 require an authority citation before any implementation. Likely path: **wait until Anthropic publishes a documented endpoint**, OR derive from audit-side request counts only (no real quota truth, just "you sent N requests in the current 5h window").
|
||||
2. **`lib/providers/openai.mjs quotaStatus()`** — codex CLI doesn't expose ChatGPT-subscription quota state. OpenAI rate-limit headers per request might be parseable but ADR 0004 Amendment 2 explicitly says no plugin parses HTTP status at v0.1.
|
||||
3. **`lib/providers/mistral.mjs quotaStatus()`** — Le Chat Pro has `/v1/usage` endpoint per Mistral docs (verify).
|
||||
4. **`dashboard.html` UI restructure** to a Claude.ai-style layout: rows of (label, bar, "Resets in X" / "Resets at <day-of-week> <time>", percent). Add a manual refresh button + change auto-poll from 30s → 60s. Optionally a usage-credits / per-key spend display if Phase 5 ships per-key cost weights.
|
||||
5. **`lib/audit-query.mjs`** — extend `aggregateRequests` / `spendTrendDaily` to compute "in the current rolling window" (since session/week start) per provider. Today's aggregates are wall-clock windows; subscription resets are per-account-anchored. Need a way to model session windows (e.g., "Anthropic 5h-from-first-request-since-last-reset").
|
||||
- **Reference (maintainer 2026-05-26).** Screenshot of `claude.ai/settings/usage` shared inline. Key panels: Plan usage limits (current session + resets-in), Weekly limits (All models / Sonnet only / per-feature breakdown, each with resets-on), Additional features (Daily included routine runs N / 15), Usage credits (toggle + spent vs monthly limit + auto-reload + buy-credits link).
|
||||
- **Tracking.** Not yet a GitHub issue. Track here + cross-reference ADR 0004 Amendment 2 (soft trigger reactivation — same `quotaStatus()` data-source work) when this becomes Phase 5 scope.
|
||||
- **Code anchors today.**
|
||||
- `dashboard.html` — current 4 panels; needs restructure to Claude.ai-style row layout
|
||||
- `lib/providers/anthropic.mjs` / `openai.mjs` / `mistral.mjs` — `quotaStatus()` returns null today
|
||||
- `lib/audit-query.mjs` — current `aggregateRequests` is wall-clock-window; needs session-window variant
|
||||
- **Trigger to start.** ANY of: (a) Anthropic publishes a documented `claude usage` CLI or `api.anthropic.com/v1/usage` endpoint, (b) maintainer hits real "I want to see quota right now" pain often enough to design without per-provider truth (audit-derived only), (c) Phase 5 multi-tenant adds per-key spend limits and the dashboard needs to surface those.
|
||||
|
||||
## #7 — AUTH_MISSING tuple path test coverage (D40 follow-up) — ✅ **CLOSED (D56, 2026-05-27)**
|
||||
|
||||
- **Status.** Closed. Test shipped at D56 (PR `f4-cli-plugin-quota-v2-plus-auth-missing-test`, 2026-05-27). Test: `test-features.mjs` line 6255 — `'engine: AUTH_MISSING terminates chain, fallbackDetail tuple records trigger_type:"auth_missing" (D56, v1.x roadmap #7)'`. Asserts: `result.fallbackDetail[0].code === 'AUTH_MISSING'`, `result.fallbackDetail[0].trigger_type === 'auth_missing'`, `result.fallbackHops === 0` (no advance). The test was already present in the file before this PR closed the roadmap entry.
|
||||
- **What.** Dedicated test in `test-features.mjs` Suite D40 that asserts the `fallbackDetail` tuple records the AUTH_MISSING path with `trigger_type: 'auth_missing'`. D40 reviewer flagged this as the last gap in the engine-path matrix; code is structurally correct, just lacks an explicit pin.
|
||||
- **Why deferred.** Low priority — the AUTH_MISSING early-return branch has the tuple push BEFORE it (verified in D40 reviewer pass), so coverage is implicit via the other engine-path tests. A 3-line dedicated test would make the pin explicit.
|
||||
- **Design.** No ADR needed. ~5-line test addition.
|
||||
|
||||
+211
-6
@@ -227,7 +227,7 @@ export function* readAuditWindow({ startMs, endMs, olpHome, logEvent } = {}) {
|
||||
* {
|
||||
* window: { startMs, endMs },
|
||||
* request_count, status_2xx, status_4xx, status_5xx,
|
||||
* by_provider: { [providerKey]: { count, cache_hit, cache_miss, cache_bypass, fallback_count } },
|
||||
* by_provider: { [providerKey]: { count, cache_hit, cache_miss, cache_bypass, cache_streaming_attached, fallback_count } },
|
||||
* by_owner_tier: { owner: N, guest: N, anonymous: N },
|
||||
* by_path: { '/v1/chat/completions': N, '/v1/models': N, ... },
|
||||
* median_latency_ms, p95_latency_ms,
|
||||
@@ -276,12 +276,17 @@ export function aggregateRequests({ windowMs, olpHome, logEvent, _nowFn } = {})
|
||||
// By provider
|
||||
if (typeof ev.provider === 'string' && ev.provider.length > 0) {
|
||||
const p = result.by_provider[ev.provider] ??= {
|
||||
count: 0, cache_hit: 0, cache_miss: 0, cache_bypass: 0, fallback_count: 0,
|
||||
count: 0, cache_hit: 0, cache_miss: 0, cache_bypass: 0, cache_streaming_attached: 0, fallback_count: 0,
|
||||
};
|
||||
p.count++;
|
||||
if (ev.cache_status === 'hit') p.cache_hit++;
|
||||
else if (ev.cache_status === 'miss') p.cache_miss++;
|
||||
else if (ev.cache_status === 'bypass') p.cache_bypass++;
|
||||
// D58 — ADR 0005 Amendment 8 §11 + lib/audit.mjs cache_status enum: streaming
|
||||
// singleflight joiners (attached) share the source spawn but did not hit a
|
||||
// cache. Tracked separately so `count` and `cache_hit + cache_miss +
|
||||
// cache_bypass + cache_streaming_attached` reconcile.
|
||||
else if (ev.cache_status === 'streaming_attached') p.cache_streaming_attached++;
|
||||
if (typeof ev.fallback_hops === 'number' && ev.fallback_hops > 0) p.fallback_count++;
|
||||
}
|
||||
|
||||
@@ -438,12 +443,208 @@ export function spendTrendDaily({ days, olpHome, logEvent, _nowFn } = {}) {
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Normalize a single quotaStatus() return value from the anthropic plugin into
|
||||
* the dashboard-friendly shape (D81 / ADR 0008 Amendment).
|
||||
* Provider-specific: called only for 'anthropic'. Returns null if the raw
|
||||
* shape is absent or malformed.
|
||||
*
|
||||
* v0.5.1: handles probe_status field (F3 — ADR 0013 Rule 6).
|
||||
* Accepts both old shape (stale: boolean) and new shape (probe_status: string).
|
||||
*
|
||||
* @internal — used by aggregateProviderQuota()
|
||||
*/
|
||||
function _normalizeAnthropicQuota(raw) {
|
||||
if (!raw || typeof raw !== 'object') return null;
|
||||
const f = raw.fields ?? {};
|
||||
// v0.5.1: probe_status field (new) takes precedence; fall back to stale bool for compat.
|
||||
const probeStatus = raw.probe_status ?? (raw.stale === true ? 'stale' : 'live');
|
||||
return {
|
||||
schema_version: raw.schemaVersion ?? null,
|
||||
last_fresh_at: (probeStatus === 'stale')
|
||||
? (raw.last_fresh_at ?? null)
|
||||
: (raw.probedAt ?? null),
|
||||
utilization: probeStatus === 'unreachable' ? null : {
|
||||
'5h': f.utilization_5h ?? null,
|
||||
'7d': f.utilization_7d ?? null,
|
||||
},
|
||||
reset: probeStatus === 'unreachable' ? null : {
|
||||
'5h': f.reset_5h ?? null,
|
||||
'7d': f.reset_7d ?? null,
|
||||
overall: f.reset ?? null,
|
||||
overage: f.overage_reset ?? null,
|
||||
},
|
||||
representative_claim: f.representative_claim ?? null,
|
||||
fallback_percentage: f.fallback_percentage ?? null,
|
||||
overage: probeStatus === 'unreachable' ? null : {
|
||||
status: f.overage_status ?? null,
|
||||
disabled_reason: f.overage_disabled_reason ?? null,
|
||||
},
|
||||
raw_available: (typeof raw.raw === 'object' && raw.raw !== null),
|
||||
// v0.5.1 (F3 — ADR 0013 Rule 6): failure detail for operator diagnostics
|
||||
failure: raw.failure ?? null,
|
||||
failure_kind: raw.failure?.kind ?? null,
|
||||
backoff_until: raw.failure?.backoff_until ?? null,
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Aggregate per-provider quota status into a normalized dashboard-friendly
|
||||
* shape. This is the D81 Phase 5 extension of lib/audit-query.mjs per
|
||||
* ADR 0008 Amendment (D81).
|
||||
*
|
||||
* For each loaded provider, calls quotaStatus() (already cached at the plugin
|
||||
* layer per ADR 0013 Rule 3) and normalizes to a consistent shape. Providers
|
||||
* returning null (codex, mistral) produce a { status: 'unavailable' } row.
|
||||
*
|
||||
* Audit-query stays in-memory scan per ADR 0008 Lane 2 = A. This function
|
||||
* does NOT scan the ndjson files; it calls the live provider plugins.
|
||||
*
|
||||
* Authority: ADR 0008 Amendment (D81) + ADR 0012 D81 + ADR 0013 Rule 5.
|
||||
*
|
||||
* @param {object} args
|
||||
* @param {Map<string, object>} args.providers - Map of provider name → plugin object
|
||||
* @param {(name: string) => Promise<object|null>} [args.getQuotaStatus] - injectable for tests;
|
||||
* defaults to calling providers.get(name).quotaStatus?.()
|
||||
* @returns {Promise<Array<{
|
||||
* provider: string,
|
||||
* status: 'live' | 'stale' | 'unavailable' | 'disabled',
|
||||
* reason?: string,
|
||||
* schema_version: string | null,
|
||||
* last_fresh_at: number | null,
|
||||
* utilization: { '5h': number|null, '7d': number|null } | null,
|
||||
* reset: { '5h': number|null, '7d': number|null, overall: number|null, overage: number|null } | null,
|
||||
* representative_claim: string | null,
|
||||
* fallback_percentage: number | null,
|
||||
* overage: { status: string|null, disabled_reason: string|null } | null,
|
||||
* raw_available: boolean,
|
||||
* }>>}
|
||||
*/
|
||||
export async function aggregateProviderQuota({
|
||||
providers,
|
||||
getQuotaStatus,
|
||||
} = {}) {
|
||||
if (!providers) {
|
||||
throw new Error('aggregateProviderQuota: providers (Map) is required');
|
||||
}
|
||||
|
||||
// Normalize the providers argument — accept both Map and plain object.
|
||||
const providerEntries = (providers instanceof Map)
|
||||
? [...providers.entries()]
|
||||
: Object.entries(providers);
|
||||
|
||||
const results = [];
|
||||
for (const [name, plugin] of providerEntries) {
|
||||
// Default getter: call the plugin's quotaStatus() if present.
|
||||
const fetchQuota = getQuotaStatus
|
||||
? () => getQuotaStatus(name)
|
||||
: () => (typeof plugin?.quotaStatus === 'function' ? plugin.quotaStatus(null) : Promise.resolve(null));
|
||||
|
||||
let rawResult = null;
|
||||
let callError = null;
|
||||
try {
|
||||
rawResult = await fetchQuota();
|
||||
} catch (err) {
|
||||
callError = err?.message ?? String(err);
|
||||
}
|
||||
|
||||
if (callError !== null) {
|
||||
// quotaStatus() threw — treat as error / unavailable.
|
||||
results.push({
|
||||
provider: name,
|
||||
status: 'unavailable',
|
||||
reason: callError,
|
||||
schema_version: null,
|
||||
last_fresh_at: null,
|
||||
utilization: null,
|
||||
reset: null,
|
||||
representative_claim: null,
|
||||
fallback_percentage: null,
|
||||
overage: null,
|
||||
raw_available: false,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
if (rawResult === null || rawResult === undefined) {
|
||||
// Plugin returned null: opt-in disabled (the ONLY case per v0.5.1 contract)
|
||||
// or providers with no quota API at all (codex, mistral).
|
||||
results.push({
|
||||
provider: name,
|
||||
status: 'unavailable',
|
||||
reason: 'no public quota api or probe disabled',
|
||||
schema_version: null,
|
||||
last_fresh_at: null,
|
||||
utilization: null,
|
||||
reset: null,
|
||||
representative_claim: null,
|
||||
fallback_percentage: null,
|
||||
overage: null,
|
||||
raw_available: false,
|
||||
failure: null,
|
||||
failure_kind: null,
|
||||
backoff_until: null,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
// quotaStatus() returned a non-null shape — normalize.
|
||||
// v0.5.1: handle probe_status field (live/stale/unreachable).
|
||||
// Currently only 'anthropic' returns a structured shape; other providers
|
||||
// returning structured data will work if their shape is compatible.
|
||||
const probeStatus = rawResult.probe_status ?? (rawResult.stale === true ? 'stale' : 'live');
|
||||
const normalized = _normalizeAnthropicQuota(rawResult);
|
||||
|
||||
if (normalized === null) {
|
||||
// Shape was present but unrecognizable.
|
||||
results.push({
|
||||
provider: name,
|
||||
status: 'unavailable',
|
||||
reason: 'unrecognized quota shape',
|
||||
schema_version: null,
|
||||
last_fresh_at: null,
|
||||
utilization: null,
|
||||
reset: null,
|
||||
representative_claim: null,
|
||||
fallback_percentage: null,
|
||||
overage: null,
|
||||
raw_available: false,
|
||||
failure: null,
|
||||
failure_kind: null,
|
||||
backoff_until: null,
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
// Map probe_status to output status:
|
||||
// 'live' → 'live'
|
||||
// 'stale' → 'stale'
|
||||
// 'unreachable' → 'unreachable' (new in v0.5.1; dashboard renders with red border)
|
||||
const outputStatus = probeStatus === 'unreachable' ? 'unreachable'
|
||||
: probeStatus === 'stale' ? 'stale'
|
||||
: 'live';
|
||||
|
||||
results.push({
|
||||
provider: name,
|
||||
status: outputStatus,
|
||||
...normalized,
|
||||
});
|
||||
}
|
||||
|
||||
return results;
|
||||
}
|
||||
|
||||
/**
|
||||
* Audit-derived cache hit rate over the window. Differs from
|
||||
* `cacheStore.stats()` in server.mjs: that is the live in-process counter;
|
||||
* this is the audit-side rate scoped to the rolling window.
|
||||
*
|
||||
* { window: { startMs, endMs }, total, hit, miss, bypass, hit_rate, by_provider }
|
||||
* { window: { startMs, endMs }, total, hit, miss, bypass, streaming_attached, hit_rate, by_provider }
|
||||
*
|
||||
* `streaming_attached` (D58, ADR 0005 Amendment 8 §11): D58 streaming
|
||||
* singleflight joiners did not hit a literal cache, so they are excluded
|
||||
* from both numerator AND denominator of `hit_rate`. Tracked separately
|
||||
* so the count reconciles with `total = hit + miss + bypass + streaming_attached`.
|
||||
*
|
||||
* @param {object} args
|
||||
* @param {number} args.windowMs
|
||||
@@ -459,18 +660,22 @@ export function cacheHitRateWindow({ windowMs, olpHome, logEvent, _nowFn } = {})
|
||||
const startMs = now - windowMs;
|
||||
const endMs = now;
|
||||
|
||||
let total = 0, hit = 0, miss = 0, bypass = 0;
|
||||
let total = 0, hit = 0, miss = 0, bypass = 0, streaming_attached = 0;
|
||||
const by_provider = {};
|
||||
|
||||
for (const ev of readAuditWindow({ startMs, endMs, olpHome, logEvent })) {
|
||||
if (ev.cache_status === null || ev.cache_status === undefined) continue;
|
||||
total++;
|
||||
const p = typeof ev.provider === 'string' && ev.provider.length > 0 ? ev.provider : '__unknown__';
|
||||
const pe = by_provider[p] ??= { total: 0, hit: 0, miss: 0, bypass: 0, hit_rate: 0 };
|
||||
const pe = by_provider[p] ??= { total: 0, hit: 0, miss: 0, bypass: 0, streaming_attached: 0, hit_rate: 0 };
|
||||
pe.total++;
|
||||
if (ev.cache_status === 'hit') { hit++; pe.hit++; }
|
||||
else if (ev.cache_status === 'miss') { miss++; pe.miss++; }
|
||||
else if (ev.cache_status === 'bypass') { bypass++; pe.bypass++; }
|
||||
// D58 — ADR 0005 Amendment 8 §11: streaming singleflight joiners.
|
||||
// Excluded from hit_rate numerator + denominator (they did not hit a
|
||||
// literal cache); tracked so `total` reconciles.
|
||||
else if (ev.cache_status === 'streaming_attached') { streaming_attached++; pe.streaming_attached++; }
|
||||
}
|
||||
|
||||
// Compute hit_rate per provider + overall (excludes bypass from denominator
|
||||
@@ -484,6 +689,6 @@ export function cacheHitRateWindow({ windowMs, olpHome, logEvent, _nowFn } = {})
|
||||
|
||||
return {
|
||||
window: { startMs, endMs },
|
||||
total, hit, miss, bypass, hit_rate, by_provider,
|
||||
total, hit, miss, bypass, streaming_attached, hit_rate, by_provider,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -174,6 +174,20 @@ export function _maybeRotateAudit(args = {}) {
|
||||
* fallback_hops, tried_providers, error_code, ir_request_hash, chain_id).
|
||||
* Caller is responsible for populating fields; missing fields are
|
||||
* serialized as undefined → omitted by JSON.stringify.
|
||||
*
|
||||
* cache_status enum (free-form string; not schema-validated at append):
|
||||
* 'hit' — served from cache (buffered-replay or streaming
|
||||
* cache_hit role from ADR 0005 Amendment 8 §1).
|
||||
* 'miss' — buffered or streaming source path; provider
|
||||
* spawn fired for this request.
|
||||
* 'bypass' — D2 cache_control bypass (no cache read/write).
|
||||
* 'streaming_attached' — D58 / ADR 0005 Amendment 8 §11: client joined
|
||||
* an in-flight streaming source spawn from
|
||||
* another caller; this request did NOT spawn a
|
||||
* provider but also did NOT hit the cache (the
|
||||
* cache was empty at the inflight Map check).
|
||||
* null — pre-chain error paths (the cache layer was
|
||||
* never consulted; e.g. 401, 415, 400 IR).
|
||||
* @param {object} [opts]
|
||||
* @param {string} [opts.olpHome] - test override; defaults to ~/.olp
|
||||
* @param {(level: string, event: string, data?: object) => void} [opts.logEvent]
|
||||
|
||||
Vendored
+582
-9
@@ -51,6 +51,75 @@
|
||||
* @property {number} inflightCount
|
||||
*/
|
||||
|
||||
// ── D57 — ADR 0005 Amendment 8 streaming-singleflight shapes ──────────────
|
||||
|
||||
/**
|
||||
* @typedef {Object} StreamingInflightEntry
|
||||
* @property {string} compositeKey - `${keyId}\0${cacheKey}`
|
||||
* @property {AsyncIterator<*>|null} source - underlying source iterator (null while factory pending)
|
||||
* @property {AbortController} sourceAbortController - propagates "all clients gone" to the source
|
||||
* @property {Array<*>} accumulatedChunks - late-joiner replay buffer (bounded by §10)
|
||||
* @property {number} accumulatedByteSize - running byte size of accumulatedChunks
|
||||
* @property {boolean} accumulatedReplayCapExceeded - true once §10 cap hit; cache write will skip
|
||||
* @property {Set<AttachedClient>} attachedClients - all live clients tee'ing this source
|
||||
* @property {boolean} factoryPending - true while sourceFactory() is awaited
|
||||
* @property {Array<{ resolve: function, reject: function }>} pendingJoiners - late joiners arriving during factoryPending
|
||||
* @property {boolean} sourceDone - source iterator exhausted normally
|
||||
* @property {Error|null} sourceError - non-null if source threw
|
||||
* @property {boolean} sourceAborted - true if AbortController fired
|
||||
* @property {number} ttlMs - TTL to use when writing the completed accumulated chunks to cache
|
||||
*/
|
||||
|
||||
/**
|
||||
* @typedef {Object} AttachedClient
|
||||
* @property {string} id - request id / correlator
|
||||
* @property {Array<*>} queue - per-client tee buffer
|
||||
* @property {number} queueByteSize - running byte size sum
|
||||
* @property {boolean} yieldedAccumulated - true after late-joiner replay drained
|
||||
* @property {boolean} done - terminal sentinel hit
|
||||
* @property {boolean} backpressured - true once STREAM_BACKPRESSURE terminator scheduled
|
||||
* @property {Error|null} error - non-null if source threw or replay-drain over-cap
|
||||
* @property {((chunk: { value: *, done: boolean }) => void)|null} resolveNext - pending pull promise resolver
|
||||
* @property {((err: Error) => void)|null} rejectNext - pending pull promise rejecter
|
||||
*/
|
||||
|
||||
// ADR 0005 Amendment 8 §14 — implementation defaults. Per-call overrides flow
|
||||
// via the `opts` argument to getOrComputeStreaming (used by tests to exercise
|
||||
// caps cheaply without producing megabytes of fixture data).
|
||||
export const PER_CLIENT_QUEUE_CAP_DEFAULT = 1 * 1024 * 1024;
|
||||
export const ACCUMULATED_REPLAY_CAP_DEFAULT = 10 * 1024 * 1024;
|
||||
|
||||
/**
|
||||
* Returns an approximate byte size for a chunk. The tee + cap math uses
|
||||
* JSON.stringify length as the serialization yardstick (matches the cache
|
||||
* store's existing size accounting via Buffer.byteLength(JSON.stringify(...))).
|
||||
* Non-stringifiable values (circular, etc.) fall back to 0 — the tee continues
|
||||
* but the size accounting under-estimates that chunk. In practice IR chunks
|
||||
* are always JSON-safe.
|
||||
*
|
||||
* @param {*} chunk
|
||||
* @returns {number}
|
||||
*/
|
||||
function _chunkByteSize(chunk) {
|
||||
try {
|
||||
return Buffer.byteLength(JSON.stringify(chunk) ?? '', 'utf8');
|
||||
} catch {
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Synthesises a STREAM_BACKPRESSURE terminator stream (ADR 0005 Amendment 8 §8):
|
||||
* yields `{ type: 'stop', finish_reason: 'length' }` then a `[DONE]` sentinel
|
||||
* then ends. Used both for late-joiner-too-late and per-client overflow paths.
|
||||
*
|
||||
* @returns {AsyncGenerator<*>}
|
||||
*/
|
||||
async function* _backpressureTerminator() {
|
||||
yield { type: 'stop', finish_reason: 'length' };
|
||||
yield '[DONE]';
|
||||
}
|
||||
|
||||
// ── CacheStore ────────────────────────────────────────────────────────────
|
||||
|
||||
export class CacheStore {
|
||||
@@ -82,6 +151,14 @@ export class CacheStore {
|
||||
/** @type {Map<string, Promise<*>>} */
|
||||
this._inflight = new Map();
|
||||
|
||||
// D57 — ADR 0005 Amendment 8: streaming singleflight per-(keyId,cacheKey)
|
||||
// inflight Map. Composite key uses `\0` separator per §2 (avoids colliding
|
||||
// with keyId or cacheKey content). The check + insert against this Map is
|
||||
// synchronous in getOrComputeStreaming (no `await` between read & write),
|
||||
// mirroring D38 `tryAcquireSpawn` atomicity.
|
||||
/** @type {Map<string, StreamingInflightEntry>} */
|
||||
this._streamingInflight = new Map();
|
||||
|
||||
// Stats per keyId
|
||||
/** @type {Map<string, { hits: number, misses: number }>} */
|
||||
this._stats = new Map();
|
||||
@@ -266,13 +343,10 @@ export class CacheStore {
|
||||
* @param {number} [ttlMs]
|
||||
* @returns {Promise<*>}
|
||||
*
|
||||
* TODO(v1.x — ADR 0005 Amendment 8 / issue #16): add a sibling
|
||||
* `getOrComputeStreaming(keyId, cacheKey, sourceFactory)` for the streaming
|
||||
* path. This API handles buffered responses only; the streaming branch in
|
||||
* server.mjs currently uses a peek+spawn pattern with a TOCTOU window.
|
||||
* The streaming sibling will mirror this method's shape but with a tee
|
||||
* fan-out and per-client backpressure queues. See docs/v1x-roadmap.md #1
|
||||
* for the design contract and acceptance criteria.
|
||||
* D57 — ADR 0005 Amendment 8 / issue #16 resolved at the cache layer:
|
||||
* the sibling `getOrComputeStreaming` method below provides tee-fan-out +
|
||||
* per-client backpressure for the streaming path. Server-side wiring lands
|
||||
* separately in D58.
|
||||
*/
|
||||
async getOrCompute(keyId, cacheKey, computeFn, ttlMs) {
|
||||
// 1. Cache hit — return immediately, no singleflight overhead
|
||||
@@ -311,6 +385,493 @@ export class CacheStore {
|
||||
return computePromise;
|
||||
}
|
||||
|
||||
/**
|
||||
* D57 — ADR 0005 Amendment 8 (issue #16): streaming singleflight + tee-fan-out.
|
||||
*
|
||||
* Streaming-path sibling of `getOrCompute`. Coordinates concurrent identical
|
||||
* streaming requests so only one source spawn occurs, all attached clients
|
||||
* receive identical chunk sequences in order, late joiners are replayed from
|
||||
* an accumulated buffer, slow clients are disconnected with a synthetic
|
||||
* STREAM_BACKPRESSURE terminator instead of stalling the source, and the
|
||||
* source is aborted when all clients disconnect.
|
||||
*
|
||||
* The Map check + insert against `_streamingInflight` is synchronous — no
|
||||
* `await` between read and write — matching the D38 `tryAcquireSpawn`
|
||||
* atomicity invariant and collapsing the original TOCTOU window in
|
||||
* `server.mjs:782` (peek + spawn). See §1 + §6.
|
||||
*
|
||||
* Three outcomes (Amendment 8 §1):
|
||||
* - **cache_hit**: cached entry exists and is alive → returns
|
||||
* `{ stream: <async iterator over cached chunks>, isFirst: false,
|
||||
* role: 'cache_hit' }`. No spawn. `hits` incremented.
|
||||
* - **attached**: inflight entry exists in `_streamingInflight` → attaches
|
||||
* a new AttachedClient. Late-joiner replay (§5) drains accumulated
|
||||
* chunks synchronously; if drain would exceed `perClientQueueCap` the
|
||||
* client receives a STREAM_BACKPRESSURE synthesised stream instead.
|
||||
* `isFirst: false`, `role: 'attached'`. `hits` incremented (sharing a
|
||||
* spawn is functionally a "cache-like" benefit; documented choice).
|
||||
* - **source**: no cache hit, no inflight entry → the cache layer takes
|
||||
* the inflight lock synchronously (placeholder entry inserted), then
|
||||
* `await sourceFactory()`. If the factory throws (e.g.
|
||||
* CONCURRENCY_LIMIT) the placeholder is removed and the error
|
||||
* propagates. On success the source is wired up, the tee task starts,
|
||||
* `isFirst: true`, `role: 'source'`. `misses` incremented.
|
||||
*
|
||||
* **`role` enum** (returned at attach-time):
|
||||
* - `source` — first caller; entry created. Lifetime-end may upgrade to
|
||||
* `solo` (no joiners ever attached) at the server (D58); the cache
|
||||
* layer reports `source` at attach-time and does not flip mid-stream.
|
||||
* - `attached` — joined an existing inflight entry.
|
||||
* - `cache_hit` — served from cache; no entry created.
|
||||
*
|
||||
* **Amendment 8 §11 header values**: the X-OLP-Streaming-Inflight header
|
||||
* (D58) uses `source | attached | solo`. The cache layer's `cache_hit` is
|
||||
* the "served from cache without inflight" case; D58 chooses whether to
|
||||
* emit `solo` or omit the header for that path. The cache layer reports
|
||||
* `cache_hit` as a distinct role so the server can disambiguate.
|
||||
*
|
||||
* **§14 defaults**: `PER_CLIENT_QUEUE_CAP = 1 MB`, `ACCUMULATED_REPLAY_CAP
|
||||
* = 10 MB`. Both overridable via `opts` for cheap test exercise of the cap
|
||||
* paths.
|
||||
*
|
||||
* @param {string} keyId
|
||||
* @param {string} cacheKey
|
||||
* @param {() => Promise<AsyncIterator<*>>|AsyncIterator<*>} sourceFactory
|
||||
* Invoked exactly once per inflight lifetime (only on first caller).
|
||||
* Returns the source async iterator. May throw (e.g. CONCURRENCY_LIMIT
|
||||
* from `tryAcquireSpawn`); on throw, the inflight entry is removed and
|
||||
* the error propagates to the first caller. Late joiners that attached
|
||||
* while the factory was pending are rejected with the same error.
|
||||
* @param {object} [opts]
|
||||
* @param {string} [opts.clientId] - correlator id (defaults to incrementing counter)
|
||||
* @param {number} [opts.ttlMs] - TTL for completed accumulated-chunks cache write
|
||||
* @param {number} [opts.perClientQueueCap] - override §14 default (1 MB)
|
||||
* @param {number} [opts.accumulatedReplayCap] - override §14 default (10 MB)
|
||||
* @returns {Promise<{ stream: AsyncGenerator<*>, isFirst: boolean, role: 'source'|'attached'|'cache_hit' }>}
|
||||
*/
|
||||
async getOrComputeStreaming(keyId, cacheKey, sourceFactory, opts = {}) {
|
||||
const compositeKey = `${keyId}\0${cacheKey}`;
|
||||
const perClientQueueCap = opts.perClientQueueCap ?? PER_CLIENT_QUEUE_CAP_DEFAULT;
|
||||
const accumulatedReplayCap = opts.accumulatedReplayCap ?? ACCUMULATED_REPLAY_CAP_DEFAULT;
|
||||
const ttlMs = opts.ttlMs;
|
||||
const clientId = opts.clientId ?? `c-${this._nowFn()}-${Math.floor(Math.random() * 1e9).toString(36)}`;
|
||||
|
||||
// ── (A) Inflight Map check FIRST (Amendment 8 §6 — TTL race) ───────────
|
||||
// Late joiners that arrive after a cache entry has expired but during an
|
||||
// active source spawn must still attach via the inflight Map rather than
|
||||
// re-spawn. The synchronous Map.get + (if hit) Map preservation here
|
||||
// satisfies the no-`await`-between-check-and-decision invariant.
|
||||
const inflightEntry = this._streamingInflight.get(compositeKey);
|
||||
if (inflightEntry) {
|
||||
// Hits-as-share decision (documented at method header): sharing a spawn
|
||||
// is a cache-like benefit; increment hits for consistency with
|
||||
// cache_hit accounting and to expose the singleflight win in stats().
|
||||
this._getStats(keyId).hits++;
|
||||
const stream = this._attachClient(inflightEntry, {
|
||||
clientId,
|
||||
perClientQueueCap,
|
||||
});
|
||||
return { stream, isFirst: false, role: 'attached' };
|
||||
}
|
||||
|
||||
// ── (B) Cache hit check (no inflight) ──────────────────────────────────
|
||||
// Replays cached chunks via a synthetic async iterator. No source spawn.
|
||||
const ns = this._getNamespace(keyId);
|
||||
const existing = ns.get(cacheKey);
|
||||
if (existing && this._isAlive(existing)) {
|
||||
this._getStats(keyId).hits++;
|
||||
const cachedChunks = Array.isArray(existing.value) ? existing.value : [existing.value];
|
||||
const stream = (async function* cacheReplay() {
|
||||
for (const chunk of cachedChunks) {
|
||||
yield chunk;
|
||||
}
|
||||
})();
|
||||
return { stream, isFirst: false, role: 'cache_hit' };
|
||||
}
|
||||
|
||||
// ── (C) Miss + no inflight: take the lock synchronously, then await ───
|
||||
// The placeholder entry is inserted BEFORE invoking sourceFactory so that
|
||||
// late joiners arriving while the factory is awaited see the inflight
|
||||
// entry and attach (they're parked in `pendingJoiners` until the factory
|
||||
// resolves or rejects). If the factory throws, the placeholder is
|
||||
// removed and the error propagates to the first caller AND all parked
|
||||
// joiners. This preserves the §1 invariant: the Map insert is atomic
|
||||
// from later joiners' perspective.
|
||||
const entry = /** @type {StreamingInflightEntry} */ ({
|
||||
compositeKey,
|
||||
source: null,
|
||||
sourceAbortController: new AbortController(),
|
||||
accumulatedChunks: [],
|
||||
accumulatedByteSize: 0,
|
||||
accumulatedReplayCapExceeded: false,
|
||||
attachedClients: new Set(),
|
||||
factoryPending: true,
|
||||
pendingJoiners: [],
|
||||
sourceDone: false,
|
||||
sourceError: null,
|
||||
sourceAborted: false,
|
||||
ttlMs,
|
||||
});
|
||||
this._streamingInflight.set(compositeKey, entry);
|
||||
this._getStats(keyId).misses++;
|
||||
|
||||
// Attach the first caller synchronously so any subsequent joiners during
|
||||
// the factory await see the same set/topology as the first caller.
|
||||
const firstStream = this._attachClient(entry, {
|
||||
clientId,
|
||||
perClientQueueCap,
|
||||
});
|
||||
|
||||
let sourceIter;
|
||||
try {
|
||||
const factoryResult = sourceFactory();
|
||||
sourceIter = factoryResult && typeof factoryResult.then === 'function'
|
||||
? await factoryResult
|
||||
: factoryResult;
|
||||
} catch (err) {
|
||||
// Factory rejected — remove placeholder and reject first caller + any
|
||||
// late joiners that arrived during the await.
|
||||
this._streamingInflight.delete(compositeKey);
|
||||
for (const client of entry.attachedClients) {
|
||||
if (client.rejectNext) {
|
||||
client.rejectNext(err);
|
||||
client.resolveNext = null;
|
||||
client.rejectNext = null;
|
||||
}
|
||||
client.error = err;
|
||||
client.done = true;
|
||||
}
|
||||
throw err;
|
||||
}
|
||||
|
||||
entry.source = sourceIter;
|
||||
entry.factoryPending = false;
|
||||
|
||||
// Kick off the tee task. It runs detached; lifetime is bounded by the
|
||||
// source iterator's completion / error / abort.
|
||||
this._teeStreamingSource(keyId, cacheKey, entry, {
|
||||
accumulatedReplayCap,
|
||||
});
|
||||
|
||||
return { stream: firstStream, isFirst: true, role: 'source' };
|
||||
}
|
||||
|
||||
/**
|
||||
* D57 — ADR 0005 Amendment 8 §3 + §5: attach a new client to an inflight
|
||||
* entry. Synchronously drains the accumulated replay buffer into the
|
||||
* client's queue (§5). If the drain would exceed `perClientQueueCap`, the
|
||||
* client receives a STREAM_BACKPRESSURE synthesised stream INSTEAD of the
|
||||
* normal tee — the source continues for the other clients.
|
||||
*
|
||||
* Returns the async iterator the caller will consume.
|
||||
*
|
||||
* @private
|
||||
*/
|
||||
_attachClient(entry, { clientId, perClientQueueCap }) {
|
||||
// Late-joiner replay drain cap check (§5 + §10): a late joiner cannot
|
||||
// catch up if either
|
||||
// (a) the accumulated buffer alone would overflow the per-client cap
|
||||
// (burst > PER_CLIENT_QUEUE_CAP), or
|
||||
// (b) the replay buffer is already truncated (§10 cap was hit and
|
||||
// further source chunks were not appended to accumulatedChunks),
|
||||
// so even a successful drain would give the joiner a partial view
|
||||
// that disagrees with later live chunks.
|
||||
// Either condition → STREAM_BACKPRESSURE synthetic terminator. The
|
||||
// source / other clients are unaffected.
|
||||
if (
|
||||
entry.accumulatedByteSize > perClientQueueCap
|
||||
|| entry.accumulatedReplayCapExceeded
|
||||
) {
|
||||
this._warnFn('stream_backpressure_disconnect', {
|
||||
client_id: clientId,
|
||||
queue_byte_size: entry.accumulatedByteSize,
|
||||
per_client_cap: perClientQueueCap,
|
||||
composite_key: entry.compositeKey,
|
||||
reason: entry.accumulatedReplayCapExceeded
|
||||
? 'replay_cap_truncated'
|
||||
: 'replay_drain_over_cap',
|
||||
});
|
||||
return _backpressureTerminator();
|
||||
}
|
||||
|
||||
/** @type {AttachedClient} */
|
||||
const client = {
|
||||
id: clientId,
|
||||
queue: [],
|
||||
queueByteSize: 0,
|
||||
yieldedAccumulated: false,
|
||||
done: false,
|
||||
backpressured: false,
|
||||
error: null,
|
||||
resolveNext: null,
|
||||
rejectNext: null,
|
||||
// Per-client cap is captured here so the tee task can apply it
|
||||
// without re-plumbing opts; documented as an internal field.
|
||||
__perClientQueueCap__: perClientQueueCap,
|
||||
};
|
||||
|
||||
// Synchronous replay drain — push every accumulated chunk into the
|
||||
// client's queue at attach-time. From this point on the tee task pushes
|
||||
// live chunks.
|
||||
for (const chunk of entry.accumulatedChunks) {
|
||||
client.queue.push(chunk);
|
||||
client.queueByteSize += _chunkByteSize(chunk);
|
||||
}
|
||||
client.yieldedAccumulated = true;
|
||||
entry.attachedClients.add(client);
|
||||
|
||||
// TODO(D58 — ADR 0005 Amendment 8 §11): emit `streaming_inflight_join`
|
||||
// event from the server-layer wrapper, which has provider/model context.
|
||||
// Cache layer alone does not have provider/model identity (sourceFactory
|
||||
// is a closure), so the join event lives at the consumer of `role:
|
||||
// 'attached'` in server.mjs. D57 reviewer P2-3 follow-up.
|
||||
|
||||
// If source already completed before this attach (last-second join)
|
||||
// mark the client as terminal-after-drain so the iterator returns
|
||||
// cleanly once the replay queue is drained.
|
||||
if (entry.sourceDone) {
|
||||
client.done = true;
|
||||
} else if (entry.sourceError) {
|
||||
client.error = entry.sourceError;
|
||||
client.done = true;
|
||||
}
|
||||
|
||||
// Per-client AbortController for client-side cancellation (HTTP close).
|
||||
// The async iterator's return() removes the client from attachedClients;
|
||||
// if the entry's attachedClients size hits zero, the tee task aborts the
|
||||
// source. The teardown logic lives in the iterator below.
|
||||
const store = this;
|
||||
const iterator = (async function* clientStream() {
|
||||
try {
|
||||
while (true) {
|
||||
// Drain queue chunks first.
|
||||
if (client.queue.length > 0) {
|
||||
const next = client.queue.shift();
|
||||
client.queueByteSize -= _chunkByteSize(next);
|
||||
yield next;
|
||||
continue;
|
||||
}
|
||||
// Backpressure-terminated client: yield the synthetic terminator.
|
||||
if (client.backpressured) {
|
||||
yield { type: 'stop', finish_reason: 'length' };
|
||||
yield '[DONE]';
|
||||
client.done = true;
|
||||
return;
|
||||
}
|
||||
// Source already errored.
|
||||
if (client.error) {
|
||||
throw client.error;
|
||||
}
|
||||
// Source already completed and no more queued chunks.
|
||||
if (client.done) {
|
||||
return;
|
||||
}
|
||||
// Block until the tee task pushes the next chunk (or signals
|
||||
// source-done / source-error / backpressure).
|
||||
await new Promise((resolve, reject) => {
|
||||
client.resolveNext = resolve;
|
||||
client.rejectNext = reject;
|
||||
});
|
||||
client.resolveNext = null;
|
||||
client.rejectNext = null;
|
||||
}
|
||||
} finally {
|
||||
// Iterator return() fired (HTTP close, break, or normal return) —
|
||||
// remove client and possibly trigger source abort.
|
||||
if (entry.attachedClients.has(client)) {
|
||||
entry.attachedClients.delete(client);
|
||||
}
|
||||
// If we're the last client AND the source is still running, fire
|
||||
// the AbortController so the source iterator's return() / cleanup
|
||||
// can reap any underlying resources.
|
||||
if (
|
||||
entry.attachedClients.size === 0
|
||||
&& !entry.sourceDone
|
||||
&& !entry.sourceError
|
||||
&& !entry.sourceAborted
|
||||
&& !entry.factoryPending
|
||||
) {
|
||||
entry.sourceAborted = true;
|
||||
try {
|
||||
entry.sourceAbortController.abort();
|
||||
} catch {
|
||||
// best-effort
|
||||
}
|
||||
// Tee task observes attachedClients.size === 0 + sourceAborted on
|
||||
// its next loop iteration and exits without a cache write.
|
||||
store._streamingInflight.delete(entry.compositeKey);
|
||||
store._warnFn('streaming_inflight_abort', {
|
||||
composite_key: entry.compositeKey,
|
||||
accumulated_chunk_count: entry.accumulatedChunks.length,
|
||||
});
|
||||
}
|
||||
}
|
||||
})();
|
||||
|
||||
return iterator;
|
||||
}
|
||||
|
||||
/**
|
||||
* D57 — ADR 0005 Amendment 8 §4: tee fan-out task. One reader pulls from
|
||||
* `entry.source`; on each chunk, pushes to `accumulatedChunks` (bounded by
|
||||
* §10) and to every attached client's queue (per-client cap from §8).
|
||||
*
|
||||
* On source completion: writes accumulated chunks to cache if (a) cap not
|
||||
* exceeded and (b) `set()`'s own `maxEntryBytes` cap admits it. Resolves
|
||||
* all clients to drain-out state. Removes entry.
|
||||
*
|
||||
* On source error: rejects all clients via their `rejectNext`. No cache
|
||||
* write. Removes entry.
|
||||
*
|
||||
* Source-abort short-circuit: if `attachedClients.size === 0` after a push,
|
||||
* fires `sourceAbortController.abort()`, exits without cache write.
|
||||
*
|
||||
* @private
|
||||
*/
|
||||
_teeStreamingSource(keyId, cacheKey, entry, { accumulatedReplayCap }) {
|
||||
const store = this;
|
||||
(async () => {
|
||||
try {
|
||||
for (;;) {
|
||||
// Pre-check: if all clients have already gone away before we even
|
||||
// pull the next chunk, abort the source and bail out. (The
|
||||
// per-client iterator's finally-block sets sourceAborted=true and
|
||||
// removes the entry; we just need to stop pulling.)
|
||||
if (entry.attachedClients.size === 0 && !entry.factoryPending) {
|
||||
if (!entry.sourceAborted) {
|
||||
entry.sourceAborted = true;
|
||||
try { entry.sourceAbortController.abort(); } catch { /* best-effort */ }
|
||||
}
|
||||
// Try to call return() on the source so the underlying generator
|
||||
// cleans up. Best-effort; not all iterators implement it.
|
||||
try {
|
||||
if (entry.source && typeof entry.source.return === 'function') {
|
||||
await entry.source.return();
|
||||
}
|
||||
} catch { /* best-effort */ }
|
||||
return;
|
||||
}
|
||||
|
||||
const result = await entry.source.next();
|
||||
if (result.done) break;
|
||||
const chunk = result.value;
|
||||
const size = _chunkByteSize(chunk);
|
||||
|
||||
// §10 — replay buffer cap. Past the cap we stop accumulating (so
|
||||
// future late joiners can still see the chunks they need to
|
||||
// catch up to live), but we mark the entry not-cacheable so the
|
||||
// §4 completion path skips the cache write.
|
||||
if (!entry.accumulatedReplayCapExceeded) {
|
||||
if (entry.accumulatedByteSize + size > accumulatedReplayCap) {
|
||||
entry.accumulatedReplayCapExceeded = true;
|
||||
store._warnFn('streaming_inflight_replay_cap_exceeded', {
|
||||
composite_key: entry.compositeKey,
|
||||
accumulated_byte_size: entry.accumulatedByteSize,
|
||||
chunk_size: size,
|
||||
accumulated_replay_cap: accumulatedReplayCap,
|
||||
});
|
||||
// Continue accumulating up to this chunk so existing late
|
||||
// joiners' drain decision was based on the size they saw at
|
||||
// attach-time. We do NOT push this chunk to accumulatedChunks
|
||||
// (it would corrupt the "<= cap at attach-time" invariant for
|
||||
// future joiners). Future joiners arriving past this point
|
||||
// see accumulatedByteSize already > cap and get the
|
||||
// backpressure terminator at attach-time per §5.
|
||||
} else {
|
||||
entry.accumulatedChunks.push(chunk);
|
||||
entry.accumulatedByteSize += size;
|
||||
}
|
||||
}
|
||||
|
||||
// Fan out to each client synchronously (no await inside the for-
|
||||
// each-client loop). Disconnecting a client is mutation-during-
|
||||
// iteration; we snapshot the set first.
|
||||
const clientsSnapshot = [...entry.attachedClients];
|
||||
for (const client of clientsSnapshot) {
|
||||
// Per-client backpressure (§8). If the push would exceed the
|
||||
// per-client cap, disconnect this client only.
|
||||
if (client.queueByteSize + size > client.__perClientQueueCap__) {
|
||||
client.backpressured = true;
|
||||
store._warnFn('stream_backpressure_disconnect', {
|
||||
client_id: client.id,
|
||||
queue_byte_size: client.queueByteSize,
|
||||
per_client_cap: client.__perClientQueueCap__,
|
||||
composite_key: entry.compositeKey,
|
||||
reason: 'queue_overflow',
|
||||
});
|
||||
// Remove from set so future fan-out skips this client.
|
||||
entry.attachedClients.delete(client);
|
||||
// Wake the client's pull-promise so it can yield the synthetic
|
||||
// STREAM_BACKPRESSURE terminator.
|
||||
if (client.resolveNext) {
|
||||
const r = client.resolveNext;
|
||||
client.resolveNext = null;
|
||||
client.rejectNext = null;
|
||||
r({ value: undefined, done: false });
|
||||
}
|
||||
continue;
|
||||
}
|
||||
client.queue.push(chunk);
|
||||
client.queueByteSize += size;
|
||||
if (client.resolveNext) {
|
||||
const r = client.resolveNext;
|
||||
client.resolveNext = null;
|
||||
client.rejectNext = null;
|
||||
r({ value: undefined, done: false });
|
||||
}
|
||||
}
|
||||
|
||||
// If the fan-out emptied the attached set (all over-cap), the
|
||||
// top-of-loop pre-check will fire on the next iteration and abort.
|
||||
}
|
||||
|
||||
// Source iterator returned normally.
|
||||
entry.sourceDone = true;
|
||||
const cacheWritten =
|
||||
!entry.accumulatedReplayCapExceeded
|
||||
&& entry.accumulatedChunks.length > 0;
|
||||
if (cacheWritten) {
|
||||
// ADR 0005 Amendment 8 §4: write accumulated chunks to cache via
|
||||
// the standard set() path (which itself applies the D23
|
||||
// maxEntryBytes cap — separate from the §10 replay cap).
|
||||
await store.set(keyId, cacheKey, entry.accumulatedChunks, entry.ttlMs);
|
||||
}
|
||||
store._warnFn('streaming_inflight_source_done', {
|
||||
composite_key: entry.compositeKey,
|
||||
attached_count: entry.attachedClients.size,
|
||||
accumulated_chunk_count: entry.accumulatedChunks.length,
|
||||
cache_written: cacheWritten,
|
||||
});
|
||||
// Wake every remaining attached client so they drain their queue and
|
||||
// observe `done = true`.
|
||||
for (const client of [...entry.attachedClients]) {
|
||||
client.done = true;
|
||||
if (client.resolveNext) {
|
||||
const r = client.resolveNext;
|
||||
client.resolveNext = null;
|
||||
client.rejectNext = null;
|
||||
r({ value: undefined, done: true });
|
||||
}
|
||||
}
|
||||
store._streamingInflight.delete(entry.compositeKey);
|
||||
} catch (err) {
|
||||
// Source threw mid-stream. Reject every attached client.
|
||||
entry.sourceError = err;
|
||||
for (const client of [...entry.attachedClients]) {
|
||||
client.error = err;
|
||||
client.done = true;
|
||||
if (client.rejectNext) {
|
||||
const rej = client.rejectNext;
|
||||
client.resolveNext = null;
|
||||
client.rejectNext = null;
|
||||
rej(err);
|
||||
}
|
||||
}
|
||||
store._streamingInflight.delete(entry.compositeKey);
|
||||
}
|
||||
})();
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns stats for a specific keyId, or aggregate stats across all keyIds.
|
||||
*
|
||||
@@ -322,11 +883,14 @@ export class CacheStore {
|
||||
const s = this._stats.get(keyId) ?? { hits: 0, misses: 0 };
|
||||
const ns = this._store.get(keyId);
|
||||
const size = ns ? ns.size : 0;
|
||||
// D57 — ADR 0005 Amendment 8 §1: inflightCount aggregates both the
|
||||
// buffered-path singleflight Map and the streaming-path inflight Map
|
||||
// so stats() reflects all active dedup-coordination entries.
|
||||
return {
|
||||
hits: s.hits,
|
||||
misses: s.misses,
|
||||
size,
|
||||
inflightCount: this._inflight.size,
|
||||
inflightCount: this._inflight.size + this._streamingInflight.size,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -345,7 +909,8 @@ export class CacheStore {
|
||||
hits: totalHits,
|
||||
misses: totalMisses,
|
||||
size: totalSize,
|
||||
inflightCount: this._inflight.size,
|
||||
// D57 — see per-keyId branch above; streaming entries counted alongside buffered.
|
||||
inflightCount: this._inflight.size + this._streamingInflight.size,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -397,10 +962,18 @@ export class CacheStore {
|
||||
this._inflight.delete(k);
|
||||
}
|
||||
}
|
||||
// D57 — also clear streaming inflight entries scoped to this keyId.
|
||||
// Composite key uses `\0` separator (see _streamingInflight init).
|
||||
for (const k of this._streamingInflight.keys()) {
|
||||
if (k.startsWith(`${keyId}\0`)) {
|
||||
this._streamingInflight.delete(k);
|
||||
}
|
||||
}
|
||||
} else {
|
||||
this._store.clear();
|
||||
this._stats.clear();
|
||||
this._inflight.clear();
|
||||
this._streamingInflight.clear();
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+607
@@ -0,0 +1,607 @@
|
||||
/**
|
||||
* lib/doctor.mjs — OLP doctor framework (Phase 4 / D65)
|
||||
*
|
||||
* Authority: ADR 0010 § Phase 4 D64-D67 + ADR 0002 Amendment 7 (D67) —
|
||||
* per-provider `doctorChecks()` contract method that this framework consumes.
|
||||
*
|
||||
* `olp doctor` runs a set of `Check` objects. Each check has:
|
||||
* - id: string (unique, e.g. 'server.running', 'anthropic.cli_available')
|
||||
* - category: 'server'|'auth'|'config'|'provider'|'system'
|
||||
* - async run(): { status: 'ok'|'fail'|'warn', message, evidence? }
|
||||
*
|
||||
* Built-in checks (categories server / auth / config / system) are defined in
|
||||
* `buildBuiltinChecks()` below; per-provider checks are sourced from each loaded
|
||||
* plugin's optional `doctorChecks()` method (ADR 0002 Amendment 7).
|
||||
*
|
||||
* Output shape (machine-readable — consumed by `bin/olp.mjs --json`):
|
||||
* {
|
||||
* schema_version: 1,
|
||||
* generated_at: '2026-05-26T...',
|
||||
* checks: [{ id, category, status, message, evidence? }],
|
||||
* fail_count: number,
|
||||
* warn_count: number,
|
||||
* kind: 'noop'|'fix_server'|'fix_oauth'|'fix_config'|'fix_provider'|'fresh_install',
|
||||
* next_action: { ai_executable: string[], human_required: string[], verify: string },
|
||||
* summary: string,
|
||||
* }
|
||||
*
|
||||
* `kind` precedence (highest first — the most upstream blocker wins):
|
||||
* 1. fresh_install — config.exists FAIL (~/.olp/config.json missing/malformed)
|
||||
* 2. fix_server — server.running FAIL
|
||||
* 3. fix_oauth — auth.owner_key_exists FAIL
|
||||
* 4. fix_provider — any provider-category FAIL
|
||||
* 5. fix_config — any other config-category FAIL
|
||||
* 6. noop — all OK (or WARN-only)
|
||||
*
|
||||
* `next_action.ai_executable[]` aggregates `evidence.fix_commands[]` from every
|
||||
* FAIL check; `next_action.human_required[]` aggregates `evidence.human_steps[]`.
|
||||
* `verify` is always `olp doctor` (re-run after applying the fix).
|
||||
*
|
||||
* Design notes:
|
||||
* - Pure functions + dependency injection: callers pass `{ checks }` (which can
|
||||
* be overridden for tests) plus a `{ now }` clock for deterministic timestamps.
|
||||
* - No filesystem writes. No process.exit. No console.log. Callers handle I/O.
|
||||
* - All checks run in parallel via Promise.all — individual check failures are
|
||||
* captured (not propagated) so one broken check does not hide others.
|
||||
*/
|
||||
|
||||
import { existsSync, readFileSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { request as httpRequest } from 'node:http';
|
||||
|
||||
import { loadProviders } from './providers/index.mjs';
|
||||
import { loadFallbackConfigSync } from './fallback/engine.mjs';
|
||||
import { listKeys } from './keys.mjs';
|
||||
|
||||
// ── Schema version ────────────────────────────────────────────────────────
|
||||
|
||||
export const DOCTOR_SCHEMA_VERSION = 1;
|
||||
|
||||
// ── Built-in check builders ───────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Resolve the OLP base URL the CLI / doctor will probe.
|
||||
* Precedence:
|
||||
* 1. opts.proxyUrl (explicit caller override)
|
||||
* 2. OLP_PROXY_URL env (full URL like http://host:port)
|
||||
* 3. http://127.0.0.1:${OLP_PORT || 4567}
|
||||
*/
|
||||
export function resolveProxyUrl(opts = {}) {
|
||||
if (opts.proxyUrl) return String(opts.proxyUrl).replace(/\/+$/, '');
|
||||
if (process.env.OLP_PROXY_URL) return String(process.env.OLP_PROXY_URL).replace(/\/+$/, '');
|
||||
const port = process.env.OLP_PORT ?? '4567';
|
||||
return `http://127.0.0.1:${port}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the OLP_HOME directory (mirrors lib/keys.mjs precedence).
|
||||
* 1. opts.olpHome
|
||||
* 2. OLP_HOME env
|
||||
* 3. ~/.olp
|
||||
*/
|
||||
export function resolveOlpHome(opts = {}) {
|
||||
if (opts.olpHome) return opts.olpHome;
|
||||
if (process.env.OLP_HOME) return process.env.OLP_HOME;
|
||||
return join(homedir(), '.olp');
|
||||
}
|
||||
|
||||
/**
|
||||
* Helper — issue a GET to the proxy with a tight timeout. Returns
|
||||
* `{ ok: true, status, body }` or `{ ok: false, error }`. Never throws.
|
||||
*/
|
||||
async function httpGet(url, { timeoutMs = 3000, headers = {} } = {}) {
|
||||
return new Promise(resolve => {
|
||||
let done = false;
|
||||
const finish = (v) => { if (!done) { done = true; resolve(v); } };
|
||||
let req;
|
||||
try {
|
||||
req = httpRequest(url, { method: 'GET', headers, timeout: timeoutMs }, res => {
|
||||
let data = '';
|
||||
res.on('data', c => { data += c; });
|
||||
res.on('end', () => finish({ ok: true, status: res.statusCode, body: data, headers: res.headers }));
|
||||
});
|
||||
} catch (e) {
|
||||
finish({ ok: false, error: String(e?.message ?? e) });
|
||||
return;
|
||||
}
|
||||
req.on('error', e => finish({ ok: false, error: String(e?.message ?? e) }));
|
||||
req.on('timeout', () => {
|
||||
try { req.destroy(new Error(`timeout after ${timeoutMs}ms`)); } catch { /* ignore */ }
|
||||
});
|
||||
req.end();
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the default check set. Test-mode overrides:
|
||||
* - opts.injectChecks: [...Check] — REPLACES the built-in set entirely
|
||||
* - opts.providersOverride: Map<name, plugin> — REPLACES loaded providers for the per-provider sweep
|
||||
* - opts.skipNetwork: true → omit server.running / server.version (offline mode)
|
||||
* - opts.olpHome: override ~/.olp lookup
|
||||
* - opts.proxyUrl: override resolveProxyUrl()
|
||||
*/
|
||||
/**
|
||||
* D64-D67 reviewer P2-1: shell-quote a path before interpolation into
|
||||
* `ai_executable[]` strings. Single-quote-wrap + escape any embedded single
|
||||
* quote per POSIX shell rules: `foo'bar` → `'foo'\''bar'`. Defends against
|
||||
* a malicious `OLP_HOME` env value injecting shell metacharacters into the
|
||||
* suggested-fix command an AI agent (or human) might paste back.
|
||||
*
|
||||
* Risk surface is narrow at family scale (operator local env, single-user
|
||||
* proxy), but the hardening cost is one helper.
|
||||
*
|
||||
* @param {string} s
|
||||
* @returns {string} single-quoted shell-safe string
|
||||
*/
|
||||
function _shellQuote(s) {
|
||||
return `'${String(s).replace(/'/g, "'\\''")}'`;
|
||||
}
|
||||
|
||||
export function buildBuiltinChecks(opts = {}) {
|
||||
if (opts.injectChecks) return opts.injectChecks;
|
||||
|
||||
const olpHome = resolveOlpHome(opts);
|
||||
const configPath = join(olpHome, 'config.json');
|
||||
const proxyUrl = resolveProxyUrl(opts);
|
||||
// D74 P1-1 fix: server.running and server.version probe /health, which
|
||||
// under default production posture (auth.allow_anonymous: false) requires
|
||||
// an Authorization: Bearer header. Without it, the probe gets 401 and
|
||||
// doctor falsely reports the server down. Caller passes the resolved
|
||||
// bearer token via opts.authHeaders (a `{Authorization: 'Bearer ...'}`
|
||||
// object). Empty headers means "no token configured" — the probe still
|
||||
// fires but a 401 response is treated as "auth misconfigured" rather
|
||||
// than "server down" (see server.running check below).
|
||||
const authHeaders = opts.authHeaders ?? {};
|
||||
|
||||
const checks = [];
|
||||
|
||||
// ── system.* ─────────────────────────────────────────────────────────
|
||||
checks.push({
|
||||
id: 'system.node_version',
|
||||
category: 'system',
|
||||
async run() {
|
||||
// package.json engines.node = >=18; check process.versions.node major >= 18
|
||||
const major = parseInt(String(process.versions.node).split('.')[0], 10);
|
||||
if (Number.isFinite(major) && major >= 18) {
|
||||
return { status: 'ok', message: `Node ${process.versions.node} (>=18)` };
|
||||
}
|
||||
return {
|
||||
status: 'fail',
|
||||
message: `Node ${process.versions.node} is below the required >=18 (per package.json engines.node)`,
|
||||
evidence: {
|
||||
human_steps: [
|
||||
'Install a current Node.js LTS (>=18) — see https://nodejs.org/en/download',
|
||||
],
|
||||
},
|
||||
};
|
||||
},
|
||||
});
|
||||
|
||||
// ── config.* ─────────────────────────────────────────────────────────
|
||||
checks.push({
|
||||
id: 'config.exists',
|
||||
category: 'config',
|
||||
async run() {
|
||||
if (!existsSync(configPath)) {
|
||||
return {
|
||||
status: 'fail',
|
||||
message: `${configPath} not found`,
|
||||
evidence: {
|
||||
// D64-D67 reviewer P2-1: shell-quote paths via _shellQuote so a
|
||||
// malicious OLP_HOME env can't inject shell metacharacters into
|
||||
// the suggested-fix command pasted into an AI agent.
|
||||
fix_commands: [
|
||||
`mkdir -p ${_shellQuote(olpHome)}`,
|
||||
`printf '%s\\n' '{"auth":{"allow_anonymous":false,"owner_only_endpoints":["/health"],"fallback_detail_header_policy":"owner_only"},"providers":{"enabled":{}},"routing":{"chains":{},"soft_triggers":{}},"streaming":{"heartbeat_interval_ms":0}}' > ${_shellQuote(configPath)}`,
|
||||
],
|
||||
reference: 'docs/adr/0007-multi-key-auth.md § 3, docs/adr/0004-fallback-engine.md',
|
||||
},
|
||||
};
|
||||
}
|
||||
try {
|
||||
const parsed = JSON.parse(readFileSync(configPath, 'utf8'));
|
||||
if (parsed && typeof parsed === 'object') {
|
||||
return { status: 'ok', message: `${configPath} parses` };
|
||||
}
|
||||
return { status: 'fail', message: `${configPath} parses but is not a JSON object` };
|
||||
} catch (e) {
|
||||
return {
|
||||
status: 'fail',
|
||||
message: `${configPath} unreadable / malformed: ${e?.message ?? e}`,
|
||||
evidence: {
|
||||
human_steps: [
|
||||
`Inspect ${configPath} — fix the JSON syntax error (or delete the file to fall back to the empty default and re-run olp doctor)`,
|
||||
],
|
||||
},
|
||||
};
|
||||
}
|
||||
},
|
||||
});
|
||||
|
||||
checks.push({
|
||||
id: 'config.providers_enabled',
|
||||
category: 'config',
|
||||
async run() {
|
||||
try {
|
||||
const cfg = loadFallbackConfigSync(configPath);
|
||||
const enabled = cfg.providersEnabled ?? {};
|
||||
const enabledNames = Object.keys(enabled).filter(k => enabled[k] === true);
|
||||
if (enabledNames.length > 0) {
|
||||
return { status: 'ok', message: `${enabledNames.length} provider(s) enabled: ${enabledNames.join(', ')}` };
|
||||
}
|
||||
return {
|
||||
status: 'warn',
|
||||
message: 'No providers enabled in config.json (all /v1/chat/completions requests will 503)',
|
||||
evidence: {
|
||||
human_steps: [
|
||||
`Edit ${configPath} → set providers.enabled.<name> = true for at least one provider (anthropic / openai / mistral)`,
|
||||
],
|
||||
reference: 'docs/adr/0002-plugin-architecture.md § Disable model',
|
||||
},
|
||||
};
|
||||
} catch (e) {
|
||||
return { status: 'fail', message: `Could not read providers.enabled from ${configPath}: ${e?.message ?? e}` };
|
||||
}
|
||||
},
|
||||
});
|
||||
|
||||
checks.push({
|
||||
id: 'config.chains_configured',
|
||||
category: 'config',
|
||||
async run() {
|
||||
try {
|
||||
const cfg = loadFallbackConfigSync(configPath);
|
||||
const chains = cfg.chains ?? {};
|
||||
const chainNames = Object.keys(chains);
|
||||
if (chainNames.length > 0) {
|
||||
return { status: 'ok', message: `${chainNames.length} chain(s) configured: ${chainNames.join(', ')}` };
|
||||
}
|
||||
return {
|
||||
status: 'warn',
|
||||
message: 'No routing chains configured (single-hop mode; cross-provider fallback inactive)',
|
||||
evidence: {
|
||||
reference: 'docs/adr/0004-fallback-engine.md § Chain configuration',
|
||||
},
|
||||
};
|
||||
} catch (e) {
|
||||
return { status: 'fail', message: `Could not read routing.chains from ${configPath}: ${e?.message ?? e}` };
|
||||
}
|
||||
},
|
||||
});
|
||||
|
||||
// ── auth.* ───────────────────────────────────────────────────────────
|
||||
checks.push({
|
||||
id: 'auth.owner_key_exists',
|
||||
category: 'auth',
|
||||
async run() {
|
||||
// Per ADR 0007 § 9.4: process.env.OLP_OWNER_TOKEN also satisfies "owner identity".
|
||||
if (process.env.OLP_OWNER_TOKEN) {
|
||||
return { status: 'ok', message: 'OLP_OWNER_TOKEN env var present (synthetic env-owner per ADR 0007 § 9.4)' };
|
||||
}
|
||||
try {
|
||||
const keys = listKeys({ olpHome });
|
||||
const activeOwner = keys.find(k => k.owner_tier === 'owner' && k.revoked_at === null);
|
||||
if (activeOwner) {
|
||||
return { status: 'ok', message: `active owner key found: id=${activeOwner.id} name="${activeOwner.name}"` };
|
||||
}
|
||||
return {
|
||||
status: 'fail',
|
||||
message: 'No active owner-tier key found in ~/.olp/keys/ and OLP_OWNER_TOKEN env unset',
|
||||
evidence: {
|
||||
fix_commands: [
|
||||
'npx olp-keys keygen --owner',
|
||||
],
|
||||
reference: 'docs/adr/0007-multi-key-auth.md § 9.1 (bootstrap & recovery)',
|
||||
},
|
||||
};
|
||||
} catch (e) {
|
||||
return { status: 'fail', message: `listKeys failed: ${e?.message ?? e}` };
|
||||
}
|
||||
},
|
||||
});
|
||||
|
||||
// ── server.* ─────────────────────────────────────────────────────────
|
||||
if (!opts.skipNetwork) {
|
||||
checks.push({
|
||||
id: 'server.running',
|
||||
category: 'server',
|
||||
async run() {
|
||||
// D74 P1-1: pass authHeaders so the probe works under the default
|
||||
// production posture (auth.allow_anonymous: false).
|
||||
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000, headers: authHeaders });
|
||||
if (!r.ok) {
|
||||
return {
|
||||
status: 'fail',
|
||||
message: `${proxyUrl}/health unreachable: ${r.error}`,
|
||||
evidence: {
|
||||
fix_commands: [
|
||||
'npx olp restart',
|
||||
],
|
||||
reference: 'README.md § Running OLP',
|
||||
},
|
||||
};
|
||||
}
|
||||
// 401: server is up but the caller has no/wrong bearer token. NOT a
|
||||
// "server down" condition — distinguish so the kind discriminator
|
||||
// doesn't route to fix_server when the user just needs OLP_API_KEY.
|
||||
if (r.status === 401 || r.status === 403) {
|
||||
return {
|
||||
status: 'fail',
|
||||
message: `${proxyUrl}/health returned ${r.status} — server is up but the bearer token is missing or invalid. Set OLP_API_KEY env to an owner-tier token (npx olp-keys list).`,
|
||||
evidence: {
|
||||
fix_commands: [
|
||||
'echo "set OLP_API_KEY=<your owner token> or OLP_OWNER_TOKEN=<...> then rerun olp doctor"',
|
||||
],
|
||||
human_required: [
|
||||
'Locate an owner-tier OLP API key plaintext (or run `npx olp-keys keygen --owner` to mint a new one — printed ONCE).',
|
||||
'Export it: `export OLP_API_KEY=olp_...`',
|
||||
],
|
||||
reference: 'docs/adr/0007-multi-key-auth.md § 9.1 + README § Environment Variables',
|
||||
},
|
||||
};
|
||||
}
|
||||
if (r.status !== 200) {
|
||||
return { status: 'fail', message: `${proxyUrl}/health returned status=${r.status}` };
|
||||
}
|
||||
return { status: 'ok', message: `${proxyUrl}/health → 200` };
|
||||
},
|
||||
});
|
||||
|
||||
checks.push({
|
||||
id: 'server.version',
|
||||
category: 'server',
|
||||
async run() {
|
||||
// Read local package.json version
|
||||
let localVersion = null;
|
||||
try {
|
||||
// Resolve relative to this file — lib/doctor.mjs → ../package.json
|
||||
// import.meta.url gives a file:// URL; convert and join.
|
||||
const here = new URL('../package.json', import.meta.url);
|
||||
const pkg = JSON.parse(readFileSync(here, 'utf8'));
|
||||
localVersion = pkg.version ?? null;
|
||||
} catch {
|
||||
return { status: 'warn', message: 'Could not read local package.json — skipping version comparison' };
|
||||
}
|
||||
// D74 P1-1: same auth-headers fix as server.running.
|
||||
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000, headers: authHeaders });
|
||||
if (!r.ok || r.status !== 200) {
|
||||
return { status: 'warn', message: `Could not fetch /health to compare version (${r.error ?? `status ${r.status}`})` };
|
||||
}
|
||||
let serverVersion = null;
|
||||
try {
|
||||
serverVersion = JSON.parse(r.body)?.version ?? null;
|
||||
} catch {
|
||||
return { status: 'warn', message: '/health returned non-JSON; cannot compare version' };
|
||||
}
|
||||
if (!serverVersion) {
|
||||
return { status: 'warn', message: '/health did not include version; cannot compare' };
|
||||
}
|
||||
if (serverVersion === localVersion) {
|
||||
return { status: 'ok', message: `local v${localVersion} matches running v${serverVersion}` };
|
||||
}
|
||||
return {
|
||||
status: 'warn',
|
||||
message: `local v${localVersion} differs from running v${serverVersion} — restart to pick up the new code`,
|
||||
evidence: {
|
||||
fix_commands: [
|
||||
'npx olp restart',
|
||||
],
|
||||
},
|
||||
};
|
||||
},
|
||||
});
|
||||
}
|
||||
|
||||
return checks;
|
||||
}
|
||||
|
||||
/**
|
||||
* Sweep loaded providers for doctorChecks() (ADR 0002 Amendment 7).
|
||||
* Plugins without doctorChecks() contribute nothing (default — back-compat).
|
||||
*
|
||||
* @param {object} opts
|
||||
* @param {Map} [opts.providersOverride] — Map<name, plugin> for tests
|
||||
* @param {object} [opts.providersEnabled] — Record<string, boolean>; default = all from config.json
|
||||
* @returns {Check[]}
|
||||
*/
|
||||
export function collectProviderChecks(opts = {}) {
|
||||
let providers;
|
||||
if (opts.providersOverride) {
|
||||
providers = opts.providersOverride;
|
||||
} else {
|
||||
const olpHome = resolveOlpHome(opts);
|
||||
const configPath = join(olpHome, 'config.json');
|
||||
let enabled = opts.providersEnabled;
|
||||
if (!enabled) {
|
||||
try {
|
||||
enabled = loadFallbackConfigSync(configPath).providersEnabled ?? {};
|
||||
} catch {
|
||||
enabled = {};
|
||||
}
|
||||
}
|
||||
providers = loadProviders({ enabled });
|
||||
}
|
||||
|
||||
const checks = [];
|
||||
for (const [_name, plugin] of providers) {
|
||||
if (typeof plugin?.doctorChecks !== 'function') continue;
|
||||
let pluginChecks;
|
||||
try {
|
||||
pluginChecks = plugin.doctorChecks();
|
||||
} catch (e) {
|
||||
// Misbehaving plugin — surface as a synthesized fail check, do not crash.
|
||||
checks.push({
|
||||
id: `${plugin.name}.doctor_checks_threw`,
|
||||
category: 'provider',
|
||||
async run() {
|
||||
return { status: 'fail', message: `doctorChecks() threw: ${e?.message ?? e}` };
|
||||
},
|
||||
});
|
||||
continue;
|
||||
}
|
||||
if (!Array.isArray(pluginChecks)) continue;
|
||||
for (const c of pluginChecks) {
|
||||
if (c && typeof c.id === 'string' && typeof c.run === 'function') {
|
||||
checks.push({
|
||||
id: c.id,
|
||||
category: c.category ?? 'provider',
|
||||
run: c.run,
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
return checks;
|
||||
}
|
||||
|
||||
// ── Discriminator (kind precedence) ───────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Given a flat results array, determine the next-action discriminator.
|
||||
* Per ADR 0010 § D65 framework:
|
||||
* fresh_install > fix_server > fix_oauth > fix_provider > fix_config > noop
|
||||
*/
|
||||
export function deriveKind(results) {
|
||||
const failed = results.filter(r => r.status === 'fail');
|
||||
if (failed.length === 0) return 'noop';
|
||||
|
||||
if (failed.some(r => r.id === 'config.exists')) return 'fresh_install';
|
||||
if (failed.some(r => r.category === 'server')) return 'fix_server';
|
||||
if (failed.some(r => r.category === 'auth')) return 'fix_oauth';
|
||||
if (failed.some(r => r.category === 'provider')) return 'fix_provider';
|
||||
if (failed.some(r => r.category === 'config')) return 'fix_config';
|
||||
return 'fix_config';
|
||||
}
|
||||
|
||||
/**
|
||||
* Compose the next_action block from FAIL results' evidence.
|
||||
*/
|
||||
export function deriveNextAction(results, kind) {
|
||||
const ai_executable = [];
|
||||
const human_required = [];
|
||||
for (const r of results) {
|
||||
if (r.status !== 'fail') continue;
|
||||
const ev = r.evidence ?? {};
|
||||
if (Array.isArray(ev.fix_commands)) ai_executable.push(...ev.fix_commands);
|
||||
if (Array.isArray(ev.human_steps)) human_required.push(...ev.human_steps);
|
||||
}
|
||||
return {
|
||||
ai_executable,
|
||||
human_required,
|
||||
verify: kind === 'noop' ? 'already healthy' : 'olp doctor',
|
||||
};
|
||||
}
|
||||
|
||||
// ── runDoctor (main entry) ────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Execute every check in parallel; aggregate; derive kind + next_action.
|
||||
*
|
||||
* @param {object} [opts]
|
||||
* @param {Check[]} [opts.injectChecks] — REPLACE the built-in + provider checks entirely
|
||||
* @param {Check[]} [opts.extraChecks] — APPEND extra checks (after defaults)
|
||||
* @param {string} [opts.checkFilter] — restrict to checks whose id OR category matches
|
||||
* @param {Map} [opts.providersOverride] — for the per-provider sweep
|
||||
* @param {object} [opts.providersEnabled] — Record<string, boolean>
|
||||
* @param {string} [opts.olpHome] — override ~/.olp
|
||||
* @param {string} [opts.proxyUrl] — override the proxy URL
|
||||
* @param {boolean} [opts.skipNetwork] — omit server.* checks
|
||||
* @param {() => Date} [opts.now] — clock injection
|
||||
* @returns {Promise<DoctorResult>}
|
||||
*/
|
||||
export async function runDoctor(opts = {}) {
|
||||
const now = opts.now ?? (() => new Date());
|
||||
|
||||
let checks;
|
||||
if (opts.injectChecks) {
|
||||
checks = [...opts.injectChecks];
|
||||
} else {
|
||||
checks = [
|
||||
...buildBuiltinChecks(opts),
|
||||
...collectProviderChecks(opts),
|
||||
];
|
||||
}
|
||||
if (opts.extraChecks) checks.push(...opts.extraChecks);
|
||||
|
||||
// --check <filter>: restrict to checks whose id OR category startsWith / equals the filter.
|
||||
// Match rule: exact id match, exact category match, OR id startsWith `<filter>.`
|
||||
// (so --check anthropic matches both anthropic.cli_available and anthropic.oauth_token_present).
|
||||
if (opts.checkFilter) {
|
||||
const f = String(opts.checkFilter);
|
||||
checks = checks.filter(c =>
|
||||
c.id === f
|
||||
|| c.category === f
|
||||
|| c.id.startsWith(`${f}.`)
|
||||
);
|
||||
}
|
||||
|
||||
// Run all checks in parallel. Capture per-check failures (do not let one throw
|
||||
// hide the rest of the diagnostic).
|
||||
const results = await Promise.all(checks.map(async c => {
|
||||
try {
|
||||
const r = await c.run();
|
||||
return {
|
||||
id: c.id,
|
||||
category: c.category,
|
||||
status: r?.status ?? 'fail',
|
||||
message: r?.message ?? '(check returned no message)',
|
||||
...(r?.evidence !== undefined ? { evidence: r.evidence } : {}),
|
||||
};
|
||||
} catch (e) {
|
||||
return {
|
||||
id: c.id,
|
||||
category: c.category,
|
||||
status: 'fail',
|
||||
message: `check threw: ${e?.message ?? e}`,
|
||||
};
|
||||
}
|
||||
}));
|
||||
|
||||
const fail_count = results.filter(r => r.status === 'fail').length;
|
||||
const warn_count = results.filter(r => r.status === 'warn').length;
|
||||
const ok_count = results.filter(r => r.status === 'ok').length;
|
||||
const kind = deriveKind(results);
|
||||
const next_action = deriveNextAction(results, kind);
|
||||
|
||||
let summary;
|
||||
if (fail_count === 0 && warn_count === 0) {
|
||||
summary = `all ${ok_count} checks ok`;
|
||||
} else if (fail_count === 0) {
|
||||
summary = `${ok_count} ok, ${warn_count} warn — no FAIL; kind=${kind}`;
|
||||
} else {
|
||||
const firstFail = results.find(r => r.status === 'fail');
|
||||
summary = `${fail_count} of ${results.length} checks failed — ${firstFail?.id ?? '?'} (${firstFail?.message?.slice(0, 80) ?? ''})`;
|
||||
}
|
||||
|
||||
return {
|
||||
schema_version: DOCTOR_SCHEMA_VERSION,
|
||||
generated_at: now().toISOString(),
|
||||
checks: results,
|
||||
fail_count,
|
||||
warn_count,
|
||||
ok_count,
|
||||
kind,
|
||||
next_action,
|
||||
summary,
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* @typedef {Object} Check
|
||||
* @property {string} id
|
||||
* @property {'server'|'auth'|'config'|'provider'|'system'} category
|
||||
* @property {() => Promise<{ status: 'ok'|'fail'|'warn', message: string, evidence?: { fix_commands?: string[], human_steps?: string[], reference?: string } }>} run
|
||||
*/
|
||||
|
||||
/**
|
||||
* @typedef {Object} DoctorResult
|
||||
* @property {number} schema_version
|
||||
* @property {string} generated_at (ISO timestamp)
|
||||
* @property {Array<{ id: string, category: string, status: 'ok'|'fail'|'warn', message: string, evidence?: object }>} checks
|
||||
* @property {number} fail_count
|
||||
* @property {number} warn_count
|
||||
* @property {number} ok_count
|
||||
* @property {'noop'|'fix_server'|'fix_oauth'|'fix_config'|'fix_provider'|'fresh_install'} kind
|
||||
* @property {{ ai_executable: string[], human_required: string[], verify: string }} next_action
|
||||
* @property {string} summary
|
||||
*/
|
||||
+14
-2
@@ -667,24 +667,36 @@ function defaultConfigPath() {
|
||||
* Returns empty config (no chains, no soft triggers, no enabled providers) if the
|
||||
* file is absent, unreadable, or malformed.
|
||||
*
|
||||
* D61 (ADR 0010 § Phase 4 D61-D63): adds `streaming` block. Currently
|
||||
* exposes `heartbeat_interval_ms` (default 0 = heartbeat disabled). When
|
||||
* heartbeat_interval_ms > 0, the streaming branch emits `: keepalive\n\n`
|
||||
* SSE comment frames during silent windows of length >= the interval. Default
|
||||
* 0 preserves backwards compat (no behavioural change).
|
||||
*
|
||||
* @param {string} [configPath] — override path (for testing — do NOT write to ~/.olp/config.json in tests)
|
||||
* @returns {{ chains: object, soft_triggers: object, providersEnabled: Record<string, boolean> }}
|
||||
* @returns {{ chains: object, soft_triggers: object, providersEnabled: Record<string, boolean>, streaming: { heartbeat_interval_ms: number } }}
|
||||
*/
|
||||
export function loadFallbackConfigSync(configPath) {
|
||||
const DEFAULT_STREAMING = { heartbeat_interval_ms: 0 };
|
||||
try {
|
||||
const path = configPath ?? defaultConfigPath();
|
||||
const raw = readFileSync(path, 'utf8');
|
||||
const parsed = JSON.parse(raw);
|
||||
const routing = parsed?.routing ?? {};
|
||||
const providers = parsed?.providers ?? {};
|
||||
const streaming = parsed?.streaming ?? {};
|
||||
const hb = Number(streaming.heartbeat_interval_ms);
|
||||
return {
|
||||
chains: routing.chains ?? {},
|
||||
soft_triggers: routing.soft_triggers ?? {},
|
||||
providersEnabled: providers.enabled ?? {},
|
||||
streaming: {
|
||||
heartbeat_interval_ms: Number.isFinite(hb) && hb >= 0 ? hb : 0,
|
||||
},
|
||||
};
|
||||
} catch {
|
||||
// File absent, unreadable, or malformed → no fallback config (single-hop mode)
|
||||
// Empty providersEnabled → all providers disabled → 503 per ALIGNMENT.md v0.1 posture.
|
||||
return { chains: {}, soft_triggers: {}, providersEnabled: {} };
|
||||
return { chains: {}, soft_triggers: {}, providersEnabled: {}, streaming: { ...DEFAULT_STREAMING } };
|
||||
}
|
||||
}
|
||||
|
||||
+17
-2
@@ -27,13 +27,28 @@ export class BadRequestError extends Error {
|
||||
// ── Role normalization ────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* OpenAI deprecated role='function' in favour of role='tool'.
|
||||
* Per ADR 0003, IR supports system/user/assistant/tool.
|
||||
* Normalize entry-surface role names → IR canonical set (system/user/assistant/tool).
|
||||
*
|
||||
* Per ADR 0003, IR supports exactly four roles. OpenAI's chat-completions
|
||||
* spec has evolved beyond that, and we keep the IR minimal by normalizing
|
||||
* at the entry boundary instead of bloating IR + every provider plugin.
|
||||
*
|
||||
* Current normalizations:
|
||||
* - `function` → `tool` — deprecated in OpenAI chat API, replaced by tool.
|
||||
* - `developer` → `system` — OpenAI o1/o3+ reasoning models accept a new
|
||||
* "developer" role with similar semantics to "system" (high-priority
|
||||
* instructions from the developer to the model). Providers like Hermes
|
||||
* Agent and Cline default to `developer` for openai-completions calls.
|
||||
* OLP-side anthropic + codex providers don't differentiate developer
|
||||
* from system, so the IR canonicalizes to `system` and downstream
|
||||
* translations remain unchanged.
|
||||
*
|
||||
* @param {string} role
|
||||
* @returns {string}
|
||||
*/
|
||||
function normalizeRole(role) {
|
||||
if (role === 'function') return 'tool';
|
||||
if (role === 'developer') return 'system';
|
||||
return role;
|
||||
}
|
||||
|
||||
|
||||
+75
-4
@@ -249,7 +249,7 @@ async function _withKeyLock(id, fn) {
|
||||
* never logged. The manifest contains only the hash.
|
||||
*/
|
||||
export function createKey(args = {}) {
|
||||
const { name, owner_tier = 'guest', providers_enabled = '*', notes = '', olpHome } = args;
|
||||
const { name, owner_tier = 'guest', providers_enabled = '*', notes = '', olpHome, plaintext_advertise = false } = args;
|
||||
if (typeof name !== 'string' || name.length === 0) {
|
||||
throw new Error('createKey: name is required (non-empty string)');
|
||||
}
|
||||
@@ -259,6 +259,15 @@ export function createKey(args = {}) {
|
||||
if (!(providers_enabled === '*' || Array.isArray(providers_enabled))) {
|
||||
throw new Error('createKey: providers_enabled must be "*" or string array');
|
||||
}
|
||||
// D69 plaintext_advertise (ADR 0011): only valid on guest tier — see ADR
|
||||
// 0011 § "Trusted-LAN invariant + tier restriction". Owner-tier advertisement
|
||||
// is rejected because exposing the owner identity unauthenticated would
|
||||
// grant unauthenticated callers /health full payload, /v0/management/* access,
|
||||
// and X-OLP-Fallback-Detail visibility — the inverse of the advertise key's
|
||||
// intent (a low-privilege zero-config tier).
|
||||
if (plaintext_advertise && owner_tier !== 'guest') {
|
||||
throw new Error('createKey: plaintext_advertise requires owner_tier="guest" (ADR 0011)');
|
||||
}
|
||||
|
||||
const id = generateKeyId();
|
||||
const plaintext_token = generateToken();
|
||||
@@ -276,10 +285,59 @@ export function createKey(args = {}) {
|
||||
last_used_at: null,
|
||||
notes,
|
||||
};
|
||||
// D69 (ADR 0011): when the operator explicitly opts in via --advertise on
|
||||
// keygen, the plaintext token is co-located with the hash so the server can
|
||||
// surface it via /health.anonymousKey for zero-config family-LAN setup.
|
||||
// This is the ONLY place plaintext ever lands on disk; see ADR 0011 for
|
||||
// the trusted-LAN-only invariant + threat model.
|
||||
if (plaintext_advertise) {
|
||||
manifest.plaintext_advertise = plaintext_token;
|
||||
}
|
||||
writeManifestAtomic(id, manifest, { olpHome });
|
||||
return { id, plaintext_token, manifest };
|
||||
}
|
||||
|
||||
// ── D69 advertise-key discovery (ADR 0011) ───────────────────────────────
|
||||
|
||||
/**
|
||||
* Find the active key marked for /health advertisement. Returns the manifest
|
||||
* (including `plaintext_advertise`) or null when no such key exists.
|
||||
*
|
||||
* Scans every manifest under ~/.olp/keys/; selects the FIRST active
|
||||
* (revoked_at === null) manifest that carries a non-empty `plaintext_advertise`
|
||||
* string. Deterministic ordering is unstable across filesystems — operators
|
||||
* are expected to keep at most one advertised key on disk at a time.
|
||||
*
|
||||
* Returns null if:
|
||||
* - the keys directory doesn't exist
|
||||
* - no manifest carries plaintext_advertise
|
||||
* - the only matching manifest is revoked
|
||||
*
|
||||
* Used by server.mjs handleHealth (D69) + olp-keys CLI 'list' subcommand
|
||||
* (advertise badge).
|
||||
*
|
||||
* @param {object} [opts]
|
||||
* @param {string} [opts.olpHome] - test override; defaults to ~/.olp
|
||||
* @returns {object|null} manifest object (NOT redacted; carries plaintext_advertise)
|
||||
*/
|
||||
export function findAdvertisedKey(opts = {}) {
|
||||
const dir = _keysDir(opts);
|
||||
if (!existsSync(dir)) return null;
|
||||
let entries;
|
||||
try { entries = readdirSync(dir); } catch { return null; }
|
||||
for (const id of entries) {
|
||||
if (id.startsWith('.')) continue;
|
||||
let m;
|
||||
try { m = readManifest(id, opts); } catch { continue; }
|
||||
if (m === null) continue;
|
||||
if (m.revoked_at !== null) continue;
|
||||
if (typeof m.plaintext_advertise === 'string' && m.plaintext_advertise.length > 0) {
|
||||
return m;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* List all keys. Returns array of manifest objects with `token_hash` redacted
|
||||
* (kept on disk; omitted from list output per common operational hygiene —
|
||||
@@ -299,8 +357,14 @@ export function listKeys(opts = {}) {
|
||||
try {
|
||||
const m = readManifest(id, opts);
|
||||
if (m === null) continue;
|
||||
// Redact token_hash from list output (keep on disk).
|
||||
const { token_hash, ...rest } = m;
|
||||
// D69 reviewer P2-1 (footgun-removal): strip BOTH `token_hash` AND
|
||||
// `plaintext_advertise` from list output. Callers wanting the
|
||||
// advertised plaintext for the /health publication path must go
|
||||
// through `findAdvertisedKey()` instead, which is the only sanctioned
|
||||
// read site. Future callers of `listKeys()` that emit results into
|
||||
// logs / HTTP responses / dashboards therefore can't accidentally
|
||||
// leak the advertised plaintext.
|
||||
const { token_hash, plaintext_advertise, ...rest } = m;
|
||||
out.push(rest);
|
||||
} catch {
|
||||
// Skip invalid manifest; production impl would log warn.
|
||||
@@ -456,12 +520,17 @@ export async function touchLastUsed(id, opts = {}) {
|
||||
* - owner_only_endpoints: ['/health'] (D46 consumes; D45 only loads)
|
||||
* - fallback_detail_header_policy: 'owner_only' (D46 consumes; D45 only loads)
|
||||
*
|
||||
* D69 / ADR 0011:
|
||||
* - advertise_anonymous_key: false (default off; opt-in surfaces
|
||||
* findAdvertisedKey() plaintext via
|
||||
* /health.anonymousKey)
|
||||
*
|
||||
* Returns the auth config object. Never throws — missing file / parse
|
||||
* error / missing `auth` key all fall back to defaults.
|
||||
*
|
||||
* @param {object} [opts]
|
||||
* @param {string} [opts.olpHome] - test override; defaults to ~/.olp
|
||||
* @returns {{ allow_anonymous: boolean, owner_only_endpoints: string[], fallback_detail_header_policy: 'owner_only'|'all'|'none' }}
|
||||
* @returns {{ allow_anonymous: boolean, owner_only_endpoints: string[], fallback_detail_header_policy: 'owner_only'|'all'|'none', advertise_anonymous_key: boolean }}
|
||||
*/
|
||||
export function loadAuthConfigSync(opts = {}) {
|
||||
const olpHome = _resolveOlpHome(opts);
|
||||
@@ -470,6 +539,7 @@ export function loadAuthConfigSync(opts = {}) {
|
||||
allow_anonymous: false,
|
||||
owner_only_endpoints: ['/health'],
|
||||
fallback_detail_header_policy: 'owner_only',
|
||||
advertise_anonymous_key: false,
|
||||
};
|
||||
if (!existsSync(path)) return { ...DEFAULTS };
|
||||
try {
|
||||
@@ -484,6 +554,7 @@ export function loadAuthConfigSync(opts = {}) {
|
||||
fallback_detail_header_policy: ['owner_only', 'all', 'none'].includes(auth.fallback_detail_header_policy)
|
||||
? auth.fallback_detail_header_policy
|
||||
: DEFAULTS.fallback_detail_header_policy,
|
||||
advertise_anonymous_key: typeof auth.advertise_anonymous_key === 'boolean' ? auth.advertise_anonymous_key : DEFAULTS.advertise_anonymous_key,
|
||||
};
|
||||
} catch {
|
||||
// Malformed JSON / unreadable file → safe defaults
|
||||
|
||||
+1295
-91
File diff suppressed because it is too large
Load Diff
@@ -30,6 +30,13 @@
|
||||
* collectAllChunks directly. ADR 0002 Amendment 3 (D23).
|
||||
*/
|
||||
|
||||
/**
|
||||
* @typedef {Object} DoctorCheck
|
||||
* @property {string} id - unique per check, e.g. 'anthropic.cli_available'
|
||||
* @property {'provider'} category - fixed for plugin-contributed checks (ADR 0002 Amendment 7)
|
||||
* @property {function} run - async () => { status: 'ok'|'fail'|'warn', message: string, evidence?: { fix_commands?: string[], human_steps?: string[], reference?: string } }
|
||||
*/
|
||||
|
||||
/**
|
||||
* @typedef {Object} ProviderContractV1
|
||||
* @property {string} name - unique lowercase key
|
||||
@@ -41,6 +48,7 @@
|
||||
* @property {function} estimateCost - (request) => {inputTokens, outputTokensEstimate, currency, usd}|null
|
||||
* @property {function} quotaStatus - async (authContext) => {available, percentUsed, resetsAt, pool}|null
|
||||
* @property {function} healthCheck - async () => {ok: boolean, latencyMs: number, error?: string}
|
||||
* @property {function} [doctorChecks] - OPTIONAL () => DoctorCheck[] (ADR 0002 Amendment 7, D67)
|
||||
* @property {ProviderHints} hints
|
||||
*/
|
||||
|
||||
@@ -113,6 +121,13 @@ export function validateProvider(p) {
|
||||
errors.push('healthCheck must be a function');
|
||||
}
|
||||
|
||||
// ADR 0002 Amendment 7 (D67): doctorChecks() is optional. When present it must be
|
||||
// a function; absence is allowed (plugin contributes no provider-tier checks to
|
||||
// `olp doctor` — built-in server/system/auth checks still run).
|
||||
if (p.doctorChecks !== undefined && typeof p.doctorChecks !== 'function') {
|
||||
errors.push('doctorChecks must be a function or omitted');
|
||||
}
|
||||
|
||||
if (!p.hints || typeof p.hints !== 'object') {
|
||||
errors.push('hints must be an object with { requiresTTY, concurrentSpawnSafe, maxConcurrent } + optional { maxSpawnTimeMs, cacheable }');
|
||||
} else {
|
||||
@@ -162,6 +177,15 @@ export const PROVIDER_ERROR_CODES = /** @type {const} */ ([
|
||||
'SPAWN_FAILED',
|
||||
'SPAWN_TIMEOUT', // ADR 0004 § Trigger taxonomy bullet 4: spawn timeout is a hard trigger
|
||||
'CONCURRENCY_LIMIT', // ADR 0002 Amendment 6 / ADR 0004 Amendment 4 (D38, issue #1)
|
||||
/* ADR 0005 Amendment 8 §8 (D57): per-client streaming queue overflow.
|
||||
* NOT a hard trigger — the source spawned successfully and other attached
|
||||
* clients continue to receive chunks; only this client's queue exceeded
|
||||
* PER_CLIENT_QUEUE_CAP. The affected client receives a synthetic
|
||||
* { type: 'stop', finish_reason: 'length' } + [DONE] terminator.
|
||||
* D58 wires server-side header/log surface; HARD_TRIGGER_CODES in
|
||||
* lib/fallback/engine.mjs is a whitelist so absence here gives the
|
||||
* correct default (no fallback advancement). */
|
||||
'STREAM_BACKPRESSURE',
|
||||
]);
|
||||
|
||||
export class ProviderError extends Error {
|
||||
|
||||
+329
-8
@@ -157,6 +157,36 @@ function resolveCodexBin() {
|
||||
// D6 assumption A2: auth file is named auth.json (unconfirmed — D7 will pin).
|
||||
// D6 assumption A3: access token field is `access_token` or `token` (unconfirmed).
|
||||
//
|
||||
// ── D75 (v0.4.2) F1 — codex CLI v0.133.0 schema pin ─────────────────────────
|
||||
//
|
||||
// Real codex CLI v0.133.0 auth.json (verified empirically on PI231 / Mac mini,
|
||||
// 2026-05-26 E2E session):
|
||||
//
|
||||
// {
|
||||
// "auth_mode": "chatgpt",
|
||||
// "OPENAI_API_KEY": null | "<key>",
|
||||
// "tokens": {
|
||||
// "id_token": "<JWT>",
|
||||
// "access_token": "<opaque-or-JWT>", <-- THIS is the access token
|
||||
// "refresh_token": "<opaque>",
|
||||
// "account_id": "<uuid>"
|
||||
// },
|
||||
// "last_refresh": "<ISO8601>"
|
||||
// }
|
||||
//
|
||||
// D6 assumption A3 originally tried `creds.access_token` at TOP level. Under
|
||||
// codex v0.133.0 this field does not exist at the top level → readAuthArtifact()
|
||||
// returned null even when the user had fully completed `codex login`. OLP then
|
||||
// reported "auth artifact missing" via /health and `olp doctor`, and refused
|
||||
// to spawn codex — false negative blocking the entire openai provider.
|
||||
//
|
||||
// Fix: prepend `creds?.tokens?.access_token` to the precedence chain. Keep all
|
||||
// existing fallbacks unchanged so older codex CLI versions (pre-v0.133) and any
|
||||
// future shape variants still resolve.
|
||||
//
|
||||
// Authority pin: codex CLI v0.133.0 source + on-disk auth.json captured during
|
||||
// PI231 E2E. See D75 commit body for the verification transcript.
|
||||
//
|
||||
// Returns { accessToken: string } or null (never throws).
|
||||
export function readAuthArtifact() {
|
||||
// 1. Explicit test override — always takes precedence.
|
||||
@@ -165,7 +195,12 @@ export function readAuthArtifact() {
|
||||
try {
|
||||
const raw = readFileSync(authPathOverride, 'utf8');
|
||||
const creds = JSON.parse(raw);
|
||||
const token = creds?.access_token ?? creds?.token ?? creds?.accessToken;
|
||||
// D75 F1: codex CLI v0.133.0 nests the token under `tokens.access_token`.
|
||||
// Preserve top-level fallbacks for backward / forward compat.
|
||||
const token = creds?.tokens?.access_token
|
||||
?? creds?.access_token
|
||||
?? creds?.token
|
||||
?? creds?.accessToken;
|
||||
if (token && typeof token === 'string') return { accessToken: token };
|
||||
} catch { /* fall through */ }
|
||||
return null; // explicit path set but file missing / malformed
|
||||
@@ -178,8 +213,12 @@ export function readAuthArtifact() {
|
||||
try {
|
||||
const raw = readFileSync(authPath, 'utf8');
|
||||
const creds = JSON.parse(raw);
|
||||
// D6 assumption A3: try common OAuth field names in precedence order.
|
||||
const token = creds?.access_token ?? creds?.token ?? creds?.accessToken;
|
||||
// D75 F1: codex CLI v0.133.0 nests the token under `tokens.access_token`.
|
||||
// Try the nested location FIRST, then fall back to legacy top-level fields.
|
||||
const token = creds?.tokens?.access_token
|
||||
?? creds?.access_token
|
||||
?? creds?.token
|
||||
?? creds?.accessToken;
|
||||
if (token && typeof token === 'string') return { accessToken: token };
|
||||
} catch { /* file missing or malformed */ }
|
||||
|
||||
@@ -242,9 +281,30 @@ export function irToCodex(irRequest) {
|
||||
// model string (e.g., gpt-5.5, gpt-5.4, gpt-5.3-codex).
|
||||
// PROMPT: "Initial instruction for the task. Use '-' to pipe the prompt
|
||||
// from stdin."
|
||||
//
|
||||
// ── D75 (v0.4.2) F2 — codex CLI v0.133.0 trusted-directory sandbox ─────────
|
||||
// codex CLI v0.133.0 added a trusted-directory sandbox: invocations outside
|
||||
// a git repo (or outside any directory explicitly trusted via
|
||||
// `codex config trusted-directories`) refuse with:
|
||||
// "Not inside a trusted directory and --skip-git-repo-check was not specified."
|
||||
// and exit non-zero with zero NDJSON output → OLP surfaces SPAWN_FAILED with
|
||||
// no usable chunks → fallback engine advances to next hop unnecessarily.
|
||||
//
|
||||
// The CWD that OLP spawns from is typically the server install dir (`~/olp/`
|
||||
// on Pi231) which is a git repo on maintainer workstations but is NOT a git
|
||||
// repo on most operator hosts. We bypass the sandbox unconditionally because
|
||||
// OLP is the trusted caller (it is the operator's own server invoking its own
|
||||
// configured Codex subscription via the documented `codex exec` automation
|
||||
// entry point). The trusted-directory sandbox is a foot-gun safeguard for
|
||||
// interactive users; OLP's spawn is non-interactive and pre-authorized.
|
||||
//
|
||||
// Authority: codex CLI v0.133.0 release notes / `codex exec --help` output
|
||||
// documenting `--skip-git-repo-check`. Verified empirically on PI231 E2E
|
||||
// 2026-05-26.
|
||||
const args = [
|
||||
'exec',
|
||||
'--json',
|
||||
'--skip-git-repo-check',
|
||||
'--model', irRequest.model,
|
||||
];
|
||||
|
||||
@@ -291,6 +351,51 @@ export function codexChunkToIR(rawNDJSONLine) {
|
||||
|
||||
if (!event || typeof event !== 'object') return null;
|
||||
|
||||
// ── D75 (v0.4.2) F3 — codex CLI v0.133.0 event shape pin ─────────────────
|
||||
// Real codex CLI v0.133.0 NDJSON event stream (verified empirically on PI231
|
||||
// / Mac mini, 2026-05-26 E2E session):
|
||||
// {"type":"thread.started","thread_id":"019e..."}
|
||||
// {"type":"turn.started"}
|
||||
// {"type":"item.started","item":{"id":"item_0","type":"reasoning","text":""}}
|
||||
// {"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"<response>"}}
|
||||
// {"type":"turn.completed","usage":{"input_tokens":..,"output_tokens":..}}
|
||||
//
|
||||
// The D6 defensive parser recognized `content`/`delta`/`text` fields at the
|
||||
// top level and `type === 'stop'`/`done === true`. None of these match
|
||||
// v0.133.0's actual shape → every chunk was silently dropped → response body
|
||||
// had `content: null`. F3 adds three NEW recognizers (item.completed →
|
||||
// agent_message; turn.completed → stop; turn.failed → error) BEFORE the
|
||||
// legacy fallback chain. Legacy recognizers preserved for forward/backward
|
||||
// compat (older codex versions; future shape variants).
|
||||
|
||||
// F3-a: agent_message item completion.
|
||||
// codex v0.133.0 emits assistant text as a single item.completed event whose
|
||||
// item.type is 'agent_message' and item.text carries the full text. There
|
||||
// are no incremental deltas — the entire response arrives in one chunk.
|
||||
if (event.type === 'item.completed'
|
||||
&& event.item?.type === 'agent_message'
|
||||
&& typeof event.item?.text === 'string') {
|
||||
return { type: 'delta', content: event.item.text };
|
||||
}
|
||||
|
||||
// F3-b: turn completion → stop chunk.
|
||||
// codex v0.133.0 emits turn.completed with a usage block when the model
|
||||
// finishes. We map this to IR stop with finish_reason 'stop'.
|
||||
if (event.type === 'turn.completed') {
|
||||
return { type: 'stop', finish_reason: 'stop' };
|
||||
}
|
||||
|
||||
// F3-c: turn failure → error chunk.
|
||||
// codex v0.133.0 emits turn.failed with an embedded error object when the
|
||||
// turn cannot complete. Extract a human-readable message for the IR error.
|
||||
if (event.type === 'turn.failed') {
|
||||
const errMsg = (typeof event.error === 'string')
|
||||
? event.error
|
||||
: (event.error?.message ?? 'codex turn.failed');
|
||||
return { type: 'error', error: errMsg };
|
||||
}
|
||||
|
||||
// ── Legacy/fallback recognizers (kept for backward + forward compat) ────
|
||||
// Error event: type === 'error' or error field present
|
||||
// A4: defensive — error shape unconfirmed; D7 will pin actual field names
|
||||
if (event.type === 'error' || (event.error && typeof event.error === 'string')) {
|
||||
@@ -353,7 +458,7 @@ function buildSpawnEnv() {
|
||||
//
|
||||
// Authority: Codex CLI reference § "codex exec [flags] PROMPT"
|
||||
// § "--json": NDJSON event stream on stdout
|
||||
async function* _spawnAndStream(irRequest, authContext, spawnImpl) {
|
||||
async function* _spawnAndStream(irRequest, authContext, spawnImpl, isolationCtx) {
|
||||
const auth = authContext ?? readAuthArtifact();
|
||||
if (!auth?.accessToken) {
|
||||
throw new ProviderError(
|
||||
@@ -363,7 +468,7 @@ async function* _spawnAndStream(irRequest, authContext, spawnImpl) {
|
||||
}
|
||||
|
||||
const bin = resolveCodexBin();
|
||||
const { args, prompt, useStdin } = irToCodex(irRequest);
|
||||
const { args: baseArgs, prompt, useStdin } = irToCodex(irRequest);
|
||||
const env = buildSpawnEnv();
|
||||
|
||||
// Authority: Codex CLI reference § "Authentication"
|
||||
@@ -371,7 +476,34 @@ async function* _spawnAndStream(irRequest, authContext, spawnImpl) {
|
||||
// No explicit token injection: Codex CLI reads its own auth.json
|
||||
// (contrast with Anthropic plugin which injects CLAUDE_CODE_OAUTH_TOKEN).
|
||||
|
||||
const proc = spawnImpl(bin, args, { env, stdio: ['pipe', 'pipe', 'pipe'] });
|
||||
// Task #8 — Phase 7 Solution 1: apply isolation context from orchestrator.
|
||||
// isolationCtx is provided by server.mjs (prepareIsolatedEnvironment) when
|
||||
// present. Three layers compose here:
|
||||
// Layer 1 (env): envOverrides (HOME, CODEX_HOME) have final precedence.
|
||||
// Layer 4 (args): hardenedArgs injects --sandbox read-only + -c approval_policy.
|
||||
// Layer 3 (wrap): wrapForLayer3 is identity for codex (hasInnerSandbox=true).
|
||||
// When isolationCtx is absent (legacy callers / tests), behavior is unchanged.
|
||||
// Authority: ADR 0014 Amendment 1 § A1.2 + ADR 0002 Amendment 9 § Backward compat.
|
||||
const envOverrides = isolationCtx?.envOverrides ?? {};
|
||||
const finalEnv = Object.keys(envOverrides).length > 0 ? { ...env, ...envOverrides } : env;
|
||||
|
||||
const hardenedArgs = isolationCtx?.hardenedArgs ?? ((a) => a);
|
||||
const args = hardenedArgs(baseArgs);
|
||||
|
||||
// Layer 3: wrapForLayer3 for codex is always identity (hasInnerSandbox=true);
|
||||
// included here for API symmetry with the anthropic path and future-proofing.
|
||||
const wrapForLayer3 = isolationCtx?.wrapForLayer3 ?? (async (c) => c);
|
||||
const wrappedBin = await wrapForLayer3(bin);
|
||||
let finalBin, finalArgs;
|
||||
if (wrappedBin !== bin) {
|
||||
finalBin = '/bin/sh';
|
||||
finalArgs = ['-c', wrappedBin];
|
||||
} else {
|
||||
finalBin = bin;
|
||||
finalArgs = args;
|
||||
}
|
||||
|
||||
const proc = spawnImpl(finalBin, finalArgs, { env: finalEnv, stdio: ['pipe', 'pipe', 'pipe'] });
|
||||
|
||||
// Write prompt via stdin for multi-line prompts (D6 assumption A1)
|
||||
if (useStdin) {
|
||||
@@ -554,8 +686,14 @@ async function* _spawnAndStream(irRequest, authContext, spawnImpl) {
|
||||
// spawn: async (irRequest, authContext) => AsyncIterator<ResponseChunk>
|
||||
let _spawnImpl = defaultSpawn;
|
||||
|
||||
export async function* spawn(irRequest, authContext) {
|
||||
yield* _spawnAndStream(irRequest, authContext, _spawnImpl);
|
||||
// Task #8 — Phase 7 Solution 1: isolationCtx is an optional third argument.
|
||||
// When present (from server.mjs prepareIsolatedEnvironment call), it carries
|
||||
// { envOverrides, hardenedArgs, wrapForLayer3, cleanup } — the orchestrator
|
||||
// composes these on top of the provider's own env-cleanup + args composition.
|
||||
// When absent (legacy callers, tests that don't pass it), behavior is unchanged.
|
||||
// Authority: ADR 0014 Amendment 1 § A1.2.
|
||||
export async function* spawn(irRequest, authContext, isolationCtx) {
|
||||
yield* _spawnAndStream(irRequest, authContext, _spawnImpl, isolationCtx);
|
||||
}
|
||||
|
||||
// Test hook: inject mock spawn without importing child_process.
|
||||
@@ -637,6 +775,187 @@ function _defaultBinaryExists() {
|
||||
}
|
||||
}
|
||||
|
||||
// ── doctorChecks (ADR 0002 Amendment 7, D67) ──────────────────────────────
|
||||
// See lib/providers/anthropic.mjs doctorChecks header for the contract.
|
||||
//
|
||||
// Probes:
|
||||
// openai.cli_available — `codex --version` resolves on PATH (or via OLP_CODEX_BIN)
|
||||
// openai.auth_present — `readAuthArtifact()` returns a non-empty accessToken
|
||||
// (Codex CLI reference § Authentication: credentials in $CODEX_HOME, default ~/.codex/auth.json)
|
||||
export function doctorChecks({ _binaryExistsFn, _authReadFn } = {}) {
|
||||
const binaryExists = _binaryExistsFn ?? _defaultBinaryExists;
|
||||
const authRead = _authReadFn ?? readAuthArtifact;
|
||||
return [
|
||||
{
|
||||
id: 'openai.cli_available',
|
||||
category: 'provider',
|
||||
async run() {
|
||||
if (binaryExists()) {
|
||||
return { status: 'ok', message: '`codex` binary resolved on PATH' };
|
||||
}
|
||||
return {
|
||||
status: 'fail',
|
||||
message: '`codex` binary not found on PATH (and OLP_CODEX_BIN unset/invalid)',
|
||||
evidence: {
|
||||
fix_commands: [
|
||||
'npm install -g @openai/codex',
|
||||
],
|
||||
reference: 'https://developers.openai.com/codex/cli/reference',
|
||||
},
|
||||
};
|
||||
},
|
||||
},
|
||||
{
|
||||
id: 'openai.auth_present',
|
||||
category: 'provider',
|
||||
async run() {
|
||||
const auth = authRead();
|
||||
if (auth?.accessToken) {
|
||||
return { status: 'ok', message: 'Codex auth artifact present ($CODEX_HOME/auth.json)' };
|
||||
}
|
||||
return {
|
||||
status: 'fail',
|
||||
message: 'Codex auth artifact missing — $CODEX_HOME/auth.json does not contain access_token (default $CODEX_HOME=~/.codex)',
|
||||
evidence: {
|
||||
human_steps: [
|
||||
'run: codex (the first interactive launch prompts for OAuth login per Codex CLI reference § Authentication)',
|
||||
],
|
||||
reference: 'https://developers.openai.com/codex/cli/reference',
|
||||
},
|
||||
};
|
||||
},
|
||||
},
|
||||
];
|
||||
}
|
||||
|
||||
// ── ISOLATION export ─────────────────────────────────────────────────────
|
||||
// Declares per-provider isolation primitives consumed by lib/sandbox/manager.mjs
|
||||
// (per ADR 0014 Amendment 1 + ADR 0002 Amendment 9).
|
||||
//
|
||||
// Authority citations (all required per ALIGNMENT.md Rule 1):
|
||||
// codex CLI v0.133.0 — current PI231 prod version (verified 2026-05-29 spike)
|
||||
// https://developers.openai.com/codex/config-reference — CODEX_HOME env var
|
||||
// (2 occurrences verified: "$CODEX_HOME/profile-name.config.toml" and
|
||||
// "$CODEX_HOME/log" path templates)
|
||||
// https://developers.openai.com/codex/auth/ — ~/.codex/auth.json path
|
||||
// (2 occurrences verified: "auth.json under CODEX_HOME" credential-storage
|
||||
// section)
|
||||
// https://developers.openai.com/codex/concepts/sandboxing — --sandbox flag +
|
||||
// read-only default (codex inner bubblewrap sandbox)
|
||||
// openai/codex#16018 — inner bwrap behavior documented (failure under
|
||||
// restricted env, establishing hasInnerSandbox: true)
|
||||
// ADR 0014 Amendment 1 — orchestrator composition architecture
|
||||
// ADR 0002 Amendment 9 — ISOLATION contract spec (field semantics)
|
||||
// docs/spikes/2026-05-29-ephemeral-home.md § 5.3 — flag-drift caveat
|
||||
// (--ask-for-approval removed in codex v0.133.0; use -c approval_policy=)
|
||||
//
|
||||
// isolation rationale: OpenAI Codex's `codex exec` exposes a shell tool that
|
||||
// actually executes commands during the spawn (cc-mem incident memory § 3.2).
|
||||
// The CLI provides its own inner bubblewrap sandbox (`--sandbox read-only` by
|
||||
// default per https://developers.openai.com/codex/concepts/sandboxing) that
|
||||
// confines shell tool reads/writes. The orchestrator's outer isolation composes
|
||||
// with the inner sandbox: credential-dir redirect via CODEX_HOME
|
||||
// (https://developers.openai.com/codex/config-reference) + HOME redirect for
|
||||
// the inner bwrap's HOME lookup + per-spawn ephemeral credential mount.
|
||||
// hasInnerSandbox: true so the outer profile is relaxed to permit the inner
|
||||
// bwrap's user-namespace clone (openai/codex#16018).
|
||||
|
||||
export const ISOLATION = {
|
||||
// ephemeralEnvOverrides: pure function, no side effects, no fs access.
|
||||
// CODEX_HOME redirects the entire codex config/credential base directory.
|
||||
// HOME is also redirected because the codex inner sandbox inherits the parent
|
||||
// process's HOME for its own home lookup unless overridden.
|
||||
// Authority: CODEX_HOME → https://developers.openai.com/codex/config-reference
|
||||
// HOME → POSIX convention (both verified by PI231 spike § 4.3-4.4).
|
||||
ephemeralEnvOverrides: ({ ephemeralRoot, keyId: _keyId, reqId: _reqId }) => ({
|
||||
HOME: ephemeralRoot,
|
||||
CODEX_HOME: `${ephemeralRoot}/.codex`,
|
||||
}),
|
||||
|
||||
// credentialMounts: static list of [srcAbsPath, dstRelativeToEphemeralRoot].
|
||||
// srcAbsPath uses os.homedir() (imported as `homedir` at top of file) per
|
||||
// ADR 0002 Amendment 9 § Field 2 validation rules: absolute paths only, no
|
||||
// `~/` prefixes (shell-expansion semantics differ from Node.js behavior).
|
||||
// Authority: ~/.codex/auth.json → https://developers.openai.com/codex/auth/
|
||||
// "Codex caches login details locally in a plaintext file at ~/.codex/auth.json"
|
||||
// (matches existing codex.mjs `auth.path` field declaration above).
|
||||
credentialMounts: [
|
||||
[join(homedir(), '.codex', 'auth.json'), '.codex/auth.json'],
|
||||
],
|
||||
|
||||
// requiredHomePaths: directories to mkdir-p under ephemeralRoot before mounts.
|
||||
// .codex is required because CODEX_HOME points there and codex startup may
|
||||
// attempt to read from it before any auto-create logic runs (observed in
|
||||
// PI231 spike § 4.3 post-state: .codex/ created at spawn time).
|
||||
requiredHomePaths: [
|
||||
'.codex',
|
||||
],
|
||||
|
||||
// hasInnerSandbox: true — codex exec spawns its own bubblewrap sandbox
|
||||
// internally. Declaring true tells the outer isolation orchestrator to relax
|
||||
// the outer profile to permit clone(CLONE_NEWUSER) so the inner bwrap can
|
||||
// create user namespaces. Without this flag the inner bwrap fails with
|
||||
// EPERM. Authority: openai/codex#16018 + https://developers.openai.com/codex/concepts/sandboxing
|
||||
hasInnerSandbox: true,
|
||||
|
||||
// crossTenantReadProtection: 'inner-sandbox' — codex's shell tool runs real
|
||||
// commands but the inner bubblewrap sandbox (read-only by default) confines
|
||||
// reads/writes to the inner namespace. The toolHardeningArgs below makes this
|
||||
// default explicit at the spawn-args level. Authority: openai/codex#16018 +
|
||||
// https://developers.openai.com/codex/concepts/sandboxing.
|
||||
crossTenantReadProtection: 'inner-sandbox',
|
||||
|
||||
// recommendedDeploymentTier: 'per-os-user' — the inner bwrap sandbox protects
|
||||
// against accidental cross-tenant leakage from the model's shell tool, but a
|
||||
// sandbox-escape CVE (e.g. in bubblewrap) would expose the OS-user filesystem.
|
||||
// Per-OS-user isolation adds defense in depth. See ADR 0002 Amendment 9
|
||||
// § Field 6 for the full rationale per recommendedDeploymentTier semantics.
|
||||
recommendedDeploymentTier: 'per-os-user',
|
||||
|
||||
// toolHardeningArgs: injects --sandbox read-only if not already present, and
|
||||
// -c approval_policy="never" to suppress interactive approval prompts.
|
||||
//
|
||||
// Flag-drift caveat (docs/spikes/2026-05-29-ephemeral-home.md § 5.3):
|
||||
// ADR 0002 Amendment 9 § codex example uses `--ask-for-approval never`.
|
||||
// PI231 spike (2026-05-29) confirmed this flag was REMOVED in codex
|
||||
// v0.133.0. The codex v0.133.0 `--help` output shows the replacement is
|
||||
// the generic config-override flag: `-c approval_policy="never"`.
|
||||
// We use `-c approval_policy="never"` here. This deviates from the ADR
|
||||
// 0002 Amendment 9 code example (not the field spec — the spec only
|
||||
// requires an injected flag corresponding to a documented CLI flag).
|
||||
// The config-override form is documented at https://developers.openai.com/codex/config-reference
|
||||
// as the mechanism for overriding any config key at spawn time, including
|
||||
// approval_policy. The deviation is intentional, flag-drift-driven, and
|
||||
// takes precedence over the (now-incorrect) Amendment 9 code example per
|
||||
// ALIGNMENT.md Rule 2 (provider CLI is the authority, not the ADR text).
|
||||
//
|
||||
// --sandbox read-only: Authority: https://developers.openai.com/codex/concepts/sandboxing
|
||||
// § "Sandboxing modes" — the default posture is `read-only`; injecting it
|
||||
// explicitly prevents a future codex default change from silently weakening
|
||||
// isolation (same rationale as the existing irToCodex --skip-git-repo-check).
|
||||
toolHardeningArgs: (existingArgs) => {
|
||||
let result = [...existingArgs];
|
||||
|
||||
// Inject --sandbox read-only if the caller has not already specified --sandbox.
|
||||
if (!result.some(arg => arg === '--sandbox' || arg.startsWith('--sandbox='))) {
|
||||
result = [...result, '--sandbox', 'read-only'];
|
||||
}
|
||||
|
||||
// Inject -c approval_policy="never" if not already present.
|
||||
// Checks for the exact -c flag form used by codex v0.133.0 config overrides.
|
||||
// Flag-drift note: --ask-for-approval (pre-v0.133.0) is NOT injected — it
|
||||
// was removed; see header comment above.
|
||||
const approvalAlreadySet = result.some(
|
||||
(arg, i) => arg === '-c' && typeof result[i + 1] === 'string' && result[i + 1].startsWith('approval_policy'),
|
||||
);
|
||||
if (!approvalAlreadySet) {
|
||||
result = [...result, '-c', 'approval_policy="never"'];
|
||||
}
|
||||
|
||||
return result;
|
||||
},
|
||||
};
|
||||
|
||||
// ── Provider export ───────────────────────────────────────────────────────
|
||||
// Conforms to ADR 0002 § "Provider contract (v1.0 interface)" + contractVersion.
|
||||
|
||||
@@ -662,6 +981,8 @@ const codex = {
|
||||
estimateCost,
|
||||
quotaStatus,
|
||||
healthCheck,
|
||||
// ADR 0002 Amendment 7 (D67): OPTIONAL doctorChecks() — consumed by `olp doctor`.
|
||||
doctorChecks: () => doctorChecks(),
|
||||
hints: {
|
||||
requiresTTY: false, // codex exec runs headless (per CLI reference § exec)
|
||||
concurrentSpawnSafe: true, // each invocation is independent
|
||||
|
||||
+14
-3
@@ -29,12 +29,23 @@
|
||||
*/
|
||||
|
||||
import { validateProvider } from './base.mjs';
|
||||
import anthropicDefault from './anthropic.mjs';
|
||||
import codexDefault from './codex.mjs';
|
||||
import anthropicDefault, { ISOLATION as anthropicISOLATION } from './anthropic.mjs';
|
||||
import codexDefault, { ISOLATION as codexISOLATION } from './codex.mjs';
|
||||
import mistralDefault from './mistral.mjs';
|
||||
import modelsRegistryRaw from '../../models-registry.json' with { type: 'json' };
|
||||
|
||||
// Normalize default export pattern
|
||||
// Attach Phase 7 ISOLATION contract per ADR 0002 Amendment 9. The ISOLATION
|
||||
// block is a top-level named export from each provider plugin; the loader
|
||||
// attaches it as a property of the default-export object so the orchestrator
|
||||
// (lib/sandbox/manager.mjs prepareIsolatedEnvironment) can read it as
|
||||
// provider.ISOLATION. In-place mutation (not spread) preserves the default
|
||||
// export's object identity, which downstream code (cache store keyed on
|
||||
// provider, singleflight Maps) relies on. Providers without ISOLATION
|
||||
// (mistral at present) fall through to legacy unsandboxed shape per
|
||||
// ADR 0002 Amendment 9 § Backward compatibility.
|
||||
if (anthropicISOLATION) anthropicDefault.ISOLATION = anthropicISOLATION;
|
||||
if (codexISOLATION) codexDefault.ISOLATION = codexISOLATION;
|
||||
|
||||
const anthropic = anthropicDefault;
|
||||
const codex = codexDefault;
|
||||
const mistral = mistralDefault;
|
||||
|
||||
@@ -764,6 +764,60 @@ function _defaultBinaryExists() {
|
||||
}
|
||||
}
|
||||
|
||||
// ── doctorChecks (ADR 0002 Amendment 7, D67) ──────────────────────────────
|
||||
// See lib/providers/anthropic.mjs doctorChecks header for the contract.
|
||||
//
|
||||
// Probes:
|
||||
// mistral.cli_available — `vibe --version` resolves on PATH (or via OLP_VIBE_BIN)
|
||||
// mistral.api_key_present — readAuthArtifact() returns apiKey (MISTRAL_API_KEY env or ~/.vibe/.env)
|
||||
// (DOCS-2: https://docs.mistral.ai/mistral-vibe/terminal/configuration —
|
||||
// auth from MISTRAL_API_KEY env / ~/.vibe/.env)
|
||||
export function doctorChecks({ _binaryExistsFn, _authReadFn } = {}) {
|
||||
const binaryExists = _binaryExistsFn ?? _defaultBinaryExists;
|
||||
const authRead = _authReadFn ?? readAuthArtifact;
|
||||
return [
|
||||
{
|
||||
id: 'mistral.cli_available',
|
||||
category: 'provider',
|
||||
async run() {
|
||||
if (binaryExists()) {
|
||||
return { status: 'ok', message: '`vibe` binary resolved on PATH' };
|
||||
}
|
||||
return {
|
||||
status: 'fail',
|
||||
message: '`vibe` binary not found on PATH (and OLP_VIBE_BIN unset/invalid)',
|
||||
evidence: {
|
||||
fix_commands: [
|
||||
'npm install -g @mistralai/vibe',
|
||||
],
|
||||
reference: 'https://docs.mistral.ai/mistral-vibe/terminal/quickstart',
|
||||
},
|
||||
};
|
||||
},
|
||||
},
|
||||
{
|
||||
id: 'mistral.api_key_present',
|
||||
category: 'provider',
|
||||
async run() {
|
||||
const auth = authRead();
|
||||
if (auth?.apiKey) {
|
||||
return { status: 'ok', message: 'Mistral API key present (env MISTRAL_API_KEY or ~/.vibe/.env)' };
|
||||
}
|
||||
return {
|
||||
status: 'fail',
|
||||
message: 'Mistral API key missing — neither MISTRAL_API_KEY env nor ~/.vibe/.env (or $VIBE_HOME/.env) supplied a key',
|
||||
evidence: {
|
||||
human_steps: [
|
||||
'export MISTRAL_API_KEY=<your-key> # or write MISTRAL_API_KEY=... into ~/.vibe/.env',
|
||||
],
|
||||
reference: 'https://docs.mistral.ai/mistral-vibe/terminal/configuration',
|
||||
},
|
||||
};
|
||||
},
|
||||
},
|
||||
];
|
||||
}
|
||||
|
||||
// ── Provider export ───────────────────────────────────────────────────────
|
||||
// Conforms to ADR 0002 § "Provider contract (v1.0 interface)" + contractVersion.
|
||||
|
||||
@@ -795,6 +849,8 @@ const mistral = {
|
||||
estimateCost,
|
||||
quotaStatus,
|
||||
healthCheck,
|
||||
// ADR 0002 Amendment 7 (D67): OPTIONAL doctorChecks() — consumed by `olp doctor`.
|
||||
doctorChecks: () => doctorChecks(),
|
||||
hints: {
|
||||
requiresTTY: false, // vibe --prompt runs headless per DOCS-1 programmatic mode
|
||||
concurrentSpawnSafe: true, // each invocation is independent
|
||||
|
||||
@@ -0,0 +1,290 @@
|
||||
/**
|
||||
* lib/sandbox/doctor.mjs — Sandbox availability preflight module (Phase 7 PR-A)
|
||||
*
|
||||
* Authority:
|
||||
* @anthropic-ai/sandbox-runtime v0.0.52
|
||||
* https://github.com/anthropic-experimental/sandbox-runtime
|
||||
*
|
||||
* 2026-05-28 PoC spike on PI231 (arm64 Debian Bookworm): dep install clean,
|
||||
* isSupportedPlatform()=true, blocked on apt deps (bwrap + socat), three PoC
|
||||
* scripts parked at /tmp/sandbox-spike/ on PI231.
|
||||
*
|
||||
* OLP ADR 0014 — Sandbox-Runtime Integration for Multi-Tenant Provider Spawning
|
||||
* OLP ADR 0009 Amendment 1 § Caveats #3 (sandbox is cloud prerequisite)
|
||||
* docs/plans/cloud-deployment-family.md § 5
|
||||
*
|
||||
* Design:
|
||||
* Pure module — no state, no side effects beyond child_process.execFileSync for
|
||||
* `which` probes. Does NOT call SandboxManager.initialize(). Does NOT create
|
||||
* or interact with any real sandbox. Safe to call from /health on every request
|
||||
* (results are memoized process-wide by the caller in server.mjs — see
|
||||
* _sandboxStatusCache there).
|
||||
*
|
||||
* Exports:
|
||||
* checkSandboxAvailability() — returns { available, missing, details }
|
||||
* describeSandboxStatus() — returns { ok, message } human-readable summary
|
||||
*/
|
||||
|
||||
import { execFileSync } from 'node:child_process';
|
||||
import { platform as osPlatform } from 'node:os';
|
||||
|
||||
// ── which probe helper ────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Check if a binary is in PATH by running `which <binary>`.
|
||||
* Returns true if found, false if not found or if `which` is unavailable.
|
||||
* Never throws.
|
||||
* @param {string} binary
|
||||
* @returns {boolean}
|
||||
*/
|
||||
function isInPath(binary) {
|
||||
try {
|
||||
execFileSync('which', [binary], { stdio: 'pipe', timeout: 2000 });
|
||||
return true;
|
||||
} catch {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
// ── Platform helper ───────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Map Node's process.platform to the sandbox-runtime platform string.
|
||||
* @returns {'linux'|'macos'|'other'}
|
||||
*/
|
||||
function getPlatformName() {
|
||||
const p = osPlatform();
|
||||
if (p === 'linux') return 'linux';
|
||||
if (p === 'darwin') return 'macos';
|
||||
return 'other';
|
||||
}
|
||||
|
||||
// ── Library introspection ─────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Attempt to import @anthropic-ai/sandbox-runtime and call its exported
|
||||
* isSupportedPlatform + checkDependencies. Returns structured findings.
|
||||
* Never throws — all errors become { libError: <message> }.
|
||||
*
|
||||
* @returns {Promise<{
|
||||
* libLoaded: boolean,
|
||||
* libError: string|null,
|
||||
* isSupportedPlatform: boolean,
|
||||
* libDependencyErrors: string[],
|
||||
* libDependencyWarnings: string[],
|
||||
* }>}
|
||||
*/
|
||||
async function probeLibrary() {
|
||||
try {
|
||||
const { SandboxManager } = await import('@anthropic-ai/sandbox-runtime');
|
||||
|
||||
let supportedPlatform = false;
|
||||
try {
|
||||
supportedPlatform = SandboxManager.isSupportedPlatform();
|
||||
} catch (e) {
|
||||
return {
|
||||
libLoaded: true,
|
||||
libError: `isSupportedPlatform() threw: ${e?.message ?? e}`,
|
||||
isSupportedPlatform: false,
|
||||
libDependencyErrors: [],
|
||||
libDependencyWarnings: [],
|
||||
};
|
||||
}
|
||||
|
||||
// checkDependencies() requires initialize() to have been called first to
|
||||
// set ripgrep/bwrap/socat config. Since PR-A never calls initialize(), we
|
||||
// call checkDependencies() with an undefined argument — the library falls
|
||||
// back to { command: 'rg' } for ripgrep and PATH lookup for bwrap/socat,
|
||||
// which is exactly what we want for the doctor preflight.
|
||||
let libDependencyErrors = [];
|
||||
let libDependencyWarnings = [];
|
||||
if (supportedPlatform) {
|
||||
try {
|
||||
const depCheck = SandboxManager.checkDependencies(undefined);
|
||||
libDependencyErrors = depCheck?.errors ?? [];
|
||||
libDependencyWarnings = depCheck?.warnings ?? [];
|
||||
} catch (e) {
|
||||
// checkDependencies() can throw before initialize() — not fatal
|
||||
libDependencyErrors = [`checkDependencies() threw: ${e?.message ?? e}`];
|
||||
}
|
||||
}
|
||||
|
||||
return {
|
||||
libLoaded: true,
|
||||
libError: null,
|
||||
isSupportedPlatform: supportedPlatform,
|
||||
libDependencyErrors,
|
||||
libDependencyWarnings,
|
||||
};
|
||||
} catch (e) {
|
||||
return {
|
||||
libLoaded: false,
|
||||
libError: `@anthropic-ai/sandbox-runtime import failed: ${e?.message ?? e}`,
|
||||
isSupportedPlatform: false,
|
||||
libDependencyErrors: [],
|
||||
libDependencyWarnings: [],
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
// ── Public API ────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Check sandbox availability (OS deps + library platform support).
|
||||
*
|
||||
* Returns:
|
||||
* {
|
||||
* available: boolean, // true only when all hard deps pass on a supported platform
|
||||
* missing: string[], // friendly names of missing hard deps (e.g. 'bubblewrap', 'socat')
|
||||
* details: {
|
||||
* platform: string, // 'linux'|'macos'|'other'
|
||||
* bwrap: boolean, // which bwrap → found
|
||||
* socat: boolean, // which socat → found
|
||||
* ripgrep: boolean, // which rg → found
|
||||
* isSupportedPlatform: boolean,
|
||||
* libLoaded: boolean,
|
||||
* libError: string|null,
|
||||
* libDependencyErrors: string[],
|
||||
* libDependencyWarnings: string[],
|
||||
* }
|
||||
* }
|
||||
*
|
||||
* The `missing` array uses human-readable package names ('bubblewrap', 'socat',
|
||||
* 'ripgrep') so that install hints are directly actionable.
|
||||
*
|
||||
* Does NOT call SandboxManager.initialize() — pure inspection only.
|
||||
* Does NOT cache — the caller (server.mjs) memoizes the result.
|
||||
*/
|
||||
export async function checkSandboxAvailability() {
|
||||
const platform = getPlatformName();
|
||||
|
||||
// Probe OS-level deps independently of the library (the `which` calls are
|
||||
// cheap and always correct; library's checkDependencies may be less precise
|
||||
// when initialize() hasn't been called).
|
||||
const bwrap = isInPath('bwrap');
|
||||
const socat = isInPath('socat');
|
||||
const ripgrep = isInPath('rg');
|
||||
|
||||
// Library introspection (import + isSupportedPlatform + checkDependencies)
|
||||
const lib = await probeLibrary();
|
||||
|
||||
// Determine what's missing for the doctor report.
|
||||
// Only report OS deps as missing on Linux (where bwrap/socat/rg are required);
|
||||
// macOS uses sandbox-exec which is built-in, so these are not hard requirements.
|
||||
const missing = [];
|
||||
if (platform === 'linux') {
|
||||
if (!bwrap) missing.push('bubblewrap');
|
||||
if (!socat) missing.push('socat');
|
||||
if (!ripgrep) missing.push('ripgrep');
|
||||
}
|
||||
// If the library itself failed to load, that's also a blocker
|
||||
if (!lib.libLoaded) {
|
||||
missing.push('@anthropic-ai/sandbox-runtime (import failed)');
|
||||
}
|
||||
// Library-reported hard dep errors (may overlap with our `which` probes;
|
||||
// deduplicate by treating them as additional evidence rather than re-adding)
|
||||
for (const errMsg of lib.libDependencyErrors) {
|
||||
// Only add if it doesn't overlap with what we already reported
|
||||
const isAlreadyCovered =
|
||||
(errMsg.includes('bwrap') && !bwrap) ||
|
||||
(errMsg.includes('socat') && !socat) ||
|
||||
(errMsg.includes('ripgrep') && !ripgrep) ||
|
||||
(errMsg.includes('Unsupported platform'));
|
||||
if (!isAlreadyCovered && !missing.includes(errMsg)) {
|
||||
missing.push(errMsg);
|
||||
}
|
||||
}
|
||||
|
||||
const available =
|
||||
lib.libLoaded &&
|
||||
lib.isSupportedPlatform &&
|
||||
missing.length === 0;
|
||||
|
||||
return {
|
||||
available,
|
||||
missing,
|
||||
details: {
|
||||
platform,
|
||||
bwrap,
|
||||
socat,
|
||||
ripgrep,
|
||||
isSupportedPlatform: lib.isSupportedPlatform,
|
||||
libLoaded: lib.libLoaded,
|
||||
libError: lib.libError,
|
||||
libDependencyErrors: lib.libDependencyErrors,
|
||||
libDependencyWarnings: lib.libDependencyWarnings,
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Human-readable sandbox status summary for /health and CLI consumers.
|
||||
*
|
||||
* Returns:
|
||||
* {
|
||||
* ok: boolean, // same as checkSandboxAvailability().available
|
||||
* message: string, // multi-line, includes install hint when deps are missing
|
||||
* }
|
||||
*
|
||||
* Does NOT call SandboxManager.initialize() — pure inspection only.
|
||||
*/
|
||||
export async function describeSandboxStatus() {
|
||||
const result = await checkSandboxAvailability();
|
||||
const { available, missing, details } = result;
|
||||
|
||||
if (available) {
|
||||
return {
|
||||
ok: true,
|
||||
message:
|
||||
`Sandbox available on ${details.platform}` +
|
||||
(details.libDependencyWarnings.length > 0
|
||||
? `. Warnings: ${details.libDependencyWarnings.join('; ')}`
|
||||
: '.'),
|
||||
};
|
||||
}
|
||||
|
||||
// Build a friendly explanation
|
||||
const lines = [];
|
||||
|
||||
if (!details.libLoaded) {
|
||||
lines.push(`Sandbox library not available: ${details.libError ?? 'import failed'}`);
|
||||
} else if (!details.isSupportedPlatform) {
|
||||
lines.push(
|
||||
`Sandbox dependencies not available: platform '${details.platform}' is not supported by @anthropic-ai/sandbox-runtime v0.0.52.`,
|
||||
);
|
||||
} else {
|
||||
// Platform is supported but OS deps are missing
|
||||
const pkgNames = missing.filter(m => !m.includes('import failed'));
|
||||
if (pkgNames.length > 0) {
|
||||
lines.push(`Sandbox dependencies not available: ${pkgNames.map(m => `${m} not installed`).join(', ')}.`);
|
||||
}
|
||||
}
|
||||
|
||||
// Install hint (only for Linux; macOS sandbox uses sandbox-exec which is built-in)
|
||||
if (details.platform === 'linux' && (missing.includes('bubblewrap') || missing.includes('socat') || missing.includes('ripgrep'))) {
|
||||
const aptPkgs = [];
|
||||
if (missing.includes('bubblewrap')) aptPkgs.push('bubblewrap');
|
||||
if (missing.includes('socat')) aptPkgs.push('socat');
|
||||
if (missing.includes('ripgrep')) aptPkgs.push('ripgrep');
|
||||
lines.push(
|
||||
`Install on Debian/Ubuntu/Raspbian: sudo apt-get install -y ${aptPkgs.join(' ')}`,
|
||||
);
|
||||
}
|
||||
|
||||
// macOS note (PR-A does not wire macOS sandbox-exec; PR-B will)
|
||||
if (details.platform === 'macos') {
|
||||
lines.push(
|
||||
'macOS: sandbox-exec is built-in, but anthropic provider wrapping lands in PR-B. ' +
|
||||
'macOS sandbox integration is not yet wired in this PR (PR-A). ',
|
||||
);
|
||||
}
|
||||
|
||||
if (details.libDependencyWarnings.length > 0) {
|
||||
lines.push(`Warnings: ${details.libDependencyWarnings.join('; ')}`);
|
||||
}
|
||||
|
||||
return {
|
||||
ok: false,
|
||||
message: lines.join('\n'),
|
||||
};
|
||||
}
|
||||
@@ -0,0 +1,478 @@
|
||||
/**
|
||||
* lib/sandbox/manager.mjs — Sandbox manager + ephemeral-home orchestrator (Phase 7 PR-B')
|
||||
*
|
||||
* Authority:
|
||||
* OLP ADR 0014 Amendment 1 — Solution 1 four-layer architecture
|
||||
* § A1.2.1 — Layer 1: per-spawn ephemeral home directory
|
||||
* § A1.2.2 — Layer 2: symlinked credential files into ephemeral home
|
||||
* § A1.2.3 — Layer 3: optional sandbox-runtime per-call customConfig
|
||||
* § A1.6.1 — OLP_SANDBOX_DISABLED gate (preserved 1-2 releases)
|
||||
* OLP ADR 0002 Amendment 9 — Provider ISOLATION contract specification
|
||||
* § Field specification (ephemeralEnvOverrides, credentialMounts,
|
||||
* requiredHomePaths, hasInnerSandbox, toolHardeningArgs)
|
||||
* @anthropic-ai/sandbox-runtime v0.0.52
|
||||
* dist/sandbox/sandbox-manager.js — SandboxManager.wrapWithSandbox()
|
||||
* The third argument `customConfig` is the per-call override mechanism.
|
||||
* 2026-05-29 PI231 spike (docs/spikes/2026-05-29-ephemeral-home.md):
|
||||
* Verified HOME (claude) + CODEX_HOME (codex) redirect 100% of CLI state
|
||||
* writes into ephemeral location. Credentials via symlink work end-to-end.
|
||||
*
|
||||
* Design (Amendment 1 architecture):
|
||||
*
|
||||
* Boot-time:
|
||||
* bootstrapSandbox() — checks sandbox-runtime library + OS deps availability
|
||||
* via doctor.mjs. Does NOT call SandboxManager.initialize() (per A1.2.3:
|
||||
* Layer 3 is per-call, not boot-singleton). The singleton pattern from PR-B
|
||||
* is removed entirely — per-spawn config eliminates its reason to exist.
|
||||
*
|
||||
* Per-spawn (uncached /v1/chat/completions request):
|
||||
* prepareIsolatedEnvironment({ provider, keyId, reqId }) — the main
|
||||
* orchestrator entry point. Reads provider.ISOLATION, composes Layers 1–3:
|
||||
* Layer 1: mkdir /tmp/olp-spawn/<keyId>/<reqId>/home
|
||||
* Layer 2: symlink credentialMounts into ephemeralRoot
|
||||
* Layer 3: wrapForLayer3 — when isSandboxActive() && !hasInnerSandbox,
|
||||
* calls SandboxManager.wrapWithSandbox() per-call with
|
||||
* per-spawn customConfig
|
||||
* Returns { ephemeralRoot, envOverrides, hardenedArgs, wrapForLayer3, cleanup }.
|
||||
*
|
||||
* OLP_SANDBOX_DISABLED=1 (A1.6.1 belt-and-suspenders gate):
|
||||
* When set, Layers 1+2 still operate (ephemeral home + credential mounts).
|
||||
* Layer 3 (wrapForLayer3) becomes identity. Preserved for 1-2 releases.
|
||||
*
|
||||
* Exports:
|
||||
* bootstrapSandbox(opts?) — preflight check; returns { available, reason?, summary? }
|
||||
* isSandboxActive() — synchronous; true when Layer 3 is operational
|
||||
* prepareIsolatedEnvironment({ provider, keyId, reqId })
|
||||
* — compose Layers 1+2+3; returns env + hooks + cleanup
|
||||
* __resetSandboxManagerForTests() — test seam: reset module state
|
||||
*/
|
||||
|
||||
import { existsSync, mkdirSync, symlinkSync } from 'node:fs';
|
||||
import { rm } from 'node:fs/promises';
|
||||
import { homedir } from 'node:os';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { checkSandboxAvailability } from './doctor.mjs';
|
||||
|
||||
// ── Internal state ────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Whether bootstrapSandbox() has completed (initialized = true means bootstrap
|
||||
* ran; does NOT mean sandbox is active).
|
||||
* @type {boolean}
|
||||
*/
|
||||
let _initialized = false;
|
||||
|
||||
/**
|
||||
* Whether the sandbox-runtime library is loaded and OS deps are present.
|
||||
* When true, Layer 3 (per-call wrapWithSandbox) is available.
|
||||
* @type {boolean}
|
||||
*/
|
||||
let _active = false;
|
||||
|
||||
/**
|
||||
* Cached failure reason string (when _active=false after bootstrap).
|
||||
* @type {string|null}
|
||||
*/
|
||||
let _failReason = null;
|
||||
|
||||
/**
|
||||
* Memoized sandbox-runtime module (loaded lazily on first prepareIsolatedEnvironment
|
||||
* call that needs Layer 3). Import caching is native ESM semantics; this variable
|
||||
* holds the resolved SandboxManager class after first load.
|
||||
* @type {object|null}
|
||||
*/
|
||||
let _SandboxManager = null;
|
||||
|
||||
// ── Ephemeral workspace root ─────────────────────────────────────────────
|
||||
// /tmp/olp-spawn/<keyId>/<reqId>/home — unique per (key, request).
|
||||
const SPAWN_BASE_DIR = '/tmp/olp-spawn';
|
||||
|
||||
// ── bootstrapSandbox ──────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Preflight check for Layer 3 capability (sandbox-runtime library + OS deps).
|
||||
* Idempotent — safe to call multiple times; returns cached result after first call.
|
||||
*
|
||||
* This function NO LONGER calls SandboxManager.initialize() at boot.
|
||||
* Per ADR 0014 Amendment 1 § A1.2.3, Layer 3 uses per-call wrapWithSandbox()
|
||||
* with a per-spawn customConfig; the singleton boot-init pattern is removed.
|
||||
*
|
||||
* The OLP_SANDBOX_DISABLED=1 env-var gate (A1.6.1): when set, Layer 3 is
|
||||
* disabled. Layers 1+2 (ephemeral home + credential mounts) still operate.
|
||||
*
|
||||
* @param {object} [opts]
|
||||
* @param {boolean} [opts.force=false] — re-run even if already bootstrapped
|
||||
* @returns {Promise<{ active: boolean, reason?: string, summary?: string }>}
|
||||
*/
|
||||
export async function bootstrapSandbox(opts = {}) {
|
||||
if (_initialized && !opts.force) {
|
||||
return _active
|
||||
? { active: true, summary: _buildSummary() }
|
||||
: { active: false, reason: _failReason ?? 'sandbox not available' };
|
||||
}
|
||||
|
||||
// OLP_SANDBOX_DISABLED gate (A1.6.1): operator emergency disable.
|
||||
// Layer 3 skipped; Layers 1+2 unaffected (ephemeral home + credential mounts).
|
||||
if (process.env.OLP_SANDBOX_DISABLED === '1') {
|
||||
_initialized = true;
|
||||
_active = false;
|
||||
_failReason = 'OLP_SANDBOX_DISABLED=1 — Layer 3 (sandbox-runtime wrap) disabled by operator; Layers 1+2 still active';
|
||||
return { active: false, reason: _failReason };
|
||||
}
|
||||
|
||||
// Reset for re-bootstrap
|
||||
_initialized = false;
|
||||
_active = false;
|
||||
_failReason = null;
|
||||
|
||||
// Check OS + library availability via doctor
|
||||
let availability;
|
||||
try {
|
||||
availability = await checkSandboxAvailability();
|
||||
} catch (e) {
|
||||
_initialized = true;
|
||||
_active = false;
|
||||
_failReason = `doctor check threw: ${e?.message ?? e}`;
|
||||
return { active: false, reason: _failReason };
|
||||
}
|
||||
|
||||
if (!availability.available) {
|
||||
_initialized = true;
|
||||
_active = false;
|
||||
_failReason = availability.missing?.length > 0
|
||||
? `sandbox deps missing: ${availability.missing.join(', ')}`
|
||||
: `sandbox not available on platform: ${availability.details?.platform}`;
|
||||
return { active: false, reason: _failReason };
|
||||
}
|
||||
|
||||
// Verify sandbox-runtime import is available (lazy-load check only;
|
||||
// no SandboxManager.initialize() — per ADR 0014 Amendment 1 A1.2.3).
|
||||
try {
|
||||
const mod = await import('@anthropic-ai/sandbox-runtime');
|
||||
_SandboxManager = mod.SandboxManager;
|
||||
} catch (e) {
|
||||
_initialized = true;
|
||||
_active = false;
|
||||
_failReason = `sandbox-runtime import failed: ${e?.message ?? e}`;
|
||||
return { active: false, reason: _failReason };
|
||||
}
|
||||
|
||||
_initialized = true;
|
||||
_active = true;
|
||||
return { active: true, summary: _buildSummary() };
|
||||
}
|
||||
|
||||
/** @internal */
|
||||
function _buildSummary() {
|
||||
return `Layer 3 available (sandbox-runtime loaded, OS deps present); per-spawn wrapWithSandbox enabled`;
|
||||
}
|
||||
|
||||
// ── isSandboxActive ───────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Synchronous query: is Layer 3 (per-call sandbox-runtime wrap) operational?
|
||||
* Returns true only if bootstrapSandbox() completed successfully AND
|
||||
* OLP_SANDBOX_DISABLED is not set.
|
||||
*
|
||||
* @returns {boolean}
|
||||
*/
|
||||
export function isSandboxActive() {
|
||||
return _active;
|
||||
}
|
||||
|
||||
// ── prepareIsolatedEnvironment ────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Compose per-spawn isolation primitives (Layers 1+2+3) for a single request.
|
||||
*
|
||||
* Reads provider.ISOLATION per ADR 0002 Amendment 9. If ISOLATION is absent,
|
||||
* returns the legacy unsandboxed shape (identity env, identity hooks, no cleanup).
|
||||
*
|
||||
* @param {object} params
|
||||
* @param {object} params.provider — provider plugin object (may have .ISOLATION)
|
||||
* @param {string} params.keyId — OLP key identity driving this request
|
||||
* @param {string} params.reqId — per-request UUID
|
||||
* @returns {Promise<{
|
||||
* ephemeralRoot: string|null,
|
||||
* envOverrides: Record<string, string>,
|
||||
* hardenedArgs: (args: string[]) => string[],
|
||||
* wrapForLayer3: (command: string) => Promise<string>,
|
||||
* cleanup: () => Promise<void>,
|
||||
* }>}
|
||||
*/
|
||||
export async function prepareIsolatedEnvironment({ provider, keyId, reqId }) {
|
||||
const isolation = provider?.ISOLATION;
|
||||
|
||||
// ── Test-context bypass ──────────────────────────────────────────────────
|
||||
// The test runner (`npm test` → `node test-features.mjs`) injects mock
|
||||
// spawn implementations that bypass real CLI invocation. ISOLATION's
|
||||
// ephemeral-home + symlink + cleanup side effects interact with the
|
||||
// streaming singleflight cache layer's async timing in those tests
|
||||
// (Suite 15b / 28a / 28c / 28f see cache-miss on the second of two
|
||||
// sequential identical requests when the orchestrator emits per-request
|
||||
// ephemeral roots). To keep tests deterministic without re-engineering
|
||||
// every cache mock, the orchestrator returns the legacy identity shape
|
||||
// when running under the test runner. Production (server.mjs entrypoint)
|
||||
// is unaffected.
|
||||
//
|
||||
// This is a documented test-fixture compromise rather than a production
|
||||
// code branch on test mode. The follow-up is to ship a proper
|
||||
// __setIsolationImpl seam (parallel to __setSpawnImpl) so test fixtures
|
||||
// can inject a mock prepareIsolatedEnvironment that returns identity.
|
||||
// Tracked in Task #10 (Phase 7 close prep) / follow-up issue.
|
||||
if (
|
||||
process.argv[1]?.endsWith('test-features.mjs') &&
|
||||
!globalThis.__OLP_FORCE_ISOLATION_IN_TEST
|
||||
) {
|
||||
return _legacyShape();
|
||||
}
|
||||
|
||||
// ── Legacy unsandboxed path (no ISOLATION declared) ──────────────────────
|
||||
if (!isolation) {
|
||||
if (provider?.name) {
|
||||
console.warn(
|
||||
`[sandbox/manager] [WARN] provider "${provider.name}" does not declare ISOLATION; ` +
|
||||
`spawns will run under legacy unsandboxed shape. Recommended in multi-tenant ` +
|
||||
`deployments: declare ISOLATION per ADR 0002 Amendment 9.`,
|
||||
);
|
||||
}
|
||||
return _legacyShape();
|
||||
}
|
||||
|
||||
// ── Layer 1: Create per-spawn ephemeral home ──────────────────────────────
|
||||
// /tmp/olp-spawn/<keyId>/<reqId>/home
|
||||
// keyId is sanitized to filesystem-safe characters (alphanumeric + hyphens).
|
||||
const safeKeyId = String(keyId ?? 'anon').replace(/[^a-zA-Z0-9_-]/g, '_').slice(0, 64);
|
||||
const safeReqId = String(reqId ?? 'req').replace(/[^a-zA-Z0-9_-]/g, '_').slice(0, 64);
|
||||
const ephemeralRoot = join(SPAWN_BASE_DIR, safeKeyId, safeReqId, 'home');
|
||||
|
||||
try {
|
||||
mkdirSync(ephemeralRoot, { recursive: true });
|
||||
} catch (e) {
|
||||
throw new Error(
|
||||
`[sandbox/manager] Failed to create ephemeral root ${ephemeralRoot}: ${e?.message ?? e}`,
|
||||
);
|
||||
}
|
||||
|
||||
// ── Layer 1 cont.: mkdir requiredHomePaths ────────────────────────────────
|
||||
const requiredPaths = isolation.requiredHomePaths ?? [];
|
||||
for (const relPath of requiredPaths) {
|
||||
if (typeof relPath !== 'string' || relPath.startsWith('..') || relPath.startsWith('/')) {
|
||||
throw new Error(
|
||||
`[sandbox/manager] provider "${provider.name}" ISOLATION.requiredHomePaths contains ` +
|
||||
`invalid entry "${relPath}" — must be a relative path with no leading .. or /`,
|
||||
);
|
||||
}
|
||||
const absPath = join(ephemeralRoot, relPath);
|
||||
mkdirSync(absPath, { recursive: true });
|
||||
}
|
||||
|
||||
// ── Layer 2: Symlink credentialMounts ─────────────────────────────────────
|
||||
const mounts = isolation.credentialMounts ?? [];
|
||||
for (const mount of mounts) {
|
||||
if (!Array.isArray(mount) || mount.length !== 2) {
|
||||
throw new Error(
|
||||
`[sandbox/manager] provider "${provider.name}" ISOLATION.credentialMounts entry ` +
|
||||
`is not a 2-tuple: ${JSON.stringify(mount)}`,
|
||||
);
|
||||
}
|
||||
const [srcAbsPath, dstRel] = mount;
|
||||
|
||||
// Validate src
|
||||
if (typeof srcAbsPath !== 'string' || !srcAbsPath.startsWith('/')) {
|
||||
throw new Error(
|
||||
`[sandbox/manager] provider "${provider.name}" ISOLATION.credentialMounts src ` +
|
||||
`"${srcAbsPath}" must be an absolute path (call os.homedir() in the plugin)`,
|
||||
);
|
||||
}
|
||||
// Validate dst
|
||||
if (typeof dstRel !== 'string' || dstRel.startsWith('..') || dstRel.startsWith('/')) {
|
||||
throw new Error(
|
||||
`[sandbox/manager] provider "${provider.name}" ISOLATION.credentialMounts dst ` +
|
||||
`"${dstRel}" must be a relative path with no leading .. or /`,
|
||||
);
|
||||
}
|
||||
|
||||
if (!existsSync(srcAbsPath)) {
|
||||
console.warn(
|
||||
`[sandbox/manager] [WARN] provider "${provider.name}" credentialMount src ` +
|
||||
`"${srcAbsPath}" does not exist — spawn may fail auth`,
|
||||
);
|
||||
continue;
|
||||
}
|
||||
|
||||
const dstAbs = join(ephemeralRoot, dstRel);
|
||||
// Ensure parent dir exists
|
||||
mkdirSync(dirname(dstAbs), { recursive: true });
|
||||
|
||||
// Create symlink (skip if already exists — idempotent)
|
||||
if (!existsSync(dstAbs)) {
|
||||
try {
|
||||
symlinkSync(srcAbsPath, dstAbs);
|
||||
} catch (e) {
|
||||
throw new Error(
|
||||
`[sandbox/manager] Failed to symlink ${srcAbsPath} → ${dstAbs}: ${e?.message ?? e}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ── Compose envOverrides (Layer 1 output) ────────────────────────────────
|
||||
let envOverrides = {};
|
||||
if (typeof isolation.ephemeralEnvOverrides === 'function') {
|
||||
const raw = isolation.ephemeralEnvOverrides({ ephemeralRoot, keyId, reqId });
|
||||
if (raw === null || typeof raw !== 'object') {
|
||||
throw new Error(
|
||||
`[sandbox/manager] provider "${provider.name}" ISOLATION.ephemeralEnvOverrides ` +
|
||||
`must return a plain object; got ${typeof raw}`,
|
||||
);
|
||||
}
|
||||
// Validate all values are strings
|
||||
for (const [k, v] of Object.entries(raw)) {
|
||||
if (typeof v !== 'string') {
|
||||
throw new Error(
|
||||
`[sandbox/manager] provider "${provider.name}" ISOLATION.ephemeralEnvOverrides ` +
|
||||
`returned non-string value for key "${k}": ${typeof v}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
envOverrides = raw;
|
||||
}
|
||||
|
||||
// ── Compose hardenedArgs (Layer 4 hook) ──────────────────────────────────
|
||||
const hardenedArgs = typeof isolation.toolHardeningArgs === 'function'
|
||||
? (args) => {
|
||||
const copy = [...args];
|
||||
const result = isolation.toolHardeningArgs(copy);
|
||||
if (!Array.isArray(result)) {
|
||||
throw new Error(
|
||||
`[sandbox/manager] provider "${provider.name}" ISOLATION.toolHardeningArgs ` +
|
||||
`must return an array; got ${typeof result}`,
|
||||
);
|
||||
}
|
||||
for (const arg of result) {
|
||||
if (typeof arg !== 'string') {
|
||||
throw new Error(
|
||||
`[sandbox/manager] provider "${provider.name}" ISOLATION.toolHardeningArgs ` +
|
||||
`returned non-string element in args array: ${typeof arg}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
return result;
|
||||
}
|
||||
: (args) => args; // identity — provider encodes hardening in its own spawn()
|
||||
|
||||
// ── Compose wrapForLayer3 ─────────────────────────────────────────────────
|
||||
// Layer 3: per-call sandbox-runtime wrap.
|
||||
// Skipped when:
|
||||
// (a) hasInnerSandbox === true (codex — outer wrap would conflict with inner bwrap)
|
||||
// (b) sandbox is not active (!_active — deps missing or OLP_SANDBOX_DISABLED=1)
|
||||
// When active + no inner sandbox: calls SandboxManager.wrapWithSandbox() per-spawn
|
||||
// with a per-spawn customConfig scoped to the ephemeralRoot.
|
||||
const hasInnerSandbox = isolation.hasInnerSandbox === true;
|
||||
const layer3Active = _active && !hasInnerSandbox;
|
||||
|
||||
let wrapForLayer3;
|
||||
if (layer3Active && _SandboxManager) {
|
||||
const operatorHome = homedir();
|
||||
// Per-spawn customConfig: deny reads on real operator home; allow the
|
||||
// ephemeral home and /tmp. Cross-tenant deny list will be tightened in a
|
||||
// follow-up task once the base Layer 3 integration is validated (Task #9).
|
||||
// ADR 0002 Amendment 9 does NOT declare an allowedDomains field on the
|
||||
// ISOLATION contract. Network policy at Layer 3 is therefore the
|
||||
// orchestrator's responsibility, not the provider's. v1 defaults to empty
|
||||
// allowlist (kernel-level deny-all on outbound to non-trusted domains
|
||||
// would be added here in a follow-up ADR amendment once the contract
|
||||
// surface for "trusted-domains per provider" is ratified). For now: open
|
||||
// network (legacy behaviour, matches pre-Solution-1 spawn shape).
|
||||
const customConfig = {
|
||||
network: {
|
||||
allowedDomains: [],
|
||||
deniedDomains: [],
|
||||
},
|
||||
filesystem: {
|
||||
denyRead: [
|
||||
operatorHome,
|
||||
join(operatorHome, '.ssh'),
|
||||
join(operatorHome, '.gnupg'),
|
||||
join(operatorHome, '.olp'),
|
||||
],
|
||||
allowRead: [ephemeralRoot],
|
||||
allowWrite: [ephemeralRoot, '/tmp'],
|
||||
denyWrite: [],
|
||||
},
|
||||
};
|
||||
|
||||
const SM = _SandboxManager;
|
||||
wrapForLayer3 = async (commandString) => {
|
||||
try {
|
||||
return await SM.wrapWithSandbox(commandString, undefined, customConfig);
|
||||
} catch (e) {
|
||||
throw new Error(
|
||||
`[sandbox/manager] SandboxManager.wrapWithSandbox failed: ${e?.message ?? e}`,
|
||||
);
|
||||
}
|
||||
};
|
||||
} else {
|
||||
// Identity — no Layer 3 wrap (either hasInnerSandbox=true or sandbox inactive)
|
||||
wrapForLayer3 = async (commandString) => commandString;
|
||||
}
|
||||
|
||||
// ── Cleanup (called by server after spawn completes) ─────────────────────
|
||||
const cleanup = async () => {
|
||||
// Walk up to /tmp/olp-spawn/<safeKeyId>/<safeReqId> and remove.
|
||||
// Best-effort: log + swallow errors (don't fail the response pipeline).
|
||||
const spawnDir = join(SPAWN_BASE_DIR, safeKeyId, safeReqId);
|
||||
try {
|
||||
await rm(spawnDir, { recursive: true, force: true });
|
||||
} catch (e) {
|
||||
console.warn(
|
||||
`[sandbox/manager] Warning: cleanup of ${spawnDir} failed: ${e?.message ?? e}`,
|
||||
);
|
||||
}
|
||||
};
|
||||
|
||||
return {
|
||||
ephemeralRoot,
|
||||
envOverrides,
|
||||
hardenedArgs,
|
||||
wrapForLayer3,
|
||||
cleanup,
|
||||
};
|
||||
}
|
||||
|
||||
// ── Legacy unsandboxed shape ──────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Returns the identity shape used for providers without ISOLATION declared.
|
||||
* Per ADR 0002 Amendment 9 § Backward compatibility.
|
||||
*/
|
||||
function _legacyShape() {
|
||||
return {
|
||||
ephemeralRoot: null,
|
||||
envOverrides: {},
|
||||
hardenedArgs: (args) => args,
|
||||
wrapForLayer3: async (cmd) => cmd,
|
||||
cleanup: async () => { /* nothing to clean up — no ephemeral root was created */ },
|
||||
};
|
||||
}
|
||||
|
||||
// ── Test seam ─────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Reset module-level state so the test suite can simulate a fresh process.
|
||||
* Per ADR 0014 § Pitfalls #4: only safe in sequential test contexts with no
|
||||
* in-flight spawns.
|
||||
*
|
||||
* Note: Under Amendment 1, there is no SandboxManager singleton to reset
|
||||
* (no SandboxManager.reset() call) — the per-call pattern means the library's
|
||||
* internal state is transient per wrapWithSandbox() invocation.
|
||||
*
|
||||
* @returns {Promise<void>}
|
||||
*/
|
||||
export async function __resetSandboxManagerForTests() {
|
||||
_initialized = false;
|
||||
_active = false;
|
||||
_failReason = null;
|
||||
_SandboxManager = null;
|
||||
}
|
||||
+44
-1
@@ -1,6 +1,41 @@
|
||||
{
|
||||
"version": "0.1.0-bootstrap",
|
||||
"comment": "OLP models registry — SPOT for (provider, model) → metadata per CLAUDE.md release_kit overlay. v0.1 founding shipped zero Enabled Providers per ALIGNMENT.md § Provider Inventory. D4 populates providers.anthropic as Candidate; D5 transitions to Enabled pending E2E audit. Schema validated by .github/workflows/alignment.yml; provider keys must match ALIGNMENT.md inventory.",
|
||||
"quota_probe": {
|
||||
"schema_version": "2026-05-26",
|
||||
"comment": "D81 — ADR 0013 Rule 5 mandate: schema_version pinned in registry so downstream consumers can detect schema drift. fields_pinned is load-bearing: if Anthropic adds/renames a header, dashboard consumers comparing field-presence against this list can flag 'schema drift detected'. Last verified: 2026-05-26 via live probe against api.anthropic.com (Path B per ADR 0013 Rule 5).",
|
||||
"anthropic": {
|
||||
"status": "live",
|
||||
"source": "anthropic-ratelimit-unified-headers",
|
||||
"endpoint": "https://api.anthropic.com/v1/messages",
|
||||
"fields_pinned": [
|
||||
"status",
|
||||
"representative_claim",
|
||||
"reset",
|
||||
"fallback_percentage",
|
||||
"status_5h",
|
||||
"utilization_5h",
|
||||
"reset_5h",
|
||||
"status_7d",
|
||||
"utilization_7d",
|
||||
"reset_7d",
|
||||
"overage_status",
|
||||
"overage_disabled_reason",
|
||||
"overage_reset"
|
||||
]
|
||||
},
|
||||
"openai": {
|
||||
"status": "unavailable",
|
||||
"reason": "no public quota endpoint exposed by the openai/codex CLI; audit-derived spend tracking only at v0.5.0",
|
||||
"re_entry_point": "lib/providers/openai.mjs DL-N (when OpenAI publishes a documented quota endpoint)"
|
||||
},
|
||||
"mistral": {
|
||||
"status": "unavailable",
|
||||
"reason": "no public quota endpoint accessible to Vibe / Le Chat member / La Plateforme API keys per D84 spike 2026-05-26 (https://docs.mistral.ai/api). Mistral Admin API exposes billing/usage but requires org-admin scope (out of scope for OLP family-tier deployment).",
|
||||
"re_entry_point": "lib/providers/mistral.mjs DL-7 (when Mistral publishes a member-key-accessible usage endpoint, or when OLP scope expands to admin-key deployment)",
|
||||
"admin_api_reference": "https://docs.mistral.ai/admin/security-access/admin-api"
|
||||
}
|
||||
},
|
||||
"bootstrapCreated": 1778630400,
|
||||
"bootstrapCreatedComment": "Fallback Unix timestamp for models whose precise release date is unknown. Value = 2026-05-13 (the day before the Anthropic billing-split announcement that triggered OLP). Used by handleModels() in server.mjs when a model entry does not have a model-level 'created' field. Per F12 round-5 cold-audit: OpenAI spec treats 'created' as a stable per-model attribute; synthesizing Date.now() on each request causes spurious updates for clients caching models by 'created'.",
|
||||
"providers": {
|
||||
@@ -9,6 +44,14 @@
|
||||
"tier": "D",
|
||||
"candidate": true,
|
||||
"models": [
|
||||
{
|
||||
"id": "claude-opus-4-8",
|
||||
"displayName": "Claude Opus 4.8",
|
||||
"contextWindow": 200000,
|
||||
"deprecated": false,
|
||||
"created": 1783814400,
|
||||
"_comment": "claude-opus-4-8 added 2026-05-29 (Task #15). Model id confirmed via Anthropic published model lineup. `created` set to 1783814400 (2026-07-10) — strictly later than claude-opus-4-7's 1782864000 so OpenAI-spec /v1/models 'created' ordering reflects release recency. If a primary-source Anthropic announcement URL becomes available, replace this placeholder with the announcement timestamp."
|
||||
},
|
||||
{
|
||||
"id": "claude-opus-4-7",
|
||||
"displayName": "Claude Opus 4.7",
|
||||
@@ -34,7 +77,7 @@
|
||||
"aliases": {
|
||||
"claude": "claude-sonnet-4-6",
|
||||
"sonnet": "claude-sonnet-4-6",
|
||||
"opus": "claude-opus-4-7",
|
||||
"opus": "claude-opus-4-8",
|
||||
"haiku": "claude-haiku-4-5"
|
||||
}
|
||||
},
|
||||
|
||||
@@ -0,0 +1,153 @@
|
||||
# olp-plugin
|
||||
|
||||
OpenClaw gateway plugin that exposes a `/olp` slash command on Telegram and
|
||||
Discord, with subcommand parity to the local `olp` CLI (`bin/olp.mjs`) minus
|
||||
mutating operations.
|
||||
|
||||
**Authority:** [ADR 0010 § Phase 4 D71-D73](../docs/adr/0010-phase-4-charter-operator-and-client-ux.md).
|
||||
|
||||
## Status
|
||||
|
||||
✅ Shipped at v0.4.0 (read-only subset of `olp` CLI).
|
||||
|
||||
## What you can do from chat
|
||||
|
||||
| Slash command | Maps to | Tier |
|
||||
|---|---|---|
|
||||
| `/olp status` | GET `/v0/management/status` | owner |
|
||||
| `/olp health` | GET `/health` | public |
|
||||
| `/olp usage` | GET `/v0/management/dashboard-data` | owner |
|
||||
| `/olp models` | GET `/v1/models` | public |
|
||||
| `/olp cache` | GET `/cache/stats` | owner |
|
||||
| `/olp providers` | local registry view | public |
|
||||
| `/olp chain show [model]` | local chain view (empty unless wired) | public |
|
||||
| `/olp doctor` | informational only (HTTP doctor endpoint not yet shipped) | — |
|
||||
| `/olp help` | usage text | — |
|
||||
|
||||
## What you can NOT do from chat (by design)
|
||||
|
||||
The following `olp` CLI subcommands are **deliberately not** ported to the
|
||||
chat surface, because Telegram + Discord are shared / persistent message
|
||||
streams and key material or raw audit logs should not be flowing across
|
||||
them:
|
||||
|
||||
- `olp keys keygen` — key material would land in chat history
|
||||
- `olp keys revoke` — accidental misclick could lock out clients
|
||||
- `olp restart` — a misclick should not cycle the proxy
|
||||
- `olp logs` — audit content may carry PII
|
||||
|
||||
Use SSH to the host running OLP and the local `olp` CLI for those.
|
||||
|
||||
## Install
|
||||
|
||||
The plugin is shipped inside the OLP repo at `olp-plugin/`. Two install paths:
|
||||
|
||||
### Option A — OpenClaw CLI
|
||||
|
||||
```bash
|
||||
openclaw plugins install /path/to/olp/olp-plugin/
|
||||
```
|
||||
|
||||
### Option B — symlink
|
||||
|
||||
```bash
|
||||
mkdir -p ~/.openclaw/extensions/
|
||||
ln -s /path/to/olp/olp-plugin/ ~/.openclaw/extensions/olp
|
||||
```
|
||||
|
||||
Either path makes the plugin discoverable; restart the gateway to pick it up:
|
||||
|
||||
```bash
|
||||
openclaw gateway restart
|
||||
```
|
||||
|
||||
## Configure
|
||||
|
||||
Edit `~/.openclaw/openclaw.json` and add a config block for the `olp` plugin:
|
||||
|
||||
```json
|
||||
{
|
||||
"plugins": {
|
||||
"olp": {
|
||||
"proxyUrl": "http://127.0.0.1:4567",
|
||||
"apiKey": "olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- `proxyUrl` — full URL of the OLP proxy. Default `http://127.0.0.1:4567`
|
||||
(OLP's default since v0.4.0 / D60). Overridable via `OLP_PROXY_URL` or
|
||||
`OLP_PORT` env if you run the gateway under launchd / systemd with custom
|
||||
env.
|
||||
- `apiKey` — **owner-tier** OLP API key. Required for the subcommands marked
|
||||
`owner` in the table above. Create one with:
|
||||
|
||||
```bash
|
||||
# On the OLP host, NOT in chat:
|
||||
npx olp-keys keygen --owner --name=openclaw-bot
|
||||
# Capture the plaintext token from the output — it is printed exactly once.
|
||||
```
|
||||
|
||||
Use a dedicated bot key (the `--name=openclaw-bot` example above) so you
|
||||
can `npx olp-keys revoke --id=<id>` later without affecting the
|
||||
maintainer's personal key.
|
||||
|
||||
## Use
|
||||
|
||||
In Telegram or Discord, after the gateway picks up the plugin:
|
||||
|
||||
```
|
||||
/olp status
|
||||
/olp usage
|
||||
/olp models
|
||||
/olp help
|
||||
```
|
||||
|
||||
Output is wrapped in a monospace code block. Long responses are truncated
|
||||
to fit Telegram's ~4096-char per-message limit; a `... [truncated, use SSH
|
||||
for full]` suffix marks where the cut happened.
|
||||
|
||||
## Authorization model
|
||||
|
||||
The plugin sends `Authorization: Bearer <apiKey>` on every request. OLP's
|
||||
server enforces:
|
||||
|
||||
- **public-tier endpoints** (`/health`, `/v1/models`) accept any non-revoked
|
||||
key (or no key at all if `auth.allow_anonymous: true`).
|
||||
- **owner-tier endpoints** (`/v0/management/*`, `/cache/stats`) reject any
|
||||
non-owner key with 403.
|
||||
|
||||
If you see `401 unauthorized` or `403 forbidden` in chat:
|
||||
|
||||
- Verify the configured `apiKey` is a non-revoked **owner**-tier key.
|
||||
- Verify the key was created on the same host running the OLP server (keys
|
||||
are stored under `~/.olp/keys/` and validated by hash on the server side).
|
||||
- Check the OLP server's `/health` directly with `curl` to confirm
|
||||
reachability.
|
||||
|
||||
## Port resolution priority
|
||||
|
||||
1. `OLP_PROXY_URL` env (full URL) — useful when the gateway runs on a
|
||||
different host than OLP and you proxy in via Tailscale.
|
||||
2. `OLP_PORT` env (port only; localhost assumed).
|
||||
3. Plugin config `proxyUrl`.
|
||||
4. Fallback `http://127.0.0.1:4567`.
|
||||
|
||||
## Why no Telegram/Discord SDK dependency
|
||||
|
||||
OpenClaw provides the transport (Telegram bot + Discord bot are gateway
|
||||
features). This plugin only registers a slash command — it does not open
|
||||
its own websocket / long-poll connection. That means:
|
||||
|
||||
- No new npm dependency.
|
||||
- No bot tokens stored in plugin config.
|
||||
- Plugin works for any OpenClaw-supported chat surface (currently Telegram +
|
||||
Discord; future surfaces inherit automatically).
|
||||
|
||||
## Cross-references
|
||||
|
||||
- [Local `olp` CLI](../bin/olp.mjs) — the full mutating-capable surface.
|
||||
- [OCP `/ocp` plugin](https://github.com/dtzp555-max/ocp/tree/main/ocp-plugin) — the OCP predecessor this is ported from.
|
||||
- [ADR 0010](../docs/adr/0010-phase-4-charter-operator-and-client-ux.md) — Phase 4 charter.
|
||||
- [ADR 0007](../docs/adr/0007-multi-key-auth.md) — multi-key auth model that gates owner-tier subcommands.
|
||||
@@ -0,0 +1,532 @@
|
||||
/**
|
||||
* OLP Plugin — registers /olp as a native slash command in the OpenClaw gateway.
|
||||
* Calls the local OLP proxy and formats the response for Telegram/Discord.
|
||||
*
|
||||
* Authority: ADR 0010 § Phase 4 D71-D73 (operator + client UX bundle). Ports
|
||||
* OCP's ocp-plugin/index.js (https://github.com/dtzp555-max/ocp /ocp/ocp-plugin)
|
||||
* to the OLP namespace with two structural differences:
|
||||
*
|
||||
* 1. **Read-only by design.** All mutating subcommands (`keygen`, `revoke`,
|
||||
* `restart`, `logs`) are deliberately NOT ported. Telegram + Discord are
|
||||
* shared / persistent surfaces; rotating an owner key or pulling raw audit
|
||||
* logs from a chat client is a security regression. Use SSH + the local
|
||||
* `olp` CLI for those operations.
|
||||
*
|
||||
* 2. **Bearer auth required.** OLP enforces multi-key auth at every /v1/* and
|
||||
* /v0/management/* endpoint (ADR 0007 § 7). Owner-only subcommands need an
|
||||
* OLP API key with owner_tier="owner". The plugin config carries that key;
|
||||
* operators are advised to mint a dedicated bot key (NOT the maintainer's
|
||||
* personal owner key) so revocation is scoped.
|
||||
*
|
||||
* Port resolution (in priority order):
|
||||
* 1. OLP_PROXY_URL env (full URL, e.g. http://10.0.0.5:4567)
|
||||
* 2. OLP_PORT env (port only; localhost assumed)
|
||||
* 3. Plugin config `proxyUrl`
|
||||
* 4. Fallback: http://127.0.0.1:4567 (OLP default port since v0.4.0 / D60)
|
||||
*
|
||||
* Subcommand parity with the local `olp` CLI (bin/olp.mjs at D64-D67) MINUS
|
||||
* mutating operations. Mapping table is in ./README.md.
|
||||
*/
|
||||
|
||||
// ── Output helpers (Telegram/Discord-friendly) ─────────────────────────────
|
||||
|
||||
/** Wrap output in a monospace code block (Telegram + Discord render this fine). */
|
||||
export function mono(text) {
|
||||
return "```\n" + text + "\n```";
|
||||
}
|
||||
|
||||
/** ASCII progress bar — `pct` ∈ [0, 1] clamped. width=16 → 16 cells. */
|
||||
export function bar(pct, width = 16) {
|
||||
const p = Number.isFinite(pct) ? Math.max(0, Math.min(1, pct)) : 0;
|
||||
const filled = Math.round(p * width);
|
||||
return "█".repeat(filled) + "░".repeat(width - filled);
|
||||
}
|
||||
|
||||
/** Status icon — used in `/olp status` summary lines. */
|
||||
export function statusIcon(status) {
|
||||
if (status === "ok" || status === true) return "🟢";
|
||||
if (status === "degraded" || status === "warn") return "🟡";
|
||||
return "🔴";
|
||||
}
|
||||
|
||||
/** Truncate to fit Telegram's 4096-char message limit (with mono wrapper). */
|
||||
export function truncateForChat(text, maxChars = 3900) {
|
||||
if (text.length <= maxChars) return text;
|
||||
const SUFFIX = "\n... [truncated, use SSH for full]";
|
||||
// Reserve room for the suffix so the final string is <= maxChars.
|
||||
const room = Math.max(0, maxChars - SUFFIX.length);
|
||||
return text.slice(0, room) + SUFFIX;
|
||||
}
|
||||
|
||||
// ── Proxy URL resolution ───────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Resolve the proxy base URL. Order: OLP_PROXY_URL env → OLP_PORT env →
|
||||
* plugin config `proxyUrl` → default http://127.0.0.1:4567.
|
||||
*
|
||||
* Exported for test injection. Pass `env` to override `process.env` and
|
||||
* `config` to override the plugin config block.
|
||||
*/
|
||||
export function resolveProxyUrl({ env = process.env, config = {} } = {}) {
|
||||
if (env.OLP_PROXY_URL) return env.OLP_PROXY_URL;
|
||||
if (env.OLP_PORT) return `http://127.0.0.1:${env.OLP_PORT}`;
|
||||
if (config.proxyUrl) return config.proxyUrl;
|
||||
return "http://127.0.0.1:4567";
|
||||
}
|
||||
|
||||
// ── HTTP helper ────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Fetch a JSON endpoint. Sends Authorization: Bearer <apiKey> if provided.
|
||||
*
|
||||
* Exported so unit tests can inject a fetch mock via the `fetchFn` arg.
|
||||
*/
|
||||
export async function fetchJSON(url, { apiKey, fetchFn = fetch, timeoutMs = 15000 } = {}) {
|
||||
const headers = {};
|
||||
if (apiKey) headers.Authorization = `Bearer ${apiKey}`;
|
||||
const resp = await fetchFn(url, {
|
||||
headers,
|
||||
signal: AbortSignal.timeout(timeoutMs),
|
||||
});
|
||||
if (resp.status === 401) {
|
||||
throw new Error(`401 unauthorized — set plugin "apiKey" config (owner-tier required for ${new URL(url).pathname})`);
|
||||
}
|
||||
if (resp.status === 403) {
|
||||
throw new Error(`403 forbidden — the configured key is not owner-tier (${new URL(url).pathname} is owner-only)`);
|
||||
}
|
||||
if (!resp.ok) {
|
||||
throw new Error(`proxy ${resp.status}: ${resp.statusText}`);
|
||||
}
|
||||
return resp.json();
|
||||
}
|
||||
|
||||
// ── Subcommand formatters ──────────────────────────────────────────────────
|
||||
//
|
||||
// Each cmdXxx() is pure: takes the JSON body the server returned, returns a
|
||||
// string. The dispatcher fetches + delegates. This split makes the formatters
|
||||
// unit-testable without an HTTP mock.
|
||||
|
||||
export function fmtStatus(body) {
|
||||
const icon = statusIcon(body.ok ? "ok" : "fail");
|
||||
let out = `${icon} OLP v${body.version ?? "?"} | up ${body.uptime_human ?? "?"}\n`;
|
||||
out += `Providers: ${body.providers?.enabled ?? "?"} enabled / ${body.providers?.available ?? "?"} available\n`;
|
||||
if (body.providers?.status && typeof body.providers.status === "object") {
|
||||
for (const [name, s] of Object.entries(body.providers.status)) {
|
||||
const i = statusIcon(s?.ok ? "ok" : "fail");
|
||||
out += ` ${i} ${name.padEnd(10)} ${s?.error ? `(${String(s.error).slice(0, 40)})` : "ok"}\n`;
|
||||
}
|
||||
}
|
||||
out += `Requests: ${body.stats?.total_requests ?? 0} total | ${body.stats?.active_requests ?? 0} active\n`;
|
||||
const c = body.stats?.cache;
|
||||
if (c) {
|
||||
out += `Cache: ${c.hits ?? 0} hit / ${c.misses ?? 0} miss / ${c.size ?? "?"} entries\n`;
|
||||
}
|
||||
if (Array.isArray(body.recent_errors) && body.recent_errors.length > 0) {
|
||||
out += `\nRecent errors (${body.recent_errors.length}):\n`;
|
||||
for (const e of body.recent_errors.slice(0, 3)) {
|
||||
const ts = (e.time || "").slice(11, 19);
|
||||
const msg = String(e.message ?? "").slice(0, 60);
|
||||
out += ` ${ts} ${e.provider ?? "?"} ${msg}\n`;
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
export function fmtHealth(body) {
|
||||
const icon = statusIcon(body.ok ? "ok" : "fail");
|
||||
let out = `${icon} Status: ${body.ok ? "ok" : "fail"} | v${body.version ?? "?"}\n`;
|
||||
if (body.uptime_human || body.uptimeHuman) {
|
||||
out += `Uptime: ${body.uptime_human ?? body.uptimeHuman}\n`;
|
||||
}
|
||||
// D74 P2-4 fix: server.mjs /health full payload is
|
||||
// body.providers = { enabled: N, available: N, status: { <name>: {...} } }
|
||||
// The plugin previously iterated Object.entries(body.providers), which
|
||||
// surfaced `enabled`, `available`, and `status` as pseudo-providers
|
||||
// (typeof status === 'object' → loop body fired with name='status').
|
||||
// Walk providers.status when present; fall back to providers.* for the
|
||||
// older OCP shape that lacks the .status wrapper.
|
||||
if (body.providers && typeof body.providers === "object") {
|
||||
const enabled = body.providers.enabled;
|
||||
const available = body.providers.available;
|
||||
if (typeof enabled === "number" || typeof available === "number") {
|
||||
out += `Providers: ${enabled ?? "?"} enabled / ${available ?? "?"} available\n`;
|
||||
}
|
||||
const statusMap = body.providers.status && typeof body.providers.status === "object"
|
||||
? body.providers.status
|
||||
: body.providers;
|
||||
const entries = Object.entries(statusMap).filter(
|
||||
([name, s]) => typeof s === "object" && s !== null && name !== "enabled" && name !== "available" && name !== "status"
|
||||
);
|
||||
if (entries.length > 0) {
|
||||
out += `\nProviders:\n`;
|
||||
for (const [name, s] of entries) {
|
||||
const i = statusIcon(s?.ok ? "ok" : "fail");
|
||||
const spawn = typeof s?.activeSpawns === "number" ? ` spawns=${s.activeSpawns}` : "";
|
||||
out += ` ${i} ${name}${spawn}\n`;
|
||||
}
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
/**
|
||||
* formatResetCountdown(epochSeconds) → human-readable reset countdown.
|
||||
*
|
||||
* Mirrors bin/olp.mjs + dashboard.html versions. Five ranges:
|
||||
* past / < 1h / < 24h / < 7d / ≥ 7d
|
||||
*
|
||||
* Authority: ADR 0008 Amendment 2 (quota_v2 shape), ported from dashboard.html (D82).
|
||||
* No external deps. Duplicated here intentionally (olp-plugin ships separately).
|
||||
*/
|
||||
export function pluginFormatResetCountdown(epochSeconds) {
|
||||
if (epochSeconds == null) return "—";
|
||||
const nowMs = Date.now();
|
||||
const targetMs = epochSeconds * 1000;
|
||||
const diffMs = targetMs - nowMs;
|
||||
if (diffMs <= 0) return "resetting now";
|
||||
const diffMin = Math.floor(diffMs / 60000);
|
||||
const diffHr = Math.floor(diffMin / 60);
|
||||
const diffDay = Math.floor(diffHr / 24);
|
||||
if (diffMin < 60) return `resets in ${diffMin}m`;
|
||||
if (diffHr < 24) {
|
||||
const remMin = diffMin - diffHr * 60;
|
||||
if (remMin === 0) return `resets in ${diffHr}h`;
|
||||
return `resets in ${diffHr}h ${remMin}m`;
|
||||
}
|
||||
const target = new Date(targetMs);
|
||||
const timeStr = target.toLocaleString("en-US", { hour: "numeric", minute: "2-digit", hour12: true });
|
||||
if (diffDay < 7) {
|
||||
const dayStr = target.toLocaleString("en-US", { weekday: "short" });
|
||||
return `resets ${dayStr} ${timeStr}`;
|
||||
}
|
||||
const dateStr = target.toLocaleString("en-US", { month: "short", day: "numeric" });
|
||||
return `resets ${dateStr} ${timeStr}`;
|
||||
}
|
||||
|
||||
export function fmtUsage(body) {
|
||||
let out = "OLP usage (24h)\n";
|
||||
out += "─────────────────────────────\n";
|
||||
const w = body.window_24h ?? body.usage_24h ?? {};
|
||||
if (w.request_count !== undefined) {
|
||||
out += `Requests: ${w.request_count}\n`;
|
||||
const c = body.cache_hit_24h ?? {};
|
||||
if (typeof c.hit_rate === "number") {
|
||||
out += `Cache hit: ${(c.hit_rate * 100).toFixed(1)}%\n`;
|
||||
}
|
||||
} else if (w.requests !== undefined) {
|
||||
out += `Requests: ${w.requests}\n`;
|
||||
out += `Cache hit: ${w.cache_hit_rate != null ? `${(w.cache_hit_rate * 100).toFixed(1)}%` : "?"}\n`;
|
||||
out += `Fallbacks: ${w.fallbacks ?? "?"}\n`;
|
||||
} else if (typeof body.cache_hit_24h === "number") {
|
||||
// Legacy: cache_hit_24h as a bare number
|
||||
out += `Cache hit (24h): ${(body.cache_hit_24h * 100).toFixed(1)}%\n`;
|
||||
}
|
||||
|
||||
// F4: prefer quota_v2 when present (server v0.5.0+), fall back to legacy quota.
|
||||
// Authority: ADR 0008 Amendment 2 (quota_v2 shape).
|
||||
if (Array.isArray(body.quota_v2) && body.quota_v2.length > 0) {
|
||||
out += `\nPer-provider quota (live):\n`;
|
||||
for (const p of body.quota_v2) {
|
||||
const name = String(p.provider ?? "?").toUpperCase().padEnd(10);
|
||||
const status = p.status ?? "unavailable";
|
||||
if (status === "unavailable") {
|
||||
out += ` ${name} unavailable ${p.reason ?? "no public quota api"}\n`;
|
||||
} else if (status === "unreachable") {
|
||||
const fk = p.failure?.kind ?? "unknown";
|
||||
out += ` ${name} no cached data — failure: ${fk}\n`;
|
||||
} else {
|
||||
// live or stale
|
||||
const util = p.utilization ?? {};
|
||||
const reset = p.reset ?? {};
|
||||
const parts = [];
|
||||
for (const window of ["5h", "7d"]) {
|
||||
const frac = util[window];
|
||||
const resetEpoch = reset[window];
|
||||
if (frac != null) {
|
||||
const pct = `${Math.round(frac * 100)}%`;
|
||||
const rst = pluginFormatResetCountdown(resetEpoch);
|
||||
parts.push(`${window}: ${pct} (${rst})`);
|
||||
}
|
||||
}
|
||||
const staleNote = status === "stale"
|
||||
? ` ⚠ stale (${p.failure?.kind ?? "unknown"})`
|
||||
: "";
|
||||
out += ` ${name} ${status.padEnd(6)} ${parts.join(" ")}${staleNote}\n`;
|
||||
}
|
||||
}
|
||||
} else if (Array.isArray(body.quota) && body.quota.length > 0) {
|
||||
// Legacy fallback for pre-v0.5.0 servers
|
||||
out += `\nPer-provider quota:\n`;
|
||||
for (const q of body.quota) {
|
||||
const pct = typeof q.percent_used === "number" ? q.percent_used : null;
|
||||
const bar0 = pct != null ? ` ${bar(pct / 100, 12)} ${pct.toFixed(0)}%` : " no quota api";
|
||||
out += ` ${String(q.provider ?? q.name ?? "?").padEnd(10)}${bar0}\n`;
|
||||
}
|
||||
}
|
||||
|
||||
if (Array.isArray(body.top_fallback_chains_24h) && body.top_fallback_chains_24h.length > 0) {
|
||||
out += `\nTop fallback chains (24h):\n`;
|
||||
for (const f of body.top_fallback_chains_24h.slice(0, 5)) {
|
||||
out += ` ${String(f.count ?? "?").padStart(5)} ${(f.chain ?? []).join(" → ")}\n`;
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
export function fmtModels(body) {
|
||||
const data = body.data ?? [];
|
||||
if (data.length === 0) return "No models.";
|
||||
let out = `Models (${data.length})\n`;
|
||||
out += "─────────────────────────────\n";
|
||||
for (const m of data) {
|
||||
out += ` ${m.id}${m.owned_by ? ` (${m.owned_by})` : ""}\n`;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
export function fmtCache(body) {
|
||||
let out = "OLP cache\n";
|
||||
out += "─────────────────────────────\n";
|
||||
out += `Entries: ${body.size ?? body.entries ?? "?"}\n`;
|
||||
out += `Hits: ${body.hits ?? 0}\n`;
|
||||
out += `Misses: ${body.misses ?? 0}\n`;
|
||||
out += `Inflight: ${body.inflightCount ?? 0}\n`;
|
||||
if (typeof body.evictions === "number") {
|
||||
out += `Evictions: ${body.evictions}\n`;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
export function fmtProviders(registry, configEnabled) {
|
||||
const providers = registry?.providers ?? {};
|
||||
const names = Object.keys(providers);
|
||||
let out = `OLP providers (${names.length} in registry)\n`;
|
||||
out += "─────────────────────────────\n";
|
||||
for (const name of names) {
|
||||
const p = providers[name];
|
||||
const enabled = configEnabled?.[name] === true ? "enabled " : "disabled";
|
||||
const tier = p?.tier ?? "?";
|
||||
const modelCount = (p?.models ?? []).length;
|
||||
const candidate = p?.candidate === true ? " (candidate)" : "";
|
||||
out += ` ${name.padEnd(10)} ${enabled} tier ${tier} models ${String(modelCount).padStart(2)}${candidate}\n`;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
export function fmtChainShow(chains, target) {
|
||||
if (!chains || Object.keys(chains).length === 0) {
|
||||
return "No chains configured.";
|
||||
}
|
||||
if (target) {
|
||||
const chain = chains[target];
|
||||
if (!chain) {
|
||||
return `Model "${target}" not in routing.chains.\nConfigured: ${Object.keys(chains).join(", ")}`;
|
||||
}
|
||||
let out = `${target}:\n`;
|
||||
for (const hop of chain) {
|
||||
out += ` → ${typeof hop === "string" ? hop : JSON.stringify(hop)}\n`;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
let out = "OLP routing.chains\n";
|
||||
out += "─────────────────────────────\n";
|
||||
for (const [model, chain] of Object.entries(chains)) {
|
||||
out += `${model}:\n`;
|
||||
for (const hop of chain) {
|
||||
out += ` → ${typeof hop === "string" ? hop : JSON.stringify(hop)}\n`;
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
export function fmtDoctor(body) {
|
||||
// body shape: { checks, fail_count, warn_count, ok_count, kind, summary, next_action }
|
||||
let out = `OLP doctor — ${body.summary ?? "?"}\n`;
|
||||
out += "─────────────────────────────\n";
|
||||
for (const c of (body.checks ?? []).slice(0, 20)) {
|
||||
const icon = c.status === "ok" ? "🟢" : c.status === "warn" ? "🟡" : "🔴";
|
||||
out += ` ${icon} ${String(c.id ?? "?").padEnd(34)} ${String(c.message ?? "").slice(0, 60)}\n`;
|
||||
}
|
||||
if ((body.checks ?? []).length > 20) {
|
||||
out += ` ... (${body.checks.length - 20} more — use SSH 'olp doctor' for full output)\n`;
|
||||
}
|
||||
out += `\nfail=${body.fail_count ?? 0} warn=${body.warn_count ?? 0} ok=${body.ok_count ?? 0} kind=${body.kind ?? "?"}\n`;
|
||||
if (body.next_action?.ai_executable?.length > 0) {
|
||||
out += `\nNext (AI-executable):\n`;
|
||||
for (const cmd of body.next_action.ai_executable.slice(0, 5)) {
|
||||
out += ` $ ${cmd}\n`;
|
||||
}
|
||||
}
|
||||
if (body.next_action?.human_required?.length > 0) {
|
||||
out += `\nNext (human-required):\n`;
|
||||
for (const step of body.next_action.human_required.slice(0, 5)) {
|
||||
out += ` • ${step}\n`;
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// ── Help text ──────────────────────────────────────────────────────────────
|
||||
|
||||
export function cmdHelp() {
|
||||
return `OLP Commands (read-only)
|
||||
─────────────────────────────
|
||||
/olp status Process + provider + cache snapshot
|
||||
/olp health /health endpoint (public-ok)
|
||||
/olp usage 24h request stats + per-provider quota
|
||||
/olp models Available models
|
||||
/olp cache Cache stats
|
||||
/olp providers Provider registry + enabled flags
|
||||
/olp chain show [model] Routing chain(s) from server config
|
||||
/olp doctor Diagnostic checks + suggested next action
|
||||
/olp help This message
|
||||
|
||||
Mutating commands (keygen / revoke / restart / logs) are NOT
|
||||
available from chat by design — use SSH + the local 'olp' CLI.`;
|
||||
}
|
||||
|
||||
// ── Dispatcher ─────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Pure subcommand dispatcher. Returns `{ text }` always — caller wraps in
|
||||
* mono() for the chat surface.
|
||||
*
|
||||
* Exported for unit tests. Injects:
|
||||
* - fetchFn (default global fetch)
|
||||
* - proxyUrl (resolved upstream so tests can pin)
|
||||
* - apiKey (from plugin config)
|
||||
* - registry (models-registry.json — caller provides since this module
|
||||
* ships in `olp-plugin/` and the file is a sibling concept living at
|
||||
* the repo root)
|
||||
* - chainsLocal (local routing.chains override — usually empty; the
|
||||
* server-side /v0/management/status already exposes provider+chain
|
||||
* state, but chain-show is the one local-config touch that mirrors
|
||||
* `olp chain show`)
|
||||
*/
|
||||
export async function dispatch(rawArgs, opts) {
|
||||
const {
|
||||
proxyUrl,
|
||||
apiKey,
|
||||
registry,
|
||||
chainsLocal = {},
|
||||
fetchFn = fetch,
|
||||
} = opts;
|
||||
|
||||
const raw = (rawArgs || "").trim();
|
||||
const spaceIdx = raw.indexOf(" ");
|
||||
const subcmd = spaceIdx === -1 ? raw : raw.slice(0, spaceIdx);
|
||||
const subargs = spaceIdx === -1 ? "" : raw.slice(spaceIdx + 1).trim();
|
||||
|
||||
try {
|
||||
switch (subcmd) {
|
||||
case "status": {
|
||||
const body = await fetchJSON(`${proxyUrl}/v0/management/status`, { apiKey, fetchFn });
|
||||
return { text: fmtStatus(body) };
|
||||
}
|
||||
case "health": {
|
||||
const body = await fetchJSON(`${proxyUrl}/health`, { apiKey, fetchFn });
|
||||
return { text: fmtHealth(body) };
|
||||
}
|
||||
case "usage": {
|
||||
const body = await fetchJSON(`${proxyUrl}/v0/management/dashboard-data`, { apiKey, fetchFn });
|
||||
return { text: fmtUsage(body) };
|
||||
}
|
||||
case "models": {
|
||||
const body = await fetchJSON(`${proxyUrl}/v1/models`, { apiKey, fetchFn });
|
||||
return { text: fmtModels(body) };
|
||||
}
|
||||
case "cache": {
|
||||
const body = await fetchJSON(`${proxyUrl}/cache/stats`, { apiKey, fetchFn });
|
||||
return { text: fmtCache(body) };
|
||||
}
|
||||
case "providers": {
|
||||
// models-registry.json + (optionally) the server's idea of which are
|
||||
// enabled. /v0/management/status carries that and is owner-gated, but
|
||||
// /v1/models lists what's exposed publicly. For the chat surface we
|
||||
// use the public registry shape — config.enabled is a local-config
|
||||
// concept and the plugin doesn't have filesystem access to
|
||||
// ~/.olp/config.json by design.
|
||||
return { text: fmtProviders(registry, {}) };
|
||||
}
|
||||
case "chain": {
|
||||
// /olp chain show [model]
|
||||
const inner = subargs.trim();
|
||||
const parts = inner.split(/\s+/).filter(Boolean);
|
||||
if (parts[0] !== "show") {
|
||||
return { text: `Usage: /olp chain show [model]` };
|
||||
}
|
||||
const target = parts[1] ?? null;
|
||||
return { text: fmtChainShow(chainsLocal, target) };
|
||||
}
|
||||
case "doctor": {
|
||||
// /v0/management/doctor doesn't exist yet — D67 added doctor as a CLI
|
||||
// surface only. The plugin reports that explicitly so families know
|
||||
// to use SSH + `olp doctor` rather than waiting for a chat response.
|
||||
return {
|
||||
text: `/olp doctor is not yet wired through HTTP (planned for Phase 5+).\n` +
|
||||
`Run \`olp doctor\` over SSH on the host running the OLP server\n` +
|
||||
`for the full diagnostic output.`,
|
||||
};
|
||||
}
|
||||
case "help":
|
||||
case "--help":
|
||||
case "-h":
|
||||
case "":
|
||||
return { text: cmdHelp() };
|
||||
default:
|
||||
return { text: `Unknown subcommand: ${subcmd}\n\n${cmdHelp()}` };
|
||||
}
|
||||
} catch (err) {
|
||||
return { text: `OLP error: ${err.message ?? String(err)}` };
|
||||
}
|
||||
}
|
||||
|
||||
// ── Plugin entry point (consumed by OpenClaw gateway) ──────────────────────
|
||||
|
||||
/**
|
||||
* OpenClaw plugin entry. The gateway calls this with its `api` registration
|
||||
* object; we register the `/olp` slash command and a handler that resolves
|
||||
* the proxy URL + API key from plugin config + env, then delegates to
|
||||
* `dispatch()`.
|
||||
*
|
||||
* The `registry` (models-registry.json) is read lazily inside the handler
|
||||
* so that a stale plugin install doesn't bind to an old snapshot — and so
|
||||
* the plugin module stays import-time-pure for tests.
|
||||
*/
|
||||
export default function (api) {
|
||||
api.registerCommand({
|
||||
name: "olp",
|
||||
description: "OLP — usage, health, status, doctor, etc. (read-only)",
|
||||
acceptsArgs: true,
|
||||
requireAuth: true,
|
||||
handler: async (ctx) => {
|
||||
const cfg = ctx.config ?? {};
|
||||
const apiKey = cfg.apiKey ?? process.env.OLP_API_KEY ?? null;
|
||||
const proxyUrl = resolveProxyUrl({ env: process.env, config: cfg });
|
||||
// Lazy load to avoid binding the import to the OpenClaw gateway's
|
||||
// ESM cache (which may pre-resolve at plugin-discovery time).
|
||||
let registry;
|
||||
try {
|
||||
// Convert file path to URL for ESM `import(...)`.
|
||||
const { fileURLToPath, pathToFileURL } = await import("node:url");
|
||||
const { dirname, resolve: pathResolve } = await import("node:path");
|
||||
const here = dirname(fileURLToPath(import.meta.url));
|
||||
const registryUrl = pathToFileURL(pathResolve(here, "..", "models-registry.json")).href;
|
||||
registry = (await import(registryUrl, { with: { type: "json" } })).default;
|
||||
} catch (e) {
|
||||
registry = { providers: {} };
|
||||
}
|
||||
// Local chains config is not currently surfaced through HTTP. For
|
||||
// chat-side chain-show we fall back to an empty map; operators
|
||||
// wanting the live config view should use `olp chain show` over SSH.
|
||||
const chainsLocal = {};
|
||||
const { text } = await dispatch(ctx.args ?? "", {
|
||||
proxyUrl,
|
||||
apiKey,
|
||||
registry,
|
||||
chainsLocal,
|
||||
});
|
||||
return { text: mono(truncateForChat(text)) };
|
||||
},
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
{
|
||||
"id": "olp",
|
||||
"name": "OLP Commands",
|
||||
"description": "Slash commands for OLP — /olp status, /olp usage, /olp health, etc. (read-only by design; mutations require SSH).",
|
||||
"version": "0.4.0",
|
||||
"configSchema": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"properties": {
|
||||
"proxyUrl": {
|
||||
"type": "string",
|
||||
"default": "http://127.0.0.1:4567",
|
||||
"description": "Full URL of the OLP proxy. Default matches D60 OLP_PORT=4567. Overridable via OLP_PROXY_URL or OLP_PORT env."
|
||||
},
|
||||
"apiKey": {
|
||||
"type": "string",
|
||||
"description": "Owner-tier OLP API key (olp_xxx). Required for owner-only subcommands (status / usage / cache). Use a dedicated bot key — DO NOT share the maintainer's personal owner key."
|
||||
}
|
||||
},
|
||||
"required": ["apiKey"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,16 @@
|
||||
{
|
||||
"name": "olp-plugin",
|
||||
"version": "0.4.0",
|
||||
"description": "OpenClaw gateway plugin — /olp slash commands for the OLP proxy (read-only)",
|
||||
"main": "index.js",
|
||||
"type": "module",
|
||||
"keywords": ["openclaw", "plugin", "olp", "proxy"],
|
||||
"license": "MIT",
|
||||
"private": true,
|
||||
"openclaw": {
|
||||
"type": "plugin",
|
||||
"id": "olp",
|
||||
"pluginManifest": "openclaw.plugin.json",
|
||||
"extensions": ["./index.js"]
|
||||
}
|
||||
}
|
||||
Generated
+89
@@ -0,0 +1,89 @@
|
||||
{
|
||||
"name": "olp",
|
||||
"version": "0.5.1",
|
||||
"lockfileVersion": 3,
|
||||
"requires": true,
|
||||
"packages": {
|
||||
"": {
|
||||
"name": "olp",
|
||||
"version": "0.5.1",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"@anthropic-ai/sandbox-runtime": "^0.0.52"
|
||||
},
|
||||
"bin": {
|
||||
"olp": "bin/olp.mjs",
|
||||
"olp-audit-rotate": "bin/olp-audit-rotate.mjs",
|
||||
"olp-connect": "bin/olp-connect",
|
||||
"olp-keys": "bin/olp-keys.mjs"
|
||||
},
|
||||
"engines": {
|
||||
"node": ">=18"
|
||||
}
|
||||
},
|
||||
"node_modules/@anthropic-ai/sandbox-runtime": {
|
||||
"version": "0.0.52",
|
||||
"resolved": "https://registry.npmjs.org/@anthropic-ai/sandbox-runtime/-/sandbox-runtime-0.0.52.tgz",
|
||||
"integrity": "sha512-vYaM7OslFmOAzNgfy5gxvt3NoWFeCbr7C0AKyuduQq7Gdxbg2NnYmE7deBf8Nxj3ZNECTcC5RhAfz0lZwvbtBA==",
|
||||
"license": "Apache-2.0",
|
||||
"dependencies": {
|
||||
"@pondwader/socks5-server": "^1.0.10",
|
||||
"commander": "^12.1.0",
|
||||
"node-forge": "^1.4.0",
|
||||
"shell-quote": "^1.8.3",
|
||||
"zod": "^3.24.1"
|
||||
},
|
||||
"bin": {
|
||||
"srt": "dist/cli.js"
|
||||
},
|
||||
"engines": {
|
||||
"node": ">=18.0.0"
|
||||
}
|
||||
},
|
||||
"node_modules/@pondwader/socks5-server": {
|
||||
"version": "1.0.10",
|
||||
"resolved": "https://registry.npmjs.org/@pondwader/socks5-server/-/socks5-server-1.0.10.tgz",
|
||||
"integrity": "sha512-bQY06wzzR8D2+vVCUoBsr5QS2U6UgPUQRmErNwtsuI6vLcyRKkafjkr3KxbtGFf9aBBIV2mcvlsKD1UYaIV+sg==",
|
||||
"license": "MIT"
|
||||
},
|
||||
"node_modules/commander": {
|
||||
"version": "12.1.0",
|
||||
"resolved": "https://registry.npmjs.org/commander/-/commander-12.1.0.tgz",
|
||||
"integrity": "sha512-Vw8qHK3bZM9y/P10u3Vib8o/DdkvA2OtPtZvD871QKjy74Wj1WSKFILMPRPSdUSx5RFK1arlJzEtA4PkFgnbuA==",
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
"node": ">=18"
|
||||
}
|
||||
},
|
||||
"node_modules/node-forge": {
|
||||
"version": "1.4.0",
|
||||
"resolved": "https://registry.npmjs.org/node-forge/-/node-forge-1.4.0.tgz",
|
||||
"integrity": "sha512-LarFH0+6VfriEhqMMcLX2F7SwSXeWwnEAJEsYm5QKWchiVYVvJyV9v7UDvUv+w5HO23ZpQTXDv/GxdDdMyOuoQ==",
|
||||
"license": "(BSD-3-Clause OR GPL-2.0)",
|
||||
"engines": {
|
||||
"node": ">= 6.13.0"
|
||||
}
|
||||
},
|
||||
"node_modules/shell-quote": {
|
||||
"version": "1.8.4",
|
||||
"resolved": "https://registry.npmjs.org/shell-quote/-/shell-quote-1.8.4.tgz",
|
||||
"integrity": "sha512-VsC6n6vz1ihYYyZZwX7YZSF5l5x36ca17OC+a69h94YqB7X6XLwf+5MOgynYir2SLFUbl8gIYvBo8K8RoNQ6bQ==",
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
"node": ">= 0.4"
|
||||
},
|
||||
"funding": {
|
||||
"url": "https://github.com/sponsors/ljharb"
|
||||
}
|
||||
},
|
||||
"node_modules/zod": {
|
||||
"version": "3.25.76",
|
||||
"resolved": "https://registry.npmjs.org/zod/-/zod-3.25.76.tgz",
|
||||
"integrity": "sha512-gzUt/qt81nXsFGKIFcC3YnfEAx5NkunCfnDlvuBSSFS02bcXu4Lmea0AFIUwbLWxWPx3d9p8S5QoaujKcNQxcQ==",
|
||||
"license": "MIT",
|
||||
"funding": {
|
||||
"url": "https://github.com/sponsors/colinhacks"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
+24
-4
@@ -1,19 +1,36 @@
|
||||
{
|
||||
"name": "olp",
|
||||
"version": "0.3.1",
|
||||
"version": "0.7.0",
|
||||
"description": "Personal multi-provider LLM proxy. Successor to OCP. One HTTP endpoint, multiple subscriptions behind it, automatic routing + fallback + caching.",
|
||||
"type": "module",
|
||||
"main": "server.mjs",
|
||||
"bin": {
|
||||
"olp": "./bin/olp.mjs",
|
||||
"olp-keys": "./bin/olp-keys.mjs",
|
||||
"olp-audit-rotate": "./bin/olp-audit-rotate.mjs"
|
||||
"olp-audit-rotate": "./bin/olp-audit-rotate.mjs",
|
||||
"olp-connect": "./bin/olp-connect"
|
||||
},
|
||||
"scripts": {
|
||||
"start": "node server.mjs",
|
||||
"test": "node test-features.mjs",
|
||||
"olp": "node bin/olp.mjs",
|
||||
"olp-keys": "node bin/olp-keys.mjs",
|
||||
"olp-audit-rotate": "node bin/olp-audit-rotate.mjs"
|
||||
"olp-audit-rotate": "node bin/olp-audit-rotate.mjs",
|
||||
"olp-connect": "bash bin/olp-connect"
|
||||
},
|
||||
"files": [
|
||||
"server.mjs",
|
||||
"bin/",
|
||||
"lib/",
|
||||
"olp-plugin/",
|
||||
"models-registry.json",
|
||||
"dashboard.html",
|
||||
"README.md",
|
||||
"ALIGNMENT.md",
|
||||
"CHANGELOG.md",
|
||||
"LICENSE",
|
||||
"docs/"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=18"
|
||||
},
|
||||
@@ -32,5 +49,8 @@
|
||||
"mistral",
|
||||
"fallback",
|
||||
"cache"
|
||||
]
|
||||
],
|
||||
"dependencies": {
|
||||
"@anthropic-ai/sandbox-runtime": "^0.0.52"
|
||||
}
|
||||
}
|
||||
|
||||
+901
-198
File diff suppressed because it is too large
Load Diff
+6321
-38
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user