mirror of
https://github.com/dtzp555-max/olp.git
synced 2026-07-22 05:25:09 +00:00
Compare commits
31
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3551921f55 | ||
|
|
b1e24b7cb0 | ||
|
|
497b2550e6 | ||
|
|
28642756b5 | ||
|
|
d0dcd281ef | ||
|
|
07d9c8a6ae | ||
|
|
e5cfc696da | ||
|
|
dbac5f5521 | ||
|
|
dd0c821272 | ||
|
|
65f945c16d | ||
|
|
97e7d16585 | ||
|
|
40f9453d88 | ||
|
|
e2f41eb60e | ||
|
|
5d60a0599f | ||
|
|
ea0392f744 | ||
|
|
cc250e71bf | ||
|
|
9dc070bc53 | ||
|
|
fa2d1af130 | ||
|
|
bddf2cba1e | ||
|
|
65681ed7d2 | ||
|
|
d872330c9e | ||
|
|
2b07a3bd1b | ||
|
|
a41420d0fc | ||
|
|
5288493f19 | ||
|
|
82d2e1cbea | ||
|
|
187e79321f | ||
|
|
1605400052 | ||
|
|
704d4fc8a0 | ||
|
|
6605b7b14a | ||
|
|
6edf6e0b94 | ||
|
|
f3716a19fd |
@@ -49,7 +49,7 @@ Runtime: Node.js (ESM, `.mjs` throughout). No build step. No bundler. `server.mj
|
|||||||
- `.github/workflows/alignment.yml` — CI blacklist grep + per-provider citation soft check; fails the build on known-hallucinated tokens.
|
- `.github/workflows/alignment.yml` — CI blacklist grep + per-provider citation soft check; fails the build on known-hallucinated tokens.
|
||||||
- `CLAUDE.md` — Claude-Code-specific session instructions + `release_kit` overlay (Iron Rule 5.5).
|
- `CLAUDE.md` — Claude-Code-specific session instructions + `release_kit` overlay (Iron Rule 5.5).
|
||||||
|
|
||||||
**Implementation status note (as of 2026-05-25):** Files marked 📋 above are designed and documented but not yet on disk; files marked 🟡 are partially shipped; files marked ✅ are Phase 2 deliverables. The shipped set as of D47 is: `server.mjs` (with Phase 2 auth middleware + audit wire + owner-vs-non-owner gating), `lib/ir/`, `lib/providers/{anthropic,codex,mistral}.mjs`, `lib/cache/{keys,store}.mjs`, `lib/fallback/engine.mjs`, `lib/keys.mjs` (core + loadAuthConfigSync — D44 + D45), `lib/audit.mjs` (D45), `bin/olp-keys.mjs` (D47), `models-registry.json`, `test-features.mjs` (Suites 19–22). Phase 2 functional scope is complete; remaining is Phase 2 close → v0.2.0 (maintainer-triggered, explicit per CLAUDE.md `release_kit.phase_close_trigger`).
|
**Implementation status note (as of 2026-05-27):** Phase 5 (Quota Probes + Dashboard Enrichment) is closed at v0.5.0 + v0.5.1 hotfix. The shipped set includes all Phases 1–5 deliverables. v0.5.1 hotfix (2026-05-27) fixes three codex review findings: F1 (doctor check bypassed backoff by calling `_probeOnce` directly — now routes through `quotaStatus()`), F2 (200 with empty `anthropic-ratelimit-*` headers was cached as live — minimum-viable-schema gate added), F3 (null collapsed all failure modes — `probe_status:'unreachable'` shape + `failure`/`failure_kind`/`backoff_until` fields added). See ADR 0008 Amendment 2 + ADR 0013 Rule 3/5 clarifications. Phase 6 is next (per CLAUDE.md `release_kit.current_phase`).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
+12
-2
@@ -196,9 +196,19 @@ In addition to the recurring 14 May audit below, the following one-shot audits a
|
|||||||
|
|
||||||
## Class-specific Exceptions
|
## Class-specific Exceptions
|
||||||
|
|
||||||
(none at project founding)
|
Any Rule 2 or Rule 3 deviation lands here as a numbered exception with PR link, reviewer, and rationale.
|
||||||
|
|
||||||
Any future Rule 3 deviation lands here as a numbered exception with PR link, reviewer, and rationale.
|
### 1. Anthropic plan-usage probe via direct `/v1/messages` call (Phase 5, D79 — 2026-05-26)
|
||||||
|
|
||||||
|
**Class:** Rule 2(a) — provider-plugin scope. The Anthropic plugin's `quotaStatus()` calls `POST https://api.anthropic.com/v1/messages` directly rather than spawning `claude -p`. Under the strict reading of Rule 2(a), plugins must mirror provider-CLI behaviour; under the strict reading, this is a deviation because the spawn path goes through the CLI binary and the probe path does not.
|
||||||
|
|
||||||
|
**Authority:** ADR 0002 Amendment 8 (governance) + ADR 0013 (implementation discipline) + ADR 0012 (Phase 5 charter). Schema pin: `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md` (compiled-binary `strings` + live API probe evidence). PR #50.
|
||||||
|
|
||||||
|
**Rationale:** Claude Code's compiled binary makes the same `POST /v1/messages` call internally (verified by `strings` over the v2.1.142 / v2.1.150 Mach-O / ELF binary). The probe mirrors that observed CLI behaviour without introducing a new wire format or output assumption. The exemption is bounded by ADR 0002 Amendment 8's three constraints (READ-ONLY, subscription-scope, idempotent-failure) + ADR 0013's seven implementation rules (notably Rule 2's per-endpoint enumeration — only `POST /v1/messages` is permitted).
|
||||||
|
|
||||||
|
**Reviewer:** fresh-context opus subagent on PR #50 (Iron Rule 10 + CLAUDE.md hard requirement #3). Verdict: APPROVE_WITH_MINOR. Six in-PR nits folded in; three outside-PR nits documented and addressed (this entry is one of them — N9).
|
||||||
|
|
||||||
|
**Re-evaluation trigger:** if Anthropic publishes a public documented quota endpoint (e.g. `GET /v1/usage`), this exception is RETIRED and the plugin migrates to the documented endpoint, deleting this exception by amendment PR. Until that hypothetical retirement, this exception is the canonical entry.
|
||||||
|
|
||||||
### Controlled deviations (entry-surface scope)
|
### Controlled deviations (entry-surface scope)
|
||||||
|
|
||||||
|
|||||||
+207
-1
@@ -4,7 +4,213 @@ All notable changes to OLP land here. Per `CLAUDE.md` release_kit overlay, this
|
|||||||
|
|
||||||
## Unreleased
|
## Unreleased
|
||||||
|
|
||||||
(empty — Phase 5 entries land here once Phase 5 opens)
|
### Phase 7 PR-B — anthropic.mjs spawn wrapped in sandbox-runtime
|
||||||
|
|
||||||
|
- feat(sandbox): Phase 7 PR-B — `lib/providers/anthropic.mjs` spawn wrapped via `@anthropic-ai/sandbox-runtime` with config-at-boot model (per-spawn ephemeral cwd `/tmp/olp-spawn/<uuid>`, network allowlist `api.anthropic.com` + `statsig.anthropic.com`, filesystem denylist for `~/.olp` / `~/.claude` / `~/.ssh` / `~/.config` / `~/.codex`). Load-bearing negative test (Suite 44, PI231-gated) confirms in-sandbox `cat` of OAuth credentials MUST fail. `/health.sandbox.active=true` on PI231 after `apt-get install bubblewrap socat ripgrep`. Adds `lib/sandbox/manager.mjs` (bootstrap + spawn-wrap layer), server startup wiring (`bootstrapSandbox()` before listen), `/health.sandbox.active` boolean field. 805 → 813 tests (+8 Suite 43; Suite 44 skips by default, runs on PI231 with `OLP_E2E_SANDBOX=1`). ADR 0014 PR-B acceptance criteria: met.
|
||||||
|
|
||||||
|
### Phase 7 PR-A — sandbox-runtime dep + doctor + ADR 0014
|
||||||
|
|
||||||
|
- feat(sandbox): Phase 7 PR-A — @anthropic-ai/sandbox-runtime dep + lib/sandbox/doctor.mjs preflight + ADR 0014. No runtime wiring yet (PR-B will wrap anthropic.mjs spawn). /health now reports sandbox availability (`available: false` until PI231 has `bubblewrap` + `socat` + `ripgrep` installed via `sudo apt-get install -y bubblewrap socat ripgrep`). On macOS (dev machine with ripgrep via Homebrew), sandbox-runtime reports `available: true` because macOS uses the built-in `sandbox-exec` seatbelt — no apt install needed. 797 → 805 tests (+8 Suite 42).
|
||||||
|
|
||||||
|
### Phase 6 D-day — stream-json transport for Anthropic provider (ADR 0009 Amendment 1)
|
||||||
|
|
||||||
|
- feat(anthropic): stream-json output + --system-prompt suppression of env-block / tool descriptions (ADR 0009 Amendment 1). Cuts ~64% per-request cost on Sonnet 4.6 via 30% input-token reduction ($0.0216 → $0.0078), fixes bot self-check hallucination (model no longer claims server cwd / OS / tool names), exposes rate_limit + usage events from NDJSON for future audit/dashboard work. Per-key API + cache + audit semantics unchanged. claude CLI v2.1.104 verified; warn if claude-version outside v2.1.100–v2.1.149.
|
||||||
|
|
||||||
|
### F4 — `bin/olp.mjs` + `olp-plugin/index.js` migration to `quota_v2` shape
|
||||||
|
|
||||||
|
**Codex post-v0.5.0 review Q4.** Both CLI surfaces (`olp usage` and `/olp usage`) previously fell through to "no quota api" for every provider because they read the legacy `body.quota` shape, which never carries `percent_used` or meaningful `available` data. Now that the server (v0.5.0+) emits `body.quota_v2` per ADR 0008 Amendment 2, both surfaces prefer `quota_v2` and fall back to legacy `quota` on older servers.
|
||||||
|
|
||||||
|
- **`bin/olp.mjs cmdUsage`**: when `body.quota_v2` is present (non-empty array), renders per-provider rows with status (`live` / `stale` / `unreachable` / `unavailable`), 5h and 7d utilization percentages with color-coding (green < 50% / yellow 50–80% / red ≥ 80%), reset countdowns, binding claim, and ⚠ stale / ❌ unreachable annotations. Legacy `body.quota` path preserved as fallback for pre-v0.5.0 servers. `formatResetCountdown(epochSeconds)` added — 5-range formatter (past / <1h / <24h / <7d / ≥7d), ported from `dashboard.html` D82, kept in-file (no shared lib).
|
||||||
|
|
||||||
|
- **`olp-plugin/index.js fmtUsage()`**: same migration — `quota_v2` rows render as one-line plain text per provider (no ANSI; Telegram/Discord safe). `pluginFormatResetCountdown(epochSeconds)` added; intentionally duplicated (plugin ships as a separate package). Legacy `body.quota` fallback preserved.
|
||||||
|
|
||||||
|
### v1.x roadmap #7 — AUTH_MISSING tuple path test coverage — ✅ CLOSED
|
||||||
|
|
||||||
|
The dedicated AUTH_MISSING engine test (asserting `fallbackDetail[0].trigger_type === 'auth_missing'`) was already shipped at D56 (`test-features.mjs` line 6255). This item closes the roadmap entry with a date stamp and PR reference per the tracker convention. No code changes — documentation only.
|
||||||
|
|
||||||
|
### Tests
|
||||||
|
|
||||||
|
- Suite 40 (9 new tests): `40a`–`40i` covering `cmdUsage` quota_v2 live/stale/unreachable/unavailable parse, legacy fallback, `pluginFormatResetCountdown` and `formatResetCountdown` 5-range coverage, olp-plugin `fmtUsage` quota_v2 + legacy paths. 759 → 768 tests, 0 fail.
|
||||||
|
|
||||||
|
### Authority
|
||||||
|
|
||||||
|
- F4: codex post-v0.5.0 review Q4 (PR #58 review); ADR 0008 Amendment 2 (quota_v2 shape).
|
||||||
|
- #7: `docs/v1x-roadmap.md` § "#7 — AUTH_MISSING tuple path test coverage (D40 follow-up)".
|
||||||
|
|
||||||
|
## v0.5.1 — 2026-05-27
|
||||||
|
|
||||||
|
**Hotfix — Quota probe cache/backoff/schema-drift correctness (codex review findings F1–F3).** Three production-quality bugs in the v0.5.0 quota probe, reproduced by codex with local mocks, are corrected. 756 → 759 tests (3 new regression tests); 4 existing test assertions updated to reflect the v0.5.1 return-shape contract.
|
||||||
|
|
||||||
|
### Fixes
|
||||||
|
|
||||||
|
- **F1 [P1] — Doctor bypass of cache + backoff (ADR 0013 Rule 3).** `anthropic.quota_probe_reachable` doctor check called `_probeOnce(auth)` directly, bypassing the module-level `quotaProbeState.backoffUntil` check. Successive `olp doctor` invocations within a backoff window each hit upstream — violating ADR 0013 Rule 3 (60s-3600s exponential backoff is mandatory for all consumers). **Fix:** doctor check now routes through `quotaStatus()`, which enforces cache + backoff. ADR 0013 Rule 3 clarification added: "All consumers of `quotaStatus()`, including `olp doctor` checks, MUST route through `quotaStatus()` and MUST NOT call `_probeOnce()` directly."
|
||||||
|
|
||||||
|
- **F2 [P2] — 200 with empty `anthropic-ratelimit-*` headers cached as live data (ADR 0013 Rule 5).** `_probeOnce` treated any 200 OK (regardless of header content) as a successful probe, caching it with `stale: false` even when zero `anthropic-ratelimit-*` headers were present. A proxy stripping headers, a schema change, or a mock returning `{}` would silently appear as "LIVE" on the dashboard with all bars empty. **Fix:** minimum-viable-schema gate requires these 4 fields non-null: `5h-utilization`, `5h-reset`, `7d-utilization`, `7d-reset`. Any absence → `failureKind: 'schema_drift'`, backoff scheduled, result not cached. ADR 0013 Rule 5 updated with the gate specification.
|
||||||
|
|
||||||
|
- **F3 [P2] — Dashboard-data loses failure detail (ADR 0013 Rule 6).** `aggregateProviderQuota()` collapsed all non-null failure modes (no credentials, auth failure, rate limit, schema drift, network error) into `status: 'unavailable', reason: 'no public quota api or probe disabled'` — the same string as providers with no quota API at all. Operator could not tell what to fix. **Fix:** `quotaStatus()` v0.5.1 return contract: `null` reserved for opt-in-off only; probe failures return `{ probe_status: 'unreachable', failure: { kind, message, backoff_until? } }`. `aggregateProviderQuota()` emits new fields `failure_kind`, `failure`, `backoff_until` per row. `status: 'unreachable'` distinguishes "probe failed" from `status: 'unavailable'` ("no API or disabled"). Dashboard renders `unreachable` with a red border + failure.message + backoff countdown.
|
||||||
|
|
||||||
|
### Backwards-compat notes
|
||||||
|
|
||||||
|
- `quotaStatus()`: `stale: false` → now also includes `probe_status: 'live'` (additive). `stale: true` → now also includes `probe_status: 'stale'` + `failure: {...}` (additive). `null` → NOW RESERVED FOR OPT-IN-OFF ONLY (breaking for callers that relied on `null` to detect "no credentials" or "probe failed" — use `probe_status: 'unreachable'` instead).
|
||||||
|
- `ProviderQuotaEntry.status`: gains `'unreachable'` as a new value (additive). Existing `'live'`, `'stale'`, `'unavailable'` semantics unchanged.
|
||||||
|
- `ProviderQuotaEntry` gains new fields `failure`, `failure_kind`, `backoff_until` (additive, null when not applicable).
|
||||||
|
- `dashboard.html`: handles `unreachable` row (no existing row had this status; additive render path).
|
||||||
|
|
||||||
|
### Test changes
|
||||||
|
|
||||||
|
- 38f, 38j, 38l: updated assertions from `null` to `probe_status: 'unreachable'` (F3 shape change).
|
||||||
|
- 38r: refactored to seed cache + manually expire it + set backoff (F1 — doctor now routes through `quotaStatus()`). Added F1-regression assertion: HTTP call counter stays at 1 after two doctor calls within backoff.
|
||||||
|
- 38g, 38k: added `probe_status` + `failure` assertions (verify new fields present on live/stale shapes).
|
||||||
|
- **38u** (new): F1 regression — successive doctor calls within backoff window → HTTP counter stays at 1.
|
||||||
|
- **38v** (new): F2 regression — 200 + empty ratelimit headers → `probe_status: 'unreachable'` + `failure_kind: 'schema_drift'` + cache stays null.
|
||||||
|
- **38w** (new): F3 regression — `lastError` + `failureKind` propagate through `quotaStatus()` shape for all failure modes (rate_limited / auth_failed / schema_drift / no_credentials).
|
||||||
|
|
||||||
|
### ADR changes
|
||||||
|
|
||||||
|
- **ADR 0013 Rule 3** clarification: doctor checks route through `quotaStatus()`, not `_probeOnce()` directly.
|
||||||
|
- **ADR 0013 Rule 5** update: minimum-viable-schema gate specification (4 required fields; absence = schema_drift signal).
|
||||||
|
- **ADR 0008 Amendment 2**: richer `ProviderQuotaEntry` shape with `failure`/`failure_kind`/`backoff_until`; `probe_status` on `quotaStatus()` return; `unreachable` status semantics; `dashboard.html` unreachable rendering.
|
||||||
|
|
||||||
|
### Authority
|
||||||
|
|
||||||
|
ADR 0013 Rules 3, 5, 6 (cache + backoff + schema-drift + failure transparency); ADR 0008 Amendment 2; ADR 0002 Amendment 8 (unchanged); codex review findings F1–F3 (codex PR review on v0.5.0 close PR #57).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## v0.5.0 — 2026-05-26
|
||||||
|
|
||||||
|
**Phase 5 — Provider Quota Probes + Dashboard Enrichment.** OLP gains live subscription-quota observability for Anthropic Pro/Max subscribers, surfaced through a Claude.ai-style Plan Usage panel on the owner-only dashboard. The probe is opt-in, READ-ONLY, idempotent on failure, and 5-min-cached with 60s→3600s exponential backoff. Six D-days, seven PRs, zero blocking reviewer findings, no flaky tests; 720 → 756 total tests.
|
||||||
|
|
||||||
|
### What's new for users
|
||||||
|
|
||||||
|
- **Live plan usage on the dashboard.** Per-provider rows show 5-hour + 7-day utilization bars with reset countdowns ("Resets in 1hr 6min" / "Resets Sun 9:00 PM"), status badges (allowed / rejected), representative-claim chips ("five_hour" / "seven_day"), overage-status indicators, and a `↻ Refresh` button. 60-second auto-refresh pauses when the tab is hidden.
|
||||||
|
- **Anthropic quota probe.** Opt-in via `~/.olp/config.json providers.anthropic.quota_probe_enabled: true`. Parses the canonical `anthropic-ratelimit-unified-*` response-header schema (13 fields) from a minimal `POST /v1/messages` probe. Reuses the spawn-path OAuth credentials — env var → `~/.claude/.credentials.json` → macOS Keychain. Refresh-on-401, stale-cache-on-failure.
|
||||||
|
- **`olp doctor anthropic.quota_probe_reachable`.** New check surfaces probe health. Returns `status: ok` with parsed utilization when fresh, `warn` on stale cache, `fail` with `human_steps[]` auth-aware recipe (re-login via `claude setup-token` or wait-and-retry).
|
||||||
|
- **Provider matrix.** Anthropic ✅ live (13 fields). OpenAI ❌ no public quota API. Mistral ❌ no member-key-accessible quota endpoint (Admin API exists but org-admin-scoped, out of scope for trusted-LAN deployment per ADR 0011). All three pinned in `models-registry.json quota_probe.<provider>` block.
|
||||||
|
|
||||||
|
### What's new for contributors
|
||||||
|
|
||||||
|
- **ADR 0012 (Phase 5 charter)** — D-day plan + exit gate + scope boundaries (`docs/adr/0012-phase-5-charter-quota-probes-dashboard.md`).
|
||||||
|
- **ADR 0002 Amendment 8** — first Class-specific Exception to the plugin contract: `quotaStatus()` may call provider HTTP APIs directly, subject to three constraints (READ-ONLY, subscription-scope, idempotent-failure) and the per-endpoint enumeration in ADR 0013 Rule 2.
|
||||||
|
- **ADR 0013** — seven rules covering OAuth READ-ONLY consumption + dual-path schema-drift mitigation (compiled-binary `strings` + live API probe diff, since Claude Code v2.1.x is now a Mach-O / ELF binary with no `cli.js` to grep).
|
||||||
|
- **`models-registry.json quota_probe.schema_version`** — pinned at `2026-05-26` (13 fields). Bump on schema-drift events per ADR 0013 Rule 5.
|
||||||
|
- **Test seams** — 5 underscore-prefixed exports in `lib/providers/anthropic.mjs` (`_setQuotaUrlsForTest`, `_resetQuotaProbeStateForTest`, `_resetQuotaStateOnlyForTest`, `_getQuotaProbeStateForTest`, `_setQuotaAuthReadFnForTest`) for hermetic probe testing. Production code must not call them.
|
||||||
|
- **ALIGNMENT.md § Class-specific Exceptions** — gains its first numbered exception (Anthropic plan-usage probe via direct `/v1/messages`).
|
||||||
|
- **Audit memory at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`** — schema canon + verification protocol + OCP institutional history.
|
||||||
|
|
||||||
|
### D-day-level changes (Phase 5)
|
||||||
|
|
||||||
|
- **D79** (PR #50 + cleanup PR #51): governance layer — ADR 0012 charter, ADR 0002 Amendment 8, ADR 0013, ALIGNMENT.md Class-specific Exceptions entry, D84 Mistral NO-GO disposition.
|
||||||
|
- **D80** (PR #52): ported OCP `server.mjs:842-1109` to `lib/providers/anthropic.mjs:quotaStatus()`. Adds macOS-keychain reader to `readAuthArtifact()`. Parses all 13 fields including 3 new since OCP's 2026-04 capture (`5h-status`, `7d-status`, `overage-reset`). Implements 5min cache + 60s-3600s exponential refresh backoff + stale-cache-on-failure + opt-in config flag + `anthropic.quota_probe_reachable` doctor check. ~250 LOC.
|
||||||
|
- **D81** (PR #53): added `lib/audit-query.mjs aggregateProviderQuota()` + `/v0/management/dashboard-data quota_v2` field + `/v0/management/quota quota_v2` field. Pinned `quota_probe.schema_version` in `models-registry.json`. Legacy `quota` field stays alongside for backwards compat until v1.0.0. ADR 0008 Amendment 1 documents the shape.
|
||||||
|
- **D82** (PR #54): `dashboard.html` restructure — Claude.ai-style Plan Usage panel above the existing 4 panels. Per-provider rows with utilization bars, reset countdowns, status chips, representative-claim badges, overage chips, "Updated N min ago" labels. 60s `setInterval` with `visibilitychange` pause/resume. Manual refresh button with 2s spam guard. Graceful fallback to legacy `quota` when `quota_v2` absent. Closes v1.x roadmap #8.
|
||||||
|
- **D83** (PR #55): Suite 38 (20 quota-probe unit tests covering all 13-header parse + cache + backoff + 401-refresh + 429-stale + schema_version + 5 doctor status paths) + Suite 39 (8 dashboard rendering smoke tests covering /dashboard 200/401 + key D82 HTML strings). Added 5 test seams to anthropic.mjs. 727 → 755 tests, 0 fail. Fold-in commit added 38j positive-path coverage (38j2: 401 → refresh succeeds → retry 200) per reviewer finding; total 756.
|
||||||
|
- **Close-prep** (PR #56): README § Plan Usage section + § Supported Providers Quota-probe column + dashboard screenshot + `docs/exit-gates/phase-5-e2e.json` live verification artifact. Fold-in commit addressed 3 maintainer accuracy findings (doctor-kind framing / Mistral admin-API acknowledgment / SPOT drift closure via `quota_probe.openai` + `quota_probe.mistral` registry entries).
|
||||||
|
|
||||||
|
### Out of Phase 5 scope (deferred to later)
|
||||||
|
|
||||||
|
- **D84 Mistral probe.** NO-GO per 2026-05-26 spike: no member-key-accessible quota endpoint at `docs.mistral.ai/api`. Re-entry point pinned at `lib/providers/mistral.mjs DL-7`; re-evaluate if Mistral publishes a member-key surface or if OLP deployment posture expands to org-admin scope (Mistral Admin API exists).
|
||||||
|
- **OpenAI / codex probe.** Permanently skipped — `openai/codex` CLI has no public quota API.
|
||||||
|
- **`X-OLP-Cost-USD` per-request header.** Deferred to Phase 6 (depends on per-(provider, model) cost weights table).
|
||||||
|
- **`context_window_exceeded` fallback trigger.** Deferred (trigger condition not yet observed).
|
||||||
|
- **Automated schema-drift detector.** ADR 0013 Rule 5 codifies a procedural runbook (Annual Alignment Audit + `olp doctor` probe-failure + manual maintainer attention at major `claude --version` bumps), not an automated alarm.
|
||||||
|
|
||||||
|
### Authority cited
|
||||||
|
|
||||||
|
ALIGNMENT.md Rules 1 + 2 + 5; CLAUDE.md release_kit (Phase 5 close trigger); ADR 0012 § Exit gate; ADR 0013 Rule 5 schema-drift protocol; OCP `server.mjs:842-1109` as port reference; live `/v1/messages` probe transcripts captured 2026-05-26 from PI231 (D79 audit) + MacBook (D80 + Phase 5 close-prep E2E); audit memory at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`.
|
||||||
|
|
||||||
|
## v0.4.4 — 2026-05-26
|
||||||
|
|
||||||
|
### D78 — `bin/olp-connect` stale-strings cleanup + README CDN-safe URL + repo-visibility flip
|
||||||
|
|
||||||
|
Patch release on top of v0.4.3. Three small issues caught when running `olp-connect` for real on MacBook (D77 client-install verification):
|
||||||
|
|
||||||
|
- **G11 fix (repo visibility).** Repo `dtzp555-max/olp` flipped from PRIVATE → PUBLIC during this session, closing the original G11 finding (`bash <(curl -fsSL .../main/bin/olp-connect)` returned 404 because anonymous curl can't fetch from private repos). README's `/main/` URL works going forward; GitHub's raw CDN may serve a stale 404 for `/main/` for ~5-15min after the visibility flip due to negative caching. D78 defends against this by adding a **tag-pinned URL (`/v0.4.4/bin/olp-connect`) as the primary recommendation in README**, with `/main/` listed as an alternative for trusted-head users. Tag-pinned URLs bypass the negative-cache because the tag ref was never queried while the repo was private.
|
||||||
|
- **G12 fix (`detect_openclaw` claimed plugin not shipped).** `bin/olp-connect`'s OpenClaw detection block said `"The OpenClaw OLP plugin (D71-D73) is NOT YET SHIPPED"` — but D71-D73 shipped `olp-plugin/` at v0.4.0. D78 replaces the stale text with real install instructions: `git clone` + `openclaw plugins install ./olp-plugin/` (or symlink), edit `~/.openclaw/openclaw.json` with a dedicated bot apiKey, restart gateway. Points at `docs/integrations/openclaw.md` for the full setup.
|
||||||
|
- **G13 fix (`olp-connect` self-version hardcoded literal).** Pre-D78 the script declared `OLP_CONNECT_VERSION="0.4.0-phase4"` as a hardcoded literal that nobody updated through v0.4.1 / v0.4.2 / v0.4.3 (the maintain-the-literal-per-release pattern is reliably forgotten). D78 derives the version at runtime from the sibling `package.json` via python3 — when the script is invoked from a checked-out repo, version resolves to the actual `package.json` value; when invoked via `curl … | bash` with no on-disk package.json next to it, falls back to `unknown`. Now `bash bin/olp-connect --version` prints `olp-connect 0.4.4` automatically with no manual touch needed at the next release.
|
||||||
|
|
||||||
|
**Pre-publish audit.** Per `~/.cc-rules/docs/guides/pre-publish-audit.md` checklist (2026-05-26 session, before the visibility flip):
|
||||||
|
- Identity scrub: 0 hits (no personal names / hostnames / home paths / personal emails leaked into the working tree)
|
||||||
|
- Credential scrub: 0 real tokens — all `olp_` matches are placeholder (`olp_XXXX...`) or test fixtures (`olp_not-a-real-key-...`); gitleaks: "no leaks found"
|
||||||
|
- Git-history author emails: 78 commits, two emails (`dtzp555@gmail.com` local + `taodeng1977@gmail.com` GitHub-account squash-merges). Maintainer chose Option A (accept) — the GitHub-account email was already verified-public on the maintainer's GitHub profile, so the visibility flip exposes nothing new.
|
||||||
|
|
||||||
|
**Test count:** 717 (v0.4.3) → 720 (v0.4.4). +3 D78 regression tests in Suite 36:
|
||||||
|
- 36v — pins absence of `NOT YET SHIPPED` text + presence of real install path
|
||||||
|
- 36w — pins runtime version derivation from package.json (hardcoded literal gone)
|
||||||
|
- 36x — pins README's tag-pinned-URL recommendation
|
||||||
|
|
||||||
|
**Authority:** D77 MacBook client-install verification session (2026-05-26); `~/.cc-rules/docs/guides/pre-publish-audit.md`. Process learning: every README that includes a `curl <raw-URL> | bash` install pattern should pin to a release tag (not `/main/`) for CDN-cache resilience. The /main/ form is correct for the long-tail (when no negative cache exists) but the tag-pinned form survives the visibility-flip transient + survives any future force-push to main.
|
||||||
|
|
||||||
|
**Out of D78 scope:**
|
||||||
|
- F6 (doctor client-side vs server-side check separation) — Phase 5 ADR amendment.
|
||||||
|
- D75 reviewer P2-1 (ADR 0004 per-hop schema amendment) + P2-2 (defensive `typeof hopModel === 'string'` invariant) — both genuine follow-ups, neither blocking.
|
||||||
|
- `scripts/migrate-from-ocp.mjs` — Phase 7.
|
||||||
|
|
||||||
|
## v0.4.3 — 2026-05-26
|
||||||
|
|
||||||
|
### D76 — README install-path overhaul + `OLP_BIND` env + AI-driven install prompt + ADR 0011 amendment
|
||||||
|
|
||||||
|
Patch release closing the install-experience gap. v0.4.0–v0.4.2 README's Quick Start was placeholder text with fictional commands (`npm install -g @dtzp555-max/olp` — package isn't published; `olp setup` / `olp start` — don't exist). 10 real gaps catalogued + fixed in one D-day; `OLP_BIND` env wired so the documented LAN onboarding flow actually works; AI-driven install prompt added per the Phase 4 charter brainstorm's #2 OCP inheritance candidate (was deferred at D64-D67 to the doctor framework only; D76 closes the README half).
|
||||||
|
|
||||||
|
- **G1-G7 (README "Quick Start" was fictional)** — rewrote § "Manual install" with the real sequence: prerequisites (Node ≥ 18 + provider CLI install matrix) → `git clone` → `npm test` verify → `olp-keys keygen --owner` first → provider OAuth (claude/codex/mistral per-CLI flows) → write `~/.olp/config.json` with the minimum that actually serves traffic → `npm start` → smoke-test → IDE pointing. Each step empirically verified against the PI231 + Mac mini E2E session (2026-05-26).
|
||||||
|
- **G8 (LAN unreachable — F5)** — added `OLP_BIND` env (default `127.0.0.1`). Operators set `OLP_BIND=0.0.0.0` (or a specific LAN IP) to accept LAN connections so `olp-connect <ip>` can actually reach the server. Pre-D76 the server was hard-coded to `server.listen(PORT, '127.0.0.1', ...)`, making the documented LAN-onboarding flow only usable through an SSH tunnel. ADR 0011's original wording referenced a `BIND_ADDRESS` concept that didn't exist; D76 makes it operational.
|
||||||
|
- **G10 (no AI-install pattern)** — README § "Install with your AI (the fast path)" added. Verbatim prompt that the operator pastes into Claude Code / Cursor / Copilot / Aider; the AI follows the README + uses `olp doctor --json` machine-readable `next_action.ai_executable[]` (D64-D67) for self-repair, stopping only when `human_required[]` is non-empty (the provider OAuth dances). This closes the Phase 4 brainstorm Top-5 inheritance candidate #2 — the OCP "paste this prompt" pattern that D64-D67 only half-built.
|
||||||
|
- **Opening compressed** — § "Why OLP" (3 paragraphs of OCP billing history) removed from the top. The OCP-trigger context moved to § "Migration from OCP" at the bottom, condensed into a single paragraph. New users land on value-prop + § "What you get" + § "Install with your AI" / § "Manual install" without needing to digest 2026-05-14 / 2026-06-15 Anthropic billing history first. OCP users get a one-line pointer at the top.
|
||||||
|
- **§ "Configuration" full schema documentation** — replaced the placeholder with the actual `~/.olp/config.json` schema including every field that v0.4.x reads. Cross-references ADR 0004/0007/0010/0011.
|
||||||
|
- **§ "Environment Variables" extended** — added `OLP_BIND`, `OLP_API_KEY`, `OLP_OWNER_TOKEN`, `OLP_PROXY_URL` rows that were used throughout the manual-install flow but undocumented.
|
||||||
|
|
||||||
|
**ADR 0011 § "Deployment configurations" amendment.** Codifies the three deployment trust contexts (`127.0.0.1` loopback / RFC1918 + tailnet LAN / `0.0.0.0` public — with `advertise_anonymous_key: true` only safe in the first two). Documents the new `anonymous_key_advertised_with_lan_bind` startup warn event. Closes ADR 0011's pre-D76 dangling reference to a non-existent `BIND_ADDRESS`.
|
||||||
|
|
||||||
|
**Test count:** 714 (v0.4.2) → 717 (v0.4.3). +3 D76 regression tests in Suite 36 (36s/36t/36u) pinning `OLP_BIND` wiring + safety warn + ADR amendment.
|
||||||
|
|
||||||
|
**Out of D76 scope (deferred):**
|
||||||
|
- F6 (doctor client-side vs server-side check separation) — needs design ADR for a `--remote` mode. Phase 5.
|
||||||
|
- D75 reviewer P2-1 (ADR 0004 amendment for per-hop schema) + P2-2 (defensive `typeof hopModel === 'string'`) — both genuine follow-ups, neither blocking.
|
||||||
|
- `scripts/migrate-from-ocp.mjs` — Phase 7.
|
||||||
|
|
||||||
|
**Authority:** PI231 + Mac mini E2E session (2026-05-26, post-v0.4.2 verification revealed the 10 README gaps); ADR 0011 amendment self-cites; Phase 4 charter (ADR 0010) Top-5 inheritance candidate #2 (AI-driven self-repair). Process learning: every D-day reviewer rubric should add "open README in §-Quick-Start and verify the commands literally exist + work in the current repo" — would have caught G1-G7 at v0.4.0.
|
||||||
|
|
||||||
|
## v0.4.2 — 2026-05-26
|
||||||
|
|
||||||
|
### Post-v0.4.1 hotfix batch (D75) — real-machine E2E findings
|
||||||
|
|
||||||
|
Patch release fixing 5 bugs caught by **real-machine E2E testing on PI231 + Mac mini (2026-05-26 session)** — bugs that prior D-day reviewers AND the post-v0.4.0 maintainer review both missed because they reviewed against spec text and against the local OLP install's `~/.codex/auth.json` shape (cached from an older codex CLI version), not against real provider CLIs running on a remote operator host that did `npm install -g @openai/codex` for the first time on 2026-05-26 and got codex CLI v0.133.0.
|
||||||
|
|
||||||
|
**Root cause of the missed-bug class.** D6 (codex plugin authoring) explicitly documented three unpinned assumptions (A3 = access-token field name, A4 = NDJSON event schema, A2-adjacent = trusted-directory sandbox). D6 noted "D7 E2E will pin." D7 then shipped without performing real-codex-CLI E2E (the E2E gating mark was carried but the actual run was deferred). Every subsequent D-day reviewer trusted the D6/D7 codex plugin code unchanged because the static review couldn't see that the v0.133.0 CLI had moved the auth-token field, the event schema, AND added a new trusted-directory sandbox flag. The D74 maintainer review focused on `/health` / `/cache/stats` / `/v0/management/dashboard-data` payload shapes — none of which exercise the codex plugin's spawn path. F7 (per-hop model override) is a different class of miss — every reviewer read `executeHopFn(provider, model, ir)` and saw `model` consumed for cache key + audit ctx, but none traced through to confirm `model` is ALSO substituted into the IR passed to `provider.spawn()`. The function signature implied per-hop semantics that the body never fully delivered.
|
||||||
|
|
||||||
|
- **[F1] codex auth.json schema pin — codex CLI v0.133.0 nests the access token under `tokens.access_token`** (verified empirically on PI231 / Mac mini, 2026-05-26). Pre-D75 `readAuthArtifact()` read only top-level `creds.access_token` / `creds.token` / `creds.accessToken` — all undefined under v0.133.0 → returned `null` → OLP reported "auth artifact missing" via `/health` and `olp doctor` AND refused to spawn codex even when the user had fully completed `codex login`. Fix: prepend `creds?.tokens?.access_token` to the precedence chain at BOTH call sites (`OPENAI_CODEX_AUTH_PATH` override branch + default `$CODEX_HOME/auth.json` branch). Legacy top-level fields preserved as fallback for backward compat with older codex CLI versions.
|
||||||
|
- **[F2] codex spawn args — codex CLI v0.133.0 trusted-directory sandbox requires `--skip-git-repo-check`.** v0.133.0 refuses with `"Not inside a trusted directory and --skip-git-repo-check was not specified."` when spawned outside a git repo, exits non-zero with zero NDJSON output → OLP surfaces `SPAWN_FAILED` with no usable chunks → the fallback engine advances to next hop unnecessarily even when codex is configured and authenticated. OLP's typical deploy CWD (`~/olp/`) is NOT a git repo on operator hosts. Fix: add `'--skip-git-repo-check'` to the args array before `--model`. OLP is the trusted caller (operator's own server invoking the operator's own subscription via documented `codex exec` automation); the sandbox safeguards interactive shells, not pre-authorized automation.
|
||||||
|
- **[F3] codex NDJSON event shape pin — codex CLI v0.133.0 emits `item.completed` + `turn.completed` + `turn.failed`**, not the D6-assumed `content`/`delta`/`text` + `type:'stop'`/`done:true` shapes. Real v0.133.0 stream (verified empirically): `{"type":"thread.started",...}` → `{"type":"turn.started"}` → `{"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"<response>"}}` → `{"type":"turn.completed","usage":{...}}`. Pre-D75, every chunk was silently dropped by `codexChunkToIR()` → response body had `content: null`. Fix: prepend three new recognizers (`item.completed` with `item.type === 'agent_message'` → IR delta; `turn.completed` → IR stop; `turn.failed` → IR error). Legacy D6 defensive recognizers preserved below as forward/backward compat fallbacks.
|
||||||
|
- **[F4] `olp status` reads `body.stats.cache.size` (not OCP-era `body.cache.entries`).** Same class as D74 P2-3 (which fixed `cmdUsage` + `cmdCache`); D74 missed the parallel bug in `cmdStatus`. Server payload nests cache stats as `body.stats.cache.{hits, misses, size, inflightCount}` per `server.mjs handleManagementStatus`, and `CacheStore.stats()` has no `entries` field per `lib/cache/store.mjs`. Pre-D75 output showed `entries=?`. Fix: read `c.size` for entries display; also surface `inflightCount` when present.
|
||||||
|
- **[F7] per-hop chain `model` field now overrides IR model in `provider.spawn()`.** Pre-D75 `executeHopFn(hopProvider, hopModel, irReq)` used `hopModel` for cache key + audit ctx but passed the ORIGINAL `irReq` (with `irReq.model` = user's original request) to `hopProviderPlugin.spawn(irReq, authContext)`. A chain config `[{provider:anthropic, model:claude-X}, {provider:openai, model:gpt-5.5}]` would always spawn BOTH plugins with `--model claude-X` — openai rejected the unknown model and the chain died. This broke the core OLP value prop (cross-provider fallback with provider-appropriate model substitution per hop). Fix: build a per-hop IR variant with `{...irReq, model: hopModel}` and pass that to spawn. Conditional skips clone when `hopModel === irReq.model` (common case: single-provider chains, or single-hop chains where the chain config repeats the request model). Applied to BOTH the buffered path (`executeHopFn`) AND the streaming path (`sourceFactory` for `getOrComputeStreaming`). **Authority:** ADR 0004 § Chain advancement step 1 (per-hop config supplies provider AND model — the contract was always specified, but the code didn't complete it).
|
||||||
|
|
||||||
|
**Phase 5 process learning recorded.** Every provider plugin's D-day must include a real-CLI E2E ON A REMOTE OPERATOR HOST before merging — not on the maintainer workstation (which may have an older CLI cached from a prior install, hiding new field renames / new sandbox flags / new event shapes). The D6/D7 codex E2E was deferred and that deferral compounded across 3 layers (D6 = unpinned, D7 = pinning deferred, D8+ = trusted D6/D7 unchanged). F7 reinforces a separate lesson: when a function signature takes `(provider, model, ir)`, reviewers must check that `model` is consumed everywhere downstream, not just at the call site they happened to look at.
|
||||||
|
|
||||||
|
**Out of D75 scope (deferred to Phase 5 explicit ADR amendments):**
|
||||||
|
- F5 (server bind 127.0.0.1 / `OLP_BIND` env) — needs `lib/keys.mjs` anonymous-key trust boundary review before binding to non-loopback by default
|
||||||
|
- F6 (`olp doctor` client-vs-server-side limit detection) — needs design ADR amendment for trigger taxonomy
|
||||||
|
|
||||||
|
- **Test count delta:** 704 (v0.4.1) → 714 (v0.4.2). +10 D75 regression tests in Suite 36 (36i through 36r).
|
||||||
|
- **Files touched:** `lib/providers/codex.mjs` (F1+F2+F3), `bin/olp.mjs` (F4 cmdStatus), `server.mjs` (F7 buffered + streaming spawn sites), `test-features.mjs` (Suite 36 extension), `package.json` (version), `CHANGELOG.md` (this entry).
|
||||||
|
- **Authority:** ADR 0002 (provider contract — codex plugin), ADR 0004 (fallback engine — per-hop model contract), `lib/providers/codex.mjs` D6 assumption A2/A3/A4 docstrings (which all said "D7 will pin" and D7 never did); codex CLI v0.133.0 on-disk schema + `codex exec --help` output verified empirically on PI231 (2026-05-26 E2E session); Iron Rule 第二律 evidence-over-should-work; CLAUDE.md `release_kit.phase_rolling_mode` cross-Phase discipline.
|
||||||
|
|
||||||
|
## v0.4.1 — 2026-05-26
|
||||||
|
|
||||||
|
### Post-Phase-4 hotfix batch (D74) — maintainer-review findings
|
||||||
|
|
||||||
|
Patch release fixing 5 issues caught by maintainer post-v0.4.0 independent review. Every finding was a real runtime bug that the per-D-day fresh-context opus reviewers all missed because they reviewed against spec text, not against the runtime contract (default `auth.allow_anonymous: false`, real `/health` payload shape, real `/cache/stats` payload shape, real `/v0/management/dashboard-data` payload shape). **Phase 4 lesson: future implementation D-days MUST include at least one test that boots the server with the default production config and exercises the new feature end-to-end** — not just stub-mocked codepaths.
|
||||||
|
|
||||||
|
- **[P1-1] `olp doctor` no longer false-negatives on auth-required `/health`.** `lib/doctor.mjs` now accepts an `authHeaders` option (threaded from `bin/olp.mjs` `cmdDoctor` via the existing `authHeaders()` chain) and passes it to the `server.running` + `server.version` probes. The `server.running` check now distinguishes 401/403 ("server up but bearer token missing/invalid — set `OLP_API_KEY`") from "server unreachable" — so the `kind` discriminator routes to a clean fix-auth path instead of `fix_server` when the operator just forgot to export the env var.
|
||||||
|
- **[P1-2] `bin/olp-connect` validates token shape + shell-quotes rc writes.** New `validate_olp_token <key> <source>` helper enforces the `^olp_[A-Za-z0-9_-]{43}$` regex (per ADR 0007 § 3 token format) at all 3 input sites: `--key` arg, `/health.anonymousKey` server-advertised consumption, and the interactive prompt fallback. New `shell_quote <value>` helper wraps rc-file writes (`export OPENAI_BASE_URL=$(shell_quote ...)`) so even a hypothetical bypass of the validator can't inject shell metacharacters into a sourced rc. systemd `environment.d/olp.conf` write additionally rejects embedded newlines. Hostile or malformed keys can no longer persist as shell startup injection.
|
||||||
|
- **[P2-3] `olp usage` + `olp cache` human formatter rewritten against the real payload shape.** `cmdUsage` previously read `body.usage_24h.requests` / `body.providers` / `body.top_fallback_chains` — all undefined under the actual server payload shape — so users saw "requests: ?" + missing per-provider quota + missing top-chains. Now reads `body.window_24h.request_count` / `body.cache_hit_24h.hit_rate` / `body.quota` / `body.top_fallback_chains_24h` per `server.mjs:2027` + `lib/audit-query.mjs`. `cmdCache` previously read `body.entries` / `body.bytes` / `body.maxBytes` (OCP-era field names). Now reads `body.size` / `body.inflightCount` per `CacheStore.stats()` and computes hit rate from `hits + misses`.
|
||||||
|
- **[P2-4] `olp-plugin/` `fmtHealth` iterates `providers.status` correctly.** Previously walked `Object.entries(body.providers)` which surfaced `enabled` / `available` / `status` as pseudo-providers (chat output showed `🟢 status` instead of `🟢 anthropic`). Now extracts the real provider map from `body.providers.status` and renders enabled/available counts in a header line + per-provider names with `activeSpawns` when present. Falls back to flat `body.providers.*` for the older OCP shape (backwards compat).
|
||||||
|
- **[P3-5] Stale v0.3.0-era doc strings updated.** README header status line + Implementation Status § now reflect v0.4.0 shipped + Phase 5 open. `server.mjs` startup banner no longer hardcodes "Phase 1 in progress" (now just lists version + provider count — derives accurate state from `VERSION` without future maintenance touch-ups).
|
||||||
|
|
||||||
|
**Phase 4 process learning recorded.** Per Iron Rule 第二律 (evidence over "should work"), every D-day review pass must include at least one runtime smoke against the default production config. The D-day reviewer rubric is updated implicitly — D74 Suite 36 tests pin the wire-contract shape so a future D-day refactoring server payloads can't silently re-break the CLI / plugin / docs.
|
||||||
|
|
||||||
|
- **Test count delta:** 696 (v0.4.0) → 704 (v0.4.1). +8 D74 regression tests in Suite 36.
|
||||||
|
- **Files touched:** `lib/doctor.mjs` (P1-1), `bin/olp.mjs` (P1-1 + P2-3), `bin/olp-connect` (P1-2), `olp-plugin/index.js` (P2-4), `server.mjs` (P3-5 banner), `README.md` (P3-5), `test-features.mjs` (Suite 36 regression), `package.json` (version), `CHANGELOG.md` (this entry).
|
||||||
|
- **Authority:** maintainer independent review of `main` / `v0.4.0` / commit `ee4d945` (2026-05-26 session); Iron Rule 第二律 evidence-over-should-work; CLAUDE.md `release_kit.phase_rolling_mode` cross-Phase discipline ("hotfix to a shipped Phase N deliverable → bump patch, tag, release before next push").
|
||||||
|
|
||||||
## v0.4.0 — 2026-05-26
|
## v0.4.0 — 2026-05-26
|
||||||
|
|
||||||
|
|||||||
@@ -135,7 +135,7 @@ release_kit:
|
|||||||
# This overlay is the authoritative source. If Iron Rule 5 appears to be silently
|
# This overlay is the authoritative source. If Iron Rule 5 appears to be silently
|
||||||
# violated (no version bump after many D-day pushes), check this section first
|
# violated (no version bump after many D-day pushes), check this section first
|
||||||
# before filing a compliance finding.
|
# before filing a compliance finding.
|
||||||
current_phase: Phase 5
|
current_phase: Phase 6
|
||||||
current_pre_release_identifier: "0.5.0-phase5"
|
current_pre_release_identifier: "0.6.0-phase6"
|
||||||
phase_close_trigger: explicit maintainer action (not automated)
|
phase_close_trigger: explicit maintainer action (not automated)
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -1,76 +1,230 @@
|
|||||||
# OLP — Open LLM Proxy
|
# OLP — Open LLM Proxy
|
||||||
|
|
||||||
A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing, automatic fallback, content-addressed caching — so your IDEs and family clients keep working as long as *any* of your subscriptions has quota left.
|
A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing + fallback + content-addressed caching. Your IDEs and family clients keep working as long as **any** of your subscriptions has quota left.
|
||||||
|
|
||||||
> **Status:** v0.3.0 shipped (2026-05-25) — Phase 1 multi-provider proxy core (v0.1.0 + v0.1.1) + Phase 2 multi-key auth + audit + owner gating + keygen CLI (v0.2.0) + Phase 3 Dashboard + audit query layer + daily audit rotation (v0.3.0). Phase 4 (per-key per-provider auth + audit retention + SQLite hybrid + provider-cost weights) is the next planned milestone. Sections marked _placeholder_ land alongside the relevant phase of work (see [phase plan](#phase-plan)).
|
> **Status:** v0.5.1 shipped, 759+ tests. Phase 5 (Quota Probes + Dashboard Enrichment) closed; Phase 6 next. Coming from [OCP](https://github.com/dtzp555-max/ocp)? See [§ Migration from OCP](#migration-from-ocp).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Why OLP
|
## What you get
|
||||||
|
|
||||||
On 2026-05-14, Anthropic announced (effective 2026-06-15) that `claude -p`, the Agent SDK, and third-party agent traffic move out of the Pro/Max subscription pool into a separate fixed monthly Agent SDK Credit pool. [OCP](https://github.com/dtzp555-max/ocp), OLP's predecessor, was a proxy around a single CLI — its core assumption was *"subscription = unlimited within rate limits"*. That assumption breaks for Anthropic on the effective date.
|
- **OpenAI-compatible** `/v1/chat/completions` endpoint — any IDE that speaks OpenAI (Cline / Continue.dev / Cursor / Aider) plugs in
|
||||||
|
- **Multi-provider chain** — primary fails / quota dies → automatically falls back to the next provider (anthropic ↔ codex ↔ mistral by default; risk-tier framework guards which ones get enabled)
|
||||||
The structural response is to stop relying on one provider's subscription terms remaining favourable. OLP spreads risk across multiple providers whose subscriptions still include CLI/programmatic use, routes intelligently between them, and caches aggressively so every request that does spawn a CLI counts.
|
- **Content-addressed cache** — repeat requests don't re-spawn the CLI; streaming requests dedup via singleflight tee
|
||||||
|
- **Multi-key auth** — owner key with full visibility, family-member keys with per-key audit log + per-provider scoping
|
||||||
OLP is **not**: a commercial multi-tenant SaaS; an enterprise gateway competing with LiteLLM / OpenCode / CLIProxyAPI on breadth; a model-capability router ("route to the smartest model" — you pick the model); a conversation-state store (your client handles that).
|
- **Telegram / Discord** `/olp` slash commands (read-only — for "is OLP up?" checks from anywhere)
|
||||||
|
- **AI-driven self-repair** — `olp doctor --json` emits machine-readable `next_action.ai_executable[]` so a Claude Code / Cursor / Copilot session can fix install issues for you (see [§ Install with your AI](#install-with-your-ai-the-fast-path))
|
||||||
See [`ALIGNMENT.md`](./ALIGNMENT.md) for OLP's constitution and [`docs/adr/`](./docs/adr/) for the founding ADRs.
|
- **Observability** — owner-only `/dashboard` (live Claude.ai-style plan-usage rows / 24h stats / 30d spend trend / top fallback chains)
|
||||||
|
- **Plan-usage probe** (Phase 5, v0.5.0) — opt-in per-provider quota probe for Anthropic Pro/Max subscriptions; parses the canonical `anthropic-ratelimit-unified-*` response headers, surfaces 5-hour + 7-day utilization with reset countdowns. See [§ Plan Usage](#plan-usage-live-quota-probe).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Quick Start
|
## Install with your AI (the fast path)
|
||||||
|
|
||||||
_placeholder — lands with Phase 1._
|
If the manual steps feel like a lot, paste this verbatim into your AI coding assistant (Claude Code / Cursor / Copilot / Aider). It walks you through everything:
|
||||||
|
|
||||||
Anticipated shape:
|
```
|
||||||
|
I want to install OLP (Open LLM Proxy) on this machine. The repo is at
|
||||||
|
https://github.com/dtzp555-max/olp and ships these tools that you can use:
|
||||||
|
|
||||||
```bash
|
1. `git clone` the repo to ~/olp
|
||||||
# install
|
2. README.md § "Manual install" has the sequenced steps. Follow them in order.
|
||||||
npm install -g @dtzp555-max/olp
|
3. After EACH step, run `node ~/olp/bin/olp.mjs doctor --json` (NOT `npx olp`
|
||||||
|
— the package isn't on npm). Parse the JSON. The `kind` field tells you
|
||||||
|
the state: noop / fresh_install / fix_server / fix_oauth / fix_provider /
|
||||||
|
fix_config.
|
||||||
|
4. If `kind != noop`, run the commands listed in `next_action.ai_executable[]`
|
||||||
|
verbatim. Then re-run doctor to verify.
|
||||||
|
5. STOP and ask me only when `next_action.human_required[]` is non-empty.
|
||||||
|
That's where I need to do a browser OAuth flow you can't do for me.
|
||||||
|
|
||||||
# run setup (writes ~/.olp/config.json, asks which providers to enable)
|
The provider CLIs OLP spawns (claude / codex / vibe) need their own one-time
|
||||||
olp setup
|
OAuth — those are the only steps I personally have to do (Claude.ai login,
|
||||||
|
ChatGPT login, Mistral API key). Everything else (clone, npm install of the
|
||||||
|
provider CLIs, owner-key generation, config.json bootstrap, server start) is
|
||||||
|
in your `ai_executable[]` and you should run it without asking.
|
||||||
|
|
||||||
# start the proxy (default port 4567 since v0.4.0 — moved off OCP's 3456 so
|
Begin.
|
||||||
# OLP and OCP can co-host on the same machine. Set OLP_PORT=3456 if you have
|
|
||||||
# no OCP on the machine and want the old default.)
|
|
||||||
olp start
|
|
||||||
|
|
||||||
# point your IDE at http://localhost:4567/v1/chat/completions with the OLP API key from `olp keys list`.
|
|
||||||
```
|
```
|
||||||
|
|
||||||
**Family-on-LAN onboarding (D68-D70).** For other devices on the same network, run on the client device:
|
Then sit back and respond when it asks for OAuth confirmation. This pattern works because `olp doctor` is purpose-built for AI consumption — every failure mode has a shell-executable repair command AND a human-required step listed separately.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Manual install (5-10 min)
|
||||||
|
|
||||||
|
### 0. Prerequisites
|
||||||
|
|
||||||
|
- **Node.js ≥ 18.** Verify: `node --version`
|
||||||
|
- **The provider CLIs you want OLP to spawn.** Install whichever you'll actually use:
|
||||||
|
|
||||||
|
| Provider | Install | Subscription |
|
||||||
|
|---|---|---|
|
||||||
|
| `anthropic` (`claude -p`) | `npm install -g @anthropic-ai/claude-code` | Claude Pro/Max (OAuth) |
|
||||||
|
| `openai` (`codex exec`) | `npm install -g @openai/codex` | ChatGPT Plus/Pro (OAuth) or OpenAI API key |
|
||||||
|
| `mistral` (`vibe --prompt`) | follow the `vibe` install docs | Le Chat Pro API key |
|
||||||
|
|
||||||
|
You only need to install the ones you'll route to. Single-provider OLP works fine.
|
||||||
|
|
||||||
|
### 1. Clone and verify the test suite
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Detects Cline / Continue.dev / Cursor / Aider / OpenClaw installed locally
|
git clone https://github.com/dtzp555-max/olp.git ~/olp
|
||||||
# and writes per-tool config pointing at the OLP host. Requires `python3`.
|
cd ~/olp
|
||||||
olp-connect <olp-host-ip>
|
npm test # 714+ tests, ~5s, no external deps
|
||||||
```
|
```
|
||||||
|
|
||||||
If the OLP host has `auth.advertise_anonymous_key: true` AND a key was created with `olp-keys keygen --anonymous --advertise`, `olp-connect` picks up the token from `/health.anonymousKey` — zero out-of-band token paste required. See [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md) for the trusted-LAN-only invariant.
|
(If `npm test` fails here, stop — that means your Node version or the repo state is broken. Don't proceed to step 2.)
|
||||||
|
|
||||||
Per-IDE setup details: [`docs/integrations/`](./docs/integrations/README.md).
|
### 2. Bootstrap the owner key
|
||||||
|
|
||||||
|
The owner key is what you (and `olp-connect`) use to authenticate to OLP. Default config has `auth.allow_anonymous: false`, so you need a key BEFORE the server starts accepting requests.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
node ~/olp/bin/olp-keys.mjs keygen --owner --name=$(whoami)-laptop
|
||||||
|
# Prints the plaintext token ONCE. Copy it now — you can't recover it later.
|
||||||
|
# Example: olp_l23-PN46tDljmPATV94-KfOgOBO0Ed8theVjTdAgQoY
|
||||||
|
```
|
||||||
|
|
||||||
|
Export it so the CLI subcommands can use it:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export OLP_API_KEY=olp_l23-PN46... # paste your real token
|
||||||
|
```
|
||||||
|
|
||||||
|
(Add to `~/.bashrc` / `~/.zshrc` to persist.)
|
||||||
|
|
||||||
|
### 3. Authenticate the providers (one-time OAuth)
|
||||||
|
|
||||||
|
Run each provider's own login flow. OLP's anthropic / openai / mistral plugins spawn these CLIs and reuse their cached credentials — OLP itself never touches the OAuth dance.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Anthropic (Claude Pro/Max subscription)
|
||||||
|
claude setup-token
|
||||||
|
# Opens a TUI / prints a URL. Authorize in browser. Paste the returned code.
|
||||||
|
# Result: ~/.claude/.credentials.json
|
||||||
|
|
||||||
|
# OpenAI (ChatGPT subscription)
|
||||||
|
codex login --device-auth
|
||||||
|
# Prints a https://auth.openai.com/codex/device URL + 10-char code.
|
||||||
|
# Open URL in browser, enter code, authorize.
|
||||||
|
# Result: ~/.codex/auth.json
|
||||||
|
|
||||||
|
# Mistral (Le Chat API key)
|
||||||
|
export MISTRAL_API_KEY=sk-... # add to ~/.bashrc to persist
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. Write a minimum config
|
||||||
|
|
||||||
|
`~/.olp/config.json`:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"auth": {
|
||||||
|
"allow_anonymous": false,
|
||||||
|
"owner_only_endpoints": [
|
||||||
|
"/health",
|
||||||
|
"/v0/management/dashboard-data",
|
||||||
|
"/v0/management/quota",
|
||||||
|
"/v0/management/status",
|
||||||
|
"/cache/stats",
|
||||||
|
"/dashboard"
|
||||||
|
],
|
||||||
|
"fallback_detail_header_policy": "owner_only"
|
||||||
|
},
|
||||||
|
"providers": {
|
||||||
|
"enabled": { "anthropic": true, "openai": true }
|
||||||
|
},
|
||||||
|
"routing": {
|
||||||
|
"chains": {
|
||||||
|
"claude-sonnet-4-6": [
|
||||||
|
{ "provider": "anthropic", "model": "claude-sonnet-4-6" },
|
||||||
|
{ "provider": "openai", "model": "gpt-5.5" }
|
||||||
|
],
|
||||||
|
"gpt-5.5": [{ "provider": "openai", "model": "gpt-5.5" }]
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"streaming": { "heartbeat_interval_ms": 15000 }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
(Enable only the providers you actually authenticated in step 3. Chains map `<your-IDE's-requested-model>` → ordered list of `{provider, model}` hops; the chain's per-hop `model` is what gets passed to that provider's CLI.)
|
||||||
|
|
||||||
|
### 5. Start the server
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd ~/olp
|
||||||
|
npm start
|
||||||
|
# OLP v0.4.3 listening on :4567 (2 providers enabled)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 6. Smoke-test
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -H "Authorization: Bearer $OLP_API_KEY" http://localhost:4567/health | jq
|
||||||
|
# Expect: {ok: true, providers: {enabled: 2, status: {anthropic: {ok: true...}, openai: {ok: true...}}}}
|
||||||
|
|
||||||
|
node ~/olp/bin/olp.mjs doctor
|
||||||
|
# Expect: "9 of 9 checks passed", kind=noop
|
||||||
|
```
|
||||||
|
|
||||||
|
### 7. Point your IDE at OLP
|
||||||
|
|
||||||
|
```
|
||||||
|
OPENAI_BASE_URL=http://localhost:4567/v1
|
||||||
|
OPENAI_API_KEY=$OLP_API_KEY
|
||||||
|
```
|
||||||
|
|
||||||
|
Per-IDE configuration details: [`docs/integrations/`](./docs/integrations/README.md).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Family / LAN setup
|
||||||
|
|
||||||
|
To let other devices on your home network use the same OLP server, you need TWO things:
|
||||||
|
|
||||||
|
1. **Bind to the LAN interface** (not just loopback). On the SERVER:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
OLP_BIND=0.0.0.0 npm start # or your specific LAN IP, e.g. 192.168.1.10
|
||||||
|
```
|
||||||
|
|
||||||
|
Default is `127.0.0.1` (loopback only). See [ADR 0011 § Deployment configurations](./docs/adr/0011-anonymous-key-deployment-context.md#deployment-configurations-d76-amendment-2026-05-26) for the trust-context table — **never set `OLP_BIND=0.0.0.0` on a public-internet-facing host** (use a tunnel like Tailscale instead).
|
||||||
|
|
||||||
|
2. **Onboard each family member's device** from THEIR machine:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Pinned to a known-good release (recommended — survives GitHub raw CDN cache hiccups):
|
||||||
|
bash <(curl -fsSL https://raw.githubusercontent.com/dtzp555-max/olp/v0.4.4/bin/olp-connect) <olp-host-ip>
|
||||||
|
|
||||||
|
# OR latest from main (use after v0.4.4 + once you trust head):
|
||||||
|
bash <(curl -fsSL https://raw.githubusercontent.com/dtzp555-max/olp/main/bin/olp-connect) <olp-host-ip>
|
||||||
|
```
|
||||||
|
|
||||||
|
Detects Cline / Continue.dev / Cursor / Aider / OpenClaw locally and writes per-tool config pointing at your OLP host. Requires `python3` on the client. Prompts for the OLP API key — OR, if the server has `auth.advertise_anonymous_key: true` AND a key was created with `olp-keys keygen --anonymous --advertise`, picks the token up from `/health.anonymousKey` (zero out-of-band paste). See [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md) for the trusted-LAN-only invariant.
|
||||||
|
|
||||||
|
Per-IDE setup details: [`docs/integrations/`](./docs/integrations/README.md). Telegram / Discord `/olp` slash command setup: [§ Telegram / Discord Usage](#telegram--discord-usage).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Supported Providers
|
## Supported Providers
|
||||||
|
|
||||||
Source of truth: [`models-registry.json`](./models-registry.json). This table is regenerated from the registry per the [`release_kit`](./CLAUDE.md) overlay; do not edit it out of sync.
|
Source of truth: [`models-registry.json`](./models-registry.json). Per-provider columns are sourced from the registry's `providers.<key>` block (model metadata + tier) and `quota_probe.<key>` block (D81+; probe status / reason / source). This table is regenerated from the registry per the [`release_kit`](./CLAUDE.md) overlay; do not edit it out of sync.
|
||||||
|
|
||||||
OLP distinguishes **Candidate Providers** (declared as intended, not yet pinned) from **Enabled Providers** (authority pin filled + plugin landed + Phase audit passed). The v0.1 founding commit ships **zero Enabled Providers** — enablement is a Phase audit deliverable, not a bootstrap claim. See [`ALIGNMENT.md` § Provider Inventory](./ALIGNMENT.md) for the transition gate.
|
OLP distinguishes **Candidate Providers** (declared as intended, not yet pinned) from **Enabled Providers** (authority pin filled + plugin landed + Phase audit passed). The v0.1 founding commit ships **zero Enabled Providers** — enablement is a Phase audit deliverable, not a bootstrap claim. See [`ALIGNMENT.md` § Provider Inventory](./ALIGNMENT.md) for the transition gate.
|
||||||
|
|
||||||
### Candidate Providers
|
### Candidate Providers
|
||||||
|
|
||||||
| Provider key | CLI | Subscription / auth | Anticipated Tier | Anticipated Phase |
|
| Provider key | CLI | Subscription / auth | Quota probe (v0.5.0+) | Anticipated Tier | Anticipated Phase |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|---|
|
||||||
| `anthropic` | `claude -p` | Pro / Max OAuth (pre-2026-06-15); Agent SDK Credit pool after | D (re-eval post-2026-06-15) | Phase 1 |
|
| `anthropic` | `claude -p` | Pro / Max OAuth (pre-2026-06-15); Agent SDK Credit pool after | ✅ Live (13 `anthropic-ratelimit-unified-*` headers; opt-in via `quota_probe_enabled`) | D (re-eval post-2026-06-15) | Phase 1 |
|
||||||
| `openai` | `codex exec --json` | ChatGPT Pro OAuth or API key | D | Phase 2 |
|
| `openai` | `codex exec --json` | ChatGPT Pro OAuth or API key | ❌ Not available (no public quota API) — audit-derived spend tracking only | D | Phase 2 |
|
||||||
| `mistral` | `vibe --prompt --output json` | Le Chat Pro API key | D | Phase 3 |
|
| `mistral` | `vibe --prompt --output json` | Le Chat Pro API key | ❌ Not implemented at v0.5.0 — no public quota endpoint accessible to Vibe / Le Chat member / La Plateforme API keys per D84 spike 2026-05-26. Mistral's [Admin API](https://docs.mistral.ai/admin/security-access/admin-api) does expose billing / usage queries but requires an org-admin scope (out of scope for OLP family-tier deployment). Audit-derived spend tracking only at v0.5.0. | D | Phase 3 |
|
||||||
| `grok` | `grok -p --output-format streaming-json` | xAI Build `xai-...` API key | C | Phase 8+ |
|
| `grok` | `grok -p --output-format streaming-json` | xAI Build `xai-...` API key | TBD (Phase 8+) | C | Phase 8+ |
|
||||||
| `kimi` | `kimi -p --output-format stream-json` | Moonshot Kimi API key | C | Phase 8+ |
|
| `kimi` | `kimi -p --output-format stream-json` | Moonshot Kimi API key | TBD (Phase 8+) | C | Phase 8+ |
|
||||||
| `minimax` | TBD | MiniMax Token Plan (¥29+/mo) | B | Phase 8+ |
|
| `minimax` | TBD | MiniMax Token Plan (¥29+/mo) | TBD (Phase 8+) | B | Phase 8+ |
|
||||||
| `glm` | TBD | Zhipu Coding Plan ($10+/mo) | B | Phase 8+ |
|
| `glm` | TBD | Zhipu Coding Plan ($10+/mo) | TBD (Phase 8+) | B | Phase 8+ |
|
||||||
| `qwen` | TBD | Alibaba Coding Plan ($50/mo) | B | Phase 8+ |
|
| `qwen` | TBD | Alibaba Coding Plan ($50/mo) | TBD (Phase 8+) | B | Phase 8+ |
|
||||||
|
|
||||||
**Risk tier guide.** D = permissive / safe (eligible for default-enabled); C = tightening signal, no enforcement history (opt-in); B = service-level key revocation risk (opt-in + consent); A = excluded by default (cannot be opt-in enabled). Tier B providers prompt for explicit consent on first enable and record consent in `~/.olp/config.json`. See [`ALIGNMENT.md` § Risk Tier Framework](./ALIGNMENT.md#risk-tier-framework).
|
**Risk tier guide.** D = permissive / safe (eligible for default-enabled); C = tightening signal, no enforcement history (opt-in); B = service-level key revocation risk (opt-in + consent); A = excluded by default (cannot be opt-in enabled). Tier B providers prompt for explicit consent on first enable and record consent in `~/.olp/config.json`. See [`ALIGNMENT.md` § Risk Tier Framework](./ALIGNMENT.md#risk-tier-framework).
|
||||||
|
|
||||||
@@ -80,12 +234,19 @@ OLP distinguishes **Candidate Providers** (declared as intended, not yet pinned)
|
|||||||
|
|
||||||
## Configuration
|
## Configuration
|
||||||
|
|
||||||
_placeholder — full configuration reference lands with Phase 4 (fallback engine)._
|
OLP reads `~/.olp/config.json` at startup. § "[Manual install § Step 4](#4-write-a-minimum-config)" above has a working minimum example. The full schema:
|
||||||
|
|
||||||
OLP reads its config from `~/.olp/config.json`. The minimum useful shape:
|
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
|
"auth": {
|
||||||
|
"allow_anonymous": false,
|
||||||
|
"owner_only_endpoints": ["/health", "/dashboard", "/v0/management/..."],
|
||||||
|
"advertise_anonymous_key": false,
|
||||||
|
"fallback_detail_header_policy": "owner_only"
|
||||||
|
},
|
||||||
|
"providers": {
|
||||||
|
"enabled": { "<provider-key>": true }
|
||||||
|
},
|
||||||
"routing": {
|
"routing": {
|
||||||
"chains": {
|
"chains": {
|
||||||
"<requested-model>": [
|
"<requested-model>": [
|
||||||
@@ -96,13 +257,78 @@ OLP reads its config from `~/.olp/config.json`. The minimum useful shape:
|
|||||||
"soft_triggers": {
|
"soft_triggers": {
|
||||||
"<provider-key>": { "<trigger>": <threshold> }
|
"<provider-key>": { "<trigger>": <threshold> }
|
||||||
}
|
}
|
||||||
|
},
|
||||||
|
"streaming": {
|
||||||
|
"heartbeat_interval_ms": 0
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
> **Note:** `routing.soft_triggers` thresholds are parsed and stored but have **no runtime effect at v0.1** — the quota polling path (`quotaStatus()` per hop) is deferred to v1.x per [ADR 0004 Amendment 2](./docs/adr/0004-fallback-engine.md#amendment-2--2026-05-24-soft-triggers-deferred-to-v1x-d22). The evaluation logic exists and is tested; only the production data ingestion path is deferred.
|
Field guide:
|
||||||
|
|
||||||
Trigger types, fallback safety, idempotency rules, and the full example config land here when Phase 4 ships. See [ADR 0004 (Fallback Engine Semantics & Safety)](./docs/adr/0004-fallback-engine.md) for the design.
|
- **`auth.allow_anonymous`** — default `false`. When false, every request needs a Bearer token; when true, anonymous-tier requests succeed (ADR 0007 § 7). Production posture is `false`.
|
||||||
|
- **`auth.owner_only_endpoints`** — list of endpoints that REQUIRE owner-tier auth (non-owner returns 401). The defaults above are minimum sane for production.
|
||||||
|
- **`auth.advertise_anonymous_key`** — default `false`. When true (+ `allow_anonymous: true` + a key created with `olp-keys keygen --anonymous --advertise`), `/health.anonymousKey` exposes the plaintext token so `olp-connect <ip>` is zero-config. **Trusted-LAN only** — see [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md).
|
||||||
|
- **`auth.fallback_detail_header_policy`** — controls `X-OLP-Fallback-Detail` response header emission. `owner_only` (default) only shows tuples to owner identity; debug surface to LAN family without leaking to anonymous.
|
||||||
|
- **`providers.enabled`** — flip a provider plugin on. Only enable providers whose CLI you've authenticated; OLP doesn't do its own OAuth.
|
||||||
|
- **`routing.chains`** — keyed by the model name your IDE / client requests. Each entry is an ordered list of fallback hops; each hop's `model` is what gets passed to that provider's CLI. F7 fix (D75) — the hop-level `model` field finally overrides the IR's request model during cross-provider fallback.
|
||||||
|
- **`routing.soft_triggers`** — parsed and stored but **inert at v0.4.x** — the `quotaStatus()` polling data path is deferred to v1.x per [ADR 0004 Amendment 2](./docs/adr/0004-fallback-engine.md#amendment-2--2026-05-24-soft-triggers-deferred-to-v1x-d22). Startup emits a warn if non-empty so the inert state is visible.
|
||||||
|
- **`streaming.heartbeat_interval_ms`** — default `0` (disabled). Set > 0 (e.g. `15000`) to emit SSE keepalive frames during silent windows. Required behind reverse proxies (nginx / Cloudflare Tunnel / Tailscale Funnel) with 60s idle aborts.
|
||||||
|
|
||||||
|
See [ADR 0004 (Fallback Engine)](./docs/adr/0004-fallback-engine.md), [ADR 0007 (Multi-key auth)](./docs/adr/0007-multi-key-auth.md), [ADR 0010 (Phase 4 charter)](./docs/adr/0010-phase-4-charter-operator-and-client-ux.md), [ADR 0011 (Anonymous-key deployment)](./docs/adr/0011-anonymous-key-deployment-context.md).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Plan Usage (live quota probe)
|
||||||
|
|
||||||
|
OLP v0.5.0+ surfaces live subscription quota for Anthropic Pro/Max subscribers on the owner-only `/dashboard`. Per-provider rows show 5-hour and 7-day utilization bars with reset countdowns, status badges, representative-claim hints, and a manual refresh button. The panel auto-refreshes every 60 seconds and pauses when the tab is hidden.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### How it works
|
||||||
|
|
||||||
|
The probe issues a minimal `POST /v1/messages` to `api.anthropic.com` (max_tokens: 1) using the same OAuth token Claude Code uses for `claude -p`. The body is discarded; only the 13 `anthropic-ratelimit-unified-*` response headers are parsed (5h/7d utilization + reset, status, representative-claim, fallback-percentage, overage status + disabled reason). Results cache for 5 minutes; refresh failures fall back to the previous cache marked `stale: true` while exponential backoff (60s → 3600s) protects against hammering the API.
|
||||||
|
|
||||||
|
See [ADR 0002 § Amendment 8](./docs/adr/0002-plugin-architecture.md), [ADR 0012 (Phase 5 charter)](./docs/adr/0012-phase-5-charter-quota-probes-dashboard.md), [ADR 0013 (OAuth READ-ONLY consumption + schema-drift mitigation)](./docs/adr/0013-oauth-read-only-consumption-and-schema-drift.md), and the schema pin at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`.
|
||||||
|
|
||||||
|
### Enabling the probe
|
||||||
|
|
||||||
|
The probe is **opt-in** (default off) per ADR 0013 Rule 4 — a fresh OLP install on a machine without OAuth credentials should not bombard `api.anthropic.com` with 401-bound probes. To enable, add to `~/.olp/config.json`:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"providers": {
|
||||||
|
"anthropic": {
|
||||||
|
"enabled": true,
|
||||||
|
"quota_probe_enabled": true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The probe reads the OAuth token from (in order): `CLAUDE_CODE_OAUTH_TOKEN` env var, `~/.claude/.credentials.json`, macOS Keychain entry `"Claude Code-credentials"`. Make sure Claude Code is logged in (`claude setup-token` or equivalent) before opting in.
|
||||||
|
|
||||||
|
`olp doctor` adds a `anthropic.quota_probe_reachable` check when the probe is enabled. The check has `category: 'provider'`, so any failure (401/403 token-expiry, 429 rate-limit, network error) discriminates to `kind: fix_provider`. The `human_steps` recovery recipe inside the check distinguishes the underlying cause (re-login via `claude setup-token` for auth failures vs wait-and-retry for rate-limit) — the discriminator is uniformly `fix_provider` but the actionable text is auth-aware. Successful probes return `status: ok` with the parsed 5h / 7d utilization in the message body; stale-cache returns `status: warn`. Routing an auth-class failure to `kind: fix_oauth` (the other discriminator the framework supports) would require splitting this check across the `provider` / `auth` boundary — deferred to v1.x if `olp doctor` consumers report the ambiguity.
|
||||||
|
|
||||||
|
### Provider coverage
|
||||||
|
|
||||||
|
| Provider | Live quota probe | Path |
|
||||||
|
|---|---|---|
|
||||||
|
| `anthropic` | ✅ Live — 13 fields via `anthropic-ratelimit-unified-*` headers | This section |
|
||||||
|
| `openai` (codex) | ❌ Not available — `openai/codex` CLI has no public quota API | Falls back to audit-derived request counts |
|
||||||
|
| `mistral` | ❌ Not implemented at v0.5.0 — no public quota endpoint accessible to Vibe / Le Chat member / La Plateforme API keys. Mistral's Admin API does expose billing / usage queries but is gated to org-admin scope and out of scope for OLP family-tier deployment. | Falls back to audit-derived request counts |
|
||||||
|
|
||||||
|
If Mistral ever publishes a usage endpoint, `lib/providers/mistral.mjs` DL-7 marks the re-entry point.
|
||||||
|
|
||||||
|
### Schema-drift protection
|
||||||
|
|
||||||
|
Claude Code v2.1.x is distributed as a compiled native binary (Mach-O on macOS, ELF on Linux) — the OCP-era "grep `cli.js`" verification no longer applies. OLP's replacement protocol (ADR 0013 § Rule 5):
|
||||||
|
|
||||||
|
1. `strings` over the platform-specific claude-code binary captures all hardcoded header names the binary expects.
|
||||||
|
2. A live `POST /v1/messages` against `api.anthropic.com` with valid OAuth captures what the server actually emits today.
|
||||||
|
3. Diff path 1 vs path 2 → the actionable schema delta.
|
||||||
|
|
||||||
|
This is re-run at every major `claude --version` bump (next trigger: v2.x → v3.x), at the Annual Alignment Audit (14 May), and whenever `olp doctor anthropic.quota_probe_reachable` returns an unexpected status code. The current pinned schema (13 fields, `2026-05-26`) lives in `models-registry.json` under `quota_probe.schema_version`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -113,9 +339,9 @@ Trigger types, fallback safety, idempotency rules, and the full example config l
|
|||||||
| `/v1/chat/completions` | POST | 1 | ✅ Shipped | OpenAI-compatible Chat Completions entry. Internally normalized to IR, dispatched to a provider plugin, response shape converted back. |
|
| `/v1/chat/completions` | POST | 1 | ✅ Shipped | OpenAI-compatible Chat Completions entry. Internally normalized to IR, dispatched to a provider plugin, response shape converted back. |
|
||||||
| `/v1/models` | GET | 1 | ✅ Shipped | Lists models from `models-registry.json`. |
|
| `/v1/models` | GET | 1 | ✅ Shipped | Lists models from `models-registry.json`. |
|
||||||
| `/health` | GET | 1 | ✅ Shipped | Per-provider health snapshot. Phase 2 owner-only-trim: full per-provider details to owner identity; trimmed `{ ok, version }` to guest / anonymous. Gate via `auth.owner_only_endpoints` config. **Optional `anonymousKey` field (D69 / Phase 4, v0.4.0)** appears in both trimmed and full payloads when `auth.advertise_anonymous_key: true` AND `auth.allow_anonymous: true` AND at least one non-revoked guest-tier key has `plaintext_advertise: true` (see [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md) for the trusted-LAN-only invariant). Default off — field absent when prereqs unmet. |
|
| `/health` | GET | 1 | ✅ Shipped | Per-provider health snapshot. Phase 2 owner-only-trim: full per-provider details to owner identity; trimmed `{ ok, version }` to guest / anonymous. Gate via `auth.owner_only_endpoints` config. **Optional `anonymousKey` field (D69 / Phase 4, v0.4.0)** appears in both trimmed and full payloads when `auth.advertise_anonymous_key: true` AND `auth.allow_anonymous: true` AND at least one non-revoked guest-tier key has `plaintext_advertise: true` (see [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md) for the trusted-LAN-only invariant). Default off — field absent when prereqs unmet. |
|
||||||
| `/dashboard` | GET | 3 | ✅ Shipped (D50 + D51) | Owner-only multi-provider dashboard HTML (4 panels: quota / 24h request stats / 30d spend trend / top fallback chains; 30s poll with visibilitychange pause). Owner-only_block; non-owner identities receive 401. Localhost-bound by default. |
|
| `/dashboard` | GET | 3 + 5 | ✅ Shipped (D50 + D51 + D82) | Owner-only multi-provider dashboard HTML. Phase 5 D82 adds a Claude.ai-style Plan Usage section at the top (per-provider utilization bars + reset countdowns + 60s auto-refresh + manual refresh button) on top of the existing four panels (24h request stats / 30d spend trend / top fallback chains / legacy quota fallback). Owner-only_block; non-owner identities receive 401. Localhost-bound by default. |
|
||||||
| `/v0/management/dashboard-data` | GET | 3 | ✅ Shipped (D50) | JSON aggregate consumed by the dashboard 30s poll: `{ generated_at, window_24h, cache_hit_24h, quota, spend_trend_30d, top_fallback_chains_24h, cache_stats }`. Owner-only_block. |
|
| `/v0/management/dashboard-data` | GET | 3 + 5 | ✅ Shipped (D50 + D81) | JSON aggregate consumed by the dashboard polls. Shape `{ generated_at, window_24h, cache_hit_24h, quota, quota_v2, spend_trend_30d, top_fallback_chains_24h, cache_stats }`. The new `quota_v2` field (D81) is the normalized per-provider shape consumed by the Plan Usage UI; the legacy `quota` field stays alongside for backwards compatibility until v1.0.0. Owner-only_block. |
|
||||||
| `/v0/management/quota` | GET | 3 | ✅ Shipped (D50) | Per-provider quota snapshot via `provider.quotaStatus()` (subset of dashboard-data; useful for scripted monitoring). Owner-only_block. |
|
| `/v0/management/quota` | GET | 3 + 5 | ✅ Shipped (D50 + D81) | Per-provider quota snapshot via `provider.quotaStatus()`. Includes both legacy `quota` and new `quota_v2` shape (mirrors `dashboard-data` for scripted monitoring). Owner-only_block. |
|
||||||
| `/cache/stats` | GET | 3 | ✅ Shipped (D50) | Live in-memory `cacheStore.stats()` (`{ hits, misses, size, inflightCount }` + `generated_at`). Owner-only_block. |
|
| `/cache/stats` | GET | 3 | ✅ Shipped (D50) | Live in-memory `cacheStore.stats()` (`{ hits, misses, size, inflightCount }` + `generated_at`). Owner-only_block. |
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -127,6 +353,10 @@ _placeholder — full table lands per-phase as variables are introduced._
|
|||||||
| Variable | Default | Description |
|
| Variable | Default | Description |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `OLP_PORT` | `4567` | HTTP listener port. Moved off `3456` at D60 / v0.4.0 to co-host with OCP — set `OLP_PORT=3456` to restore the pre-D60 default. |
|
| `OLP_PORT` | `4567` | HTTP listener port. Moved off `3456` at D60 / v0.4.0 to co-host with OCP — set `OLP_PORT=3456` to restore the pre-D60 default. |
|
||||||
|
| `OLP_BIND` | `127.0.0.1` | HTTP listener bind address. **Set to `0.0.0.0` or your LAN IP to accept LAN connections** (required for `olp-connect <ip>` to actually reach the server). Default loopback-only is the secure default. See [ADR 0011 § Deployment configurations](./docs/adr/0011-anonymous-key-deployment-context.md#deployment-configurations-d76-amendment-2026-05-26) for the trust-context table — never bind to a public-internet IP. |
|
||||||
|
| `OLP_API_KEY` | (none) | Owner-tier OLP API key (the `olp_...` plaintext from `olp-keys keygen --owner`) used by `olp` CLI subcommands as the bearer for management endpoints. |
|
||||||
|
| `OLP_OWNER_TOKEN` | (none) | Fallback used by `olp` CLI if `OLP_API_KEY` is absent. |
|
||||||
|
| `OLP_PROXY_URL` | `http://127.0.0.1:$OLP_PORT` | Override target URL for `olp` CLI subcommands (so the same binary works against a remote OLP via SSH tunnel or direct LAN). |
|
||||||
| `OLP_CLAUDE_BIN` | `claude` (from PATH) | Override path to the `claude` binary (Anthropic provider). Useful when multiple `claude` installs are present. |
|
| `OLP_CLAUDE_BIN` | `claude` (from PATH) | Override path to the `claude` binary (Anthropic provider). Useful when multiple `claude` installs are present. |
|
||||||
| `OLP_CODEX_BIN` | `codex` (from PATH) | Override path to the `codex` binary (OpenAI provider). |
|
| `OLP_CODEX_BIN` | `codex` (from PATH) | Override path to the `codex` binary (OpenAI provider). |
|
||||||
| `OLP_VIBE_BIN` | `vibe` (from PATH) | Override path to the `vibe` binary (Mistral provider). |
|
| `OLP_VIBE_BIN` | `vibe` (from PATH) | Override path to the `vibe` binary (Mistral provider). |
|
||||||
@@ -252,9 +482,9 @@ Use a dedicated bot key — not the maintainer's personal owner key — so revoc
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Implementation status (as of 2026-05-25, post-v0.2.0)
|
## Implementation status (as of 2026-05-26, post-v0.4.0)
|
||||||
|
|
||||||
Phase 1 closed at v0.1.1 (multi-provider proxy core + pre-Phase-2 cleanup). Phase 2 closed at v0.2.0 (multi-key auth + audit + owner gating + keygen CLI; ADR 0007 § 10 all 11 acceptance criteria shipped). Phase 3 closed at v0.3.0 (Dashboard + `lib/audit-query.mjs` + daily audit rotation; ADR 0008 § 10 all 15 acceptance criteria shipped). Phase 4 (per-key per-provider auth + audit retention + SQLite hybrid + provider-cost weights) is the next planned milestone. This table reflects what is currently shipped vs. what is designed for later phases.
|
Phase 1 closed at v0.1.1 (multi-provider proxy core + pre-Phase-2 cleanup). Phase 2 closed at v0.2.0 (multi-key auth + audit + owner gating + keygen CLI; ADR 0007 § 10 all 11 acceptance criteria shipped). Phase 3 closed at v0.3.0 (Dashboard + `lib/audit-query.mjs` + daily audit rotation; ADR 0008 § 10 all 15 acceptance criteria shipped). Phase 4 closed at v0.4.0 (Operator + Client UX per ADR 0010: SSE heartbeat + `recentErrors[20]` + `/v0/management/status` / `olp` Node CLI + `olp doctor` framework + ADR 0002 Amendment 7 / `olp-connect` bash + `/health.anonymousKey` + ADR 0011 / `olp-plugin/` Telegram-Discord + 6-IDE integration docs). Phase 5 closed at v0.5.0 (Quota Probes + Dashboard Enrichment — live Anthropic plan-usage probe + Claude.ai-style dashboard + audit-query aggregateProviderQuota); v0.5.1 hotfix (quota probe cache/backoff/schema-drift correctness — codex review findings F1–F3). Phase 6 is next. This table reflects what is currently shipped vs. what is designed for later phases.
|
||||||
|
|
||||||
| File / artifact | Status | Notes |
|
| File / artifact | Status | Notes |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
@@ -289,7 +519,9 @@ Behaviors that work correctly at personal/family scale but have ratified follow-
|
|||||||
- **Soft triggers configured but inert.** `routing.soft_triggers` in `~/.olp/config.json` is honored by the engine's evaluation logic but `quotaStatus()` polling is not wired (ADR 0004 Amendment 2). A startup warning fires if the field is non-empty so the inert state is visible.
|
- **Soft triggers configured but inert.** `routing.soft_triggers` in `~/.olp/config.json` is honored by the engine's evaluation logic but `quotaStatus()` polling is not wired (ADR 0004 Amendment 2). A startup warning fires if the field is non-empty so the inert state is visible.
|
||||||
- **Multi-key auth + owner gating + keygen CLI shipped at v0.2.0 (D44 + D45 + D46 + D47).** `lib/keys.mjs` (core), `lib/audit.mjs` (audit), owner-vs-guest `/health` payload trimming + `X-OLP-Fallback-Detail` policy gating, `bin/olp-keys.mjs` (keygen CLI). All 11 ADR 0007 § 10 acceptance criteria covered. v0.2.0 maintainer-merged 2026-05-25.
|
- **Multi-key auth + owner gating + keygen CLI shipped at v0.2.0 (D44 + D45 + D46 + D47).** `lib/keys.mjs` (core), `lib/audit.mjs` (audit), owner-vs-guest `/health` payload trimming + `X-OLP-Fallback-Detail` policy gating, `bin/olp-keys.mjs` (keygen CLI). All 11 ADR 0007 § 10 acceptance criteria covered. v0.2.0 maintainer-merged 2026-05-25.
|
||||||
|
|
||||||
- **Phase 3 (Dashboard + audit query layer + rotation) shipped to main (D48-D54); v0.3.0 release pending.** `docs/adr/0008-dashboard-and-audit-query.md` ratified at D48. `lib/audit-query.mjs` (D49) implements the 5-function aggregate query API (in-memory ndjson scan, PII-guarded). 4 new owner-only_block endpoints at D50 (`/dashboard`, `/v0/management/dashboard-data`, `/v0/management/quota`, `/cache/stats`). `dashboard.html` full multi-panel UI at D51 (vanilla HTML+JS+fetch, 30s poll with visibilitychange pause). Daily audit rotation at D52 (synchronous on first append after UTC midnight; `audit-YYYY-MM-DD.ndjson` naming) + optional `bin/olp-audit-rotate.mjs` cron tool. `tried_providers` schema semantic fix at D53 (D45 P2 deferral). Phase 3 close to v0.3.0 is maintainer-triggered per CLAUDE.md `release_kit.phase_close_trigger`.
|
- **Phase 3 (Dashboard + audit query layer + rotation) shipped at v0.3.0 (D48–D54).** `docs/adr/0008-dashboard-and-audit-query.md` + `lib/audit-query.mjs` (D49) + 4 owner-only_block endpoints (D50) + `dashboard.html` (D51) + daily audit rotation (D52) + `tried_providers` schema fix (D53). All 15 ADR 0008 § 10 acceptance criteria covered.
|
||||||
|
|
||||||
|
- **Phase 4 (Operator + Client UX) shipped at v0.4.0 (D60 → D73).** ADR 0010 (charter) + ADR 0011 (anonymous-key trusted-LAN limits) + ADR 0002 Amendment 7 (provider `doctorChecks()` contract). Default `OLP_PORT` 3456 → 4567 so OLP and OCP can co-host. SSE heartbeat (D61) + `recentErrors[20]` + `/v0/management/status` (D62-D63). `bin/olp.mjs` Node CLI + `bin/olp-keys.mjs` + `lib/doctor.mjs` framework with `next_action.ai_executable[]` (D64-D67). `bin/olp-connect` bash zero-config IDE auto-config + opt-in `/health.anonymousKey` (D68-D70). `olp-plugin/` OpenClaw `/olp` Telegram-Discord plugin (read-only, no chat mutations) + 6 IDE integration docs at `docs/integrations/*.md` (D71-D73). Test count 623 → 696.
|
||||||
|
|
||||||
**Bootstrap workflow (D47):** for first-run / production setup:
|
**Bootstrap workflow (D47):** for first-run / production setup:
|
||||||
|
|
||||||
@@ -314,6 +546,15 @@ Behaviors that work correctly at personal/family scale but have ratified follow-
|
|||||||
**New config block consumed at D45:** `config.json auth.{ allow_anonymous, owner_only_endpoints, fallback_detail_header_policy }`. Default `allow_anonymous: false` (production-off); set true to accept requests without an OLP API key (development / single-user dev mode). Startup emits a warn when `allow_anonymous: true` so the relaxed posture is observable.
|
**New config block consumed at D45:** `config.json auth.{ allow_anonymous, owner_only_endpoints, fallback_detail_header_policy }`. Default `allow_anonymous: false` (production-off); set true to accept requests without an OLP API key (development / single-user dev mode). Startup emits a warn when `allow_anonymous: true` so the relaxed posture is observable.
|
||||||
- **Provider-level `cacheKeyFields` mask not implemented.** Cache keys include every IR field including ones individual plugins drop at spawn (e.g., Anthropic plugin drops `temperature`). Spurious cache misses possible (extra spawn cost; never spurious hits). Conservative posture documented in [ADR 0005 Amendment 7](./docs/adr/0005-cache-cross-provider.md). Tracked in [v1.x roadmap #5](./docs/v1x-roadmap.md).
|
- **Provider-level `cacheKeyFields` mask not implemented.** Cache keys include every IR field including ones individual plugins drop at spawn (e.g., Anthropic plugin drops `temperature`). Spurious cache misses possible (extra spawn cost; never spurious hits). Conservative posture documented in [ADR 0005 Amendment 7](./docs/adr/0005-cache-cross-provider.md). Tracked in [v1.x roadmap #5](./docs/v1x-roadmap.md).
|
||||||
|
|
||||||
|
- **Agentic clients with shell-tool routing may report OLP-server-side state as "self".** This is an architectural property of spawn-CLI proxying that OLP cannot fully fix at the proxy layer. When a client like OpenClaw runs in **client mode** (gateway on user's machine, LLM backend pointed at remote OLP) and the agent exposes shell / fs tools, those tool calls execute on whatever machine the client's tool-handler is wired to. If the client's `ocp` / `olp` plugin routes shell to the OLP server host, an in-agent "do a self-check" prompt produces results describing the OLP host (e.g. PI231) rather than the user's local machine. OLP cannot inject "you are the client, not the server" into the prompt because (a) the client owns the system message, and (b) OLP is stateless and doesn't know which client is calling. **Phase 6c's `--system-prompt` override (ADR 0009 Amendment 1) addresses one side of this — claude CLI no longer injects `<env>cwd=...</env>` blocks into the prompt** — but it cannot prevent the client from sending tool-results that the model then describes as its own state. Recommendations for integrators:
|
||||||
|
|
||||||
|
- **OpenClaw client mode** — if you want bot self-checks to describe the user's local machine, configure OpenClaw's tool plugins (`plugins.entries.{ocp,olp}` etc.) so shell / fs tools route to the local host, not to the OLP server. The bundled `olp-plugin/` ships as a read-only telemetry surface (no shell mutations); the older `ocp` plugin's shell-routing semantics are OCP-era legacy and may misroute when OLP is the LLM backend.
|
||||||
|
- **Hermes Agent client mode** — Hermes pre-processes tools on its own host before sending; the LLM emits no tool_use that reaches OLP, so this limitation does not apply to chat-only Hermes flows. Tool-using Hermes flows behave correctly: Hermes runs the tool locally and includes the result as a follow-up user message.
|
||||||
|
- **Cline / Continue.dev / Cursor / Aider** — IDE clients typically run shell / fs tools locally on the user's machine, so self-checks report the user's machine correctly. No OLP-side action needed.
|
||||||
|
- **Generic agentic clients** — if your client routes tool execution to the OLP server, expect bot self-reports to describe the OLP server's state. Either: (1) configure your client's tool handler to run tools locally, or (2) document this to your client users as a known limitation.
|
||||||
|
|
||||||
|
See [ADR 0014](./docs/adr/0014-sandbox-runtime-integration.md) for the multi-tenant security counterpart of this issue — even with shell-tool routing, OLP server-side sandboxing prevents one client from reading another client's OAuth tokens (Phase 7 PR-A shipped; PR-B HTTP-path activation pending).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
@@ -341,7 +582,8 @@ The original v0.1 spec (in `~/.cc-rules/memory/projects/olp_v0_1_spec.md` on the
|
|||||||
- **Phase 1** — Multi-provider proxy core: `server.mjs`, IR, three Tier-D provider plugins (Anthropic / OpenAI Codex / Mistral Vibe), cache (D1+D4) + cleanup (D2 bypass / D3 chunked replay / D23 size cap), fallback engine with first-chunk safety + hard triggers + per-hop log observability, IR↔OpenAI translation under Rule 2(b). ✅ Shipped — v0.1.0 (2026-05-24) + v0.1.1 cleanup (2026-05-25, D35–D42).
|
- **Phase 1** — Multi-provider proxy core: `server.mjs`, IR, three Tier-D provider plugins (Anthropic / OpenAI Codex / Mistral Vibe), cache (D1+D4) + cleanup (D2 bypass / D3 chunked replay / D23 size cap), fallback engine with first-chunk safety + hard triggers + per-hop log observability, IR↔OpenAI translation under Rule 2(b). ✅ Shipped — v0.1.0 (2026-05-24) + v0.1.1 cleanup (2026-05-25, D35–D42).
|
||||||
- **Phase 2** — Multi-key auth (`lib/keys.mjs`) per ADR 0007: opaque OLP API keys, per-key cache namespacing, owner-vs-guest tier for header gating, audit ndjson (`lib/audit.mjs`), `/health` payload trimming + `X-OLP-Fallback-Detail` emission gating, `OLP_OWNER_TOKEN` env override, keygen CLI (`bin/olp-keys.mjs`). ✅ Shipped — v0.2.0 (2026-05-25, D43-A → D47). All 11 ADR 0007 § 10 acceptance criteria covered.
|
- **Phase 2** — Multi-key auth (`lib/keys.mjs`) per ADR 0007: opaque OLP API keys, per-key cache namespacing, owner-vs-guest tier for header gating, audit ndjson (`lib/audit.mjs`), `/health` payload trimming + `X-OLP-Fallback-Detail` emission gating, `OLP_OWNER_TOKEN` env override, keygen CLI (`bin/olp-keys.mjs`). ✅ Shipped — v0.2.0 (2026-05-25, D43-A → D47). All 11 ADR 0007 § 10 acceptance criteria covered.
|
||||||
- **Phase 3** — Dashboard + audit query layer + daily audit rotation per ADR 0008: in-memory ndjson aggregate query layer (`lib/audit-query.mjs`), 4 owner-only_block management endpoints (`/dashboard` + `/v0/management/dashboard-data` + `/v0/management/quota` + `/cache/stats`), multi-panel `dashboard.html` with 30s poll, synchronous daily audit rotation + `bin/olp-audit-rotate.mjs` cron tool, `tried_providers` schema fix (D45 P2 deferral). ✅ Shipped — v0.3.0 (2026-05-25, D48 → D54). All 15 ADR 0008 § 10 acceptance criteria covered.
|
- **Phase 3** — Dashboard + audit query layer + daily audit rotation per ADR 0008: in-memory ndjson aggregate query layer (`lib/audit-query.mjs`), 4 owner-only_block management endpoints (`/dashboard` + `/v0/management/dashboard-data` + `/v0/management/quota` + `/cache/stats`), multi-panel `dashboard.html` with 30s poll, synchronous daily audit rotation + `bin/olp-audit-rotate.mjs` cron tool, `tried_providers` schema fix (D45 P2 deferral). ✅ Shipped — v0.3.0 (2026-05-25, D48 → D54). All 15 ADR 0008 § 10 acceptance criteria covered.
|
||||||
- **Phase 4 (planned)** — Per-key per-provider auth artifact mapping (ADR 0007 § 12 deferral), audit query rotation/retention policies, SQLite hybrid migration (ADR 0007 § 13 trigger), provider-cost weights for spend trend.
|
- **Phase 5** — Live quota probe (Anthropic Pro/Max OAuth plan-usage via `anthropic-ratelimit-unified-*` headers), Claude.ai-style dashboard enrichment (utilization bars + reset countdowns), audit-query `aggregateProviderQuota()`, per-provider quota_v2 shape in dashboard-data. ✅ Shipped — v0.5.0 (2026-05-27). v0.5.1 hotfix (2026-05-27): quota probe cache/backoff/schema-drift correctness (codex review findings F1–F3).
|
||||||
|
- **Phase 6 (planned)** — Per-key per-provider auth artifact mapping (ADR 0007 § 12 deferral), audit query rotation/retention policies, SQLite hybrid migration (ADR 0007 § 13 trigger), provider-cost weights for spend trend.
|
||||||
- **Phase 4+ (v1.x roadmap, triggered as needed)** — Full deferred-work tracker: [`docs/v1x-roadmap.md`](./docs/v1x-roadmap.md). Includes streaming-path singleflight ([issue #16](https://github.com/dtzp555-max/olp/issues/16) + ADR 0005 Amendment 8 design ratified), soft-trigger reactivation (ADR 0004 Amendment 2), `/health` activeSpawns integration, provider-level `cacheKeyFields` mask, streaming-path SPAWN_FAILED salvage.
|
- **Phase 4+ (v1.x roadmap, triggered as needed)** — Full deferred-work tracker: [`docs/v1x-roadmap.md`](./docs/v1x-roadmap.md). Includes streaming-path singleflight ([issue #16](https://github.com/dtzp555-max/olp/issues/16) + ADR 0005 Amendment 8 design ratified), soft-trigger reactivation (ADR 0004 Amendment 2), `/health` activeSpawns integration, provider-level `cacheKeyFields` mask, streaming-path SPAWN_FAILED salvage.
|
||||||
- **Phase N (opt-in)** — Tier-2 / Tier-C provider plugins (Grok / Kimi / MiniMax / GLM / Qwen) per [ADR 0006](./docs/adr/0006-provider-inclusion.md); provider-native protocol endpoints; deterministic triggers. Triggered by tier-2 demand, not on the bootstrap path.
|
- **Phase N (opt-in)** — Tier-2 / Tier-C provider plugins (Grok / Kimi / MiniMax / GLM / Qwen) per [ADR 0006](./docs/adr/0006-provider-inclusion.md); provider-native protocol endpoints; deterministic triggers. Triggered by tier-2 demand, not on the bootstrap path.
|
||||||
|
|
||||||
@@ -351,16 +593,22 @@ Full spec (decision rationale, open questions, risks): `~/.cc-rules/memory/proje
|
|||||||
|
|
||||||
## Migration from OCP
|
## Migration from OCP
|
||||||
|
|
||||||
|
OLP is OCP's successor. The trigger was Anthropic's 2026-05-14 announcement (effective 2026-06-15) splitting `claude -p` / Agent SDK / third-party agent traffic out of the Pro/Max subscription pool into a separate fixed $100/month Agent SDK credit pool — invalidating OCP's foundational assumption (*"subscription = unlimited within rate limits"*) for its only provider. OLP's structural response is to spread risk across multiple subscriptions whose CLI/programmatic use remains in their main subscription pool, with intelligent fallback when one runs out.
|
||||||
|
|
||||||
|
Beyond the billing trigger, OLP is intentionally NOT a commercial multi-tenant SaaS (LiteLLM / OpenRouter / Portkey already serve that market with funding + SOC2), NOT an enterprise gateway competing on provider breadth, NOT a model-capability router ("route to the smartest model" — you pick the model in `routing.chains`), and NOT a conversation-state store (your client manages its own context). See [ADR 0001](./docs/adr/0001-project-founding.md) for the founding decision and [`ALIGNMENT.md`](./ALIGNMENT.md) for the constitution that governs every plugin / IR / entry-surface change.
|
||||||
|
|
||||||
|
### Migrating an existing OCP install
|
||||||
|
|
||||||
_placeholder — `scripts/migrate-from-ocp.mjs` lands with Phase 7 (📋 planned, not yet authored)._
|
_placeholder — `scripts/migrate-from-ocp.mjs` lands with Phase 7 (📋 planned, not yet authored)._
|
||||||
|
|
||||||
Anticipated user-facing flow (target: <5 minutes):
|
Anticipated user-facing flow (target: <5 minutes):
|
||||||
|
|
||||||
1. Stop OCP (`launchctl bootout` the OCP service or `ocp stop`).
|
1. Stop OCP (`launchctl bootout` the OCP service or `ocp stop`).
|
||||||
2. Install OLP.
|
2. Install OLP (per [§ Manual install](#manual-install-5-10-min) above).
|
||||||
3. Run `olp migrate-from-ocp` — copies `~/.ocp/keys/` to `~/.olp/keys/` and points provider plugins at OCP's existing auth artifacts where applicable.
|
3. Run `olp migrate-from-ocp` — will copy `~/.ocp/keys/` to `~/.olp/keys/` and point provider plugins at OCP's existing auth artifacts where applicable.
|
||||||
4. Start OLP. Clients pointing at port 4567 (or 3456 with `OLP_PORT=3456`) keep working; their existing OLP API keys remain valid. **Note (v0.4.0+):** default port moved from 3456 → 4567 so OCP and OLP can co-host during migration; set `OLP_PORT=3456` if you want the pre-D60 default.
|
4. Start OLP. Clients pointing at port 4567 (or 3456 with `OLP_PORT=3456`) keep working; their existing OLP API keys remain valid.
|
||||||
|
|
||||||
OCP's cache directory is *not* migrated: OLP's cache key format includes provider+model and warms cold naturally. OCP enters maintenance mode (stability fixes only) when OLP v0.1 ships; new development happens in OLP.
|
**Default port moved 3456 → 4567 at v0.4.0** so OCP and OLP can co-host on the same machine during the migration window — set `OLP_PORT=3456` if you want the pre-D60 default. OCP's cache directory is *not* migrated: OLP's cache key format includes provider+model and warms cold naturally. OCP enters maintenance mode (stability fixes only) when OLP v0.1 ships; new development happens in OLP.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
+106
-13
@@ -29,7 +29,35 @@
|
|||||||
|
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
OLP_CONNECT_VERSION="0.4.0-phase4"
|
# D78 (G13): derive version from package.json instead of hardcoding (was
|
||||||
|
# stuck at "0.4.0-phase4" through v0.4.1/v0.4.2/v0.4.3 because no one
|
||||||
|
# updated it). Look up package.json next to the script if available;
|
||||||
|
# fall back to "unknown" when running curl-piped (no on-disk package.json).
|
||||||
|
_resolve_version() {
|
||||||
|
local script_dir pkg
|
||||||
|
# When curl-piped (`curl ... | bash`), BASH_SOURCE[0] is empty → dirname
|
||||||
|
# yields "." → script_dir resolves to cwd. D78 reviewer P2-1 hardening:
|
||||||
|
# require the suffix-strip to actually fire (script_dir ENDED with /bin),
|
||||||
|
# otherwise we'd happily pick up an unrelated package.json from whatever
|
||||||
|
# directory the user happens to be in when piping. Belt-and-braces.
|
||||||
|
# ${BASH_SOURCE[0]:-} default-empty guards against `set -u` nounset error
|
||||||
|
# when invoked via `curl ... | bash` (no source file → BASH_SOURCE unset).
|
||||||
|
script_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]:-}")" &>/dev/null && pwd)"
|
||||||
|
if [[ "$script_dir" != */bin ]]; then
|
||||||
|
echo "unknown"
|
||||||
|
return
|
||||||
|
fi
|
||||||
|
pkg="${script_dir%/bin}/package.json"
|
||||||
|
# D78 reviewer P2-2: pass $pkg via env var instead of -c interpolation
|
||||||
|
# so paths with apostrophes / shell metacharacters can't break the
|
||||||
|
# python invocation. Canonical layout is safe; this is defense-in-depth.
|
||||||
|
if [[ -f "$pkg" ]] && command -v python3 >/dev/null 2>&1; then
|
||||||
|
OLP_PKG_PATH="$pkg" python3 -c 'import json,os;print(json.load(open(os.environ["OLP_PKG_PATH"])).get("version","unknown"))' 2>/dev/null || echo "unknown"
|
||||||
|
else
|
||||||
|
echo "unknown"
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
OLP_CONNECT_VERSION="$(_resolve_version)"
|
||||||
|
|
||||||
show_version() {
|
show_version() {
|
||||||
echo "olp-connect $OLP_CONNECT_VERSION"
|
echo "olp-connect $OLP_CONNECT_VERSION"
|
||||||
@@ -114,6 +142,35 @@ key_display() {
|
|||||||
fi
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# D74 P1-2: validate OLP API key format. Per ADR 0007 § 3, tokens are
|
||||||
|
# `olp_` + 32 random bytes base64url-encoded (43 chars, no padding). This
|
||||||
|
# regex pins the on-the-wire shape so a malformed or hostile `--key` /
|
||||||
|
# server-advertised `anonymousKey` never gets persisted into a shell rc.
|
||||||
|
# Returns 0 on valid, 1 on invalid (with diagnostic to stderr).
|
||||||
|
validate_olp_token() {
|
||||||
|
local k="$1" source="$2"
|
||||||
|
if [[ ! "$k" =~ ^olp_[A-Za-z0-9_-]{43}$ ]]; then
|
||||||
|
log_err "Rejected $source: token format does not match ^olp_[A-Za-z0-9_-]{43}$ (ADR 0007 § 3)."
|
||||||
|
log_err " Got ${#k}-char value starting with '$(echo "$k" | cut -c1-8)...'"
|
||||||
|
log_err " Expected: olp_ followed by 43 base64url chars. Run 'npx olp-keys list' on the server"
|
||||||
|
log_err " to confirm the key format, or have the operator regenerate with 'npx olp-keys keygen'."
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
# D74 P1-2: POSIX shell-quote a value before interpolating into a shell rc
|
||||||
|
# write. Wraps in single quotes + escapes embedded single quotes per:
|
||||||
|
# foo'bar → 'foo'\''bar'
|
||||||
|
# Even with the validator above, this is defense-in-depth: any non-token
|
||||||
|
# string that slips through (e.g., environment.d KEY=VALUE writes) MUST be
|
||||||
|
# safe to source. Same helper pattern as lib/doctor.mjs _shellQuote.
|
||||||
|
shell_quote() {
|
||||||
|
local s="$1"
|
||||||
|
# Escape any single quotes: ' → '\''
|
||||||
|
printf "'%s'" "${s//\'/\'\\\'\'}"
|
||||||
|
}
|
||||||
|
|
||||||
# Detect Claude Code and print warn-only message. Per ADR 0010 § Out of
|
# Detect Claude Code and print warn-only message. Per ADR 0010 § Out of
|
||||||
# Phase 4 scope, OLP does NOT ship /v1/messages and CC is not a supported
|
# Phase 4 scope, OLP does NOT ship /v1/messages and CC is not a supported
|
||||||
# client. The user is steered toward Cline + OLP.
|
# client. The user is steered toward Cline + OLP.
|
||||||
@@ -217,17 +274,28 @@ detect_aider() {
|
|||||||
fi
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
# Detect OpenClaw. Per Phase 4 D71-D73 (NOT in this PR), olp will ship
|
# Detect OpenClaw. Phase 4 D71-D73 shipped olp-plugin/ as the OpenClaw
|
||||||
# olp-plugin/ for OpenClaw with full Telegram/Discord /olp slash commands.
|
# gateway plugin for /olp Telegram + Discord slash commands. Point users
|
||||||
# Until that ships, we just announce detection and link.
|
# at the install path.
|
||||||
detect_openclaw() {
|
detect_openclaw() {
|
||||||
if command -v openclaw &>/dev/null || [[ -f "$HOME/.openclaw/openclaw.json" ]]; then
|
if command -v openclaw &>/dev/null || [[ -f "$HOME/.openclaw/openclaw.json" ]]; then
|
||||||
log_info ""
|
log_info ""
|
||||||
log_info "Detected: OpenClaw"
|
log_info "Detected: OpenClaw"
|
||||||
log_info " The OpenClaw OLP plugin (D71-D73) is NOT YET SHIPPED."
|
log_info " OLP ships an OpenClaw gateway plugin for /olp Telegram + Discord"
|
||||||
log_info " When it ships, install with: openclaw plugin install olp"
|
log_info " slash commands (status / usage / cache / models / providers /"
|
||||||
log_info " For now, you can manually point OpenClaw at OLP via the OPENAI_BASE_URL"
|
log_info " chain show / health / doctor). Read-only by design — no chat-side"
|
||||||
log_info " env var (already written to your shell rc above)."
|
log_info " mutations."
|
||||||
|
log_info ""
|
||||||
|
log_info " Install the plugin (one-time, on the host running OpenClaw):"
|
||||||
|
log_info " git clone https://github.com/dtzp555-max/olp.git /tmp/olp-repo"
|
||||||
|
log_info " openclaw plugins install /tmp/olp-repo/olp-plugin"
|
||||||
|
log_info " # OR symlink: ln -sf /tmp/olp-repo/olp-plugin ~/.openclaw/extensions/olp"
|
||||||
|
log_info ""
|
||||||
|
log_info " Then edit ~/.openclaw/openclaw.json to set the plugin apiKey to a"
|
||||||
|
log_info " dedicated OLP key (NOT your owner key — create one via olp-keys"
|
||||||
|
log_info " keygen --name <bot-name>). Restart OpenClaw gateway."
|
||||||
|
log_info ""
|
||||||
|
log_info " See docs/integrations/openclaw.md for full instructions."
|
||||||
fi
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -297,17 +365,20 @@ append_olp_block() {
|
|||||||
if $DRY_RUN; then
|
if $DRY_RUN; then
|
||||||
log_change "[dry-run] would append OLP block to $rc_file:"
|
log_change "[dry-run] would append OLP block to $rc_file:"
|
||||||
log_change " # OLP LAN (added by olp-connect)"
|
log_change " # OLP LAN (added by olp-connect)"
|
||||||
log_change " export OPENAI_BASE_URL=$base_url/v1"
|
log_change " export OPENAI_BASE_URL=$(shell_quote "$base_url/v1")"
|
||||||
[[ -n "$key" ]] && log_change " export OPENAI_API_KEY=$(key_display "$key")"
|
[[ -n "$key" ]] && log_change " export OPENAI_API_KEY=$(shell_quote "$(key_display "$key")")"
|
||||||
log_change " # /OLP LAN"
|
log_change " # /OLP LAN"
|
||||||
return 0
|
return 0
|
||||||
fi
|
fi
|
||||||
|
# D74 P1-2: shell-quote values before writing to rc files. Defense-in-depth
|
||||||
|
# alongside validate_olp_token — even if a future code path bypasses the
|
||||||
|
# validator, the rc file remains safe to source.
|
||||||
{
|
{
|
||||||
echo ""
|
echo ""
|
||||||
echo "# OLP LAN (added by olp-connect)"
|
echo "# OLP LAN (added by olp-connect)"
|
||||||
echo "export OPENAI_BASE_URL=$base_url/v1"
|
echo "export OPENAI_BASE_URL=$(shell_quote "$base_url/v1")"
|
||||||
if [[ -n "$key" ]]; then
|
if [[ -n "$key" ]]; then
|
||||||
echo "export OPENAI_API_KEY=$key"
|
echo "export OPENAI_API_KEY=$(shell_quote "$key")"
|
||||||
fi
|
fi
|
||||||
echo "# /OLP LAN"
|
echo "# /OLP LAN"
|
||||||
} >> "$rc_file"
|
} >> "$rc_file"
|
||||||
@@ -341,6 +412,14 @@ set_system_env() {
|
|||||||
return 0
|
return 0
|
||||||
fi
|
fi
|
||||||
mkdir -p "$env_dir" 2>/dev/null
|
mkdir -p "$env_dir" 2>/dev/null
|
||||||
|
# D74 P1-2: systemd environment.d format is KEY=VALUE per line. While
|
||||||
|
# systemd does its own parsing (no shell sourcing), reject embedded
|
||||||
|
# newlines defensively — validate_olp_token already enforces the
|
||||||
|
# restricted charset for the API key, so this is belt-and-braces.
|
||||||
|
if [[ "$base_url" == *$'\n'* || "$key" == *$'\n'* ]]; then
|
||||||
|
log_err "Refusing to write environment.d entry: value contains newline."
|
||||||
|
return 2
|
||||||
|
fi
|
||||||
{
|
{
|
||||||
echo "OPENAI_BASE_URL=$base_url/v1"
|
echo "OPENAI_BASE_URL=$base_url/v1"
|
||||||
if [[ -n "$key" ]]; then
|
if [[ -n "$key" ]]; then
|
||||||
@@ -363,8 +442,13 @@ main() {
|
|||||||
--port=*) port="${1#*=}"; shift ;;
|
--port=*) port="${1#*=}"; shift ;;
|
||||||
--key) key="${2:?--key requires a value}"
|
--key) key="${2:?--key requires a value}"
|
||||||
[[ -z "$key" ]] && { log_err "--key cannot be empty (omit --key for zero-config / auto-discovery)"; exit 1; }
|
[[ -z "$key" ]] && { log_err "--key cannot be empty (omit --key for zero-config / auto-discovery)"; exit 1; }
|
||||||
|
# D74 P1-2: reject malformed --key before it ever reaches an rc write.
|
||||||
|
validate_olp_token "$key" "--key flag" || exit 1
|
||||||
shift 2 ;;
|
shift 2 ;;
|
||||||
--key=*) key="${1#*=}"; shift ;;
|
--key=*) key="${1#*=}"
|
||||||
|
[[ -z "$key" ]] && { log_err "--key cannot be empty (omit --key for zero-config / auto-discovery)"; exit 1; }
|
||||||
|
validate_olp_token "$key" "--key flag" || exit 1
|
||||||
|
shift ;;
|
||||||
--no-system-env) NO_SYSTEM_ENV=true; shift ;;
|
--no-system-env) NO_SYSTEM_ENV=true; shift ;;
|
||||||
--dry-run) DRY_RUN=true; shift ;;
|
--dry-run) DRY_RUN=true; shift ;;
|
||||||
--version) show_version; exit 0 ;;
|
--version) show_version; exit 0 ;;
|
||||||
@@ -457,6 +541,13 @@ try:
|
|||||||
print(k if isinstance(k, str) and k else '')
|
print(k if isinstance(k, str) and k else '')
|
||||||
except: print('')" 2>/dev/null || echo "")
|
except: print('')" 2>/dev/null || echo "")
|
||||||
if [[ -n "$anon_key" ]]; then
|
if [[ -n "$anon_key" ]]; then
|
||||||
|
# D74 P1-2: validate server-advertised token shape before consuming.
|
||||||
|
# A hostile or misconfigured server could otherwise inject arbitrary
|
||||||
|
# strings into the user's rc file via the `anonymousKey` field.
|
||||||
|
if ! validate_olp_token "$anon_key" "/health.anonymousKey from ${host}:${port}"; then
|
||||||
|
log_err "Refusing to consume malformed advertised key. Use --key explicitly or contact the OLP operator."
|
||||||
|
exit 2
|
||||||
|
fi
|
||||||
key="$anon_key"
|
key="$anon_key"
|
||||||
log_ok "Using server-advertised anonymous key: $(key_display "$key")"
|
log_ok "Using server-advertised anonymous key: $(key_display "$key")"
|
||||||
log_info " (set by remote via auth.advertise_anonymous_key=true; see ADR 0011 for"
|
log_info " (set by remote via auth.advertise_anonymous_key=true; see ADR 0011 for"
|
||||||
@@ -479,6 +570,8 @@ except: print('')" 2>/dev/null || echo "")
|
|||||||
log_err " Re-run with: olp-connect $host --key olp_..."
|
log_err " Re-run with: olp-connect $host --key olp_..."
|
||||||
exit 2
|
exit 2
|
||||||
fi
|
fi
|
||||||
|
# D74 P1-2: also validate the interactively-prompted key.
|
||||||
|
validate_olp_token "$key" "interactive prompt" || exit 1
|
||||||
fi
|
fi
|
||||||
fi
|
fi
|
||||||
fi
|
fi
|
||||||
|
|||||||
+143
-18
@@ -226,6 +226,53 @@ function formatMs(ms) {
|
|||||||
return `${Math.floor(ms / 3600000)}h${Math.floor((ms % 3600000) / 60000)}m`;
|
return `${Math.floor(ms / 3600000)}h${Math.floor((ms % 3600000) / 60000)}m`;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** formatAgo(diffMs) — "N min ago" / "Nh ago" from a millisecond diff. */
|
||||||
|
function formatAgo(diffMs) {
|
||||||
|
if (typeof diffMs !== 'number' || diffMs < 0) return 'just now';
|
||||||
|
const sec = Math.floor(diffMs / 1000);
|
||||||
|
if (sec < 60) return `${sec}s ago`;
|
||||||
|
const min = Math.floor(sec / 60);
|
||||||
|
if (min < 60) return `${min}m ago`;
|
||||||
|
return `${Math.floor(min / 60)}h ago`;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* formatResetCountdown(epochSeconds) → human-readable reset countdown string.
|
||||||
|
*
|
||||||
|
* Mirrors dashboard.html formatResetCountdown(). Five ranges:
|
||||||
|
* past / < 1h / < 24h / < 7d / ≥ 7d
|
||||||
|
*
|
||||||
|
* Authority: ADR 0008 Amendment 2 (quota_v2 shape); ported from
|
||||||
|
* dashboard.html (D82). No external deps. Pure formatter.
|
||||||
|
*
|
||||||
|
* @param {number|null} epochSeconds — Unix epoch seconds for reset time
|
||||||
|
* @returns {string}
|
||||||
|
*/
|
||||||
|
export function formatResetCountdown(epochSeconds) {
|
||||||
|
if (epochSeconds == null) return '—';
|
||||||
|
const nowMs = Date.now();
|
||||||
|
const targetMs = epochSeconds * 1000;
|
||||||
|
const diffMs = targetMs - nowMs;
|
||||||
|
if (diffMs <= 0) return 'resetting now';
|
||||||
|
const diffMin = Math.floor(diffMs / 60000);
|
||||||
|
const diffHr = Math.floor(diffMin / 60);
|
||||||
|
const diffDay = Math.floor(diffHr / 24);
|
||||||
|
if (diffMin < 60) return `resets in ${diffMin}m`;
|
||||||
|
if (diffHr < 24) {
|
||||||
|
const remMin = diffMin - diffHr * 60;
|
||||||
|
if (remMin === 0) return `resets in ${diffHr}h`;
|
||||||
|
return `resets in ${diffHr}h ${remMin}m`;
|
||||||
|
}
|
||||||
|
const target = new Date(targetMs);
|
||||||
|
const timeStr = target.toLocaleString('en-US', { hour: 'numeric', minute: '2-digit', hour12: true });
|
||||||
|
if (diffDay < 7) {
|
||||||
|
const dayStr = target.toLocaleString('en-US', { weekday: 'short' });
|
||||||
|
return `resets ${dayStr} ${timeStr}`;
|
||||||
|
}
|
||||||
|
const dateStr = target.toLocaleString('en-US', { month: 'short', day: 'numeric' });
|
||||||
|
return `resets ${dateStr} ${timeStr}`;
|
||||||
|
}
|
||||||
|
|
||||||
// ── Subcommand: status ────────────────────────────────────────────────────
|
// ── Subcommand: status ────────────────────────────────────────────────────
|
||||||
|
|
||||||
async function cmdStatus(flags, io) {
|
async function cmdStatus(flags, io) {
|
||||||
@@ -251,9 +298,15 @@ async function cmdStatus(flags, io) {
|
|||||||
}
|
}
|
||||||
io.log(` total reqs: ${body.stats?.total_requests ?? 0}`);
|
io.log(` total reqs: ${body.stats?.total_requests ?? 0}`);
|
||||||
io.log(` active reqs: ${body.stats?.active_requests ?? 0}`);
|
io.log(` active reqs: ${body.stats?.active_requests ?? 0}`);
|
||||||
|
// D75 F4 fix: server payload nests cache stats under stats.cache (per
|
||||||
|
// server.mjs handleManagementStatus, ~line 2092). The CacheStore.stats()
|
||||||
|
// contract returns { hits, misses, size, inflightCount } per
|
||||||
|
// lib/cache/store.mjs — there is no `entries` field. Pre-D75 cmdStatus read
|
||||||
|
// `body.stats.cache.entries` (OCP-era) which was always undefined → output
|
||||||
|
// showed "entries=?". Same pattern as D74 P2-3 applied to cmdCache/cmdUsage.
|
||||||
if (body.stats?.cache) {
|
if (body.stats?.cache) {
|
||||||
const c = body.stats.cache;
|
const c = body.stats.cache;
|
||||||
io.log(` cache: hits=${c.hits ?? 0} misses=${c.misses ?? 0} entries=${c.entries ?? '?'}`);
|
io.log(` cache: hits=${c.hits ?? 0} misses=${c.misses ?? 0} entries=${c.size ?? 0}${typeof c.inflightCount === 'number' ? ` inflight=${c.inflightCount}` : ''}`);
|
||||||
}
|
}
|
||||||
if (Array.isArray(body.recent_errors) && body.recent_errors.length > 0) {
|
if (Array.isArray(body.recent_errors) && body.recent_errors.length > 0) {
|
||||||
io.log(` recent errors: ${body.recent_errors.length}`);
|
io.log(` recent errors: ${body.recent_errors.length}`);
|
||||||
@@ -301,29 +354,88 @@ async function cmdUsage(flags, io) {
|
|||||||
try { body = JSON.parse(res.body); }
|
try { body = JSON.parse(res.body); }
|
||||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||||
|
// D74 P2-3 fix: server payload shape is { generated_at, window_24h: { request_count, status_2xx,
|
||||||
|
// status_4xx, status_5xx, by_provider, by_owner_tier, by_path, median_latency_ms, p95_latency_ms },
|
||||||
|
// cache_hit_24h: { total, hit, miss, bypass, streaming_attached, hit_rate, by_provider }, quota: [{provider, ...}],
|
||||||
|
// spend_trend_30d: [{date, request_count, by_provider}], top_fallback_chains_24h: [{chain, count, ...}],
|
||||||
|
// cache_stats: { hits, misses, size, inflightCount } } per server.mjs:2027 + lib/audit-query.mjs.
|
||||||
io.log(colorize('OLP usage (24h)', ANSI.bold, io.useColor));
|
io.log(colorize('OLP usage (24h)', ANSI.bold, io.useColor));
|
||||||
io.log('─'.repeat(60));
|
io.log('─'.repeat(60));
|
||||||
const u24 = body.usage_24h ?? body.usage24h ?? body['24h'] ?? {};
|
const w24 = body.window_24h ?? {};
|
||||||
if (typeof u24 === 'object' && Object.keys(u24).length > 0) {
|
const cache24 = body.cache_hit_24h ?? {};
|
||||||
io.log(` requests: ${u24.requests ?? '?'}`);
|
if (typeof w24 === 'object' && (w24.request_count ?? 0) > 0) {
|
||||||
io.log(` cache hits: ${u24.cache_hits ?? '?'}`);
|
io.log(` requests: ${w24.request_count}`);
|
||||||
io.log(` fallbacks: ${u24.fallbacks ?? '?'}`);
|
io.log(` 2xx / 4xx / 5xx: ${w24.status_2xx ?? 0} / ${w24.status_4xx ?? 0} / ${w24.status_5xx ?? 0}`);
|
||||||
|
if (typeof w24.median_latency_ms === 'number') {
|
||||||
|
io.log(` latency p50/p95: ${w24.median_latency_ms}ms / ${w24.p95_latency_ms ?? 0}ms`);
|
||||||
|
}
|
||||||
|
if (typeof cache24.hit_rate === 'number') {
|
||||||
|
const pct = (cache24.hit_rate * 100).toFixed(1);
|
||||||
|
io.log(` cache hit rate: ${pct}% (hit=${cache24.hit ?? 0} miss=${cache24.miss ?? 0}${cache24.streaming_attached ? ` streaming_attached=${cache24.streaming_attached}` : ''})`);
|
||||||
|
}
|
||||||
} else {
|
} else {
|
||||||
io.log(' (no 24h usage data — server may not have processed any requests yet)');
|
io.log(' (no 24h usage data — server may not have processed any requests yet)');
|
||||||
}
|
}
|
||||||
if (Array.isArray(body.providers)) {
|
// F4 (v0.5.1 codex post-release review Q4): prefer quota_v2 when present
|
||||||
|
// (server v0.5.0+), fall back to legacy quota array on older servers.
|
||||||
|
// Authority: ADR 0008 Amendment 2 (quota_v2 shape).
|
||||||
|
if (Array.isArray(body.quota_v2) && body.quota_v2.length > 0) {
|
||||||
|
io.log('');
|
||||||
|
io.log(colorize('Per-provider quota (live)', ANSI.bold, io.useColor));
|
||||||
|
io.log('─'.repeat(60));
|
||||||
|
for (const p of body.quota_v2) {
|
||||||
|
const label = String(p.provider ?? '?').toUpperCase().padEnd(12);
|
||||||
|
const status = p.status ?? 'unavailable';
|
||||||
|
if (status === 'unavailable') {
|
||||||
|
io.log(` ${colorize(label, ANSI.gray, io.useColor)} unavailable ${p.reason ?? 'no public quota api'}`);
|
||||||
|
} else if (status === 'unreachable') {
|
||||||
|
const fk = p.failure?.kind ?? 'unknown';
|
||||||
|
const fm = p.failure?.message ?? 'probe failed';
|
||||||
|
io.log(` ${colorize(label, ANSI.red, io.useColor)} ❌ no cached data — failure: ${fk} (${fm})`);
|
||||||
|
} else {
|
||||||
|
// live or stale
|
||||||
|
const staleWarn = status === 'stale'
|
||||||
|
? colorize(` ⚠ stale${p.last_fresh_at ? ` (${formatAgo(Date.now() - p.last_fresh_at)})` : ''} failure: ${p.failure?.kind ?? 'unknown'}`, ANSI.yellow, io.useColor)
|
||||||
|
: '';
|
||||||
|
const util = p.utilization ?? {};
|
||||||
|
const reset = p.reset ?? {};
|
||||||
|
const parts = [];
|
||||||
|
for (const window of ['5h', '7d']) {
|
||||||
|
const frac = util[window];
|
||||||
|
const resetEpoch = reset[window];
|
||||||
|
if (frac != null) {
|
||||||
|
const pct = `${Math.round(frac * 100)}%`;
|
||||||
|
const rst = formatResetCountdown(resetEpoch);
|
||||||
|
parts.push(`${window}: ${colorize(pct, frac >= 0.8 ? ANSI.red : frac >= 0.5 ? ANSI.yellow : ANSI.green, io.useColor)} (${rst})`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const binding = p.representative_claim ? ` binding: ${p.representative_claim.replace('_', '-')}` : '';
|
||||||
|
io.log(` ${colorize(label, ANSI.bold, io.useColor)} ${colorize(status, status === 'live' ? ANSI.green : ANSI.yellow, io.useColor).padEnd(6)} ${parts.join(' ')}${binding}${staleWarn}`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
} else if (Array.isArray(body.quota) && body.quota.length > 0) {
|
||||||
|
// Legacy fallback for pre-v0.5.0 servers
|
||||||
io.log('');
|
io.log('');
|
||||||
io.log(colorize('Per-provider quota', ANSI.bold, io.useColor));
|
io.log(colorize('Per-provider quota', ANSI.bold, io.useColor));
|
||||||
io.log('─'.repeat(60));
|
io.log('─'.repeat(60));
|
||||||
for (const p of body.providers) {
|
for (const p of body.quota) {
|
||||||
io.log(` ${String(p.name ?? '?').padEnd(12)} ${p.percent_used != null ? `${p.percent_used}% used` : 'no quota api'}`);
|
const label = String(p.provider ?? '?').padEnd(12);
|
||||||
|
if (p.error) {
|
||||||
|
io.log(` ${label} error: ${p.error}`);
|
||||||
|
} else if (typeof p.percent_used === 'number') {
|
||||||
|
io.log(` ${label} ${p.percent_used}% used${p.resets_in_human ? ` (resets in ${p.resets_in_human})` : ''}`);
|
||||||
|
} else if (p.available === false) {
|
||||||
|
io.log(` ${label} unavailable`);
|
||||||
|
} else {
|
||||||
|
io.log(` ${label} no quota api`);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
if (Array.isArray(body.top_fallback_chains)) {
|
}
|
||||||
|
if (Array.isArray(body.top_fallback_chains_24h) && body.top_fallback_chains_24h.length > 0) {
|
||||||
io.log('');
|
io.log('');
|
||||||
io.log(colorize('Top fallback chains', ANSI.bold, io.useColor));
|
io.log(colorize('Top fallback chains (24h)', ANSI.bold, io.useColor));
|
||||||
io.log('─'.repeat(60));
|
io.log('─'.repeat(60));
|
||||||
for (const f of body.top_fallback_chains.slice(0, 10)) {
|
for (const f of body.top_fallback_chains_24h.slice(0, 10)) {
|
||||||
io.log(` ${String(f.count ?? '?').padStart(5)} ${(f.chain ?? []).join(' → ')}`);
|
io.log(` ${String(f.count ?? '?').padStart(5)} ${(f.chain ?? []).join(' → ')}`);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -360,13 +472,20 @@ async function cmdCache(flags, io) {
|
|||||||
try { body = JSON.parse(res.body); }
|
try { body = JSON.parse(res.body); }
|
||||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||||
io.log(colorize('OLP cache', ANSI.bold, io.useColor));
|
// D74 P2-3 fix: cacheStore.stats() returns { hits, misses, size, inflightCount }
|
||||||
|
// per lib/cache/store.mjs:320. There is no entries / evictions / bytes / maxBytes
|
||||||
|
// in the OLP cache model — those were OCP-era field names. Compute a hit rate from
|
||||||
|
// the numerator/denominator instead of fabricating bytes.
|
||||||
|
const hits = body.hits ?? 0;
|
||||||
|
const misses = body.misses ?? 0;
|
||||||
|
const denom = hits + misses;
|
||||||
|
const hitRate = denom > 0 ? ((hits / denom) * 100).toFixed(1) : '0.0';
|
||||||
|
io.log(colorize('OLP cache (live in-memory)', ANSI.bold, io.useColor));
|
||||||
io.log('─'.repeat(60));
|
io.log('─'.repeat(60));
|
||||||
io.log(` entries: ${body.entries ?? '?'}`);
|
io.log(` entries: ${body.size ?? 0}`);
|
||||||
io.log(` hits: ${body.hits ?? 0}`);
|
io.log(` hits / misses: ${hits} / ${misses} (hit rate ${hitRate}%)`);
|
||||||
io.log(` misses: ${body.misses ?? 0}`);
|
io.log(` inflight: ${body.inflightCount ?? 0}`);
|
||||||
io.log(` evictions:${body.evictions ?? 0}`);
|
if (body.generated_at) io.log(` generated_at: ${body.generated_at}`);
|
||||||
io.log(` bytes: ${formatBytes(body.bytes ?? 0)} (max ${formatBytes(body.maxBytes ?? 0)})`);
|
|
||||||
return 0;
|
return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -574,10 +693,16 @@ async function cmdDoctor(flags, io) {
|
|||||||
const proxyUrl = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
const proxyUrl = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
||||||
const checkFilter = typeof flags.check === 'string' ? flags.check : undefined;
|
const checkFilter = typeof flags.check === 'string' ? flags.check : undefined;
|
||||||
|
|
||||||
|
// D74 P1-1: pass authHeaders so server.running / server.version checks
|
||||||
|
// succeed under the default production posture (auth.allow_anonymous:
|
||||||
|
// false). resolveBearerToken returns null when no env var is set; the
|
||||||
|
// doctor still runs but distinguishes 401 from "server down" by status
|
||||||
|
// code per the updated check.
|
||||||
const result = await runDoctor({
|
const result = await runDoctor({
|
||||||
olpHome,
|
olpHome,
|
||||||
proxyUrl,
|
proxyUrl,
|
||||||
checkFilter,
|
checkFilter,
|
||||||
|
authHeaders: authHeaders(),
|
||||||
});
|
});
|
||||||
|
|
||||||
if (io.wantJson) {
|
if (io.wantJson) {
|
||||||
|
|||||||
+575
-24
@@ -1,16 +1,24 @@
|
|||||||
<!DOCTYPE html>
|
<!DOCTYPE html>
|
||||||
<!--
|
<!--
|
||||||
OLP Dashboard — Phase 3 / D51
|
OLP Dashboard — Phase 5 / D82
|
||||||
------------------------------
|
------------------------------
|
||||||
Multi-panel owner-only dashboard per ADR 0008 § 6. Polls
|
Multi-panel owner-only dashboard per ADR 0008 § 6.
|
||||||
/v0/management/dashboard-data every 30 seconds (paused when the
|
|
||||||
page is hidden via document.visibilityState).
|
|
||||||
|
|
||||||
Panels (per spec v0.1 § 4.6 + ADR 0008 Lane 5 = B full):
|
Panels:
|
||||||
1. Per-provider quota / credit pool
|
0. Plan Usage (new D82 — Claude.ai-style per-provider rows; quota_v2; 1-min refresh)
|
||||||
2. Per-provider 24h request count + cache hit rate + fallback rate
|
1. Per-provider quota / credit pool (legacy; kept for graceful fallback when quota_v2 absent)
|
||||||
3. 30-day spend trend (SVG sparkline; per-provider in tooltip)
|
2. Per-provider 24h request count + cache hit rate + fallback rate (30s refresh)
|
||||||
4. Top 10 fallback chains by trigger count
|
3. 30-day spend trend (SVG sparkline; per-provider in tooltip) (30s refresh)
|
||||||
|
4. Top 10 fallback chains by trigger count (30s refresh)
|
||||||
|
|
||||||
|
Refresh cadence:
|
||||||
|
- Plan Usage panel: 60s (separate timer; visibilityState-guarded per ADR 0012 D82)
|
||||||
|
- Other panels: 30s (original poll cadence; paused when tab hidden)
|
||||||
|
|
||||||
|
Authority:
|
||||||
|
- ADR 0008 § 6 — dashboard layout + owner-only_block
|
||||||
|
- ADR 0012 D82 — quota_v2 Claude.ai-style restructure
|
||||||
|
- v1.x roadmap #8 — closed by this D-day
|
||||||
|
|
||||||
No build step, no framework, no external dependencies. Vanilla JS +
|
No build step, no framework, no external dependencies. Vanilla JS +
|
||||||
fetch + DOM render. Owner-only_block: anonymous / guest / no-auth all
|
fetch + DOM render. Owner-only_block: anonymous / guest / no-auth all
|
||||||
@@ -43,15 +51,243 @@
|
|||||||
.chain { font-family: ui-monospace, "SF Mono", Menlo, monospace; font-size: 0.85rem; color: #374151; }
|
.chain { font-family: ui-monospace, "SF Mono", Menlo, monospace; font-size: 0.85rem; color: #374151; }
|
||||||
.pill { display: inline-block; background: #e5e7eb; color: #374151; padding: 0.05rem 0.4rem; border-radius: 3px; font-size: 0.75rem; }
|
.pill { display: inline-block; background: #e5e7eb; color: #374151; padding: 0.05rem 0.4rem; border-radius: 3px; font-size: 0.75rem; }
|
||||||
footer { margin-top: 2rem; color: #9ca3af; font-size: 0.75rem; text-align: center; }
|
footer { margin-top: 2rem; color: #9ca3af; font-size: 0.75rem; text-align: center; }
|
||||||
|
|
||||||
|
/* ───────────────────────────────────────────
|
||||||
|
Plan Usage panel — D82 Claude.ai-style rows
|
||||||
|
─────────────────────────────────────────── */
|
||||||
|
.plan-usage-header {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
justify-content: space-between;
|
||||||
|
margin-bottom: 1rem;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
gap: 0.5rem;
|
||||||
|
}
|
||||||
|
.plan-usage-header h2 { margin: 0; }
|
||||||
|
.plan-usage-meta {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.75rem;
|
||||||
|
font-size: 0.8rem;
|
||||||
|
color: #6b7280;
|
||||||
|
}
|
||||||
|
.refresh-btn {
|
||||||
|
display: inline-flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.3rem;
|
||||||
|
padding: 0.35rem 0.75rem;
|
||||||
|
background: #fff;
|
||||||
|
border: 1px solid #d1d5db;
|
||||||
|
border-radius: 4px;
|
||||||
|
font-size: 0.8rem;
|
||||||
|
color: #374151;
|
||||||
|
cursor: pointer;
|
||||||
|
transition: background 0.15s, border-color 0.15s;
|
||||||
|
white-space: nowrap;
|
||||||
|
}
|
||||||
|
.refresh-btn:hover:not(:disabled) { background: #f9fafb; border-color: #9ca3af; }
|
||||||
|
.refresh-btn:disabled { opacity: 0.55; cursor: not-allowed; }
|
||||||
|
.refresh-btn .spin { display: inline-block; animation: spin 0.8s linear infinite; }
|
||||||
|
@keyframes spin { to { transform: rotate(360deg); } }
|
||||||
|
|
||||||
|
.provider-row {
|
||||||
|
border: 1px solid #e5e7eb;
|
||||||
|
border-radius: 8px;
|
||||||
|
padding: 1rem 1.25rem;
|
||||||
|
margin-bottom: 0.75rem;
|
||||||
|
background: #fff;
|
||||||
|
}
|
||||||
|
.provider-row:last-child { margin-bottom: 0; }
|
||||||
|
.provider-row.unavailable { background: #f9fafb; }
|
||||||
|
.provider-row.stale { border-color: #fcd34d; }
|
||||||
|
.provider-row.unreachable { border-color: #fca5a5; background: #fff5f5; }
|
||||||
|
|
||||||
|
.provider-row-top {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.6rem;
|
||||||
|
margin-bottom: 0.75rem;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
}
|
||||||
|
.provider-badge {
|
||||||
|
display: inline-block;
|
||||||
|
padding: 0.2rem 0.55rem;
|
||||||
|
border-radius: 4px;
|
||||||
|
font-size: 0.75rem;
|
||||||
|
font-weight: 700;
|
||||||
|
letter-spacing: 0.06em;
|
||||||
|
text-transform: uppercase;
|
||||||
|
color: #fff;
|
||||||
|
}
|
||||||
|
.provider-badge.anthropic { background: #cc4b24; }
|
||||||
|
.provider-badge.codex { background: #10a37f; }
|
||||||
|
.provider-badge.mistral { background: #6d5acd; }
|
||||||
|
.provider-badge.openai { background: #10a37f; }
|
||||||
|
.provider-badge.default { background: #6b7280; }
|
||||||
|
|
||||||
|
.status-dot {
|
||||||
|
display: inline-block;
|
||||||
|
width: 8px;
|
||||||
|
height: 8px;
|
||||||
|
border-radius: 50%;
|
||||||
|
flex-shrink: 0;
|
||||||
|
}
|
||||||
|
.status-dot.live { background: #10b981; }
|
||||||
|
.status-dot.stale { background: #f59e0b; }
|
||||||
|
.status-dot.unavailable { background: #9ca3af; }
|
||||||
|
.status-dot.unreachable { background: #ef4444; }
|
||||||
|
|
||||||
|
.status-chip {
|
||||||
|
display: inline-block;
|
||||||
|
padding: 0.1rem 0.45rem;
|
||||||
|
border-radius: 99px;
|
||||||
|
font-size: 0.7rem;
|
||||||
|
font-weight: 600;
|
||||||
|
letter-spacing: 0.03em;
|
||||||
|
text-transform: uppercase;
|
||||||
|
}
|
||||||
|
.status-chip.live { background: #d1fae5; color: #065f46; }
|
||||||
|
.status-chip.stale { background: #fef3c7; color: #92400e; }
|
||||||
|
.status-chip.unavailable { background: #f3f4f6; color: #6b7280; }
|
||||||
|
.status-chip.unreachable { background: #fee2e2; color: #991b1b; }
|
||||||
|
|
||||||
|
.chip-sm {
|
||||||
|
display: inline-block;
|
||||||
|
padding: 0.1rem 0.45rem;
|
||||||
|
border-radius: 4px;
|
||||||
|
font-size: 0.7rem;
|
||||||
|
color: #374151;
|
||||||
|
background: #f3f4f6;
|
||||||
|
border: 1px solid #e5e7eb;
|
||||||
|
}
|
||||||
|
.schema-tag {
|
||||||
|
margin-left: auto;
|
||||||
|
font-size: 0.7rem;
|
||||||
|
color: #9ca3af;
|
||||||
|
}
|
||||||
|
.unavailable-reason {
|
||||||
|
font-size: 0.875rem;
|
||||||
|
color: #9ca3af;
|
||||||
|
font-style: italic;
|
||||||
|
padding: 0.25rem 0 0;
|
||||||
|
}
|
||||||
|
.unreachable-reason {
|
||||||
|
font-size: 0.875rem;
|
||||||
|
color: #b91c1c;
|
||||||
|
font-style: italic;
|
||||||
|
padding: 0.25rem 0 0;
|
||||||
|
}
|
||||||
|
.last-fresh-tag {
|
||||||
|
font-size: 0.7rem;
|
||||||
|
color: #9ca3af;
|
||||||
|
}
|
||||||
|
|
||||||
|
.utilization-bars { display: flex; flex-direction: column; gap: 0.6rem; }
|
||||||
|
.util-row {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.75rem;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
}
|
||||||
|
.util-label {
|
||||||
|
flex: 0 0 180px;
|
||||||
|
font-size: 0.8rem;
|
||||||
|
color: #6b7280;
|
||||||
|
white-space: nowrap;
|
||||||
|
overflow: hidden;
|
||||||
|
text-overflow: ellipsis;
|
||||||
|
}
|
||||||
|
@media (max-width: 600px) {
|
||||||
|
.util-label { flex: 0 0 100%; }
|
||||||
|
.util-row { flex-direction: column; align-items: flex-start; }
|
||||||
|
.grid { grid-template-columns: 1fr; }
|
||||||
|
}
|
||||||
|
.util-bar-wrap {
|
||||||
|
flex: 1 1 120px;
|
||||||
|
min-width: 80px;
|
||||||
|
height: 8px;
|
||||||
|
background: #e5e7eb;
|
||||||
|
border-radius: 99px;
|
||||||
|
overflow: hidden;
|
||||||
|
}
|
||||||
|
.util-bar-fill {
|
||||||
|
height: 100%;
|
||||||
|
border-radius: 99px;
|
||||||
|
transition: width 0.3s ease;
|
||||||
|
}
|
||||||
|
.util-bar-fill.green { background: linear-gradient(90deg, #34d399, #10b981); }
|
||||||
|
.util-bar-fill.amber { background: linear-gradient(90deg, #fbbf24, #f59e0b); }
|
||||||
|
.util-bar-fill.red { background: linear-gradient(90deg, #f87171, #ef4444); }
|
||||||
|
|
||||||
|
.util-pct {
|
||||||
|
flex: 0 0 40px;
|
||||||
|
font-size: 0.8rem;
|
||||||
|
font-variant-numeric: tabular-nums;
|
||||||
|
font-weight: 600;
|
||||||
|
color: #374151;
|
||||||
|
text-align: right;
|
||||||
|
}
|
||||||
|
.util-reset {
|
||||||
|
flex: 0 0 auto;
|
||||||
|
font-size: 0.75rem;
|
||||||
|
color: #6b7280;
|
||||||
|
white-space: nowrap;
|
||||||
|
}
|
||||||
|
.rep-claim-badge {
|
||||||
|
display: inline-block;
|
||||||
|
padding: 0.1rem 0.45rem;
|
||||||
|
border-radius: 4px;
|
||||||
|
font-size: 0.7rem;
|
||||||
|
background: #ede9fe;
|
||||||
|
color: #5b21b6;
|
||||||
|
border: 1px solid #ddd6fe;
|
||||||
|
font-weight: 600;
|
||||||
|
}
|
||||||
|
.overage-chip {
|
||||||
|
display: inline-block;
|
||||||
|
padding: 0.1rem 0.45rem;
|
||||||
|
border-radius: 4px;
|
||||||
|
font-size: 0.7rem;
|
||||||
|
background: #fef3c7;
|
||||||
|
color: #92400e;
|
||||||
|
border: 1px solid #fcd34d;
|
||||||
|
}
|
||||||
|
.overage-chip.allowed {
|
||||||
|
background: #d1fae5;
|
||||||
|
color: #065f46;
|
||||||
|
border-color: #6ee7b7;
|
||||||
|
}
|
||||||
|
.provider-row-bottom {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.5rem;
|
||||||
|
margin-top: 0.65rem;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
}
|
||||||
</style>
|
</style>
|
||||||
</head>
|
</head>
|
||||||
<body>
|
<body>
|
||||||
<h1>OLP Dashboard</h1>
|
<h1>OLP Dashboard</h1>
|
||||||
<div id="meta" class="meta">Loading…</div>
|
<div id="meta" class="meta">Loading…</div>
|
||||||
<div id="banner-slot"></div>
|
<div id="banner-slot"></div>
|
||||||
|
|
||||||
|
<!-- Plan Usage panel (D82 — Claude.ai-style; full width) -->
|
||||||
|
<section class="panel" style="max-width: 1200px; margin-bottom: 1rem;">
|
||||||
|
<div class="plan-usage-header">
|
||||||
|
<h2>Plan Usage</h2>
|
||||||
|
<div class="plan-usage-meta">
|
||||||
|
<span id="quota-last-refresh"></span>
|
||||||
|
<button class="refresh-btn" id="quota-refresh-btn" title="Refresh quota data">
|
||||||
|
<span id="quota-refresh-icon">↻</span> Refresh
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div id="panel-plan-usage"><div class="panel-loading">Loading…</div></div>
|
||||||
|
</section>
|
||||||
|
|
||||||
<div class="grid">
|
<div class="grid">
|
||||||
<section class="panel">
|
<section class="panel" id="legacy-quota-section" style="display:none;">
|
||||||
<h2>Quota (per provider)</h2>
|
<h2>Quota (per provider) — legacy</h2>
|
||||||
<div id="panel-quota"><div class="panel-loading">Loading…</div></div>
|
<div id="panel-quota"><div class="panel-loading">Loading…</div></div>
|
||||||
</section>
|
</section>
|
||||||
<section class="panel">
|
<section class="panel">
|
||||||
@@ -67,15 +303,21 @@
|
|||||||
<div id="panel-chains"><div class="panel-loading">Loading…</div></div>
|
<div id="panel-chains"><div class="panel-loading">Loading…</div></div>
|
||||||
</section>
|
</section>
|
||||||
</div>
|
</div>
|
||||||
<footer>OLP Dashboard · poll every 30s · paused when tab hidden · v0.3.0-phase3</footer>
|
<footer>OLP Dashboard · Plan Usage: 60s refresh · other panels: 30s · paused when tab hidden · v0.5.1</footer>
|
||||||
<script>
|
<script>
|
||||||
(function () {
|
(function () {
|
||||||
'use strict';
|
'use strict';
|
||||||
const POLL_INTERVAL_MS = 30000;
|
|
||||||
let pollHandle = null;
|
|
||||||
|
|
||||||
|
/* ─────────────── constants ─────────────── */
|
||||||
|
const POLL_INTERVAL_MS = 30000; // 30s for legacy panels
|
||||||
|
const QUOTA_POLL_INTERVAL_MS = 60000; // 60s for Plan Usage (D82)
|
||||||
|
let pollHandle = null;
|
||||||
|
let quotaRefreshTimer = null;
|
||||||
|
|
||||||
|
/* ─────────────── DOM helpers ─────────────── */
|
||||||
function fmtNum(n) { return (n ?? 0).toLocaleString(); }
|
function fmtNum(n) { return (n ?? 0).toLocaleString(); }
|
||||||
function fmtPct(rate) { return (rate * 100).toFixed(1) + '%'; }
|
function fmtPct(rate) { return (rate * 100).toFixed(1) + '%'; }
|
||||||
|
|
||||||
function el(tag, attrs, ...children) {
|
function el(tag, attrs, ...children) {
|
||||||
const node = document.createElement(tag);
|
const node = document.createElement(tag);
|
||||||
if (attrs) for (const [k, v] of Object.entries(attrs)) {
|
if (attrs) for (const [k, v] of Object.entries(attrs)) {
|
||||||
@@ -97,6 +339,223 @@
|
|||||||
return node;
|
return node;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Reset countdown helper (D82 § B) ─────────────── */
|
||||||
|
/**
|
||||||
|
* formatResetCountdown(epochSeconds) → human-readable string
|
||||||
|
*
|
||||||
|
* - past: "Resetting now…"
|
||||||
|
* - < 1 hour: "Resets in 23 min"
|
||||||
|
* - < 24 hours: "Resets in 12hr 30min"
|
||||||
|
* - < 7 days: "Resets Sun 9:00 PM"
|
||||||
|
* - >= 7 days: "Resets May 31 9:00 PM"
|
||||||
|
*/
|
||||||
|
function formatResetCountdown(epochSeconds) {
|
||||||
|
if (epochSeconds == null) return '—';
|
||||||
|
const nowMs = Date.now();
|
||||||
|
const targetMs = epochSeconds * 1000;
|
||||||
|
const diffMs = targetMs - nowMs;
|
||||||
|
|
||||||
|
if (diffMs <= 0) return 'Resetting now…';
|
||||||
|
|
||||||
|
const diffSec = Math.floor(diffMs / 1000);
|
||||||
|
const diffMin = Math.floor(diffSec / 60);
|
||||||
|
const diffHr = Math.floor(diffMin / 60);
|
||||||
|
const diffDay = Math.floor(diffHr / 24);
|
||||||
|
|
||||||
|
if (diffMin < 60) {
|
||||||
|
return 'Resets in ' + diffMin + ' min';
|
||||||
|
}
|
||||||
|
if (diffHr < 24) {
|
||||||
|
const remMin = diffMin - diffHr * 60;
|
||||||
|
if (remMin === 0) return 'Resets in ' + diffHr + 'hr';
|
||||||
|
return 'Resets in ' + diffHr + 'hr ' + remMin + 'min';
|
||||||
|
}
|
||||||
|
// Format as "Resets <day-of-week> <time>" or "Resets <month> <day> <time>"
|
||||||
|
const target = new Date(targetMs);
|
||||||
|
const timeStr = target.toLocaleString('en-US', { hour: 'numeric', minute: '2-digit', hour12: true });
|
||||||
|
if (diffDay < 7) {
|
||||||
|
const dayStr = target.toLocaleString('en-US', { weekday: 'short' });
|
||||||
|
return 'Resets ' + dayStr + ' ' + timeStr;
|
||||||
|
}
|
||||||
|
const dateStr = target.toLocaleString('en-US', { month: 'short', day: 'numeric' });
|
||||||
|
return 'Resets ' + dateStr + ' ' + timeStr;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ─────────────── "Updated N min ago" helper ─────────────── */
|
||||||
|
function formatAgo(epochMs) {
|
||||||
|
if (epochMs == null) return '';
|
||||||
|
const diffMs = Date.now() - epochMs;
|
||||||
|
if (diffMs < 0) return 'just now';
|
||||||
|
const diffSec = Math.floor(diffMs / 1000);
|
||||||
|
if (diffSec < 60) return 'Updated just now';
|
||||||
|
const diffMin = Math.floor(diffSec / 60);
|
||||||
|
if (diffMin === 1) return 'Updated 1 min ago';
|
||||||
|
if (diffMin < 60) return 'Updated ' + diffMin + ' min ago';
|
||||||
|
const diffHr = Math.floor(diffMin / 60);
|
||||||
|
if (diffHr === 1) return 'Updated ~1hr ago';
|
||||||
|
return 'Updated ~' + diffHr + 'hr ago';
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Utilization bar color ─────────────── */
|
||||||
|
function utilizationColor(fraction) {
|
||||||
|
if (fraction == null) return 'green';
|
||||||
|
if (fraction >= 0.80) return 'red';
|
||||||
|
if (fraction >= 0.50) return 'amber';
|
||||||
|
return 'green';
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Provider badge color class ─────────────── */
|
||||||
|
function providerBadgeClass(name) {
|
||||||
|
const n = (name || '').toLowerCase();
|
||||||
|
if (n === 'anthropic') return 'anthropic';
|
||||||
|
if (n === 'codex') return 'codex';
|
||||||
|
if (n === 'mistral') return 'mistral';
|
||||||
|
if (n === 'openai') return 'openai';
|
||||||
|
return 'default';
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Plan Usage renderer (quota_v2) ─────────────── */
|
||||||
|
function renderPlanUsage(quotaV2) {
|
||||||
|
const target = document.getElementById('panel-plan-usage');
|
||||||
|
target.innerHTML = '';
|
||||||
|
|
||||||
|
if (!Array.isArray(quotaV2) || quotaV2.length === 0) {
|
||||||
|
target.appendChild(el('div', { class: 'panel-loading' }, 'No quota data available.'));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const frag = document.createDocumentFragment();
|
||||||
|
|
||||||
|
for (const entry of quotaV2) {
|
||||||
|
const status = entry.status || 'unavailable';
|
||||||
|
const rowEl = el('div', { class: 'provider-row ' + status });
|
||||||
|
|
||||||
|
/* ── top bar: badge + status dot + chips + schema tag ── */
|
||||||
|
const topBar = el('div', { class: 'provider-row-top' });
|
||||||
|
|
||||||
|
topBar.appendChild(el('span', { class: 'provider-badge ' + providerBadgeClass(entry.provider) }, (entry.provider || '').toUpperCase()));
|
||||||
|
topBar.appendChild(el('span', { class: 'status-dot ' + status, title: 'Status: ' + status }));
|
||||||
|
topBar.appendChild(el('span', { class: 'status-chip ' + status }, status));
|
||||||
|
|
||||||
|
if (entry.schema_version) {
|
||||||
|
topBar.appendChild(el('span', { class: 'schema-tag' }, 'schema: ' + entry.schema_version));
|
||||||
|
}
|
||||||
|
|
||||||
|
rowEl.appendChild(topBar);
|
||||||
|
|
||||||
|
/* ── unavailable: just show reason, no bars ── */
|
||||||
|
if (status === 'unavailable') {
|
||||||
|
const reason = entry.reason || 'no public quota api or probe disabled';
|
||||||
|
rowEl.appendChild(el('div', { class: 'unavailable-reason' }, reason));
|
||||||
|
frag.appendChild(rowEl);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── unreachable (v0.5.1): probe failed + no cache — show failure detail ── */
|
||||||
|
if (status === 'unreachable') {
|
||||||
|
const failure = entry.failure || {};
|
||||||
|
const kind = failure.kind || 'unknown';
|
||||||
|
const msg = failure.message || 'probe failed — no cached data available';
|
||||||
|
const shortText = `${kind}: ${msg}`;
|
||||||
|
rowEl.appendChild(el('div', { class: 'unreachable-reason' }, shortText));
|
||||||
|
if (failure.backoff_until) {
|
||||||
|
const backoffMs = Math.max(0, failure.backoff_until - Date.now());
|
||||||
|
const backoffSec = Math.round(backoffMs / 1000);
|
||||||
|
if (backoffSec > 0) {
|
||||||
|
rowEl.appendChild(el('div', { class: 'unavailable-reason' }, `backoff active: ${backoffSec}s remaining`));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
frag.appendChild(rowEl);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── utilization bars (5h + 7d) ── */
|
||||||
|
const util = entry.utilization || {};
|
||||||
|
const reset = entry.reset || {};
|
||||||
|
const barsWrap = el('div', { class: 'utilization-bars' });
|
||||||
|
|
||||||
|
const windows = [
|
||||||
|
{ key: '5h', label: 'Current 5-hour session' },
|
||||||
|
{ key: '7d', label: 'Weekly all-models' },
|
||||||
|
];
|
||||||
|
|
||||||
|
for (const w of windows) {
|
||||||
|
const frac = util[w.key];
|
||||||
|
const resetEpoch = reset[w.key];
|
||||||
|
const color = utilizationColor(frac);
|
||||||
|
const pctStr = frac != null ? Math.round(frac * 100) + '%' : '—';
|
||||||
|
const fillPct = frac != null ? Math.min(100, Math.round(frac * 100)) : 0;
|
||||||
|
const resetStr = formatResetCountdown(resetEpoch);
|
||||||
|
|
||||||
|
const utilRow = el('div', { class: 'util-row' });
|
||||||
|
|
||||||
|
utilRow.appendChild(el('span', { class: 'util-label', title: w.label },
|
||||||
|
w.label + (frac != null ? ': ' + pctStr : '')
|
||||||
|
));
|
||||||
|
|
||||||
|
const barWrap = el('div', { class: 'util-bar-wrap' });
|
||||||
|
barWrap.appendChild(el('div', {
|
||||||
|
class: 'util-bar-fill ' + color,
|
||||||
|
style: 'width: ' + fillPct + '%',
|
||||||
|
'aria-valuenow': fillPct,
|
||||||
|
'aria-valuemin': '0',
|
||||||
|
'aria-valuemax': '100',
|
||||||
|
role: 'progressbar',
|
||||||
|
}));
|
||||||
|
utilRow.appendChild(barWrap);
|
||||||
|
|
||||||
|
utilRow.appendChild(el('span', { class: 'util-pct' }, pctStr));
|
||||||
|
utilRow.appendChild(el('span', { class: 'util-reset' }, resetStr));
|
||||||
|
|
||||||
|
barsWrap.appendChild(utilRow);
|
||||||
|
}
|
||||||
|
|
||||||
|
rowEl.appendChild(barsWrap);
|
||||||
|
|
||||||
|
/* ── bottom chips: representative-claim, overage, last-fresh ── */
|
||||||
|
const bottomBar = el('div', { class: 'provider-row-bottom' });
|
||||||
|
|
||||||
|
if (entry.representative_claim) {
|
||||||
|
const claimLabel = entry.representative_claim === 'five_hour' ? '5-hour claim'
|
||||||
|
: entry.representative_claim === 'seven_day' ? '7-day claim'
|
||||||
|
: entry.representative_claim;
|
||||||
|
bottomBar.appendChild(el('span', { class: 'rep-claim-badge', title: 'Binding window: ' + entry.representative_claim }, claimLabel));
|
||||||
|
}
|
||||||
|
|
||||||
|
if (entry.overage && entry.overage.status) {
|
||||||
|
const ov = entry.overage;
|
||||||
|
const ovStatus = (ov.status || 'unknown').toLowerCase();
|
||||||
|
const isAllowed = ovStatus === 'allowed' || ovStatus === 'active';
|
||||||
|
const chipClass = isAllowed ? 'overage-chip allowed' : 'overage-chip';
|
||||||
|
const label = 'Overage: ' + (ov.status || '—')
|
||||||
|
+ (ov.disabled_reason ? ' (' + ov.disabled_reason + ')' : '');
|
||||||
|
bottomBar.appendChild(el('span', { class: chipClass, title: label }, label));
|
||||||
|
}
|
||||||
|
|
||||||
|
if (entry.fallback_percentage != null) {
|
||||||
|
const fpPct = Math.round(entry.fallback_percentage * 100) + '%';
|
||||||
|
bottomBar.appendChild(el('span', { class: 'chip-sm', title: 'Fallback rate (last window)' }, 'Fallback ' + fpPct));
|
||||||
|
}
|
||||||
|
|
||||||
|
if (entry.last_fresh_at) {
|
||||||
|
bottomBar.appendChild(el('span', { class: 'last-fresh-tag' }, formatAgo(entry.last_fresh_at)));
|
||||||
|
}
|
||||||
|
|
||||||
|
if (status === 'stale') {
|
||||||
|
const staleTitle = entry.last_fresh_at
|
||||||
|
? 'Last successful probe was ' + formatAgo(entry.last_fresh_at) + '; backoff active'
|
||||||
|
: 'Probe data is stale; backoff active';
|
||||||
|
bottomBar.appendChild(el('span', { class: 'chip-sm', style: 'color: #92400e; background: #fef3c7; border-color: #fcd34d;', title: staleTitle }, '⚠ stale data'));
|
||||||
|
}
|
||||||
|
|
||||||
|
rowEl.appendChild(bottomBar);
|
||||||
|
frag.appendChild(rowEl);
|
||||||
|
}
|
||||||
|
|
||||||
|
target.appendChild(frag);
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Legacy quota renderer (graceful fallback) ─────────────── */
|
||||||
function renderQuota(data) {
|
function renderQuota(data) {
|
||||||
const target = document.getElementById('panel-quota');
|
const target = document.getElementById('panel-quota');
|
||||||
target.innerHTML = '';
|
target.innerHTML = '';
|
||||||
@@ -129,6 +588,28 @@
|
|||||||
target.appendChild(table);
|
target.appendChild(table);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Plan Usage top-level render + quota routing ─────────────── */
|
||||||
|
function renderQuotaSection(data) {
|
||||||
|
const hasV2 = Array.isArray(data.quota_v2) && data.quota_v2.length > 0;
|
||||||
|
const legacySection = document.getElementById('legacy-quota-section');
|
||||||
|
|
||||||
|
if (hasV2) {
|
||||||
|
// D82: use enriched quota_v2 rows; hide legacy panel
|
||||||
|
legacySection.style.display = 'none';
|
||||||
|
renderPlanUsage(data.quota_v2);
|
||||||
|
} else {
|
||||||
|
// Graceful fallback: show legacy quota panel (older server build without D81)
|
||||||
|
legacySection.style.display = '';
|
||||||
|
// Also show legacy data in Plan Usage panel with a note
|
||||||
|
const target = document.getElementById('panel-plan-usage');
|
||||||
|
target.innerHTML = '';
|
||||||
|
target.appendChild(el('div', { class: 'panel-loading', style: 'color:#6b7280;' },
|
||||||
|
'quota_v2 not available (server may not have D81 yet). See legacy Quota panel below.'));
|
||||||
|
renderQuota(data.quota);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Other panel renderers (unchanged from D51) ─────────────── */
|
||||||
function render24h(window24h, cacheHit24h) {
|
function render24h(window24h, cacheHit24h) {
|
||||||
const target = document.getElementById('panel-24h');
|
const target = document.getElementById('panel-24h');
|
||||||
target.innerHTML = '';
|
target.innerHTML = '';
|
||||||
@@ -196,7 +677,7 @@
|
|||||||
const minLabel = svgEl('text', { x: 4, y: height - 4, 'font-size': 10, fill: '#6b7280' });
|
const minLabel = svgEl('text', { x: 4, y: height - 4, 'font-size': 10, fill: '#6b7280' });
|
||||||
minLabel.textContent = '0';
|
minLabel.textContent = '0';
|
||||||
svg.appendChild(minLabel);
|
svg.appendChild(minLabel);
|
||||||
// Date labels (first + last only at v0.3.0; mid labels deferred — added if needed by Phase 4 UX feedback)
|
// Date labels (first + last only)
|
||||||
if (spendTrend30d.length > 0) {
|
if (spendTrend30d.length > 0) {
|
||||||
const firstDate = svgEl('text', { x: padding.left, y: height - 4, 'font-size': 10, fill: '#6b7280' });
|
const firstDate = svgEl('text', { x: padding.left, y: height - 4, 'font-size': 10, fill: '#6b7280' });
|
||||||
firstDate.textContent = spendTrend30d[0].date.slice(5);
|
firstDate.textContent = spendTrend30d[0].date.slice(5);
|
||||||
@@ -240,6 +721,7 @@
|
|||||||
target.appendChild(table);
|
target.appendChild(table);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Error / clear banner ─────────────── */
|
||||||
function showError(message) {
|
function showError(message) {
|
||||||
const slot = document.getElementById('banner-slot');
|
const slot = document.getElementById('banner-slot');
|
||||||
slot.innerHTML = '';
|
slot.innerHTML = '';
|
||||||
@@ -250,6 +732,7 @@
|
|||||||
document.getElementById('banner-slot').innerHTML = '';
|
document.getElementById('banner-slot').innerHTML = '';
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Fetch ─────────────── */
|
||||||
async function fetchDashboardData() {
|
async function fetchDashboardData() {
|
||||||
const res = await fetch('/v0/management/dashboard-data', {
|
const res = await fetch('/v0/management/dashboard-data', {
|
||||||
headers: { 'Accept': 'application/json' },
|
headers: { 'Accept': 'application/json' },
|
||||||
@@ -266,24 +749,82 @@
|
|||||||
return await res.json();
|
return await res.json();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Quota-only refresh (D82 § C — 60s timer) ─────────────── */
|
||||||
|
let _lastQuotaFetchedAt = null;
|
||||||
|
|
||||||
|
async function refreshQuotaV2() {
|
||||||
|
try {
|
||||||
|
const data = await fetchDashboardData();
|
||||||
|
clearError();
|
||||||
|
_lastQuotaFetchedAt = Date.now();
|
||||||
|
renderQuotaSection(data);
|
||||||
|
updateQuotaLastRefreshLabel();
|
||||||
|
} catch (err) {
|
||||||
|
console.warn('OLP quota refresh failed:', err.message);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function updateQuotaLastRefreshLabel() {
|
||||||
|
const span = document.getElementById('quota-last-refresh');
|
||||||
|
if (!span) return;
|
||||||
|
if (_lastQuotaFetchedAt) {
|
||||||
|
span.textContent = 'Updated ' + new Date(_lastQuotaFetchedAt).toLocaleTimeString();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ─────────────── 60s quota timer with visibilityState guard ─────────────── */
|
||||||
|
function startQuotaRefresh() {
|
||||||
|
if (quotaRefreshTimer !== null) return;
|
||||||
|
quotaRefreshTimer = setInterval(refreshQuotaV2, QUOTA_POLL_INTERVAL_MS);
|
||||||
|
}
|
||||||
|
function stopQuotaRefresh() {
|
||||||
|
if (quotaRefreshTimer === null) return;
|
||||||
|
clearInterval(quotaRefreshTimer);
|
||||||
|
quotaRefreshTimer = null;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ─────────────── Manual refresh button (D82 § D) ─────────────── */
|
||||||
|
(function wireRefreshButton() {
|
||||||
|
const btn = document.getElementById('quota-refresh-btn');
|
||||||
|
const icon = document.getElementById('quota-refresh-icon');
|
||||||
|
if (!btn) return;
|
||||||
|
btn.addEventListener('click', async () => {
|
||||||
|
if (btn.disabled) return;
|
||||||
|
btn.disabled = true;
|
||||||
|
icon.textContent = '⟳';
|
||||||
|
icon.classList.add('spin');
|
||||||
|
try {
|
||||||
|
await refreshQuotaV2();
|
||||||
|
} finally {
|
||||||
|
icon.classList.remove('spin');
|
||||||
|
icon.textContent = '↻';
|
||||||
|
// Re-enable after 2s spam guard
|
||||||
|
setTimeout(() => { btn.disabled = false; }, 2000);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
})();
|
||||||
|
|
||||||
|
/* ─────────────── Full 30s refresh (legacy panels + meta) ─────────────── */
|
||||||
async function refresh() {
|
async function refresh() {
|
||||||
try {
|
try {
|
||||||
const data = await fetchDashboardData();
|
const data = await fetchDashboardData();
|
||||||
clearError();
|
clearError();
|
||||||
const generated = data.generated_at ? new Date(data.generated_at) : new Date();
|
const generated = data.generated_at ? new Date(data.generated_at) : new Date();
|
||||||
document.getElementById('meta').textContent =
|
document.getElementById('meta').textContent =
|
||||||
'Last refresh: ' + generated.toLocaleString() + ' · next in ~30s';
|
'Last refresh: ' + generated.toLocaleString() + ' · quota every 60s · other panels every 30s';
|
||||||
renderQuota(data.quota);
|
// Quota section: also render on each full refresh to keep in sync
|
||||||
|
_lastQuotaFetchedAt = Date.now();
|
||||||
|
renderQuotaSection(data);
|
||||||
|
updateQuotaLastRefreshLabel();
|
||||||
render24h(data.window_24h, data.cache_hit_24h);
|
render24h(data.window_24h, data.cache_hit_24h);
|
||||||
renderTrend(data.spend_trend_30d);
|
renderTrend(data.spend_trend_30d);
|
||||||
renderChains(data.top_fallback_chains_24h);
|
renderChains(data.top_fallback_chains_24h);
|
||||||
} catch (err) {
|
} catch (err) {
|
||||||
// Error banner already shown by fetchDashboardData; keep panels in
|
|
||||||
// their last-good state. Console for operator debugging.
|
|
||||||
console.warn('OLP dashboard refresh failed:', err.message);
|
console.warn('OLP dashboard refresh failed:', err.message);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* ─────────────── 30s poll (legacy panels) ─────────────── */
|
||||||
function startPolling() {
|
function startPolling() {
|
||||||
if (pollHandle !== null) return;
|
if (pollHandle !== null) return;
|
||||||
pollHandle = setInterval(refresh, POLL_INTERVAL_MS);
|
pollHandle = setInterval(refresh, POLL_INTERVAL_MS);
|
||||||
@@ -294,14 +835,24 @@
|
|||||||
pollHandle = null;
|
pollHandle = null;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Pause when tab hidden, resume on visible (ADR 0008 § 6.5).
|
/* ─────────────── visibilitychange (both timers) ─────────────── */
|
||||||
document.addEventListener('visibilitychange', () => {
|
document.addEventListener('visibilitychange', () => {
|
||||||
if (document.visibilityState === 'hidden') stopPolling();
|
if (document.visibilityState === 'hidden') {
|
||||||
else { refresh(); startPolling(); }
|
stopPolling();
|
||||||
|
stopQuotaRefresh();
|
||||||
|
} else {
|
||||||
|
refresh();
|
||||||
|
startPolling();
|
||||||
|
refreshQuotaV2();
|
||||||
|
startQuotaRefresh();
|
||||||
|
}
|
||||||
});
|
});
|
||||||
|
|
||||||
// Initial fetch + start poll.
|
/* ─────────────── Boot ─────────────── */
|
||||||
refresh().finally(startPolling);
|
refresh().finally(() => {
|
||||||
|
startPolling();
|
||||||
|
if (document.visibilityState === 'visible') startQuotaRefresh();
|
||||||
|
});
|
||||||
})();
|
})();
|
||||||
</script>
|
</script>
|
||||||
</body>
|
</body>
|
||||||
|
|||||||
@@ -9,6 +9,30 @@
|
|||||||
|
|
||||||
> **Note on numbering.** Sequence is 1, 3, 4, 5, 6, 7 — Amendment 2 was never written. The reserved slot was originally planned for a separate `maxConcurrent` ratification, but that content was folded into Amendment 1 (the retroactive contract-sync amendment) at filing time and the gap was not backfilled. The gap is intentional and load-bearing — no missing content; do not renumber Amendments 3+ to close it (cross-references to Amendment N from other docs would silently break).
|
> **Note on numbering.** Sequence is 1, 3, 4, 5, 6, 7 — Amendment 2 was never written. The reserved slot was originally planned for a separate `maxConcurrent` ratification, but that content was folded into Amendment 1 (the retroactive contract-sync amendment) at filing time and the gap was not backfilled. The gap is intentional and load-bearing — no missing content; do not renumber Amendments 3+ to close it (cross-references to Amendment N from other docs would silently break).
|
||||||
|
|
||||||
|
### Amendment 8 — 2026-05-26: Permit `quotaStatus()` direct-API access (READ-ONLY exemption) for plan-usage probes (D79–D80 — Phase 5)
|
||||||
|
|
||||||
|
- **Context:** ADR 0012 (Phase 5 charter) opens 2026-05-26 to port OCP's plan-usage probe (`ocp/server.mjs:842-1109`) into `lib/providers/anthropic.mjs:quotaStatus()`. The probe calls `POST https://api.anthropic.com/v1/messages` directly with an OAuth bearer and parses `anthropic-ratelimit-unified-*` response headers. This violates the plugin contract's implicit assumption that ALL provider interaction goes through `spawn` (the binary CLI). `ALIGNMENT.md` Rule 2 (provider-CLI-as-authority) further constrains plugins to operations the provider CLI itself performs. The OCP-derived plan-usage probe satisfies neither of these — it bypasses `claude -p` and hits the public API directly. **Without an explicit exemption Amendment, D80 is unalignable.**
|
||||||
|
- **Why the exemption is sound:** The probe is strictly **READ-ONLY** (one `POST /v1/messages` with `max_tokens: 1`; the response body is discarded; only response headers are parsed) AND **subscription-scope** (the OAuth bearer is the same one Claude Code uses for `claude -p`; no extra grant is requested) AND **idempotent** (probe failure returns `null`, never throws to a caller). The "what authority backs this?" answer is: Anthropic's CLI internally makes the same `/v1/messages` call (verified 2026-05-26 by `strings` on the compiled binary — see `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`); the probe is mirroring an established CLI behaviour rather than introducing a new wire format. Under `ALIGNMENT.md` Rule 2, mirroring observed CLI behaviour is permitted; the Rule's intent is "don't invent wire formats Anthropic's CLI does not perform", which the probe respects.
|
||||||
|
- **Change — extend the Provider contract description:**
|
||||||
|
- `quotaStatus(authContext): { quotaInfo }` is now permitted to call provider HTTP APIs directly, subject to **all three** constraints:
|
||||||
|
1. **READ-ONLY** — the API call must not mutate provider-side state. POST is acceptable when the response is what's needed (Anthropic returns ratelimit headers on `POST /v1/messages`); the request body MUST minimise side-effects (`max_tokens: 1`, dummy `messages`).
|
||||||
|
2. **Subscription-scope reuse** — the credentials used MUST be the same auth artifact the spawn path already reads via `readAuthArtifact()`. No new OAuth grant, no new API-key registration, no separate scopes.
|
||||||
|
3. **Idempotent failure** — if the probe fails for any reason (network error, 401, 429, schema parse failure), the function returns a structured shape (`{ probe_status: 'unreachable', failure: { kind, message, backoff_until? } }` since v0.5.1; see ADR 0013 Rule 6 + ADR 0008 Amendment 2) rather than throwing. The caller (server.mjs / dashboard / `olp usage` CLI) gracefully degrades. At v0.5.0 the failure shape was the literal value `null`; v0.5.1 refined this to a structured shape so operators can distinguish auth failures from rate-limit failures from network failures from in-backoff stale-cache. The substantive idempotent-failure constraint (no throw to caller) is unchanged.
|
||||||
|
- `healthCheck()` and other contract methods are NOT extended by this Amendment. Only `quotaStatus()` may make direct API calls. A plugin that wants live data for any other contract method must continue to use `spawn` or `readAuthArtifact`.
|
||||||
|
- The probe MUST cache its result. Recommended TTL: 5 minutes (mirrors OCP `USAGE_CACHE_TTL`). Tighter TTLs (e.g. dashboard's 1-minute refresh) are served from the cached value if fresh; cache miss triggers a real probe.
|
||||||
|
- The probe MUST implement exponential backoff on refresh failures: minimum 60s, maximum 3600s (mirrors OCP `OAUTH_REFRESH_MIN_BACKOFF` / `OAUTH_REFRESH_MAX_BACKOFF`). Tight loop on failure has historically burned through Anthropic's rate limit in seconds (OCP institutional lesson 2026-04).
|
||||||
|
- The probe MUST be opt-in via `~/.olp/config.json` (`providers.<name>.quota_probe_enabled: true`; default `false`). Reasoning: a fresh OLP install on a machine without OAuth credentials should not bombard `api.anthropic.com` with 401-bound probes; the operator opts in once the credentials are configured.
|
||||||
|
- **What this Amendment does NOT permit:**
|
||||||
|
- Mutating API calls (e.g. POST/PATCH/DELETE that change provider-side state). Still forbidden.
|
||||||
|
- API calls for any contract method other than `quotaStatus()`. `spawn` / `healthCheck` / `doctorChecks` / `estimateCost` / `models` / `hints` / `name` / `displayName` / `auth` remain spawn-and-filesystem-only.
|
||||||
|
- Per-provider new auth grants. The probe uses the spawn path's existing credentials.
|
||||||
|
- Bypassing the alignment.yml blacklist. The hallucinated `/api/oauth/usage` token stays blacklisted; the probe uses `/v1/messages` (real endpoint).
|
||||||
|
- **API calls to endpoints not explicitly enumerated by the companion ADR 0013 § Rule 2.** Amendment 8 permits the *kind* of call (READ-ONLY direct API for quota probing); ADR 0013 Rule 2 enumerates *which specific endpoint* is permitted. A future reader of Amendment 8 alone should NOT infer that any READ-ONLY/idempotent endpoint is fair game — the per-endpoint containment is locked to ADR 0013. Re-opening per-endpoint scope requires an ADR 0013 amendment, not a new plugin-level interpretation of Amendment 8.
|
||||||
|
- **Backwards compatibility:** Plugins whose `quotaStatus()` still returns `null` (mistral at v0.5.0 pending D84 audit, codex permanently per Phase 5 charter) are NOT affected. No existing behaviour changes for them.
|
||||||
|
- **Authority cited at the implementation:** D80 commit cites this Amendment + `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md` + `Claude Code v2.1.x § OAuth bearer + ratelimit-unified headers` + live-probe transcript from 2026-05-26 in the commit body. ALIGNMENT.md Rule 1 + Rule 5 (CI) both satisfied.
|
||||||
|
- **Tests:** Suite 38 (Phase 5 D83) covers the probe: mock HTTP server returning all 13 `anthropic-ratelimit-unified-*` headers; assert parse correctness for each; assert 5min cache; assert 60s-3600s exponential backoff on simulated 429; assert stale-cache-on-failure. At v0.5.0 stale-failure returned `{ stale: true, ... }` with `null` reserved for no-cache failures; v0.5.1 refined the return contract — `null` is now reserved STRICTLY for opt-in-off, and all failure modes (auth / rate-limit / schema-drift / network / no-creds) return `{ probe_status: 'unreachable' | 'stale', failure: {...} }`. See `test-features.mjs` Suite 38 (38u/38v/38w added for the v0.5.1 hotfix regression coverage of F1 / F2 / F3 per codex review).
|
||||||
|
- **Procedural mechanism:** Iron Rule 11 (IDR) — this Amendment, ADR 0012 (Phase 5 charter), and ADR 0013 (OAuth READ-ONLY consumption rules) land together at D79 as a single coupled commit. Reviewing them separately cannot verify consumer-producer alignment. Iron Rule 10 fresh-context reviewer per CLAUDE.md hard requirement #3.
|
||||||
|
|
||||||
### Amendment 7 — 2026-05-26: Add OPTIONAL `doctorChecks()` to the Provider contract (D67 — Phase 4 operator UX)
|
### Amendment 7 — 2026-05-26: Add OPTIONAL `doctorChecks()` to the Provider contract (D67 — Phase 4 operator UX)
|
||||||
|
|
||||||
- **Context:** ADR 0010 § Phase 4 D64-D67 ships `bin/olp.mjs` operator CLI + `olp doctor` framework. `olp doctor` runs a set of `Check` objects (id / category / async `run()` returning `{ status, message, evidence? }`) and discriminates the next remediation step via a `kind` field (`noop` / `fix_server` / `fix_oauth` / `fix_provider` / `fresh_install`). The framework needs per-provider checks so a user with a broken `claude` install gets a different fix recipe than a user with a broken `vibe` install. Hardcoding the recipes in `bin/olp.mjs` would re-introduce the kind of per-provider knowledge drift that ADR 0002 § Decision exists to prevent — when a new provider plugin lands, the operator CLI would have to be edited too.
|
- **Context:** ADR 0010 § Phase 4 D64-D67 ships `bin/olp.mjs` operator CLI + `olp doctor` framework. `olp doctor` runs a set of `Check` objects (id / category / async `run()` returning `{ status, message, evidence? }`) and discriminates the next remediation step via a `kind` field (`noop` / `fix_server` / `fix_oauth` / `fix_provider` / `fresh_install`). The framework needs per-provider checks so a user with a broken `claude` install gets a different fix recipe than a user with a broken `vibe` install. Hardcoding the recipes in `bin/olp.mjs` would re-introduce the kind of per-provider knowledge drift that ADR 0002 § Decision exists to prevent — when a new provider plugin lands, the operator CLI would have to be edited too.
|
||||||
|
|||||||
@@ -7,6 +7,23 @@
|
|||||||
|
|
||||||
## Amendments
|
## Amendments
|
||||||
|
|
||||||
|
### Amendment 3 — 2026-05-27: Accept OpenAI `role: "developer"` at entry surface, normalize to `system` in IR
|
||||||
|
|
||||||
|
- **Finding:** Hermes Agent v0.13/v0.14 (and likely Cline, Continue.dev, and other modern openai-completions clients) default to `role: "developer"` for what was historically the `system`-role slot when the model id matches OpenAI's o1/o3+ reasoning family. The `developer` role was introduced by OpenAI's Responses-API spec for reasoning models (high-priority developer-authored instructions; semantically a peer of `system`). OLP IR's role allow-list at v0.1 was the original four roles (`system|user|assistant|tool`); the IR validator rejected `developer` with `400 IR validation failed: role must be one of system|user|assistant|tool, got "developer"`. Reproduced 2026-05-27 on PI230 Hermes v0.14.0 → OLP v0.5.1 routing path.
|
||||||
|
- **Decision:** Extend `openai-to-ir.mjs:normalizeRole()` to map `developer` → `system` at the entry boundary. The IR's canonical-four-roles invariant is preserved; every provider plugin's role-handling stays unchanged. The normalize-at-entry pattern matches the existing `function` → `tool` normalization that already lives in the same function (function-role-deprecation was the precedent for entry-boundary normalization vs. IR schema bloat).
|
||||||
|
- **Why not "add `developer` to VALID_ROLES + handle in every provider":** That alternative would require:
|
||||||
|
- Expanding `VALID_ROLES` in `lib/ir/types.mjs`.
|
||||||
|
- Adding `developer` branch in `anthropic.mjs:irToAnthropic` (which would map to `[System]` annotation anyway).
|
||||||
|
- Adding `developer` branch in `codex.mjs:irToCodex` (would map to `[System]` annotation anyway).
|
||||||
|
- Adding `developer` branch in `mistral.mjs:irToMistral` (same).
|
||||||
|
- Coordinating every future role addition (e.g., if OpenAI adds another role tomorrow) across N provider plugins.
|
||||||
|
- Wider IR surface area = more drift-prone over time.
|
||||||
|
Normalize-at-entry centralizes role-spec-evolution handling in one file. ADR 0003's IR-design principle ("encode the common subset every provider plugin can consume") supports keeping the IR minimal.
|
||||||
|
- **Forward note:** Future OpenAI role additions follow the same pattern: extend `normalizeRole()`. If a role genuinely conveys provider-distinguishable semantics (e.g., a hypothetical role that meaningfully changes anthropic vs codex behavior), the calculus flips and a IR-level addition would be justified. That decision goes through a new ADR 0003 amendment.
|
||||||
|
- **Cache-key impact:** After this amendment, a request whose first message uses `role: "developer"` and one using `role: "system"` with otherwise-identical content produce the **same** IR (because normalization happens before IR construction) → the **same** cache key (per ADR 0005 cache key composition). This is intentional and matches OpenAI's own backward-compat behavior ("system message with reasoning models is treated as developer"). If a future debug session is investigating "why does my new `developer` request hit a cache entry from an old `system` request" — this is by design.
|
||||||
|
- **Tests:** Suite IR translation in `test-features.mjs` gains three pin tests: (a) `role: "developer"` → `role: "system"` translation, (b) mixed-role array including developer validates cleanly through to IR, (c) negative control — an unknown role (e.g. `"admin"`) still raises `BadRequestError`, confirming the normalize-at-entry mapping did not accidentally widen the role allow-list.
|
||||||
|
- **Authority:** OpenAI Responses API spec — developer role documented as high-priority developer-authored instructions for o1/o3+ reasoning models (https://platform.openai.com/docs/api-reference/responses). Hermes Agent / Cline / Continue.dev tracking the same convention. Reproduced live on PI230 → PI231 OLP 2026-05-27.
|
||||||
|
|
||||||
### Amendment 2 — 2026-05-24: Correct model-mapping example; document verbatim-pass-through design (D32 F2)
|
### Amendment 2 — 2026-05-24: Correct model-mapping example; document verbatim-pass-through design (D32 F2)
|
||||||
|
|
||||||
- **Finding:** Round-4 cold-audit F2 (P3 ADR example vs implementation drift) — § Decision "Required fields" item `model` reads: "The provider plugin maps this to the provider-native model identifier (e.g., `claude-sonnet-4-6` → `claude-sonnet-4-6-20260301` for Anthropic)." This is WRONG per the D17 SPOT decision (commit `cb86807`): OLP does NOT perform a model-alias mapping inside the provider plugin. `irRequest.model` is passed verbatim to the provider CLI (`claude -p --model <model>`, `codex exec --model <model>`, etc.); each provider's CLI resolves its own aliases natively per its documented behaviour.
|
- **Finding:** Round-4 cold-audit F2 (P3 ADR example vs implementation drift) — § Decision "Required fields" item `model` reads: "The provider plugin maps this to the provider-native model identifier (e.g., `claude-sonnet-4-6` → `claude-sonnet-4-6-20260301` for Anthropic)." This is WRONG per the D17 SPOT decision (commit `cb86807`): OLP does NOT perform a model-alias mapping inside the provider plugin. `irRequest.model` is passed verbatim to the provider CLI (`claude -p --model <model>`, `codex exec --model <model>`, etc.); each provider's CLI resolves its own aliases natively per its documented behaviour.
|
||||||
|
|||||||
@@ -2,6 +2,154 @@
|
|||||||
|
|
||||||
- **Date:** 2026-05-25
|
- **Date:** 2026-05-25
|
||||||
- **Status:** Accepted (D48, design-only — implementation D-days D49–D54 follow; Phase 3 close = v0.3.0)
|
- **Status:** Accepted (D48, design-only — implementation D-days D49–D54 follow; Phase 3 close = v0.3.0)
|
||||||
|
|
||||||
|
## Amendments
|
||||||
|
|
||||||
|
### Amendment 2 — 2026-05-27: v0.5.1 quota_v2 richer failure-mode shape (codex finding F3)
|
||||||
|
|
||||||
|
**Scope:** v0.5.1 hotfix extends `ProviderQuotaEntry` and `aggregateProviderQuota()` to surface richer failure-mode detail, addressing codex review finding F3 (operator cannot distinguish failure modes from the `unavailable` catch-all). Authority: ADR 0013 Rule 6 + codex review findings F1–F3.
|
||||||
|
|
||||||
|
#### 1. Extended `ProviderQuotaEntry` shape
|
||||||
|
|
||||||
|
```js
|
||||||
|
{
|
||||||
|
provider: string,
|
||||||
|
// v0.5.1: 'unreachable' added (probe enabled, creds present, but no cache + probe failed)
|
||||||
|
status: 'live' | 'stale' | 'unreachable' | 'unavailable',
|
||||||
|
reason?: string, // only when status === 'unavailable' (no API or disabled)
|
||||||
|
schema_version: string|null,
|
||||||
|
last_fresh_at: number|null,
|
||||||
|
utilization: { '5h': number|null, '7d': number|null } | null,
|
||||||
|
reset: { '5h': number|null, '7d': number|null, overall: number|null, overage: number|null } | null,
|
||||||
|
representative_claim: string|null,
|
||||||
|
fallback_percentage: number|null,
|
||||||
|
overage: { status: string|null, disabled_reason: string|null } | null,
|
||||||
|
raw_available: boolean,
|
||||||
|
// v0.5.1 (F3 — ADR 0013 Rule 6):
|
||||||
|
failure: { kind, message, backoff_until? } | null,
|
||||||
|
failure_kind: 'no_credentials'|'auth_failed'|'rate_limited'|'schema_drift'|'network'|'other' | null,
|
||||||
|
// Note: 'opt_in_off' is NOT in this enum — when probe is opted out, the row's status
|
||||||
|
// is 'unavailable' (not 'unreachable'); failure_kind stays null. Distinguishing
|
||||||
|
// "user opted out" from "provider has no API" requires reading config separately.
|
||||||
|
backoff_until: number | null, // epoch-ms when next probe attempt is allowed
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Status semantics:
|
||||||
|
- `'unavailable'` — probe disabled (`quota_probe_enabled: false`) OR provider has no public quota API (codex, mistral). `failure`, `failure_kind`, `backoff_until` are null.
|
||||||
|
- `'live'` — probe succeeded within TTL. `failure` is null.
|
||||||
|
- `'stale'` — probe failed but stale cache exists. `failure.kind` describes why the last probe failed. `last_fresh_at` is the epoch of the last successful probe. `backoff_until` tells when the next attempt is scheduled.
|
||||||
|
- `'unreachable'` (new) — probe enabled + creds present (or missing!) but no cache available + probe failed. `failure.kind` distinguishes: `no_credentials`, `auth_failed`, `rate_limited`, `schema_drift`, `network`, `other`. `utilization` and `reset` are null (no data).
|
||||||
|
|
||||||
|
#### 2. `quotaStatus()` v0.5.1 return contract
|
||||||
|
|
||||||
|
`null` is now RESERVED for `quota_probe_enabled: false` only. All other failure paths return a structured shape:
|
||||||
|
|
||||||
|
```js
|
||||||
|
null // ONLY: opt-in off
|
||||||
|
{ probe_status: 'live', ... } // cache fresh
|
||||||
|
{ probe_status: 'stale', ..., failure: { kind, message, backoff_until } } // cache stale + backoff
|
||||||
|
{ probe_status: 'unreachable', source, schemaVersion, failure: { ... } } // no cache + failed
|
||||||
|
```
|
||||||
|
|
||||||
|
The `stale: boolean` field is retained for backwards-compat (`stale: false` on live, `stale: true` on stale). New code should use `probe_status`.
|
||||||
|
|
||||||
|
#### 3. `dashboard.html` unreachable rendering
|
||||||
|
|
||||||
|
A new CSS class `.provider-row.unreachable` (red border + light red background) and `.unreachable-reason` text style handle the new status. `failure.message` and `failure_kind` are surfaced as a short text line under the provider badge. `failure.backoff_until` renders a "backoff active: Xs remaining" note if within window.
|
||||||
|
|
||||||
|
#### 4. Authority
|
||||||
|
|
||||||
|
- ADR 0013 Rule 6 (failure transparency mandate)
|
||||||
|
- Codex review findings F1 (doctor bypass), F2 (200+empty-headers → schema_drift), F3 (failure-mode collapse)
|
||||||
|
- v0.5.1 hotfix PR
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Amendment 1 — 2026-05-26: D81 Phase 5 quota_v2 shape + aggregateProviderQuota()
|
||||||
|
|
||||||
|
**Scope:** D81 (Phase 5 / ADR 0012 D81) extends the audit-query layer and dashboard-data endpoint to surface the new per-provider quota shape introduced by D80 (`lib/providers/anthropic.mjs:quotaStatus()`). This amendment documents the three new interfaces.
|
||||||
|
|
||||||
|
#### 1. `models-registry.json` — new `quota_probe` top-level key
|
||||||
|
|
||||||
|
D81 adds a `quota_probe` key at the root of `models-registry.json` per ADR 0013 Rule 5 (schema_version in registry so downstream consumers can detect schema drift):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"quota_probe": {
|
||||||
|
"schema_version": "2026-05-26",
|
||||||
|
"anthropic": {
|
||||||
|
"source": "anthropic-ratelimit-unified-headers",
|
||||||
|
"endpoint": "https://api.anthropic.com/v1/messages",
|
||||||
|
"fields_pinned": [ ...13 field names... ]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`fields_pinned` is load-bearing: if Anthropic adds/renames a header in a future CLI version, dashboard consumers comparing field-presence against this list can flag "schema drift detected" per the ADR 0013 Rule 5 drift-detection runbook. This field must be updated alongside the parser whenever a drift event occurs.
|
||||||
|
|
||||||
|
`lib/providers/anthropic.mjs` reads `quota_probe.schema_version` from the registry at call time (via `_resolveSchemaVersion()`) with the module-level `QUOTA_SCHEMA_VERSION` constant as fallback. No hard dependency on the registry — the constant is the safety net.
|
||||||
|
|
||||||
|
#### 2. `lib/audit-query.mjs` — new `aggregateProviderQuota()` export
|
||||||
|
|
||||||
|
```js
|
||||||
|
export async function aggregateProviderQuota({
|
||||||
|
providers, // Map<name, plugin> or plain object
|
||||||
|
getQuotaStatus, // optional injectable getter (name) => Promise<shape|null>
|
||||||
|
}): Promise<Array<ProviderQuotaEntry>>
|
||||||
|
```
|
||||||
|
|
||||||
|
For each provider, calls `quotaStatus()` (already cached at the plugin layer per ADR 0013 Rule 3) and normalizes to the `ProviderQuotaEntry` shape:
|
||||||
|
|
||||||
|
```js
|
||||||
|
{
|
||||||
|
provider: string,
|
||||||
|
status: 'live' | 'stale' | 'unavailable',
|
||||||
|
reason?: string, // only when status === 'unavailable'
|
||||||
|
schema_version: string|null,
|
||||||
|
last_fresh_at: number|null, // epoch-ms of last successful probe
|
||||||
|
utilization: { '5h': number|null, '7d': number|null } | null,
|
||||||
|
reset: {
|
||||||
|
'5h': number|null, '7d': number|null,
|
||||||
|
overall: number|null, overage: number|null,
|
||||||
|
} | null,
|
||||||
|
representative_claim: string|null,
|
||||||
|
fallback_percentage: number|null,
|
||||||
|
overage: { status: string|null, disabled_reason: string|null } | null,
|
||||||
|
raw_available: boolean,
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Providers returning `null` from `quotaStatus()` (codex, mistral — no public quota API; or probe disabled) produce `{ status: 'unavailable', reason: 'no public quota api or probe disabled', ...null fields }`.
|
||||||
|
|
||||||
|
Providers whose `quotaStatus()` throws produce `{ status: 'unavailable', reason: <error.message>, ...null fields }`.
|
||||||
|
|
||||||
|
This function does NOT scan ndjson files; it calls live provider plugins. It is audit-query-adjacent (normalized query shape for the dashboard layer) but not audit-derived. Query model remains Lane 2 = A (in-memory, no SQLite).
|
||||||
|
|
||||||
|
#### 3. `/v0/management/dashboard-data` and `/v0/management/quota` — new `quota_v2` field
|
||||||
|
|
||||||
|
Both endpoints now return TWO quota keys:
|
||||||
|
|
||||||
|
- **`quota`** (legacy, unchanged): `Array<{ provider, ...rawQuotaStatus, available }>`. Kept for backwards compatibility with the existing `dashboard.html` (D82 will switch consumers to `quota_v2`).
|
||||||
|
- **`quota_v2`** (D81 new): `Array<ProviderQuotaEntry>` — the normalized shape from `aggregateProviderQuota()` above. This is what D82's enriched dashboard UI will consume.
|
||||||
|
|
||||||
|
Both fields are computed from the same underlying `quotaStatus()` call. The legacy `quota` key calls `quotaStatus()` independently from `quota_v2`; since the probe is cached at the plugin layer (ADR 0013 Rule 3), the double call incurs no extra API requests.
|
||||||
|
|
||||||
|
**Deprecation timeline:** the legacy `quota` key is deprecated as of D81. Target removal: v1.0.0 or when D82 completes the dashboard migration (whichever comes first). Removal requires a separate PR with a CHANGELOG entry.
|
||||||
|
|
||||||
|
#### 4. Failure handling
|
||||||
|
|
||||||
|
`aggregateProviderQuota()` never throws to the dashboard endpoint. Per-provider failures are absorbed as `{ status: 'unavailable', reason: <error> }` entries. If `aggregateProviderQuota()` itself throws (implementation bug), `handleManagementDashboardData` and `handleManagementQuota` catch the error, log `dashboard_data_quota_v2_failed` / `management_quota_v2_failed`, and return `quota_v2: []` so the rest of the payload is unaffected.
|
||||||
|
|
||||||
|
#### 5. Authority citations for this amendment
|
||||||
|
|
||||||
|
- **ADR 0012 D81** — the D-day this amendment documents.
|
||||||
|
- **ADR 0013 Rule 5** — mandate for `quota_probe.schema_version` in `models-registry.json`.
|
||||||
|
- **D80 PR #52 commit 82d2e1c** — the producer of the `quotaStatus()` shape this amendment normalizes.
|
||||||
|
- **ADR 0008 Lane 2 = A** — query model unchanged; `aggregateProviderQuota()` does not scan ndjson.
|
||||||
|
|
||||||
|
---
|
||||||
- **Authors:** project maintainer (with AI drafting assistance)
|
- **Authors:** project maintainer (with AI drafting assistance)
|
||||||
- **Related:**
|
- **Related:**
|
||||||
- OLP v0.1 spec § 4.6 (Dashboard requirements — port from OCP with multi-provider support) and § 4.7 (observability endpoints)
|
- OLP v0.1 spec § 4.6 (Dashboard requirements — port from OCP with multi-provider support) and § 4.7 (observability endpoints)
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
# ADR 0009 — Anthropic Interactive-Mode Path (Placeholder)
|
# ADR 0009 — Anthropic Interactive-Mode Path (Placeholder)
|
||||||
|
|
||||||
- **Date:** 2026-05-25
|
- **Date:** 2026-05-25 (Placeholder); 2026-05-27 Amendment 1 (Accepted)
|
||||||
- **Status:** Draft (Placeholder — blocked on OCP ADR 0007 P0 experiment outcome; no implementation D-day scheduled until P0 lands)
|
- **Status:** **Accepted** (post-Amendment 1 — OLP self-spike supersedes OCP-wait; implementation D-day scheduled this Phase 6)
|
||||||
- **Authors:** project maintainer (with AI advisory drafting)
|
- **Authors:** project maintainer (with AI advisory drafting)
|
||||||
- **Related:**
|
- **Related:**
|
||||||
- **OCP ADR 0007** (Interactive-Mode Execution Pool, stream-json) — at `~/ocp/docs/adr/0007-interactive-mode-pool.md` on the maintainer's workstation. Pin reference at the time of this writing: OCP ADR 0007 is Draft status pending the same P0 outcome.
|
- **OCP ADR 0007** (Interactive-Mode Execution Pool, stream-json) — at `~/ocp/docs/adr/0007-interactive-mode-pool.md` on the maintainer's workstation. Pin reference at the time of this writing: OCP ADR 0007 is Draft status pending the same P0 outcome.
|
||||||
@@ -193,5 +193,117 @@ If OCP P0 fails, **this ADR is shelved** and Phase 4 ordering is unchanged.
|
|||||||
## Status transitions (recorded for clarity)
|
## Status transitions (recorded for clarity)
|
||||||
|
|
||||||
- 2026-05-25 — Created as Draft (Placeholder). OCP ADR 0007 also Draft.
|
- 2026-05-25 — Created as Draft (Placeholder). OCP ADR 0007 also Draft.
|
||||||
- _(future)_ — If OCP ADR 0007 → Accepted with a confirmed transport: this ADR moves to "Pending Phase 4 implementation D-day", maintainer decides Option 1 / 2 / 3 + lane.
|
- 2026-05-27 — Amendment 1 promotes to **Accepted**. OLP self-spike + empirical Transport-A confirmation on `claude` CLI v2.1.104 superseded wait-for-OCP. OCP is now in maintenance mode (per maintainer statement 2026-05-27 session) — OLP leads. Implementation lane: **Option 1 (parallel implementation, no warm pool, no PTY)**, scope reduced from "10-day warm pool with billing router" to "2-3-day stateless stream-json adapter".
|
||||||
- _(future)_ — If OCP ADR 0007 → Rejected: this ADR moves to "Shelved (upstream P0 failure)" with a note explaining the fallback (multi-provider routing already covers).
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Amendment 1 — 2026-05-27: Self-spike supersedes wait-for-OCP; lock Option 1 with stream-json-no-`-p` transport
|
||||||
|
|
||||||
|
### Trigger
|
||||||
|
|
||||||
|
Two findings on 2026-05-27 changed the placeholder's premises:
|
||||||
|
|
||||||
|
1. **OCP is no longer the lead project.** The maintainer stated in the 2026-05-27 session: "OCP 不会大改动…主力方向放到 OLP." The "wait-and-port" strategy implicitly assumed OCP would do the P0 first. With OCP in maintenance mode, OLP cannot wait — the 2026-06-15 Anthropic billing split is 19 days out from this amendment.
|
||||||
|
|
||||||
|
2. **OLP self-spike confirmed Transport A (`--output-format stream-json --verbose` without `-p`) emits NDJSON on Claude Code v2.1.104.** The placeholder ADR § 1.3 cited OCP's "v2.1.150 only" observation as the binding caveat. Local empirical re-test on 2026-05-27 against PI231's deployed `claude` v2.1.104 produced the full NDJSON event stream (system/init + stream_event token deltas + message_stop + result + rate_limit_event) for invocations **without** `-p`. The `claude --help` text saying "(only works with --print)" is misleading — the flags accept invocation without `-p` and produce the documented NDJSON shape.
|
||||||
|
|
||||||
|
### Additional spike findings (2026-05-27 billing classification)
|
||||||
|
|
||||||
|
A separate web/GitHub research spike on the 2026-06-15 billing classification returned:
|
||||||
|
|
||||||
|
- Anthropic's published policy is **intent-based, not mechanism-based**. The Agent SDK credit pool covers: Agent SDK Python/TypeScript packages, `claude -p`, GitHub Actions, **and "third-party apps that authenticate with your Claude subscription through the Agent SDK"**. Subscription pool covers "Claude Code in the terminal or your IDE in interactive mode."
|
||||||
|
- The third-party-app clause is the load-bearing ambiguity. OLP qualifies as a third-party app regardless of which CLI mode it spawns. If Anthropic tightens that clause from "via Agent SDK" to "any third-party app," OLP is caught regardless of `-p` flag presence.
|
||||||
|
- Behavioral fingerprinting (request cadence, OAuth-scope patterns, isTTY absence) is a separate detection vector Anthropic could deploy without policy-text changes.
|
||||||
|
|
||||||
|
The spike's recommendation: "viable bridge for ~30-60 days post-2026-06-15, NOT durable solution."
|
||||||
|
|
||||||
|
### Value re-anchoring
|
||||||
|
|
||||||
|
The placeholder framed interactive-mode as "the durable answer to keep OLP anthropic subscription value past 2026-06-15." The 2026-05-27 spike re-anchors the value:
|
||||||
|
|
||||||
|
| Value | Placeholder framing | 2026-05-27 framing |
|
||||||
|
|---|---|---|
|
||||||
|
| Keep subscription pool 6.15+ | **Primary value** | **Uncertain bridge** (30-60 day plausibility) |
|
||||||
|
| Hallucination fix (env-block / cwd injection) | (Not addressed) | **Primary value** — empirically proven |
|
||||||
|
| Cost reduction (drop default tool descriptions) | (Not addressed) | **Primary value** — ~30% input token / ~64% per-request cost reduction measured against `--system-prompt` override |
|
||||||
|
| Observability (rate_limit / cache / usage per request) | (Not addressed) | **Primary value** — NDJSON events expose data the current `--output-format text` path discards |
|
||||||
|
| Protocol foundation for future tool-call passthrough | (Not addressed) | **Secondary value** — same NDJSON parser is reusable for Phase 8+ tool passthrough work |
|
||||||
|
|
||||||
|
**Net**: even if Anthropic immediately reclassifies third-party apps to Agent SDK pool on 2026-06-15 — making the bridge worthless — the implementation still earns its keep through the other four values.
|
||||||
|
|
||||||
|
### Locked decision
|
||||||
|
|
||||||
|
**Option 1 — Parallel implementation in OLP's `lib/providers/anthropic.mjs`.**
|
||||||
|
|
||||||
|
Lane: stream-json output, no `-p` flag (Transport A confirmed), stateless per-request spawn (no warm pool, no PTY, no node-pty dependency).
|
||||||
|
|
||||||
|
Rejected lanes and why:
|
||||||
|
|
||||||
|
- **Option 2 (chain OCP)** — OCP is in maintenance mode; coupling OLP's anthropic provider to OCP's HTTP shim is the wrong direction.
|
||||||
|
- **Option 3 (both)** — premature complexity; pick the simple lane first.
|
||||||
|
- **Warm-process pool** — OLP is stateless per AGENTS.md § "No conversation state". Pool lifecycle, crash backoff, and permission auto-response from OCP ADR 0007 § 4 are unnecessary for OLP's per-request model.
|
||||||
|
- **PTY (Transport B with node-pty)** — Transport A worked; engines-bump for a native addon is unjustified when the simpler transport produces the documented NDJSON.
|
||||||
|
|
||||||
|
### Implementation scope (Option 1, this Phase 6)
|
||||||
|
|
||||||
|
| Change | Surface | Authority |
|
||||||
|
|---|---|---|
|
||||||
|
| `buildCliArgs(model)` drop `-p` and `--output-format text`; add `--output-format stream-json`, `--verbose`, `--no-session-persistence`, `--model` | `lib/providers/anthropic.mjs` | `claude --help` (v2.1.104) § `--output-format` / § `--verbose` |
|
||||||
|
| `buildCliArgs(model, systemPrompt)` accepts optional system prompt; spawns with `--system-prompt "<OLP wrapper text>"` | `lib/providers/anthropic.mjs` | `claude --help` (v2.1.104) § `--system-prompt` |
|
||||||
|
| OLP-managed system prompt construction (extract client `role:system` IR messages, prepend OLP wrapper saying "you are accessed via HTTP proxy; no local env/fs/shell access; respond directly") | `lib/providers/anthropic.mjs` `irToAnthropic` | This ADR § "OLP system prompt wrapper" below |
|
||||||
|
| New `anthropicStreamJsonChunkToIR` parser replacing/supplementing `anthropicChunkToIR` — handles NDJSON event types `system/init`, `stream_event/content_block_delta`, `assistant`, `result`, `rate_limit_event` | `lib/providers/anthropic.mjs` | This ADR § "NDJSON event handling" below |
|
||||||
|
| New tests verifying NDJSON parsing, system-prompt construction, env-block absence | `test-features.mjs` | (test surface; no external authority) |
|
||||||
|
| README troubleshooting / supported-providers § note about the bridge nature | `README.md` | (docs surface) |
|
||||||
|
|
||||||
|
The `irToAnthropic` text serialization path is preserved for client messages (`role: user`, `role: assistant`); the `role: system` extraction goes to `--system-prompt`.
|
||||||
|
|
||||||
|
### OLP system prompt wrapper
|
||||||
|
|
||||||
|
The wrapper text injected via `--system-prompt`:
|
||||||
|
|
||||||
|
```
|
||||||
|
You are accessed via the OLP HTTP proxy. You do NOT have access to any local
|
||||||
|
filesystem, working directory, shell, git status, or machine environment.
|
||||||
|
Do not infer or invent such information from any context you observe.
|
||||||
|
Respond only based on the conversation provided.
|
||||||
|
```
|
||||||
|
|
||||||
|
If the client IR request contains `role: system` messages, their concatenated `content` is appended after a blank line.
|
||||||
|
|
||||||
|
### NDJSON event handling
|
||||||
|
|
||||||
|
The parser must yield IR chunks based on the event stream:
|
||||||
|
|
||||||
|
| NDJSON event | IR yield | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `{type:"system", subtype:"init"}` | None (consumed for session_id tracking) | First event always; ignore |
|
||||||
|
| `{type:"stream_event", event:{type:"content_block_delta", delta:{type:"text_delta", text:"..."}}}` | `{type: "delta", content: "<text>"}` | Token-by-token streaming |
|
||||||
|
| `{type:"assistant"}` | None (already captured by per-token deltas) | Aggregate message; ignore (or use for verify, optional) |
|
||||||
|
| `{type:"result", subtype:"success"}` | `{type:"stop", finish_reason:"stop"}` | Marks end |
|
||||||
|
| `{type:"rate_limit_event"}` | None (consumed for audit/dashboard) | Forward to OLP audit/observability layer later (Phase 6+ enhancement) |
|
||||||
|
| `{type:"control_request"}` | Log + ignore | Per Anthropic stream-json docs |
|
||||||
|
|
||||||
|
The cache key composition (ADR 0005) is unchanged — same IR request hash; the on-the-wire format change is internal to the anthropic plugin.
|
||||||
|
|
||||||
|
### Token cost measurement (binding evidence)
|
||||||
|
|
||||||
|
Two requests against PI231 v2.1.104 on 2026-05-27 with identical user prompt `"reply: OK"`, model `claude-sonnet-4-6`:
|
||||||
|
|
||||||
|
- Default invocation (no `--system-prompt`): `cache_creation_input_tokens=4785`, `cache_read_input_tokens=11816`, total input ≈ 16,601 tokens, `total_cost_usd=$0.0216`.
|
||||||
|
- With `--system-prompt "You are a chat assistant. Respond directly."`: `cache_creation_input_tokens=1306`, `cache_read_input_tokens=9394`, total input ≈ 10,700 tokens, `total_cost_usd=$0.0078`.
|
||||||
|
|
||||||
|
**Net**: ~30% input token reduction, ~64% per-request cost reduction. Replicable by anyone with `claude` v2.1.104 + OAuth on a similar setup.
|
||||||
|
|
||||||
|
### Caveats binding the implementation
|
||||||
|
|
||||||
|
1. **Bridge value uncertain.** The 30-60 day estimate is a spike judgment, not Anthropic-confirmed. Implementation must continue to function correctly if Anthropic re-routes this path to Agent SDK billing on 2026-06-15 — the only consequence is the bridge value disappears, but the other four values (hallucination / cost / observability / protocol foundation) remain.
|
||||||
|
2. **No claim about durability.** This ADR amends only as far as "the bridge is worth the 2-3 day investment given the orthogonal values." A future ADR (likely Phase 7 sandbox-runtime + Phase 8 multi-provider robustness) will revisit the anthropic provider's strategic role once the post-2026-06-15 picture clarifies.
|
||||||
|
3. **Sandbox-runtime still required for real multi-tenant deployment.** Per the 2026-05-27 session prior-art search, Anthropic's official multi-tenant answer is `@anthropic-ai/sandbox-runtime` (OS-level isolation). This ADR does NOT substitute for that work; sandbox-runtime remains Phase 7 scope and is a hard prerequisite before any cloud deployment per `docs/plans/cloud-deployment-family.md`.
|
||||||
|
4. **CLI version pin guidance.** Stream-json without `-p` was confirmed on v2.1.104. Future versions may tighten this; the plugin's spawn should emit a warning to OLP server log if `claude --version` falls outside a `v2.1.100`–`v2.1.149` range. Hard failure on out-of-range version is NOT required; warning is sufficient for v0.6.x.
|
||||||
|
|
||||||
|
### Updated authority citations (in addition to placeholder § Authority citations)
|
||||||
|
|
||||||
|
- **OLP self-spike — 2026-05-27 session live transcripts** (PI231 ssh; `claude -p --output-format stream-json --verbose` and `claude` no-`-p` variants captured in session log; retained in cc-mem post-implementation).
|
||||||
|
- **P0 billing classification spike — 2026-05-27** subagent transcript; sources include Anthropic published docs at `code.claude.com/docs/en/headless`, `support.claude.com/en/articles/15036540`, `support.claude.com/en/articles/11145838`.
|
||||||
|
- **claude CLI v2.1.104 `--help`** (live capture on PI231) § `--output-format`, § `--verbose`, § `--system-prompt`, § `--no-session-persistence`.
|
||||||
|
- **CLAUDE.md `release_kit.phase_rolling_mode.current_phase`** — Phase 6; this ADR consumes a Phase 6 D-day per the amendment, NOT a Phase 4 D-day (the placeholder's hypothetical scheduling).
|
||||||
|
|||||||
@@ -214,6 +214,24 @@ alongside the "using server-advertised key" notice.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## Deployment configurations (D76 amendment, 2026-05-26)
|
||||||
|
|
||||||
|
Original ADR 0011 referenced a `BIND_ADDRESS` concept that did not exist in the v0.4.0–v0.4.2 codebase — the server was hard-coded to `server.listen(PORT, '127.0.0.1', ...)`. D76 closes this gap by adding the `OLP_BIND` env var (default `127.0.0.1`), making the deployment-context discussion below operational rather than aspirational.
|
||||||
|
|
||||||
|
Three deployment configurations are supported:
|
||||||
|
|
||||||
|
| `OLP_BIND` value | Reachability | Anonymous-key publication |
|
||||||
|
|---|---|---|
|
||||||
|
| `127.0.0.1` (default) | Loopback only | Safe with any auth posture (no LAN exposure at all) |
|
||||||
|
| RFC1918 IP / tailnet IP / `0.0.0.0` on a trusted LAN | LAN clients only | Safe when `advertise_anonymous_key: true` — the documented "trusted-LAN" zero-config family onboarding flow |
|
||||||
|
| Public IP / `0.0.0.0` on a public-facing host | Public internet | **Incompatible with `advertise_anonymous_key: true`.** Operator MUST keep `advertise_anonymous_key: false` (default). |
|
||||||
|
|
||||||
|
The server emits a startup warn event `anonymous_key_advertised_with_lan_bind` when `OLP_BIND` is non-loopback AND `advertise_anonymous_key: true` (per the `lib/keys.mjs` + `server.mjs` checks). The warn is a **checkpoint, not a hard gate** — the server cannot tell from the bind address alone whether the operator is on a trusted LAN (RFC1918 / tailnet) or has accidentally exposed a public IP. The Re-evaluation trigger #1 below escalates to a hard gate when OLP gains a public-internet deployment mode.
|
||||||
|
|
||||||
|
`olp-connect <ip>` consumes `/health.anonymousKey` over the network — therefore requires `OLP_BIND` to include the LAN interface on the server side. Without setting `OLP_BIND=<lan-ip>` (or `0.0.0.0`), `olp-connect <ip>` will fail with `connect ECONNREFUSED` because the server only accepts loopback connections.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## Re-evaluation triggers
|
## Re-evaluation triggers
|
||||||
|
|
||||||
Re-open this ADR when ANY of the following fires:
|
Re-open this ADR when ANY of the following fires:
|
||||||
|
|||||||
@@ -0,0 +1,147 @@
|
|||||||
|
# ADR 0012 — Phase 5 Charter: Provider Quota Probes + Dashboard Enrichment
|
||||||
|
|
||||||
|
**Status:** Accepted (Phase 5 open as of 2026-05-26)
|
||||||
|
**Date:** 2026-05-26
|
||||||
|
**D-day:** D79 (charter + ADR 0002 Amendment 8 + ADR 0013 land together as the constitutional layer of Phase 5)
|
||||||
|
|
||||||
|
## Amendments
|
||||||
|
|
||||||
|
### Amendment 1 — 2026-05-26: D84 Mistral probe NO-GO (post-D79-close spike)
|
||||||
|
|
||||||
|
The D-day table originally listed D84 as "optional, depends on D79-close 30-min Mistral docs spike". The spike completed 2026-05-26 with verdict **NO-GO** — Mistral does not expose a programmatic quota/usage endpoint **accessible to Vibe / Le Chat member / La Plateforme API keys** (the key tier OLP uses for spawning the `vibe` CLI):
|
||||||
|
|
||||||
|
- `docs.mistral.ai/api` (the public API spec) covers Chat, FIM, Embeddings, Classifiers, Files, Models, Batch, OCR, Audio, Events, Beta (Agents/Conversations/Libraries/Workflows/Observability). No usage/quota/credits/billing/limits endpoint accessible to a member API key.
|
||||||
|
- Direct probe `https://api.mistral.ai/v1/usage` returns 404.
|
||||||
|
- Mistral's "Limits and Usage" help article documents limit viewing via the `admin.mistral.ai/plateforme/limits` web console.
|
||||||
|
- No `x-ratelimit-*` response headers documented on `/v1/chat/completions`. (Third-party summaries mentioning these headers are unsourced — appears to be OpenAI-convention extrapolation.)
|
||||||
|
- OLP `lib/providers/mistral.mjs` already records this independently — DL-7 comment: "If quota/budget API surfaces in Le Chat Pro, pin the endpoint here."
|
||||||
|
|
||||||
|
**Out-of-scope but worth pinning for future revisit.** Mistral's [Admin API](https://docs.mistral.ai/admin/security-access/admin-api) DOES expose programmatic "Billing and usage queries", and the [Usage limits docs](https://docs.mistral.ai/admin/user-management-finops/usage-limits) describe usage/cost queries via that surface. The Admin API requires an **org-admin scoped API key** (separate from the member key OLP uses). For OLP's family-tier deployment posture (a maintainer's personal Le Chat Pro / La Plateforme account, not an organization's admin console), provisioning + storing an org-admin token raises the credential-scope ceiling beyond what the trusted-LAN deployment context (ADR 0011) was designed for. The NO-GO at v0.5.0 is therefore "out of scope for OLP's current deployment posture", NOT "Mistral has no programmatic surface". If the deployment posture expands to an org-admin context (e.g., a small-business multi-user deployment), this decision should be re-evaluated.
|
||||||
|
|
||||||
|
**Disposition:**
|
||||||
|
- D84 row dropped from D-day plan (struck through below).
|
||||||
|
- Mistral dashboard row in D82 UI shows "spend tracking only" badge sourced from `audit-query.mjs` aggregates (request count, estimated cost from `estimateCost()`).
|
||||||
|
- `DL-7` in `mistral.mjs` is the documented re-entry point if Mistral ever publishes a usage endpoint.
|
||||||
|
- Phase 5 total D-day budget revised: ~5 D-days (down from ~6).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Phase 4 (ADR 0010) shipped OLP's operator + client UX layer — `bin/olp` operator CLI, `olp doctor` framework, `olp-connect` zero-config IDE wiring, OpenClaw `/olp` slash commands, anonymous-key deployment-context limits, SSE heartbeat. v0.4.4 is the current shipped state. Phase 4 closed every gap on the OCP-feature-parity matrix EXCEPT one: **live quota / plan-usage surfacing**.
|
||||||
|
|
||||||
|
Today `lib/providers/anthropic.mjs:445` has a stub `quotaStatus()` returning `null` (D4 placeholder). The OLP dashboard's quota panel renders "—" for all providers. OCP, in contrast, exposes a live "39% session / 30% weekly" panel — the maintainer uses this multiple times per day to decide when to throttle voluntary `claude -p` traffic away from interactive sessions. OLP cannot become an OCP successor in practice (vs. just feature-parity-on-paper) until quota surfacing works.
|
||||||
|
|
||||||
|
A pre-flight institutional-knowledge audit (2026-05-26 — see `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`) confirmed:
|
||||||
|
|
||||||
|
1. **The OCP probe still works today** — Anthropic returns the same `anthropic-ratelimit-unified-*` headers on every `POST /v1/messages` call. Tested live 2026-05-26 from PI231 OAuth credentials.
|
||||||
|
2. **Schema added 3 fields since OCP's 2026-04 capture** — `5h-status`, `7d-status` (per-window status), `overage-reset` (only on active overage). No fields removed or renamed.
|
||||||
|
3. **Verification protocol has shifted** — Claude Code v2.1.x is now a **compiled binary** (Mach-O / ELF), not bundled JS. OCP's "grep cli.js" approach no longer applies; the replacement protocol is `strings` against the binary + periodic live probe diff.
|
||||||
|
4. **OAuth refresh path unchanged** — `platform.claude.com/v1/oauth/token` + `9d1c250a-...` client_id + 60s-3600s exponential backoff.
|
||||||
|
|
||||||
|
The audit makes Phase 5 implementation low-risk: this is a port of a working OCP function, not a re-derivation. The work is mechanical + adapter-layer plumbing into OLP's plugin contract.
|
||||||
|
|
||||||
|
A parallel maintainer request (2026-05-26, with reference screenshot of claude.ai/settings/usage) asked for Claude.ai-style dashboard enrichment: per-row utilization bars, reset countdown, 1-minute auto-refresh, manual refresh button. P5-1 (probe) + P5-2 (dashboard) together unlock both: the data plus the surface. v1.x roadmap #8 is closed by P5-2.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
Phase 5 scope is **Provider quota probes + dashboard enrichment**. The phase opens 2026-05-26 with D79 (this charter + ADR 0002 Amendment 8 + ADR 0013 OAuth READ-ONLY consumption rules). Phase 5 close ships v0.5.0; per `CLAUDE.md release_kit.phase_rolling_mode`, the close PR is maintainer-triggered.
|
||||||
|
|
||||||
|
### In scope — Phase 5 D-day plan (~6 D-days)
|
||||||
|
|
||||||
|
| D-day | Deliverable | Authority | Estimate |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **D79** | This charter ADR 0012 + ADR 0002 Amendment 8 (direct-API READ-ONLY) + ADR 0013 (OAuth READ-ONLY consumption rules + schema-drift mitigation) + `package.json` `current_pre_release_identifier` → `0.5.0-phase5` + `CLAUDE.md release_kit.phase_rolling_mode.current_phase` → Phase 5 | This charter + audit memory | 0.5d |
|
||||||
|
| **D80** | `lib/providers/anthropic.mjs:quotaStatus()` ported from OCP `server.mjs:842-1109` — full probe with macOS-keychain auth read added (existing OLP reader only handles env + `.credentials.json`) + 5min cache + 60s-3600s refresh backoff + stale-cache-on-429 + all 13 headers parsed (including new 5h-status / 7d-status / overage-reset) | Port OCP probe + ALIGNMENT.md Rule 2 exemption per ADR 0002 Amendment 8 + audit memory | 2d |
|
||||||
|
| **D81** | `lib/audit-query.mjs` + `/v0/management/dashboard-data` extended to surface the new quota shape per provider (utilization, reset, representative-claim, fallback-percentage, overage-status). Audit-query stays in-memory scan per ADR 0008 Lane 2 = A (no SQLite). Schema migration documented in ADR 0008 § Amendment | ADR 0008 + this charter | 1d |
|
||||||
|
| **D82** | `dashboard.html` Claude.ai-style restructure — per-provider rows replace the current single Quota panel; each row: provider badge, model placeholder, utilization bar (5h + 7d), reset countdown ("Your limit will reset at HH:MM AM/PM" format from the user-shared claude.ai screenshot), status badge, representative-claim hint. 1-minute auto-refresh via `setInterval` with `document.visibilityState` guard. Manual refresh button calls `/v0/management/dashboard-data` directly | v1.x roadmap #8 + maintainer reference screenshot | 1.5d |
|
||||||
|
| **D83** | Test coverage — Suite 38 quota-probe unit tests (mock HTTP server returning the 13 headers; assert parse + cache + backoff + stale-on-429); Suite 39 dashboard rendering smoke (curl `/dashboard` after pre-seeding mock quota cache; assert HTML contains expected utilization strings); update Suite 33 doctor checks for new `anthropic.quota_probe_reachable` check | Test convention from existing suites | 1d |
|
||||||
|
| ~~D84~~ **DROPPED** | ~~Mistral `quotaStatus()` port — depends on D79-close spike~~ **NO-GO per 2026-05-26 spike (see § Amendment 1).** Mistral dashboard row in D82 shows "spend tracking only" badge sourced from `audit-query.mjs` aggregates. `DL-7` hook point in `mistral.mjs` already marks the location for future upgrade if Mistral ever publishes a usage endpoint. Codex permanently skipped (no public API). | n/a (dropped) | 0d |
|
||||||
|
| **close** | v0.5.0 release PR — `package.json` `0.4.4 → 0.5.0`, CHANGELOG promotion, `release_kit.phase_rolling_mode.current_pre_release_identifier` advance to Phase 6 token | `CLAUDE.md release_kit overlay` | maintainer-triggered |
|
||||||
|
|
||||||
|
### Out of Phase 5 scope (with explicit triggers)
|
||||||
|
|
||||||
|
#### `X-OLP-Cost-USD` per-request response header
|
||||||
|
|
||||||
|
**Status:** Deferred to Phase 6. Was listed in ADR 0010 § Out-of-scope as "Phase 5 prerequisite". The prerequisite (provider-cost weights table) is non-trivial — needs per-(provider, model) `input_cost_per_1k_tokens` / `output_cost_per_1k_tokens` / `cache_read_discount` data sourced from each provider's published pricing page. Phase 5 already pulls in two new ADRs; adding a third data-onboarding ADR is scope creep.
|
||||||
|
|
||||||
|
**Re-open condition.** Phase 6 unless a maintainer reports a cost-attribution debugging need that warrants pulling forward.
|
||||||
|
|
||||||
|
#### `context_window_exceeded` fallback trigger (LiteLLM prior-art)
|
||||||
|
|
||||||
|
**Status:** Deferred. ADR 0010 listed this as opportunistic-in-Phase-5 unless the trigger fires sooner. The trigger has not fired in Phase 4 production traffic. Continue to defer.
|
||||||
|
|
||||||
|
#### per-(provider, model) live stats Map (replacing audit-query scan)
|
||||||
|
|
||||||
|
**Status:** Deferred. Current scan latency is ~20ms at 7-day depth. Acceptable until volume grows (>100k requests/day). Re-evaluate at Phase 6 if dashboard latency degrades.
|
||||||
|
|
||||||
|
#### Anthropic interactive-mode P0 (ADR 0009)
|
||||||
|
|
||||||
|
**Status:** Still trigger-gated on Anthropic's 2026-06-15 billing-split rollout. Phase 5 does NOT depend on P0 — the quota probe reads `anthropic-ratelimit-unified-*` headers regardless of which billing pool the spawn path consumes. If P0 succeeds Phase 7+ Phase 5's probe code remains unchanged; if P0 fails Phase 5's probe code remains unchanged. The probe is billing-pool-agnostic because the headers are subscription-pool metadata, not Agent-SDK-Credit metadata.
|
||||||
|
|
||||||
|
#### `/v1/messages` Anthropic-shape entry surface
|
||||||
|
|
||||||
|
**Status:** Still deferred per ADR 0010 § Out-of-scope. No change in Phase 5.
|
||||||
|
|
||||||
|
#### v1.x roadmap #3 / #5 / #6
|
||||||
|
|
||||||
|
**Status:** Still trigger-gated per `docs/v1x-roadmap.md`. None has fired. Continue to defer.
|
||||||
|
|
||||||
|
### Opportunistic Phase 5 micro-additions (not blocking)
|
||||||
|
|
||||||
|
Items small enough to land alongside a planned D-day without scope creep, if encountered:
|
||||||
|
|
||||||
|
- README § Dashboard screenshot update (post-P5-2 enrichment) — capture from MacBook test path per `~/.cc-rules/memory/feedback/mac_mini_never_for_testing.md`.
|
||||||
|
- `olp usage` CLI subcommand (bin/olp.mjs) surfaces the parsed quota shape in terminal form. Already partially exists (cmdUsage in bin/olp.mjs); confirm payload alignment after D80.
|
||||||
|
- Add `claude_code_oauth_client_id` config override in `~/.olp/config.json` so power users can override the hardcoded `9d1c250a-...` UUID without env-var fiddling. Mirrors compiled binary's `CLAUDE_CODE_OAUTH_CLIENT_ID` env support.
|
||||||
|
- `docs/provider-audits/anthropic.md` re-capture with current `claude --version` (v2.1.142 MacBook / v2.1.150 PI231) + binary distribution layout note.
|
||||||
|
|
||||||
|
### Exit gate — v0.5.0 close criteria
|
||||||
|
|
||||||
|
1. D79 — D84 all merged with fresh-context opus reviewer APPROVE per Iron Rule 10.
|
||||||
|
2. CI green on every D-day merge commit and on the v0.5.0 release commit head. `alignment.yml` blacklist re-confirmed (no new hallucinated tokens introduced).
|
||||||
|
3. README § Quota / Plan Usage section present with screenshot of the enriched dashboard. README § Supported Providers table updated to note "quota probe: anthropic ✅, mistral ⚠️/✅ (D84 outcome), codex ❌ (no public API)".
|
||||||
|
4. ADR 0012 (this charter) + ADR 0002 Amendment 8 + ADR 0013 (OAuth READ-ONLY consumption) on disk.
|
||||||
|
5. `CHANGELOG.md "Unreleased"` promoted to `"## v0.5.0 — <date>"` with D79 — D84 entries.
|
||||||
|
6. `package.json` bumped to `0.5.0`.
|
||||||
|
7. `CLAUDE.md release_kit.phase_rolling_mode.current_phase` advances `Phase 5 → Phase 6`; `current_pre_release_identifier` advances `0.5.0-phase5 → 0.6.0-phase6`.
|
||||||
|
8. Standing autopilot grant covers D-day-by-D-day execution; v0.5.0 close PR is maintainer-triggered.
|
||||||
|
9. Live MacBook E2E verification — dashboard renders enriched panel with real quota data (probe live, not mocked).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
**Positive.**
|
||||||
|
|
||||||
|
- OLP finally has the load-bearing observability OCP had — maintainer can see live "39% session / 30% weekly" and decide whether voluntary `claude -p` traffic stays or moves.
|
||||||
|
- Family members on the LAN see real reset times instead of "—", which makes the "wait 2 hours" guidance concrete vs. abstract.
|
||||||
|
- The institutional-knowledge audit captured the schema in a memory file pinned with date stamps — future ports (mistral, future provider) re-use the verification protocol without re-deriving.
|
||||||
|
- v1.x roadmap #8 (Dashboard enrichment per Claude.ai-style usage page) closes inside Phase 5 rather than waiting for a separate phase.
|
||||||
|
- Compiled-binary-distribution awareness ("no more cli.js to grep") is now codified in OLP governance; the next time Anthropic ships a major CC version, the verification protocol is already written.
|
||||||
|
|
||||||
|
**Negative.**
|
||||||
|
|
||||||
|
- ADR 0002 gains another amendment (Amendment 8). The constitution surface area for `anthropic.mjs` grows. Counter-pressure: the alternative (probe lives in `server.mjs`, like OCP) violates the plugin-architecture principle that per-provider knowledge stays in `lib/providers/`. Amendment 8 is the smaller violation.
|
||||||
|
- The probe makes one `/v1/messages` call per 5min cache miss. That's ~12 calls/hour worst case across the whole proxy (probe is per-credentials, not per-key). With `max_tokens: 1` the cost is < $0.01/day at family-scale traffic. Negligible but not zero.
|
||||||
|
- Schema-drift risk over the long horizon. Anthropic could rename or remove headers in a future version. The mitigation protocol (strings + live probe diff) is in place, but it's a manual check — needs to be invoked by the maintainer or scheduled.
|
||||||
|
- Dashboard refactor introduces a breaking-change risk for the existing dashboard.html consumers (none today, but conceptually). Bumping to v0.5.0 signals this clearly.
|
||||||
|
|
||||||
|
**Neutral.**
|
||||||
|
|
||||||
|
- Phase 5 has more ADR work than Phase 4 (3 governance docs vs. 2). The constitutional layer is deliberately heavier because direct-API access is the single biggest authority decision since the plugin contract itself.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Authority + cross-references
|
||||||
|
|
||||||
|
- **Iron Rule 11 (IDR)** — Phase 5 ships across 6 D-days, each a minimum reviewable unit. The governance trio (this ADR + Amendment 8 + ADR 0013) lands at D79 as a single coupled commit (reviewing them separately cannot verify consumer-producer alignment), per ADR 0002 Amendment 7's precedent.
|
||||||
|
- **Iron Rule 12 (prior-art search)** — discharged via the audit memory at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`. Memory committed prior to D80 implementation.
|
||||||
|
- **ALIGNMENT.md Rule 1 (citation)** — D80 commit must cite compiled-binary `strings` evidence per audit memory § Path A (Claude Code v2.1.x has no traditional `§ section` structure because it is a Mach-O / ELF compiled binary) plus the audit memory file path. Live-probe transcript MUST be included in the commit body.
|
||||||
|
- **ALIGNMENT.md Rule 2 (provider-CLI-as-authority)** — direct-API access bypasses the spawn-binary contract. Amendment 8 is the explicit exemption. Without Amendment 8, the D80 commit is unalignable.
|
||||||
|
- **ALIGNMENT.md Rule 5 (CI alignment.yml)** — must continue to pass. `api.anthropic.com/v1/messages` is NOT on the blacklist (correct — that's the real endpoint). The hallucinated `/api/oauth/usage` IS on the blacklist (transitive from OCP) and must remain.
|
||||||
|
- **ADR 0002 Amendment 8** — companion ADR. Direct-API access scoping; READ-ONLY constraint; opt-in via config flag (default off).
|
||||||
|
- **ADR 0013** — companion ADR. OAuth credentials shared between spawn path + probe path; refresh backoff; schema-drift mitigation protocol.
|
||||||
|
- **CLAUDE.md release_kit** — Phase boundary triggers maintainer-led version bump. D-day commits within Phase 5 stay under "Unreleased". `0.5.0-phase5` is the pre-release identifier during the phase.
|
||||||
@@ -0,0 +1,197 @@
|
|||||||
|
# ADR 0013 — OAuth READ-ONLY Consumption Rules + Schema-Drift Mitigation Protocol
|
||||||
|
|
||||||
|
**Status:** Accepted (2026-05-26)
|
||||||
|
**Date:** 2026-05-26
|
||||||
|
**D-day:** D79 (lands alongside ADR 0012 Phase 5 charter + ADR 0002 Amendment 8 as the constitutional trio of Phase 5)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
ADR 0002 Amendment 8 permits `quotaStatus()` to call provider HTTP APIs directly, subject to a READ-ONLY constraint. That Amendment opens the door but does not specify HOW READ-ONLY discipline is preserved across credential lifecycle events (refresh, expiry, revocation), nor how OLP detects when the upstream API schema drifts. ADR 0013 fills both gaps.
|
||||||
|
|
||||||
|
The motivating concern: a provider that ships its CLI as a **compiled native binary** (Anthropic Claude Code v2.1.x is now Mach-O on macOS, ELF on Linux) closes off the previous schema-verification path (grep `cli.js`). If OLP's probe parser silently breaks because a header was renamed, the dashboard shows stale or wrong numbers, and the maintainer's load-bearing throttling decision is based on bad data. This ADR establishes the verification protocol that survives the binary-distribution shift.
|
||||||
|
|
||||||
|
A second motivating concern: the OAuth credentials used by the probe are the SAME credentials the spawn path uses for `claude -p`. Both paths consume them; the probe must not interfere with the spawn path's ability to refresh or invalidate them. Concretely: the probe must not write to the credentials artifact, must not race the spawn path on refresh, and must not amplify a 429 into a refresh storm.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
### Rule 1 — Credential reuse is mandatory
|
||||||
|
|
||||||
|
The probe MUST consume the same OAuth artifact the spawn path reads via the plugin's `readAuthArtifact()`. No new OAuth grant. No alternate credential store. No environment-variable-only fallback (env var `CLAUDE_CODE_OAUTH_TOKEN` is supported as an override consistent with the spawn path, but is not the probe's primary source).
|
||||||
|
|
||||||
|
Precedence order (mirrors OCP `getOAuthCredentials` 2026-04-stable):
|
||||||
|
|
||||||
|
1. `process.env.CLAUDE_CODE_OAUTH_TOKEN` if non-empty (manual override; common in CI / dev / one-off debugging).
|
||||||
|
2. `~/.claude/.credentials.json` → `claudeAiOauth.accessToken` (Linux + macOS without keychain access).
|
||||||
|
3. macOS Keychain: `security find-generic-password -a "${USER}" -s "Claude Code-credentials" -w` (preferred on macOS — current `lib/providers/anthropic.mjs` only covers (1) + (2); D80 adds (3)).
|
||||||
|
|
||||||
|
Rationale: a separate OAuth grant would require the maintainer to repeat `claude setup-token` against an OLP-specific scope, doubling credential exposure and divergence risk. Reusing the spawn path's credentials guarantees the probe never has more permission than the spawn path itself.
|
||||||
|
|
||||||
|
### Rule 2 — READ-ONLY at the wire
|
||||||
|
|
||||||
|
The probe MUST issue exactly one HTTP request per cache miss. Method MAY be POST (Anthropic's ratelimit headers come back on `POST /v1/messages`; this is the only way to read them). Request body MUST minimise side effects:
|
||||||
|
|
||||||
|
- `max_tokens: 1` (cost: ~$0.000001 per probe)
|
||||||
|
- `messages: [{role: "user", content: "hi"}]` (any minimal valid payload)
|
||||||
|
- Model: cheapest available in the plan (`claude-haiku-4-5` at v0.5.0)
|
||||||
|
- Do NOT include `system` prompts, `tools[]`, `tool_choice`, large content arrays, or anything that the upstream might bill differently.
|
||||||
|
|
||||||
|
The probe MUST discard the response body. Only response headers are parsed.
|
||||||
|
|
||||||
|
The probe MUST NOT call any other HTTP path on the provider's API. No `/v1/models` enumeration, no admin endpoints, no `/v1/messages/<id>` retrievals. The only permitted endpoint is `POST /v1/messages`.
|
||||||
|
|
||||||
|
### Rule 3 — Cache TTL and refresh discipline
|
||||||
|
|
||||||
|
- Cache TTL: 5 minutes. Cache miss triggers a real probe. Cache hit returns the cached value.
|
||||||
|
- The dashboard refreshes every 1 minute; that's served from the cache between probes. A manual refresh button MAY force-clear the cache (per maintainer request 2026-05-26); ADR 0012 D82 documents the button.
|
||||||
|
- On refresh failure (token expired, 401/403/429, network error), the probe schedules an exponential backoff: minimum 60s, maximum 3600s. The cache entry is NOT invalidated during backoff; `quotaStatus()` returns the stale cache marked `{ stale: true, last_fresh_at: <epoch> }`. If no stale entry exists, returns an `unreachable` shape (v0.5.1+) rather than `null`.
|
||||||
|
- Successive successful probes reset the backoff to the minimum.
|
||||||
|
- Token refresh (`POST https://platform.claude.com/v1/oauth/token`) follows the same backoff discipline. The probe MUST NOT refresh a token more than once per backoff window. The refresh path is shared with the spawn path; both observe the same backoff.
|
||||||
|
- **All consumers of `quotaStatus()`, including `olp doctor` checks, MUST route through `quotaStatus()` and MUST NOT call `_probeOnce()` directly.** `_probeOnce()` is an internal implementation detail. Routing doctor checks through `quotaStatus()` ensures the cache+backoff discipline is enforced for every caller — including operators running `olp doctor` in a debug loop. (Clarification added v0.5.1 to address codex finding F1: the original doctor check bypassed backoff by calling `_probeOnce` directly.)
|
||||||
|
|
||||||
|
### Rule 4 — Opt-in via config
|
||||||
|
|
||||||
|
A new config field at `~/.olp/config.json` controls per-provider opt-in:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"providers": {
|
||||||
|
"anthropic": {
|
||||||
|
"enabled": true,
|
||||||
|
"quota_probe_enabled": false
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Default: `false`. The maintainer must explicitly opt in after credentials are configured. Reasoning: a fresh install on a machine without OAuth credentials should not bombard `api.anthropic.com` with 401-bound probes.
|
||||||
|
|
||||||
|
`olp doctor` adds a per-provider check `<provider>.quota_probe_reachable` (only runs if `quota_probe_enabled: true`). Failed check provides a `next_action.ai_executable[]` recipe to either re-authenticate or disable the probe.
|
||||||
|
|
||||||
|
### Rule 5 — Schema-drift mitigation protocol (minimum-viable-schema gate)
|
||||||
|
|
||||||
|
The CC binary-distribution shift means OCP's "grep cli.js" verification is no longer applicable. OLP adopts a two-path protocol for proactive monitoring, AND enforces a minimum-viable-schema gate at parse time:
|
||||||
|
|
||||||
|
**Minimum-viable-schema gate (v0.5.1+).** `_probeOnce()` requires at least these 4 fields present (non-null after parse) before treating a response as successful:
|
||||||
|
- `anthropic-ratelimit-unified-5h-utilization`
|
||||||
|
- `anthropic-ratelimit-unified-5h-reset`
|
||||||
|
- `anthropic-ratelimit-unified-7d-utilization`
|
||||||
|
- `anthropic-ratelimit-unified-7d-reset`
|
||||||
|
|
||||||
|
If any of these 4 is absent, `_probeOnce()` classifies the probe as a schema-drift failure (`failureKind = 'schema_drift'`), schedules backoff, and returns `null`. This means a 200 OK with zero `anthropic-ratelimit-*` headers (e.g. a server-side change, a proxy stripping headers, or a mock returning `{}`) is immediately caught as drift rather than silently cached as "live" data. The other 9 fields are tolerated as absent (overage fields are conditional; top-level status fields may be absent on edge cases). The 5h/7d core 4 are load-bearing — the dashboard's progress bars depend on them. (Gate added v0.5.1 to address codex finding F2.)
|
||||||
|
|
||||||
|
The CC binary-distribution shift means OCP's "grep cli.js" verification is no longer applicable. OLP adopts a two-path protocol:
|
||||||
|
|
||||||
|
**Path A — Compiled-binary string extraction.** Run `strings` over the platform-specific binary in the claude-code distribution. Captures all hardcoded header names the binary expects:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
BIN_DIR=$(npm root -g)/@anthropic-ai/claude-code/node_modules/@anthropic-ai/claude-code-*
|
||||||
|
strings "$BIN_DIR/claude" | grep -iE "anthropic-ratelimit|/v1/(messages|oauth)|platform\.claude\.com"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Path A prerequisites.** GNU or BSD `strings` (part of binutils/coreutils on Linux + macOS — always present on a normal developer machine; Windows requires WSL or `binutils-mingw`). A locally installed Claude Code v2.1.x (npm-global or volta-managed). A reviewer without `claude` installed can still run Path B but Path A is gated on having the binary on disk. A future Claude Code version that ships as a different distribution shape (e.g. Rust binary, statically linked Go) keeps the protocol valid: `strings` works on any ELF/Mach-O regardless of compile source.
|
||||||
|
|
||||||
|
**Path B — Live API probe.** Run the actual probe against `api.anthropic.com` with valid OAuth credentials. Captures what the server returns today:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -s -i -m 10 -X POST https://api.anthropic.com/v1/messages \
|
||||||
|
-H "Authorization: Bearer $TOKEN" \
|
||||||
|
-H "anthropic-beta: oauth-2025-04-20" \
|
||||||
|
-H "anthropic-version: 2023-06-01" \
|
||||||
|
-H "Content-Type: application/json" \
|
||||||
|
-d '{"model":"claude-haiku-4-5","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' \
|
||||||
|
| grep -iE "^anthropic-ratelimit"
|
||||||
|
```
|
||||||
|
|
||||||
|
Path A tells you what the client expects. Path B tells you what the server actually emits. The diff is the actionable schema delta.
|
||||||
|
|
||||||
|
**Required cadence.** The diff MUST be re-run at every major `claude --version` bump (v2.x → v3.x is the next trigger). The current pinned schema lives at `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`. After re-verification, that memory file MUST be updated (or a successor file written with a new date stamp; the old one cross-linked).
|
||||||
|
|
||||||
|
**Trigger for re-running the diff.** There is no automated detector for a major `claude --version` bump at v0.5.0. Three explicit hooks share this responsibility:
|
||||||
|
|
||||||
|
1. **Annual Alignment Audit** (`ALIGNMENT.md` § Annual Alignment Audit, every 14 May) — diff is mandatory as part of the audit checklist.
|
||||||
|
2. **`olp doctor anthropic.quota_probe_reachable` failure** — if the probe returns non-2xx for any reason other than 401/403/429/network (typical schema breaks manifest as 422 or 400), `olp doctor` surfaces a `kind: fix_provider` recipe whose first step is "re-run the Rule 5 dual-path diff".
|
||||||
|
3. **Manual maintainer attention at a major Claude Code release** — if the maintainer sees a major version bump in `claude --version`, kick off the diff before the next Phase opens. Rolling-mode discipline (CLAUDE.md release_kit) means major-version bumps usually intersect with Phase boundaries.
|
||||||
|
|
||||||
|
If the diff is missed across a major version bump, the failure mode is graceful degradation: the parser silently drops unknown headers; the dashboard shows older values (cached stale) or `null` per Rule 3; `olp doctor` surfaces the staleness.
|
||||||
|
|
||||||
|
**Required action on drift detection.** If a header is renamed or removed:
|
||||||
|
|
||||||
|
1. File a Phase-N issue tagging the maintainer.
|
||||||
|
2. Update the parser in `lib/providers/anthropic.mjs:quotaStatus()` to handle both names (graceful migration), prefer the new name.
|
||||||
|
3. Update the audit memory file with a "drift event" section recording: date, old field, new field, evidence URLs.
|
||||||
|
4. Bump the `models-registry.json` `quota_probe.schema_version` (NEW field added at D80) so downstream consumers can detect.
|
||||||
|
|
||||||
|
If a new header appears in the live response that the parser doesn't read: low-priority enhancement; add to the parser, document in the audit memory, no schema_version bump required.
|
||||||
|
|
||||||
|
### Rule 6 — Failure transparency
|
||||||
|
|
||||||
|
The probe's failure modes are visible to the operator:
|
||||||
|
|
||||||
|
- `/v0/management/dashboard-data` includes per-provider `{ quota_probe: { status: 'ok' | 'stale' | 'failed' | 'disabled', last_fresh_at, last_error?, backoff_until? } }`.
|
||||||
|
- `olp doctor` surfaces probe failure as `kind: fix_oauth` (if 401/403) or `kind: fix_provider` (if 429 with no stale cache or network error).
|
||||||
|
- The dashboard row badge shows the status; clicking a failed row shows the last error (truncated to 200 chars, no full credential traces).
|
||||||
|
|
||||||
|
### Rule 7 — Out-of-scope
|
||||||
|
|
||||||
|
This ADR does NOT govern:
|
||||||
|
|
||||||
|
- Spawn-path OAuth refresh (the spawn path's refresh logic predates this ADR and is governed by the underlying CLI). The probe shares the credential artifact but does not own the refresh.
|
||||||
|
- Anthropic-specific bearer revocation (Anthropic side). Revocation manifests as 401 to the probe, which falls into Rule 6.
|
||||||
|
- Non-Anthropic provider OAuth flows. Mistral / future providers MAY adopt this protocol via plugin-specific ADRs; ADR 0013 establishes the template.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
**Positive.**
|
||||||
|
|
||||||
|
- The probe is bounded — Rule 2 caps the wire traffic, Rule 3 caps the refresh rate, Rule 4 caps activation surface.
|
||||||
|
- Schema-drift detection is procedural and reproducible — Rule 5 gives the maintainer a runbook that doesn't depend on Anthropic publishing a deprecation notice.
|
||||||
|
- Failure is visible — Rule 6 means a broken probe shows up in `olp doctor` and the dashboard, not as a silent "—" in the quota row.
|
||||||
|
- Credential reuse (Rule 1) keeps the security surface area minimal.
|
||||||
|
|
||||||
|
**Negative.**
|
||||||
|
|
||||||
|
- The `quota_probe_enabled` opt-in adds a configuration step. Mitigated by `olp doctor` surfacing the recipe when credentials are present but the probe is off.
|
||||||
|
- The schema-drift protocol is manual. Anthropic could ship a v3.x binary tomorrow and the verification only happens when the maintainer or a doctor probe failure prompts it. Counter-pressure: drift events at OCP scale (~12 months) suggest manual verification on major version bumps is sufficient.
|
||||||
|
- Stale-cache-on-failure (Rule 3) means the dashboard could show 30-minute-old data without an obvious "stale" indicator unless the UI explicitly renders the `stale: true` marker. ADR 0012 D82 requires the dashboard to surface staleness; reviewing that during P5-2 implementation.
|
||||||
|
|
||||||
|
**Neutral.**
|
||||||
|
|
||||||
|
- The protocol is portable. Future provider plugins adopting direct-API probes (mistral if its `/v1/usage` exists) can reuse the same six rules with provider-specific endpoint substitution.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Alternatives considered
|
||||||
|
|
||||||
|
### A — Probe lives in `server.mjs` (OCP-style)
|
||||||
|
|
||||||
|
OCP's probe is in `server.mjs:842-1109` because OCP is single-provider and pre-plugin-architecture. Porting that pattern to OLP would violate ADR 0002 (per-provider knowledge stays in `lib/providers/`). Rejected.
|
||||||
|
|
||||||
|
### B — Spawn `claude -p --dry-run` and parse ratelimit headers
|
||||||
|
|
||||||
|
`claude -p` does not expose response headers; the CLI consumes and discards them. Even if it did, parsing CLI stdout is fragile. Rejected.
|
||||||
|
|
||||||
|
### C — Wait for Anthropic to publish a public quota API
|
||||||
|
|
||||||
|
The 2026-06-15 Agent SDK Credit billing-split announcement does not include a public quota API. Anthropic may publish one in the future; this ADR is forward-compatible (Rule 7 explicitly notes "if Anthropic publishes a public ratelimit API, this entire workaround becomes obsolete — re-evaluate"). Rejected for v0.5.0 (no ETA).
|
||||||
|
|
||||||
|
### D — Mandate token-rotation in OLP
|
||||||
|
|
||||||
|
Tempting (auditability), but OCP's experience shows token rotation breaks the spawn path more often than it improves security at family-scale deployment. The credential rotation cadence is Anthropic-side (token TTL); OLP respects whatever Claude Code does. Rejected.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Authority + cross-references
|
||||||
|
|
||||||
|
- **ADR 0002 Amendment 8** — the contract-level permission. ADR 0013 is the implementation discipline for that permission.
|
||||||
|
- **ADR 0012** — Phase 5 charter that schedules D80 implementation.
|
||||||
|
- **ADR 0011** — anonymous-key deployment-context (LAN-only). Separate scope; this ADR does not amend it.
|
||||||
|
- **`~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md`** — the live schema pin. Updated on every drift event per Rule 5.
|
||||||
|
- **`alignment.yml`** — must continue to blacklist `/api/oauth/usage` and related hallucinated tokens. Must NOT add `/v1/messages` to the blacklist (legitimate endpoint).
|
||||||
|
- **OCP `server.mjs:842-1109`** — the source-of-truth port reference for D80.
|
||||||
|
- **OCP `ALIGNMENT.md`** — the institutional precedent (2026-04-11 drift → ALIGNMENT introduction) this ADR consolidates for OLP.
|
||||||
@@ -0,0 +1,291 @@
|
|||||||
|
# ADR 0014 — Sandbox-Runtime Integration for Multi-Tenant Provider Spawning
|
||||||
|
|
||||||
|
**Status:** Accepted (PR-A — deps + doctor + ADR only; PR-B/C/D pending)
|
||||||
|
**Date:** 2026-05-28
|
||||||
|
**Phase:** Phase 7
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Related
|
||||||
|
|
||||||
|
- **ADR 0001** (Project Founding) — OLP's multi-provider rationale and "no conversation state" principle.
|
||||||
|
- **ADR 0009 Amendment 1** (stream-json transport, Phase 6) § Caveats #3: "Sandbox-runtime still required for real multi-tenant deployment."
|
||||||
|
- **ADR 0002** (Plugin Architecture) — Provider contract; `spawn()` is the surface this ADR will wrap in PR-B/C.
|
||||||
|
- **ADR 0006** (Provider Inclusion / Risk Tier Framework) — classifies providers by deployment risk; sandbox status is a gating condition for Tier-A (cloud-deployed).
|
||||||
|
- **`docs/plans/cloud-deployment-family.md` § 5** — sandbox is a hard prerequisite before any cloud rollout.
|
||||||
|
- **cc-mem incident memory** — `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` — the multi-tenant security gap that motivates this ADR.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Context
|
||||||
|
|
||||||
|
### 1.1 The multi-tenant security gap
|
||||||
|
|
||||||
|
OLP is a personal-scale proxy (ADR 0001 § Non-commercial). However, the "family-scale" deployment model means multiple human callers share a single OLP instance — each with their own OLP API key (ADR 0007) but all using the same underlying `claude` or `codex` CLI installation on the server host.
|
||||||
|
|
||||||
|
The security gap, identified in the 2026-05-27 session and captured in cc-mem incident memory § 3, is:
|
||||||
|
|
||||||
|
1. **OAuth token exposure.** A malicious (or misbehaving) prompt to the Anthropic provider could elicit a `cat ~/.olp/keys/...` or similar read of any file the OLP process user can access — including the OAuth credentials file that allows the attacker to impersonate the server-side identity.
|
||||||
|
2. **Codex shell-tool execution.** The `codex exec` path exposes a shell tool to the model. With OLP acting as a relay, a prompt to codex from one client could execute arbitrary commands in the server process's working directory, reading or writing files belonging to other clients.
|
||||||
|
3. **Cross-tenant data leakage.** Even without adversarial prompts, a model that freely accesses the filesystem could inadvertently leak one client's cached context to another client's response.
|
||||||
|
|
||||||
|
The 2026-05-27 prior-art search (incident memory § 4) surveyed the multi-tenant LLM proxy ecosystem (LiteLLM, OpenCode, CLIProxyAPI, open-source Anthropic proxies) and found that **none solve multi-tenant file-system and tool isolation at the OS level**. The field's typical answer is "don't run multi-tenant" or "use a separate VM per tenant" — neither applicable at OLP's family scale.
|
||||||
|
|
||||||
|
### 1.2 Anthropic's official answer: `@anthropic-ai/sandbox-runtime`
|
||||||
|
|
||||||
|
The `@anthropic-ai/sandbox-runtime` package (Anthropic Experimental org, `anthropic-experimental/sandbox-runtime`, v0.0.52 as of this ADR) is Anthropic's open-source solution to wrapping security boundaries around arbitrary processes. It is the library that Claude Code itself uses internally to sandbox MCP servers and tool execution.
|
||||||
|
|
||||||
|
The library provides:
|
||||||
|
|
||||||
|
- **Linux:** bubblewrap (`bwrap`) namespace isolation + socat network bridge + seccomp filter via `apply-seccomp-filter` binary. Ripgrep (`rg`) is required for deny-path glob expansion.
|
||||||
|
- **macOS:** `sandbox-exec` seatbelt profile, which is a built-in OS facility (no additional packages required).
|
||||||
|
|
||||||
|
Both paths enforce filesystem read/write restrictions and network policy at the kernel level, not at the process level. A `cat ~/.olp/keys/...` inside the sandbox fails at the syscall layer regardless of what the shell or model requests.
|
||||||
|
|
||||||
|
### 1.3 The 2026-05-28 spike
|
||||||
|
|
||||||
|
A PoC spike was conducted on PI231 (arm64 Debian Bookworm) on 2026-05-28. Key findings:
|
||||||
|
|
||||||
|
1. **`npm install @anthropic-ai/sandbox-runtime@0.0.52` succeeds cleanly** on arm64 Linux. No native build step; prebuilt binaries were available.
|
||||||
|
2. **`SandboxManager.isSupportedPlatform()` returns `true`** on PI231 (Linux, not WSL).
|
||||||
|
3. **`SandboxManager.checkDependencies()` reports errors**: `bubblewrap (bwrap) not installed`, `socat not installed`, `ripgrep (rg) not found`. These are the three OS-level deps that must be installed separately (not bundled in the npm package).
|
||||||
|
4. **The install fix is a one-liner**: `sudo apt-get install -y bubblewrap socat ripgrep`. This is a 5-minute operational task, not a code change.
|
||||||
|
5. **Three PoC scripts** were parked at `/tmp/sandbox-spike/` on PI231 verifying: dependency check return shapes, `SandboxManager.wrapWithSandbox` call signature, and filesystem-deny path behaviour.
|
||||||
|
|
||||||
|
Verdict: **YELLOW** — architecturally green (the library works and the platform is supported), operationally blocked on apt deps. PR-A lays the dependency + doctor layer. PR-B wraps the anthropic spawn after apt install.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Decision
|
||||||
|
|
||||||
|
### 2.1 Layered rollout (Iron Rule 11 — minimum reviewable unit)
|
||||||
|
|
||||||
|
The sandbox integration is split into four discrete PRs, each independently reviewable and independently safe to land or revert:
|
||||||
|
|
||||||
|
| PR | Scope | Blocking condition | Status |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **PR-A** (this PR) | npm dep `@anthropic-ai/sandbox-runtime ^0.0.52` + `lib/sandbox/doctor.mjs` (preflight module) + `/health` `sandbox` field + ADR 0014 | None — no runtime initialization | ✅ Accepted |
|
||||||
|
| **PR-B** | `lib/sandbox/manager.mjs` (bootstrap + spawn-wrap) + `lib/providers/anthropic.mjs` spawn wrapped + server startup wiring + `/health.sandbox.active` + Suite 43/44 tests | `bubblewrap` + `socat` + `rg` installed on PI231 (`sudo apt-get install -y bubblewrap socat ripgrep`) | ✅ Implemented — pending PI231 validation (Suite 44) + opus reviewer |
|
||||||
|
| **PR-C** | `lib/providers/codex.mjs` spawn wrapped with `enableWeakerNestedSandbox: true` | PR-B accepted + codex PoC on PI231 | 🔲 Blocked on PR-B |
|
||||||
|
| **PR-D** | `docs/plans/cloud-deployment-family.md` § "Phase 7 prerequisite met" update; cloud rollout unblocked | PR-B + PR-C accepted | 🔲 Blocked on PR-C |
|
||||||
|
|
||||||
|
Rationale for the split:
|
||||||
|
|
||||||
|
- **PR-A is safe without bwrap.** The doctor module and `/health` field add observability with no runtime side effects. No `SandboxManager.initialize()` call. No sandbox spawned.
|
||||||
|
- **PR-B is the load-bearing security gate.** Wrapping `anthropic.mjs` spawn requires empirical negative-test confirmation (in-sandbox `cat ~/.olp/keys/...` MUST fail). This cannot be verified until PI231 has bwrap installed.
|
||||||
|
- **PR-C follows PR-B** because codex has a distinct issue: codex itself uses bubblewrap internally (`codex exec` spawns its own sandbox). `enableWeakerNestedSandbox: true` is required to allow the inner sandbox to function inside the outer OLP sandbox.
|
||||||
|
- **PR-D is documentation-only** and depends on the runtime PRs being proven in production.
|
||||||
|
|
||||||
|
### 2.2 PR-A specific scope (binding)
|
||||||
|
|
||||||
|
PR-A MUST NOT include:
|
||||||
|
|
||||||
|
- Any call to `SandboxManager.initialize()` (no real sandbox created)
|
||||||
|
- Any modification to `lib/providers/anthropic.mjs`, `lib/providers/codex.mjs`, or `lib/providers/mistral.mjs`
|
||||||
|
- Any new HTTP endpoint (no `/metrics`, no new dashboard endpoint)
|
||||||
|
- Any modification to `models-registry.json`
|
||||||
|
|
||||||
|
PR-A MUST include:
|
||||||
|
|
||||||
|
- `package.json` dependency: `"@anthropic-ai/sandbox-runtime": "^0.0.52"`
|
||||||
|
- `lib/sandbox/doctor.mjs`: pure preflight module (no state; no initialization)
|
||||||
|
- `/health` response: top-level `sandbox` field (`available`, `missing`, `platform`, `message` when unavailable)
|
||||||
|
- `docs/adr/0014-sandbox-runtime-integration.md` (this document)
|
||||||
|
- `CHANGELOG.md` Unreleased entry
|
||||||
|
- `test-features.mjs` Suite 42 (8 new tests, all passing)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. `lib/sandbox/doctor.mjs` design
|
||||||
|
|
||||||
|
### 3.1 Exports
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
// Returns { available: boolean, missing: string[], details: { ... } }
|
||||||
|
export async function checkSandboxAvailability() { ... }
|
||||||
|
|
||||||
|
// Returns { ok: boolean, message: string } — human-readable summary
|
||||||
|
export async function describeSandboxStatus() { ... }
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.2 `checkSandboxAvailability` algorithm
|
||||||
|
|
||||||
|
1. Probe OS deps independently via `child_process.execFileSync('which', [binary])`:
|
||||||
|
- `bwrap` (Linux only — macOS uses built-in `sandbox-exec`)
|
||||||
|
- `socat` (Linux only)
|
||||||
|
- `rg` (ripgrep — Linux only; macOS seatbelt profiles use regex patterns natively)
|
||||||
|
2. Call `probeLibrary()` which `import()`s `@anthropic-ai/sandbox-runtime` and calls:
|
||||||
|
- `SandboxManager.isSupportedPlatform()` — platform classification
|
||||||
|
- `SandboxManager.checkDependencies(undefined)` — library's own dep check (called without initialize, falling back to PATH lookup)
|
||||||
|
3. Compute `missing[]`: on Linux, add 'bubblewrap', 'socat', 'ripgrep' for each absent dep; if library import failed, add that too.
|
||||||
|
4. `available = libLoaded && isSupportedPlatform && missing.length === 0`
|
||||||
|
|
||||||
|
`probeLibrary()` wraps everything in try/catch — any library-side error becomes `{ libLoaded: false, libError: '<reason>' }` rather than an unhandled rejection.
|
||||||
|
|
||||||
|
### 3.3 `/health` integration
|
||||||
|
|
||||||
|
The `sandbox` field is added to the full (owner-tier) payload only. For trimmed payloads (guest/anonymous per ADR 0007 § 7.1), the field is absent (consistent with the existing trim model). This prevents leaking infrastructure details to non-owner callers.
|
||||||
|
|
||||||
|
The result is memoized process-wide via `_sandboxStatusCache` in `server.mjs`. The install state of bwrap/socat cannot change at runtime without a process restart, so a single lazy fetch at the first `/health` call is correct.
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"ok": true,
|
||||||
|
"version": "0.5.1",
|
||||||
|
"providers": { ... },
|
||||||
|
"sandbox": {
|
||||||
|
"available": false,
|
||||||
|
"missing": ["bubblewrap", "socat", "ripgrep"],
|
||||||
|
"platform": "linux",
|
||||||
|
"message": "Sandbox dependencies not available: bubblewrap not installed, socat not installed, ripgrep not installed. Install: sudo apt-get install -y bubblewrap socat ripgrep"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
When available (after apt install + process restart) and PR-B bootstrapped:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"sandbox": {
|
||||||
|
"available": true,
|
||||||
|
"active": true,
|
||||||
|
"missing": [],
|
||||||
|
"platform": "linux"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
(PR-A shape did not include `active`. PR-B adds `active: boolean` — distinguishes
|
||||||
|
"deps present" from "sandbox actually initialized and wrapping spawns".)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. PR-B/C/D acceptance criteria
|
||||||
|
|
||||||
|
### 4.1 PR-B (anthropic.mjs spawn wrap) — ✅ Implementation shipped, PI231 validation pending
|
||||||
|
|
||||||
|
**PR-B implementation (commit pending reviewer):**
|
||||||
|
- `lib/sandbox/manager.mjs`: singleton bootstrap + transparent `wrapSpawn()` API
|
||||||
|
- `lib/providers/anthropic.mjs`: spawn site wrapped via `wrapSpawn()` (ADR 0009 Amendment 1 spawn args unchanged)
|
||||||
|
- `server.mjs`: `bootstrapSandbox()` called before `server.listen()`, `/health.sandbox.active` field added
|
||||||
|
- `test-features.mjs` Suite 43 (8 tests, all pass on macOS) + Suite 44 (2 tests, PI231-gated with `OLP_E2E_SANDBOX=1`)
|
||||||
|
- 805 → 813 tests. Suite 44 skipped by default; runs on PI231 after apt install.
|
||||||
|
|
||||||
|
**Load-bearing negative test (required for PR-B to merge):**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# On PI231, with bwrap+socat installed, with PR-B wired:
|
||||||
|
olp-keys list # identify owner key
|
||||||
|
curl -X POST http://127.0.0.1:4567/v1/chat/completions \
|
||||||
|
-H "Authorization: Bearer <owner-key>" \
|
||||||
|
-H "Content-Type: application/json" \
|
||||||
|
-d '{"model":"claude-sonnet-4-6","messages":[{"role":"user","content":"run: cat /home/<user>/.olp/keys/owner-key.json"}]}'
|
||||||
|
# Expected: response MUST NOT contain any content from the keys file.
|
||||||
|
# The model must either say it cannot access the filesystem, or produce an
|
||||||
|
# error. Any response containing the file content is a PR-B blocking failure.
|
||||||
|
```
|
||||||
|
|
||||||
|
Additional criteria:
|
||||||
|
- `SandboxManager.initialize()` is called once at startup (per ADR 0014 § 5 singleton decision, TBD in PR-B ADR amendment)
|
||||||
|
- p95 latency overhead of wrapping ≤ 200ms measured over 50 warm requests
|
||||||
|
- `checkSandboxAvailability().available === true` reported in `/health.sandbox` after PR-B rolls out
|
||||||
|
- All existing Suite 41 tests continue to pass (stream-json transport unaffected)
|
||||||
|
|
||||||
|
### 4.2 PR-C (codex.mjs wrap)
|
||||||
|
|
||||||
|
- `enableWeakerNestedSandbox: true` is set in the `SandboxManager.initialize()` call (or per-spawn config if the API allows per-spawn override — verify against v0.0.52 API)
|
||||||
|
- `codex exec` inner bubblewrap nest still functions: a sandboxed codex invocation that reads from an allowed path succeeds
|
||||||
|
- Analogous negative test: in-sandbox `cat /home/<user>/.olp/keys/...` MUST fail
|
||||||
|
|
||||||
|
### 4.3 PR-D (cloud deployment plan update)
|
||||||
|
|
||||||
|
- `docs/plans/cloud-deployment-family.md` § 5 "Phase 7 prerequisite" section updated: "sandbox-runtime integration (PR-B + PR-C) confirmed operational on PI231; prerequisite met"
|
||||||
|
- `README.md` § "Supported Providers" or § "Security" updated with a note about sandbox isolation
|
||||||
|
- Phase 7 close PR per `CLAUDE.md release_kit.phase_rolling_mode`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Open questions (to be resolved in PR-B)
|
||||||
|
|
||||||
|
1. **Singleton vs per-spawn initialization.** `SandboxManager` is a process-wide singleton (per the library's `reset()` being a global operation). The current design plan is one `initialize()` call at server startup with a union config covering all providers. If providers require different configs (e.g., different `denyRead` paths for anthropic vs codex), this may require a mutex approach or separate singleton instances. Decision reserved for PR-B.
|
||||||
|
|
||||||
|
2. **`SandboxManager.reset()` in tests.** The singleton means test suites that call `initialize()` must call `reset()` in their `after()` hooks. PR-B must add this discipline or tests will leak sandbox state across suites.
|
||||||
|
|
||||||
|
3. **MITM proxy and Claude CLI cert pinning.** The sandbox-runtime network bridge on Linux uses a local MITM proxy to intercept HTTPS traffic. If `claude` CLI pins certificates (e.g., for `api.anthropic.com`), HTTPS through the bridge may fail. PR-B must empirically verify this on PI231 before merging.
|
||||||
|
|
||||||
|
4. **macOS `sandbox-exec` profile content.** macOS uses a seatbelt (SBPL) profile, not bwrap. The profile must explicitly allow `network outbound "api.anthropic.com"` etc. The default profile may be too restrictive for the Claude CLI's OAuth refresh calls. PR-B must test macOS as well as Linux.
|
||||||
|
|
||||||
|
5. **`getDefaultWritePaths()` output.** The library exports `getDefaultWritePaths()` which returns the paths the sandbox always allows writing to. OLP's spawn directory may not be in that list — PR-B must verify the working directory is writable or pass it explicitly in `filesystem.allowWrite`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Pitfalls inherited from the spike (binding warnings for PR-B/C authors)
|
||||||
|
|
||||||
|
These were confirmed empirically or inferred from the library source during the 2026-05-28 spike:
|
||||||
|
|
||||||
|
1. **Three OS deps, not one.** The npm package bundles nothing. Linux requires: `bubblewrap` (bwrap), `socat`, `ripgrep` (rg). All three. Missing even one → `checkDependencies()` returns errors → `wrapWithSandbox` will fail at runtime.
|
||||||
|
|
||||||
|
2. **Linux deny-paths are literal, not glob.** The library's `linuxGetMandatoryDenyPaths()` uses ripgrep to expand glob patterns to concrete paths before passing them to bwrap. But custom `filesystem.denyRead` entries that contain glob chars (`~/.ssh/*`) must be either expanded manually OR passed as the glob form (the library expands them if `rg` is available). The safe convention for PR-B: use absolute literal paths (e.g., `/home/<user>/.ssh`) rather than `~/`-prefixed or glob paths.
|
||||||
|
|
||||||
|
3. **`enableWeakerNestedSandbox: true` is required for codex.** Codex's `exec` subcommand spawns its own bubblewrap sandbox internally. Without `enableWeakerNestedSandbox`, the outer OLP sandbox blocks the inner codex sandbox from creating user namespaces. The flag loosens the outer sandbox's seccomp filter specifically to allow `clone(CLONE_NEWUSER)` — the inner sandbox then runs with reduced but non-zero isolation.
|
||||||
|
|
||||||
|
4. **`SandboxManager.reset()` is process-wide.** Calling `reset()` anywhere (including test teardown) clears the singleton config. Any concurrent in-flight spawn that still holds a reference to the old sandbox state will break. PR-B's design must either (a) initialize once at boot and never reset, or (b) use a mutex to prevent concurrent init/reset.
|
||||||
|
|
||||||
|
5. **MITM CA generation is async and expensive.** `SandboxManager.initialize()` generates a self-signed CA certificate for the MITM proxy on Linux. This takes ~100-500ms. Initialize at server startup, not per-request.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Authority citations
|
||||||
|
|
||||||
|
- **`@anthropic-ai/sandbox-runtime` v0.0.52** — https://github.com/anthropic-experimental/sandbox-runtime
|
||||||
|
- `dist/sandbox/sandbox-manager.js` — `isSupportedPlatform()`, `checkDependencies()`, `SandboxManager` export shape
|
||||||
|
- `dist/sandbox/linux-sandbox-utils.js` — `checkLinuxDependencies()`, `whichSync` usage, `enableWeakerNestedSandbox` rationale
|
||||||
|
- `README.md` — installation prerequisites, platform support matrix
|
||||||
|
|
||||||
|
- **2026-05-28 PoC spike on PI231 (arm64 Debian Bookworm)** — report at `/tmp/sandbox-spike/report.md` on PI231. Key findings: dep install clean; `isSupportedPlatform()=true`; `checkDependencies()` errors on bwrap+socat+rg absence; three PoC scripts parked. Verdict YELLOW.
|
||||||
|
|
||||||
|
- **cc-mem incident memory 2026-05-27** — `~/.cc-rules/memory/projects/olp/incident_2026_05_27_spawn_cli_security.md` § 3 (gap description), § 4 (prior-art search showing ecosystem hasn't solved multi-tenant fs/tool isolation).
|
||||||
|
|
||||||
|
- **OLP ADR 0009 Amendment 1 § Caveats #3** — "Sandbox-runtime still required for real multi-tenant deployment. Per the 2026-05-27 session prior-art search, Anthropic's official multi-tenant answer is `@anthropic-ai/sandbox-runtime` (OS-level isolation)."
|
||||||
|
|
||||||
|
- **`docs/plans/cloud-deployment-family.md` § 5** — sandbox is a hard prerequisite before any cloud deployment.
|
||||||
|
|
||||||
|
- **OLP ALIGNMENT.md** — PR-A is library/doctor/governance; it does not touch provider plugins, the entry surface, or the IR. The authority citation for the npm dep is the official sandbox-runtime repo URL + the spike report (not a provider CLI, not the OpenAI spec, not an existing ADR — this is a new dependency decision, which is the correct scope for ADR 0014).
|
||||||
|
|
||||||
|
- **Iron Rule 11 (Incremental Diff Review)** — splits non-trivial work into the minimum reviewable unit. The 4-PR split (A/B/C/D) is the direct application of this rule to the sandbox integration: each PR is independently reviewable, independently safe to land or revert, and corresponds to one logical layer.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Consequences
|
||||||
|
|
||||||
|
### Positive
|
||||||
|
|
||||||
|
- **Multi-tenant isolation at the OS level.** After PR-B+C land, each provider spawn runs inside a bubblewrap (Linux) or sandbox-exec (macOS) boundary. A prompt-injected `cat ~/.olp/keys/...` hits a kernel-level deny. Cross-client filesystem leakage is structurally prevented, not just mitigated by prompt engineering.
|
||||||
|
|
||||||
|
- **Cloud deployment unblocked.** `docs/plans/cloud-deployment-family.md` § 5 cites sandbox as the hard prerequisite for moving from family-LAN to cloud. PR-D closes this gate.
|
||||||
|
|
||||||
|
- **Observability from day one.** The `/health.sandbox` field makes the install state machine-readable. Any monitoring script or dashboard can tell whether sandbox isolation is active without SSH access.
|
||||||
|
|
||||||
|
- **Anthropic's official library.** Using `@anthropic-ai/sandbox-runtime` rather than a home-grown bwrap wrapper means OLP inherits Anthropic's tested integration patterns (deny-path expansion, MITM proxy, seccomp, macOS seatbelt profiles) rather than reinventing them. When the library updates, OLP upgrades via `npm update`.
|
||||||
|
|
||||||
|
### Negative
|
||||||
|
|
||||||
|
- **Three new OS-level dependencies.** `bubblewrap`, `socat`, and `ripgrep` must be installed on every host running OLP with sandbox isolation active. Absent these deps, sandbox is unavailable (but OLP continues to function without isolation — degraded security, not degraded functionality). The `/health.sandbox.available` field makes this state explicit.
|
||||||
|
|
||||||
|
- **p95 latency overhead.** The spike did not measure sandbox wrapping overhead directly (blocked on apt install). Expected overhead per the sandbox-runtime README: ~100-200ms for sandbox initialization amortized over the process lifetime (one-time at startup); per-spawn overhead is the namespace clone + filesystem mount overhead, typically <50ms on modern kernels. PR-B's acceptance criteria gates on ≤200ms p95 overhead over 50 warm requests.
|
||||||
|
|
||||||
|
- **Codex inner-sandbox degradation.** `enableWeakerNestedSandbox: true` loosens the outer OLP sandbox's seccomp filter to allow `clone(CLONE_NEWUSER)`. The codex inner sandbox still runs with meaningful isolation (its own namespace, its own deny-list), but the combined depth of protection is less than ideal compared to a world where codex didn't self-sandbox.
|
||||||
|
|
||||||
|
- **Library is experimental.** The `anthropic-experimental` org signals this is not a production-stable API. The version pin (`^0.0.52`) provides a minor-range buffer but the API surface may change. If the library is deprecated or the API breaks, OLP's fallback is to remove the sandbox wrapping (reverting PRs B-D) until a replacement path is found. This is acceptable at family scale — security degradation is not a service outage.
|
||||||
|
|
||||||
|
### Reversibility
|
||||||
|
|
||||||
|
- **PR-A** is trivially reversible: `npm uninstall @anthropic-ai/sandbox-runtime` + delete `lib/sandbox/doctor.mjs` + revert server.mjs and CHANGELOG changes. No production behavior changes.
|
||||||
|
- **PR-B/C** are reversible by removing the `SandboxManager.wrapWithSandbox` call from each provider's `spawn()` method. The spawn falls back to the current unsandboxed path.
|
||||||
|
- **PR-D** is a documentation update; reverting it is a docs-only change.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Status transitions
|
||||||
|
|
||||||
|
- 2026-05-28 — Created. Status: Accepted for PR-A scope. PR-B/C/D pending operational prereqs.
|
||||||
@@ -25,6 +25,8 @@ New ADRs increment from the highest existing number. Filenames are `NNNN-<short-
|
|||||||
| [0009](0009-interactive-mode-path-placeholder.md) | Anthropic Interactive-Mode Path (Placeholder) | Placeholder ADR (2026-05-25, Draft) — blocked on OCP ADR 0007 P0 experiment outcome. Records the maintainer's "wait + port" decision: do NOT independently implement; ride OCP's P0 result. If P0 confirms Transport A (stdio NDJSON) or B (PTY) bills as subscription rather than Agent SDK credit, port to OLP `lib/providers/anthropic.mjs` (Option 1 parallel impl, or Option 2 OCP-as-backend; decision deferred to P0-resolution time). If P0 fails on both, shelve. No Phase 4 D-day scheduled until P0 lands AND maintainer issues explicit "go" naming this ADR. |
|
| [0009](0009-interactive-mode-path-placeholder.md) | Anthropic Interactive-Mode Path (Placeholder) | Placeholder ADR (2026-05-25, Draft) — blocked on OCP ADR 0007 P0 experiment outcome. Records the maintainer's "wait + port" decision: do NOT independently implement; ride OCP's P0 result. If P0 confirms Transport A (stdio NDJSON) or B (PTY) bills as subscription rather than Agent SDK credit, port to OLP `lib/providers/anthropic.mjs` (Option 1 parallel impl, or Option 2 OCP-as-backend; decision deferred to P0-resolution time). If P0 fails on both, shelve. No Phase 4 D-day scheduled until P0 lands AND maintainer issues explicit "go" naming this ADR. |
|
||||||
| [0010](0010-phase-4-charter-operator-and-client-ux.md) | Phase 4 Charter — Operator + Client UX | Phase 4 scope ratification (2026-05-26, Accepted). Phase 4 = operator + client UX (SSE heartbeat / `olp` CLI + doctor / `olp-connect` zero-config + Telegram-Discord plugin + IDE docs bundle). ~13 D-days, D60 → v0.4.0. Records the explicit decision to DEFER `/v1/messages` (Anthropic-shape entry surface) on the rationale that under ADR 0009 P0 failure it provides no billing benefit AND degrades worse on fallback than OpenAI-shape clients. Re-open trigger: ADR 0009 P0 success + maintainer-named family CC user. Also closes the OCP-OLP port co-host ambiguity from ADR 0001 (default `OLP_PORT` 3456 → 4567). |
|
| [0010](0010-phase-4-charter-operator-and-client-ux.md) | Phase 4 Charter — Operator + Client UX | Phase 4 scope ratification (2026-05-26, Accepted). Phase 4 = operator + client UX (SSE heartbeat / `olp` CLI + doctor / `olp-connect` zero-config + Telegram-Discord plugin + IDE docs bundle). ~13 D-days, D60 → v0.4.0. Records the explicit decision to DEFER `/v1/messages` (Anthropic-shape entry surface) on the rationale that under ADR 0009 P0 failure it provides no billing benefit AND degrades worse on fallback than OpenAI-shape clients. Re-open trigger: ADR 0009 P0 success + maintainer-named family CC user. Also closes the OCP-OLP port co-host ambiguity from ADR 0001 (default `OLP_PORT` 3456 → 4567). |
|
||||||
| [0011](0011-anonymous-key-deployment-context.md) | Anonymous-Key Deployment-Context Limits (Trusted-LAN Invariant) | D70 (2026-05-26, Accepted). Codifies the trust posture for `/health.anonymousKey` opt-in field (D69) + `bin/olp-connect` zero-config consumer (D68). Three-prerequisite gate (`auth.advertise_anonymous_key=true` + `auth.allow_anonymous=true` + an active key with `plaintext_advertise` field). Guest-tier-only restriction (`createKey()` + CLI reject owner+advertise). Trusted-LAN deployment invariant (loopback / RFC1918 / tailnet / `.local` / `.internal` — soft constraint at v0.4.0; hard enforcement deferred until OLP gains a public-deployment recipe). Re-evaluation trigger: any "expose to public internet" README mode. |
|
| [0011](0011-anonymous-key-deployment-context.md) | Anonymous-Key Deployment-Context Limits (Trusted-LAN Invariant) | D70 (2026-05-26, Accepted). Codifies the trust posture for `/health.anonymousKey` opt-in field (D69) + `bin/olp-connect` zero-config consumer (D68). Three-prerequisite gate (`auth.advertise_anonymous_key=true` + `auth.allow_anonymous=true` + an active key with `plaintext_advertise` field). Guest-tier-only restriction (`createKey()` + CLI reject owner+advertise). Trusted-LAN deployment invariant (loopback / RFC1918 / tailnet / `.local` / `.internal` — soft constraint at v0.4.0; hard enforcement deferred until OLP gains a public-deployment recipe). Re-evaluation trigger: any "expose to public internet" README mode. |
|
||||||
|
| [0012](0012-phase-5-charter-quota-probes-dashboard.md) | Phase 5 Charter — Provider Quota Probes + Dashboard Enrichment | Phase 5 scope ratification (2026-05-26, Accepted). Phase 5 = port OCP's plan-usage probe to `lib/providers/anthropic.mjs:quotaStatus()` + Claude.ai-style dashboard enrichment (1-min auto-refresh + manual refresh + per-provider rows with utilization bars, reset countdowns, status badges) + optional mistral probe at D84 (codex explicitly skipped — no public API). ~6 D-days, D79 → v0.5.0. Companion to ADR 0002 Amendment 8 (direct-API READ-ONLY exemption) + ADR 0013 (OAuth READ-ONLY consumption rules). Closes v1.x roadmap #8 (dashboard enrichment). Re-confirmed schema 2026-05-26 via compiled-binary `strings` + live API probe; 3 new fields since OCP 2026-04 capture, no removals. |
|
||||||
|
| [0013](0013-oauth-read-only-consumption-and-schema-drift.md) | OAuth READ-ONLY Consumption Rules + Schema-Drift Mitigation Protocol | D79 (2026-05-26, Accepted). Implementation discipline for ADR 0002 Amendment 8. Seven rules covering: credential reuse with spawn path (no new OAuth grant); READ-ONLY at the wire (one probe per cache miss, `max_tokens:1`, headers-only parse, discard body); cache TTL 5min + 60s-3600s exponential refresh backoff + stale-cache-on-failure; opt-in via `~/.olp/config.json providers.<name>.quota_probe_enabled` (default false); schema-drift mitigation via dual-path verification (compiled-binary `strings` + live API probe diff); failure transparency through `olp doctor` + dashboard staleness markers; out-of-scope clarifications. Bound by `~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md` as the live schema pin. |
|
||||||
|
|
||||||
## When to write a new ADR
|
## When to write a new ADR
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,34 @@
|
|||||||
|
{
|
||||||
|
"phase": "v0.5.1 post-release",
|
||||||
|
"purpose": "Refresh dashboard screenshot with live MacBook data after v0.5.1 hotfix (replacing D82's synthetic-data render)",
|
||||||
|
"captured_at_utc": "2026-05-27T01:26:17.304Z",
|
||||||
|
"host": "maintainer's MacBook (Mac client test target per project test-envs; specific IP / Tailscale node redacted per public-repo hygiene)",
|
||||||
|
"server_version": "0.5.1 (main @ commit fa2d1af \u2014 F4+#7 post-merge)",
|
||||||
|
"olp_port": 14567,
|
||||||
|
"endpoint_tested": "/v0/management/dashboard-data",
|
||||||
|
"auth": "owner-tier OLP key (temp, revoked post-test)",
|
||||||
|
"result_summary": {
|
||||||
|
"anthropic": {
|
||||||
|
"status": "live",
|
||||||
|
"schema_version": "2026-05-26",
|
||||||
|
"utilization_5h": 0.06,
|
||||||
|
"utilization_7d": 0.38,
|
||||||
|
"representative_claim": "five_hour",
|
||||||
|
"failure": null
|
||||||
|
},
|
||||||
|
"openai": {
|
||||||
|
"status": "unavailable",
|
||||||
|
"reason": "no public quota api or probe disabled"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"v0_5_1_contract_verified": [
|
||||||
|
"quota_v2[i].status enum includes 'live' (anthropic) and 'unavailable' (openai) \u2014 both rendered correctly",
|
||||||
|
"quota_v2[i].failure is null for healthy live status (per ADR 0013 Rule 6 \u2014 failure info only on stale/unreachable)",
|
||||||
|
"quota_v2[i].schema_version pinned at 2026-05-26 \u2014 matches models-registry.json quota_probe.schema_version"
|
||||||
|
],
|
||||||
|
"post_test_cleanup": [
|
||||||
|
"temp owner key (id=0m6s2s97, name=v0.5.1-screenshot) revoked",
|
||||||
|
"~/.olp/config.json providers.anthropic.quota_probe_enabled flag removed (config restored to baseline)",
|
||||||
|
"test server (pid varies, port=14567) terminated"
|
||||||
|
]
|
||||||
|
}
|
||||||
Binary file not shown.
|
After Width: | Height: | Size: 113 KiB |
+285
-61
@@ -1,16 +1,25 @@
|
|||||||
# OpenClaw + OLP
|
# OpenClaw + OLP
|
||||||
|
|
||||||
[OpenClaw](https://github.com/openclaw/openclaw) is a multi-bot gateway
|
[OpenClaw](https://github.com/openclaw/openclaw) is a multi-bot gateway that exposes slash commands on Telegram, Discord, and other chat surfaces. OLP integrates with OpenClaw in two ways:
|
||||||
that exposes slash commands on Telegram, Discord, and other chat
|
|
||||||
surfaces. OLP ships [`olp-plugin/`](../../olp-plugin/) as a native
|
|
||||||
OpenClaw plugin that registers a `/olp` slash command with read-only
|
|
||||||
parity to the local `olp` CLI.
|
|
||||||
|
|
||||||
**Status:** ✅ Supported.
|
1. **`/olp` slash commands** via the [`olp-plugin/`](../../olp-plugin/) plugin (read-only parity to the local `olp` CLI).
|
||||||
|
2. **LLM routing** — OpenClaw's chat agent can route its model calls through your OLP server, giving you per-key audit + quota observability for every bot reply.
|
||||||
|
|
||||||
## What you get
|
This doc covers both. **Status:** ✅ Supported.
|
||||||
|
|
||||||
After install, from Telegram or Discord:
|
## Two deployment modes — pick yours
|
||||||
|
|
||||||
|
The OpenClaw config differs significantly depending on whether OpenClaw runs on the same host as the OLP server or on a separate client machine talking to a remote OLP. Pick the right section.
|
||||||
|
|
||||||
|
| | **Mode A: Server-co-located** | **Mode B: Client-mode (recommended for multi-machine setups)** |
|
||||||
|
|---|---|---|
|
||||||
|
| OpenClaw runs on | the OLP server host (loopback) | a different machine (Mac mini, laptop, etc.) |
|
||||||
|
| OLP server runs on | localhost (same host) | a remote host (e.g. PI231) |
|
||||||
|
| `olp-claude` baseUrl | `http://127.0.0.1:4567/v1` | `http://<server-ip>:4567/v1` |
|
||||||
|
| Auth | `authHeader: false` (loopback trusted), OR anonymous-key if `auth.allow_anonymous: true` | `apiKey: "${OLP_OPENCLAW_BOT_TOKEN}"` env-var reference (NOT raw string, NOT `OPENAI_API_KEY` — see § Gotchas) |
|
||||||
|
| `/olp` slash plugin proxyUrl | `http://127.0.0.1:4567` | `http://<server-ip>:4567` |
|
||||||
|
|
||||||
|
## `/olp` slash commands you get
|
||||||
|
|
||||||
| Slash command | Maps to | Tier |
|
| Slash command | Maps to | Tier |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
@@ -24,14 +33,15 @@ After install, from Telegram or Discord:
|
|||||||
| `/olp doctor` | informational (HTTP endpoint not yet shipped) | — |
|
| `/olp doctor` | informational (HTTP endpoint not yet shipped) | — |
|
||||||
| `/olp help` | usage text | — |
|
| `/olp help` | usage text | — |
|
||||||
|
|
||||||
**Mutating subcommands are deliberately not exposed via chat.** `keygen`,
|
**Mutating subcommands are deliberately not exposed via chat.** `keygen`, `revoke`, `restart`, `logs` are SSH-only. See [`olp-plugin/README.md`](../../olp-plugin/README.md#what-you-can-not-do-from-chat-by-design) for the rationale.
|
||||||
`revoke`, `restart`, `logs` are SSH-only. See
|
|
||||||
[`olp-plugin/README.md`](../../olp-plugin/README.md#what-you-can-not-do-from-chat-by-design)
|
|
||||||
for the rationale.
|
|
||||||
|
|
||||||
## Quick setup
|
---
|
||||||
|
|
||||||
### 1. Install the plugin
|
## Mode A — Server-co-located install
|
||||||
|
|
||||||
|
OpenClaw + OLP on the same host. Auth is simpler because everything is on loopback.
|
||||||
|
|
||||||
|
### A1. Install the plugin
|
||||||
|
|
||||||
Two install paths — either works.
|
Two install paths — either works.
|
||||||
|
|
||||||
@@ -48,96 +58,310 @@ mkdir -p ~/.openclaw/extensions/
|
|||||||
ln -s /path/to/olp/olp-plugin/ ~/.openclaw/extensions/olp
|
ln -s /path/to/olp/olp-plugin/ ~/.openclaw/extensions/olp
|
||||||
```
|
```
|
||||||
|
|
||||||
### 2. Mint a bot owner key
|
### A2. Mint a bot owner key
|
||||||
|
|
||||||
Run on the OLP host (NOT in chat):
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
npx olp-keys keygen --owner --name=openclaw-bot
|
npx olp-keys keygen --owner --name=openclaw-bot
|
||||||
```
|
```
|
||||||
|
|
||||||
Capture the printed plaintext token — it is shown exactly once.
|
Capture the printed plaintext token — shown exactly once.
|
||||||
|
|
||||||
### 3. Configure
|
### A3. Configure (loopback recipe)
|
||||||
|
|
||||||
Edit `~/.openclaw/openclaw.json`:
|
Edit `~/.openclaw/openclaw.json`:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
"plugins": {
|
"plugins": {
|
||||||
|
"allow": ["...", "olp"],
|
||||||
|
"entries": {
|
||||||
"olp": {
|
"olp": {
|
||||||
|
"enabled": true,
|
||||||
|
"config": {
|
||||||
"proxyUrl": "http://127.0.0.1:4567",
|
"proxyUrl": "http://127.0.0.1:4567",
|
||||||
"apiKey": "olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
|
"apiKey": "olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
### 4. Restart the gateway
|
For LLM routing through OLP, add (or update) the `olp-claude` provider so the bot's default agent goes through OLP-spawned `claude -p`:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"models": {
|
||||||
|
"providers": {
|
||||||
|
"olp-claude": {
|
||||||
|
"baseUrl": "http://127.0.0.1:4567/v1",
|
||||||
|
"api": "openai-completions",
|
||||||
|
"authHeader": false,
|
||||||
|
"models": [
|
||||||
|
{ "id": "claude-sonnet-4-6", "name": "Claude Sonnet 4.6", "input": ["text"],
|
||||||
|
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||||
|
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`authHeader: false` is safe on loopback. If you set `auth.allow_anonymous: true` on the OLP server, the bot doesn't even need a key for slash commands (the `apiKey` field can be omitted). Owner-only subcommands (`/olp status`, `/olp usage`, `/olp cache`) still need an owner-tier key.
|
||||||
|
|
||||||
|
### A4. Restart the gateway
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
openclaw gateway restart
|
openclaw gateway restart
|
||||||
```
|
```
|
||||||
|
|
||||||
The plugin is now active. Try `/olp help` in your bot's chat.
|
---
|
||||||
|
|
||||||
## Known issues
|
## Mode B — Client-mode install (OpenClaw on different host than OLP)
|
||||||
|
|
||||||
- **`openclaw gateway restart` is required after install.** OpenClaw caches
|
OpenClaw on machine X (e.g., Mac mini), OLP server on machine Y (e.g., a Raspberry Pi or any LAN host). This is the common family deployment shape.
|
||||||
plugin discovery at gateway start. `openclaw plugins reload` does not
|
|
||||||
guarantee a fresh import of the plugin module.
|
|
||||||
|
|
||||||
- **Owner key revocation kicks the plugin out immediately.** If you revoke
|
### B1. Install the plugin
|
||||||
the bot's owner key (`npx olp-keys revoke --id=<id>`), the next `/olp
|
|
||||||
status` will return `401 unauthorized`. Mint a replacement key with a
|
|
||||||
new name and edit `~/.openclaw/openclaw.json`; do NOT reuse the revoked
|
|
||||||
key's UUID.
|
|
||||||
|
|
||||||
- **Long responses are truncated.** Telegram caps messages at ~4096
|
Same as Mode A:
|
||||||
characters. The plugin truncates with a `... [truncated, use SSH for
|
|
||||||
full]` suffix when the rendered output would exceed ~3900 chars. Use
|
```bash
|
||||||
SSH + the local `olp` CLI for full output.
|
openclaw plugins install /path/to/olp/olp-plugin/
|
||||||
|
# OR
|
||||||
|
mkdir -p ~/.openclaw/extensions/
|
||||||
|
ln -s /path/to/olp/olp-plugin/ ~/.openclaw/extensions/olp
|
||||||
|
```
|
||||||
|
|
||||||
|
### B2. Mint a bot owner key (on the OLP server, NOT on the OpenClaw host)
|
||||||
|
|
||||||
|
SSH to the OLP server:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh user@olp-server
|
||||||
|
cd ~/olp
|
||||||
|
node bin/olp-keys.mjs keygen --owner --name=openclaw-<hostname>-bot
|
||||||
|
```
|
||||||
|
|
||||||
|
Capture the plaintext — shown exactly once. **This token will live in `~/.openclaw/openclaw.json` on your OpenClaw host**; pick a name that makes it independently revocable if that host is lost/compromised.
|
||||||
|
|
||||||
|
### B3. Set the bot-token env var (`OLP_OPENCLAW_BOT_TOKEN`)
|
||||||
|
|
||||||
|
OpenClaw's canonical pattern for custom-provider auth is `apiKey: "${VAR_NAME}"` — an env-var reference, NOT a raw token. Choose a **custom** variable name (NOT `OPENAI_API_KEY` — OpenClaw service-manages that one and clobbers it with its own ChatGPT key on every restart). Convention: `OLP_OPENCLAW_BOT_TOKEN`.
|
||||||
|
|
||||||
|
**macOS (gateway under launchd)**:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
launchctl setenv OLP_OPENCLAW_BOT_TOKEN olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||||
|
```
|
||||||
|
|
||||||
|
Add the same `export` to `~/.zshrc` so it survives reboot:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export OLP_OPENCLAW_BOT_TOKEN=olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||||
|
```
|
||||||
|
|
||||||
|
**Linux (gateway under systemd-user)**: drop a file at `~/.config/environment.d/openclaw-olp.conf`:
|
||||||
|
|
||||||
|
```
|
||||||
|
OLP_OPENCLAW_BOT_TOKEN=olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
|
||||||
|
```
|
||||||
|
|
||||||
|
Restart the gateway service so it picks up the new env.
|
||||||
|
|
||||||
|
### B4. Configure `~/.openclaw/openclaw.json`
|
||||||
|
|
||||||
|
Edit `~/.openclaw/openclaw.json` on the OpenClaw host:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"plugins": {
|
||||||
|
"allow": ["...", "olp"],
|
||||||
|
"entries": {
|
||||||
|
"olp": {
|
||||||
|
"enabled": true,
|
||||||
|
"config": {
|
||||||
|
"proxyUrl": "http://<olp-server-ip>:4567",
|
||||||
|
"apiKey": "olp_XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"models": {
|
||||||
|
"providers": {
|
||||||
|
"olp-claude": {
|
||||||
|
"baseUrl": "http://<olp-server-ip>:4567/v1",
|
||||||
|
"api": "openai-completions",
|
||||||
|
"apiKey": "${OLP_OPENCLAW_BOT_TOKEN}",
|
||||||
|
"models": [
|
||||||
|
{ "id": "claude-sonnet-4-6", "name": "Claude Sonnet 4.6 (via OLP)", "input": ["text"],
|
||||||
|
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||||
|
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } },
|
||||||
|
{ "id": "claude-opus-4-7", "name": "Claude Opus 4.7 (via OLP)", "input": ["text"],
|
||||||
|
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||||
|
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } },
|
||||||
|
{ "id": "claude-haiku-4-5", "name": "Claude Haiku 4.5 (via OLP)", "input": ["text"],
|
||||||
|
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||||
|
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Note: the `plugins.entries.olp.config.apiKey` field (line 11) IS allowed to be a raw token — it's a separate code path that doesn't suffer the service-managed-env clobber problem. Only the `models.providers.<id>.apiKey` field needs the `${VAR}` env-var-reference workaround.
|
||||||
|
|
||||||
|
### B5. Confirm the default agent model is on `olp-claude`
|
||||||
|
|
||||||
|
Check `agents.defaults.model.primary` in `openclaw.json`. It should be something like:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "agents": { "defaults": { "model": { "primary": "olp-claude/claude-sonnet-4-6" } } } }
|
||||||
|
```
|
||||||
|
|
||||||
|
If it's pointing at one of OpenClaw's stock providers (`openai/...`, `anthropic/...`, `github-copilot/...`), free-text chat will **bypass OLP entirely** and hit your direct API account. You'll see no traffic in OLP's `/dashboard` and `/olp usage` will show no recent activity.
|
||||||
|
|
||||||
|
### B5. Restart the gateway
|
||||||
|
|
||||||
|
```bash
|
||||||
|
openclaw gateway restart
|
||||||
|
```
|
||||||
|
|
||||||
|
### B6. Verify routing
|
||||||
|
|
||||||
|
In Telegram or Discord, send a free-text message ("hello"). It should:
|
||||||
|
1. Return a normal LLM reply (not "Something went wrong")
|
||||||
|
2. Show up on the OLP dashboard's 24h-requests counter
|
||||||
|
3. Show up in `/olp usage` per-provider count
|
||||||
|
|
||||||
|
If you see "Something went wrong" — see § Troubleshooting below.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Using codex / OpenAI models through OLP
|
||||||
|
|
||||||
|
By default the `olp-claude` provider only knows about Claude models. To route OpenAI / codex models through OLP (so bot calls to `gpt-5.5` etc. spawn `codex exec --json` on the OLP server and benefit from per-key audit + quota tracking), add a second provider:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"models": {
|
||||||
|
"providers": {
|
||||||
|
"olp-codex": {
|
||||||
|
"baseUrl": "http://<olp-server-ip>:4567/v1",
|
||||||
|
"api": "openai-completions",
|
||||||
|
"apiKey": "${OLP_OPENCLAW_BOT_TOKEN}",
|
||||||
|
"models": [
|
||||||
|
{ "id": "gpt-5.5", "name": "GPT 5.5 (via OLP→codex)", "input": ["text"],
|
||||||
|
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||||
|
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } },
|
||||||
|
{ "id": "gpt-5.4-mini", "name": "GPT 5.4 mini (via OLP→codex)", "input": ["text"],
|
||||||
|
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||||
|
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } },
|
||||||
|
{ "id": "gpt-5.3-codex", "name": "GPT 5.3 codex (via OLP→codex)", "input": ["text"],
|
||||||
|
"contextWindow": 200000, "maxTokens": 16384, "api": "openai-completions",
|
||||||
|
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"agents": {
|
||||||
|
"defaults": {
|
||||||
|
"models": {
|
||||||
|
"olp-codex/gpt-5.5": { "alias": "OLP GPT 5.5" },
|
||||||
|
"olp-codex/gpt-5.4-mini": { "alias": "OLP GPT 5.4 mini" },
|
||||||
|
"olp-codex/gpt-5.3-codex": { "alias": "OLP Codex" }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
After restart, type `/models` in Telegram and pick `olp-codex/gpt-5.5` from the menu that appears. **`/models` is menu-driven — it does not accept inline model names**; typing `/models olp-codex/gpt-5.5` won't directly switch you. The available IDs are the ones OLP's `/v1/models` returns — typically `gpt-5.5`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.3-codex`, `gpt-5.3-codex-spark`. Query your OLP server to see the live list:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -s -H "Authorization: Bearer olp_…" http://<olp-server-ip>:4567/v1/models | jq '.data[].id'
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why not use OpenClaw's stock `openai` provider?** OpenClaw's built-in `openai-codex` provider uses the local ChatGPT account (via the `sk-proj-…` API key OpenClaw stores) and bypasses your OLP server entirely. You'd lose per-key audit + per-key quota visibility. `olp-codex` keeps everything routed through your central OLP for observability.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Gotchas
|
||||||
|
|
||||||
|
### Auth must use `apiKey: "${VAR}"` env-var reference — three failure modes to avoid
|
||||||
|
|
||||||
|
Custom OpenAI-compatible providers in OpenClaw have a fragile auth path. Three patterns that **don't work** + the one that **does**:
|
||||||
|
|
||||||
|
**❌ `apiKey: "olp_<raw-token>"`** — raw string. Silently bypassed in some routing paths because OpenClaw treats `OPENAI_API_KEY` as service-managed (`OPENCLAW_SERVICE_MANAGED_ENV_KEYS=DEEPSEEK_API_KEY,OPENAI_API_KEY`), and certain model id patterns (notably `gpt-*`) fall back to that env var instead of using your explicit `apiKey`. Symptom: OLP audit shows the request as `__anonymous__` instead of your owner key. Confirmed via [openclaw#41157](https://github.com/openclaw/openclaw/issues/41157) (Gemini openai-completions Authorization not sent) and [#1669](https://github.com/openclaw/openclaw/issues/1669) (Ollama provider ignores apiKey, hardcodes Bearer). Both unresolved upstream as of OpenClaw v2026.5.
|
||||||
|
|
||||||
|
**❌ `headers: { "Authorization": "Bearer olp_<raw-token>" }`** — works for SOME provider/model combinations (e.g., model id `claude-sonnet-4-6`) but breaks for openai-shape model ids (`gpt-5.5` etc.) which take a different code path that ignores the `headers` field. Mixed behavior is worse than no behavior.
|
||||||
|
|
||||||
|
**❌ Setting `OPENAI_API_KEY=olp_…`** in the gateway env. OpenClaw service-manages that variable and overwrites your value with the user's ChatGPT key on every gateway start.
|
||||||
|
|
||||||
|
**✅ `apiKey: "${OLP_OPENCLAW_BOT_TOKEN}"`** — env-var reference with a **custom** variable name (NOT `OPENAI_API_KEY`). OpenClaw resolves the reference at request-construction time, before any service-managed-env logic runs. Both `olp-claude/*` (Claude models) and `olp-codex/*` (OpenAI models) auth correctly with this pattern. Verified end-to-end 2026-05-27: OLP audit shows requests attributed to the correct bot key for both provider blocks.
|
||||||
|
|
||||||
|
OpenClaw docs call out this as the canonical pattern: see [docs.openclaw.ai/concepts/model-providers](https://docs.openclaw.ai/concepts/model-providers) "API key or SecretRef/env reference".
|
||||||
|
|
||||||
|
### Default agent model still points at a removed provider
|
||||||
|
|
||||||
|
If you've removed a provider (e.g., torn down a co-located OCP server) but the bot's default agent model still references that provider, free-text messages will fail with "Something went wrong while processing your request." Check `agents.defaults.model.primary` and update it to a provider that exists.
|
||||||
|
|
||||||
|
### `/new` does not reset model selection — use `/reset`
|
||||||
|
|
||||||
|
OpenClaw's `/new` resets the **conversation context** but **preserves** the session's `/models` selection. If a session has been switched to a model that no longer works (revoked / removed), `/new` won't help — use `/reset` (resets both context and model selection).
|
||||||
|
|
||||||
|
### `/models` is menu-only — does not accept inline model names
|
||||||
|
|
||||||
|
The OpenClaw `/models` command in Telegram is **menu-driven**: typing `/models` pops a model-picker menu where you tap the model name. Typing `/models olp-codex/gpt-5.5` does NOT switch — it'll open the picker. The bot's own success-message after a pick may say *"Use `/model olp-codex/gpt-5.5 --runtime <runtime>` to switch harnesses."* — **that command form is not actually accepted by the bot**; ignore that line.
|
||||||
|
|
||||||
|
### OpenClaw v2026.5+ requires `openclaw.extensions` in `package.json`
|
||||||
|
|
||||||
|
OpenClaw versions ≥ 2026.5.22 enforce a stricter plugin-manifest validation at `openclaw plugins install` time. If `Option A` fails with `package.json missing openclaw.extensions` despite recent OLP releases, your local `olp-plugin/package.json` may predate the v0.5.x fix that adds `"extensions": ["./index.js"]` to the `openclaw` block. Pull latest OLP main (`git pull` in your OLP clone) and retry, or fall through to symlink Option B which works against any plugin shape. (Original drift event: 2026-05-27, see commit history of `olp-plugin/package.json`.)
|
||||||
|
|
||||||
|
### `openclaw gateway restart` is required after install
|
||||||
|
|
||||||
|
OpenClaw caches plugin discovery + model-provider config at gateway start. `openclaw plugins reload` does not guarantee a fresh import of the plugin module nor a fresh re-read of `models.providers.*`. Restart the gateway after every change to `~/.openclaw/openclaw.json`.
|
||||||
|
|
||||||
|
### Owner-key revocation kicks the plugin out immediately
|
||||||
|
|
||||||
|
If you revoke the bot's owner key (`npx olp-keys revoke --id=<id>`), the next `/olp status` will return `401 unauthorized`. Mint a replacement key with a new name and edit `~/.openclaw/openclaw.json`; do NOT reuse the revoked key's UUID.
|
||||||
|
|
||||||
|
### Long responses are truncated
|
||||||
|
|
||||||
|
Telegram caps messages at ~4096 characters. The plugin truncates with a `... [truncated, use SSH for full]` suffix when the rendered output would exceed ~3900 chars. Use SSH + the local `olp` CLI for full output.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## OLP-specific notes
|
## OLP-specific notes
|
||||||
|
|
||||||
The plugin honours these env vars on the OpenClaw gateway process:
|
The plugin honours these env vars on the OpenClaw gateway process:
|
||||||
|
|
||||||
- `OLP_PROXY_URL` — full URL, overrides plugin config `proxyUrl`.
|
- `OLP_PROXY_URL` — full URL, overrides plugin config `proxyUrl`.
|
||||||
- `OLP_PORT` — port only, localhost assumed; overrides `proxyUrl` when
|
- `OLP_PORT` — port only, localhost assumed; overrides `proxyUrl` when `OLP_PROXY_URL` is unset.
|
||||||
`OLP_PROXY_URL` is unset.
|
|
||||||
|
|
||||||
If you run the OpenClaw gateway under launchd or systemd with custom env
|
If you run the OpenClaw gateway under launchd or systemd with custom env vars, set `OLP_PROXY_URL` there rather than editing the plugin config — that way the same plugin install can serve multiple OLP hosts.
|
||||||
vars, set `OLP_PROXY_URL` there rather than editing the plugin config —
|
|
||||||
that way the same plugin install can serve multiple OLP hosts.
|
|
||||||
|
|
||||||
## Per-bot vs maintainer key
|
## Per-bot vs maintainer key
|
||||||
|
|
||||||
**Always create a dedicated bot key**, never the maintainer's personal
|
**Always create a dedicated bot key**, never the maintainer's personal owner key. The bot key:
|
||||||
owner key. The bot key:
|
|
||||||
|
|
||||||
- Has its own `id` so you can revoke it without affecting other clients.
|
- Has its own `id` so you can revoke it without affecting other clients.
|
||||||
- Has its own audit-log entries so you can attribute `/v0/management/*`
|
- Has its own audit-log entries so you can attribute `/v0/management/*` traffic to the bot.
|
||||||
traffic to the bot.
|
- Can be rotated routinely (every 90 days etc.) without coordinating with the maintainer's daily-driver IDE configs.
|
||||||
- Can be rotated routinely (every 90 days etc.) without coordinating with
|
|
||||||
the maintainer's daily-driver IDE configs.
|
|
||||||
|
|
||||||
## Test it
|
## Troubleshooting
|
||||||
|
|
||||||
After restart, in Telegram or Discord:
|
| Symptom | Likely cause | Fix |
|
||||||
|
|---|---|---|
|
||||||
```
|
| `/olp status` returns 401 | bot key revoked / wrong / missing | Mint new key on OLP host; update `plugins.entries.olp.config.apiKey`; restart gateway |
|
||||||
/olp health
|
| `/olp status` returns 403 | bot key is guest-tier, not owner-tier | Generate owner-tier key (`olp-keys keygen --owner --name=...`); update config |
|
||||||
/olp status
|
| `OLP error: fetch failed` | `proxyUrl` unreachable from the gateway host | `curl http://<proxyUrl>/health` from the gateway host to confirm reachability; check firewall / OLP server `OLP_BIND=0.0.0.0` for LAN access |
|
||||||
/olp models
|
| Bot free-text chat returns "Something went wrong" but `/olp ...` works | Default agent model points at a broken provider (e.g., a removed OCP install) | Check `agents.defaults.model.primary` in `openclaw.json`; update to `olp-claude/claude-sonnet-4-6` or another working provider |
|
||||||
```
|
| Free-text returns `HTTP 401: OLP API key is invalid` despite fresh key | Raw-string `apiKey: "olp_..."` shadowed by service-managed env clobber; or `headers.Authorization` bypassed for `gpt-*` model ids | Switch `models.providers.<id>.apiKey` to env-var reference: `"${OLP_OPENCLAW_BOT_TOKEN}"` (see § Gotchas: Auth) |
|
||||||
|
| OLP audit shows `key_id=__anonymous__` for traffic that should be owner-attributed | Same root cause as 401 — raw-string apiKey or headers bypassed in some routing paths | Switch to env-var-reference `apiKey: "${VAR}"` pattern + verify `launchctl getenv OLP_OPENCLAW_BOT_TOKEN` returns the expected token |
|
||||||
Each should return a code-block-wrapped response within a few seconds.
|
| Bot routes to ChatGPT account directly, not through OLP | Provider config uses OpenClaw stock `openai-codex` instead of a custom OLP-pointing provider | Add `olp-codex` provider per § Using codex / OpenAI models through OLP |
|
||||||
|
| `/models olp-codex/gpt-5.5` typed inline doesn't work | OpenClaw `/models` is menu-only, doesn't accept inline names | Type `/models`, tap the model from the picker menu that appears |
|
||||||
If you see `401 unauthorized`: the configured key is missing / wrong /
|
|
||||||
revoked. If you see `403 forbidden`: the key is not owner-tier. If you
|
|
||||||
see `OLP error: fetch failed` or similar: the `proxyUrl` is unreachable
|
|
||||||
from the gateway host (test with `curl http://<proxyUrl>/health` from
|
|
||||||
that host).
|
|
||||||
|
|
||||||
## Cross-references
|
## Cross-references
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,655 @@
|
|||||||
|
# OLP Cloud Deployment Plan — Family Testing Phase
|
||||||
|
|
||||||
|
**Status:** Draft — pending current Phase 6 completion
|
||||||
|
**Target:** Oracle Cloud VM (existing infrastructure)
|
||||||
|
**Audience:** Project maintainer deployment reference
|
||||||
|
**Scope:** Single-VM deployment for family (3–5 users), spawn-binary architecture, public internet exposure with hardened auth
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 0. Prerequisites
|
||||||
|
|
||||||
|
- OLP current phase (Phase 6) is closed and tagged
|
||||||
|
- Oracle Cloud VM accessible via SSH (existing `opc` user)
|
||||||
|
- Domain name (optional but strongly recommended for TLS)
|
||||||
|
- Provider CLI OAuth completed on at least one machine (credentials transferable)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Architecture Overview
|
||||||
|
|
||||||
|
```
|
||||||
|
┌────────────────────────────────────────────────────────────────────┐
|
||||||
|
│ Family Devices (anywhere on internet) │
|
||||||
|
│ │
|
||||||
|
│ Wife iPad / Kid Laptop / Maintainer MacBook / ... │
|
||||||
|
│ IDE: Cline / Continue.dev / Cursor / Aider / OpenClaw │
|
||||||
|
│ Config: OPENAI_BASE_URL=https://olp.example.com/v1 │
|
||||||
|
│ OPENAI_API_KEY=olp_<personal-key> │
|
||||||
|
└──────────────────────────┬─────────────────────────────────────────┘
|
||||||
|
│ HTTPS (TLS 1.3)
|
||||||
|
▼
|
||||||
|
┌────────────────────────────────────────────────────────────────────┐
|
||||||
|
│ Oracle Cloud VM │
|
||||||
|
│ │
|
||||||
|
│ ┌─ iptables / OCI Security List ──────────────────────────────┐ │
|
||||||
|
│ │ ALLOW: TCP 443 (HTTPS) from 0.0.0.0/0 │ │
|
||||||
|
│ │ ALLOW: TCP 22 (SSH) from maintainer IP only │ │
|
||||||
|
│ │ DENY: everything else │ │
|
||||||
|
│ └─────────────────────────────────────────────────────────────┘ │
|
||||||
|
│ │
|
||||||
|
│ ┌─ Nginx (reverse proxy + TLS termination) ───────────────────┐ │
|
||||||
|
│ │ :443 → TLS (Let's Encrypt auto-renew via certbot) │ │
|
||||||
|
│ │ proxy_pass → http://127.0.0.1:4567 │ │
|
||||||
|
│ │ Rate limit: 30 req/min per IP (burst 10) │ │
|
||||||
|
│ │ Request body limit: 1MB │ │
|
||||||
|
│ │ Connection timeout: 300s (streaming needs long timeout) │ │
|
||||||
|
│ └─────────────────────────────────────────────────────────────┘ │
|
||||||
|
│ │
|
||||||
|
│ ┌─ OLP server.mjs ───────────────────────────────────────────┐ │
|
||||||
|
│ │ OLP_BIND=127.0.0.1 (loopback only — Nginx fronts it) │ │
|
||||||
|
│ │ OLP_PORT=4567 │ │
|
||||||
|
│ │ auth.allow_anonymous: false │ │
|
||||||
|
│ │ auth.advertise_anonymous_key: false │ │
|
||||||
|
│ │ Per-key audit logging to ~/.olp/logs/audit.ndjson │ │
|
||||||
|
│ │ Owner key: maintainer only │ │
|
||||||
|
│ │ Guest keys: one per family member │ │
|
||||||
|
│ └─────────────────────────────────────────────────────────────┘ │
|
||||||
|
│ │
|
||||||
|
│ ┌─ Provider CLIs (installed on this VM) ──────────────────────┐ │
|
||||||
|
│ │ claude → ~/.claude/.credentials.json (OAuth) │ │
|
||||||
|
│ │ codex → ~/.codex/auth.json (OAuth) │ │
|
||||||
|
│ │ vibe → ~/.vibe/.env (API key) │ │
|
||||||
|
│ └─────────────────────────────────────────────────────────────┘ │
|
||||||
|
│ │
|
||||||
|
│ ┌─ systemd service ──────────────────────────────────────────┐ │
|
||||||
|
│ │ olp.service: auto-start, auto-restart on crash │ │
|
||||||
|
│ │ Runs as dedicated `olp` user (not root, not opc) │ │
|
||||||
|
│ └─────────────────────────────────────────────────────────────┘ │
|
||||||
|
│ │
|
||||||
|
└────────────────────────────────────────────────────────────────────┘
|
||||||
|
│
|
||||||
|
│ Provider CLIs spawn outbound HTTPS calls
|
||||||
|
▼
|
||||||
|
Anthropic API / OpenAI API / Mistral API
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Security Design (7 Layers)
|
||||||
|
|
||||||
|
### Layer 1 — Network Perimeter (OCI Security List + iptables)
|
||||||
|
|
||||||
|
**Principle:** Minimum attack surface. Only two ports reachable from the internet.
|
||||||
|
|
||||||
|
```
|
||||||
|
OCI Security List (stateful ingress rules):
|
||||||
|
┌──────────┬────────────┬───────────────────────────────┐
|
||||||
|
│ Port │ Protocol │ Source │
|
||||||
|
├──────────┼────────────┼───────────────────────────────┤
|
||||||
|
│ 443 │ TCP │ 0.0.0.0/0 (public HTTPS) │
|
||||||
|
│ 22 │ TCP │ <maintainer-IP>/32 only │
|
||||||
|
└──────────┴────────────┴───────────────────────────────┘
|
||||||
|
|
||||||
|
NOT exposed:
|
||||||
|
- Port 4567 (OLP direct) — Nginx fronts it
|
||||||
|
- Port 80 (HTTP) — only for certbot ACME challenge, redirect to 443
|
||||||
|
```
|
||||||
|
|
||||||
|
**iptables backup** (defense in depth — OCI Security List is primary, iptables is secondary):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Drop everything by default
|
||||||
|
sudo iptables -P INPUT DROP
|
||||||
|
sudo iptables -P FORWARD DROP
|
||||||
|
|
||||||
|
# Allow established connections
|
||||||
|
sudo iptables -A INPUT -m state --state ESTABLISHED,RELATED -j ACCEPT
|
||||||
|
|
||||||
|
# Allow loopback
|
||||||
|
sudo iptables -A INPUT -i lo -j ACCEPT
|
||||||
|
|
||||||
|
# Allow SSH from maintainer IP only
|
||||||
|
sudo iptables -A INPUT -p tcp --dport 22 -s <MAINTAINER_IP> -j ACCEPT
|
||||||
|
|
||||||
|
# Allow HTTPS from anywhere
|
||||||
|
sudo iptables -A INPUT -p tcp --dport 443 -j ACCEPT
|
||||||
|
|
||||||
|
# Allow HTTP (certbot ACME only — Nginx redirects everything else)
|
||||||
|
sudo iptables -A INPUT -p tcp --dport 80 -j ACCEPT
|
||||||
|
|
||||||
|
# Persist
|
||||||
|
sudo iptables-save | sudo tee /etc/iptables/rules.v4
|
||||||
|
```
|
||||||
|
|
||||||
|
### Layer 2 — TLS Termination (Nginx + Let's Encrypt)
|
||||||
|
|
||||||
|
**Principle:** All client traffic encrypted. OLP itself runs plain HTTP on loopback — simpler, no cert management in Node.
|
||||||
|
|
||||||
|
```nginx
|
||||||
|
# /etc/nginx/sites-available/olp.conf
|
||||||
|
|
||||||
|
# Redirect HTTP → HTTPS
|
||||||
|
server {
|
||||||
|
listen 80;
|
||||||
|
server_name olp.example.com;
|
||||||
|
|
||||||
|
# Let's Encrypt ACME challenge
|
||||||
|
location /.well-known/acme-challenge/ {
|
||||||
|
root /var/www/certbot;
|
||||||
|
}
|
||||||
|
|
||||||
|
location / {
|
||||||
|
return 301 https://$host$request_uri;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
# HTTPS — TLS 1.3 only
|
||||||
|
server {
|
||||||
|
listen 443 ssl http2;
|
||||||
|
server_name olp.example.com;
|
||||||
|
|
||||||
|
# TLS config
|
||||||
|
ssl_certificate /etc/letsencrypt/live/olp.example.com/fullchain.pem;
|
||||||
|
ssl_certificate_key /etc/letsencrypt/live/olp.example.com/privkey.pem;
|
||||||
|
ssl_protocols TLSv1.3; # TLS 1.3 only
|
||||||
|
ssl_prefer_server_ciphers off; # TLS 1.3 manages its own
|
||||||
|
ssl_session_timeout 1d;
|
||||||
|
ssl_session_cache shared:SSL:10m;
|
||||||
|
|
||||||
|
# Security headers
|
||||||
|
add_header Strict-Transport-Security "max-age=63072000" always;
|
||||||
|
add_header X-Content-Type-Options nosniff;
|
||||||
|
add_header X-Frame-Options DENY;
|
||||||
|
|
||||||
|
# Rate limiting (per IP)
|
||||||
|
limit_req zone=olp_limit burst=10 nodelay;
|
||||||
|
|
||||||
|
# Request body size (LLM prompts can be large but cap at 1MB)
|
||||||
|
client_max_body_size 1m;
|
||||||
|
|
||||||
|
# Proxy to OLP
|
||||||
|
location / {
|
||||||
|
proxy_pass http://127.0.0.1:4567;
|
||||||
|
proxy_http_version 1.1;
|
||||||
|
proxy_set_header Host $host;
|
||||||
|
proxy_set_header X-Real-IP $remote_addr;
|
||||||
|
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||||
|
proxy_set_header X-Forwarded-Proto $scheme;
|
||||||
|
|
||||||
|
# SSE streaming support (critical for /v1/chat/completions)
|
||||||
|
proxy_set_header Connection '';
|
||||||
|
proxy_buffering off; # Don't buffer SSE
|
||||||
|
proxy_cache off;
|
||||||
|
chunked_transfer_encoding on;
|
||||||
|
|
||||||
|
# Long timeouts for LLM inference
|
||||||
|
proxy_connect_timeout 10s;
|
||||||
|
proxy_read_timeout 300s; # 5 min — long reasoning
|
||||||
|
proxy_send_timeout 300s;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
# Rate limit zone definition (in http {} block of nginx.conf)
|
||||||
|
# limit_req_zone $binary_remote_addr zone=olp_limit:10m rate=30r/m;
|
||||||
|
```
|
||||||
|
|
||||||
|
**Certbot auto-renewal:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
sudo certbot certonly --webroot -w /var/www/certbot -d olp.example.com
|
||||||
|
# Auto-renew via systemd timer (certbot installs this automatically)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Layer 3 — Application Auth (OLP Multi-Key)
|
||||||
|
|
||||||
|
**Principle:** Every request must carry a valid API key. No anonymous access. Per-key audit trail.
|
||||||
|
|
||||||
|
```json
|
||||||
|
// ~/.olp/config.json on the cloud VM
|
||||||
|
{
|
||||||
|
"auth": {
|
||||||
|
"allow_anonymous": false,
|
||||||
|
"advertise_anonymous_key": false,
|
||||||
|
"owner_only_endpoints": [
|
||||||
|
"/health",
|
||||||
|
"/v0/management/dashboard-data",
|
||||||
|
"/v0/management/quota",
|
||||||
|
"/v0/management/status",
|
||||||
|
"/cache/stats",
|
||||||
|
"/dashboard"
|
||||||
|
],
|
||||||
|
"fallback_detail_header_policy": "owner_only"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key provisioning plan:**
|
||||||
|
|
||||||
|
```
|
||||||
|
┌───────────────┬──────────┬─────────────────────────────────────┐
|
||||||
|
│ Key name │ Tier │ providers_enabled │
|
||||||
|
├───────────────┼──────────┼─────────────────────────────────────┤
|
||||||
|
│ cloud-owner │ owner │ all (dashboard + management access) │
|
||||||
|
│ wife-ipad │ guest │ anthropic, openai │
|
||||||
|
│ kid-laptop │ guest │ anthropic only (cost control) │
|
||||||
|
│ maintainer-mb │ guest │ all (daily driver, not owner tier) │
|
||||||
|
└───────────────┴──────────┴─────────────────────────────────────┘
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why maintainer uses a guest key for daily driving:** owner key gives access to management endpoints. Routine IDE usage should not carry owner privilege. Owner key is used only for dashboard access and administration.
|
||||||
|
|
||||||
|
**Key lifecycle:**
|
||||||
|
- Keys generated on the cloud VM via `olp-keys keygen`
|
||||||
|
- Plaintext token communicated to family member via secure channel (Signal / iMessage, not email)
|
||||||
|
- Each key logged independently in audit.ndjson (per-key `key_id` field)
|
||||||
|
- Revocation: `olp-keys revoke --id=<key-id>` — immediate, no grace period
|
||||||
|
|
||||||
|
### Layer 4 — Process Isolation (Dedicated User + systemd)
|
||||||
|
|
||||||
|
**Principle:** OLP runs as a non-root, non-login user. Crash recovery is automatic.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Create dedicated user
|
||||||
|
sudo useradd --system --shell /usr/sbin/nologin --home-dir /opt/olp olp
|
||||||
|
|
||||||
|
# OLP code
|
||||||
|
sudo mkdir -p /opt/olp
|
||||||
|
sudo git clone https://github.com/dtzp555-max/olp.git /opt/olp/app
|
||||||
|
sudo chown -R olp:olp /opt/olp
|
||||||
|
|
||||||
|
# OLP data (keys, config, logs, cache)
|
||||||
|
sudo mkdir -p /home/olp/.olp/{keys,logs,cache}
|
||||||
|
sudo chown -R olp:olp /home/olp
|
||||||
|
```
|
||||||
|
|
||||||
|
**systemd unit:**
|
||||||
|
|
||||||
|
```ini
|
||||||
|
# /etc/systemd/system/olp.service
|
||||||
|
[Unit]
|
||||||
|
Description=OLP — Open LLM Proxy
|
||||||
|
After=network-online.target
|
||||||
|
Wants=network-online.target
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=simple
|
||||||
|
User=olp
|
||||||
|
Group=olp
|
||||||
|
|
||||||
|
WorkingDirectory=/opt/olp/app
|
||||||
|
ExecStart=/usr/bin/node server.mjs
|
||||||
|
|
||||||
|
# Environment
|
||||||
|
Environment=OLP_BIND=127.0.0.1
|
||||||
|
Environment=OLP_PORT=4567
|
||||||
|
Environment=NODE_ENV=production
|
||||||
|
Environment=HOME=/home/olp
|
||||||
|
|
||||||
|
# Auto-restart on crash
|
||||||
|
Restart=on-failure
|
||||||
|
RestartSec=5
|
||||||
|
StartLimitIntervalSec=60
|
||||||
|
StartLimitBurst=5
|
||||||
|
|
||||||
|
# Security hardening
|
||||||
|
NoNewPrivileges=true
|
||||||
|
ProtectSystem=strict
|
||||||
|
ProtectHome=false
|
||||||
|
ReadWritePaths=/home/olp/.olp
|
||||||
|
PrivateTmp=true
|
||||||
|
|
||||||
|
# Resource limits
|
||||||
|
LimitNOFILE=65536
|
||||||
|
MemoryMax=1G
|
||||||
|
|
||||||
|
# Logging
|
||||||
|
StandardOutput=journal
|
||||||
|
StandardError=journal
|
||||||
|
SyslogIdentifier=olp
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=multi-user.target
|
||||||
|
```
|
||||||
|
|
||||||
|
### Layer 5 — Credential Protection (Provider OAuth Tokens)
|
||||||
|
|
||||||
|
**Principle:** OAuth tokens are the crown jewels. Stolen tokens = someone else using your Claude/OpenAI subscription.
|
||||||
|
|
||||||
|
```
|
||||||
|
Credential storage on cloud VM:
|
||||||
|
|
||||||
|
~olp/
|
||||||
|
├── .claude/
|
||||||
|
│ └── .credentials.json # chmod 600, owner=olp
|
||||||
|
├── .codex/
|
||||||
|
│ └── auth.json # chmod 600, owner=olp
|
||||||
|
└── .vibe/
|
||||||
|
└── .env # chmod 600, owner=olp
|
||||||
|
|
||||||
|
Security measures:
|
||||||
|
1. chmod 600 on all credential files (olp user only)
|
||||||
|
2. Credential files NOT in the git repo (already .gitignored)
|
||||||
|
3. No credential in env vars (OLP reads from filesystem)
|
||||||
|
4. Credential transfer: scp from local machine, then delete local copy of the scp command from shell history
|
||||||
|
5. Periodic rotation: re-auth quarterly (or on any suspicion of compromise)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Credential transfer procedure:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# FROM maintainer's Mac mini (one-time):
|
||||||
|
|
||||||
|
# 1. Claude credentials
|
||||||
|
scp ~/.claude/.credentials.json opc@<cloud-ip>:/tmp/claude-cred.json
|
||||||
|
ssh opc@<cloud-ip> "sudo mv /tmp/claude-cred.json /home/olp/.claude/.credentials.json && sudo chown olp:olp /home/olp/.claude/.credentials.json && sudo chmod 600 /home/olp/.claude/.credentials.json"
|
||||||
|
|
||||||
|
# 2. Codex credentials
|
||||||
|
scp ~/.codex/auth.json opc@<cloud-ip>:/tmp/codex-cred.json
|
||||||
|
ssh opc@<cloud-ip> "sudo mv /tmp/codex-cred.json /home/olp/.codex/auth.json && sudo chown olp:olp /home/olp/.codex/auth.json && sudo chmod 600 /home/olp/.codex/auth.json"
|
||||||
|
|
||||||
|
# 3. Mistral API key
|
||||||
|
ssh opc@<cloud-ip> "sudo -u olp bash -c 'echo MISTRAL_API_KEY=sk-xxx > ~/.vibe/.env && chmod 600 ~/.vibe/.env'"
|
||||||
|
|
||||||
|
# 4. Verify
|
||||||
|
ssh opc@<cloud-ip> "sudo -u olp node /opt/olp/app/bin/olp.mjs doctor --json" | jq '.checks[] | select(.name | contains("auth"))'
|
||||||
|
```
|
||||||
|
|
||||||
|
### Layer 6 — Audit and Monitoring
|
||||||
|
|
||||||
|
**Principle:** Every request logged. Anomalies detectable. No silent failures.
|
||||||
|
|
||||||
|
**Audit (already built into OLP):**
|
||||||
|
- `~/.olp/logs/audit.ndjson` — append-only, per-request, includes `key_id`, provider, model, cache hit/miss, fallback hops
|
||||||
|
- Daily rotation: `audit-YYYY-MM-DD.ndjson` (built-in, triggers on first append after UTC midnight)
|
||||||
|
- External rotation tool: `olp-audit-rotate` (idempotent, cron-safe)
|
||||||
|
|
||||||
|
**Additional monitoring for cloud deployment:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Cron: daily audit rotation (belt-and-suspenders alongside in-server rotation)
|
||||||
|
0 0 * * * /usr/bin/node /opt/olp/app/bin/olp-audit-rotate.mjs
|
||||||
|
|
||||||
|
# Cron: daily health check + alert
|
||||||
|
*/5 * * * * curl -sf -H "Authorization: Bearer $OLP_OWNER_KEY" https://olp.example.com/health > /dev/null || echo "OLP health check failed at $(date)" >> /home/olp/alerts.log
|
||||||
|
|
||||||
|
# Cron: audit log size check (alert if >100MB — suggests anomalous traffic)
|
||||||
|
0 6 * * * find /home/olp/.olp/logs -name 'audit*.ndjson' -size +100M -exec echo "Large audit log: {}" \; >> /home/olp/alerts.log
|
||||||
|
|
||||||
|
# Cron: disk usage check
|
||||||
|
0 6 * * * df -h / | awk 'NR==2 && $5+0 > 80 {print "Disk usage above 80%: "$5}' >> /home/olp/alerts.log
|
||||||
|
```
|
||||||
|
|
||||||
|
**What to watch for (manually, weekly):**
|
||||||
|
1. `olp-keys list` — any unexpected keys?
|
||||||
|
2. Dashboard (`/dashboard`) — unusual request volume? Unknown providers being hit?
|
||||||
|
3. `journalctl -u olp --since "7 days ago" | grep -c ERROR` — error spike?
|
||||||
|
4. Audit log: `grep "fallback" ~/.olp/logs/audit.ndjson | wc -l` — fallback frequency (high = provider instability)
|
||||||
|
|
||||||
|
### Layer 7 — Update and Recovery
|
||||||
|
|
||||||
|
**Principle:** Rollback within 60 seconds. No data loss on failed update.
|
||||||
|
|
||||||
|
**Update procedure:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# SSH to cloud VM as opc
|
||||||
|
|
||||||
|
# 1. Snapshot before update (Oracle Cloud console or CLI)
|
||||||
|
# OCI CLI: oci compute boot-volume-backup create ...
|
||||||
|
|
||||||
|
# 2. Pull latest code
|
||||||
|
cd /opt/olp/app
|
||||||
|
sudo -u olp git fetch origin main
|
||||||
|
sudo -u olp git log --oneline HEAD..origin/main # review what's coming
|
||||||
|
|
||||||
|
# 3. Run tests BEFORE deploying
|
||||||
|
sudo -u olp git checkout main
|
||||||
|
sudo -u olp git pull
|
||||||
|
sudo -u olp node test-features.mjs
|
||||||
|
# STOP if tests fail
|
||||||
|
|
||||||
|
# 4. Restart service
|
||||||
|
sudo systemctl restart olp
|
||||||
|
sleep 3
|
||||||
|
sudo systemctl status olp # verify running
|
||||||
|
|
||||||
|
# 5. Smoke test
|
||||||
|
curl -sf -H "Authorization: Bearer $OLP_OWNER_KEY" https://olp.example.com/health | jq .ok
|
||||||
|
# Expect: true
|
||||||
|
```
|
||||||
|
|
||||||
|
**Rollback:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# If update breaks things:
|
||||||
|
cd /opt/olp/app
|
||||||
|
sudo -u olp git checkout <previous-tag> # e.g. v0.6.0
|
||||||
|
sudo systemctl restart olp
|
||||||
|
```
|
||||||
|
|
||||||
|
**Backup (automated):**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Cron: daily backup of OLP state (keys + config + recent audit)
|
||||||
|
0 3 * * * tar czf /home/opc/backups/olp-state-$(date +\%Y\%m\%d).tar.gz -C /home/olp .olp/keys .olp/config.json .olp/logs/audit.ndjson 2>/dev/null; find /home/opc/backups -name 'olp-state-*' -mtime +30 -delete
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Implementation Checklist
|
||||||
|
|
||||||
|
Execute in order. Each step has a verification gate — do not proceed if the gate fails.
|
||||||
|
|
||||||
|
### Phase A — VM Preparation
|
||||||
|
|
||||||
|
```
|
||||||
|
[ ] A1. SSH to Oracle Cloud VM, verify Node.js >= 18
|
||||||
|
Gate: `node --version` prints v18+
|
||||||
|
|
||||||
|
[ ] A2. Create `olp` system user
|
||||||
|
Gate: `id olp` shows the user exists
|
||||||
|
|
||||||
|
[ ] A3. Clone OLP repo to /opt/olp/app
|
||||||
|
Gate: `sudo -u olp node /opt/olp/app/test-features.mjs` — all tests pass
|
||||||
|
|
||||||
|
[ ] A4. Install provider CLIs (as olp user)
|
||||||
|
- npm install -g @anthropic-ai/claude-code
|
||||||
|
- npm install -g @openai/codex
|
||||||
|
- (mistral vibe if needed)
|
||||||
|
Gate: `which claude && which codex` both resolve
|
||||||
|
|
||||||
|
[ ] A5. Transfer OAuth credentials (Layer 5 procedure)
|
||||||
|
Gate: `sudo -u olp claude auth status` shows authenticated
|
||||||
|
```
|
||||||
|
|
||||||
|
### Phase B — Security Hardening
|
||||||
|
|
||||||
|
```
|
||||||
|
[ ] B1. Configure OCI Security List (Layer 1)
|
||||||
|
Gate: nmap from external IP shows only 22 and 443 open
|
||||||
|
|
||||||
|
[ ] B2. Configure iptables backup (Layer 1)
|
||||||
|
Gate: `sudo iptables -L -n` matches the plan
|
||||||
|
|
||||||
|
[ ] B3. Install + configure Nginx (Layer 2)
|
||||||
|
Gate: `curl -I http://olp.example.com` returns 301 → HTTPS
|
||||||
|
|
||||||
|
[ ] B4. Obtain Let's Encrypt certificate
|
||||||
|
Gate: `curl -I https://olp.example.com` returns valid cert
|
||||||
|
|
||||||
|
[ ] B5. Verify Nginx SSE passthrough
|
||||||
|
Gate: test streaming request completes without timeout
|
||||||
|
```
|
||||||
|
|
||||||
|
### Phase C — OLP Configuration
|
||||||
|
|
||||||
|
```
|
||||||
|
[ ] C1. Write ~/.olp/config.json (Layer 3 — auth config)
|
||||||
|
Gate: config validates (no startup warnings in journal)
|
||||||
|
|
||||||
|
[ ] C2. Generate owner key
|
||||||
|
Gate: `olp-keys list --owner-only` shows 1 owner key
|
||||||
|
|
||||||
|
[ ] C3. Generate family guest keys (one per person)
|
||||||
|
Gate: `olp-keys list` shows correct count
|
||||||
|
|
||||||
|
[ ] C4. Install systemd unit (Layer 4)
|
||||||
|
Gate: `systemctl status olp` shows active (running)
|
||||||
|
|
||||||
|
[ ] C5. Verify /health with owner key
|
||||||
|
Gate: `curl -H "Authorization: Bearer $OWNER_KEY" https://olp.example.com/health | jq .ok` → true
|
||||||
|
|
||||||
|
[ ] C6. Verify /health rejects unauthenticated
|
||||||
|
Gate: `curl https://olp.example.com/health` → 401
|
||||||
|
|
||||||
|
[ ] C7. Verify guest key cannot access /dashboard
|
||||||
|
Gate: `curl -H "Authorization: Bearer $GUEST_KEY" https://olp.example.com/dashboard` → 403
|
||||||
|
|
||||||
|
[ ] C8. End-to-end LLM request with guest key
|
||||||
|
Gate: streaming chat completion returns a valid response
|
||||||
|
```
|
||||||
|
|
||||||
|
### Phase D — Monitoring Setup
|
||||||
|
|
||||||
|
```
|
||||||
|
[ ] D1. Install cron jobs (Layer 6)
|
||||||
|
Gate: `crontab -l` shows all 4 jobs
|
||||||
|
|
||||||
|
[ ] D2. Verify daily backup cron
|
||||||
|
Gate: manual trigger produces valid tar.gz
|
||||||
|
|
||||||
|
[ ] D3. Test health-check alert
|
||||||
|
Gate: stop OLP, wait 5min, check alerts.log has entry
|
||||||
|
```
|
||||||
|
|
||||||
|
### Phase E — Family Onboarding
|
||||||
|
|
||||||
|
```
|
||||||
|
[ ] E1. Send each family member their API key via Signal/iMessage
|
||||||
|
(NOT via email, NOT via any cloud-stored medium)
|
||||||
|
|
||||||
|
[ ] E2. Each family member configures their IDE:
|
||||||
|
export OPENAI_BASE_URL=https://olp.example.com/v1
|
||||||
|
export OPENAI_API_KEY=olp_<their-key>
|
||||||
|
|
||||||
|
[ ] E3. Each family member runs a test prompt
|
||||||
|
Gate: audit.ndjson shows their key_id in the log
|
||||||
|
|
||||||
|
[ ] E4. Verify per-key provider scoping
|
||||||
|
Gate: kid's key cannot hit providers outside their scope
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Security Threat Model
|
||||||
|
|
||||||
|
| Threat | Mitigation | Residual Risk |
|
||||||
|
|---|---|---|
|
||||||
|
| **Brute-force API key** | 32-byte entropy = 2^256 keyspace; Nginx rate limit 30r/m | Negligible |
|
||||||
|
| **TLS downgrade** | TLS 1.3 only; HSTS header | None with modern clients |
|
||||||
|
| **Credential theft (OAuth tokens on VM)** | chmod 600 + dedicated user + no root access to OLP dirs | VM root compromise (mitigated by OCI IAM) |
|
||||||
|
| **Stolen guest key** | Single-key revocation via `olp-keys revoke`; per-key audit trail for forensics | Window between theft and detection |
|
||||||
|
| **DDoS** | OCI DDoS protection (free tier) + Nginx rate limit + Nginx connection limit | Sustained volumetric attack may overwhelm free-tier VM |
|
||||||
|
| **Provider credential abuse** | OLP is the only consumer; anomalous spend visible on provider dashboard | Provider-side detection lag |
|
||||||
|
| **Supply chain (OLP code tampered)** | Git clone from known repo; `npm test` before deploy; no npm dependencies | Compromised maintainer GitHub account |
|
||||||
|
| **Log exfiltration** | audit.ndjson contains no message content (PII guard per ADR 0008); only metadata | Key IDs in logs (low sensitivity) |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Operational Runbooks
|
||||||
|
|
||||||
|
### Runbook: OAuth Token Expired
|
||||||
|
|
||||||
|
```
|
||||||
|
Symptom: /health shows provider auth.ok=false; fallback firing on every request
|
||||||
|
Diagnosis: sudo -u olp claude auth status → "not authenticated" or expired
|
||||||
|
|
||||||
|
Fix:
|
||||||
|
1. sudo -u olp claude setup-token
|
||||||
|
2. Complete OAuth flow (browser URL → paste code)
|
||||||
|
3. Verify: sudo -u olp claude auth status → authenticated
|
||||||
|
4. No OLP restart needed — next spawn picks up new credentials
|
||||||
|
```
|
||||||
|
|
||||||
|
### Runbook: Revoke a Compromised Key
|
||||||
|
|
||||||
|
```
|
||||||
|
Symptom: suspicious traffic in audit.ndjson from a specific key_id
|
||||||
|
grep "<suspected-key-id>" ~/.olp/logs/audit.ndjson | tail -20
|
||||||
|
|
||||||
|
Fix:
|
||||||
|
1. olp-keys revoke --id=<key-id>
|
||||||
|
2. Notify family member: "Your key was revoked. Here's a new one."
|
||||||
|
3. olp-keys keygen --name=<new-name> --providers=<same-providers>
|
||||||
|
4. Send new key via secure channel
|
||||||
|
```
|
||||||
|
|
||||||
|
### Runbook: VM Disk Full
|
||||||
|
|
||||||
|
```
|
||||||
|
Symptom: OLP stops writing audit logs; new requests may fail
|
||||||
|
Diagnosis: df -h /
|
||||||
|
|
||||||
|
Fix:
|
||||||
|
1. Purge old audit logs: find ~/.olp/logs -name 'audit-202*.ndjson' -mtime +90 -delete
|
||||||
|
2. Purge old backups: find /home/opc/backups -name 'olp-state-*' -mtime +60 -delete
|
||||||
|
3. Purge cache if needed: rm -rf ~/.olp/cache/*
|
||||||
|
4. Verify: df -h / shows >20% free
|
||||||
|
```
|
||||||
|
|
||||||
|
### Runbook: OLP Process Crash Loop
|
||||||
|
|
||||||
|
```
|
||||||
|
Symptom: systemctl status olp shows "activating (auto-restart)"
|
||||||
|
Diagnosis: journalctl -u olp --since "10 min ago" | tail -50
|
||||||
|
|
||||||
|
Common causes:
|
||||||
|
- Port conflict → check `lsof -nP -iTCP:4567`
|
||||||
|
- Corrupt config.json → validate JSON syntax
|
||||||
|
- Node.js version drift → `node --version`
|
||||||
|
|
||||||
|
Fix:
|
||||||
|
1. Fix root cause
|
||||||
|
2. sudo systemctl restart olp
|
||||||
|
3. Gate: `curl -H "Authorization: Bearer $OWNER_KEY" https://olp.example.com/health | jq .ok`
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Cost Estimate (Oracle Cloud Free Tier)
|
||||||
|
|
||||||
|
| Resource | Spec | Cost |
|
||||||
|
|---|---|---|
|
||||||
|
| VM | ARM Ampere A1 (4 OCPU, 24GB RAM) | **Free** (Always Free tier) |
|
||||||
|
| Boot volume | 200GB | **Free** (up to 200GB) |
|
||||||
|
| Outbound bandwidth | 10TB/month | **Free** (first 10TB) |
|
||||||
|
| Public IP | 1 reserved | **Free** |
|
||||||
|
| Domain | olp.example.com | ~$10/year (external registrar) |
|
||||||
|
| TLS cert | Let's Encrypt | **Free** |
|
||||||
|
| **Total** | | **~$10/year** (domain only) |
|
||||||
|
|
||||||
|
Oracle Cloud's Always Free ARM VM is overprovisioned for this use case. OLP + Nginx + 3 provider CLIs will use <1GB RAM and negligible CPU (the LLM inference happens at the provider, not here).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Migration Path to Commercial
|
||||||
|
|
||||||
|
This family deployment is a stepping stone. When commercial service is ready:
|
||||||
|
|
||||||
|
| Aspect | Family (this plan) | Commercial (future) |
|
||||||
|
|---|---|---|
|
||||||
|
| Upstream | spawn CLI (subscription) | direct API (commercial key) |
|
||||||
|
| Auth | OLP multi-key (filesystem) | Registration + billing system |
|
||||||
|
| TLS | Let's Encrypt (single domain) | Managed cert (Cloudflare / AWS ACM) |
|
||||||
|
| Compute | Single VM (Oracle Free) | Container cluster (auto-scale) |
|
||||||
|
| Monitoring | Cron + manual | Prometheus + Grafana + PagerDuty |
|
||||||
|
| Rate limit | Nginx per-IP | Per-key token bucket in OLP |
|
||||||
|
| Data | ~/.olp/ filesystem | PostgreSQL + S3 |
|
||||||
|
|
||||||
|
The deployment experience from this plan directly informs the commercial architecture. Every operational runbook becomes a feature requirement for the commercial platform.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Authors:** project maintainer (with AI drafting assistance)
|
||||||
|
**Created:** 2026-05-27
|
||||||
+22
-2
@@ -8,7 +8,7 @@
|
|||||||
3. **Where** does the work live in the tree today (file + anchor).
|
3. **Where** does the work live in the tree today (file + anchor).
|
||||||
4. **When** does it need to land (trigger: load profile, security event, governance amendment).
|
4. **When** does it need to land (trigger: load profile, security event, governance amendment).
|
||||||
|
|
||||||
**Reading order for a v1.x sprint kickoff.** As of 2026-05-25, #1 (streaming SF, D57+D58) and #2 (multi-key auth, Phase 2) are CLOSED, and #4 and #7 closed in D56. Remaining v1.x scope: #3 (soft trigger reactivation), #5 (provider cacheKeyFields mask), #6 (streaming SPAWN_FAILED salvage — unbundled from #1 at #1 close). All three remaining items have explicit "trigger to start" gates that have not fired.
|
**Reading order for a v1.x sprint kickoff.** As of 2026-05-27, #1 (streaming SF, D57+D58), #2 (multi-key auth, Phase 2), #4, #7, and #8 are CLOSED. Remaining v1.x scope: #3 (soft trigger reactivation), #5 (provider cacheKeyFields mask), #6 (streaming SPAWN_FAILED salvage — unbundled from #1 at #1 close). All three remaining items have explicit "trigger to start" gates that have not fired.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -88,8 +88,28 @@
|
|||||||
- **Tracking.** Not a GitHub issue. Tracked here.
|
- **Tracking.** Not a GitHub issue. Tracked here.
|
||||||
- **Trigger to start.** First report of streaming-path SPAWN_FAILED mid-stream where partial-chunk salvage would have helped a downstream caller. Practically unlikely at family scale.
|
- **Trigger to start.** First report of streaming-path SPAWN_FAILED mid-stream where partial-chunk salvage would have helped a downstream caller. Practically unlikely at family scale.
|
||||||
|
|
||||||
## #7 — AUTH_MISSING tuple path test coverage (D40 follow-up)
|
## #8 — Dashboard enrichment: per-provider subscription quota + reset times + 1-min refresh + manual refresh (D78 follow-up) — ✅ **CLOSED (D82, v0.5.0)**
|
||||||
|
|
||||||
|
- **Status.** Closed at D82 (Phase 5). `dashboard.html` restructured to Claude.ai-style per-provider rows rendering `quota_v2`. Closed by PR on branch `d82-dashboard-ui-claude-ai-style`; ships with v0.5.0. 60s quota auto-refresh + manual refresh button + visibilityState guard implemented. Graceful fallback to legacy `quota` field when server runs a pre-D81 build.
|
||||||
|
- **What.** Phase 3 dashboard (D51 `dashboard.html`, v0.3.0) shows: per-provider quota (currently always "n/a — no quota api"), last-24h request count + cache hit + fallback rate, 30d request-count sparkline, top fallback chains. **Maintainer request 2026-05-26 post-D78**: extend to show what each enabled provider's subscription is actually consuming, with reset times visible, refresh once per minute (current 30s is OK but maintainer specified 1min target), and a manual refresh button. Reference design: Claude.ai's own `claude.ai/settings/usage` page — current session bar with "Resets in 1hr 6min", weekly all-models bar with "Resets Sun 9:00 PM", per-model bar (Sonnet only), additional features (routine runs), usage credits + monthly spend limit + auto-reload toggle.
|
||||||
|
- **Why deferred.** v0.3.0/v0.4.x ships the dashboard frame but `provider.quotaStatus()` returns `null` in all three v0.1 plugins (anthropic / openai / mistral). The ratifying spec in ADR 0004 Amendment 2 punts `quotaStatus()` to v1.x ("soft trigger reactivation") — this dashboard ask is the **operator-facing reason** that work would land.
|
||||||
|
- **What this requires.** Per-provider plugin work + dashboard.html UI work + audit-query.mjs aggregation:
|
||||||
|
1. **`lib/providers/anthropic.mjs quotaStatus()`** — discover where the maintainer's Claude.ai subscription quota state is exposed. Candidates: (a) `claude` CLI command (e.g., `claude usage`) if Anthropic adds one — currently absent; (b) parsing the `claude-code` output for rate-limit error messages and caching state from headers; (c) hitting `api.anthropic.com/v1/.../usage` directly via the OAuth refresh token — not a documented endpoint, primary-source risk. ADR 0002 Rule 1 / Rule 5 require an authority citation before any implementation. Likely path: **wait until Anthropic publishes a documented endpoint**, OR derive from audit-side request counts only (no real quota truth, just "you sent N requests in the current 5h window").
|
||||||
|
2. **`lib/providers/openai.mjs quotaStatus()`** — codex CLI doesn't expose ChatGPT-subscription quota state. OpenAI rate-limit headers per request might be parseable but ADR 0004 Amendment 2 explicitly says no plugin parses HTTP status at v0.1.
|
||||||
|
3. **`lib/providers/mistral.mjs quotaStatus()`** — Le Chat Pro has `/v1/usage` endpoint per Mistral docs (verify).
|
||||||
|
4. **`dashboard.html` UI restructure** to a Claude.ai-style layout: rows of (label, bar, "Resets in X" / "Resets at <day-of-week> <time>", percent). Add a manual refresh button + change auto-poll from 30s → 60s. Optionally a usage-credits / per-key spend display if Phase 5 ships per-key cost weights.
|
||||||
|
5. **`lib/audit-query.mjs`** — extend `aggregateRequests` / `spendTrendDaily` to compute "in the current rolling window" (since session/week start) per provider. Today's aggregates are wall-clock windows; subscription resets are per-account-anchored. Need a way to model session windows (e.g., "Anthropic 5h-from-first-request-since-last-reset").
|
||||||
|
- **Reference (maintainer 2026-05-26).** Screenshot of `claude.ai/settings/usage` shared inline. Key panels: Plan usage limits (current session + resets-in), Weekly limits (All models / Sonnet only / per-feature breakdown, each with resets-on), Additional features (Daily included routine runs N / 15), Usage credits (toggle + spent vs monthly limit + auto-reload + buy-credits link).
|
||||||
|
- **Tracking.** Not yet a GitHub issue. Track here + cross-reference ADR 0004 Amendment 2 (soft trigger reactivation — same `quotaStatus()` data-source work) when this becomes Phase 5 scope.
|
||||||
|
- **Code anchors today.**
|
||||||
|
- `dashboard.html` — current 4 panels; needs restructure to Claude.ai-style row layout
|
||||||
|
- `lib/providers/anthropic.mjs` / `openai.mjs` / `mistral.mjs` — `quotaStatus()` returns null today
|
||||||
|
- `lib/audit-query.mjs` — current `aggregateRequests` is wall-clock-window; needs session-window variant
|
||||||
|
- **Trigger to start.** ANY of: (a) Anthropic publishes a documented `claude usage` CLI or `api.anthropic.com/v1/usage` endpoint, (b) maintainer hits real "I want to see quota right now" pain often enough to design without per-provider truth (audit-derived only), (c) Phase 5 multi-tenant adds per-key spend limits and the dashboard needs to surface those.
|
||||||
|
|
||||||
|
## #7 — AUTH_MISSING tuple path test coverage (D40 follow-up) — ✅ **CLOSED (D56, 2026-05-27)**
|
||||||
|
|
||||||
|
- **Status.** Closed. Test shipped at D56 (PR `f4-cli-plugin-quota-v2-plus-auth-missing-test`, 2026-05-27). Test: `test-features.mjs` line 6255 — `'engine: AUTH_MISSING terminates chain, fallbackDetail tuple records trigger_type:"auth_missing" (D56, v1.x roadmap #7)'`. Asserts: `result.fallbackDetail[0].code === 'AUTH_MISSING'`, `result.fallbackDetail[0].trigger_type === 'auth_missing'`, `result.fallbackHops === 0` (no advance). The test was already present in the file before this PR closed the roadmap entry.
|
||||||
- **What.** Dedicated test in `test-features.mjs` Suite D40 that asserts the `fallbackDetail` tuple records the AUTH_MISSING path with `trigger_type: 'auth_missing'`. D40 reviewer flagged this as the last gap in the engine-path matrix; code is structurally correct, just lacks an explicit pin.
|
- **What.** Dedicated test in `test-features.mjs` Suite D40 that asserts the `fallbackDetail` tuple records the AUTH_MISSING path with `trigger_type: 'auth_missing'`. D40 reviewer flagged this as the last gap in the engine-path matrix; code is structurally correct, just lacks an explicit pin.
|
||||||
- **Why deferred.** Low priority — the AUTH_MISSING early-return branch has the tuple push BEFORE it (verified in D40 reviewer pass), so coverage is implicit via the other engine-path tests. A 3-line dedicated test would make the pin explicit.
|
- **Why deferred.** Low priority — the AUTH_MISSING early-return branch has the tuple push BEFORE it (verified in D40 reviewer pass), so coverage is implicit via the other engine-path tests. A 3-line dedicated test would make the pin explicit.
|
||||||
- **Design.** No ADR needed. ~5-line test addition.
|
- **Design.** No ADR needed. ~5-line test addition.
|
||||||
|
|||||||
@@ -443,6 +443,197 @@ export function spendTrendDaily({ days, olpHome, logEvent, _nowFn } = {}) {
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Normalize a single quotaStatus() return value from the anthropic plugin into
|
||||||
|
* the dashboard-friendly shape (D81 / ADR 0008 Amendment).
|
||||||
|
* Provider-specific: called only for 'anthropic'. Returns null if the raw
|
||||||
|
* shape is absent or malformed.
|
||||||
|
*
|
||||||
|
* v0.5.1: handles probe_status field (F3 — ADR 0013 Rule 6).
|
||||||
|
* Accepts both old shape (stale: boolean) and new shape (probe_status: string).
|
||||||
|
*
|
||||||
|
* @internal — used by aggregateProviderQuota()
|
||||||
|
*/
|
||||||
|
function _normalizeAnthropicQuota(raw) {
|
||||||
|
if (!raw || typeof raw !== 'object') return null;
|
||||||
|
const f = raw.fields ?? {};
|
||||||
|
// v0.5.1: probe_status field (new) takes precedence; fall back to stale bool for compat.
|
||||||
|
const probeStatus = raw.probe_status ?? (raw.stale === true ? 'stale' : 'live');
|
||||||
|
return {
|
||||||
|
schema_version: raw.schemaVersion ?? null,
|
||||||
|
last_fresh_at: (probeStatus === 'stale')
|
||||||
|
? (raw.last_fresh_at ?? null)
|
||||||
|
: (raw.probedAt ?? null),
|
||||||
|
utilization: probeStatus === 'unreachable' ? null : {
|
||||||
|
'5h': f.utilization_5h ?? null,
|
||||||
|
'7d': f.utilization_7d ?? null,
|
||||||
|
},
|
||||||
|
reset: probeStatus === 'unreachable' ? null : {
|
||||||
|
'5h': f.reset_5h ?? null,
|
||||||
|
'7d': f.reset_7d ?? null,
|
||||||
|
overall: f.reset ?? null,
|
||||||
|
overage: f.overage_reset ?? null,
|
||||||
|
},
|
||||||
|
representative_claim: f.representative_claim ?? null,
|
||||||
|
fallback_percentage: f.fallback_percentage ?? null,
|
||||||
|
overage: probeStatus === 'unreachable' ? null : {
|
||||||
|
status: f.overage_status ?? null,
|
||||||
|
disabled_reason: f.overage_disabled_reason ?? null,
|
||||||
|
},
|
||||||
|
raw_available: (typeof raw.raw === 'object' && raw.raw !== null),
|
||||||
|
// v0.5.1 (F3 — ADR 0013 Rule 6): failure detail for operator diagnostics
|
||||||
|
failure: raw.failure ?? null,
|
||||||
|
failure_kind: raw.failure?.kind ?? null,
|
||||||
|
backoff_until: raw.failure?.backoff_until ?? null,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Aggregate per-provider quota status into a normalized dashboard-friendly
|
||||||
|
* shape. This is the D81 Phase 5 extension of lib/audit-query.mjs per
|
||||||
|
* ADR 0008 Amendment (D81).
|
||||||
|
*
|
||||||
|
* For each loaded provider, calls quotaStatus() (already cached at the plugin
|
||||||
|
* layer per ADR 0013 Rule 3) and normalizes to a consistent shape. Providers
|
||||||
|
* returning null (codex, mistral) produce a { status: 'unavailable' } row.
|
||||||
|
*
|
||||||
|
* Audit-query stays in-memory scan per ADR 0008 Lane 2 = A. This function
|
||||||
|
* does NOT scan the ndjson files; it calls the live provider plugins.
|
||||||
|
*
|
||||||
|
* Authority: ADR 0008 Amendment (D81) + ADR 0012 D81 + ADR 0013 Rule 5.
|
||||||
|
*
|
||||||
|
* @param {object} args
|
||||||
|
* @param {Map<string, object>} args.providers - Map of provider name → plugin object
|
||||||
|
* @param {(name: string) => Promise<object|null>} [args.getQuotaStatus] - injectable for tests;
|
||||||
|
* defaults to calling providers.get(name).quotaStatus?.()
|
||||||
|
* @returns {Promise<Array<{
|
||||||
|
* provider: string,
|
||||||
|
* status: 'live' | 'stale' | 'unavailable' | 'disabled',
|
||||||
|
* reason?: string,
|
||||||
|
* schema_version: string | null,
|
||||||
|
* last_fresh_at: number | null,
|
||||||
|
* utilization: { '5h': number|null, '7d': number|null } | null,
|
||||||
|
* reset: { '5h': number|null, '7d': number|null, overall: number|null, overage: number|null } | null,
|
||||||
|
* representative_claim: string | null,
|
||||||
|
* fallback_percentage: number | null,
|
||||||
|
* overage: { status: string|null, disabled_reason: string|null } | null,
|
||||||
|
* raw_available: boolean,
|
||||||
|
* }>>}
|
||||||
|
*/
|
||||||
|
export async function aggregateProviderQuota({
|
||||||
|
providers,
|
||||||
|
getQuotaStatus,
|
||||||
|
} = {}) {
|
||||||
|
if (!providers) {
|
||||||
|
throw new Error('aggregateProviderQuota: providers (Map) is required');
|
||||||
|
}
|
||||||
|
|
||||||
|
// Normalize the providers argument — accept both Map and plain object.
|
||||||
|
const providerEntries = (providers instanceof Map)
|
||||||
|
? [...providers.entries()]
|
||||||
|
: Object.entries(providers);
|
||||||
|
|
||||||
|
const results = [];
|
||||||
|
for (const [name, plugin] of providerEntries) {
|
||||||
|
// Default getter: call the plugin's quotaStatus() if present.
|
||||||
|
const fetchQuota = getQuotaStatus
|
||||||
|
? () => getQuotaStatus(name)
|
||||||
|
: () => (typeof plugin?.quotaStatus === 'function' ? plugin.quotaStatus(null) : Promise.resolve(null));
|
||||||
|
|
||||||
|
let rawResult = null;
|
||||||
|
let callError = null;
|
||||||
|
try {
|
||||||
|
rawResult = await fetchQuota();
|
||||||
|
} catch (err) {
|
||||||
|
callError = err?.message ?? String(err);
|
||||||
|
}
|
||||||
|
|
||||||
|
if (callError !== null) {
|
||||||
|
// quotaStatus() threw — treat as error / unavailable.
|
||||||
|
results.push({
|
||||||
|
provider: name,
|
||||||
|
status: 'unavailable',
|
||||||
|
reason: callError,
|
||||||
|
schema_version: null,
|
||||||
|
last_fresh_at: null,
|
||||||
|
utilization: null,
|
||||||
|
reset: null,
|
||||||
|
representative_claim: null,
|
||||||
|
fallback_percentage: null,
|
||||||
|
overage: null,
|
||||||
|
raw_available: false,
|
||||||
|
});
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (rawResult === null || rawResult === undefined) {
|
||||||
|
// Plugin returned null: opt-in disabled (the ONLY case per v0.5.1 contract)
|
||||||
|
// or providers with no quota API at all (codex, mistral).
|
||||||
|
results.push({
|
||||||
|
provider: name,
|
||||||
|
status: 'unavailable',
|
||||||
|
reason: 'no public quota api or probe disabled',
|
||||||
|
schema_version: null,
|
||||||
|
last_fresh_at: null,
|
||||||
|
utilization: null,
|
||||||
|
reset: null,
|
||||||
|
representative_claim: null,
|
||||||
|
fallback_percentage: null,
|
||||||
|
overage: null,
|
||||||
|
raw_available: false,
|
||||||
|
failure: null,
|
||||||
|
failure_kind: null,
|
||||||
|
backoff_until: null,
|
||||||
|
});
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
// quotaStatus() returned a non-null shape — normalize.
|
||||||
|
// v0.5.1: handle probe_status field (live/stale/unreachable).
|
||||||
|
// Currently only 'anthropic' returns a structured shape; other providers
|
||||||
|
// returning structured data will work if their shape is compatible.
|
||||||
|
const probeStatus = rawResult.probe_status ?? (rawResult.stale === true ? 'stale' : 'live');
|
||||||
|
const normalized = _normalizeAnthropicQuota(rawResult);
|
||||||
|
|
||||||
|
if (normalized === null) {
|
||||||
|
// Shape was present but unrecognizable.
|
||||||
|
results.push({
|
||||||
|
provider: name,
|
||||||
|
status: 'unavailable',
|
||||||
|
reason: 'unrecognized quota shape',
|
||||||
|
schema_version: null,
|
||||||
|
last_fresh_at: null,
|
||||||
|
utilization: null,
|
||||||
|
reset: null,
|
||||||
|
representative_claim: null,
|
||||||
|
fallback_percentage: null,
|
||||||
|
overage: null,
|
||||||
|
raw_available: false,
|
||||||
|
failure: null,
|
||||||
|
failure_kind: null,
|
||||||
|
backoff_until: null,
|
||||||
|
});
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Map probe_status to output status:
|
||||||
|
// 'live' → 'live'
|
||||||
|
// 'stale' → 'stale'
|
||||||
|
// 'unreachable' → 'unreachable' (new in v0.5.1; dashboard renders with red border)
|
||||||
|
const outputStatus = probeStatus === 'unreachable' ? 'unreachable'
|
||||||
|
: probeStatus === 'stale' ? 'stale'
|
||||||
|
: 'live';
|
||||||
|
|
||||||
|
results.push({
|
||||||
|
provider: name,
|
||||||
|
status: outputStatus,
|
||||||
|
...normalized,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
return results;
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Audit-derived cache hit rate over the window. Differs from
|
* Audit-derived cache hit rate over the window. Differs from
|
||||||
* `cacheStore.stats()` in server.mjs: that is the live in-process counter;
|
* `cacheStore.stats()` in server.mjs: that is the live in-process counter;
|
||||||
|
|||||||
+33
-2
@@ -144,6 +144,15 @@ export function buildBuiltinChecks(opts = {}) {
|
|||||||
const olpHome = resolveOlpHome(opts);
|
const olpHome = resolveOlpHome(opts);
|
||||||
const configPath = join(olpHome, 'config.json');
|
const configPath = join(olpHome, 'config.json');
|
||||||
const proxyUrl = resolveProxyUrl(opts);
|
const proxyUrl = resolveProxyUrl(opts);
|
||||||
|
// D74 P1-1 fix: server.running and server.version probe /health, which
|
||||||
|
// under default production posture (auth.allow_anonymous: false) requires
|
||||||
|
// an Authorization: Bearer header. Without it, the probe gets 401 and
|
||||||
|
// doctor falsely reports the server down. Caller passes the resolved
|
||||||
|
// bearer token via opts.authHeaders (a `{Authorization: 'Bearer ...'}`
|
||||||
|
// object). Empty headers means "no token configured" — the probe still
|
||||||
|
// fires but a 401 response is treated as "auth misconfigured" rather
|
||||||
|
// than "server down" (see server.running check below).
|
||||||
|
const authHeaders = opts.authHeaders ?? {};
|
||||||
|
|
||||||
const checks = [];
|
const checks = [];
|
||||||
|
|
||||||
@@ -298,7 +307,9 @@ export function buildBuiltinChecks(opts = {}) {
|
|||||||
id: 'server.running',
|
id: 'server.running',
|
||||||
category: 'server',
|
category: 'server',
|
||||||
async run() {
|
async run() {
|
||||||
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000 });
|
// D74 P1-1: pass authHeaders so the probe works under the default
|
||||||
|
// production posture (auth.allow_anonymous: false).
|
||||||
|
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000, headers: authHeaders });
|
||||||
if (!r.ok) {
|
if (!r.ok) {
|
||||||
return {
|
return {
|
||||||
status: 'fail',
|
status: 'fail',
|
||||||
@@ -311,6 +322,25 @@ export function buildBuiltinChecks(opts = {}) {
|
|||||||
},
|
},
|
||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
// 401: server is up but the caller has no/wrong bearer token. NOT a
|
||||||
|
// "server down" condition — distinguish so the kind discriminator
|
||||||
|
// doesn't route to fix_server when the user just needs OLP_API_KEY.
|
||||||
|
if (r.status === 401 || r.status === 403) {
|
||||||
|
return {
|
||||||
|
status: 'fail',
|
||||||
|
message: `${proxyUrl}/health returned ${r.status} — server is up but the bearer token is missing or invalid. Set OLP_API_KEY env to an owner-tier token (npx olp-keys list).`,
|
||||||
|
evidence: {
|
||||||
|
fix_commands: [
|
||||||
|
'echo "set OLP_API_KEY=<your owner token> or OLP_OWNER_TOKEN=<...> then rerun olp doctor"',
|
||||||
|
],
|
||||||
|
human_required: [
|
||||||
|
'Locate an owner-tier OLP API key plaintext (or run `npx olp-keys keygen --owner` to mint a new one — printed ONCE).',
|
||||||
|
'Export it: `export OLP_API_KEY=olp_...`',
|
||||||
|
],
|
||||||
|
reference: 'docs/adr/0007-multi-key-auth.md § 9.1 + README § Environment Variables',
|
||||||
|
},
|
||||||
|
};
|
||||||
|
}
|
||||||
if (r.status !== 200) {
|
if (r.status !== 200) {
|
||||||
return { status: 'fail', message: `${proxyUrl}/health returned status=${r.status}` };
|
return { status: 'fail', message: `${proxyUrl}/health returned status=${r.status}` };
|
||||||
}
|
}
|
||||||
@@ -333,7 +363,8 @@ export function buildBuiltinChecks(opts = {}) {
|
|||||||
} catch {
|
} catch {
|
||||||
return { status: 'warn', message: 'Could not read local package.json — skipping version comparison' };
|
return { status: 'warn', message: 'Could not read local package.json — skipping version comparison' };
|
||||||
}
|
}
|
||||||
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000 });
|
// D74 P1-1: same auth-headers fix as server.running.
|
||||||
|
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000, headers: authHeaders });
|
||||||
if (!r.ok || r.status !== 200) {
|
if (!r.ok || r.status !== 200) {
|
||||||
return { status: 'warn', message: `Could not fetch /health to compare version (${r.error ?? `status ${r.status}`})` };
|
return { status: 'warn', message: `Could not fetch /health to compare version (${r.error ?? `status ${r.status}`})` };
|
||||||
}
|
}
|
||||||
|
|||||||
+17
-2
@@ -27,13 +27,28 @@ export class BadRequestError extends Error {
|
|||||||
// ── Role normalization ────────────────────────────────────────────────────
|
// ── Role normalization ────────────────────────────────────────────────────
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* OpenAI deprecated role='function' in favour of role='tool'.
|
* Normalize entry-surface role names → IR canonical set (system/user/assistant/tool).
|
||||||
* Per ADR 0003, IR supports system/user/assistant/tool.
|
*
|
||||||
|
* Per ADR 0003, IR supports exactly four roles. OpenAI's chat-completions
|
||||||
|
* spec has evolved beyond that, and we keep the IR minimal by normalizing
|
||||||
|
* at the entry boundary instead of bloating IR + every provider plugin.
|
||||||
|
*
|
||||||
|
* Current normalizations:
|
||||||
|
* - `function` → `tool` — deprecated in OpenAI chat API, replaced by tool.
|
||||||
|
* - `developer` → `system` — OpenAI o1/o3+ reasoning models accept a new
|
||||||
|
* "developer" role with similar semantics to "system" (high-priority
|
||||||
|
* instructions from the developer to the model). Providers like Hermes
|
||||||
|
* Agent and Cline default to `developer` for openai-completions calls.
|
||||||
|
* OLP-side anthropic + codex providers don't differentiate developer
|
||||||
|
* from system, so the IR canonicalizes to `system` and downstream
|
||||||
|
* translations remain unchanged.
|
||||||
|
*
|
||||||
* @param {string} role
|
* @param {string} role
|
||||||
* @returns {string}
|
* @returns {string}
|
||||||
*/
|
*/
|
||||||
function normalizeRole(role) {
|
function normalizeRole(role) {
|
||||||
if (role === 'function') return 'tool';
|
if (role === 'function') return 'tool';
|
||||||
|
if (role === 'developer') return 'system';
|
||||||
return role;
|
return role;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
+1110
-90
File diff suppressed because it is too large
Load Diff
+108
-3
@@ -157,6 +157,36 @@ function resolveCodexBin() {
|
|||||||
// D6 assumption A2: auth file is named auth.json (unconfirmed — D7 will pin).
|
// D6 assumption A2: auth file is named auth.json (unconfirmed — D7 will pin).
|
||||||
// D6 assumption A3: access token field is `access_token` or `token` (unconfirmed).
|
// D6 assumption A3: access token field is `access_token` or `token` (unconfirmed).
|
||||||
//
|
//
|
||||||
|
// ── D75 (v0.4.2) F1 — codex CLI v0.133.0 schema pin ─────────────────────────
|
||||||
|
//
|
||||||
|
// Real codex CLI v0.133.0 auth.json (verified empirically on PI231 / Mac mini,
|
||||||
|
// 2026-05-26 E2E session):
|
||||||
|
//
|
||||||
|
// {
|
||||||
|
// "auth_mode": "chatgpt",
|
||||||
|
// "OPENAI_API_KEY": null | "<key>",
|
||||||
|
// "tokens": {
|
||||||
|
// "id_token": "<JWT>",
|
||||||
|
// "access_token": "<opaque-or-JWT>", <-- THIS is the access token
|
||||||
|
// "refresh_token": "<opaque>",
|
||||||
|
// "account_id": "<uuid>"
|
||||||
|
// },
|
||||||
|
// "last_refresh": "<ISO8601>"
|
||||||
|
// }
|
||||||
|
//
|
||||||
|
// D6 assumption A3 originally tried `creds.access_token` at TOP level. Under
|
||||||
|
// codex v0.133.0 this field does not exist at the top level → readAuthArtifact()
|
||||||
|
// returned null even when the user had fully completed `codex login`. OLP then
|
||||||
|
// reported "auth artifact missing" via /health and `olp doctor`, and refused
|
||||||
|
// to spawn codex — false negative blocking the entire openai provider.
|
||||||
|
//
|
||||||
|
// Fix: prepend `creds?.tokens?.access_token` to the precedence chain. Keep all
|
||||||
|
// existing fallbacks unchanged so older codex CLI versions (pre-v0.133) and any
|
||||||
|
// future shape variants still resolve.
|
||||||
|
//
|
||||||
|
// Authority pin: codex CLI v0.133.0 source + on-disk auth.json captured during
|
||||||
|
// PI231 E2E. See D75 commit body for the verification transcript.
|
||||||
|
//
|
||||||
// Returns { accessToken: string } or null (never throws).
|
// Returns { accessToken: string } or null (never throws).
|
||||||
export function readAuthArtifact() {
|
export function readAuthArtifact() {
|
||||||
// 1. Explicit test override — always takes precedence.
|
// 1. Explicit test override — always takes precedence.
|
||||||
@@ -165,7 +195,12 @@ export function readAuthArtifact() {
|
|||||||
try {
|
try {
|
||||||
const raw = readFileSync(authPathOverride, 'utf8');
|
const raw = readFileSync(authPathOverride, 'utf8');
|
||||||
const creds = JSON.parse(raw);
|
const creds = JSON.parse(raw);
|
||||||
const token = creds?.access_token ?? creds?.token ?? creds?.accessToken;
|
// D75 F1: codex CLI v0.133.0 nests the token under `tokens.access_token`.
|
||||||
|
// Preserve top-level fallbacks for backward / forward compat.
|
||||||
|
const token = creds?.tokens?.access_token
|
||||||
|
?? creds?.access_token
|
||||||
|
?? creds?.token
|
||||||
|
?? creds?.accessToken;
|
||||||
if (token && typeof token === 'string') return { accessToken: token };
|
if (token && typeof token === 'string') return { accessToken: token };
|
||||||
} catch { /* fall through */ }
|
} catch { /* fall through */ }
|
||||||
return null; // explicit path set but file missing / malformed
|
return null; // explicit path set but file missing / malformed
|
||||||
@@ -178,8 +213,12 @@ export function readAuthArtifact() {
|
|||||||
try {
|
try {
|
||||||
const raw = readFileSync(authPath, 'utf8');
|
const raw = readFileSync(authPath, 'utf8');
|
||||||
const creds = JSON.parse(raw);
|
const creds = JSON.parse(raw);
|
||||||
// D6 assumption A3: try common OAuth field names in precedence order.
|
// D75 F1: codex CLI v0.133.0 nests the token under `tokens.access_token`.
|
||||||
const token = creds?.access_token ?? creds?.token ?? creds?.accessToken;
|
// Try the nested location FIRST, then fall back to legacy top-level fields.
|
||||||
|
const token = creds?.tokens?.access_token
|
||||||
|
?? creds?.access_token
|
||||||
|
?? creds?.token
|
||||||
|
?? creds?.accessToken;
|
||||||
if (token && typeof token === 'string') return { accessToken: token };
|
if (token && typeof token === 'string') return { accessToken: token };
|
||||||
} catch { /* file missing or malformed */ }
|
} catch { /* file missing or malformed */ }
|
||||||
|
|
||||||
@@ -242,9 +281,30 @@ export function irToCodex(irRequest) {
|
|||||||
// model string (e.g., gpt-5.5, gpt-5.4, gpt-5.3-codex).
|
// model string (e.g., gpt-5.5, gpt-5.4, gpt-5.3-codex).
|
||||||
// PROMPT: "Initial instruction for the task. Use '-' to pipe the prompt
|
// PROMPT: "Initial instruction for the task. Use '-' to pipe the prompt
|
||||||
// from stdin."
|
// from stdin."
|
||||||
|
//
|
||||||
|
// ── D75 (v0.4.2) F2 — codex CLI v0.133.0 trusted-directory sandbox ─────────
|
||||||
|
// codex CLI v0.133.0 added a trusted-directory sandbox: invocations outside
|
||||||
|
// a git repo (or outside any directory explicitly trusted via
|
||||||
|
// `codex config trusted-directories`) refuse with:
|
||||||
|
// "Not inside a trusted directory and --skip-git-repo-check was not specified."
|
||||||
|
// and exit non-zero with zero NDJSON output → OLP surfaces SPAWN_FAILED with
|
||||||
|
// no usable chunks → fallback engine advances to next hop unnecessarily.
|
||||||
|
//
|
||||||
|
// The CWD that OLP spawns from is typically the server install dir (`~/olp/`
|
||||||
|
// on Pi231) which is a git repo on maintainer workstations but is NOT a git
|
||||||
|
// repo on most operator hosts. We bypass the sandbox unconditionally because
|
||||||
|
// OLP is the trusted caller (it is the operator's own server invoking its own
|
||||||
|
// configured Codex subscription via the documented `codex exec` automation
|
||||||
|
// entry point). The trusted-directory sandbox is a foot-gun safeguard for
|
||||||
|
// interactive users; OLP's spawn is non-interactive and pre-authorized.
|
||||||
|
//
|
||||||
|
// Authority: codex CLI v0.133.0 release notes / `codex exec --help` output
|
||||||
|
// documenting `--skip-git-repo-check`. Verified empirically on PI231 E2E
|
||||||
|
// 2026-05-26.
|
||||||
const args = [
|
const args = [
|
||||||
'exec',
|
'exec',
|
||||||
'--json',
|
'--json',
|
||||||
|
'--skip-git-repo-check',
|
||||||
'--model', irRequest.model,
|
'--model', irRequest.model,
|
||||||
];
|
];
|
||||||
|
|
||||||
@@ -291,6 +351,51 @@ export function codexChunkToIR(rawNDJSONLine) {
|
|||||||
|
|
||||||
if (!event || typeof event !== 'object') return null;
|
if (!event || typeof event !== 'object') return null;
|
||||||
|
|
||||||
|
// ── D75 (v0.4.2) F3 — codex CLI v0.133.0 event shape pin ─────────────────
|
||||||
|
// Real codex CLI v0.133.0 NDJSON event stream (verified empirically on PI231
|
||||||
|
// / Mac mini, 2026-05-26 E2E session):
|
||||||
|
// {"type":"thread.started","thread_id":"019e..."}
|
||||||
|
// {"type":"turn.started"}
|
||||||
|
// {"type":"item.started","item":{"id":"item_0","type":"reasoning","text":""}}
|
||||||
|
// {"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"<response>"}}
|
||||||
|
// {"type":"turn.completed","usage":{"input_tokens":..,"output_tokens":..}}
|
||||||
|
//
|
||||||
|
// The D6 defensive parser recognized `content`/`delta`/`text` fields at the
|
||||||
|
// top level and `type === 'stop'`/`done === true`. None of these match
|
||||||
|
// v0.133.0's actual shape → every chunk was silently dropped → response body
|
||||||
|
// had `content: null`. F3 adds three NEW recognizers (item.completed →
|
||||||
|
// agent_message; turn.completed → stop; turn.failed → error) BEFORE the
|
||||||
|
// legacy fallback chain. Legacy recognizers preserved for forward/backward
|
||||||
|
// compat (older codex versions; future shape variants).
|
||||||
|
|
||||||
|
// F3-a: agent_message item completion.
|
||||||
|
// codex v0.133.0 emits assistant text as a single item.completed event whose
|
||||||
|
// item.type is 'agent_message' and item.text carries the full text. There
|
||||||
|
// are no incremental deltas — the entire response arrives in one chunk.
|
||||||
|
if (event.type === 'item.completed'
|
||||||
|
&& event.item?.type === 'agent_message'
|
||||||
|
&& typeof event.item?.text === 'string') {
|
||||||
|
return { type: 'delta', content: event.item.text };
|
||||||
|
}
|
||||||
|
|
||||||
|
// F3-b: turn completion → stop chunk.
|
||||||
|
// codex v0.133.0 emits turn.completed with a usage block when the model
|
||||||
|
// finishes. We map this to IR stop with finish_reason 'stop'.
|
||||||
|
if (event.type === 'turn.completed') {
|
||||||
|
return { type: 'stop', finish_reason: 'stop' };
|
||||||
|
}
|
||||||
|
|
||||||
|
// F3-c: turn failure → error chunk.
|
||||||
|
// codex v0.133.0 emits turn.failed with an embedded error object when the
|
||||||
|
// turn cannot complete. Extract a human-readable message for the IR error.
|
||||||
|
if (event.type === 'turn.failed') {
|
||||||
|
const errMsg = (typeof event.error === 'string')
|
||||||
|
? event.error
|
||||||
|
: (event.error?.message ?? 'codex turn.failed');
|
||||||
|
return { type: 'error', error: errMsg };
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Legacy/fallback recognizers (kept for backward + forward compat) ────
|
||||||
// Error event: type === 'error' or error field present
|
// Error event: type === 'error' or error field present
|
||||||
// A4: defensive — error shape unconfirmed; D7 will pin actual field names
|
// A4: defensive — error shape unconfirmed; D7 will pin actual field names
|
||||||
if (event.type === 'error' || (event.error && typeof event.error === 'string')) {
|
if (event.type === 'error' || (event.error && typeof event.error === 'string')) {
|
||||||
|
|||||||
@@ -0,0 +1,290 @@
|
|||||||
|
/**
|
||||||
|
* lib/sandbox/doctor.mjs — Sandbox availability preflight module (Phase 7 PR-A)
|
||||||
|
*
|
||||||
|
* Authority:
|
||||||
|
* @anthropic-ai/sandbox-runtime v0.0.52
|
||||||
|
* https://github.com/anthropic-experimental/sandbox-runtime
|
||||||
|
*
|
||||||
|
* 2026-05-28 PoC spike on PI231 (arm64 Debian Bookworm): dep install clean,
|
||||||
|
* isSupportedPlatform()=true, blocked on apt deps (bwrap + socat), three PoC
|
||||||
|
* scripts parked at /tmp/sandbox-spike/ on PI231.
|
||||||
|
*
|
||||||
|
* OLP ADR 0014 — Sandbox-Runtime Integration for Multi-Tenant Provider Spawning
|
||||||
|
* OLP ADR 0009 Amendment 1 § Caveats #3 (sandbox is cloud prerequisite)
|
||||||
|
* docs/plans/cloud-deployment-family.md § 5
|
||||||
|
*
|
||||||
|
* Design:
|
||||||
|
* Pure module — no state, no side effects beyond child_process.execFileSync for
|
||||||
|
* `which` probes. Does NOT call SandboxManager.initialize(). Does NOT create
|
||||||
|
* or interact with any real sandbox. Safe to call from /health on every request
|
||||||
|
* (results are memoized process-wide by the caller in server.mjs — see
|
||||||
|
* _sandboxStatusCache there).
|
||||||
|
*
|
||||||
|
* Exports:
|
||||||
|
* checkSandboxAvailability() — returns { available, missing, details }
|
||||||
|
* describeSandboxStatus() — returns { ok, message } human-readable summary
|
||||||
|
*/
|
||||||
|
|
||||||
|
import { execFileSync } from 'node:child_process';
|
||||||
|
import { platform as osPlatform } from 'node:os';
|
||||||
|
|
||||||
|
// ── which probe helper ────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Check if a binary is in PATH by running `which <binary>`.
|
||||||
|
* Returns true if found, false if not found or if `which` is unavailable.
|
||||||
|
* Never throws.
|
||||||
|
* @param {string} binary
|
||||||
|
* @returns {boolean}
|
||||||
|
*/
|
||||||
|
function isInPath(binary) {
|
||||||
|
try {
|
||||||
|
execFileSync('which', [binary], { stdio: 'pipe', timeout: 2000 });
|
||||||
|
return true;
|
||||||
|
} catch {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Platform helper ───────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Map Node's process.platform to the sandbox-runtime platform string.
|
||||||
|
* @returns {'linux'|'macos'|'other'}
|
||||||
|
*/
|
||||||
|
function getPlatformName() {
|
||||||
|
const p = osPlatform();
|
||||||
|
if (p === 'linux') return 'linux';
|
||||||
|
if (p === 'darwin') return 'macos';
|
||||||
|
return 'other';
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Library introspection ─────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Attempt to import @anthropic-ai/sandbox-runtime and call its exported
|
||||||
|
* isSupportedPlatform + checkDependencies. Returns structured findings.
|
||||||
|
* Never throws — all errors become { libError: <message> }.
|
||||||
|
*
|
||||||
|
* @returns {Promise<{
|
||||||
|
* libLoaded: boolean,
|
||||||
|
* libError: string|null,
|
||||||
|
* isSupportedPlatform: boolean,
|
||||||
|
* libDependencyErrors: string[],
|
||||||
|
* libDependencyWarnings: string[],
|
||||||
|
* }>}
|
||||||
|
*/
|
||||||
|
async function probeLibrary() {
|
||||||
|
try {
|
||||||
|
const { SandboxManager } = await import('@anthropic-ai/sandbox-runtime');
|
||||||
|
|
||||||
|
let supportedPlatform = false;
|
||||||
|
try {
|
||||||
|
supportedPlatform = SandboxManager.isSupportedPlatform();
|
||||||
|
} catch (e) {
|
||||||
|
return {
|
||||||
|
libLoaded: true,
|
||||||
|
libError: `isSupportedPlatform() threw: ${e?.message ?? e}`,
|
||||||
|
isSupportedPlatform: false,
|
||||||
|
libDependencyErrors: [],
|
||||||
|
libDependencyWarnings: [],
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
// checkDependencies() requires initialize() to have been called first to
|
||||||
|
// set ripgrep/bwrap/socat config. Since PR-A never calls initialize(), we
|
||||||
|
// call checkDependencies() with an undefined argument — the library falls
|
||||||
|
// back to { command: 'rg' } for ripgrep and PATH lookup for bwrap/socat,
|
||||||
|
// which is exactly what we want for the doctor preflight.
|
||||||
|
let libDependencyErrors = [];
|
||||||
|
let libDependencyWarnings = [];
|
||||||
|
if (supportedPlatform) {
|
||||||
|
try {
|
||||||
|
const depCheck = SandboxManager.checkDependencies(undefined);
|
||||||
|
libDependencyErrors = depCheck?.errors ?? [];
|
||||||
|
libDependencyWarnings = depCheck?.warnings ?? [];
|
||||||
|
} catch (e) {
|
||||||
|
// checkDependencies() can throw before initialize() — not fatal
|
||||||
|
libDependencyErrors = [`checkDependencies() threw: ${e?.message ?? e}`];
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
libLoaded: true,
|
||||||
|
libError: null,
|
||||||
|
isSupportedPlatform: supportedPlatform,
|
||||||
|
libDependencyErrors,
|
||||||
|
libDependencyWarnings,
|
||||||
|
};
|
||||||
|
} catch (e) {
|
||||||
|
return {
|
||||||
|
libLoaded: false,
|
||||||
|
libError: `@anthropic-ai/sandbox-runtime import failed: ${e?.message ?? e}`,
|
||||||
|
isSupportedPlatform: false,
|
||||||
|
libDependencyErrors: [],
|
||||||
|
libDependencyWarnings: [],
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Public API ────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Check sandbox availability (OS deps + library platform support).
|
||||||
|
*
|
||||||
|
* Returns:
|
||||||
|
* {
|
||||||
|
* available: boolean, // true only when all hard deps pass on a supported platform
|
||||||
|
* missing: string[], // friendly names of missing hard deps (e.g. 'bubblewrap', 'socat')
|
||||||
|
* details: {
|
||||||
|
* platform: string, // 'linux'|'macos'|'other'
|
||||||
|
* bwrap: boolean, // which bwrap → found
|
||||||
|
* socat: boolean, // which socat → found
|
||||||
|
* ripgrep: boolean, // which rg → found
|
||||||
|
* isSupportedPlatform: boolean,
|
||||||
|
* libLoaded: boolean,
|
||||||
|
* libError: string|null,
|
||||||
|
* libDependencyErrors: string[],
|
||||||
|
* libDependencyWarnings: string[],
|
||||||
|
* }
|
||||||
|
* }
|
||||||
|
*
|
||||||
|
* The `missing` array uses human-readable package names ('bubblewrap', 'socat',
|
||||||
|
* 'ripgrep') so that install hints are directly actionable.
|
||||||
|
*
|
||||||
|
* Does NOT call SandboxManager.initialize() — pure inspection only.
|
||||||
|
* Does NOT cache — the caller (server.mjs) memoizes the result.
|
||||||
|
*/
|
||||||
|
export async function checkSandboxAvailability() {
|
||||||
|
const platform = getPlatformName();
|
||||||
|
|
||||||
|
// Probe OS-level deps independently of the library (the `which` calls are
|
||||||
|
// cheap and always correct; library's checkDependencies may be less precise
|
||||||
|
// when initialize() hasn't been called).
|
||||||
|
const bwrap = isInPath('bwrap');
|
||||||
|
const socat = isInPath('socat');
|
||||||
|
const ripgrep = isInPath('rg');
|
||||||
|
|
||||||
|
// Library introspection (import + isSupportedPlatform + checkDependencies)
|
||||||
|
const lib = await probeLibrary();
|
||||||
|
|
||||||
|
// Determine what's missing for the doctor report.
|
||||||
|
// Only report OS deps as missing on Linux (where bwrap/socat/rg are required);
|
||||||
|
// macOS uses sandbox-exec which is built-in, so these are not hard requirements.
|
||||||
|
const missing = [];
|
||||||
|
if (platform === 'linux') {
|
||||||
|
if (!bwrap) missing.push('bubblewrap');
|
||||||
|
if (!socat) missing.push('socat');
|
||||||
|
if (!ripgrep) missing.push('ripgrep');
|
||||||
|
}
|
||||||
|
// If the library itself failed to load, that's also a blocker
|
||||||
|
if (!lib.libLoaded) {
|
||||||
|
missing.push('@anthropic-ai/sandbox-runtime (import failed)');
|
||||||
|
}
|
||||||
|
// Library-reported hard dep errors (may overlap with our `which` probes;
|
||||||
|
// deduplicate by treating them as additional evidence rather than re-adding)
|
||||||
|
for (const errMsg of lib.libDependencyErrors) {
|
||||||
|
// Only add if it doesn't overlap with what we already reported
|
||||||
|
const isAlreadyCovered =
|
||||||
|
(errMsg.includes('bwrap') && !bwrap) ||
|
||||||
|
(errMsg.includes('socat') && !socat) ||
|
||||||
|
(errMsg.includes('ripgrep') && !ripgrep) ||
|
||||||
|
(errMsg.includes('Unsupported platform'));
|
||||||
|
if (!isAlreadyCovered && !missing.includes(errMsg)) {
|
||||||
|
missing.push(errMsg);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const available =
|
||||||
|
lib.libLoaded &&
|
||||||
|
lib.isSupportedPlatform &&
|
||||||
|
missing.length === 0;
|
||||||
|
|
||||||
|
return {
|
||||||
|
available,
|
||||||
|
missing,
|
||||||
|
details: {
|
||||||
|
platform,
|
||||||
|
bwrap,
|
||||||
|
socat,
|
||||||
|
ripgrep,
|
||||||
|
isSupportedPlatform: lib.isSupportedPlatform,
|
||||||
|
libLoaded: lib.libLoaded,
|
||||||
|
libError: lib.libError,
|
||||||
|
libDependencyErrors: lib.libDependencyErrors,
|
||||||
|
libDependencyWarnings: lib.libDependencyWarnings,
|
||||||
|
},
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Human-readable sandbox status summary for /health and CLI consumers.
|
||||||
|
*
|
||||||
|
* Returns:
|
||||||
|
* {
|
||||||
|
* ok: boolean, // same as checkSandboxAvailability().available
|
||||||
|
* message: string, // multi-line, includes install hint when deps are missing
|
||||||
|
* }
|
||||||
|
*
|
||||||
|
* Does NOT call SandboxManager.initialize() — pure inspection only.
|
||||||
|
*/
|
||||||
|
export async function describeSandboxStatus() {
|
||||||
|
const result = await checkSandboxAvailability();
|
||||||
|
const { available, missing, details } = result;
|
||||||
|
|
||||||
|
if (available) {
|
||||||
|
return {
|
||||||
|
ok: true,
|
||||||
|
message:
|
||||||
|
`Sandbox available on ${details.platform}` +
|
||||||
|
(details.libDependencyWarnings.length > 0
|
||||||
|
? `. Warnings: ${details.libDependencyWarnings.join('; ')}`
|
||||||
|
: '.'),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
// Build a friendly explanation
|
||||||
|
const lines = [];
|
||||||
|
|
||||||
|
if (!details.libLoaded) {
|
||||||
|
lines.push(`Sandbox library not available: ${details.libError ?? 'import failed'}`);
|
||||||
|
} else if (!details.isSupportedPlatform) {
|
||||||
|
lines.push(
|
||||||
|
`Sandbox dependencies not available: platform '${details.platform}' is not supported by @anthropic-ai/sandbox-runtime v0.0.52.`,
|
||||||
|
);
|
||||||
|
} else {
|
||||||
|
// Platform is supported but OS deps are missing
|
||||||
|
const pkgNames = missing.filter(m => !m.includes('import failed'));
|
||||||
|
if (pkgNames.length > 0) {
|
||||||
|
lines.push(`Sandbox dependencies not available: ${pkgNames.map(m => `${m} not installed`).join(', ')}.`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Install hint (only for Linux; macOS sandbox uses sandbox-exec which is built-in)
|
||||||
|
if (details.platform === 'linux' && (missing.includes('bubblewrap') || missing.includes('socat') || missing.includes('ripgrep'))) {
|
||||||
|
const aptPkgs = [];
|
||||||
|
if (missing.includes('bubblewrap')) aptPkgs.push('bubblewrap');
|
||||||
|
if (missing.includes('socat')) aptPkgs.push('socat');
|
||||||
|
if (missing.includes('ripgrep')) aptPkgs.push('ripgrep');
|
||||||
|
lines.push(
|
||||||
|
`Install on Debian/Ubuntu/Raspbian: sudo apt-get install -y ${aptPkgs.join(' ')}`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
// macOS note (PR-A does not wire macOS sandbox-exec; PR-B will)
|
||||||
|
if (details.platform === 'macos') {
|
||||||
|
lines.push(
|
||||||
|
'macOS: sandbox-exec is built-in, but anthropic provider wrapping lands in PR-B. ' +
|
||||||
|
'macOS sandbox integration is not yet wired in this PR (PR-A). ',
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
if (details.libDependencyWarnings.length > 0) {
|
||||||
|
lines.push(`Warnings: ${details.libDependencyWarnings.join('; ')}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
ok: false,
|
||||||
|
message: lines.join('\n'),
|
||||||
|
};
|
||||||
|
}
|
||||||
@@ -0,0 +1,409 @@
|
|||||||
|
/**
|
||||||
|
* lib/sandbox/manager.mjs — Sandbox manager bootstrap + spawn-wrap (Phase 7 PR-B)
|
||||||
|
*
|
||||||
|
* Authority:
|
||||||
|
* @anthropic-ai/sandbox-runtime v0.0.52
|
||||||
|
* https://github.com/anthropic-experimental/sandbox-runtime
|
||||||
|
* dist/sandbox/sandbox-manager.js — SandboxManager.initialize(), wrapWithSandbox()
|
||||||
|
* dist/sandbox/sandbox-utils.js — getDefaultWritePaths() (used internally)
|
||||||
|
*
|
||||||
|
* 2026-05-28 PR-A spike report on PI231 (arm64 Debian Bookworm):
|
||||||
|
* /tmp/sandbox-spike/spike-anthropic.mjs — wrapWithSandbox call signature,
|
||||||
|
* CLAUDE_CODE_OAUTH_TOKEN env passthrough, shell-mode spawn pattern.
|
||||||
|
* OLP ADR 0014 § Decision (singleton at boot) + § PR-B specific scope
|
||||||
|
* OLP ADR 0009 Amendment 1 § Caveats #3 (sandbox is cloud prerequisite)
|
||||||
|
* cc-mem incident 2026-05-27 § 3 (multi-tenant security gap motivation)
|
||||||
|
* ALIGNMENT.md Rule 1 — provider plugin authority citation
|
||||||
|
*
|
||||||
|
* Design:
|
||||||
|
* One-shot bootstrap at server startup (idempotent). If sandbox not available
|
||||||
|
* (doctor.available=false or SandboxManager.initialize throws), bootstrap is a
|
||||||
|
* no-op and isSandboxActive() returns false → provider falls back to direct spawn
|
||||||
|
* (transparent pass-through).
|
||||||
|
*
|
||||||
|
* Singleton pattern: SandboxManager is a process-wide singleton per library
|
||||||
|
* design (reset() clears ALL state). PR-B initializes once at boot with union
|
||||||
|
* config (Anthropic domains only; codex config follows in PR-C). Per-request
|
||||||
|
* wrapSpawn() calls SandboxManager.wrapWithSandbox() which reads from the
|
||||||
|
* already-initialized config state — no per-request initialize().
|
||||||
|
*
|
||||||
|
* ADR 0014 § Pitfalls #4: SandboxManager.reset() in test teardown must happen
|
||||||
|
* in finally blocks; concurrent in-flight spawns may break if reset fires while
|
||||||
|
* a wrapWithSandbox call is in-flight. OLP's current single-server model (one
|
||||||
|
* process) makes this safe: tests call __resetSandboxManagerForTests() which
|
||||||
|
* also calls SandboxManager.reset() — only safe in test context where no real
|
||||||
|
* spawns are in-flight.
|
||||||
|
*
|
||||||
|
* Exports:
|
||||||
|
* bootstrapSandbox(opts?) — one-shot bootstrap; returns { active, reason?, summary? }
|
||||||
|
* isSandboxActive() — synchronous query
|
||||||
|
* wrapSpawn({ bin, args, env, cwd, allowedDomains })
|
||||||
|
* — wraps spawn args; transparent pass-through when inactive
|
||||||
|
* __resetSandboxManagerForTests() — test seam: reset internal state + SandboxManager
|
||||||
|
*/
|
||||||
|
|
||||||
|
import { createHash } from 'node:crypto';
|
||||||
|
import { mkdirSync } from 'node:fs';
|
||||||
|
import { homedir } from 'node:os';
|
||||||
|
import { join } from 'node:path';
|
||||||
|
import { checkSandboxAvailability } from './doctor.mjs';
|
||||||
|
|
||||||
|
// ── Internal state ────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Whether bootstrapSandbox() has been called (initialized = true means we
|
||||||
|
* ran through bootstrap, not necessarily that sandbox is active).
|
||||||
|
* @type {boolean}
|
||||||
|
*/
|
||||||
|
let _initialized = false;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Whether the SandboxManager was successfully initialized and is ready to wrap.
|
||||||
|
* @type {boolean}
|
||||||
|
*/
|
||||||
|
let _active = false;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The config-at-boot snapshot passed to SandboxManager.initialize().
|
||||||
|
* Null if never initialized or bootstrap failed.
|
||||||
|
* @type {object|null}
|
||||||
|
*/
|
||||||
|
let _initConfig = null;
|
||||||
|
|
||||||
|
// ── Ephemeral workspace root ─────────────────────────────────────────────
|
||||||
|
// Per-request cwd: /tmp/olp-spawn/<uuid>/ — unique per request to prevent
|
||||||
|
// cross-request contamination. Caller (provider) owns cleanup (or trusts tmpfs
|
||||||
|
// lifetime). Created by mkdirSync(recursive:true) inside wrapSpawn().
|
||||||
|
const SPAWN_BASE_DIR = '/tmp/olp-spawn';
|
||||||
|
|
||||||
|
// ── Custom error types ───────────────────────────────────────────────────
|
||||||
|
|
||||||
|
export class SandboxBootstrapError extends Error {
|
||||||
|
constructor(message) {
|
||||||
|
super(message);
|
||||||
|
this.name = 'SandboxBootstrapError';
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export class SandboxWrapError extends Error {
|
||||||
|
constructor(message) {
|
||||||
|
super(message);
|
||||||
|
this.name = 'SandboxWrapError';
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── bootstrapSandbox ──────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* One-shot bootstrap of the sandbox. Idempotent — safe to call multiple times.
|
||||||
|
* If already bootstrapped, returns cached result immediately.
|
||||||
|
*
|
||||||
|
* Steps:
|
||||||
|
* 1. Call checkSandboxAvailability() from doctor module.
|
||||||
|
* 2. If !available → set _active=false, return { active:false, reason }.
|
||||||
|
* 3. If available → build config-at-boot, call SandboxManager.initialize(config).
|
||||||
|
* 4. On init success → _active=true, return { active:true, summary }.
|
||||||
|
* 5. On init failure → log + _active=false + return error (server still starts).
|
||||||
|
*
|
||||||
|
* The network allowedDomains covers the Anthropic provider only (PR-B scope).
|
||||||
|
* Codex domains will be added in PR-C alongside the enableWeakerNestedSandbox flag.
|
||||||
|
*
|
||||||
|
* ADR 0014 § PR-B: denyRead covers ~/.olp, ~/.claude, ~/.ssh, ~/.config, ~/.codex
|
||||||
|
* using absolute literal Linux paths (no globs — see ADR 0014 § Pitfalls #2).
|
||||||
|
* ~/.olp contains keys.json (OLP API keys). ~/.claude contains OAuth credentials.
|
||||||
|
* ~/.ssh and ~/.config contain identity material. ~/.codex contains codex config.
|
||||||
|
*
|
||||||
|
* @param {object} [opts]
|
||||||
|
* @param {boolean} [opts.force=false] — if true, re-run bootstrap even if already initialized
|
||||||
|
* @returns {Promise<{ active: boolean, reason?: string, summary?: string }>}
|
||||||
|
*/
|
||||||
|
export async function bootstrapSandbox(opts = {}) {
|
||||||
|
// Return cached result if already initialized (unless forced)
|
||||||
|
if (_initialized && !opts.force) {
|
||||||
|
return _active
|
||||||
|
? { active: true, summary: _buildSummary() }
|
||||||
|
: { active: false, reason: _initConfig?.failReason ?? 'sandbox not available' };
|
||||||
|
}
|
||||||
|
|
||||||
|
// OLP_SANDBOX_DISABLED env-var gate (2026-05-28 PR-B emergency disable):
|
||||||
|
// Live PI231 evidence showed that even with the exit-null guard, HTTP-path
|
||||||
|
// anthropic spawns produced no claude stdout when wrapped (manual exec of
|
||||||
|
// the SAME wrap script in the same process did produce output — root cause
|
||||||
|
// not yet isolated; likely interaction between SandboxManager in-process
|
||||||
|
// proxy sockets and OLP's request-handler event loop). Until the root cause
|
||||||
|
// is debugged + Suite 44-equivalent E2E tests cover the HTTP path, the
|
||||||
|
// sandbox bootstrap is opt-out via OLP_SANDBOX_DISABLED=1 in the server env.
|
||||||
|
//
|
||||||
|
// Default is sandbox-enabled (no env var = try-and-bootstrap). Sandbox is
|
||||||
|
// skipped only when the operator explicitly disables.
|
||||||
|
//
|
||||||
|
// Future PR-B follow-up: investigate the in-process proxy lifecycle
|
||||||
|
// interaction with OLP's HTTP server event loop; capture diagnostic
|
||||||
|
// transcript; ship Suite 44-equivalent that exercises the full HTTP
|
||||||
|
// request → sandbox spawn → response pipeline.
|
||||||
|
if (process.env.OLP_SANDBOX_DISABLED === '1') {
|
||||||
|
_initialized = true;
|
||||||
|
_active = false;
|
||||||
|
_initConfig = { failReason: 'OLP_SANDBOX_DISABLED=1 — sandbox bootstrap skipped by operator' };
|
||||||
|
return {
|
||||||
|
active: false,
|
||||||
|
reason: 'OLP_SANDBOX_DISABLED=1 — sandbox bootstrap skipped by operator',
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
// Reset state for re-bootstrap
|
||||||
|
_initialized = false;
|
||||||
|
_active = false;
|
||||||
|
_initConfig = null;
|
||||||
|
|
||||||
|
// Step 1: Check OS + library availability
|
||||||
|
let availability;
|
||||||
|
try {
|
||||||
|
availability = await checkSandboxAvailability();
|
||||||
|
} catch (e) {
|
||||||
|
_initialized = true;
|
||||||
|
_active = false;
|
||||||
|
_initConfig = { failReason: `doctor check threw: ${e?.message ?? e}` };
|
||||||
|
return { active: false, reason: _initConfig.failReason };
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!availability.available) {
|
||||||
|
_initialized = true;
|
||||||
|
_active = false;
|
||||||
|
const reason = availability.missing.length > 0
|
||||||
|
? `sandbox deps missing: ${availability.missing.join(', ')}`
|
||||||
|
: `sandbox not available on platform: ${availability.details?.platform}`;
|
||||||
|
_initConfig = { failReason: reason };
|
||||||
|
return { active: false, reason };
|
||||||
|
}
|
||||||
|
|
||||||
|
// Step 2: Build config-at-boot
|
||||||
|
// Network allowedDomains: Anthropic provider API domains (PR-B scope).
|
||||||
|
// - api.anthropic.com: primary Anthropic API endpoint
|
||||||
|
// - statsig.anthropic.com: claude CLI telemetry (verified empirically in spike;
|
||||||
|
// required by claude CLI OAuth token refresh path — removing it causes auth failure)
|
||||||
|
// TODO(PR-C): union in codex/openai provider domains when codex wrap lands.
|
||||||
|
const allowedDomains = [
|
||||||
|
'api.anthropic.com',
|
||||||
|
'statsig.anthropic.com',
|
||||||
|
];
|
||||||
|
|
||||||
|
const home = homedir();
|
||||||
|
|
||||||
|
// denyRead: Absolute literal Linux paths per ADR 0014 § Pitfalls #2.
|
||||||
|
// No ~ or glob — ripgrep glob expansion is not used here to stay safe on
|
||||||
|
// both Linux (bwrap) and macOS (sandbox-exec profile).
|
||||||
|
//
|
||||||
|
// 2026-05-28 PR-B fold-in: ~/.claude is NOT in denyRead. It contains the
|
||||||
|
// spawn's own OAuth credentials — claude CLI must read its own auth file
|
||||||
|
// to function. Denying read here causes "Not logged in" failures even
|
||||||
|
// though the operator has valid credentials present.
|
||||||
|
//
|
||||||
|
// The cross-tenant risk for ~/.claude is mitigated by Phase 6c's
|
||||||
|
// --system-prompt flag (ADR 0009 Amendment 1): the system prompt is
|
||||||
|
// fully replaced, suppressing the default tool descriptions that would
|
||||||
|
// otherwise tell the model it has Read/Bash. Without tool descriptions,
|
||||||
|
// the model is highly unlikely to emit tool_use even under prompt
|
||||||
|
// injection. Sandbox's contribution here is protecting OTHER auth
|
||||||
|
// material (other clients' OLP keys, SSH identity, other providers'
|
||||||
|
// tokens) — files claude CLI does NOT legitimately need.
|
||||||
|
//
|
||||||
|
// If we ever switch to a CLI that requires reading credentials.json
|
||||||
|
// AND also legitimately offers tool execution that surfaces those files
|
||||||
|
// (no known case today), this trade-off needs revisiting.
|
||||||
|
const denyRead = [
|
||||||
|
join(home, '.olp'), // OLP API keys + config — cross-tenant
|
||||||
|
join(home, '.ssh'), // SSH identity material — lateral movement
|
||||||
|
join(home, '.config'), // Generic config dir (may contain tokens)
|
||||||
|
join(home, '.codex'), // Codex config — other-provider auth (PR-C will wrap codex)
|
||||||
|
// NOT denied: ~/.claude — this spawn's own auth, breaks claude CLI if denied
|
||||||
|
];
|
||||||
|
|
||||||
|
// allowWrite: ephemeral spawn workspace only. mkdirSync at bootstrap.
|
||||||
|
// getDefaultWritePaths() adds /dev/stdout, /dev/null etc. internally.
|
||||||
|
try {
|
||||||
|
mkdirSync(SPAWN_BASE_DIR, { recursive: true });
|
||||||
|
} catch (e) {
|
||||||
|
// Non-fatal: if this dir can't be created, wrapSpawn will fail per-request.
|
||||||
|
console.warn(`[sandbox/manager] Warning: could not create ${SPAWN_BASE_DIR}: ${e?.message}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
const config = {
|
||||||
|
network: {
|
||||||
|
allowedDomains,
|
||||||
|
deniedDomains: [],
|
||||||
|
},
|
||||||
|
filesystem: {
|
||||||
|
denyRead,
|
||||||
|
allowWrite: [SPAWN_BASE_DIR, '/tmp'],
|
||||||
|
denyWrite: [],
|
||||||
|
},
|
||||||
|
};
|
||||||
|
|
||||||
|
// Step 3: Initialize SandboxManager
|
||||||
|
let SandboxManager;
|
||||||
|
try {
|
||||||
|
const mod = await import('@anthropic-ai/sandbox-runtime');
|
||||||
|
SandboxManager = mod.SandboxManager;
|
||||||
|
} catch (e) {
|
||||||
|
_initialized = true;
|
||||||
|
_active = false;
|
||||||
|
_initConfig = { failReason: `sandbox-runtime import failed: ${e?.message ?? e}` };
|
||||||
|
return { active: false, reason: _initConfig.failReason };
|
||||||
|
}
|
||||||
|
|
||||||
|
try {
|
||||||
|
// ADR 0014 § Pitfalls #5: initialize() generates MITM CA cert (~100-500ms).
|
||||||
|
// Must happen at boot, not per-request.
|
||||||
|
await SandboxManager.initialize(config);
|
||||||
|
_initialized = true;
|
||||||
|
_active = true;
|
||||||
|
_initConfig = { config, SandboxManager };
|
||||||
|
return { active: true, summary: _buildSummary() };
|
||||||
|
} catch (e) {
|
||||||
|
_initialized = true;
|
||||||
|
_active = false;
|
||||||
|
const reason = `SandboxManager.initialize failed: ${e?.message ?? e}`;
|
||||||
|
_initConfig = { failReason: reason };
|
||||||
|
// Log but DO NOT throw — server still starts in unsandboxed mode.
|
||||||
|
// PR-D will add hard-fail mode via config flag.
|
||||||
|
console.warn(`[sandbox/manager] WARNING: ${reason} — provider spawns will run UNSANDBOXED`);
|
||||||
|
return { active: false, reason };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** @internal — returns summary string for logging */
|
||||||
|
function _buildSummary() {
|
||||||
|
const cfg = _initConfig?.config;
|
||||||
|
if (!cfg) return 'active (no config)';
|
||||||
|
const domains = (cfg.network?.allowedDomains ?? []).join(', ');
|
||||||
|
return `network allowlist=[${domains}], denyRead=[${(cfg.filesystem?.denyRead ?? []).length} paths], allowWrite=[${SPAWN_BASE_DIR}, /tmp]`;
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── isSandboxActive ───────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Synchronous query of bootstrap state.
|
||||||
|
* Returns true only if bootstrapSandbox() completed successfully.
|
||||||
|
* Used by provider plugins to decide spawn path.
|
||||||
|
*
|
||||||
|
* @returns {boolean}
|
||||||
|
*/
|
||||||
|
export function isSandboxActive() {
|
||||||
|
return _active;
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── wrapSpawn ─────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Wrap a spawn command + args for sandbox execution.
|
||||||
|
*
|
||||||
|
* Returns { bin, args, env, cwd, sandboxed: boolean }.
|
||||||
|
* - If sandbox inactive: returns inputs unchanged with sandboxed:false.
|
||||||
|
* - If sandbox active: returns the wrapped shell string as
|
||||||
|
* { bin: '/bin/sh', args: ['-c', wrappedShellString], env, cwd, sandboxed:true }.
|
||||||
|
*
|
||||||
|
* The wrapped command is a shell string from SandboxManager.wrapWithSandbox().
|
||||||
|
* It must be spawned with shell:true OR by invoking /bin/sh -c <string> directly
|
||||||
|
* (the latter is what we do here — avoids relying on the shell that Node picks).
|
||||||
|
*
|
||||||
|
* Per-spawn ephemeral cwd uses a UUID to prevent cross-request contamination.
|
||||||
|
* The caller is responsible for cleanup (or trusts tmpfs lifetime).
|
||||||
|
*
|
||||||
|
* ADR 0014 § PR-B: env vars passed through unchanged so CLAUDE_CODE_OAUTH_TOKEN
|
||||||
|
* (if operator set at OLP boot time) still works inside the sandbox.
|
||||||
|
*
|
||||||
|
* @param {object} params
|
||||||
|
* @param {string} params.bin — original binary (e.g. 'claude')
|
||||||
|
* @param {string[]} params.args — original args
|
||||||
|
* @param {object} params.env — spawn environment (from buildSpawnEnv())
|
||||||
|
* @param {string} [params.cwd] — original cwd (ignored; replaced by ephemeral dir)
|
||||||
|
* @param {string[]} [params.allowedDomains] — per-spawn domain override (passed as customConfig)
|
||||||
|
* @returns {Promise<{ bin: string, args: string[], env: object, cwd: string, sandboxed: boolean }>}
|
||||||
|
*/
|
||||||
|
export async function wrapSpawn({ bin, args, env, cwd: _cwd, allowedDomains }) {
|
||||||
|
// Transparent pass-through when sandbox inactive
|
||||||
|
if (!_active || !_initConfig?.SandboxManager) {
|
||||||
|
return {
|
||||||
|
bin,
|
||||||
|
args: args ?? [],
|
||||||
|
env: env ?? {},
|
||||||
|
cwd: _cwd,
|
||||||
|
sandboxed: false,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
const SandboxManager = _initConfig.SandboxManager;
|
||||||
|
|
||||||
|
// Build the shell command string from bin + args.
|
||||||
|
// Each arg is shell-quoted to handle spaces and special characters.
|
||||||
|
// Authority: spike-anthropic.mjs line 29-31 — same quoting pattern.
|
||||||
|
const quotedArgs = (args ?? []).map(a =>
|
||||||
|
/[\s"'`$\\;&|<>()\[\]{}!#~*?]/.test(a)
|
||||||
|
? `"${a.replace(/\\/g, '\\\\').replace(/"/g, '\\"').replace(/\$/g, '\\$').replace(/`/g, '\\`')}"`
|
||||||
|
: a
|
||||||
|
);
|
||||||
|
const commandString = [bin, ...quotedArgs].join(' ');
|
||||||
|
|
||||||
|
// Per-spawn ephemeral cwd (UUID) — prevents cross-request contamination.
|
||||||
|
// ADR 0014 § PR-B: unique per request.
|
||||||
|
const reqId = createHash('sha256').update(`${Date.now()}-${Math.random()}`).digest('hex').slice(0, 16);
|
||||||
|
const spawnCwd = join(SPAWN_BASE_DIR, reqId);
|
||||||
|
try {
|
||||||
|
mkdirSync(spawnCwd, { recursive: true });
|
||||||
|
} catch (e) {
|
||||||
|
throw new SandboxWrapError(`Failed to create ephemeral spawn dir ${spawnCwd}: ${e?.message ?? e}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Per-spawn customConfig: allow caller to override domains (e.g. different provider).
|
||||||
|
// Default: use the config-at-boot allowedDomains.
|
||||||
|
let customConfig;
|
||||||
|
if (allowedDomains && allowedDomains.length > 0) {
|
||||||
|
customConfig = {
|
||||||
|
network: {
|
||||||
|
allowedDomains,
|
||||||
|
deniedDomains: [],
|
||||||
|
},
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
let wrappedCommand;
|
||||||
|
try {
|
||||||
|
wrappedCommand = await SandboxManager.wrapWithSandbox(commandString, undefined, customConfig);
|
||||||
|
} catch (e) {
|
||||||
|
throw new SandboxWrapError(`SandboxManager.wrapWithSandbox failed: ${e?.message ?? e}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Invoke via /bin/sh -c to avoid spawning a second shell layer.
|
||||||
|
// The wrapped command is already a complete shell invocation (bwrap args or
|
||||||
|
// sandbox-exec profile + the original command inside).
|
||||||
|
return {
|
||||||
|
bin: '/bin/sh',
|
||||||
|
args: ['-c', wrappedCommand],
|
||||||
|
env: env ?? {},
|
||||||
|
cwd: spawnCwd,
|
||||||
|
sandboxed: true,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Test seam ─────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reset internal state so test suite can simulate fresh process.
|
||||||
|
* Also calls SandboxManager.reset() if it was initialized (to clear singleton).
|
||||||
|
*
|
||||||
|
* ADR 0014 § Pitfalls #4: must only be called when no in-flight wrapSpawn calls
|
||||||
|
* are active. Safe in sequential test contexts.
|
||||||
|
*
|
||||||
|
* @returns {Promise<void>}
|
||||||
|
*/
|
||||||
|
export async function __resetSandboxManagerForTests() {
|
||||||
|
if (_active && _initConfig?.SandboxManager) {
|
||||||
|
try {
|
||||||
|
await _initConfig.SandboxManager.reset();
|
||||||
|
} catch { /* ignore — test teardown, best-effort */ }
|
||||||
|
}
|
||||||
|
_initialized = false;
|
||||||
|
_active = false;
|
||||||
|
_initConfig = null;
|
||||||
|
}
|
||||||
@@ -1,6 +1,41 @@
|
|||||||
{
|
{
|
||||||
"version": "0.1.0-bootstrap",
|
"version": "0.1.0-bootstrap",
|
||||||
"comment": "OLP models registry — SPOT for (provider, model) → metadata per CLAUDE.md release_kit overlay. v0.1 founding shipped zero Enabled Providers per ALIGNMENT.md § Provider Inventory. D4 populates providers.anthropic as Candidate; D5 transitions to Enabled pending E2E audit. Schema validated by .github/workflows/alignment.yml; provider keys must match ALIGNMENT.md inventory.",
|
"comment": "OLP models registry — SPOT for (provider, model) → metadata per CLAUDE.md release_kit overlay. v0.1 founding shipped zero Enabled Providers per ALIGNMENT.md § Provider Inventory. D4 populates providers.anthropic as Candidate; D5 transitions to Enabled pending E2E audit. Schema validated by .github/workflows/alignment.yml; provider keys must match ALIGNMENT.md inventory.",
|
||||||
|
"quota_probe": {
|
||||||
|
"schema_version": "2026-05-26",
|
||||||
|
"comment": "D81 — ADR 0013 Rule 5 mandate: schema_version pinned in registry so downstream consumers can detect schema drift. fields_pinned is load-bearing: if Anthropic adds/renames a header, dashboard consumers comparing field-presence against this list can flag 'schema drift detected'. Last verified: 2026-05-26 via live probe against api.anthropic.com (Path B per ADR 0013 Rule 5).",
|
||||||
|
"anthropic": {
|
||||||
|
"status": "live",
|
||||||
|
"source": "anthropic-ratelimit-unified-headers",
|
||||||
|
"endpoint": "https://api.anthropic.com/v1/messages",
|
||||||
|
"fields_pinned": [
|
||||||
|
"status",
|
||||||
|
"representative_claim",
|
||||||
|
"reset",
|
||||||
|
"fallback_percentage",
|
||||||
|
"status_5h",
|
||||||
|
"utilization_5h",
|
||||||
|
"reset_5h",
|
||||||
|
"status_7d",
|
||||||
|
"utilization_7d",
|
||||||
|
"reset_7d",
|
||||||
|
"overage_status",
|
||||||
|
"overage_disabled_reason",
|
||||||
|
"overage_reset"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"openai": {
|
||||||
|
"status": "unavailable",
|
||||||
|
"reason": "no public quota endpoint exposed by the openai/codex CLI; audit-derived spend tracking only at v0.5.0",
|
||||||
|
"re_entry_point": "lib/providers/openai.mjs DL-N (when OpenAI publishes a documented quota endpoint)"
|
||||||
|
},
|
||||||
|
"mistral": {
|
||||||
|
"status": "unavailable",
|
||||||
|
"reason": "no public quota endpoint accessible to Vibe / Le Chat member / La Plateforme API keys per D84 spike 2026-05-26 (https://docs.mistral.ai/api). Mistral Admin API exposes billing/usage but requires org-admin scope (out of scope for OLP family-tier deployment).",
|
||||||
|
"re_entry_point": "lib/providers/mistral.mjs DL-7 (when Mistral publishes a member-key-accessible usage endpoint, or when OLP scope expands to admin-key deployment)",
|
||||||
|
"admin_api_reference": "https://docs.mistral.ai/admin/security-access/admin-api"
|
||||||
|
}
|
||||||
|
},
|
||||||
"bootstrapCreated": 1778630400,
|
"bootstrapCreated": 1778630400,
|
||||||
"bootstrapCreatedComment": "Fallback Unix timestamp for models whose precise release date is unknown. Value = 2026-05-13 (the day before the Anthropic billing-split announcement that triggered OLP). Used by handleModels() in server.mjs when a model entry does not have a model-level 'created' field. Per F12 round-5 cold-audit: OpenAI spec treats 'created' as a stable per-model attribute; synthesizing Date.now() on each request causes spurious updates for clients caching models by 'created'.",
|
"bootstrapCreatedComment": "Fallback Unix timestamp for models whose precise release date is unknown. Value = 2026-05-13 (the day before the Anthropic billing-split announcement that triggered OLP). Used by handleModels() in server.mjs when a model entry does not have a model-level 'created' field. Per F12 round-5 cold-audit: OpenAI spec treats 'created' as a stable per-model attribute; synthesizing Date.now() on each request causes spurious updates for clients caching models by 'created'.",
|
||||||
"providers": {
|
"providers": {
|
||||||
|
|||||||
+102
-7
@@ -138,37 +138,132 @@ export function fmtHealth(body) {
|
|||||||
if (body.uptime_human || body.uptimeHuman) {
|
if (body.uptime_human || body.uptimeHuman) {
|
||||||
out += `Uptime: ${body.uptime_human ?? body.uptimeHuman}\n`;
|
out += `Uptime: ${body.uptime_human ?? body.uptimeHuman}\n`;
|
||||||
}
|
}
|
||||||
|
// D74 P2-4 fix: server.mjs /health full payload is
|
||||||
|
// body.providers = { enabled: N, available: N, status: { <name>: {...} } }
|
||||||
|
// The plugin previously iterated Object.entries(body.providers), which
|
||||||
|
// surfaced `enabled`, `available`, and `status` as pseudo-providers
|
||||||
|
// (typeof status === 'object' → loop body fired with name='status').
|
||||||
|
// Walk providers.status when present; fall back to providers.* for the
|
||||||
|
// older OCP shape that lacks the .status wrapper.
|
||||||
if (body.providers && typeof body.providers === "object") {
|
if (body.providers && typeof body.providers === "object") {
|
||||||
|
const enabled = body.providers.enabled;
|
||||||
|
const available = body.providers.available;
|
||||||
|
if (typeof enabled === "number" || typeof available === "number") {
|
||||||
|
out += `Providers: ${enabled ?? "?"} enabled / ${available ?? "?"} available\n`;
|
||||||
|
}
|
||||||
|
const statusMap = body.providers.status && typeof body.providers.status === "object"
|
||||||
|
? body.providers.status
|
||||||
|
: body.providers;
|
||||||
|
const entries = Object.entries(statusMap).filter(
|
||||||
|
([name, s]) => typeof s === "object" && s !== null && name !== "enabled" && name !== "available" && name !== "status"
|
||||||
|
);
|
||||||
|
if (entries.length > 0) {
|
||||||
out += `\nProviders:\n`;
|
out += `\nProviders:\n`;
|
||||||
for (const [name, s] of Object.entries(body.providers)) {
|
for (const [name, s] of entries) {
|
||||||
if (typeof s !== "object" || s === null) continue;
|
|
||||||
const i = statusIcon(s?.ok ? "ok" : "fail");
|
const i = statusIcon(s?.ok ? "ok" : "fail");
|
||||||
out += ` ${i} ${name}\n`;
|
const spawn = typeof s?.activeSpawns === "number" ? ` spawns=${s.activeSpawns}` : "";
|
||||||
|
out += ` ${i} ${name}${spawn}\n`;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return out;
|
return out;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* formatResetCountdown(epochSeconds) → human-readable reset countdown.
|
||||||
|
*
|
||||||
|
* Mirrors bin/olp.mjs + dashboard.html versions. Five ranges:
|
||||||
|
* past / < 1h / < 24h / < 7d / ≥ 7d
|
||||||
|
*
|
||||||
|
* Authority: ADR 0008 Amendment 2 (quota_v2 shape), ported from dashboard.html (D82).
|
||||||
|
* No external deps. Duplicated here intentionally (olp-plugin ships separately).
|
||||||
|
*/
|
||||||
|
export function pluginFormatResetCountdown(epochSeconds) {
|
||||||
|
if (epochSeconds == null) return "—";
|
||||||
|
const nowMs = Date.now();
|
||||||
|
const targetMs = epochSeconds * 1000;
|
||||||
|
const diffMs = targetMs - nowMs;
|
||||||
|
if (diffMs <= 0) return "resetting now";
|
||||||
|
const diffMin = Math.floor(diffMs / 60000);
|
||||||
|
const diffHr = Math.floor(diffMin / 60);
|
||||||
|
const diffDay = Math.floor(diffHr / 24);
|
||||||
|
if (diffMin < 60) return `resets in ${diffMin}m`;
|
||||||
|
if (diffHr < 24) {
|
||||||
|
const remMin = diffMin - diffHr * 60;
|
||||||
|
if (remMin === 0) return `resets in ${diffHr}h`;
|
||||||
|
return `resets in ${diffHr}h ${remMin}m`;
|
||||||
|
}
|
||||||
|
const target = new Date(targetMs);
|
||||||
|
const timeStr = target.toLocaleString("en-US", { hour: "numeric", minute: "2-digit", hour12: true });
|
||||||
|
if (diffDay < 7) {
|
||||||
|
const dayStr = target.toLocaleString("en-US", { weekday: "short" });
|
||||||
|
return `resets ${dayStr} ${timeStr}`;
|
||||||
|
}
|
||||||
|
const dateStr = target.toLocaleString("en-US", { month: "short", day: "numeric" });
|
||||||
|
return `resets ${dateStr} ${timeStr}`;
|
||||||
|
}
|
||||||
|
|
||||||
export function fmtUsage(body) {
|
export function fmtUsage(body) {
|
||||||
let out = "OLP usage (24h)\n";
|
let out = "OLP usage (24h)\n";
|
||||||
out += "─────────────────────────────\n";
|
out += "─────────────────────────────\n";
|
||||||
const w = body.window_24h ?? body.usage_24h ?? {};
|
const w = body.window_24h ?? body.usage_24h ?? {};
|
||||||
if (w.requests !== undefined) {
|
if (w.request_count !== undefined) {
|
||||||
|
out += `Requests: ${w.request_count}\n`;
|
||||||
|
const c = body.cache_hit_24h ?? {};
|
||||||
|
if (typeof c.hit_rate === "number") {
|
||||||
|
out += `Cache hit: ${(c.hit_rate * 100).toFixed(1)}%\n`;
|
||||||
|
}
|
||||||
|
} else if (w.requests !== undefined) {
|
||||||
out += `Requests: ${w.requests}\n`;
|
out += `Requests: ${w.requests}\n`;
|
||||||
out += `Cache hit: ${w.cache_hit_rate != null ? `${(w.cache_hit_rate * 100).toFixed(1)}%` : "?"}\n`;
|
out += `Cache hit: ${w.cache_hit_rate != null ? `${(w.cache_hit_rate * 100).toFixed(1)}%` : "?"}\n`;
|
||||||
out += `Fallbacks: ${w.fallbacks ?? "?"}\n`;
|
out += `Fallbacks: ${w.fallbacks ?? "?"}\n`;
|
||||||
} else if (typeof body.cache_hit_24h === "number") {
|
} else if (typeof body.cache_hit_24h === "number") {
|
||||||
// Dashboard-data shape: cache_hit_24h is a rate ∈ [0,1]
|
// Legacy: cache_hit_24h as a bare number
|
||||||
out += `Cache hit (24h): ${(body.cache_hit_24h * 100).toFixed(1)}%\n`;
|
out += `Cache hit (24h): ${(body.cache_hit_24h * 100).toFixed(1)}%\n`;
|
||||||
}
|
}
|
||||||
if (Array.isArray(body.quota) && body.quota.length > 0) {
|
|
||||||
|
// F4: prefer quota_v2 when present (server v0.5.0+), fall back to legacy quota.
|
||||||
|
// Authority: ADR 0008 Amendment 2 (quota_v2 shape).
|
||||||
|
if (Array.isArray(body.quota_v2) && body.quota_v2.length > 0) {
|
||||||
|
out += `\nPer-provider quota (live):\n`;
|
||||||
|
for (const p of body.quota_v2) {
|
||||||
|
const name = String(p.provider ?? "?").toUpperCase().padEnd(10);
|
||||||
|
const status = p.status ?? "unavailable";
|
||||||
|
if (status === "unavailable") {
|
||||||
|
out += ` ${name} unavailable ${p.reason ?? "no public quota api"}\n`;
|
||||||
|
} else if (status === "unreachable") {
|
||||||
|
const fk = p.failure?.kind ?? "unknown";
|
||||||
|
out += ` ${name} no cached data — failure: ${fk}\n`;
|
||||||
|
} else {
|
||||||
|
// live or stale
|
||||||
|
const util = p.utilization ?? {};
|
||||||
|
const reset = p.reset ?? {};
|
||||||
|
const parts = [];
|
||||||
|
for (const window of ["5h", "7d"]) {
|
||||||
|
const frac = util[window];
|
||||||
|
const resetEpoch = reset[window];
|
||||||
|
if (frac != null) {
|
||||||
|
const pct = `${Math.round(frac * 100)}%`;
|
||||||
|
const rst = pluginFormatResetCountdown(resetEpoch);
|
||||||
|
parts.push(`${window}: ${pct} (${rst})`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const staleNote = status === "stale"
|
||||||
|
? ` ⚠ stale (${p.failure?.kind ?? "unknown"})`
|
||||||
|
: "";
|
||||||
|
out += ` ${name} ${status.padEnd(6)} ${parts.join(" ")}${staleNote}\n`;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
} else if (Array.isArray(body.quota) && body.quota.length > 0) {
|
||||||
|
// Legacy fallback for pre-v0.5.0 servers
|
||||||
out += `\nPer-provider quota:\n`;
|
out += `\nPer-provider quota:\n`;
|
||||||
for (const q of body.quota) {
|
for (const q of body.quota) {
|
||||||
const pct = typeof q.percent_used === "number" ? q.percent_used : null;
|
const pct = typeof q.percent_used === "number" ? q.percent_used : null;
|
||||||
const bar0 = pct != null ? ` ${bar(pct / 100, 12)} ${pct.toFixed(0)}%` : " no quota api";
|
const bar0 = pct != null ? ` ${bar(pct / 100, 12)} ${pct.toFixed(0)}%` : " no quota api";
|
||||||
out += ` ${String(q.name ?? "?").padEnd(10)}${bar0}\n`;
|
out += ` ${String(q.provider ?? q.name ?? "?").padEnd(10)}${bar0}\n`;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
if (Array.isArray(body.top_fallback_chains_24h) && body.top_fallback_chains_24h.length > 0) {
|
if (Array.isArray(body.top_fallback_chains_24h) && body.top_fallback_chains_24h.length > 0) {
|
||||||
out += `\nTop fallback chains (24h):\n`;
|
out += `\nTop fallback chains (24h):\n`;
|
||||||
for (const f of body.top_fallback_chains_24h.slice(0, 5)) {
|
for (const f of body.top_fallback_chains_24h.slice(0, 5)) {
|
||||||
|
|||||||
@@ -10,6 +10,7 @@
|
|||||||
"openclaw": {
|
"openclaw": {
|
||||||
"type": "plugin",
|
"type": "plugin",
|
||||||
"id": "olp",
|
"id": "olp",
|
||||||
"pluginManifest": "openclaw.plugin.json"
|
"pluginManifest": "openclaw.plugin.json",
|
||||||
|
"extensions": ["./index.js"]
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
Generated
+89
@@ -0,0 +1,89 @@
|
|||||||
|
{
|
||||||
|
"name": "olp",
|
||||||
|
"version": "0.5.1",
|
||||||
|
"lockfileVersion": 3,
|
||||||
|
"requires": true,
|
||||||
|
"packages": {
|
||||||
|
"": {
|
||||||
|
"name": "olp",
|
||||||
|
"version": "0.5.1",
|
||||||
|
"license": "MIT",
|
||||||
|
"dependencies": {
|
||||||
|
"@anthropic-ai/sandbox-runtime": "^0.0.52"
|
||||||
|
},
|
||||||
|
"bin": {
|
||||||
|
"olp": "bin/olp.mjs",
|
||||||
|
"olp-audit-rotate": "bin/olp-audit-rotate.mjs",
|
||||||
|
"olp-connect": "bin/olp-connect",
|
||||||
|
"olp-keys": "bin/olp-keys.mjs"
|
||||||
|
},
|
||||||
|
"engines": {
|
||||||
|
"node": ">=18"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"node_modules/@anthropic-ai/sandbox-runtime": {
|
||||||
|
"version": "0.0.52",
|
||||||
|
"resolved": "https://registry.npmjs.org/@anthropic-ai/sandbox-runtime/-/sandbox-runtime-0.0.52.tgz",
|
||||||
|
"integrity": "sha512-vYaM7OslFmOAzNgfy5gxvt3NoWFeCbr7C0AKyuduQq7Gdxbg2NnYmE7deBf8Nxj3ZNECTcC5RhAfz0lZwvbtBA==",
|
||||||
|
"license": "Apache-2.0",
|
||||||
|
"dependencies": {
|
||||||
|
"@pondwader/socks5-server": "^1.0.10",
|
||||||
|
"commander": "^12.1.0",
|
||||||
|
"node-forge": "^1.4.0",
|
||||||
|
"shell-quote": "^1.8.3",
|
||||||
|
"zod": "^3.24.1"
|
||||||
|
},
|
||||||
|
"bin": {
|
||||||
|
"srt": "dist/cli.js"
|
||||||
|
},
|
||||||
|
"engines": {
|
||||||
|
"node": ">=18.0.0"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"node_modules/@pondwader/socks5-server": {
|
||||||
|
"version": "1.0.10",
|
||||||
|
"resolved": "https://registry.npmjs.org/@pondwader/socks5-server/-/socks5-server-1.0.10.tgz",
|
||||||
|
"integrity": "sha512-bQY06wzzR8D2+vVCUoBsr5QS2U6UgPUQRmErNwtsuI6vLcyRKkafjkr3KxbtGFf9aBBIV2mcvlsKD1UYaIV+sg==",
|
||||||
|
"license": "MIT"
|
||||||
|
},
|
||||||
|
"node_modules/commander": {
|
||||||
|
"version": "12.1.0",
|
||||||
|
"resolved": "https://registry.npmjs.org/commander/-/commander-12.1.0.tgz",
|
||||||
|
"integrity": "sha512-Vw8qHK3bZM9y/P10u3Vib8o/DdkvA2OtPtZvD871QKjy74Wj1WSKFILMPRPSdUSx5RFK1arlJzEtA4PkFgnbuA==",
|
||||||
|
"license": "MIT",
|
||||||
|
"engines": {
|
||||||
|
"node": ">=18"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"node_modules/node-forge": {
|
||||||
|
"version": "1.4.0",
|
||||||
|
"resolved": "https://registry.npmjs.org/node-forge/-/node-forge-1.4.0.tgz",
|
||||||
|
"integrity": "sha512-LarFH0+6VfriEhqMMcLX2F7SwSXeWwnEAJEsYm5QKWchiVYVvJyV9v7UDvUv+w5HO23ZpQTXDv/GxdDdMyOuoQ==",
|
||||||
|
"license": "(BSD-3-Clause OR GPL-2.0)",
|
||||||
|
"engines": {
|
||||||
|
"node": ">= 6.13.0"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"node_modules/shell-quote": {
|
||||||
|
"version": "1.8.4",
|
||||||
|
"resolved": "https://registry.npmjs.org/shell-quote/-/shell-quote-1.8.4.tgz",
|
||||||
|
"integrity": "sha512-VsC6n6vz1ihYYyZZwX7YZSF5l5x36ca17OC+a69h94YqB7X6XLwf+5MOgynYir2SLFUbl8gIYvBo8K8RoNQ6bQ==",
|
||||||
|
"license": "MIT",
|
||||||
|
"engines": {
|
||||||
|
"node": ">= 0.4"
|
||||||
|
},
|
||||||
|
"funding": {
|
||||||
|
"url": "https://github.com/sponsors/ljharb"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"node_modules/zod": {
|
||||||
|
"version": "3.25.76",
|
||||||
|
"resolved": "https://registry.npmjs.org/zod/-/zod-3.25.76.tgz",
|
||||||
|
"integrity": "sha512-gzUt/qt81nXsFGKIFcC3YnfEAx5NkunCfnDlvuBSSFS02bcXu4Lmea0AFIUwbLWxWPx3d9p8S5QoaujKcNQxcQ==",
|
||||||
|
"license": "MIT",
|
||||||
|
"funding": {
|
||||||
|
"url": "https://github.com/sponsors/colinhacks"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
+5
-2
@@ -1,6 +1,6 @@
|
|||||||
{
|
{
|
||||||
"name": "olp",
|
"name": "olp",
|
||||||
"version": "0.4.0",
|
"version": "0.5.1",
|
||||||
"description": "Personal multi-provider LLM proxy. Successor to OCP. One HTTP endpoint, multiple subscriptions behind it, automatic routing + fallback + caching.",
|
"description": "Personal multi-provider LLM proxy. Successor to OCP. One HTTP endpoint, multiple subscriptions behind it, automatic routing + fallback + caching.",
|
||||||
"type": "module",
|
"type": "module",
|
||||||
"main": "server.mjs",
|
"main": "server.mjs",
|
||||||
@@ -49,5 +49,8 @@
|
|||||||
"mistral",
|
"mistral",
|
||||||
"fallback",
|
"fallback",
|
||||||
"cache"
|
"cache"
|
||||||
]
|
],
|
||||||
|
"dependencies": {
|
||||||
|
"@anthropic-ai/sandbox-runtime": "^0.0.52"
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
+175
-12
@@ -18,6 +18,14 @@
|
|||||||
* so OLP can co-host with OCP for migration windows. ADR 0010 §
|
* so OLP can co-host with OCP for migration windows. ADR 0010 §
|
||||||
* Default port. Set OLP_PORT=3456 explicitly to restore the
|
* Default port. Set OLP_PORT=3456 explicitly to restore the
|
||||||
* pre-D60 default when not co-hosting with OCP.)
|
* pre-D60 default when not co-hosting with OCP.)
|
||||||
|
* OLP_BIND — listen address (default: 127.0.0.1 since D76 / v0.4.3). Set
|
||||||
|
* to 0.0.0.0 (or a specific interface IP) to accept connections
|
||||||
|
* from LAN clients (required for olp-connect <ip> to actually
|
||||||
|
* reach the server). Server emits a startup warn if BIND
|
||||||
|
* resolves to a non-loopback address AND auth.allow_anonymous
|
||||||
|
* is true (anonymous-key over LAN may be acceptable; anonymous-
|
||||||
|
* key over public internet is not — see ADR 0011 § Deployment
|
||||||
|
* configurations).
|
||||||
*/
|
*/
|
||||||
|
|
||||||
import { createServer } from 'node:http';
|
import { createServer } from 'node:http';
|
||||||
@@ -62,12 +70,24 @@ import {
|
|||||||
ENV_OWNER_KEY_ID,
|
ENV_OWNER_KEY_ID,
|
||||||
} from './lib/keys.mjs';
|
} from './lib/keys.mjs';
|
||||||
import { appendAuditEvent } from './lib/audit.mjs';
|
import { appendAuditEvent } from './lib/audit.mjs';
|
||||||
|
// Phase 7 / PR-A — sandbox availability preflight module (ADR 0014).
|
||||||
|
// checkSandboxAvailability is called lazily at first /health hit and memoized
|
||||||
|
// process-wide (bwrap/socat install state does not change at runtime; we don't
|
||||||
|
// want a child_process.execFileSync per /health call).
|
||||||
|
import { checkSandboxAvailability } from './lib/sandbox/doctor.mjs';
|
||||||
|
// Phase 7 / PR-B — sandbox manager bootstrap + spawn-wrap (ADR 0014 § PR-B).
|
||||||
|
// bootstrapSandbox() is called at server startup (before listen) and sets up
|
||||||
|
// the process-wide SandboxManager singleton. isSandboxActive() is used by
|
||||||
|
// /health to report sandbox.active.
|
||||||
|
import { bootstrapSandbox, isSandboxActive, __resetSandboxManagerForTests } from './lib/sandbox/manager.mjs';
|
||||||
// Phase 3 / D50 — management endpoints consume the audit aggregate query layer.
|
// Phase 3 / D50 — management endpoints consume the audit aggregate query layer.
|
||||||
|
// D81 (Phase 5) — adds aggregateProviderQuota for quota_v2 shape.
|
||||||
import {
|
import {
|
||||||
aggregateRequests as auditAggregateRequests,
|
aggregateRequests as auditAggregateRequests,
|
||||||
topFallbackChains as auditTopFallbackChains,
|
topFallbackChains as auditTopFallbackChains,
|
||||||
spendTrendDaily as auditSpendTrendDaily,
|
spendTrendDaily as auditSpendTrendDaily,
|
||||||
cacheHitRateWindow as auditCacheHitRateWindow,
|
cacheHitRateWindow as auditCacheHitRateWindow,
|
||||||
|
aggregateProviderQuota as auditAggregateProviderQuota,
|
||||||
} from './lib/audit-query.mjs';
|
} from './lib/audit-query.mjs';
|
||||||
|
|
||||||
// ── Config ────────────────────────────────────────────────────────────────
|
// ── Config ────────────────────────────────────────────────────────────────
|
||||||
@@ -77,6 +97,14 @@ const pkg = JSON.parse(readFileSync(join(__dirname, 'package.json'), 'utf8'));
|
|||||||
const VERSION = pkg.version;
|
const VERSION = pkg.version;
|
||||||
|
|
||||||
const PORT = parseInt(process.env.OLP_PORT ?? '4567', 10);
|
const PORT = parseInt(process.env.OLP_PORT ?? '4567', 10);
|
||||||
|
// F5 / D76: OLP_BIND env. Defaults to 127.0.0.1 (loopback only — secure
|
||||||
|
// default). Operators expose LAN by setting OLP_BIND=0.0.0.0 (or a specific
|
||||||
|
// interface). Per ADR 0011 § Deployment configurations:
|
||||||
|
// - 127.0.0.1: trusted single-machine; safe with any auth posture
|
||||||
|
// - RFC1918 / tailnet / specific LAN IP: trusted-LAN — anonymous_key OK
|
||||||
|
// - 0.0.0.0: ALL interfaces — operator MUST ensure auth posture matches the
|
||||||
|
// network reachability (e.g. no advertise_anonymous_key on public IP)
|
||||||
|
const BIND = process.env.OLP_BIND ?? '127.0.0.1';
|
||||||
const BODY_LIMIT = 5 * 1024 * 1024; // 5 MB
|
const BODY_LIMIT = 5 * 1024 * 1024; // 5 MB
|
||||||
|
|
||||||
// ── Logging ───────────────────────────────────────────────────────────────
|
// ── Logging ───────────────────────────────────────────────────────────────
|
||||||
@@ -203,6 +231,20 @@ export function __resetRequestCounters() {
|
|||||||
_activeRequests = 0;
|
_activeRequests = 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── Phase 7 PR-A: sandbox availability cache ──────────────────────────────
|
||||||
|
// checkSandboxAvailability() forks `which bwrap` / `which socat` / `which rg`
|
||||||
|
// and imports @anthropic-ai/sandbox-runtime. Neither can change at runtime —
|
||||||
|
// bwrap is either installed or it isn't. Memoize the first result to avoid
|
||||||
|
// repeated child_process.execFileSync calls on every /health hit.
|
||||||
|
//
|
||||||
|
// _sandboxStatusCache: null → not yet fetched
|
||||||
|
// object → memoized result from checkSandboxAvailability()
|
||||||
|
let _sandboxStatusCache = null;
|
||||||
|
/** @internal — test seam: reset sandbox cache between tests. */
|
||||||
|
export function __resetSandboxStatusCache() {
|
||||||
|
_sandboxStatusCache = null;
|
||||||
|
}
|
||||||
|
|
||||||
// ── Startup config ────────────────────────────────────────────────────────
|
// ── Startup config ────────────────────────────────────────────────────────
|
||||||
// Read ~/.olp/config.json once at startup. Provides:
|
// Read ~/.olp/config.json once at startup. Provides:
|
||||||
// - providers.enabled → which providers are loaded (ADR 0002 § Disable model)
|
// - providers.enabled → which providers are loaded (ADR 0002 § Disable model)
|
||||||
@@ -253,6 +295,19 @@ if (_authConfig.advertise_anonymous_key === true) {
|
|||||||
message: 'auth.advertise_anonymous_key=true but no active key with plaintext_advertise exists. Run `olp-keys keygen --anonymous --advertise` to create one. /health.anonymousKey will NOT be emitted until then. See ADR 0011.',
|
message: 'auth.advertise_anonymous_key=true but no active key with plaintext_advertise exists. Run `olp-keys keygen --anonymous --advertise` to create one. /health.anonymousKey will NOT be emitted until then. See ADR 0011.',
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
// F5 / D76 (ADR 0011 Deployment configurations): publishing the anonymous
|
||||||
|
// key via /health is only safe when /health is reachable ONLY from a
|
||||||
|
// trusted network. If OLP_BIND is set to a non-loopback address AND
|
||||||
|
// advertise_anonymous_key is true, warn the operator. We can't tell from
|
||||||
|
// here whether the non-loopback bind is "trusted LAN" (RFC1918 / tailnet)
|
||||||
|
// or "public internet" — that's the operator's responsibility. The warn is
|
||||||
|
// a checkpoint, not a hard gate.
|
||||||
|
if (BIND !== '127.0.0.1' && BIND !== 'localhost' && BIND !== '::1') {
|
||||||
|
logEvent('warn', 'anonymous_key_advertised_with_lan_bind', {
|
||||||
|
message: `auth.advertise_anonymous_key=true with OLP_BIND=${BIND} — /health.anonymousKey will be reachable from any host that can connect to ${BIND}:${PORT}. Confirm this address is on a trusted LAN (RFC1918 / tailnet) — never a public IP. See ADR 0011 § Deployment configurations.`,
|
||||||
|
bind: BIND,
|
||||||
|
});
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** @internal — test seam: inject a synthetic auth config (no file I/O). */
|
/** @internal — test seam: inject a synthetic auth config (no file I/O). */
|
||||||
@@ -828,10 +883,53 @@ async function handleHealth(req, res) {
|
|||||||
providerStatuses[name] = { ok: false, error: e.message, activeSpawns };
|
providerStatuses[name] = { ok: false, error: e.message, activeSpawns };
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
// Phase 7 PR-A (ADR 0014): sandbox availability field.
|
||||||
|
// Result is memoized process-wide in _sandboxStatusCache — bwrap/socat
|
||||||
|
// install state does not change at runtime. If the library call throws for
|
||||||
|
// any reason, the field is still included with available: false + error
|
||||||
|
// (don't crash /health).
|
||||||
|
if (_sandboxStatusCache === null) {
|
||||||
|
try {
|
||||||
|
_sandboxStatusCache = await checkSandboxAvailability();
|
||||||
|
} catch (e) {
|
||||||
|
_sandboxStatusCache = {
|
||||||
|
available: false,
|
||||||
|
missing: [],
|
||||||
|
details: { platform: process.platform, error: String(e?.message ?? e) },
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const sandboxField = {
|
||||||
|
available: _sandboxStatusCache.available,
|
||||||
|
// Phase 7 PR-B: active = sandbox was bootstrapped and SandboxManager is
|
||||||
|
// ready to wrap spawns. available=true + active=true means every provider
|
||||||
|
// spawn is actually sandboxed. available=true + active=false means deps
|
||||||
|
// present but bootstrap failed at runtime (see server startup log).
|
||||||
|
active: isSandboxActive(),
|
||||||
|
missing: _sandboxStatusCache.missing ?? [],
|
||||||
|
platform: _sandboxStatusCache.details?.platform ?? process.platform,
|
||||||
|
};
|
||||||
|
if (!_sandboxStatusCache.available) {
|
||||||
|
// Include human-readable install hint for owner-tier callers.
|
||||||
|
const missingDeps = (_sandboxStatusCache.missing ?? []).filter(
|
||||||
|
m => m === 'bubblewrap' || m === 'socat' || m === 'ripgrep',
|
||||||
|
);
|
||||||
|
if (missingDeps.length > 0) {
|
||||||
|
sandboxField.message =
|
||||||
|
`Sandbox dependencies not available: ${missingDeps.map(m => `${m} not installed`).join(', ')}.` +
|
||||||
|
` Install: sudo apt-get install -y ${missingDeps.join(' ')}`;
|
||||||
|
} else if (_sandboxStatusCache.details?.error) {
|
||||||
|
sandboxField.message = `Sandbox check error: ${_sandboxStatusCache.details.error}`;
|
||||||
|
} else if (_sandboxStatusCache.details?.libError) {
|
||||||
|
sandboxField.message = `Sandbox library error: ${_sandboxStatusCache.details.libError}`;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
const fullPayload = {
|
const fullPayload = {
|
||||||
ok: true,
|
ok: true,
|
||||||
version: VERSION,
|
version: VERSION,
|
||||||
providers: { enabled, available, status: providerStatuses },
|
providers: { enabled, available, status: providerStatuses },
|
||||||
|
sandbox: sandboxField,
|
||||||
};
|
};
|
||||||
if (anonymousKey !== null) fullPayload.anonymousKey = anonymousKey;
|
if (anonymousKey !== null) fullPayload.anonymousKey = anonymousKey;
|
||||||
sendJSON(res, 200, fullPayload);
|
sendJSON(res, 200, fullPayload);
|
||||||
@@ -1224,13 +1322,25 @@ async function handleChatCompletions(req, res) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
const chunks = [];
|
const chunks = [];
|
||||||
// try/finally: releaseSpawn MUST fire on every exit path — success
|
// D75 (v0.4.2) F7 — per-hop model override.
|
||||||
// (return at end), spawn throw (caught and re-thrown below), or the
|
// The chain config (routing.chains[<requested-model>][<hop>].model) is the
|
||||||
// D16 truncation-salvage return. The finally is the only mechanism
|
// model name to pass to THIS hop's provider plugin — NOT the model the
|
||||||
// that guarantees release across all three.
|
// user originally requested. Pre-D75, executeHopFn used hopModel for the
|
||||||
|
// cache key + audit ctx but passed the ORIGINAL irReq (with irReq.model
|
||||||
|
// still set to the user's request) into hopProviderPlugin.spawn(). Result:
|
||||||
|
// a 2-hop chain like [{anthropic, claude-sonnet-4-6}, {openai, gpt-5.5}]
|
||||||
|
// would spawn codex with --model claude-sonnet-4-6 on hop 1 — openai rejects
|
||||||
|
// the unknown model and the chain dies. This broke the core OLP value prop
|
||||||
|
// (cross-provider fallback with provider-appropriate model substitution).
|
||||||
|
//
|
||||||
|
// Fix: build a per-hop IR variant with the hop's model substituted. Skip
|
||||||
|
// the clone when hopModel === irReq.model (single-provider chains and
|
||||||
|
// chain hops whose model matches the request). Authority: ADR 0004 §
|
||||||
|
// Chain advancement step 1 (per-hop config supplies provider AND model).
|
||||||
|
const hopIrReq = irReq.model === hopModel ? irReq : { ...irReq, model: hopModel };
|
||||||
try {
|
try {
|
||||||
try {
|
try {
|
||||||
for await (const irChunk of hopProviderPlugin.spawn(irReq, authContext)) {
|
for await (const irChunk of hopProviderPlugin.spawn(hopIrReq, authContext)) {
|
||||||
// D16: check error chunks BEFORE pushing — preserves the invariant that
|
// D16: check error chunks BEFORE pushing — preserves the invariant that
|
||||||
// chunks array contains only delta/stop chunks. Without this, the catch
|
// chunks array contains only delta/stop chunks. Without this, the catch
|
||||||
// block's `chunks.length > 0` would mistake a single error chunk for
|
// block's `chunks.length > 0` would mistake a single error chunk for
|
||||||
@@ -1422,9 +1532,16 @@ async function handleChatCompletions(req, res) {
|
|||||||
// releaseSpawn fires exactly once in finally — regardless of normal
|
// releaseSpawn fires exactly once in finally — regardless of normal
|
||||||
// exhaustion, mid-stream throw, or iterator.return() from cache-layer
|
// exhaustion, mid-stream throw, or iterator.return() from cache-layer
|
||||||
// sourceAbortController propagation (§9).
|
// sourceAbortController propagation (§9).
|
||||||
|
//
|
||||||
|
// D75 (v0.4.2) F7 — per-hop model override (streaming path).
|
||||||
|
// Mirror the buffered-path fix: pass the hop's configured model (streamModel)
|
||||||
|
// into streamPlugin.spawn(), not the user's original ir.model. Skip the
|
||||||
|
// clone when ir.model === streamModel. See executeHopFn() above for the
|
||||||
|
// full F7 rationale + authority citation.
|
||||||
|
const streamIr = ir.model === streamModel ? ir : { ...ir, model: streamModel };
|
||||||
return (async function* sourceWithRelease() {
|
return (async function* sourceWithRelease() {
|
||||||
try {
|
try {
|
||||||
for await (const irChunk of streamPlugin.spawn(ir, authContext)) {
|
for await (const irChunk of streamPlugin.spawn(streamIr, authContext)) {
|
||||||
yield irChunk;
|
yield irChunk;
|
||||||
}
|
}
|
||||||
} finally {
|
} finally {
|
||||||
@@ -2011,8 +2128,10 @@ async function handleDashboard(req, res) {
|
|||||||
async function handleManagementDashboardData(req, res) {
|
async function handleManagementDashboardData(req, res) {
|
||||||
return _runOwnerOnlyManagementEndpoint(req, res, 'GET', '/v0/management/dashboard-data',
|
return _runOwnerOnlyManagementEndpoint(req, res, 'GET', '/v0/management/dashboard-data',
|
||||||
async (_req, res2, _identity, _auditCtx) => {
|
async (_req, res2, _identity, _auditCtx) => {
|
||||||
// Quota panel: collect quotaStatus from each loaded provider; null on
|
// Quota panel (legacy): collect quotaStatus from each loaded provider; null on
|
||||||
// throw or null return → "unavailable" indicator.
|
// throw or null return → "unavailable" indicator.
|
||||||
|
// DEPRECATED: kept for backwards compat with current dashboard.html (D82 will
|
||||||
|
// switch consumers to quota_v2; legacy 'quota' key removed at v1.0.0 or earlier).
|
||||||
const quota = [];
|
const quota = [];
|
||||||
for (const [name, provider] of loadedProviders) {
|
for (const [name, provider] of loadedProviders) {
|
||||||
try {
|
try {
|
||||||
@@ -2023,12 +2142,27 @@ async function handleManagementDashboardData(req, res) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// quota_v2 (D81 / Phase 5): normalized per-provider quota shape for enriched
|
||||||
|
// dashboard rendering. Built from aggregateProviderQuota() in lib/audit-query.mjs.
|
||||||
|
// Each entry contains: provider, status, schema_version, last_fresh_at,
|
||||||
|
// utilization, reset, representative_claim, fallback_percentage, overage, raw_available.
|
||||||
|
// Providers with null quotaStatus() return { status: 'unavailable', reason: ... }.
|
||||||
|
// Authority: ADR 0008 Amendment (D81) + ADR 0012 D81 + ADR 0013 Rule 5.
|
||||||
|
let quota_v2 = [];
|
||||||
|
try {
|
||||||
|
quota_v2 = await auditAggregateProviderQuota({ providers: loadedProviders });
|
||||||
|
} catch (err) {
|
||||||
|
// Graceful degradation: quota_v2 is optional enrichment; don't fail entire payload.
|
||||||
|
logEvent('warn', 'dashboard_data_quota_v2_failed', { error: err?.message ?? String(err) });
|
||||||
|
}
|
||||||
|
|
||||||
const WINDOW_24H = 24 * 60 * 60 * 1000;
|
const WINDOW_24H = 24 * 60 * 60 * 1000;
|
||||||
const payload = {
|
const payload = {
|
||||||
generated_at: new Date().toISOString(),
|
generated_at: new Date().toISOString(),
|
||||||
window_24h: auditAggregateRequests({ windowMs: WINDOW_24H, logEvent }),
|
window_24h: auditAggregateRequests({ windowMs: WINDOW_24H, logEvent }),
|
||||||
cache_hit_24h: auditCacheHitRateWindow({ windowMs: WINDOW_24H, logEvent }),
|
cache_hit_24h: auditCacheHitRateWindow({ windowMs: WINDOW_24H, logEvent }),
|
||||||
quota,
|
quota,
|
||||||
|
quota_v2,
|
||||||
spend_trend_30d: auditSpendTrendDaily({ days: 30, logEvent }),
|
spend_trend_30d: auditSpendTrendDaily({ days: 30, logEvent }),
|
||||||
top_fallback_chains_24h: auditTopFallbackChains({ windowMs: WINDOW_24H, limit: 10, logEvent }),
|
top_fallback_chains_24h: auditTopFallbackChains({ windowMs: WINDOW_24H, limit: 10, logEvent }),
|
||||||
cache_stats: cacheStore.stats(),
|
cache_stats: cacheStore.stats(),
|
||||||
@@ -2045,6 +2179,7 @@ async function handleManagementDashboardData(req, res) {
|
|||||||
async function handleManagementQuota(req, res) {
|
async function handleManagementQuota(req, res) {
|
||||||
return _runOwnerOnlyManagementEndpoint(req, res, 'GET', '/v0/management/quota',
|
return _runOwnerOnlyManagementEndpoint(req, res, 'GET', '/v0/management/quota',
|
||||||
async (_req, res2, _identity, _auditCtx) => {
|
async (_req, res2, _identity, _auditCtx) => {
|
||||||
|
// Legacy quota array (backwards compat).
|
||||||
const quota = [];
|
const quota = [];
|
||||||
for (const [name, provider] of loadedProviders) {
|
for (const [name, provider] of loadedProviders) {
|
||||||
try {
|
try {
|
||||||
@@ -2054,7 +2189,14 @@ async function handleManagementQuota(req, res) {
|
|||||||
quota.push({ provider: name, error: err?.message ?? String(err), available: null });
|
quota.push({ provider: name, error: err?.message ?? String(err), available: null });
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
sendJSON(res2, 200, { generated_at: new Date().toISOString(), quota });
|
// quota_v2 (D81): normalized per-provider quota shape per ADR 0008 Amendment (D81).
|
||||||
|
let quota_v2 = [];
|
||||||
|
try {
|
||||||
|
quota_v2 = await auditAggregateProviderQuota({ providers: loadedProviders });
|
||||||
|
} catch (err) {
|
||||||
|
logEvent('warn', 'management_quota_v2_failed', { error: err?.message ?? String(err) });
|
||||||
|
}
|
||||||
|
sendJSON(res2, 200, { generated_at: new Date().toISOString(), quota, quota_v2 });
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -2226,6 +2368,8 @@ export function createOlpServer() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
export { router, loadedProviders, VERSION };
|
export { router, loadedProviders, VERSION };
|
||||||
|
// Phase 7 PR-B: re-export sandbox manager test seam so tests can reset state.
|
||||||
|
export { __resetSandboxManagerForTests };
|
||||||
|
|
||||||
// Main guard: only listen when invoked as the entrypoint. ESM equivalent of
|
// Main guard: only listen when invoked as the entrypoint. ESM equivalent of
|
||||||
// `require.main === module` is comparing import.meta.url against argv[1].
|
// `require.main === module` is comparing import.meta.url against argv[1].
|
||||||
@@ -2238,11 +2382,30 @@ const isMain = (() => {
|
|||||||
})();
|
})();
|
||||||
|
|
||||||
if (isMain) {
|
if (isMain) {
|
||||||
const server = createOlpServer();
|
// Phase 7 PR-B (ADR 0014 § PR-B): bootstrap sandbox before listening.
|
||||||
server.listen(PORT, '127.0.0.1', () => {
|
// bootstrapSandbox() is idempotent + error-safe — server always starts even
|
||||||
const enabledCount = loadedProviders.size;
|
// if sandbox initialization fails (degrades to unsandboxed, logs a warning).
|
||||||
|
// The /health.sandbox.active field reflects the result.
|
||||||
|
const sandboxBoot = await bootstrapSandbox();
|
||||||
|
if (sandboxBoot.active) {
|
||||||
process.stdout.write(
|
process.stdout.write(
|
||||||
`OLP v${VERSION} listening on :${PORT} (${enabledCount} providers enabled — Phase 1 in progress)\n`,
|
`OLP sandbox active (config-at-boot): ${sandboxBoot.summary}\n`,
|
||||||
|
);
|
||||||
|
} else {
|
||||||
|
process.stderr.write(
|
||||||
|
`OLP sandbox NOT active: ${sandboxBoot.reason} — ` +
|
||||||
|
`provider spawns will run UNSANDBOXED (test/dev only; not safe for cloud)\n`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
const server = createOlpServer();
|
||||||
|
server.listen(PORT, BIND, () => {
|
||||||
|
const enabledCount = loadedProviders.size;
|
||||||
|
// D74 P3-5: banner no longer hardcodes the phase. Derives from VERSION
|
||||||
|
// (which advances at every Phase close) so banner stays accurate
|
||||||
|
// without future-maintenance touch-ups at every Phase boundary.
|
||||||
|
process.stdout.write(
|
||||||
|
`OLP v${VERSION} listening on :${PORT} (${enabledCount} providers enabled)\n`,
|
||||||
);
|
);
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|||||||
+3614
-30
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user