Port OCP server.mjs:842-1109 plan-usage probe to lib/providers/anthropic.mjs:quotaStatus(). ## Authority citations (CLAUDE.md hard requirement #1) 1. Schema pin: ~/.cc-rules/memory/learnings/anthropic_plan_usage_probe_schema_2026_05_26.md — 13-field canonical schema verified live 2026-05-26; 3 fields new vs OCP 2026-04 capture (5h-status, 7d-status, overage-reset). 2. OCP port source: OCP server.mjs:842-1109 — usageCache, oauthRefreshBackoff, getOAuthCredentials, refreshOAuthToken, fetchUsageFromApi, parseRateLimitHeaders. Claude Code CLI uses this same POST /v1/messages internally (verified 2026-05-26 by `strings` on @anthropic-ai/claude-code v2.1.142 / v2.1.150 Mach-O binary — see audit memory). This is observed CLI behaviour, not invention. 3. ADR 0002 Amendment 8 — READ-ONLY exemption for quotaStatus() direct-API access. Three constraints satisfied: READ-ONLY (max_tokens:1, body discarded), subscription-scope (same readAuthArtifact() creds as spawn path), idempotent-failure (returns null / stale on any error, never throws). 4. ADR 0013 — OAuth READ-ONLY consumption rules. All 7 rules satisfied: Rule 1: credential reuse via readAuthArtifact() (env → .credentials.json → keychain). Rule 2: only POST /v1/messages (no other endpoints). Body discarded; headers-only. Rule 3: 5min TTL cache; 60s–3600s exponential backoff; stale-on-failure. Rule 4: opt-in via ~/.olp/config.json providers.anthropic.quota_probe_enabled (default false). Rule 5: schema pin committed to memory file; drift detection protocol in place. Rule 6: doctor check anthropic.quota_probe_reachable surfaces probe status. Rule 7: does not govern spawn-path refresh (separate concern). 5. ADR 0012 D80 — Phase 5 charter: this commit is the D80 deliverable. ## Live probe transcript (2026-05-26 from MacBook keychain OAuth credentials) Path B verification per ADR 0013 Rule 5: curl -s -i -m 20 -X POST https://api.anthropic.com/v1/messages \ -H "Authorization: Bearer <token>" \ -H "anthropic-beta: oauth-2025-04-20" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{"model":"claude-haiku-4-5-20251001","max_tokens":1,"messages":[{"role":"user","content":"."}]}' Response (header lines only): HTTP/2 200 anthropic-ratelimit-unified-status: allowed anthropic-ratelimit-unified-5h-status: allowed anthropic-ratelimit-unified-5h-reset: 1779794400 anthropic-ratelimit-unified-5h-utilization: 0.09 anthropic-ratelimit-unified-7d-status: allowed anthropic-ratelimit-unified-7d-reset: 1780225200 anthropic-ratelimit-unified-7d-utilization: 0.32 anthropic-ratelimit-unified-representative-claim: five_hour anthropic-ratelimit-unified-fallback-percentage: 0.5 anthropic-ratelimit-unified-reset: 1779794400 anthropic-ratelimit-unified-overage-disabled-reason: org_level_disabled_until anthropic-ratelimit-unified-overage-status: rejected (no anthropic-ratelimit-unified-overage-reset — expected: only present on active overage) 12/13 fields present. overage-reset absent = no active overage (expected per audit memory). All fields parsed correctly by _parseRateLimitHeaders(). Confirmed via D80 smoke test. ## Implementation A. quotaStatus() — full probe implementation replacing D4 null stub: - _readProviderConfig('anthropic') gate (Rule 4 opt-in) - 5min module-level cache check (quotaProbeState.cache) - 60s–3600s exponential backoff check (quotaProbeState.backoffUntil / backoffMs) - readAuthArtifact() credential read (env → .credentials.json → macOS keychain) - _probeOnce() → POST /v1/messages with 4 required headers; body discarded - 401/403 → single refresh-and-retry via _refreshAccessToken() - On success: cache { fetchedAt, data } + reset backoff to MIN - On failure: _scheduleBackoff() (doubles backoffMs, caps at MAX) + return stale or null - Return shape: { probedAt, source, schemaVersion, stale, fields:{...13}, raw:{...} } B. _parseRateLimitHeaders() — all 13 fields (3 new vs OCP): - status, representative_claim, reset, fallback_percentage (aggregate) - status_5h, utilization_5h, reset_5h (5h window) - status_7d, utilization_7d, reset_7d (7d window) - overage_status, overage_disabled_reason, overage_reset (overage) - Numeric strings → numbers; missing fields → null (not 0 or "unknown") C. _refreshAccessToken() — uses Node.js built-in https (no fetch/3rd-party deps). Shared backoff state via quotaProbeState. Max one refresh per backoff window. D. _probeOnce() — uses Node.js built-in https. 15s timeout. Drains + discards body. E. _readProviderConfig() — reads ~/.olp/config.json providers.<name> block. OLP_HOME respected (same as lib/keys.mjs). Never throws; returns {} on error. F. doctorChecks() — new anthropic.quota_probe_reachable check (ADR 0013 Rule 6): - status: ok when probe disabled (returns advisory message) - status: ok when probe succeeds (shows utilization %) - status: warn when stale cache exists (probe failed but cache present) - status: fail when no cache + probe failed (fix_commands + human_steps recipe) G. docs/v1x-roadmap.md — #8 Dashboard enrichment entry (D79 follow-up) added. ## Tests - All 720 existing tests pass (npm test). - Suite 33j updated to include anthropic.quota_probe_reachable in the expected probe set (3 probes total, previously 2). - D83 (Suite 38) will add quota-probe unit tests with mock HTTP server. ## What NOT changed - dashboard.html — untouched (D82) - lib/audit-query.mjs — untouched (D81) - lib/providers/codex.mjs, mistral.mjs — untouched (D84 NO-GO per ADR 0012 Amendment 1) Co-authored-by: dtzp555 <dtzp555@gmail.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
16 KiB
OLP v1.x Roadmap — Deferred Work Tracker
Purpose. Single landing page for every Phase-1 deferral that an actual v1.x sprint must pick up. Each entry cross-references its ratifying ADR, its GitHub issue (if any), and the load-bearing code anchor so a future maintainer can resume without spelunking the commit history.
Status: Living document. Add new entries at the top. Each item should answer:
- What is deferred?
- Why was it deferred (link the ratifying ADR amendment).
- Where does the work live in the tree today (file + anchor).
- When does it need to land (trigger: load profile, security event, governance amendment).
Reading order for a v1.x sprint kickoff. As of 2026-05-25, #1 (streaming SF, D57+D58) and #2 (multi-key auth, Phase 2) are CLOSED, and #4 and #7 closed in D56. Remaining v1.x scope: #3 (soft trigger reactivation), #5 (provider cacheKeyFields mask), #6 (streaming SPAWN_FAILED salvage — unbundled from #1 at #1 close). All three remaining items have explicit "trigger to start" gates that have not fired.
#1 — Streaming-path singleflight + TOCTOU close — ✅ SHIPPED (D57 + D58, 2026-05-25)
- Status. Closed. Trigger (b) fired 2026-05-25 — maintainer "go" after v0.3.1. Shipped across three D-days:
- D57 (PR #36) — cache layer:
cacheStore.getOrComputeStreaming(keyId, cacheKey, sourceFactory, opts) → { stream, isFirst, role }with_streamingInflightMap, tee fan-out, late-joiner replay buffer, per-client backpressure (PER_CLIENT_QUEUE_CAP=1MB), replay cap (ACCUMULATED_REPLAY_CAP=10MB), AbortController propagation, synchronous Map check+insert (closes TOCTOU). Suite 27 = 12 unit tests. - D58 (PR #37) — server.mjs wiring: streaming branch swap;
tryAcquireSpawn/releaseSpawnmoved insidesourceFactoryclosure (D38 §7 coordination);CONCURRENCY_LIMITfallthrough preserved;X-OLP-Streaming-Inflight: source | attachedheader;cache_status: 'streaming_attached'audit value + audit-query gauge reconciliation;res.on('close') → stream.return()for client disconnect; D16 truncated-not-cached invariant preserved viacacheStore.deleteon stop-less exhaustion. Suite 28 = 8 HTTP integration tests. - D59 (this commit) — docs polish: README known-limitations entry inverted; this roadmap entry closed; issue #16 closed.
- D57 (PR #36) — cache layer:
- Design authority.
docs/adr/0005-cache-cross-provider.mdAmendment 8 — implemented per spec §§1–14 across D57+D58. - Tracking issue. GitHub issue #16 — CLOSED at D59 with refs to D57+D58 PRs.
- Final test count delta. 603 (v0.3.1) → 623 (v0.3.2/v0.4.0). +20 tests across the SF arc.
- Deferred sub-items (left here as future-work pointers, NOT blocking #1 closure):
X-OLP-Streaming-Inflight: solovalue not emitted on the wire (Amendment 8 §11). It's observable only post-stream via thestreaming_inflight_source_donelog event'sattached_count: 0. Future ADR amendment may expose via HTTP trailer.streaming_inflight_joinlog event from_attachClientcache-layer path (carries no provider/model context). D58 emits the event from the server-layer wrapper instead; cache-layer emission would need a provider/model plumb (TODO marker atlib/cache/store.mjs:~620).isFirstfield returned bygetOrComputeStreamingis currently unused by server.mjs (rolesupersedes). Could be removed in a future cache-layer API cleanup.
#2 — Multi-key auth (lib/keys.mjs) — PHASE 2 ACTIVE (no longer deferred)
- Status. Phase 2 active as of 2026-05-25. Design ratified at D43-B. This entry stays for cross-reference but is no longer a v1.x deferral; implementation D-days D44+ execute within Phase 2.
- What. Per-API-key identity, namespace scoping for the cache, ownership tier (owner vs guest) for header gating, and audit log of which key issued which request. Detailed scope in ADR 0007.
- Design ADR (ratified).
docs/adr/0007-multi-key-auth.md— Option 2 (filesystem manifest) + opaque token, with explicit forward path to Option 3 hybrid (SQLite-indexed mirror) when Phase 3+ Dashboard / SQL-aggregate quota work justifies. Migratable, manifest-as-SPOT. - Tracking. Not a GitHub issue. Tracked here + via ADR 0007 acceptance criteria (§ 10) which drive the D44+ test surface.
- Resolves.
X-OLP-Fallback-Detailowner-only gating (D40 / ADR 0004 Amendment 5 — currently ungated; Phase 2 re-gates per ADR 0007 § 7)./healthper-key visibility (currently anonymous-only — owner / guest / anonymous tiers per ADR 0007 § 7).
- Code anchors today (unchanged at ADR ratification; replaced by D44+ implementation).
lib/cache/store.mjs:77-79per-keyId namespace Map — wire is in place.lib/cache/store.mjs:287singleflight composition${keyId}:${cacheKey}— wire is in place.server.mjs:502, :531— the twokeyId='__anonymous__'call sites to replace.server.mjs:392—/healthhandler entry (Phase 2 gate).server.mjs:1072, :1101—X-OLP-Fallback-Detailheader-write paths (Phase 2 gate).
- Trigger (already fired). Maintainer opened Phase 2 sprint 2026-05-25.
#3 — Soft trigger reactivation (ADR 0004 Amendment 2)
- What. Per-provider
quotaStatuspolling,softThresholdcomparisons, soft-skip advancement when quota approaches limit. CurrentlyevaluateSoftTriggersalways returnsfalsebecausequotaSnapshotis never populated. - Why deferred. v0.1 hard triggers (SPAWN_FAILED / CLI_NOT_FOUND / SPAWN_TIMEOUT / CONCURRENCY_LIMIT) are sufficient for fallback advancement at personal/family scale. Soft triggers require persistent quota snapshots and a polling mechanism, which adds operational surface (timer drift, snapshot staleness, observability burden).
- Design ADR.
docs/adr/0004-fallback-engine.mdAmendment 2 — explicit v1.x deferral with mitigations (startup warning if user configures soft thresholds without runtime enforcement). - Tracking. Not a GitHub issue. Tracked here + via the startup warning in
server.mjs(the_softTriggersConfiguredwarn emission). - Blocks.
- Issue #8 (
X-OLP-Provider-Usedchain-origin semantics) — Option A (trackfirstAttemptedProvider) becomes preferable once soft triggers can fire. See ADR 0004 Amendment 6 § v1.x re-evaluation. X-OLP-Fallback-Detailtrigger_type: 'soft'path — currently dead code, becomes live with this work.
- Issue #8 (
- Code anchors today.
lib/fallback/engine.mjsevaluateSoftTriggers(returns false unconditionally at v0.1).lib/providers/base.mjsProvider.quotaStatuscontract (declared but unused at v0.1).
- Trigger to start. First quota-rate-limit event in the wild — at which point the operator would want pre-emptive advancement rather than spawn-then-fail.
#4 — /health activeSpawns integration
- What. Surface D38
getActiveSpawnCount(providerName)per-provider on the/healthendpoint at the pathproviders.status.<name>.activeSpawns. - Why deferred. D38 (issue #1) shipped the runtime enforcement and exported
getActiveSpawnCount;/healthintegration was scoped out as forward-looking polish. - Design ADR.
docs/adr/0002-plugin-architecture.mdAmendment 6 — names the target path explicitly: "/healthintegration deferred — when surfaced there will land atproviders.status.<name>.activeSpawns; not wired at D38." - Tracking. Not a GitHub issue. Tracked here.
- Code anchors today.
lib/providers/index.mjsexportsgetActiveSpawnCountalready.server.mjs handleHealth— extension point for the new field.
- Trigger to start. First time the maintainer wants per-provider concurrency visibility for capacity planning.
#5 — Provider-level cacheKeyFields (per-plugin mask)
- What. Per-plugin declaration of which IR fields are actually consumed by the underlying CLI invocation, used by
computeCacheKeyto skip fields that the plugin drops at spawn. Reduces spurious-miss rate from the v0.1 conservative-posture trade-off (Amendment 7). - Why deferred. At personal/family scale the extra spawn cost from spurious misses is negligible. The contract extension adds complexity (per-plugin field set + plumbing through
buildDefaultChain→executeHopFn→computeCacheKey). - Design ADR.
docs/adr/0005-cache-cross-provider.mdAmendment 7 § Forward path. - Tracking. Not a GitHub issue. Tracked here.
- Code anchors today.
- Plugin file headers — each lists its "fields dropped at spawn" table for human reference; the v1.x amendment makes that table machine-readable.
lib/cache/keys.mjs computeCacheKey— would acceptpluginCacheKeyMaskparameter.
- Trigger to start. First time spurious-miss rate becomes a measurable load factor.
#6 — Streaming-path SPAWN_FAILED salvage
- What. Currently the streaming branch does NOT participate in D16 salvage (the salvage-on-SPAWN_FAILED + chunks pattern that the buffered path uses). Streaming SPAWN_FAILED mid-stream → the truncation marker (D35 #10) fires, but no salvage logic captures partial chunks for downstream cache reuse.
- Why deferred. Less impactful than #1 — at most one client benefits per spawn event, and the buffered path already provides salvage for the bulk of requests. Streaming is the minority path.
- Status update post-#1 close (2026-05-25). #1 was originally bundled with #6 in the design ADR (Amendment 8). The tee architecture as implemented does NOT carry salvage semantics — D57's tee writes
accumulatedChunksto cache only on normal source completion (stop chunk seen); on SPAWN_FAILED mid-stream the cache layer rejects all clients with the error and does NOT persist partial chunks. D58 preserves D16's truncated-not-cached invariant via server-layercacheStore.deleteon stop-less exhaustion. #6 therefore remains independently deferrable. - Design ADR. Not yet ratified. The unbundling from #1 means #6 now needs its own ADR amendment when triggered.
- Tracking. Not a GitHub issue. Tracked here.
- Trigger to start. First report of streaming-path SPAWN_FAILED mid-stream where partial-chunk salvage would have helped a downstream caller. Practically unlikely at family scale.
#8 — Dashboard enrichment: per-provider subscription quota + reset times + 1-min refresh + manual refresh (D78 follow-up)
- What. Phase 3 dashboard (D51
dashboard.html, v0.3.0) shows: per-provider quota (currently always "n/a — no quota api"), last-24h request count + cache hit + fallback rate, 30d request-count sparkline, top fallback chains. Maintainer request 2026-05-26 post-D78: extend to show what each enabled provider's subscription is actually consuming, with reset times visible, refresh once per minute (current 30s is OK but maintainer specified 1min target), and a manual refresh button. Reference design: Claude.ai's ownclaude.ai/settings/usagepage — current session bar with "Resets in 1hr 6min", weekly all-models bar with "Resets Sun 9:00 PM", per-model bar (Sonnet only), additional features (routine runs), usage credits + monthly spend limit + auto-reload toggle. - Why deferred. v0.3.0/v0.4.x ships the dashboard frame but
provider.quotaStatus()returnsnullin all three v0.1 plugins (anthropic / openai / mistral). The ratifying spec in ADR 0004 Amendment 2 puntsquotaStatus()to v1.x ("soft trigger reactivation") — this dashboard ask is the operator-facing reason that work would land. - What this requires. Per-provider plugin work + dashboard.html UI work + audit-query.mjs aggregation:
lib/providers/anthropic.mjs quotaStatus()— discover where the maintainer's Claude.ai subscription quota state is exposed. Candidates: (a)claudeCLI command (e.g.,claude usage) if Anthropic adds one — currently absent; (b) parsing theclaude-codeoutput for rate-limit error messages and caching state from headers; (c) hittingapi.anthropic.com/v1/.../usagedirectly via the OAuth refresh token — not a documented endpoint, primary-source risk. ADR 0002 Rule 1 / Rule 5 require an authority citation before any implementation. Likely path: wait until Anthropic publishes a documented endpoint, OR derive from audit-side request counts only (no real quota truth, just "you sent N requests in the current 5h window").lib/providers/openai.mjs quotaStatus()— codex CLI doesn't expose ChatGPT-subscription quota state. OpenAI rate-limit headers per request might be parseable but ADR 0004 Amendment 2 explicitly says no plugin parses HTTP status at v0.1.lib/providers/mistral.mjs quotaStatus()— Le Chat Pro has/v1/usageendpoint per Mistral docs (verify).dashboard.htmlUI restructure to a Claude.ai-style layout: rows of (label, bar, "Resets in X" / "Resets atlib/audit-query.mjs— extendaggregateRequests/spendTrendDailyto compute "in the current rolling window" (since session/week start) per provider. Today's aggregates are wall-clock windows; subscription resets are per-account-anchored. Need a way to model session windows (e.g., "Anthropic 5h-from-first-request-since-last-reset").
- Reference (maintainer 2026-05-26). Screenshot of
claude.ai/settings/usageshared inline. Key panels: Plan usage limits (current session + resets-in), Weekly limits (All models / Sonnet only / per-feature breakdown, each with resets-on), Additional features (Daily included routine runs N / 15), Usage credits (toggle + spent vs monthly limit + auto-reload + buy-credits link). - Tracking. Not yet a GitHub issue. Track here + cross-reference ADR 0004 Amendment 2 (soft trigger reactivation — same
quotaStatus()data-source work) when this becomes Phase 5 scope. - Code anchors today.
dashboard.html— current 4 panels; needs restructure to Claude.ai-style row layoutlib/providers/anthropic.mjs/openai.mjs/mistral.mjs—quotaStatus()returns null todaylib/audit-query.mjs— currentaggregateRequestsis wall-clock-window; needs session-window variant
- Trigger to start. ANY of: (a) Anthropic publishes a documented
claude usageCLI orapi.anthropic.com/v1/usageendpoint, (b) maintainer hits real "I want to see quota right now" pain often enough to design without per-provider truth (audit-derived only), (c) Phase 5 multi-tenant adds per-key spend limits and the dashboard needs to surface those.
#7 — AUTH_MISSING tuple path test coverage (D40 follow-up)
- What. Dedicated test in
test-features.mjsSuite D40 that asserts thefallbackDetailtuple records the AUTH_MISSING path withtrigger_type: 'auth_missing'. D40 reviewer flagged this as the last gap in the engine-path matrix; code is structurally correct, just lacks an explicit pin. - Why deferred. Low priority — the AUTH_MISSING early-return branch has the tuple push BEFORE it (verified in D40 reviewer pass), so coverage is implicit via the other engine-path tests. A 3-line dedicated test would make the pin explicit.
- Design. No ADR needed. ~5-line test addition.
- Tracking. Not a GitHub issue. Tracked here.
- Trigger to start. Next routine test-suite hardening pass, OR when AUTH_MISSING handling is changed for any reason.
Adding a new entry
When a future D-day defers work, the deferring commit should:
- Always update this file with a new entry at the top.
- Always name the ratifying ADR amendment (or note "no ADR yet — future work needs one").
- Always name the load-bearing code anchor (
file:lineform preferred over symbolic names — the symbolic name can drift). - Always name a concrete trigger to start the work — vague triggers ("when needed") let entries rot.
- If the deferral has a GitHub issue, keep it OPEN and reference it here. If it does NOT, leave a note explaining why (e.g., "tracked here only — no external governance event filed").
The maintainer's session-startup discipline should grep this file at sprint kickoff. If an entry's "trigger to start" condition is met, it leaves this page and becomes a sprint item.