mirror of
https://github.com/dtzp555-max/olp.git
synced 2026-07-22 13:35:10 +00:00
Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6edf6e0b94 | ||
|
|
f3716a19fd |
@@ -6,6 +6,48 @@ All notable changes to OLP land here. Per `CLAUDE.md` release_kit overlay, this
|
||||
|
||||
(empty — Phase 5 entries land here once Phase 5 opens)
|
||||
|
||||
## v0.4.2 — 2026-05-26
|
||||
|
||||
### Post-v0.4.1 hotfix batch (D75) — real-machine E2E findings
|
||||
|
||||
Patch release fixing 5 bugs caught by **real-machine E2E testing on PI231 + Mac mini (2026-05-26 session)** — bugs that prior D-day reviewers AND the post-v0.4.0 maintainer review both missed because they reviewed against spec text and against the local OLP install's `~/.codex/auth.json` shape (cached from an older codex CLI version), not against real provider CLIs running on a remote operator host that did `npm install -g @openai/codex` for the first time on 2026-05-26 and got codex CLI v0.133.0.
|
||||
|
||||
**Root cause of the missed-bug class.** D6 (codex plugin authoring) explicitly documented three unpinned assumptions (A3 = access-token field name, A4 = NDJSON event schema, A2-adjacent = trusted-directory sandbox). D6 noted "D7 E2E will pin." D7 then shipped without performing real-codex-CLI E2E (the E2E gating mark was carried but the actual run was deferred). Every subsequent D-day reviewer trusted the D6/D7 codex plugin code unchanged because the static review couldn't see that the v0.133.0 CLI had moved the auth-token field, the event schema, AND added a new trusted-directory sandbox flag. The D74 maintainer review focused on `/health` / `/cache/stats` / `/v0/management/dashboard-data` payload shapes — none of which exercise the codex plugin's spawn path. F7 (per-hop model override) is a different class of miss — every reviewer read `executeHopFn(provider, model, ir)` and saw `model` consumed for cache key + audit ctx, but none traced through to confirm `model` is ALSO substituted into the IR passed to `provider.spawn()`. The function signature implied per-hop semantics that the body never fully delivered.
|
||||
|
||||
- **[F1] codex auth.json schema pin — codex CLI v0.133.0 nests the access token under `tokens.access_token`** (verified empirically on PI231 / Mac mini, 2026-05-26). Pre-D75 `readAuthArtifact()` read only top-level `creds.access_token` / `creds.token` / `creds.accessToken` — all undefined under v0.133.0 → returned `null` → OLP reported "auth artifact missing" via `/health` and `olp doctor` AND refused to spawn codex even when the user had fully completed `codex login`. Fix: prepend `creds?.tokens?.access_token` to the precedence chain at BOTH call sites (`OPENAI_CODEX_AUTH_PATH` override branch + default `$CODEX_HOME/auth.json` branch). Legacy top-level fields preserved as fallback for backward compat with older codex CLI versions.
|
||||
- **[F2] codex spawn args — codex CLI v0.133.0 trusted-directory sandbox requires `--skip-git-repo-check`.** v0.133.0 refuses with `"Not inside a trusted directory and --skip-git-repo-check was not specified."` when spawned outside a git repo, exits non-zero with zero NDJSON output → OLP surfaces `SPAWN_FAILED` with no usable chunks → the fallback engine advances to next hop unnecessarily even when codex is configured and authenticated. OLP's typical deploy CWD (`~/olp/`) is NOT a git repo on operator hosts. Fix: add `'--skip-git-repo-check'` to the args array before `--model`. OLP is the trusted caller (operator's own server invoking the operator's own subscription via documented `codex exec` automation); the sandbox safeguards interactive shells, not pre-authorized automation.
|
||||
- **[F3] codex NDJSON event shape pin — codex CLI v0.133.0 emits `item.completed` + `turn.completed` + `turn.failed`**, not the D6-assumed `content`/`delta`/`text` + `type:'stop'`/`done:true` shapes. Real v0.133.0 stream (verified empirically): `{"type":"thread.started",...}` → `{"type":"turn.started"}` → `{"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"<response>"}}` → `{"type":"turn.completed","usage":{...}}`. Pre-D75, every chunk was silently dropped by `codexChunkToIR()` → response body had `content: null`. Fix: prepend three new recognizers (`item.completed` with `item.type === 'agent_message'` → IR delta; `turn.completed` → IR stop; `turn.failed` → IR error). Legacy D6 defensive recognizers preserved below as forward/backward compat fallbacks.
|
||||
- **[F4] `olp status` reads `body.stats.cache.size` (not OCP-era `body.cache.entries`).** Same class as D74 P2-3 (which fixed `cmdUsage` + `cmdCache`); D74 missed the parallel bug in `cmdStatus`. Server payload nests cache stats as `body.stats.cache.{hits, misses, size, inflightCount}` per `server.mjs handleManagementStatus`, and `CacheStore.stats()` has no `entries` field per `lib/cache/store.mjs`. Pre-D75 output showed `entries=?`. Fix: read `c.size` for entries display; also surface `inflightCount` when present.
|
||||
- **[F7] per-hop chain `model` field now overrides IR model in `provider.spawn()`.** Pre-D75 `executeHopFn(hopProvider, hopModel, irReq)` used `hopModel` for cache key + audit ctx but passed the ORIGINAL `irReq` (with `irReq.model` = user's original request) to `hopProviderPlugin.spawn(irReq, authContext)`. A chain config `[{provider:anthropic, model:claude-X}, {provider:openai, model:gpt-5.5}]` would always spawn BOTH plugins with `--model claude-X` — openai rejected the unknown model and the chain died. This broke the core OLP value prop (cross-provider fallback with provider-appropriate model substitution per hop). Fix: build a per-hop IR variant with `{...irReq, model: hopModel}` and pass that to spawn. Conditional skips clone when `hopModel === irReq.model` (common case: single-provider chains, or single-hop chains where the chain config repeats the request model). Applied to BOTH the buffered path (`executeHopFn`) AND the streaming path (`sourceFactory` for `getOrComputeStreaming`). **Authority:** ADR 0004 § Chain advancement step 1 (per-hop config supplies provider AND model — the contract was always specified, but the code didn't complete it).
|
||||
|
||||
**Phase 5 process learning recorded.** Every provider plugin's D-day must include a real-CLI E2E ON A REMOTE OPERATOR HOST before merging — not on the maintainer workstation (which may have an older CLI cached from a prior install, hiding new field renames / new sandbox flags / new event shapes). The D6/D7 codex E2E was deferred and that deferral compounded across 3 layers (D6 = unpinned, D7 = pinning deferred, D8+ = trusted D6/D7 unchanged). F7 reinforces a separate lesson: when a function signature takes `(provider, model, ir)`, reviewers must check that `model` is consumed everywhere downstream, not just at the call site they happened to look at.
|
||||
|
||||
**Out of D75 scope (deferred to Phase 5 explicit ADR amendments):**
|
||||
- F5 (server bind 127.0.0.1 / `OLP_BIND` env) — needs `lib/keys.mjs` anonymous-key trust boundary review before binding to non-loopback by default
|
||||
- F6 (`olp doctor` client-vs-server-side limit detection) — needs design ADR amendment for trigger taxonomy
|
||||
|
||||
- **Test count delta:** 704 (v0.4.1) → 714 (v0.4.2). +10 D75 regression tests in Suite 36 (36i through 36r).
|
||||
- **Files touched:** `lib/providers/codex.mjs` (F1+F2+F3), `bin/olp.mjs` (F4 cmdStatus), `server.mjs` (F7 buffered + streaming spawn sites), `test-features.mjs` (Suite 36 extension), `package.json` (version), `CHANGELOG.md` (this entry).
|
||||
- **Authority:** ADR 0002 (provider contract — codex plugin), ADR 0004 (fallback engine — per-hop model contract), `lib/providers/codex.mjs` D6 assumption A2/A3/A4 docstrings (which all said "D7 will pin" and D7 never did); codex CLI v0.133.0 on-disk schema + `codex exec --help` output verified empirically on PI231 (2026-05-26 E2E session); Iron Rule 第二律 evidence-over-should-work; CLAUDE.md `release_kit.phase_rolling_mode` cross-Phase discipline.
|
||||
|
||||
## v0.4.1 — 2026-05-26
|
||||
|
||||
### Post-Phase-4 hotfix batch (D74) — maintainer-review findings
|
||||
|
||||
Patch release fixing 5 issues caught by maintainer post-v0.4.0 independent review. Every finding was a real runtime bug that the per-D-day fresh-context opus reviewers all missed because they reviewed against spec text, not against the runtime contract (default `auth.allow_anonymous: false`, real `/health` payload shape, real `/cache/stats` payload shape, real `/v0/management/dashboard-data` payload shape). **Phase 4 lesson: future implementation D-days MUST include at least one test that boots the server with the default production config and exercises the new feature end-to-end** — not just stub-mocked codepaths.
|
||||
|
||||
- **[P1-1] `olp doctor` no longer false-negatives on auth-required `/health`.** `lib/doctor.mjs` now accepts an `authHeaders` option (threaded from `bin/olp.mjs` `cmdDoctor` via the existing `authHeaders()` chain) and passes it to the `server.running` + `server.version` probes. The `server.running` check now distinguishes 401/403 ("server up but bearer token missing/invalid — set `OLP_API_KEY`") from "server unreachable" — so the `kind` discriminator routes to a clean fix-auth path instead of `fix_server` when the operator just forgot to export the env var.
|
||||
- **[P1-2] `bin/olp-connect` validates token shape + shell-quotes rc writes.** New `validate_olp_token <key> <source>` helper enforces the `^olp_[A-Za-z0-9_-]{43}$` regex (per ADR 0007 § 3 token format) at all 3 input sites: `--key` arg, `/health.anonymousKey` server-advertised consumption, and the interactive prompt fallback. New `shell_quote <value>` helper wraps rc-file writes (`export OPENAI_BASE_URL=$(shell_quote ...)`) so even a hypothetical bypass of the validator can't inject shell metacharacters into a sourced rc. systemd `environment.d/olp.conf` write additionally rejects embedded newlines. Hostile or malformed keys can no longer persist as shell startup injection.
|
||||
- **[P2-3] `olp usage` + `olp cache` human formatter rewritten against the real payload shape.** `cmdUsage` previously read `body.usage_24h.requests` / `body.providers` / `body.top_fallback_chains` — all undefined under the actual server payload shape — so users saw "requests: ?" + missing per-provider quota + missing top-chains. Now reads `body.window_24h.request_count` / `body.cache_hit_24h.hit_rate` / `body.quota` / `body.top_fallback_chains_24h` per `server.mjs:2027` + `lib/audit-query.mjs`. `cmdCache` previously read `body.entries` / `body.bytes` / `body.maxBytes` (OCP-era field names). Now reads `body.size` / `body.inflightCount` per `CacheStore.stats()` and computes hit rate from `hits + misses`.
|
||||
- **[P2-4] `olp-plugin/` `fmtHealth` iterates `providers.status` correctly.** Previously walked `Object.entries(body.providers)` which surfaced `enabled` / `available` / `status` as pseudo-providers (chat output showed `🟢 status` instead of `🟢 anthropic`). Now extracts the real provider map from `body.providers.status` and renders enabled/available counts in a header line + per-provider names with `activeSpawns` when present. Falls back to flat `body.providers.*` for the older OCP shape (backwards compat).
|
||||
- **[P3-5] Stale v0.3.0-era doc strings updated.** README header status line + Implementation Status § now reflect v0.4.0 shipped + Phase 5 open. `server.mjs` startup banner no longer hardcodes "Phase 1 in progress" (now just lists version + provider count — derives accurate state from `VERSION` without future maintenance touch-ups).
|
||||
|
||||
**Phase 4 process learning recorded.** Per Iron Rule 第二律 (evidence over "should work"), every D-day review pass must include at least one runtime smoke against the default production config. The D-day reviewer rubric is updated implicitly — D74 Suite 36 tests pin the wire-contract shape so a future D-day refactoring server payloads can't silently re-break the CLI / plugin / docs.
|
||||
|
||||
- **Test count delta:** 696 (v0.4.0) → 704 (v0.4.1). +8 D74 regression tests in Suite 36.
|
||||
- **Files touched:** `lib/doctor.mjs` (P1-1), `bin/olp.mjs` (P1-1 + P2-3), `bin/olp-connect` (P1-2), `olp-plugin/index.js` (P2-4), `server.mjs` (P3-5 banner), `README.md` (P3-5), `test-features.mjs` (Suite 36 regression), `package.json` (version), `CHANGELOG.md` (this entry).
|
||||
- **Authority:** maintainer independent review of `main` / `v0.4.0` / commit `ee4d945` (2026-05-26 session); Iron Rule 第二律 evidence-over-should-work; CLAUDE.md `release_kit.phase_rolling_mode` cross-Phase discipline ("hotfix to a shipped Phase N deliverable → bump patch, tag, release before next push").
|
||||
|
||||
## v0.4.0 — 2026-05-26
|
||||
|
||||
### Phase 4 — Operator + Client UX (D60 → D73)
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing, automatic fallback, content-addressed caching — so your IDEs and family clients keep working as long as *any* of your subscriptions has quota left.
|
||||
|
||||
> **Status:** v0.3.0 shipped (2026-05-25) — Phase 1 multi-provider proxy core (v0.1.0 + v0.1.1) + Phase 2 multi-key auth + audit + owner gating + keygen CLI (v0.2.0) + Phase 3 Dashboard + audit query layer + daily audit rotation (v0.3.0). Phase 4 (per-key per-provider auth + audit retention + SQLite hybrid + provider-cost weights) is the next planned milestone. Sections marked _placeholder_ land alongside the relevant phase of work (see [phase plan](#phase-plan)).
|
||||
> **Status:** v0.4.0 shipped (2026-05-26) — Phase 1 multi-provider proxy core (v0.1.0 + v0.1.1) + Phase 2 multi-key auth + audit + owner gating + keygen CLI (v0.2.0) + Phase 3 Dashboard + audit query layer + daily audit rotation (v0.3.0) + Phase 4 Operator + Client UX (v0.4.0): SSE heartbeat / `olp` Node CLI + `olp doctor` framework / `olp-connect` zero-config LAN setup / `/health.anonymousKey` opt-in / `/olp` Telegram-Discord plugin / 6-IDE integration docs. Phase 5 scope is open — candidates per ADR 0010 § Out-of-Phase-4-scope: `/v1/messages` (gated on ADR 0009 P0 outcome + named family CC user), context-window-exceeded fallback trigger, per-(provider, model) live stats. Sections marked _placeholder_ land alongside the relevant phase of work (see [phase plan](#phase-plan)).
|
||||
|
||||
---
|
||||
|
||||
@@ -252,9 +252,9 @@ Use a dedicated bot key — not the maintainer's personal owner key — so revoc
|
||||
|
||||
---
|
||||
|
||||
## Implementation status (as of 2026-05-25, post-v0.2.0)
|
||||
## Implementation status (as of 2026-05-26, post-v0.4.0)
|
||||
|
||||
Phase 1 closed at v0.1.1 (multi-provider proxy core + pre-Phase-2 cleanup). Phase 2 closed at v0.2.0 (multi-key auth + audit + owner gating + keygen CLI; ADR 0007 § 10 all 11 acceptance criteria shipped). Phase 3 closed at v0.3.0 (Dashboard + `lib/audit-query.mjs` + daily audit rotation; ADR 0008 § 10 all 15 acceptance criteria shipped). Phase 4 (per-key per-provider auth + audit retention + SQLite hybrid + provider-cost weights) is the next planned milestone. This table reflects what is currently shipped vs. what is designed for later phases.
|
||||
Phase 1 closed at v0.1.1 (multi-provider proxy core + pre-Phase-2 cleanup). Phase 2 closed at v0.2.0 (multi-key auth + audit + owner gating + keygen CLI; ADR 0007 § 10 all 11 acceptance criteria shipped). Phase 3 closed at v0.3.0 (Dashboard + `lib/audit-query.mjs` + daily audit rotation; ADR 0008 § 10 all 15 acceptance criteria shipped). Phase 4 closed at v0.4.0 (Operator + Client UX per ADR 0010: SSE heartbeat + `recentErrors[20]` + `/v0/management/status` / `olp` Node CLI + `olp doctor` framework + ADR 0002 Amendment 7 / `olp-connect` bash + `/health.anonymousKey` + ADR 0011 / `olp-plugin/` Telegram-Discord + 6-IDE integration docs). Phase 5 scope is open — candidates per ADR 0010 § Out-of-Phase-4-scope. This table reflects what is currently shipped vs. what is designed for later phases.
|
||||
|
||||
| File / artifact | Status | Notes |
|
||||
|---|---|---|
|
||||
@@ -289,7 +289,9 @@ Behaviors that work correctly at personal/family scale but have ratified follow-
|
||||
- **Soft triggers configured but inert.** `routing.soft_triggers` in `~/.olp/config.json` is honored by the engine's evaluation logic but `quotaStatus()` polling is not wired (ADR 0004 Amendment 2). A startup warning fires if the field is non-empty so the inert state is visible.
|
||||
- **Multi-key auth + owner gating + keygen CLI shipped at v0.2.0 (D44 + D45 + D46 + D47).** `lib/keys.mjs` (core), `lib/audit.mjs` (audit), owner-vs-guest `/health` payload trimming + `X-OLP-Fallback-Detail` policy gating, `bin/olp-keys.mjs` (keygen CLI). All 11 ADR 0007 § 10 acceptance criteria covered. v0.2.0 maintainer-merged 2026-05-25.
|
||||
|
||||
- **Phase 3 (Dashboard + audit query layer + rotation) shipped to main (D48-D54); v0.3.0 release pending.** `docs/adr/0008-dashboard-and-audit-query.md` ratified at D48. `lib/audit-query.mjs` (D49) implements the 5-function aggregate query API (in-memory ndjson scan, PII-guarded). 4 new owner-only_block endpoints at D50 (`/dashboard`, `/v0/management/dashboard-data`, `/v0/management/quota`, `/cache/stats`). `dashboard.html` full multi-panel UI at D51 (vanilla HTML+JS+fetch, 30s poll with visibilitychange pause). Daily audit rotation at D52 (synchronous on first append after UTC midnight; `audit-YYYY-MM-DD.ndjson` naming) + optional `bin/olp-audit-rotate.mjs` cron tool. `tried_providers` schema semantic fix at D53 (D45 P2 deferral). Phase 3 close to v0.3.0 is maintainer-triggered per CLAUDE.md `release_kit.phase_close_trigger`.
|
||||
- **Phase 3 (Dashboard + audit query layer + rotation) shipped at v0.3.0 (D48–D54).** `docs/adr/0008-dashboard-and-audit-query.md` + `lib/audit-query.mjs` (D49) + 4 owner-only_block endpoints (D50) + `dashboard.html` (D51) + daily audit rotation (D52) + `tried_providers` schema fix (D53). All 15 ADR 0008 § 10 acceptance criteria covered.
|
||||
|
||||
- **Phase 4 (Operator + Client UX) shipped at v0.4.0 (D60 → D73).** ADR 0010 (charter) + ADR 0011 (anonymous-key trusted-LAN limits) + ADR 0002 Amendment 7 (provider `doctorChecks()` contract). Default `OLP_PORT` 3456 → 4567 so OLP and OCP can co-host. SSE heartbeat (D61) + `recentErrors[20]` + `/v0/management/status` (D62-D63). `bin/olp.mjs` Node CLI + `bin/olp-keys.mjs` + `lib/doctor.mjs` framework with `next_action.ai_executable[]` (D64-D67). `bin/olp-connect` bash zero-config IDE auto-config + opt-in `/health.anonymousKey` (D68-D70). `olp-plugin/` OpenClaw `/olp` Telegram-Discord plugin (read-only, no chat mutations) + 6 IDE integration docs at `docs/integrations/*.md` (D71-D73). Test count 623 → 696.
|
||||
|
||||
**Bootstrap workflow (D47):** for first-run / production setup:
|
||||
|
||||
|
||||
+59
-5
@@ -114,6 +114,35 @@ key_display() {
|
||||
fi
|
||||
}
|
||||
|
||||
# D74 P1-2: validate OLP API key format. Per ADR 0007 § 3, tokens are
|
||||
# `olp_` + 32 random bytes base64url-encoded (43 chars, no padding). This
|
||||
# regex pins the on-the-wire shape so a malformed or hostile `--key` /
|
||||
# server-advertised `anonymousKey` never gets persisted into a shell rc.
|
||||
# Returns 0 on valid, 1 on invalid (with diagnostic to stderr).
|
||||
validate_olp_token() {
|
||||
local k="$1" source="$2"
|
||||
if [[ ! "$k" =~ ^olp_[A-Za-z0-9_-]{43}$ ]]; then
|
||||
log_err "Rejected $source: token format does not match ^olp_[A-Za-z0-9_-]{43}$ (ADR 0007 § 3)."
|
||||
log_err " Got ${#k}-char value starting with '$(echo "$k" | cut -c1-8)...'"
|
||||
log_err " Expected: olp_ followed by 43 base64url chars. Run 'npx olp-keys list' on the server"
|
||||
log_err " to confirm the key format, or have the operator regenerate with 'npx olp-keys keygen'."
|
||||
return 1
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# D74 P1-2: POSIX shell-quote a value before interpolating into a shell rc
|
||||
# write. Wraps in single quotes + escapes embedded single quotes per:
|
||||
# foo'bar → 'foo'\''bar'
|
||||
# Even with the validator above, this is defense-in-depth: any non-token
|
||||
# string that slips through (e.g., environment.d KEY=VALUE writes) MUST be
|
||||
# safe to source. Same helper pattern as lib/doctor.mjs _shellQuote.
|
||||
shell_quote() {
|
||||
local s="$1"
|
||||
# Escape any single quotes: ' → '\''
|
||||
printf "'%s'" "${s//\'/\'\\\'\'}"
|
||||
}
|
||||
|
||||
# Detect Claude Code and print warn-only message. Per ADR 0010 § Out of
|
||||
# Phase 4 scope, OLP does NOT ship /v1/messages and CC is not a supported
|
||||
# client. The user is steered toward Cline + OLP.
|
||||
@@ -297,17 +326,20 @@ append_olp_block() {
|
||||
if $DRY_RUN; then
|
||||
log_change "[dry-run] would append OLP block to $rc_file:"
|
||||
log_change " # OLP LAN (added by olp-connect)"
|
||||
log_change " export OPENAI_BASE_URL=$base_url/v1"
|
||||
[[ -n "$key" ]] && log_change " export OPENAI_API_KEY=$(key_display "$key")"
|
||||
log_change " export OPENAI_BASE_URL=$(shell_quote "$base_url/v1")"
|
||||
[[ -n "$key" ]] && log_change " export OPENAI_API_KEY=$(shell_quote "$(key_display "$key")")"
|
||||
log_change " # /OLP LAN"
|
||||
return 0
|
||||
fi
|
||||
# D74 P1-2: shell-quote values before writing to rc files. Defense-in-depth
|
||||
# alongside validate_olp_token — even if a future code path bypasses the
|
||||
# validator, the rc file remains safe to source.
|
||||
{
|
||||
echo ""
|
||||
echo "# OLP LAN (added by olp-connect)"
|
||||
echo "export OPENAI_BASE_URL=$base_url/v1"
|
||||
echo "export OPENAI_BASE_URL=$(shell_quote "$base_url/v1")"
|
||||
if [[ -n "$key" ]]; then
|
||||
echo "export OPENAI_API_KEY=$key"
|
||||
echo "export OPENAI_API_KEY=$(shell_quote "$key")"
|
||||
fi
|
||||
echo "# /OLP LAN"
|
||||
} >> "$rc_file"
|
||||
@@ -341,6 +373,14 @@ set_system_env() {
|
||||
return 0
|
||||
fi
|
||||
mkdir -p "$env_dir" 2>/dev/null
|
||||
# D74 P1-2: systemd environment.d format is KEY=VALUE per line. While
|
||||
# systemd does its own parsing (no shell sourcing), reject embedded
|
||||
# newlines defensively — validate_olp_token already enforces the
|
||||
# restricted charset for the API key, so this is belt-and-braces.
|
||||
if [[ "$base_url" == *$'\n'* || "$key" == *$'\n'* ]]; then
|
||||
log_err "Refusing to write environment.d entry: value contains newline."
|
||||
return 2
|
||||
fi
|
||||
{
|
||||
echo "OPENAI_BASE_URL=$base_url/v1"
|
||||
if [[ -n "$key" ]]; then
|
||||
@@ -363,8 +403,13 @@ main() {
|
||||
--port=*) port="${1#*=}"; shift ;;
|
||||
--key) key="${2:?--key requires a value}"
|
||||
[[ -z "$key" ]] && { log_err "--key cannot be empty (omit --key for zero-config / auto-discovery)"; exit 1; }
|
||||
# D74 P1-2: reject malformed --key before it ever reaches an rc write.
|
||||
validate_olp_token "$key" "--key flag" || exit 1
|
||||
shift 2 ;;
|
||||
--key=*) key="${1#*=}"; shift ;;
|
||||
--key=*) key="${1#*=}"
|
||||
[[ -z "$key" ]] && { log_err "--key cannot be empty (omit --key for zero-config / auto-discovery)"; exit 1; }
|
||||
validate_olp_token "$key" "--key flag" || exit 1
|
||||
shift ;;
|
||||
--no-system-env) NO_SYSTEM_ENV=true; shift ;;
|
||||
--dry-run) DRY_RUN=true; shift ;;
|
||||
--version) show_version; exit 0 ;;
|
||||
@@ -457,6 +502,13 @@ try:
|
||||
print(k if isinstance(k, str) and k else '')
|
||||
except: print('')" 2>/dev/null || echo "")
|
||||
if [[ -n "$anon_key" ]]; then
|
||||
# D74 P1-2: validate server-advertised token shape before consuming.
|
||||
# A hostile or misconfigured server could otherwise inject arbitrary
|
||||
# strings into the user's rc file via the `anonymousKey` field.
|
||||
if ! validate_olp_token "$anon_key" "/health.anonymousKey from $remote_host"; then
|
||||
log_err "Refusing to consume malformed advertised key. Use --key explicitly or contact the OLP operator."
|
||||
exit 2
|
||||
fi
|
||||
key="$anon_key"
|
||||
log_ok "Using server-advertised anonymous key: $(key_display "$key")"
|
||||
log_info " (set by remote via auth.advertise_anonymous_key=true; see ADR 0011 for"
|
||||
@@ -479,6 +531,8 @@ except: print('')" 2>/dev/null || echo "")
|
||||
log_err " Re-run with: olp-connect $host --key olp_..."
|
||||
exit 2
|
||||
fi
|
||||
# D74 P1-2: also validate the interactively-prompted key.
|
||||
validate_olp_token "$key" "interactive prompt" || exit 1
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
+58
-18
@@ -251,9 +251,15 @@ async function cmdStatus(flags, io) {
|
||||
}
|
||||
io.log(` total reqs: ${body.stats?.total_requests ?? 0}`);
|
||||
io.log(` active reqs: ${body.stats?.active_requests ?? 0}`);
|
||||
// D75 F4 fix: server payload nests cache stats under stats.cache (per
|
||||
// server.mjs handleManagementStatus, ~line 2092). The CacheStore.stats()
|
||||
// contract returns { hits, misses, size, inflightCount } per
|
||||
// lib/cache/store.mjs — there is no `entries` field. Pre-D75 cmdStatus read
|
||||
// `body.stats.cache.entries` (OCP-era) which was always undefined → output
|
||||
// showed "entries=?". Same pattern as D74 P2-3 applied to cmdCache/cmdUsage.
|
||||
if (body.stats?.cache) {
|
||||
const c = body.stats.cache;
|
||||
io.log(` cache: hits=${c.hits ?? 0} misses=${c.misses ?? 0} entries=${c.entries ?? '?'}`);
|
||||
io.log(` cache: hits=${c.hits ?? 0} misses=${c.misses ?? 0} entries=${c.size ?? 0}${typeof c.inflightCount === 'number' ? ` inflight=${c.inflightCount}` : ''}`);
|
||||
}
|
||||
if (Array.isArray(body.recent_errors) && body.recent_errors.length > 0) {
|
||||
io.log(` recent errors: ${body.recent_errors.length}`);
|
||||
@@ -301,29 +307,50 @@ async function cmdUsage(flags, io) {
|
||||
try { body = JSON.parse(res.body); }
|
||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||
// D74 P2-3 fix: server payload shape is { generated_at, window_24h: { request_count, status_2xx,
|
||||
// status_4xx, status_5xx, by_provider, by_owner_tier, by_path, median_latency_ms, p95_latency_ms },
|
||||
// cache_hit_24h: { total, hit, miss, bypass, streaming_attached, hit_rate, by_provider }, quota: [{provider, ...}],
|
||||
// spend_trend_30d: [{date, request_count, by_provider}], top_fallback_chains_24h: [{chain, count, ...}],
|
||||
// cache_stats: { hits, misses, size, inflightCount } } per server.mjs:2027 + lib/audit-query.mjs.
|
||||
io.log(colorize('OLP usage (24h)', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
const u24 = body.usage_24h ?? body.usage24h ?? body['24h'] ?? {};
|
||||
if (typeof u24 === 'object' && Object.keys(u24).length > 0) {
|
||||
io.log(` requests: ${u24.requests ?? '?'}`);
|
||||
io.log(` cache hits: ${u24.cache_hits ?? '?'}`);
|
||||
io.log(` fallbacks: ${u24.fallbacks ?? '?'}`);
|
||||
const w24 = body.window_24h ?? {};
|
||||
const cache24 = body.cache_hit_24h ?? {};
|
||||
if (typeof w24 === 'object' && (w24.request_count ?? 0) > 0) {
|
||||
io.log(` requests: ${w24.request_count}`);
|
||||
io.log(` 2xx / 4xx / 5xx: ${w24.status_2xx ?? 0} / ${w24.status_4xx ?? 0} / ${w24.status_5xx ?? 0}`);
|
||||
if (typeof w24.median_latency_ms === 'number') {
|
||||
io.log(` latency p50/p95: ${w24.median_latency_ms}ms / ${w24.p95_latency_ms ?? 0}ms`);
|
||||
}
|
||||
if (typeof cache24.hit_rate === 'number') {
|
||||
const pct = (cache24.hit_rate * 100).toFixed(1);
|
||||
io.log(` cache hit rate: ${pct}% (hit=${cache24.hit ?? 0} miss=${cache24.miss ?? 0}${cache24.streaming_attached ? ` streaming_attached=${cache24.streaming_attached}` : ''})`);
|
||||
}
|
||||
} else {
|
||||
io.log(' (no 24h usage data — server may not have processed any requests yet)');
|
||||
}
|
||||
if (Array.isArray(body.providers)) {
|
||||
if (Array.isArray(body.quota) && body.quota.length > 0) {
|
||||
io.log('');
|
||||
io.log(colorize('Per-provider quota', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const p of body.providers) {
|
||||
io.log(` ${String(p.name ?? '?').padEnd(12)} ${p.percent_used != null ? `${p.percent_used}% used` : 'no quota api'}`);
|
||||
for (const p of body.quota) {
|
||||
const label = String(p.provider ?? '?').padEnd(12);
|
||||
if (p.error) {
|
||||
io.log(` ${label} error: ${p.error}`);
|
||||
} else if (typeof p.percent_used === 'number') {
|
||||
io.log(` ${label} ${p.percent_used}% used${p.resets_in_human ? ` (resets in ${p.resets_in_human})` : ''}`);
|
||||
} else if (p.available === false) {
|
||||
io.log(` ${label} unavailable`);
|
||||
} else {
|
||||
io.log(` ${label} no quota api`);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (Array.isArray(body.top_fallback_chains)) {
|
||||
if (Array.isArray(body.top_fallback_chains_24h) && body.top_fallback_chains_24h.length > 0) {
|
||||
io.log('');
|
||||
io.log(colorize('Top fallback chains', ANSI.bold, io.useColor));
|
||||
io.log(colorize('Top fallback chains (24h)', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
for (const f of body.top_fallback_chains.slice(0, 10)) {
|
||||
for (const f of body.top_fallback_chains_24h.slice(0, 10)) {
|
||||
io.log(` ${String(f.count ?? '?').padStart(5)} ${(f.chain ?? []).join(' → ')}`);
|
||||
}
|
||||
}
|
||||
@@ -360,13 +387,20 @@ async function cmdCache(flags, io) {
|
||||
try { body = JSON.parse(res.body); }
|
||||
catch { io.errln('Error: server returned non-JSON body'); return 2; }
|
||||
if (io.wantJson) { io.emitJson(body); return 0; }
|
||||
io.log(colorize('OLP cache', ANSI.bold, io.useColor));
|
||||
// D74 P2-3 fix: cacheStore.stats() returns { hits, misses, size, inflightCount }
|
||||
// per lib/cache/store.mjs:320. There is no entries / evictions / bytes / maxBytes
|
||||
// in the OLP cache model — those were OCP-era field names. Compute a hit rate from
|
||||
// the numerator/denominator instead of fabricating bytes.
|
||||
const hits = body.hits ?? 0;
|
||||
const misses = body.misses ?? 0;
|
||||
const denom = hits + misses;
|
||||
const hitRate = denom > 0 ? ((hits / denom) * 100).toFixed(1) : '0.0';
|
||||
io.log(colorize('OLP cache (live in-memory)', ANSI.bold, io.useColor));
|
||||
io.log('─'.repeat(60));
|
||||
io.log(` entries: ${body.entries ?? '?'}`);
|
||||
io.log(` hits: ${body.hits ?? 0}`);
|
||||
io.log(` misses: ${body.misses ?? 0}`);
|
||||
io.log(` evictions:${body.evictions ?? 0}`);
|
||||
io.log(` bytes: ${formatBytes(body.bytes ?? 0)} (max ${formatBytes(body.maxBytes ?? 0)})`);
|
||||
io.log(` entries: ${body.size ?? 0}`);
|
||||
io.log(` hits / misses: ${hits} / ${misses} (hit rate ${hitRate}%)`);
|
||||
io.log(` inflight: ${body.inflightCount ?? 0}`);
|
||||
if (body.generated_at) io.log(` generated_at: ${body.generated_at}`);
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -574,10 +608,16 @@ async function cmdDoctor(flags, io) {
|
||||
const proxyUrl = resolveProxyUrl({ proxyUrl: flags['proxy-url'] });
|
||||
const checkFilter = typeof flags.check === 'string' ? flags.check : undefined;
|
||||
|
||||
// D74 P1-1: pass authHeaders so server.running / server.version checks
|
||||
// succeed under the default production posture (auth.allow_anonymous:
|
||||
// false). resolveBearerToken returns null when no env var is set; the
|
||||
// doctor still runs but distinguishes 401 from "server down" by status
|
||||
// code per the updated check.
|
||||
const result = await runDoctor({
|
||||
olpHome,
|
||||
proxyUrl,
|
||||
checkFilter,
|
||||
authHeaders: authHeaders(),
|
||||
});
|
||||
|
||||
if (io.wantJson) {
|
||||
|
||||
+33
-2
@@ -144,6 +144,15 @@ export function buildBuiltinChecks(opts = {}) {
|
||||
const olpHome = resolveOlpHome(opts);
|
||||
const configPath = join(olpHome, 'config.json');
|
||||
const proxyUrl = resolveProxyUrl(opts);
|
||||
// D74 P1-1 fix: server.running and server.version probe /health, which
|
||||
// under default production posture (auth.allow_anonymous: false) requires
|
||||
// an Authorization: Bearer header. Without it, the probe gets 401 and
|
||||
// doctor falsely reports the server down. Caller passes the resolved
|
||||
// bearer token via opts.authHeaders (a `{Authorization: 'Bearer ...'}`
|
||||
// object). Empty headers means "no token configured" — the probe still
|
||||
// fires but a 401 response is treated as "auth misconfigured" rather
|
||||
// than "server down" (see server.running check below).
|
||||
const authHeaders = opts.authHeaders ?? {};
|
||||
|
||||
const checks = [];
|
||||
|
||||
@@ -298,7 +307,9 @@ export function buildBuiltinChecks(opts = {}) {
|
||||
id: 'server.running',
|
||||
category: 'server',
|
||||
async run() {
|
||||
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000 });
|
||||
// D74 P1-1: pass authHeaders so the probe works under the default
|
||||
// production posture (auth.allow_anonymous: false).
|
||||
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000, headers: authHeaders });
|
||||
if (!r.ok) {
|
||||
return {
|
||||
status: 'fail',
|
||||
@@ -311,6 +322,25 @@ export function buildBuiltinChecks(opts = {}) {
|
||||
},
|
||||
};
|
||||
}
|
||||
// 401: server is up but the caller has no/wrong bearer token. NOT a
|
||||
// "server down" condition — distinguish so the kind discriminator
|
||||
// doesn't route to fix_server when the user just needs OLP_API_KEY.
|
||||
if (r.status === 401 || r.status === 403) {
|
||||
return {
|
||||
status: 'fail',
|
||||
message: `${proxyUrl}/health returned ${r.status} — server is up but the bearer token is missing or invalid. Set OLP_API_KEY env to an owner-tier token (npx olp-keys list).`,
|
||||
evidence: {
|
||||
fix_commands: [
|
||||
'echo "set OLP_API_KEY=<your owner token> or OLP_OWNER_TOKEN=<...> then rerun olp doctor"',
|
||||
],
|
||||
human_required: [
|
||||
'Locate an owner-tier OLP API key plaintext (or run `npx olp-keys keygen --owner` to mint a new one — printed ONCE).',
|
||||
'Export it: `export OLP_API_KEY=olp_...`',
|
||||
],
|
||||
reference: 'docs/adr/0007-multi-key-auth.md § 9.1 + README § Environment Variables',
|
||||
},
|
||||
};
|
||||
}
|
||||
if (r.status !== 200) {
|
||||
return { status: 'fail', message: `${proxyUrl}/health returned status=${r.status}` };
|
||||
}
|
||||
@@ -333,7 +363,8 @@ export function buildBuiltinChecks(opts = {}) {
|
||||
} catch {
|
||||
return { status: 'warn', message: 'Could not read local package.json — skipping version comparison' };
|
||||
}
|
||||
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000 });
|
||||
// D74 P1-1: same auth-headers fix as server.running.
|
||||
const r = await httpGet(`${proxyUrl}/health`, { timeoutMs: 3000, headers: authHeaders });
|
||||
if (!r.ok || r.status !== 200) {
|
||||
return { status: 'warn', message: `Could not fetch /health to compare version (${r.error ?? `status ${r.status}`})` };
|
||||
}
|
||||
|
||||
+108
-3
@@ -157,6 +157,36 @@ function resolveCodexBin() {
|
||||
// D6 assumption A2: auth file is named auth.json (unconfirmed — D7 will pin).
|
||||
// D6 assumption A3: access token field is `access_token` or `token` (unconfirmed).
|
||||
//
|
||||
// ── D75 (v0.4.2) F1 — codex CLI v0.133.0 schema pin ─────────────────────────
|
||||
//
|
||||
// Real codex CLI v0.133.0 auth.json (verified empirically on PI231 / Mac mini,
|
||||
// 2026-05-26 E2E session):
|
||||
//
|
||||
// {
|
||||
// "auth_mode": "chatgpt",
|
||||
// "OPENAI_API_KEY": null | "<key>",
|
||||
// "tokens": {
|
||||
// "id_token": "<JWT>",
|
||||
// "access_token": "<opaque-or-JWT>", <-- THIS is the access token
|
||||
// "refresh_token": "<opaque>",
|
||||
// "account_id": "<uuid>"
|
||||
// },
|
||||
// "last_refresh": "<ISO8601>"
|
||||
// }
|
||||
//
|
||||
// D6 assumption A3 originally tried `creds.access_token` at TOP level. Under
|
||||
// codex v0.133.0 this field does not exist at the top level → readAuthArtifact()
|
||||
// returned null even when the user had fully completed `codex login`. OLP then
|
||||
// reported "auth artifact missing" via /health and `olp doctor`, and refused
|
||||
// to spawn codex — false negative blocking the entire openai provider.
|
||||
//
|
||||
// Fix: prepend `creds?.tokens?.access_token` to the precedence chain. Keep all
|
||||
// existing fallbacks unchanged so older codex CLI versions (pre-v0.133) and any
|
||||
// future shape variants still resolve.
|
||||
//
|
||||
// Authority pin: codex CLI v0.133.0 source + on-disk auth.json captured during
|
||||
// PI231 E2E. See D75 commit body for the verification transcript.
|
||||
//
|
||||
// Returns { accessToken: string } or null (never throws).
|
||||
export function readAuthArtifact() {
|
||||
// 1. Explicit test override — always takes precedence.
|
||||
@@ -165,7 +195,12 @@ export function readAuthArtifact() {
|
||||
try {
|
||||
const raw = readFileSync(authPathOverride, 'utf8');
|
||||
const creds = JSON.parse(raw);
|
||||
const token = creds?.access_token ?? creds?.token ?? creds?.accessToken;
|
||||
// D75 F1: codex CLI v0.133.0 nests the token under `tokens.access_token`.
|
||||
// Preserve top-level fallbacks for backward / forward compat.
|
||||
const token = creds?.tokens?.access_token
|
||||
?? creds?.access_token
|
||||
?? creds?.token
|
||||
?? creds?.accessToken;
|
||||
if (token && typeof token === 'string') return { accessToken: token };
|
||||
} catch { /* fall through */ }
|
||||
return null; // explicit path set but file missing / malformed
|
||||
@@ -178,8 +213,12 @@ export function readAuthArtifact() {
|
||||
try {
|
||||
const raw = readFileSync(authPath, 'utf8');
|
||||
const creds = JSON.parse(raw);
|
||||
// D6 assumption A3: try common OAuth field names in precedence order.
|
||||
const token = creds?.access_token ?? creds?.token ?? creds?.accessToken;
|
||||
// D75 F1: codex CLI v0.133.0 nests the token under `tokens.access_token`.
|
||||
// Try the nested location FIRST, then fall back to legacy top-level fields.
|
||||
const token = creds?.tokens?.access_token
|
||||
?? creds?.access_token
|
||||
?? creds?.token
|
||||
?? creds?.accessToken;
|
||||
if (token && typeof token === 'string') return { accessToken: token };
|
||||
} catch { /* file missing or malformed */ }
|
||||
|
||||
@@ -242,9 +281,30 @@ export function irToCodex(irRequest) {
|
||||
// model string (e.g., gpt-5.5, gpt-5.4, gpt-5.3-codex).
|
||||
// PROMPT: "Initial instruction for the task. Use '-' to pipe the prompt
|
||||
// from stdin."
|
||||
//
|
||||
// ── D75 (v0.4.2) F2 — codex CLI v0.133.0 trusted-directory sandbox ─────────
|
||||
// codex CLI v0.133.0 added a trusted-directory sandbox: invocations outside
|
||||
// a git repo (or outside any directory explicitly trusted via
|
||||
// `codex config trusted-directories`) refuse with:
|
||||
// "Not inside a trusted directory and --skip-git-repo-check was not specified."
|
||||
// and exit non-zero with zero NDJSON output → OLP surfaces SPAWN_FAILED with
|
||||
// no usable chunks → fallback engine advances to next hop unnecessarily.
|
||||
//
|
||||
// The CWD that OLP spawns from is typically the server install dir (`~/olp/`
|
||||
// on Pi231) which is a git repo on maintainer workstations but is NOT a git
|
||||
// repo on most operator hosts. We bypass the sandbox unconditionally because
|
||||
// OLP is the trusted caller (it is the operator's own server invoking its own
|
||||
// configured Codex subscription via the documented `codex exec` automation
|
||||
// entry point). The trusted-directory sandbox is a foot-gun safeguard for
|
||||
// interactive users; OLP's spawn is non-interactive and pre-authorized.
|
||||
//
|
||||
// Authority: codex CLI v0.133.0 release notes / `codex exec --help` output
|
||||
// documenting `--skip-git-repo-check`. Verified empirically on PI231 E2E
|
||||
// 2026-05-26.
|
||||
const args = [
|
||||
'exec',
|
||||
'--json',
|
||||
'--skip-git-repo-check',
|
||||
'--model', irRequest.model,
|
||||
];
|
||||
|
||||
@@ -291,6 +351,51 @@ export function codexChunkToIR(rawNDJSONLine) {
|
||||
|
||||
if (!event || typeof event !== 'object') return null;
|
||||
|
||||
// ── D75 (v0.4.2) F3 — codex CLI v0.133.0 event shape pin ─────────────────
|
||||
// Real codex CLI v0.133.0 NDJSON event stream (verified empirically on PI231
|
||||
// / Mac mini, 2026-05-26 E2E session):
|
||||
// {"type":"thread.started","thread_id":"019e..."}
|
||||
// {"type":"turn.started"}
|
||||
// {"type":"item.started","item":{"id":"item_0","type":"reasoning","text":""}}
|
||||
// {"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"<response>"}}
|
||||
// {"type":"turn.completed","usage":{"input_tokens":..,"output_tokens":..}}
|
||||
//
|
||||
// The D6 defensive parser recognized `content`/`delta`/`text` fields at the
|
||||
// top level and `type === 'stop'`/`done === true`. None of these match
|
||||
// v0.133.0's actual shape → every chunk was silently dropped → response body
|
||||
// had `content: null`. F3 adds three NEW recognizers (item.completed →
|
||||
// agent_message; turn.completed → stop; turn.failed → error) BEFORE the
|
||||
// legacy fallback chain. Legacy recognizers preserved for forward/backward
|
||||
// compat (older codex versions; future shape variants).
|
||||
|
||||
// F3-a: agent_message item completion.
|
||||
// codex v0.133.0 emits assistant text as a single item.completed event whose
|
||||
// item.type is 'agent_message' and item.text carries the full text. There
|
||||
// are no incremental deltas — the entire response arrives in one chunk.
|
||||
if (event.type === 'item.completed'
|
||||
&& event.item?.type === 'agent_message'
|
||||
&& typeof event.item?.text === 'string') {
|
||||
return { type: 'delta', content: event.item.text };
|
||||
}
|
||||
|
||||
// F3-b: turn completion → stop chunk.
|
||||
// codex v0.133.0 emits turn.completed with a usage block when the model
|
||||
// finishes. We map this to IR stop with finish_reason 'stop'.
|
||||
if (event.type === 'turn.completed') {
|
||||
return { type: 'stop', finish_reason: 'stop' };
|
||||
}
|
||||
|
||||
// F3-c: turn failure → error chunk.
|
||||
// codex v0.133.0 emits turn.failed with an embedded error object when the
|
||||
// turn cannot complete. Extract a human-readable message for the IR error.
|
||||
if (event.type === 'turn.failed') {
|
||||
const errMsg = (typeof event.error === 'string')
|
||||
? event.error
|
||||
: (event.error?.message ?? 'codex turn.failed');
|
||||
return { type: 'error', error: errMsg };
|
||||
}
|
||||
|
||||
// ── Legacy/fallback recognizers (kept for backward + forward compat) ────
|
||||
// Error event: type === 'error' or error field present
|
||||
// A4: defensive — error shape unconfirmed; D7 will pin actual field names
|
||||
if (event.type === 'error' || (event.error && typeof event.error === 'string')) {
|
||||
|
||||
+25
-5
@@ -138,12 +138,32 @@ export function fmtHealth(body) {
|
||||
if (body.uptime_human || body.uptimeHuman) {
|
||||
out += `Uptime: ${body.uptime_human ?? body.uptimeHuman}\n`;
|
||||
}
|
||||
// D74 P2-4 fix: server.mjs /health full payload is
|
||||
// body.providers = { enabled: N, available: N, status: { <name>: {...} } }
|
||||
// The plugin previously iterated Object.entries(body.providers), which
|
||||
// surfaced `enabled`, `available`, and `status` as pseudo-providers
|
||||
// (typeof status === 'object' → loop body fired with name='status').
|
||||
// Walk providers.status when present; fall back to providers.* for the
|
||||
// older OCP shape that lacks the .status wrapper.
|
||||
if (body.providers && typeof body.providers === "object") {
|
||||
out += `\nProviders:\n`;
|
||||
for (const [name, s] of Object.entries(body.providers)) {
|
||||
if (typeof s !== "object" || s === null) continue;
|
||||
const i = statusIcon(s?.ok ? "ok" : "fail");
|
||||
out += ` ${i} ${name}\n`;
|
||||
const enabled = body.providers.enabled;
|
||||
const available = body.providers.available;
|
||||
if (typeof enabled === "number" || typeof available === "number") {
|
||||
out += `Providers: ${enabled ?? "?"} enabled / ${available ?? "?"} available\n`;
|
||||
}
|
||||
const statusMap = body.providers.status && typeof body.providers.status === "object"
|
||||
? body.providers.status
|
||||
: body.providers;
|
||||
const entries = Object.entries(statusMap).filter(
|
||||
([name, s]) => typeof s === "object" && s !== null && name !== "enabled" && name !== "available" && name !== "status"
|
||||
);
|
||||
if (entries.length > 0) {
|
||||
out += `\nProviders:\n`;
|
||||
for (const [name, s] of entries) {
|
||||
const i = statusIcon(s?.ok ? "ok" : "fail");
|
||||
const spawn = typeof s?.activeSpawns === "number" ? ` spawns=${s.activeSpawns}` : "";
|
||||
out += ` ${i} ${name}${spawn}\n`;
|
||||
}
|
||||
}
|
||||
}
|
||||
return out;
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "olp",
|
||||
"version": "0.4.0",
|
||||
"version": "0.4.2",
|
||||
"description": "Personal multi-provider LLM proxy. Successor to OCP. One HTTP endpoint, multiple subscriptions behind it, automatic routing + fallback + caching.",
|
||||
"type": "module",
|
||||
"main": "server.mjs",
|
||||
|
||||
+29
-7
@@ -1224,13 +1224,25 @@ async function handleChatCompletions(req, res) {
|
||||
}
|
||||
|
||||
const chunks = [];
|
||||
// try/finally: releaseSpawn MUST fire on every exit path — success
|
||||
// (return at end), spawn throw (caught and re-thrown below), or the
|
||||
// D16 truncation-salvage return. The finally is the only mechanism
|
||||
// that guarantees release across all three.
|
||||
// D75 (v0.4.2) F7 — per-hop model override.
|
||||
// The chain config (routing.chains[<requested-model>][<hop>].model) is the
|
||||
// model name to pass to THIS hop's provider plugin — NOT the model the
|
||||
// user originally requested. Pre-D75, executeHopFn used hopModel for the
|
||||
// cache key + audit ctx but passed the ORIGINAL irReq (with irReq.model
|
||||
// still set to the user's request) into hopProviderPlugin.spawn(). Result:
|
||||
// a 2-hop chain like [{anthropic, claude-sonnet-4-6}, {openai, gpt-5.5}]
|
||||
// would spawn codex with --model claude-sonnet-4-6 on hop 1 — openai rejects
|
||||
// the unknown model and the chain dies. This broke the core OLP value prop
|
||||
// (cross-provider fallback with provider-appropriate model substitution).
|
||||
//
|
||||
// Fix: build a per-hop IR variant with the hop's model substituted. Skip
|
||||
// the clone when hopModel === irReq.model (single-provider chains and
|
||||
// chain hops whose model matches the request). Authority: ADR 0004 §
|
||||
// Chain advancement step 1 (per-hop config supplies provider AND model).
|
||||
const hopIrReq = irReq.model === hopModel ? irReq : { ...irReq, model: hopModel };
|
||||
try {
|
||||
try {
|
||||
for await (const irChunk of hopProviderPlugin.spawn(irReq, authContext)) {
|
||||
for await (const irChunk of hopProviderPlugin.spawn(hopIrReq, authContext)) {
|
||||
// D16: check error chunks BEFORE pushing — preserves the invariant that
|
||||
// chunks array contains only delta/stop chunks. Without this, the catch
|
||||
// block's `chunks.length > 0` would mistake a single error chunk for
|
||||
@@ -1422,9 +1434,16 @@ async function handleChatCompletions(req, res) {
|
||||
// releaseSpawn fires exactly once in finally — regardless of normal
|
||||
// exhaustion, mid-stream throw, or iterator.return() from cache-layer
|
||||
// sourceAbortController propagation (§9).
|
||||
//
|
||||
// D75 (v0.4.2) F7 — per-hop model override (streaming path).
|
||||
// Mirror the buffered-path fix: pass the hop's configured model (streamModel)
|
||||
// into streamPlugin.spawn(), not the user's original ir.model. Skip the
|
||||
// clone when ir.model === streamModel. See executeHopFn() above for the
|
||||
// full F7 rationale + authority citation.
|
||||
const streamIr = ir.model === streamModel ? ir : { ...ir, model: streamModel };
|
||||
return (async function* sourceWithRelease() {
|
||||
try {
|
||||
for await (const irChunk of streamPlugin.spawn(ir, authContext)) {
|
||||
for await (const irChunk of streamPlugin.spawn(streamIr, authContext)) {
|
||||
yield irChunk;
|
||||
}
|
||||
} finally {
|
||||
@@ -2241,8 +2260,11 @@ if (isMain) {
|
||||
const server = createOlpServer();
|
||||
server.listen(PORT, '127.0.0.1', () => {
|
||||
const enabledCount = loadedProviders.size;
|
||||
// D74 P3-5: banner no longer hardcodes the phase. Derives from VERSION
|
||||
// (which advances at every Phase close) so banner stays accurate
|
||||
// without future-maintenance touch-ups at every Phase boundary.
|
||||
process.stdout.write(
|
||||
`OLP v${VERSION} listening on :${PORT} (${enabledCount} providers enabled — Phase 1 in progress)\n`,
|
||||
`OLP v${VERSION} listening on :${PORT} (${enabledCount} providers enabled)\n`,
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
@@ -14791,3 +14791,574 @@ describe('Suite 35 — D71 olp-plugin/ smoke tests (ADR 0010 § Phase 4 D71-D73)
|
||||
assert.match(r.text, /not owner-tier/);
|
||||
});
|
||||
});
|
||||
|
||||
// ── Suite 36: D74 v0.4.1 hotfix regression tests (maintainer-review findings) ─
|
||||
//
|
||||
// These tests pin the FIVE issues maintainer caught in independent post-v0.4.0
|
||||
// review that previous D-day reviewers all missed (because they reviewed against
|
||||
// spec text, not runtime contracts). Each test exercises a real codepath that,
|
||||
// pre-D74, would have produced the documented wrong behavior.
|
||||
|
||||
import { writeFileSync as _writeFileSyncS36, readFileSync as _readFileSyncS36, mkdtempSync as _mkdtempSyncS36, rmSync as _rmSyncS36 } from 'node:fs';
|
||||
import { tmpdir as _tmpdirS36 } from 'node:os';
|
||||
import { join as _joinS36 } from 'node:path';
|
||||
import { spawnSync as _spawnSyncS36 } from 'node:child_process';
|
||||
|
||||
describe('Suite 36 — D74 v0.4.1 hotfix regression (maintainer review findings)', () => {
|
||||
it('36a (P1-1) — runDoctor accepts authHeaders + passes them to server.running probe', async () => {
|
||||
// Pre-D74: lib/doctor.mjs called httpGet(`${proxyUrl}/health`) with no
|
||||
// headers, so under default `auth.allow_anonymous: false` the probe got
|
||||
// 401 and reported "server down" — false negative. D74 threads
|
||||
// authHeaders through buildBuiltinChecks. This test confirms the wire.
|
||||
const { buildBuiltinChecks } = await import('./lib/doctor.mjs');
|
||||
const tmpHome = _mkdtempSyncS36(_joinS36(_tmpdirS36(), 'olp-d74a-'));
|
||||
const checks = buildBuiltinChecks({
|
||||
olpHome: tmpHome,
|
||||
proxyUrl: 'http://127.0.0.1:1', // closed port — probe fails with ECONN
|
||||
authHeaders: { Authorization: 'Bearer olp_FAKE' },
|
||||
});
|
||||
const serverRunning = checks.find(c => c.id === 'server.running');
|
||||
assert.ok(serverRunning, 'server.running check must exist');
|
||||
// We can't easily intercept httpGet here without DI, but the check at
|
||||
// least accepts the opts and runs. Real wire-level coverage is 36b below.
|
||||
_rmSyncS36(tmpHome, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
it('36b (P1-1) — server.running distinguishes 401 (server up, bad token) from "server down"', async () => {
|
||||
__setAuthConfig({ allow_anonymous: false, owner_only_endpoints: ['/health'] });
|
||||
const serverMod = await import('./server.mjs');
|
||||
const s = serverMod.createOlpServer();
|
||||
let port36b = 28300 + Math.floor(Math.random() * 100);
|
||||
await new Promise((resolve, reject) => {
|
||||
s.listen(port36b, '127.0.0.1', resolve);
|
||||
s.once('error', e => {
|
||||
if (e.code === 'EADDRINUSE') { port36b++; s.listen(port36b, '127.0.0.1', resolve); s.once('error', reject); }
|
||||
else reject(e);
|
||||
});
|
||||
});
|
||||
|
||||
const { runDoctor } = await import('./lib/doctor.mjs');
|
||||
try {
|
||||
const result = await runDoctor({
|
||||
olpHome: process.env.OLP_HOME,
|
||||
proxyUrl: `http://127.0.0.1:${port36b}`,
|
||||
// intentionally NO authHeaders — simulates user without OLP_API_KEY
|
||||
});
|
||||
const sr = result.checks.find(c => c.id === 'server.running');
|
||||
assert.ok(sr, 'server.running present');
|
||||
assert.equal(sr.status, 'fail', 'fails because no token');
|
||||
assert.match(sr.message, /401|bearer token/i,
|
||||
`must mention 401 / bearer token specifically — got: ${sr.message}`);
|
||||
assert.ok(!/unreachable/i.test(sr.message),
|
||||
'"unreachable" would falsely imply server is down (P1-1 bug)');
|
||||
} finally {
|
||||
await new Promise(r => s.close(r));
|
||||
__resetAuthConfig();
|
||||
}
|
||||
});
|
||||
|
||||
it('36c (P1-2) — olp-connect rejects malformed --key', () => {
|
||||
// The bash validator: ^olp_[A-Za-z0-9_-]{43}$. Anything else is rejected
|
||||
// before it can touch a shell rc file or environment.d entry.
|
||||
const result = _spawnSyncS36('bash', [
|
||||
_joinS36(import.meta.dirname ?? process.cwd(), 'bin/olp-connect'),
|
||||
'127.0.0.1',
|
||||
'--port', '1',
|
||||
'--key', 'not-an-olp-token',
|
||||
'--dry-run',
|
||||
], { encoding: 'utf8', timeout: 5000 });
|
||||
assert.equal(result.status, 1, `expected exit 1 for malformed --key, got ${result.status}; stderr=${result.stderr.slice(0, 300)}`);
|
||||
assert.match(result.stderr + result.stdout, /token format does not match/i,
|
||||
'must surface the validator error');
|
||||
});
|
||||
|
||||
it('36d (P1-2) — olp-connect accepts a properly-formed olp_ token', () => {
|
||||
// Pad to 43 base64url chars after olp_. Don't actually hit the server —
|
||||
// use --port 1 (closed) which makes the connectivity probe fail at
|
||||
// the connect step (exit 2), but only AFTER the --key validator passes.
|
||||
const goodKey = 'olp_' + 'A'.repeat(43);
|
||||
const result = _spawnSyncS36('bash', [
|
||||
_joinS36(import.meta.dirname ?? process.cwd(), 'bin/olp-connect'),
|
||||
'127.0.0.1',
|
||||
'--port', '1',
|
||||
'--key', goodKey,
|
||||
'--dry-run',
|
||||
], { encoding: 'utf8', timeout: 5000 });
|
||||
// Validator passes → reaches connectivity step → exits 2 (or whatever
|
||||
// bash propagation gives) but NOT exit 1 from the validator.
|
||||
assert.notEqual(result.status, 1, `expected validator to pass; got exit 1 with stderr=${result.stderr.slice(0, 300)}`);
|
||||
assert.ok(!/token format does not match/i.test(result.stderr + result.stdout),
|
||||
'validator must NOT trigger for a well-formed token');
|
||||
});
|
||||
|
||||
it('36e (P2-3) — CacheStore.stats() returns {hits, misses, size, inflightCount} shape', async () => {
|
||||
// Real cacheStore.stats() returns { hits, misses, size, inflightCount } per
|
||||
// lib/cache/store.mjs. Pre-D74 cmdCache read body.entries / body.bytes /
|
||||
// body.maxBytes — all undefined → output showed "entries: ?" and "bytes: 0".
|
||||
// Pin the shape so a future server-side rename can't reintroduce the bug.
|
||||
// Use the CacheStore class directly (no HTTP) since the field names ARE the
|
||||
// contract — server.mjs handleCacheStats just sendJSON(s, 200, cacheStore.stats()).
|
||||
const { CacheStore: CS36 } = await import('./lib/cache/store.mjs');
|
||||
const cs = new CS36();
|
||||
const stats = cs.stats();
|
||||
assert.ok('size' in stats, 'CacheStore.stats() shape: size field present');
|
||||
assert.ok('hits' in stats, 'CacheStore.stats() shape: hits field present');
|
||||
assert.ok('misses' in stats, 'CacheStore.stats() shape: misses field present');
|
||||
assert.ok('inflightCount' in stats, 'CacheStore.stats() shape: inflightCount field present');
|
||||
assert.ok(!('entries' in stats), 'must NOT have OCP-era entries field');
|
||||
assert.ok(!('bytes' in stats), 'must NOT have OCP-era bytes field');
|
||||
assert.ok(!('maxBytes' in stats), 'must NOT have OCP-era maxBytes field');
|
||||
// Also pin that cmdCache reads body.size (not body.entries) by grepping the source
|
||||
const cliSrc = _readFileSyncS36(_joinS36(import.meta.dirname ?? process.cwd(), 'bin/olp.mjs'), 'utf8');
|
||||
assert.ok(/body\.size\b/.test(cliSrc), 'cmdCache must read body.size for entries display');
|
||||
assert.ok(/body\.inflightCount\b/.test(cliSrc), 'cmdCache must read body.inflightCount');
|
||||
assert.ok(!/body\.entries\b/.test(cliSrc), 'cmdCache must NOT read body.entries (OCP-era field)');
|
||||
});
|
||||
|
||||
it('36f (P2-3) — dashboard-data payload + cmdUsage source pin the wire-contract field names', () => {
|
||||
// server.mjs handleManagementDashboardData (~line 2027) builds the payload
|
||||
// inline. Pin by grepping the server source: the keys MUST be window_24h /
|
||||
// cache_hit_24h / quota / spend_trend_30d / top_fallback_chains_24h /
|
||||
// cache_stats. Pre-D74 cmdUsage read body.usage_24h.requests / body.providers
|
||||
// / body.top_fallback_chains — none of which exist in the actual payload.
|
||||
const serverSrc = _readFileSyncS36(_joinS36(import.meta.dirname ?? process.cwd(), 'server.mjs'), 'utf8');
|
||||
assert.ok(/window_24h:\s*auditAggregateRequests/.test(serverSrc),
|
||||
'handleManagementDashboardData must emit window_24h field');
|
||||
assert.ok(/cache_hit_24h:\s*auditCacheHitRateWindow/.test(serverSrc),
|
||||
'handleManagementDashboardData must emit cache_hit_24h field');
|
||||
assert.ok(/top_fallback_chains_24h:/.test(serverSrc),
|
||||
'handleManagementDashboardData must emit top_fallback_chains_24h field (NOT top_fallback_chains)');
|
||||
assert.ok(/cache_stats:\s*cacheStore\.stats/.test(serverSrc),
|
||||
'handleManagementDashboardData must emit cache_stats field');
|
||||
// Now pin cmdUsage reads the matching keys
|
||||
const cliSrc = _readFileSyncS36(_joinS36(import.meta.dirname ?? process.cwd(), 'bin/olp.mjs'), 'utf8');
|
||||
assert.ok(/body\.window_24h\b/.test(cliSrc), 'cmdUsage must read body.window_24h');
|
||||
assert.ok(/body\.cache_hit_24h\b/.test(cliSrc), 'cmdUsage must read body.cache_hit_24h');
|
||||
assert.ok(/body\.top_fallback_chains_24h\b/.test(cliSrc),
|
||||
'cmdUsage must read body.top_fallback_chains_24h (NOT body.top_fallback_chains)');
|
||||
assert.ok(/body\.quota\b/.test(cliSrc), 'cmdUsage must read body.quota');
|
||||
// Pre-D74 stale reads must be GONE from the CLI source
|
||||
assert.ok(!/body\.usage_24h\.requests\b/.test(cliSrc),
|
||||
'cmdUsage must NOT read body.usage_24h.requests (OCP-era field)');
|
||||
});
|
||||
|
||||
it('36g (P2-4) — olp-plugin fmtHealth iterates providers.status (not providers.*)', async () => {
|
||||
// Real /health full payload: providers: { enabled, available, status: { anthropic: {...} } }
|
||||
// Pre-D74 fmtHealth iterated Object.entries(body.providers) and the body
|
||||
// contained `status: {...}` so the loop printed a pseudo-provider named "status".
|
||||
const { fmtHealth } = await import('./olp-plugin/index.js');
|
||||
const realServerShape = {
|
||||
ok: true,
|
||||
version: '0.4.1',
|
||||
uptime_human: '1m 3s',
|
||||
providers: {
|
||||
enabled: 2,
|
||||
available: 3,
|
||||
status: {
|
||||
anthropic: { ok: true, activeSpawns: 0 },
|
||||
openai: { ok: false, error: 'oauth-missing', activeSpawns: 0 },
|
||||
},
|
||||
},
|
||||
};
|
||||
const out = fmtHealth(realServerShape);
|
||||
assert.match(out, /anthropic/, 'must list anthropic by name');
|
||||
assert.match(out, /openai/, 'must list openai by name');
|
||||
// assert.notMatch isn't available on assert/strict in some Node versions;
|
||||
// use the explicit negation pattern to stay portable.
|
||||
assert.ok(!/^\s*[🟢🔴⚪]\s*status\s*$/m.test(out),
|
||||
'must NOT list "status" as a pseudo-provider (the P2-4 bug)');
|
||||
assert.ok(!/^\s*[🟢🔴⚪]\s*enabled\s*$/m.test(out),
|
||||
'must NOT list "enabled" as a pseudo-provider');
|
||||
assert.ok(!/^\s*[🟢🔴⚪]\s*available\s*$/m.test(out),
|
||||
'must NOT list "available" as a pseudo-provider');
|
||||
// Also confirm the new enabled/available header
|
||||
assert.match(out, /2 enabled \/ 3 available/);
|
||||
});
|
||||
|
||||
it('36h (P3-5) — server startup banner does not hardcode a stale phase', () => {
|
||||
// Pre-D74 banner literal said "Phase 1 in progress" forever.
|
||||
const serverSrc = _readFileSyncS36(_joinS36(import.meta.dirname ?? process.cwd(), 'server.mjs'), 'utf8');
|
||||
assert.ok(!/Phase 1 in progress/.test(serverSrc),
|
||||
'startup banner must not hardcode "Phase 1 in progress" (stale post v0.4.0)');
|
||||
assert.ok(!/Phase 2 in progress/.test(serverSrc));
|
||||
assert.ok(!/Phase 3 in progress/.test(serverSrc));
|
||||
});
|
||||
|
||||
// ── D75 (v0.4.2) extension — real-machine E2E findings F1, F2, F3, F4, F7 ──
|
||||
// Bugs caught on PI231 + Mac mini 2026-05-26 E2E session. Pre-v0.4.0 reviewers
|
||||
// and the post-v0.4.0 maintainer review all missed these because they reviewed
|
||||
// against spec text, not against real provider CLIs running on a remote machine.
|
||||
|
||||
it('36i (F1) — codex readAuthArtifact() recognises v0.133.0 nested tokens.access_token', () => {
|
||||
// Real codex CLI v0.133.0 auth.json shape: { auth_mode, OPENAI_API_KEY,
|
||||
// tokens: { id_token, access_token, refresh_token, account_id }, last_refresh }
|
||||
// Pre-D75 readAuthArtifact read only top-level access_token → returned null →
|
||||
// OLP falsely reported "auth artifact missing" for fully-logged-in users.
|
||||
const tmpDir = _mkdtempSyncS36(_joinS36(_tmpdirS36(), 'olp-d75-f1-nested-'));
|
||||
const authPath = _joinS36(tmpDir, 'auth.json');
|
||||
_writeFileSyncS36(authPath, JSON.stringify({
|
||||
auth_mode: 'chatgpt',
|
||||
OPENAI_API_KEY: null,
|
||||
tokens: {
|
||||
id_token: 'header.body.sig',
|
||||
access_token: 'eyJ-real-nested-access-token-d75-36i',
|
||||
refresh_token: 'refresh-opaque',
|
||||
account_id: '00000000-0000-0000-0000-000000000000',
|
||||
},
|
||||
last_refresh: '2026-05-26T00:00:00Z',
|
||||
}));
|
||||
const savedPath = process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = authPath;
|
||||
try {
|
||||
const result = codexReadAuthArtifact();
|
||||
assert.ok(result, 'must return non-null for v0.133 nested-shape auth.json');
|
||||
assert.equal(result.accessToken, 'eyJ-real-nested-access-token-d75-36i',
|
||||
'must extract token from tokens.access_token (v0.133.0 nested shape)');
|
||||
} finally {
|
||||
if (savedPath !== undefined) process.env.OPENAI_CODEX_AUTH_PATH = savedPath;
|
||||
else delete process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
_rmSyncS36(tmpDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it('36j (F1) — codex readAuthArtifact() still reads legacy top-level access_token', () => {
|
||||
// Backwards compat: older codex CLI versions (pre-v0.133) wrote the token
|
||||
// at top level. D75 fix preserves that fallback so upgrades don't regress.
|
||||
const tmpDir = _mkdtempSyncS36(_joinS36(_tmpdirS36(), 'olp-d75-f1-legacy-'));
|
||||
const authPath = _joinS36(tmpDir, 'auth.json');
|
||||
_writeFileSyncS36(authPath, JSON.stringify({
|
||||
access_token: 'legacy-top-level-token-d75-36j',
|
||||
}));
|
||||
const savedPath = process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = authPath;
|
||||
try {
|
||||
const result = codexReadAuthArtifact();
|
||||
assert.ok(result, 'must return non-null for legacy top-level shape');
|
||||
assert.equal(result.accessToken, 'legacy-top-level-token-d75-36j',
|
||||
'must fall back to top-level access_token when tokens.access_token absent');
|
||||
} finally {
|
||||
if (savedPath !== undefined) process.env.OPENAI_CODEX_AUTH_PATH = savedPath;
|
||||
else delete process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
_rmSyncS36(tmpDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it('36k (F1) — codex readAuthArtifact() returns null for malformed auth.json', () => {
|
||||
// Malformed JSON / missing token / wrong structure → null (not throw).
|
||||
const tmpDir = _mkdtempSyncS36(_joinS36(_tmpdirS36(), 'olp-d75-f1-malformed-'));
|
||||
const savedPath = process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
try {
|
||||
// Case A: unparseable JSON
|
||||
const badJsonPath = _joinS36(tmpDir, 'bad.json');
|
||||
_writeFileSyncS36(badJsonPath, 'not-json-at-all{{{');
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = badJsonPath;
|
||||
assert.equal(codexReadAuthArtifact(), null, 'unparseable JSON → null');
|
||||
|
||||
// Case B: valid JSON but no token field anywhere
|
||||
const noTokenPath = _joinS36(tmpDir, 'no-token.json');
|
||||
_writeFileSyncS36(noTokenPath, JSON.stringify({ auth_mode: 'chatgpt', tokens: {} }));
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = noTokenPath;
|
||||
assert.equal(codexReadAuthArtifact(), null, 'no token in tokens.access_token or top-level → null');
|
||||
|
||||
// Case C: tokens field present but access_token wrong type
|
||||
const wrongTypePath = _joinS36(tmpDir, 'wrong-type.json');
|
||||
_writeFileSyncS36(wrongTypePath, JSON.stringify({
|
||||
tokens: { access_token: 42 },
|
||||
}));
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = wrongTypePath;
|
||||
assert.equal(codexReadAuthArtifact(), null, 'non-string access_token → null');
|
||||
} finally {
|
||||
if (savedPath !== undefined) process.env.OPENAI_CODEX_AUTH_PATH = savedPath;
|
||||
else delete process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
_rmSyncS36(tmpDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it('36l (F2) — irToCodex() includes --skip-git-repo-check before --model', () => {
|
||||
// codex CLI v0.133.0 sandboxes non-git-repo CWDs by default with:
|
||||
// "Not inside a trusted directory and --skip-git-repo-check was not specified."
|
||||
// OLP servers typically deploy outside a git repo on the operator host. The
|
||||
// --skip-git-repo-check flag must be in args, before --model, so the spawn
|
||||
// succeeds without depending on the install-time CWD.
|
||||
const { args } = irToCodex({
|
||||
irVersion: IR_VERSION,
|
||||
model: 'gpt-5.5',
|
||||
stream: false,
|
||||
messages: [{ role: 'user', content: 'hello' }],
|
||||
});
|
||||
assert.ok(args.includes('--skip-git-repo-check'),
|
||||
'irToCodex().args must include --skip-git-repo-check (F2 fix)');
|
||||
const skipIdx = args.indexOf('--skip-git-repo-check');
|
||||
const modelIdx = args.indexOf('--model');
|
||||
assert.ok(skipIdx > 0 && modelIdx > 0, 'both flags must be present');
|
||||
assert.ok(skipIdx < modelIdx,
|
||||
'--skip-git-repo-check must come before --model (cleaner readability + matches codex v0.133 docs ordering)');
|
||||
// Confirm exec is still first and --json still adjacent
|
||||
assert.equal(args[0], 'exec');
|
||||
assert.equal(args[1], '--json');
|
||||
assert.equal(args[2], '--skip-git-repo-check');
|
||||
assert.equal(args[3], '--model');
|
||||
assert.equal(args[4], 'gpt-5.5');
|
||||
});
|
||||
|
||||
it('36m (F3) — codexChunkToIR recognises v0.133.0 item.completed agent_message', () => {
|
||||
// Real codex v0.133.0 emits the assistant text as one item.completed event:
|
||||
// {"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"hello world"}}
|
||||
// Pre-D75 codexChunkToIR returned null for this shape → response body had
|
||||
// content: null because the chunk was silently dropped.
|
||||
const result = codexChunkToIR(JSON.stringify({
|
||||
type: 'item.completed',
|
||||
item: { id: 'item_0', type: 'agent_message', text: 'hello world from codex' },
|
||||
}));
|
||||
assert.deepEqual(result, { type: 'delta', content: 'hello world from codex' });
|
||||
});
|
||||
|
||||
it('36n (F3) — codexChunkToIR recognises v0.133.0 turn.completed → stop', () => {
|
||||
// turn.completed marks the end of a codex turn. Map to IR stop.
|
||||
const result = codexChunkToIR(JSON.stringify({
|
||||
type: 'turn.completed',
|
||||
usage: { input_tokens: 12, output_tokens: 7 },
|
||||
}));
|
||||
assert.deepEqual(result, { type: 'stop', finish_reason: 'stop' });
|
||||
});
|
||||
|
||||
it('36o (F3) — codexChunkToIR recognises v0.133.0 turn.failed → error', () => {
|
||||
// turn.failed: { type: 'turn.failed', error: { message: 'X' } }
|
||||
const result1 = codexChunkToIR(JSON.stringify({
|
||||
type: 'turn.failed',
|
||||
error: { message: 'rate limit exceeded' },
|
||||
}));
|
||||
assert.equal(result1?.type, 'error');
|
||||
assert.equal(result1.error, 'rate limit exceeded');
|
||||
// turn.failed with string error
|
||||
const result2 = codexChunkToIR(JSON.stringify({
|
||||
type: 'turn.failed',
|
||||
error: 'simple string error',
|
||||
}));
|
||||
assert.equal(result2?.type, 'error');
|
||||
assert.equal(result2.error, 'simple string error');
|
||||
// turn.failed with no error payload → falls back to default message
|
||||
const result3 = codexChunkToIR(JSON.stringify({ type: 'turn.failed' }));
|
||||
assert.equal(result3?.type, 'error');
|
||||
assert.match(result3.error, /turn\.failed/);
|
||||
});
|
||||
|
||||
it('36p (F4) — cmdStatus reads body.stats.cache.size (not OCP-era body.cache.entries)', () => {
|
||||
// Server payload (server.mjs handleManagementStatus ~line 2092):
|
||||
// { ok, version, ..., stats: { total_requests, active_requests, cache: { hits, misses, size, inflightCount } } }
|
||||
// Pre-D75 cmdStatus formatter read `body.cache.entries` (OCP-era) → always
|
||||
// undefined → output showed "entries=?". Pin the wire contract by grepping
|
||||
// the CLI source so a future server-side rename can't silently re-break.
|
||||
const cliSrc = _readFileSyncS36(_joinS36(import.meta.dirname ?? process.cwd(), 'bin/olp.mjs'), 'utf8');
|
||||
// Must contain the cmdStatus block that reads body.stats.cache (not body.cache directly)
|
||||
assert.ok(/cmdStatus\b/.test(cliSrc), 'cmdStatus function must exist');
|
||||
// The status block must read body.stats?.cache (the actual server payload nesting)
|
||||
assert.ok(/body\.stats\?\.cache/.test(cliSrc) || /body\.stats\.cache/.test(cliSrc),
|
||||
'cmdStatus must read body.stats.cache for status display');
|
||||
// Specifically, the "entries=" output must derive from c.size (matches CacheStore.stats() shape)
|
||||
// Find the status cache line and verify it uses c.size
|
||||
const statusBlock = cliSrc.match(/cmdStatus[\s\S]+?(?=async function cmd|^\}\s*$)/m);
|
||||
assert.ok(statusBlock, 'cmdStatus block extractable');
|
||||
assert.ok(/c\.size/.test(statusBlock[0]),
|
||||
'cmdStatus cache line must use c.size (matches CacheStore.stats() shape)');
|
||||
// And must NOT use the OCP-era c.entries
|
||||
assert.ok(!/c\.entries/.test(statusBlock[0]),
|
||||
'cmdStatus cache line must NOT read c.entries (OCP-era field that does not exist in CacheStore.stats())');
|
||||
// Also pin the CacheStore.stats() shape contract (re-affirms 36e but for status path)
|
||||
const cs = new CacheStore();
|
||||
const stats = cs.stats();
|
||||
assert.ok('size' in stats, 'CacheStore.stats() exposes size');
|
||||
assert.ok(!('entries' in stats), 'CacheStore.stats() does NOT expose entries');
|
||||
});
|
||||
|
||||
it('36q (F7) — per-hop chain `model` field overrides IR model on fallback hop', async () => {
|
||||
// Pre-D75: executeHopFn(provider, model, irReq) used `model` for cache key
|
||||
// + audit but passed irReq UNCHANGED to provider.spawn() → the chain
|
||||
// [{anthropic, claude-X}, {openai, gpt-5.5}]
|
||||
// would spawn codex with --model claude-X (the user's original request)
|
||||
// instead of gpt-5.5 (the chain's per-hop config). This broke cross-provider
|
||||
// fallback wherever the chain assigned different models per provider.
|
||||
//
|
||||
// Strategy: inject a mock openai provider that captures the IR it receives.
|
||||
// Force the anthropic hop to fail with SPAWN_FAILED so the chain advances
|
||||
// to openai. Assert the mock saw model === hop's configured model.
|
||||
|
||||
const { __setAuthConfig: setAC36q, __resetAuthConfig: resetAC36q,
|
||||
__setFallbackConfig: setFC36q, __resetFallbackConfig: resetFC36q,
|
||||
createOlpServer, loadedProviders } = await import('./server.mjs');
|
||||
|
||||
setAC36q({ allow_anonymous: true });
|
||||
|
||||
// Capture for mock openai provider
|
||||
let capturedIR = null;
|
||||
const mockOpenAI = {
|
||||
name: 'openai',
|
||||
displayName: 'Mock OpenAI for F7',
|
||||
contractVersion: '1.0',
|
||||
models: ['gpt-5.5'],
|
||||
auth: { type: 'oauth', storage: 'file', path: '/tmp/fake', refresh: 'auto' },
|
||||
async * spawn(ir, _authContext) {
|
||||
// Snapshot the model field as received
|
||||
capturedIR = { model: ir.model };
|
||||
yield { type: 'delta', role: 'assistant', content: 'mock-openai-response' };
|
||||
yield { type: 'stop', finish_reason: 'stop' };
|
||||
},
|
||||
estimateCost: () => null,
|
||||
quotaStatus: async () => null,
|
||||
healthCheck: async () => ({ ok: true, latencyMs: 0 }),
|
||||
hints: { requiresTTY: false, concurrentSpawnSafe: true, maxConcurrent: 4, cacheable: true },
|
||||
};
|
||||
|
||||
// Mock anthropic provider that always fails immediately with SPAWN_FAILED
|
||||
const mockAnthropic = {
|
||||
name: 'anthropic',
|
||||
displayName: 'Mock Anthropic for F7',
|
||||
contractVersion: '1.0',
|
||||
models: ['claude-sonnet-4-6'],
|
||||
auth: { type: 'oauth', storage: 'file', path: '/tmp/fake', refresh: 'auto' },
|
||||
async * spawn(_ir, _authContext) {
|
||||
throw new ProviderError('mock anthropic always fails (F7 test)', 'SPAWN_FAILED');
|
||||
// eslint-disable-next-line no-unreachable
|
||||
yield { type: 'stop' };
|
||||
},
|
||||
estimateCost: () => null,
|
||||
quotaStatus: async () => null,
|
||||
healthCheck: async () => ({ ok: true, latencyMs: 0 }),
|
||||
hints: { requiresTTY: false, concurrentSpawnSafe: true, maxConcurrent: 4, cacheable: true },
|
||||
};
|
||||
|
||||
// Save + clear loadedProviders, install the mocks
|
||||
const savedProviders = new Map(loadedProviders);
|
||||
loadedProviders.clear();
|
||||
loadedProviders.set('anthropic', mockAnthropic);
|
||||
loadedProviders.set('openai', mockOpenAI);
|
||||
|
||||
// 2-hop chain with DIFFERENT models per hop — the bug only manifests when
|
||||
// hopModel !== requested model on the hop that actually serves.
|
||||
setFC36q({
|
||||
chains: {
|
||||
'claude-sonnet-4-6': [
|
||||
{ provider: 'anthropic', model: 'claude-sonnet-4-6' },
|
||||
{ provider: 'openai', model: 'gpt-5.5' },
|
||||
],
|
||||
},
|
||||
soft_triggers: {},
|
||||
});
|
||||
|
||||
const srv = createOlpServer();
|
||||
let port = 28400 + Math.floor(Math.random() * 100);
|
||||
await new Promise((resolve, reject) => {
|
||||
srv.listen(port, '127.0.0.1', resolve);
|
||||
srv.once('error', e => {
|
||||
if (e.code === 'EADDRINUSE') { port++; srv.listen(port, '127.0.0.1', resolve); srv.once('error', reject); }
|
||||
else reject(e);
|
||||
});
|
||||
});
|
||||
|
||||
try {
|
||||
const r = await fetch({
|
||||
port,
|
||||
method: 'POST',
|
||||
path: '/v1/chat/completions',
|
||||
body: {
|
||||
model: 'claude-sonnet-4-6',
|
||||
messages: [{ role: 'user', content: 'F7 test' }],
|
||||
},
|
||||
});
|
||||
assert.equal(r.status, 200, `expected 200, got ${r.status}: ${r.body.slice(0, 300)}`);
|
||||
assert.equal(r.headers['x-olp-provider-used'], 'openai',
|
||||
'openai must be the serving hop (anthropic failed)');
|
||||
// The F7 assertion: the openai hop received the chain's configured model,
|
||||
// NOT the user's original request model.
|
||||
assert.ok(capturedIR, 'mock openai must have been invoked');
|
||||
assert.equal(capturedIR.model, 'gpt-5.5',
|
||||
`F7: openai hop must receive its chain-configured model 'gpt-5.5', got '${capturedIR.model}'`);
|
||||
} finally {
|
||||
await new Promise(r => srv.close(r));
|
||||
resetFC36q();
|
||||
// Restore original providers
|
||||
loadedProviders.clear();
|
||||
for (const [k, v] of savedProviders) loadedProviders.set(k, v);
|
||||
resetAC36q();
|
||||
}
|
||||
});
|
||||
|
||||
it('36r (F7) — per-hop model override preserved on streaming single-hop path', async () => {
|
||||
// The streaming path (server.mjs ~line 1390-1434) has its own spawn call site
|
||||
// inside sourceFactory that previously also passed `ir` unchanged. F7 fix
|
||||
// wraps it with the same { ...ir, model: streamModel } substitution. This
|
||||
// test exercises a single-hop streaming request and asserts the mock saw
|
||||
// the chain's configured model.
|
||||
|
||||
const { __setAuthConfig: setAC36r, __resetAuthConfig: resetAC36r,
|
||||
__setFallbackConfig: setFC36r, __resetFallbackConfig: resetFC36r,
|
||||
createOlpServer, loadedProviders } = await import('./server.mjs');
|
||||
|
||||
setAC36r({ allow_anonymous: true });
|
||||
|
||||
let capturedModel = null;
|
||||
const mockOpenAIStream = {
|
||||
name: 'openai',
|
||||
displayName: 'Mock Codex for F7 streaming',
|
||||
contractVersion: '1.0',
|
||||
models: ['gpt-5.5'],
|
||||
auth: { type: 'oauth', storage: 'file', path: '/tmp/fake', refresh: 'auto' },
|
||||
async * spawn(ir, _authContext) {
|
||||
capturedModel = ir.model;
|
||||
yield { type: 'delta', role: 'assistant', content: 'streamed-content' };
|
||||
yield { type: 'stop', finish_reason: 'stop' };
|
||||
},
|
||||
estimateCost: () => null,
|
||||
quotaStatus: async () => null,
|
||||
healthCheck: async () => ({ ok: true, latencyMs: 0 }),
|
||||
hints: { requiresTTY: false, concurrentSpawnSafe: true, maxConcurrent: 4, cacheable: true },
|
||||
};
|
||||
|
||||
const savedProviders = new Map(loadedProviders);
|
||||
loadedProviders.clear();
|
||||
loadedProviders.set('openai', mockOpenAIStream);
|
||||
|
||||
// Single-hop chain with a DIFFERENT model than the request — the bug is
|
||||
// visible only when the user requests one model name and the chain remaps it.
|
||||
setFC36r({
|
||||
chains: {
|
||||
'gpt-aliased': [
|
||||
{ provider: 'openai', model: 'gpt-5.5' },
|
||||
],
|
||||
},
|
||||
soft_triggers: {},
|
||||
});
|
||||
|
||||
const srv = createOlpServer();
|
||||
let port = 28500 + Math.floor(Math.random() * 100);
|
||||
await new Promise((resolve, reject) => {
|
||||
srv.listen(port, '127.0.0.1', resolve);
|
||||
srv.once('error', e => {
|
||||
if (e.code === 'EADDRINUSE') { port++; srv.listen(port, '127.0.0.1', resolve); srv.once('error', reject); }
|
||||
else reject(e);
|
||||
});
|
||||
});
|
||||
|
||||
try {
|
||||
const r = await fetch({
|
||||
port,
|
||||
method: 'POST',
|
||||
path: '/v1/chat/completions',
|
||||
body: {
|
||||
model: 'gpt-aliased',
|
||||
stream: true,
|
||||
messages: [{ role: 'user', content: 'F7 streaming test' }],
|
||||
},
|
||||
});
|
||||
assert.equal(r.status, 200, `expected 200, got ${r.status}: ${r.body.slice(0, 200)}`);
|
||||
assert.equal(capturedModel, 'gpt-5.5',
|
||||
`F7 streaming: openai must receive chain-configured model 'gpt-5.5', got '${capturedModel}'`);
|
||||
} finally {
|
||||
await new Promise(r => srv.close(r));
|
||||
resetFC36r();
|
||||
loadedProviders.clear();
|
||||
for (const [k, v] of savedProviders) loadedProviders.set(k, v);
|
||||
resetAC36r();
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user