mirror of
https://github.com/dtzp555-max/olp.git
synced 2026-07-22 13:35:10 +00:00
Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6605b7b14a | ||
|
|
6edf6e0b94 |
@@ -6,6 +6,54 @@ All notable changes to OLP land here. Per `CLAUDE.md` release_kit overlay, this
|
||||
|
||||
(empty — Phase 5 entries land here once Phase 5 opens)
|
||||
|
||||
## v0.4.3 — 2026-05-26
|
||||
|
||||
### D76 — README install-path overhaul + `OLP_BIND` env + AI-driven install prompt + ADR 0011 amendment
|
||||
|
||||
Patch release closing the install-experience gap. v0.4.0–v0.4.2 README's Quick Start was placeholder text with fictional commands (`npm install -g @dtzp555-max/olp` — package isn't published; `olp setup` / `olp start` — don't exist). 10 real gaps catalogued + fixed in one D-day; `OLP_BIND` env wired so the documented LAN onboarding flow actually works; AI-driven install prompt added per the Phase 4 charter brainstorm's #2 OCP inheritance candidate (was deferred at D64-D67 to the doctor framework only; D76 closes the README half).
|
||||
|
||||
- **G1-G7 (README "Quick Start" was fictional)** — rewrote § "Manual install" with the real sequence: prerequisites (Node ≥ 18 + provider CLI install matrix) → `git clone` → `npm test` verify → `olp-keys keygen --owner` first → provider OAuth (claude/codex/mistral per-CLI flows) → write `~/.olp/config.json` with the minimum that actually serves traffic → `npm start` → smoke-test → IDE pointing. Each step empirically verified against the PI231 + Mac mini E2E session (2026-05-26).
|
||||
- **G8 (LAN unreachable — F5)** — added `OLP_BIND` env (default `127.0.0.1`). Operators set `OLP_BIND=0.0.0.0` (or a specific LAN IP) to accept LAN connections so `olp-connect <ip>` can actually reach the server. Pre-D76 the server was hard-coded to `server.listen(PORT, '127.0.0.1', ...)`, making the documented LAN-onboarding flow only usable through an SSH tunnel. ADR 0011's original wording referenced a `BIND_ADDRESS` concept that didn't exist; D76 makes it operational.
|
||||
- **G10 (no AI-install pattern)** — README § "Install with your AI (the fast path)" added. Verbatim prompt that the operator pastes into Claude Code / Cursor / Copilot / Aider; the AI follows the README + uses `olp doctor --json` machine-readable `next_action.ai_executable[]` (D64-D67) for self-repair, stopping only when `human_required[]` is non-empty (the provider OAuth dances). This closes the Phase 4 brainstorm Top-5 inheritance candidate #2 — the OCP "paste this prompt" pattern that D64-D67 only half-built.
|
||||
- **Opening compressed** — § "Why OLP" (3 paragraphs of OCP billing history) removed from the top. The OCP-trigger context moved to § "Migration from OCP" at the bottom, condensed into a single paragraph. New users land on value-prop + § "What you get" + § "Install with your AI" / § "Manual install" without needing to digest 2026-05-14 / 2026-06-15 Anthropic billing history first. OCP users get a one-line pointer at the top.
|
||||
- **§ "Configuration" full schema documentation** — replaced the placeholder with the actual `~/.olp/config.json` schema including every field that v0.4.x reads. Cross-references ADR 0004/0007/0010/0011.
|
||||
- **§ "Environment Variables" extended** — added `OLP_BIND`, `OLP_API_KEY`, `OLP_OWNER_TOKEN`, `OLP_PROXY_URL` rows that were used throughout the manual-install flow but undocumented.
|
||||
|
||||
**ADR 0011 § "Deployment configurations" amendment.** Codifies the three deployment trust contexts (`127.0.0.1` loopback / RFC1918 + tailnet LAN / `0.0.0.0` public — with `advertise_anonymous_key: true` only safe in the first two). Documents the new `anonymous_key_advertised_with_lan_bind` startup warn event. Closes ADR 0011's pre-D76 dangling reference to a non-existent `BIND_ADDRESS`.
|
||||
|
||||
**Test count:** 714 (v0.4.2) → 717 (v0.4.3). +3 D76 regression tests in Suite 36 (36s/36t/36u) pinning `OLP_BIND` wiring + safety warn + ADR amendment.
|
||||
|
||||
**Out of D76 scope (deferred):**
|
||||
- F6 (doctor client-side vs server-side check separation) — needs design ADR for a `--remote` mode. Phase 5.
|
||||
- D75 reviewer P2-1 (ADR 0004 amendment for per-hop schema) + P2-2 (defensive `typeof hopModel === 'string'`) — both genuine follow-ups, neither blocking.
|
||||
- `scripts/migrate-from-ocp.mjs` — Phase 7.
|
||||
|
||||
**Authority:** PI231 + Mac mini E2E session (2026-05-26, post-v0.4.2 verification revealed the 10 README gaps); ADR 0011 amendment self-cites; Phase 4 charter (ADR 0010) Top-5 inheritance candidate #2 (AI-driven self-repair). Process learning: every D-day reviewer rubric should add "open README in §-Quick-Start and verify the commands literally exist + work in the current repo" — would have caught G1-G7 at v0.4.0.
|
||||
|
||||
## v0.4.2 — 2026-05-26
|
||||
|
||||
### Post-v0.4.1 hotfix batch (D75) — real-machine E2E findings
|
||||
|
||||
Patch release fixing 5 bugs caught by **real-machine E2E testing on PI231 + Mac mini (2026-05-26 session)** — bugs that prior D-day reviewers AND the post-v0.4.0 maintainer review both missed because they reviewed against spec text and against the local OLP install's `~/.codex/auth.json` shape (cached from an older codex CLI version), not against real provider CLIs running on a remote operator host that did `npm install -g @openai/codex` for the first time on 2026-05-26 and got codex CLI v0.133.0.
|
||||
|
||||
**Root cause of the missed-bug class.** D6 (codex plugin authoring) explicitly documented three unpinned assumptions (A3 = access-token field name, A4 = NDJSON event schema, A2-adjacent = trusted-directory sandbox). D6 noted "D7 E2E will pin." D7 then shipped without performing real-codex-CLI E2E (the E2E gating mark was carried but the actual run was deferred). Every subsequent D-day reviewer trusted the D6/D7 codex plugin code unchanged because the static review couldn't see that the v0.133.0 CLI had moved the auth-token field, the event schema, AND added a new trusted-directory sandbox flag. The D74 maintainer review focused on `/health` / `/cache/stats` / `/v0/management/dashboard-data` payload shapes — none of which exercise the codex plugin's spawn path. F7 (per-hop model override) is a different class of miss — every reviewer read `executeHopFn(provider, model, ir)` and saw `model` consumed for cache key + audit ctx, but none traced through to confirm `model` is ALSO substituted into the IR passed to `provider.spawn()`. The function signature implied per-hop semantics that the body never fully delivered.
|
||||
|
||||
- **[F1] codex auth.json schema pin — codex CLI v0.133.0 nests the access token under `tokens.access_token`** (verified empirically on PI231 / Mac mini, 2026-05-26). Pre-D75 `readAuthArtifact()` read only top-level `creds.access_token` / `creds.token` / `creds.accessToken` — all undefined under v0.133.0 → returned `null` → OLP reported "auth artifact missing" via `/health` and `olp doctor` AND refused to spawn codex even when the user had fully completed `codex login`. Fix: prepend `creds?.tokens?.access_token` to the precedence chain at BOTH call sites (`OPENAI_CODEX_AUTH_PATH` override branch + default `$CODEX_HOME/auth.json` branch). Legacy top-level fields preserved as fallback for backward compat with older codex CLI versions.
|
||||
- **[F2] codex spawn args — codex CLI v0.133.0 trusted-directory sandbox requires `--skip-git-repo-check`.** v0.133.0 refuses with `"Not inside a trusted directory and --skip-git-repo-check was not specified."` when spawned outside a git repo, exits non-zero with zero NDJSON output → OLP surfaces `SPAWN_FAILED` with no usable chunks → the fallback engine advances to next hop unnecessarily even when codex is configured and authenticated. OLP's typical deploy CWD (`~/olp/`) is NOT a git repo on operator hosts. Fix: add `'--skip-git-repo-check'` to the args array before `--model`. OLP is the trusted caller (operator's own server invoking the operator's own subscription via documented `codex exec` automation); the sandbox safeguards interactive shells, not pre-authorized automation.
|
||||
- **[F3] codex NDJSON event shape pin — codex CLI v0.133.0 emits `item.completed` + `turn.completed` + `turn.failed`**, not the D6-assumed `content`/`delta`/`text` + `type:'stop'`/`done:true` shapes. Real v0.133.0 stream (verified empirically): `{"type":"thread.started",...}` → `{"type":"turn.started"}` → `{"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"<response>"}}` → `{"type":"turn.completed","usage":{...}}`. Pre-D75, every chunk was silently dropped by `codexChunkToIR()` → response body had `content: null`. Fix: prepend three new recognizers (`item.completed` with `item.type === 'agent_message'` → IR delta; `turn.completed` → IR stop; `turn.failed` → IR error). Legacy D6 defensive recognizers preserved below as forward/backward compat fallbacks.
|
||||
- **[F4] `olp status` reads `body.stats.cache.size` (not OCP-era `body.cache.entries`).** Same class as D74 P2-3 (which fixed `cmdUsage` + `cmdCache`); D74 missed the parallel bug in `cmdStatus`. Server payload nests cache stats as `body.stats.cache.{hits, misses, size, inflightCount}` per `server.mjs handleManagementStatus`, and `CacheStore.stats()` has no `entries` field per `lib/cache/store.mjs`. Pre-D75 output showed `entries=?`. Fix: read `c.size` for entries display; also surface `inflightCount` when present.
|
||||
- **[F7] per-hop chain `model` field now overrides IR model in `provider.spawn()`.** Pre-D75 `executeHopFn(hopProvider, hopModel, irReq)` used `hopModel` for cache key + audit ctx but passed the ORIGINAL `irReq` (with `irReq.model` = user's original request) to `hopProviderPlugin.spawn(irReq, authContext)`. A chain config `[{provider:anthropic, model:claude-X}, {provider:openai, model:gpt-5.5}]` would always spawn BOTH plugins with `--model claude-X` — openai rejected the unknown model and the chain died. This broke the core OLP value prop (cross-provider fallback with provider-appropriate model substitution per hop). Fix: build a per-hop IR variant with `{...irReq, model: hopModel}` and pass that to spawn. Conditional skips clone when `hopModel === irReq.model` (common case: single-provider chains, or single-hop chains where the chain config repeats the request model). Applied to BOTH the buffered path (`executeHopFn`) AND the streaming path (`sourceFactory` for `getOrComputeStreaming`). **Authority:** ADR 0004 § Chain advancement step 1 (per-hop config supplies provider AND model — the contract was always specified, but the code didn't complete it).
|
||||
|
||||
**Phase 5 process learning recorded.** Every provider plugin's D-day must include a real-CLI E2E ON A REMOTE OPERATOR HOST before merging — not on the maintainer workstation (which may have an older CLI cached from a prior install, hiding new field renames / new sandbox flags / new event shapes). The D6/D7 codex E2E was deferred and that deferral compounded across 3 layers (D6 = unpinned, D7 = pinning deferred, D8+ = trusted D6/D7 unchanged). F7 reinforces a separate lesson: when a function signature takes `(provider, model, ir)`, reviewers must check that `model` is consumed everywhere downstream, not just at the call site they happened to look at.
|
||||
|
||||
**Out of D75 scope (deferred to Phase 5 explicit ADR amendments):**
|
||||
- F5 (server bind 127.0.0.1 / `OLP_BIND` env) — needs `lib/keys.mjs` anonymous-key trust boundary review before binding to non-loopback by default
|
||||
- F6 (`olp doctor` client-vs-server-side limit detection) — needs design ADR amendment for trigger taxonomy
|
||||
|
||||
- **Test count delta:** 704 (v0.4.1) → 714 (v0.4.2). +10 D75 regression tests in Suite 36 (36i through 36r).
|
||||
- **Files touched:** `lib/providers/codex.mjs` (F1+F2+F3), `bin/olp.mjs` (F4 cmdStatus), `server.mjs` (F7 buffered + streaming spawn sites), `test-features.mjs` (Suite 36 extension), `package.json` (version), `CHANGELOG.md` (this entry).
|
||||
- **Authority:** ADR 0002 (provider contract — codex plugin), ADR 0004 (fallback engine — per-hop model contract), `lib/providers/codex.mjs` D6 assumption A2/A3/A4 docstrings (which all said "D7 will pin" and D7 never did); codex CLI v0.133.0 on-disk schema + `codex exec --help` output verified empirically on PI231 (2026-05-26 E2E session); Iron Rule 第二律 evidence-over-should-work; CLAUDE.md `release_kit.phase_rolling_mode` cross-Phase discipline.
|
||||
|
||||
## v0.4.1 — 2026-05-26
|
||||
|
||||
### Post-Phase-4 hotfix batch (D74) — maintainer-review findings
|
||||
|
||||
@@ -1,55 +1,204 @@
|
||||
# OLP — Open LLM Proxy
|
||||
|
||||
A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing, automatic fallback, content-addressed caching — so your IDEs and family clients keep working as long as *any* of your subscriptions has quota left.
|
||||
A personal- and family-scale multi-provider LLM proxy. One HTTP endpoint, many subscriptions behind it, automatic routing + fallback + content-addressed caching. Your IDEs and family clients keep working as long as **any** of your subscriptions has quota left.
|
||||
|
||||
> **Status:** v0.4.0 shipped (2026-05-26) — Phase 1 multi-provider proxy core (v0.1.0 + v0.1.1) + Phase 2 multi-key auth + audit + owner gating + keygen CLI (v0.2.0) + Phase 3 Dashboard + audit query layer + daily audit rotation (v0.3.0) + Phase 4 Operator + Client UX (v0.4.0): SSE heartbeat / `olp` Node CLI + `olp doctor` framework / `olp-connect` zero-config LAN setup / `/health.anonymousKey` opt-in / `/olp` Telegram-Discord plugin / 6-IDE integration docs. Phase 5 scope is open — candidates per ADR 0010 § Out-of-Phase-4-scope: `/v1/messages` (gated on ADR 0009 P0 outcome + named family CC user), context-window-exceeded fallback trigger, per-(provider, model) live stats. Sections marked _placeholder_ land alongside the relevant phase of work (see [phase plan](#phase-plan)).
|
||||
> **Status:** v0.4.3 shipped, 714+ tests. Phase 4 (Operator + Client UX) closed; Phase 5 scope is open. Coming from [OCP](https://github.com/dtzp555-max/ocp)? See [§ Migration from OCP](#migration-from-ocp).
|
||||
|
||||
---
|
||||
|
||||
## Why OLP
|
||||
## What you get
|
||||
|
||||
On 2026-05-14, Anthropic announced (effective 2026-06-15) that `claude -p`, the Agent SDK, and third-party agent traffic move out of the Pro/Max subscription pool into a separate fixed monthly Agent SDK Credit pool. [OCP](https://github.com/dtzp555-max/ocp), OLP's predecessor, was a proxy around a single CLI — its core assumption was *"subscription = unlimited within rate limits"*. That assumption breaks for Anthropic on the effective date.
|
||||
|
||||
The structural response is to stop relying on one provider's subscription terms remaining favourable. OLP spreads risk across multiple providers whose subscriptions still include CLI/programmatic use, routes intelligently between them, and caches aggressively so every request that does spawn a CLI counts.
|
||||
|
||||
OLP is **not**: a commercial multi-tenant SaaS; an enterprise gateway competing with LiteLLM / OpenCode / CLIProxyAPI on breadth; a model-capability router ("route to the smartest model" — you pick the model); a conversation-state store (your client handles that).
|
||||
|
||||
See [`ALIGNMENT.md`](./ALIGNMENT.md) for OLP's constitution and [`docs/adr/`](./docs/adr/) for the founding ADRs.
|
||||
- **OpenAI-compatible** `/v1/chat/completions` endpoint — any IDE that speaks OpenAI (Cline / Continue.dev / Cursor / Aider) plugs in
|
||||
- **Multi-provider chain** — primary fails / quota dies → automatically falls back to the next provider (anthropic ↔ codex ↔ mistral by default; risk-tier framework guards which ones get enabled)
|
||||
- **Content-addressed cache** — repeat requests don't re-spawn the CLI; streaming requests dedup via singleflight tee
|
||||
- **Multi-key auth** — owner key with full visibility, family-member keys with per-key audit log + per-provider scoping
|
||||
- **Telegram / Discord** `/olp` slash commands (read-only — for "is OLP up?" checks from anywhere)
|
||||
- **AI-driven self-repair** — `olp doctor --json` emits machine-readable `next_action.ai_executable[]` so a Claude Code / Cursor / Copilot session can fix install issues for you (see [§ Install with your AI](#install-with-your-ai-the-fast-path))
|
||||
- **Observability** — owner-only `/dashboard` (quota / 24h stats / 30d spend trend / top fallback chains)
|
||||
|
||||
---
|
||||
|
||||
## Quick Start
|
||||
## Install with your AI (the fast path)
|
||||
|
||||
_placeholder — lands with Phase 1._
|
||||
If the manual steps feel like a lot, paste this verbatim into your AI coding assistant (Claude Code / Cursor / Copilot / Aider). It walks you through everything:
|
||||
|
||||
Anticipated shape:
|
||||
```
|
||||
I want to install OLP (Open LLM Proxy) on this machine. The repo is at
|
||||
https://github.com/dtzp555-max/olp and ships these tools that you can use:
|
||||
|
||||
```bash
|
||||
# install
|
||||
npm install -g @dtzp555-max/olp
|
||||
1. `git clone` the repo to ~/olp
|
||||
2. README.md § "Manual install" has the sequenced steps. Follow them in order.
|
||||
3. After EACH step, run `node ~/olp/bin/olp.mjs doctor --json` (NOT `npx olp`
|
||||
— the package isn't on npm). Parse the JSON. The `kind` field tells you
|
||||
the state: noop / fresh_install / fix_server / fix_oauth / fix_provider /
|
||||
fix_config.
|
||||
4. If `kind != noop`, run the commands listed in `next_action.ai_executable[]`
|
||||
verbatim. Then re-run doctor to verify.
|
||||
5. STOP and ask me only when `next_action.human_required[]` is non-empty.
|
||||
That's where I need to do a browser OAuth flow you can't do for me.
|
||||
|
||||
# run setup (writes ~/.olp/config.json, asks which providers to enable)
|
||||
olp setup
|
||||
The provider CLIs OLP spawns (claude / codex / vibe) need their own one-time
|
||||
OAuth — those are the only steps I personally have to do (Claude.ai login,
|
||||
ChatGPT login, Mistral API key). Everything else (clone, npm install of the
|
||||
provider CLIs, owner-key generation, config.json bootstrap, server start) is
|
||||
in your `ai_executable[]` and you should run it without asking.
|
||||
|
||||
# start the proxy (default port 4567 since v0.4.0 — moved off OCP's 3456 so
|
||||
# OLP and OCP can co-host on the same machine. Set OLP_PORT=3456 if you have
|
||||
# no OCP on the machine and want the old default.)
|
||||
olp start
|
||||
|
||||
# point your IDE at http://localhost:4567/v1/chat/completions with the OLP API key from `olp keys list`.
|
||||
Begin.
|
||||
```
|
||||
|
||||
**Family-on-LAN onboarding (D68-D70).** For other devices on the same network, run on the client device:
|
||||
Then sit back and respond when it asks for OAuth confirmation. This pattern works because `olp doctor` is purpose-built for AI consumption — every failure mode has a shell-executable repair command AND a human-required step listed separately.
|
||||
|
||||
---
|
||||
|
||||
## Manual install (5-10 min)
|
||||
|
||||
### 0. Prerequisites
|
||||
|
||||
- **Node.js ≥ 18.** Verify: `node --version`
|
||||
- **The provider CLIs you want OLP to spawn.** Install whichever you'll actually use:
|
||||
|
||||
| Provider | Install | Subscription |
|
||||
|---|---|---|
|
||||
| `anthropic` (`claude -p`) | `npm install -g @anthropic-ai/claude-code` | Claude Pro/Max (OAuth) |
|
||||
| `openai` (`codex exec`) | `npm install -g @openai/codex` | ChatGPT Plus/Pro (OAuth) or OpenAI API key |
|
||||
| `mistral` (`vibe --prompt`) | follow the `vibe` install docs | Le Chat Pro API key |
|
||||
|
||||
You only need to install the ones you'll route to. Single-provider OLP works fine.
|
||||
|
||||
### 1. Clone and verify the test suite
|
||||
|
||||
```bash
|
||||
# Detects Cline / Continue.dev / Cursor / Aider / OpenClaw installed locally
|
||||
# and writes per-tool config pointing at the OLP host. Requires `python3`.
|
||||
olp-connect <olp-host-ip>
|
||||
git clone https://github.com/dtzp555-max/olp.git ~/olp
|
||||
cd ~/olp
|
||||
npm test # 714+ tests, ~5s, no external deps
|
||||
```
|
||||
|
||||
If the OLP host has `auth.advertise_anonymous_key: true` AND a key was created with `olp-keys keygen --anonymous --advertise`, `olp-connect` picks up the token from `/health.anonymousKey` — zero out-of-band token paste required. See [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md) for the trusted-LAN-only invariant.
|
||||
(If `npm test` fails here, stop — that means your Node version or the repo state is broken. Don't proceed to step 2.)
|
||||
|
||||
Per-IDE setup details: [`docs/integrations/`](./docs/integrations/README.md).
|
||||
### 2. Bootstrap the owner key
|
||||
|
||||
The owner key is what you (and `olp-connect`) use to authenticate to OLP. Default config has `auth.allow_anonymous: false`, so you need a key BEFORE the server starts accepting requests.
|
||||
|
||||
```bash
|
||||
node ~/olp/bin/olp-keys.mjs keygen --owner --name=$(whoami)-laptop
|
||||
# Prints the plaintext token ONCE. Copy it now — you can't recover it later.
|
||||
# Example: olp_l23-PN46tDljmPATV94-KfOgOBO0Ed8theVjTdAgQoY
|
||||
```
|
||||
|
||||
Export it so the CLI subcommands can use it:
|
||||
|
||||
```bash
|
||||
export OLP_API_KEY=olp_l23-PN46... # paste your real token
|
||||
```
|
||||
|
||||
(Add to `~/.bashrc` / `~/.zshrc` to persist.)
|
||||
|
||||
### 3. Authenticate the providers (one-time OAuth)
|
||||
|
||||
Run each provider's own login flow. OLP's anthropic / openai / mistral plugins spawn these CLIs and reuse their cached credentials — OLP itself never touches the OAuth dance.
|
||||
|
||||
```bash
|
||||
# Anthropic (Claude Pro/Max subscription)
|
||||
claude setup-token
|
||||
# Opens a TUI / prints a URL. Authorize in browser. Paste the returned code.
|
||||
# Result: ~/.claude/.credentials.json
|
||||
|
||||
# OpenAI (ChatGPT subscription)
|
||||
codex login --device-auth
|
||||
# Prints a https://auth.openai.com/codex/device URL + 10-char code.
|
||||
# Open URL in browser, enter code, authorize.
|
||||
# Result: ~/.codex/auth.json
|
||||
|
||||
# Mistral (Le Chat API key)
|
||||
export MISTRAL_API_KEY=sk-... # add to ~/.bashrc to persist
|
||||
```
|
||||
|
||||
### 4. Write a minimum config
|
||||
|
||||
`~/.olp/config.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"auth": {
|
||||
"allow_anonymous": false,
|
||||
"owner_only_endpoints": [
|
||||
"/health",
|
||||
"/v0/management/dashboard-data",
|
||||
"/v0/management/quota",
|
||||
"/v0/management/status",
|
||||
"/cache/stats",
|
||||
"/dashboard"
|
||||
],
|
||||
"fallback_detail_header_policy": "owner_only"
|
||||
},
|
||||
"providers": {
|
||||
"enabled": { "anthropic": true, "openai": true }
|
||||
},
|
||||
"routing": {
|
||||
"chains": {
|
||||
"claude-sonnet-4-6": [
|
||||
{ "provider": "anthropic", "model": "claude-sonnet-4-6" },
|
||||
{ "provider": "openai", "model": "gpt-5.5" }
|
||||
],
|
||||
"gpt-5.5": [{ "provider": "openai", "model": "gpt-5.5" }]
|
||||
}
|
||||
},
|
||||
"streaming": { "heartbeat_interval_ms": 15000 }
|
||||
}
|
||||
```
|
||||
|
||||
(Enable only the providers you actually authenticated in step 3. Chains map `<your-IDE's-requested-model>` → ordered list of `{provider, model}` hops; the chain's per-hop `model` is what gets passed to that provider's CLI.)
|
||||
|
||||
### 5. Start the server
|
||||
|
||||
```bash
|
||||
cd ~/olp
|
||||
npm start
|
||||
# OLP v0.4.3 listening on :4567 (2 providers enabled)
|
||||
```
|
||||
|
||||
### 6. Smoke-test
|
||||
|
||||
```bash
|
||||
curl -H "Authorization: Bearer $OLP_API_KEY" http://localhost:4567/health | jq
|
||||
# Expect: {ok: true, providers: {enabled: 2, status: {anthropic: {ok: true...}, openai: {ok: true...}}}}
|
||||
|
||||
node ~/olp/bin/olp.mjs doctor
|
||||
# Expect: "9 of 9 checks passed", kind=noop
|
||||
```
|
||||
|
||||
### 7. Point your IDE at OLP
|
||||
|
||||
```
|
||||
OPENAI_BASE_URL=http://localhost:4567/v1
|
||||
OPENAI_API_KEY=$OLP_API_KEY
|
||||
```
|
||||
|
||||
Per-IDE configuration details: [`docs/integrations/`](./docs/integrations/README.md).
|
||||
|
||||
---
|
||||
|
||||
## Family / LAN setup
|
||||
|
||||
To let other devices on your home network use the same OLP server, you need TWO things:
|
||||
|
||||
1. **Bind to the LAN interface** (not just loopback). On the SERVER:
|
||||
|
||||
```bash
|
||||
OLP_BIND=0.0.0.0 npm start # or your specific LAN IP, e.g. 192.168.1.10
|
||||
```
|
||||
|
||||
Default is `127.0.0.1` (loopback only). See [ADR 0011 § Deployment configurations](./docs/adr/0011-anonymous-key-deployment-context.md#deployment-configurations-d76-amendment-2026-05-26) for the trust-context table — **never set `OLP_BIND=0.0.0.0` on a public-internet-facing host** (use a tunnel like Tailscale instead).
|
||||
|
||||
2. **Onboard each family member's device** from THEIR machine:
|
||||
|
||||
```bash
|
||||
bash <(curl -fsSL https://raw.githubusercontent.com/dtzp555-max/olp/main/bin/olp-connect) <olp-host-ip>
|
||||
```
|
||||
|
||||
Detects Cline / Continue.dev / Cursor / Aider / OpenClaw locally and writes per-tool config pointing at your OLP host. Requires `python3` on the client. Prompts for the OLP API key — OR, if the server has `auth.advertise_anonymous_key: true` AND a key was created with `olp-keys keygen --anonymous --advertise`, picks the token up from `/health.anonymousKey` (zero out-of-band paste). See [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md) for the trusted-LAN-only invariant.
|
||||
|
||||
Per-IDE setup details: [`docs/integrations/`](./docs/integrations/README.md). Telegram / Discord `/olp` slash command setup: [§ Telegram / Discord Usage](#telegram--discord-usage).
|
||||
|
||||
---
|
||||
|
||||
@@ -80,12 +229,19 @@ OLP distinguishes **Candidate Providers** (declared as intended, not yet pinned)
|
||||
|
||||
## Configuration
|
||||
|
||||
_placeholder — full configuration reference lands with Phase 4 (fallback engine)._
|
||||
|
||||
OLP reads its config from `~/.olp/config.json`. The minimum useful shape:
|
||||
OLP reads `~/.olp/config.json` at startup. § "[Manual install § Step 4](#4-write-a-minimum-config)" above has a working minimum example. The full schema:
|
||||
|
||||
```json
|
||||
{
|
||||
"auth": {
|
||||
"allow_anonymous": false,
|
||||
"owner_only_endpoints": ["/health", "/dashboard", "/v0/management/..."],
|
||||
"advertise_anonymous_key": false,
|
||||
"fallback_detail_header_policy": "owner_only"
|
||||
},
|
||||
"providers": {
|
||||
"enabled": { "<provider-key>": true }
|
||||
},
|
||||
"routing": {
|
||||
"chains": {
|
||||
"<requested-model>": [
|
||||
@@ -96,13 +252,25 @@ OLP reads its config from `~/.olp/config.json`. The minimum useful shape:
|
||||
"soft_triggers": {
|
||||
"<provider-key>": { "<trigger>": <threshold> }
|
||||
}
|
||||
},
|
||||
"streaming": {
|
||||
"heartbeat_interval_ms": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
> **Note:** `routing.soft_triggers` thresholds are parsed and stored but have **no runtime effect at v0.1** — the quota polling path (`quotaStatus()` per hop) is deferred to v1.x per [ADR 0004 Amendment 2](./docs/adr/0004-fallback-engine.md#amendment-2--2026-05-24-soft-triggers-deferred-to-v1x-d22). The evaluation logic exists and is tested; only the production data ingestion path is deferred.
|
||||
Field guide:
|
||||
|
||||
Trigger types, fallback safety, idempotency rules, and the full example config land here when Phase 4 ships. See [ADR 0004 (Fallback Engine Semantics & Safety)](./docs/adr/0004-fallback-engine.md) for the design.
|
||||
- **`auth.allow_anonymous`** — default `false`. When false, every request needs a Bearer token; when true, anonymous-tier requests succeed (ADR 0007 § 7). Production posture is `false`.
|
||||
- **`auth.owner_only_endpoints`** — list of endpoints that REQUIRE owner-tier auth (non-owner returns 401). The defaults above are minimum sane for production.
|
||||
- **`auth.advertise_anonymous_key`** — default `false`. When true (+ `allow_anonymous: true` + a key created with `olp-keys keygen --anonymous --advertise`), `/health.anonymousKey` exposes the plaintext token so `olp-connect <ip>` is zero-config. **Trusted-LAN only** — see [ADR 0011](./docs/adr/0011-anonymous-key-deployment-context.md).
|
||||
- **`auth.fallback_detail_header_policy`** — controls `X-OLP-Fallback-Detail` response header emission. `owner_only` (default) only shows tuples to owner identity; debug surface to LAN family without leaking to anonymous.
|
||||
- **`providers.enabled`** — flip a provider plugin on. Only enable providers whose CLI you've authenticated; OLP doesn't do its own OAuth.
|
||||
- **`routing.chains`** — keyed by the model name your IDE / client requests. Each entry is an ordered list of fallback hops; each hop's `model` is what gets passed to that provider's CLI. F7 fix (D75) — the hop-level `model` field finally overrides the IR's request model during cross-provider fallback.
|
||||
- **`routing.soft_triggers`** — parsed and stored but **inert at v0.4.x** — the `quotaStatus()` polling data path is deferred to v1.x per [ADR 0004 Amendment 2](./docs/adr/0004-fallback-engine.md#amendment-2--2026-05-24-soft-triggers-deferred-to-v1x-d22). Startup emits a warn if non-empty so the inert state is visible.
|
||||
- **`streaming.heartbeat_interval_ms`** — default `0` (disabled). Set > 0 (e.g. `15000`) to emit SSE keepalive frames during silent windows. Required behind reverse proxies (nginx / Cloudflare Tunnel / Tailscale Funnel) with 60s idle aborts.
|
||||
|
||||
See [ADR 0004 (Fallback Engine)](./docs/adr/0004-fallback-engine.md), [ADR 0007 (Multi-key auth)](./docs/adr/0007-multi-key-auth.md), [ADR 0010 (Phase 4 charter)](./docs/adr/0010-phase-4-charter-operator-and-client-ux.md), [ADR 0011 (Anonymous-key deployment)](./docs/adr/0011-anonymous-key-deployment-context.md).
|
||||
|
||||
---
|
||||
|
||||
@@ -127,6 +295,10 @@ _placeholder — full table lands per-phase as variables are introduced._
|
||||
| Variable | Default | Description |
|
||||
|---|---|---|
|
||||
| `OLP_PORT` | `4567` | HTTP listener port. Moved off `3456` at D60 / v0.4.0 to co-host with OCP — set `OLP_PORT=3456` to restore the pre-D60 default. |
|
||||
| `OLP_BIND` | `127.0.0.1` | HTTP listener bind address. **Set to `0.0.0.0` or your LAN IP to accept LAN connections** (required for `olp-connect <ip>` to actually reach the server). Default loopback-only is the secure default. See [ADR 0011 § Deployment configurations](./docs/adr/0011-anonymous-key-deployment-context.md#deployment-configurations-d76-amendment-2026-05-26) for the trust-context table — never bind to a public-internet IP. |
|
||||
| `OLP_API_KEY` | (none) | Owner-tier OLP API key (the `olp_...` plaintext from `olp-keys keygen --owner`) used by `olp` CLI subcommands as the bearer for management endpoints. |
|
||||
| `OLP_OWNER_TOKEN` | (none) | Fallback used by `olp` CLI if `OLP_API_KEY` is absent. |
|
||||
| `OLP_PROXY_URL` | `http://127.0.0.1:$OLP_PORT` | Override target URL for `olp` CLI subcommands (so the same binary works against a remote OLP via SSH tunnel or direct LAN). |
|
||||
| `OLP_CLAUDE_BIN` | `claude` (from PATH) | Override path to the `claude` binary (Anthropic provider). Useful when multiple `claude` installs are present. |
|
||||
| `OLP_CODEX_BIN` | `codex` (from PATH) | Override path to the `codex` binary (OpenAI provider). |
|
||||
| `OLP_VIBE_BIN` | `vibe` (from PATH) | Override path to the `vibe` binary (Mistral provider). |
|
||||
@@ -353,16 +525,22 @@ Full spec (decision rationale, open questions, risks): `~/.cc-rules/memory/proje
|
||||
|
||||
## Migration from OCP
|
||||
|
||||
OLP is OCP's successor. The trigger was Anthropic's 2026-05-14 announcement (effective 2026-06-15) splitting `claude -p` / Agent SDK / third-party agent traffic out of the Pro/Max subscription pool into a separate fixed $100/month Agent SDK credit pool — invalidating OCP's foundational assumption (*"subscription = unlimited within rate limits"*) for its only provider. OLP's structural response is to spread risk across multiple subscriptions whose CLI/programmatic use remains in their main subscription pool, with intelligent fallback when one runs out.
|
||||
|
||||
Beyond the billing trigger, OLP is intentionally NOT a commercial multi-tenant SaaS (LiteLLM / OpenRouter / Portkey already serve that market with funding + SOC2), NOT an enterprise gateway competing on provider breadth, NOT a model-capability router ("route to the smartest model" — you pick the model in `routing.chains`), and NOT a conversation-state store (your client manages its own context). See [ADR 0001](./docs/adr/0001-project-founding.md) for the founding decision and [`ALIGNMENT.md`](./ALIGNMENT.md) for the constitution that governs every plugin / IR / entry-surface change.
|
||||
|
||||
### Migrating an existing OCP install
|
||||
|
||||
_placeholder — `scripts/migrate-from-ocp.mjs` lands with Phase 7 (📋 planned, not yet authored)._
|
||||
|
||||
Anticipated user-facing flow (target: <5 minutes):
|
||||
|
||||
1. Stop OCP (`launchctl bootout` the OCP service or `ocp stop`).
|
||||
2. Install OLP.
|
||||
3. Run `olp migrate-from-ocp` — copies `~/.ocp/keys/` to `~/.olp/keys/` and points provider plugins at OCP's existing auth artifacts where applicable.
|
||||
4. Start OLP. Clients pointing at port 4567 (or 3456 with `OLP_PORT=3456`) keep working; their existing OLP API keys remain valid. **Note (v0.4.0+):** default port moved from 3456 → 4567 so OCP and OLP can co-host during migration; set `OLP_PORT=3456` if you want the pre-D60 default.
|
||||
2. Install OLP (per [§ Manual install](#manual-install-5-10-min) above).
|
||||
3. Run `olp migrate-from-ocp` — will copy `~/.ocp/keys/` to `~/.olp/keys/` and point provider plugins at OCP's existing auth artifacts where applicable.
|
||||
4. Start OLP. Clients pointing at port 4567 (or 3456 with `OLP_PORT=3456`) keep working; their existing OLP API keys remain valid.
|
||||
|
||||
OCP's cache directory is *not* migrated: OLP's cache key format includes provider+model and warms cold naturally. OCP enters maintenance mode (stability fixes only) when OLP v0.1 ships; new development happens in OLP.
|
||||
**Default port moved 3456 → 4567 at v0.4.0** so OCP and OLP can co-host on the same machine during the migration window — set `OLP_PORT=3456` if you want the pre-D60 default. OCP's cache directory is *not* migrated: OLP's cache key format includes provider+model and warms cold naturally. OCP enters maintenance mode (stability fixes only) when OLP v0.1 ships; new development happens in OLP.
|
||||
|
||||
---
|
||||
|
||||
|
||||
+7
-1
@@ -251,9 +251,15 @@ async function cmdStatus(flags, io) {
|
||||
}
|
||||
io.log(` total reqs: ${body.stats?.total_requests ?? 0}`);
|
||||
io.log(` active reqs: ${body.stats?.active_requests ?? 0}`);
|
||||
// D75 F4 fix: server payload nests cache stats under stats.cache (per
|
||||
// server.mjs handleManagementStatus, ~line 2092). The CacheStore.stats()
|
||||
// contract returns { hits, misses, size, inflightCount } per
|
||||
// lib/cache/store.mjs — there is no `entries` field. Pre-D75 cmdStatus read
|
||||
// `body.stats.cache.entries` (OCP-era) which was always undefined → output
|
||||
// showed "entries=?". Same pattern as D74 P2-3 applied to cmdCache/cmdUsage.
|
||||
if (body.stats?.cache) {
|
||||
const c = body.stats.cache;
|
||||
io.log(` cache: hits=${c.hits ?? 0} misses=${c.misses ?? 0} entries=${c.entries ?? '?'}`);
|
||||
io.log(` cache: hits=${c.hits ?? 0} misses=${c.misses ?? 0} entries=${c.size ?? 0}${typeof c.inflightCount === 'number' ? ` inflight=${c.inflightCount}` : ''}`);
|
||||
}
|
||||
if (Array.isArray(body.recent_errors) && body.recent_errors.length > 0) {
|
||||
io.log(` recent errors: ${body.recent_errors.length}`);
|
||||
|
||||
@@ -214,6 +214,24 @@ alongside the "using server-advertised key" notice.
|
||||
|
||||
---
|
||||
|
||||
## Deployment configurations (D76 amendment, 2026-05-26)
|
||||
|
||||
Original ADR 0011 referenced a `BIND_ADDRESS` concept that did not exist in the v0.4.0–v0.4.2 codebase — the server was hard-coded to `server.listen(PORT, '127.0.0.1', ...)`. D76 closes this gap by adding the `OLP_BIND` env var (default `127.0.0.1`), making the deployment-context discussion below operational rather than aspirational.
|
||||
|
||||
Three deployment configurations are supported:
|
||||
|
||||
| `OLP_BIND` value | Reachability | Anonymous-key publication |
|
||||
|---|---|---|
|
||||
| `127.0.0.1` (default) | Loopback only | Safe with any auth posture (no LAN exposure at all) |
|
||||
| RFC1918 IP / tailnet IP / `0.0.0.0` on a trusted LAN | LAN clients only | Safe when `advertise_anonymous_key: true` — the documented "trusted-LAN" zero-config family onboarding flow |
|
||||
| Public IP / `0.0.0.0` on a public-facing host | Public internet | **Incompatible with `advertise_anonymous_key: true`.** Operator MUST keep `advertise_anonymous_key: false` (default). |
|
||||
|
||||
The server emits a startup warn event `anonymous_key_advertised_with_lan_bind` when `OLP_BIND` is non-loopback AND `advertise_anonymous_key: true` (per the `lib/keys.mjs` + `server.mjs` checks). The warn is a **checkpoint, not a hard gate** — the server cannot tell from the bind address alone whether the operator is on a trusted LAN (RFC1918 / tailnet) or has accidentally exposed a public IP. The Re-evaluation trigger #1 below escalates to a hard gate when OLP gains a public-internet deployment mode.
|
||||
|
||||
`olp-connect <ip>` consumes `/health.anonymousKey` over the network — therefore requires `OLP_BIND` to include the LAN interface on the server side. Without setting `OLP_BIND=<lan-ip>` (or `0.0.0.0`), `olp-connect <ip>` will fail with `connect ECONNREFUSED` because the server only accepts loopback connections.
|
||||
|
||||
---
|
||||
|
||||
## Re-evaluation triggers
|
||||
|
||||
Re-open this ADR when ANY of the following fires:
|
||||
|
||||
+108
-3
@@ -157,6 +157,36 @@ function resolveCodexBin() {
|
||||
// D6 assumption A2: auth file is named auth.json (unconfirmed — D7 will pin).
|
||||
// D6 assumption A3: access token field is `access_token` or `token` (unconfirmed).
|
||||
//
|
||||
// ── D75 (v0.4.2) F1 — codex CLI v0.133.0 schema pin ─────────────────────────
|
||||
//
|
||||
// Real codex CLI v0.133.0 auth.json (verified empirically on PI231 / Mac mini,
|
||||
// 2026-05-26 E2E session):
|
||||
//
|
||||
// {
|
||||
// "auth_mode": "chatgpt",
|
||||
// "OPENAI_API_KEY": null | "<key>",
|
||||
// "tokens": {
|
||||
// "id_token": "<JWT>",
|
||||
// "access_token": "<opaque-or-JWT>", <-- THIS is the access token
|
||||
// "refresh_token": "<opaque>",
|
||||
// "account_id": "<uuid>"
|
||||
// },
|
||||
// "last_refresh": "<ISO8601>"
|
||||
// }
|
||||
//
|
||||
// D6 assumption A3 originally tried `creds.access_token` at TOP level. Under
|
||||
// codex v0.133.0 this field does not exist at the top level → readAuthArtifact()
|
||||
// returned null even when the user had fully completed `codex login`. OLP then
|
||||
// reported "auth artifact missing" via /health and `olp doctor`, and refused
|
||||
// to spawn codex — false negative blocking the entire openai provider.
|
||||
//
|
||||
// Fix: prepend `creds?.tokens?.access_token` to the precedence chain. Keep all
|
||||
// existing fallbacks unchanged so older codex CLI versions (pre-v0.133) and any
|
||||
// future shape variants still resolve.
|
||||
//
|
||||
// Authority pin: codex CLI v0.133.0 source + on-disk auth.json captured during
|
||||
// PI231 E2E. See D75 commit body for the verification transcript.
|
||||
//
|
||||
// Returns { accessToken: string } or null (never throws).
|
||||
export function readAuthArtifact() {
|
||||
// 1. Explicit test override — always takes precedence.
|
||||
@@ -165,7 +195,12 @@ export function readAuthArtifact() {
|
||||
try {
|
||||
const raw = readFileSync(authPathOverride, 'utf8');
|
||||
const creds = JSON.parse(raw);
|
||||
const token = creds?.access_token ?? creds?.token ?? creds?.accessToken;
|
||||
// D75 F1: codex CLI v0.133.0 nests the token under `tokens.access_token`.
|
||||
// Preserve top-level fallbacks for backward / forward compat.
|
||||
const token = creds?.tokens?.access_token
|
||||
?? creds?.access_token
|
||||
?? creds?.token
|
||||
?? creds?.accessToken;
|
||||
if (token && typeof token === 'string') return { accessToken: token };
|
||||
} catch { /* fall through */ }
|
||||
return null; // explicit path set but file missing / malformed
|
||||
@@ -178,8 +213,12 @@ export function readAuthArtifact() {
|
||||
try {
|
||||
const raw = readFileSync(authPath, 'utf8');
|
||||
const creds = JSON.parse(raw);
|
||||
// D6 assumption A3: try common OAuth field names in precedence order.
|
||||
const token = creds?.access_token ?? creds?.token ?? creds?.accessToken;
|
||||
// D75 F1: codex CLI v0.133.0 nests the token under `tokens.access_token`.
|
||||
// Try the nested location FIRST, then fall back to legacy top-level fields.
|
||||
const token = creds?.tokens?.access_token
|
||||
?? creds?.access_token
|
||||
?? creds?.token
|
||||
?? creds?.accessToken;
|
||||
if (token && typeof token === 'string') return { accessToken: token };
|
||||
} catch { /* file missing or malformed */ }
|
||||
|
||||
@@ -242,9 +281,30 @@ export function irToCodex(irRequest) {
|
||||
// model string (e.g., gpt-5.5, gpt-5.4, gpt-5.3-codex).
|
||||
// PROMPT: "Initial instruction for the task. Use '-' to pipe the prompt
|
||||
// from stdin."
|
||||
//
|
||||
// ── D75 (v0.4.2) F2 — codex CLI v0.133.0 trusted-directory sandbox ─────────
|
||||
// codex CLI v0.133.0 added a trusted-directory sandbox: invocations outside
|
||||
// a git repo (or outside any directory explicitly trusted via
|
||||
// `codex config trusted-directories`) refuse with:
|
||||
// "Not inside a trusted directory and --skip-git-repo-check was not specified."
|
||||
// and exit non-zero with zero NDJSON output → OLP surfaces SPAWN_FAILED with
|
||||
// no usable chunks → fallback engine advances to next hop unnecessarily.
|
||||
//
|
||||
// The CWD that OLP spawns from is typically the server install dir (`~/olp/`
|
||||
// on Pi231) which is a git repo on maintainer workstations but is NOT a git
|
||||
// repo on most operator hosts. We bypass the sandbox unconditionally because
|
||||
// OLP is the trusted caller (it is the operator's own server invoking its own
|
||||
// configured Codex subscription via the documented `codex exec` automation
|
||||
// entry point). The trusted-directory sandbox is a foot-gun safeguard for
|
||||
// interactive users; OLP's spawn is non-interactive and pre-authorized.
|
||||
//
|
||||
// Authority: codex CLI v0.133.0 release notes / `codex exec --help` output
|
||||
// documenting `--skip-git-repo-check`. Verified empirically on PI231 E2E
|
||||
// 2026-05-26.
|
||||
const args = [
|
||||
'exec',
|
||||
'--json',
|
||||
'--skip-git-repo-check',
|
||||
'--model', irRequest.model,
|
||||
];
|
||||
|
||||
@@ -291,6 +351,51 @@ export function codexChunkToIR(rawNDJSONLine) {
|
||||
|
||||
if (!event || typeof event !== 'object') return null;
|
||||
|
||||
// ── D75 (v0.4.2) F3 — codex CLI v0.133.0 event shape pin ─────────────────
|
||||
// Real codex CLI v0.133.0 NDJSON event stream (verified empirically on PI231
|
||||
// / Mac mini, 2026-05-26 E2E session):
|
||||
// {"type":"thread.started","thread_id":"019e..."}
|
||||
// {"type":"turn.started"}
|
||||
// {"type":"item.started","item":{"id":"item_0","type":"reasoning","text":""}}
|
||||
// {"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"<response>"}}
|
||||
// {"type":"turn.completed","usage":{"input_tokens":..,"output_tokens":..}}
|
||||
//
|
||||
// The D6 defensive parser recognized `content`/`delta`/`text` fields at the
|
||||
// top level and `type === 'stop'`/`done === true`. None of these match
|
||||
// v0.133.0's actual shape → every chunk was silently dropped → response body
|
||||
// had `content: null`. F3 adds three NEW recognizers (item.completed →
|
||||
// agent_message; turn.completed → stop; turn.failed → error) BEFORE the
|
||||
// legacy fallback chain. Legacy recognizers preserved for forward/backward
|
||||
// compat (older codex versions; future shape variants).
|
||||
|
||||
// F3-a: agent_message item completion.
|
||||
// codex v0.133.0 emits assistant text as a single item.completed event whose
|
||||
// item.type is 'agent_message' and item.text carries the full text. There
|
||||
// are no incremental deltas — the entire response arrives in one chunk.
|
||||
if (event.type === 'item.completed'
|
||||
&& event.item?.type === 'agent_message'
|
||||
&& typeof event.item?.text === 'string') {
|
||||
return { type: 'delta', content: event.item.text };
|
||||
}
|
||||
|
||||
// F3-b: turn completion → stop chunk.
|
||||
// codex v0.133.0 emits turn.completed with a usage block when the model
|
||||
// finishes. We map this to IR stop with finish_reason 'stop'.
|
||||
if (event.type === 'turn.completed') {
|
||||
return { type: 'stop', finish_reason: 'stop' };
|
||||
}
|
||||
|
||||
// F3-c: turn failure → error chunk.
|
||||
// codex v0.133.0 emits turn.failed with an embedded error object when the
|
||||
// turn cannot complete. Extract a human-readable message for the IR error.
|
||||
if (event.type === 'turn.failed') {
|
||||
const errMsg = (typeof event.error === 'string')
|
||||
? event.error
|
||||
: (event.error?.message ?? 'codex turn.failed');
|
||||
return { type: 'error', error: errMsg };
|
||||
}
|
||||
|
||||
// ── Legacy/fallback recognizers (kept for backward + forward compat) ────
|
||||
// Error event: type === 'error' or error field present
|
||||
// A4: defensive — error shape unconfirmed; D7 will pin actual field names
|
||||
if (event.type === 'error' || (event.error && typeof event.error === 'string')) {
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "olp",
|
||||
"version": "0.4.1",
|
||||
"version": "0.4.3",
|
||||
"description": "Personal multi-provider LLM proxy. Successor to OCP. One HTTP endpoint, multiple subscriptions behind it, automatic routing + fallback + caching.",
|
||||
"type": "module",
|
||||
"main": "server.mjs",
|
||||
|
||||
+55
-7
@@ -18,6 +18,14 @@
|
||||
* so OLP can co-host with OCP for migration windows. ADR 0010 §
|
||||
* Default port. Set OLP_PORT=3456 explicitly to restore the
|
||||
* pre-D60 default when not co-hosting with OCP.)
|
||||
* OLP_BIND — listen address (default: 127.0.0.1 since D76 / v0.4.3). Set
|
||||
* to 0.0.0.0 (or a specific interface IP) to accept connections
|
||||
* from LAN clients (required for olp-connect <ip> to actually
|
||||
* reach the server). Server emits a startup warn if BIND
|
||||
* resolves to a non-loopback address AND auth.allow_anonymous
|
||||
* is true (anonymous-key over LAN may be acceptable; anonymous-
|
||||
* key over public internet is not — see ADR 0011 § Deployment
|
||||
* configurations).
|
||||
*/
|
||||
|
||||
import { createServer } from 'node:http';
|
||||
@@ -77,6 +85,14 @@ const pkg = JSON.parse(readFileSync(join(__dirname, 'package.json'), 'utf8'));
|
||||
const VERSION = pkg.version;
|
||||
|
||||
const PORT = parseInt(process.env.OLP_PORT ?? '4567', 10);
|
||||
// F5 / D76: OLP_BIND env. Defaults to 127.0.0.1 (loopback only — secure
|
||||
// default). Operators expose LAN by setting OLP_BIND=0.0.0.0 (or a specific
|
||||
// interface). Per ADR 0011 § Deployment configurations:
|
||||
// - 127.0.0.1: trusted single-machine; safe with any auth posture
|
||||
// - RFC1918 / tailnet / specific LAN IP: trusted-LAN — anonymous_key OK
|
||||
// - 0.0.0.0: ALL interfaces — operator MUST ensure auth posture matches the
|
||||
// network reachability (e.g. no advertise_anonymous_key on public IP)
|
||||
const BIND = process.env.OLP_BIND ?? '127.0.0.1';
|
||||
const BODY_LIMIT = 5 * 1024 * 1024; // 5 MB
|
||||
|
||||
// ── Logging ───────────────────────────────────────────────────────────────
|
||||
@@ -253,6 +269,19 @@ if (_authConfig.advertise_anonymous_key === true) {
|
||||
message: 'auth.advertise_anonymous_key=true but no active key with plaintext_advertise exists. Run `olp-keys keygen --anonymous --advertise` to create one. /health.anonymousKey will NOT be emitted until then. See ADR 0011.',
|
||||
});
|
||||
}
|
||||
// F5 / D76 (ADR 0011 Deployment configurations): publishing the anonymous
|
||||
// key via /health is only safe when /health is reachable ONLY from a
|
||||
// trusted network. If OLP_BIND is set to a non-loopback address AND
|
||||
// advertise_anonymous_key is true, warn the operator. We can't tell from
|
||||
// here whether the non-loopback bind is "trusted LAN" (RFC1918 / tailnet)
|
||||
// or "public internet" — that's the operator's responsibility. The warn is
|
||||
// a checkpoint, not a hard gate.
|
||||
if (BIND !== '127.0.0.1' && BIND !== 'localhost' && BIND !== '::1') {
|
||||
logEvent('warn', 'anonymous_key_advertised_with_lan_bind', {
|
||||
message: `auth.advertise_anonymous_key=true with OLP_BIND=${BIND} — /health.anonymousKey will be reachable from any host that can connect to ${BIND}:${PORT}. Confirm this address is on a trusted LAN (RFC1918 / tailnet) — never a public IP. See ADR 0011 § Deployment configurations.`,
|
||||
bind: BIND,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
/** @internal — test seam: inject a synthetic auth config (no file I/O). */
|
||||
@@ -1224,13 +1253,25 @@ async function handleChatCompletions(req, res) {
|
||||
}
|
||||
|
||||
const chunks = [];
|
||||
// try/finally: releaseSpawn MUST fire on every exit path — success
|
||||
// (return at end), spawn throw (caught and re-thrown below), or the
|
||||
// D16 truncation-salvage return. The finally is the only mechanism
|
||||
// that guarantees release across all three.
|
||||
// D75 (v0.4.2) F7 — per-hop model override.
|
||||
// The chain config (routing.chains[<requested-model>][<hop>].model) is the
|
||||
// model name to pass to THIS hop's provider plugin — NOT the model the
|
||||
// user originally requested. Pre-D75, executeHopFn used hopModel for the
|
||||
// cache key + audit ctx but passed the ORIGINAL irReq (with irReq.model
|
||||
// still set to the user's request) into hopProviderPlugin.spawn(). Result:
|
||||
// a 2-hop chain like [{anthropic, claude-sonnet-4-6}, {openai, gpt-5.5}]
|
||||
// would spawn codex with --model claude-sonnet-4-6 on hop 1 — openai rejects
|
||||
// the unknown model and the chain dies. This broke the core OLP value prop
|
||||
// (cross-provider fallback with provider-appropriate model substitution).
|
||||
//
|
||||
// Fix: build a per-hop IR variant with the hop's model substituted. Skip
|
||||
// the clone when hopModel === irReq.model (single-provider chains and
|
||||
// chain hops whose model matches the request). Authority: ADR 0004 §
|
||||
// Chain advancement step 1 (per-hop config supplies provider AND model).
|
||||
const hopIrReq = irReq.model === hopModel ? irReq : { ...irReq, model: hopModel };
|
||||
try {
|
||||
try {
|
||||
for await (const irChunk of hopProviderPlugin.spawn(irReq, authContext)) {
|
||||
for await (const irChunk of hopProviderPlugin.spawn(hopIrReq, authContext)) {
|
||||
// D16: check error chunks BEFORE pushing — preserves the invariant that
|
||||
// chunks array contains only delta/stop chunks. Without this, the catch
|
||||
// block's `chunks.length > 0` would mistake a single error chunk for
|
||||
@@ -1422,9 +1463,16 @@ async function handleChatCompletions(req, res) {
|
||||
// releaseSpawn fires exactly once in finally — regardless of normal
|
||||
// exhaustion, mid-stream throw, or iterator.return() from cache-layer
|
||||
// sourceAbortController propagation (§9).
|
||||
//
|
||||
// D75 (v0.4.2) F7 — per-hop model override (streaming path).
|
||||
// Mirror the buffered-path fix: pass the hop's configured model (streamModel)
|
||||
// into streamPlugin.spawn(), not the user's original ir.model. Skip the
|
||||
// clone when ir.model === streamModel. See executeHopFn() above for the
|
||||
// full F7 rationale + authority citation.
|
||||
const streamIr = ir.model === streamModel ? ir : { ...ir, model: streamModel };
|
||||
return (async function* sourceWithRelease() {
|
||||
try {
|
||||
for await (const irChunk of streamPlugin.spawn(ir, authContext)) {
|
||||
for await (const irChunk of streamPlugin.spawn(streamIr, authContext)) {
|
||||
yield irChunk;
|
||||
}
|
||||
} finally {
|
||||
@@ -2239,7 +2287,7 @@ const isMain = (() => {
|
||||
|
||||
if (isMain) {
|
||||
const server = createOlpServer();
|
||||
server.listen(PORT, '127.0.0.1', () => {
|
||||
server.listen(PORT, BIND, () => {
|
||||
const enabledCount = loadedProviders.size;
|
||||
// D74 P3-5: banner no longer hardcodes the phase. Derives from VERSION
|
||||
// (which advances at every Phase close) so banner stays accurate
|
||||
|
||||
@@ -14983,4 +14983,419 @@ describe('Suite 36 — D74 v0.4.1 hotfix regression (maintainer review findings)
|
||||
assert.ok(!/Phase 2 in progress/.test(serverSrc));
|
||||
assert.ok(!/Phase 3 in progress/.test(serverSrc));
|
||||
});
|
||||
|
||||
// ── D75 (v0.4.2) extension — real-machine E2E findings F1, F2, F3, F4, F7 ──
|
||||
// Bugs caught on PI231 + Mac mini 2026-05-26 E2E session. Pre-v0.4.0 reviewers
|
||||
// and the post-v0.4.0 maintainer review all missed these because they reviewed
|
||||
// against spec text, not against real provider CLIs running on a remote machine.
|
||||
|
||||
it('36i (F1) — codex readAuthArtifact() recognises v0.133.0 nested tokens.access_token', () => {
|
||||
// Real codex CLI v0.133.0 auth.json shape: { auth_mode, OPENAI_API_KEY,
|
||||
// tokens: { id_token, access_token, refresh_token, account_id }, last_refresh }
|
||||
// Pre-D75 readAuthArtifact read only top-level access_token → returned null →
|
||||
// OLP falsely reported "auth artifact missing" for fully-logged-in users.
|
||||
const tmpDir = _mkdtempSyncS36(_joinS36(_tmpdirS36(), 'olp-d75-f1-nested-'));
|
||||
const authPath = _joinS36(tmpDir, 'auth.json');
|
||||
_writeFileSyncS36(authPath, JSON.stringify({
|
||||
auth_mode: 'chatgpt',
|
||||
OPENAI_API_KEY: null,
|
||||
tokens: {
|
||||
id_token: 'header.body.sig',
|
||||
access_token: 'eyJ-real-nested-access-token-d75-36i',
|
||||
refresh_token: 'refresh-opaque',
|
||||
account_id: '00000000-0000-0000-0000-000000000000',
|
||||
},
|
||||
last_refresh: '2026-05-26T00:00:00Z',
|
||||
}));
|
||||
const savedPath = process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = authPath;
|
||||
try {
|
||||
const result = codexReadAuthArtifact();
|
||||
assert.ok(result, 'must return non-null for v0.133 nested-shape auth.json');
|
||||
assert.equal(result.accessToken, 'eyJ-real-nested-access-token-d75-36i',
|
||||
'must extract token from tokens.access_token (v0.133.0 nested shape)');
|
||||
} finally {
|
||||
if (savedPath !== undefined) process.env.OPENAI_CODEX_AUTH_PATH = savedPath;
|
||||
else delete process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
_rmSyncS36(tmpDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it('36j (F1) — codex readAuthArtifact() still reads legacy top-level access_token', () => {
|
||||
// Backwards compat: older codex CLI versions (pre-v0.133) wrote the token
|
||||
// at top level. D75 fix preserves that fallback so upgrades don't regress.
|
||||
const tmpDir = _mkdtempSyncS36(_joinS36(_tmpdirS36(), 'olp-d75-f1-legacy-'));
|
||||
const authPath = _joinS36(tmpDir, 'auth.json');
|
||||
_writeFileSyncS36(authPath, JSON.stringify({
|
||||
access_token: 'legacy-top-level-token-d75-36j',
|
||||
}));
|
||||
const savedPath = process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = authPath;
|
||||
try {
|
||||
const result = codexReadAuthArtifact();
|
||||
assert.ok(result, 'must return non-null for legacy top-level shape');
|
||||
assert.equal(result.accessToken, 'legacy-top-level-token-d75-36j',
|
||||
'must fall back to top-level access_token when tokens.access_token absent');
|
||||
} finally {
|
||||
if (savedPath !== undefined) process.env.OPENAI_CODEX_AUTH_PATH = savedPath;
|
||||
else delete process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
_rmSyncS36(tmpDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it('36k (F1) — codex readAuthArtifact() returns null for malformed auth.json', () => {
|
||||
// Malformed JSON / missing token / wrong structure → null (not throw).
|
||||
const tmpDir = _mkdtempSyncS36(_joinS36(_tmpdirS36(), 'olp-d75-f1-malformed-'));
|
||||
const savedPath = process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
try {
|
||||
// Case A: unparseable JSON
|
||||
const badJsonPath = _joinS36(tmpDir, 'bad.json');
|
||||
_writeFileSyncS36(badJsonPath, 'not-json-at-all{{{');
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = badJsonPath;
|
||||
assert.equal(codexReadAuthArtifact(), null, 'unparseable JSON → null');
|
||||
|
||||
// Case B: valid JSON but no token field anywhere
|
||||
const noTokenPath = _joinS36(tmpDir, 'no-token.json');
|
||||
_writeFileSyncS36(noTokenPath, JSON.stringify({ auth_mode: 'chatgpt', tokens: {} }));
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = noTokenPath;
|
||||
assert.equal(codexReadAuthArtifact(), null, 'no token in tokens.access_token or top-level → null');
|
||||
|
||||
// Case C: tokens field present but access_token wrong type
|
||||
const wrongTypePath = _joinS36(tmpDir, 'wrong-type.json');
|
||||
_writeFileSyncS36(wrongTypePath, JSON.stringify({
|
||||
tokens: { access_token: 42 },
|
||||
}));
|
||||
process.env.OPENAI_CODEX_AUTH_PATH = wrongTypePath;
|
||||
assert.equal(codexReadAuthArtifact(), null, 'non-string access_token → null');
|
||||
} finally {
|
||||
if (savedPath !== undefined) process.env.OPENAI_CODEX_AUTH_PATH = savedPath;
|
||||
else delete process.env.OPENAI_CODEX_AUTH_PATH;
|
||||
_rmSyncS36(tmpDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it('36l (F2) — irToCodex() includes --skip-git-repo-check before --model', () => {
|
||||
// codex CLI v0.133.0 sandboxes non-git-repo CWDs by default with:
|
||||
// "Not inside a trusted directory and --skip-git-repo-check was not specified."
|
||||
// OLP servers typically deploy outside a git repo on the operator host. The
|
||||
// --skip-git-repo-check flag must be in args, before --model, so the spawn
|
||||
// succeeds without depending on the install-time CWD.
|
||||
const { args } = irToCodex({
|
||||
irVersion: IR_VERSION,
|
||||
model: 'gpt-5.5',
|
||||
stream: false,
|
||||
messages: [{ role: 'user', content: 'hello' }],
|
||||
});
|
||||
assert.ok(args.includes('--skip-git-repo-check'),
|
||||
'irToCodex().args must include --skip-git-repo-check (F2 fix)');
|
||||
const skipIdx = args.indexOf('--skip-git-repo-check');
|
||||
const modelIdx = args.indexOf('--model');
|
||||
assert.ok(skipIdx > 0 && modelIdx > 0, 'both flags must be present');
|
||||
assert.ok(skipIdx < modelIdx,
|
||||
'--skip-git-repo-check must come before --model (cleaner readability + matches codex v0.133 docs ordering)');
|
||||
// Confirm exec is still first and --json still adjacent
|
||||
assert.equal(args[0], 'exec');
|
||||
assert.equal(args[1], '--json');
|
||||
assert.equal(args[2], '--skip-git-repo-check');
|
||||
assert.equal(args[3], '--model');
|
||||
assert.equal(args[4], 'gpt-5.5');
|
||||
});
|
||||
|
||||
it('36m (F3) — codexChunkToIR recognises v0.133.0 item.completed agent_message', () => {
|
||||
// Real codex v0.133.0 emits the assistant text as one item.completed event:
|
||||
// {"type":"item.completed","item":{"id":"item_0","type":"agent_message","text":"hello world"}}
|
||||
// Pre-D75 codexChunkToIR returned null for this shape → response body had
|
||||
// content: null because the chunk was silently dropped.
|
||||
const result = codexChunkToIR(JSON.stringify({
|
||||
type: 'item.completed',
|
||||
item: { id: 'item_0', type: 'agent_message', text: 'hello world from codex' },
|
||||
}));
|
||||
assert.deepEqual(result, { type: 'delta', content: 'hello world from codex' });
|
||||
});
|
||||
|
||||
it('36n (F3) — codexChunkToIR recognises v0.133.0 turn.completed → stop', () => {
|
||||
// turn.completed marks the end of a codex turn. Map to IR stop.
|
||||
const result = codexChunkToIR(JSON.stringify({
|
||||
type: 'turn.completed',
|
||||
usage: { input_tokens: 12, output_tokens: 7 },
|
||||
}));
|
||||
assert.deepEqual(result, { type: 'stop', finish_reason: 'stop' });
|
||||
});
|
||||
|
||||
it('36o (F3) — codexChunkToIR recognises v0.133.0 turn.failed → error', () => {
|
||||
// turn.failed: { type: 'turn.failed', error: { message: 'X' } }
|
||||
const result1 = codexChunkToIR(JSON.stringify({
|
||||
type: 'turn.failed',
|
||||
error: { message: 'rate limit exceeded' },
|
||||
}));
|
||||
assert.equal(result1?.type, 'error');
|
||||
assert.equal(result1.error, 'rate limit exceeded');
|
||||
// turn.failed with string error
|
||||
const result2 = codexChunkToIR(JSON.stringify({
|
||||
type: 'turn.failed',
|
||||
error: 'simple string error',
|
||||
}));
|
||||
assert.equal(result2?.type, 'error');
|
||||
assert.equal(result2.error, 'simple string error');
|
||||
// turn.failed with no error payload → falls back to default message
|
||||
const result3 = codexChunkToIR(JSON.stringify({ type: 'turn.failed' }));
|
||||
assert.equal(result3?.type, 'error');
|
||||
assert.match(result3.error, /turn\.failed/);
|
||||
});
|
||||
|
||||
it('36p (F4) — cmdStatus reads body.stats.cache.size (not OCP-era body.cache.entries)', () => {
|
||||
// Server payload (server.mjs handleManagementStatus ~line 2092):
|
||||
// { ok, version, ..., stats: { total_requests, active_requests, cache: { hits, misses, size, inflightCount } } }
|
||||
// Pre-D75 cmdStatus formatter read `body.cache.entries` (OCP-era) → always
|
||||
// undefined → output showed "entries=?". Pin the wire contract by grepping
|
||||
// the CLI source so a future server-side rename can't silently re-break.
|
||||
const cliSrc = _readFileSyncS36(_joinS36(import.meta.dirname ?? process.cwd(), 'bin/olp.mjs'), 'utf8');
|
||||
// Must contain the cmdStatus block that reads body.stats.cache (not body.cache directly)
|
||||
assert.ok(/cmdStatus\b/.test(cliSrc), 'cmdStatus function must exist');
|
||||
// The status block must read body.stats?.cache (the actual server payload nesting)
|
||||
assert.ok(/body\.stats\?\.cache/.test(cliSrc) || /body\.stats\.cache/.test(cliSrc),
|
||||
'cmdStatus must read body.stats.cache for status display');
|
||||
// Specifically, the "entries=" output must derive from c.size (matches CacheStore.stats() shape)
|
||||
// Find the status cache line and verify it uses c.size
|
||||
const statusBlock = cliSrc.match(/cmdStatus[\s\S]+?(?=async function cmd|^\}\s*$)/m);
|
||||
assert.ok(statusBlock, 'cmdStatus block extractable');
|
||||
assert.ok(/c\.size/.test(statusBlock[0]),
|
||||
'cmdStatus cache line must use c.size (matches CacheStore.stats() shape)');
|
||||
// And must NOT use the OCP-era c.entries
|
||||
assert.ok(!/c\.entries/.test(statusBlock[0]),
|
||||
'cmdStatus cache line must NOT read c.entries (OCP-era field that does not exist in CacheStore.stats())');
|
||||
// Also pin the CacheStore.stats() shape contract (re-affirms 36e but for status path)
|
||||
const cs = new CacheStore();
|
||||
const stats = cs.stats();
|
||||
assert.ok('size' in stats, 'CacheStore.stats() exposes size');
|
||||
assert.ok(!('entries' in stats), 'CacheStore.stats() does NOT expose entries');
|
||||
});
|
||||
|
||||
it('36q (F7) — per-hop chain `model` field overrides IR model on fallback hop', async () => {
|
||||
// Pre-D75: executeHopFn(provider, model, irReq) used `model` for cache key
|
||||
// + audit but passed irReq UNCHANGED to provider.spawn() → the chain
|
||||
// [{anthropic, claude-X}, {openai, gpt-5.5}]
|
||||
// would spawn codex with --model claude-X (the user's original request)
|
||||
// instead of gpt-5.5 (the chain's per-hop config). This broke cross-provider
|
||||
// fallback wherever the chain assigned different models per provider.
|
||||
//
|
||||
// Strategy: inject a mock openai provider that captures the IR it receives.
|
||||
// Force the anthropic hop to fail with SPAWN_FAILED so the chain advances
|
||||
// to openai. Assert the mock saw model === hop's configured model.
|
||||
|
||||
const { __setAuthConfig: setAC36q, __resetAuthConfig: resetAC36q,
|
||||
__setFallbackConfig: setFC36q, __resetFallbackConfig: resetFC36q,
|
||||
createOlpServer, loadedProviders } = await import('./server.mjs');
|
||||
|
||||
setAC36q({ allow_anonymous: true });
|
||||
|
||||
// Capture for mock openai provider
|
||||
let capturedIR = null;
|
||||
const mockOpenAI = {
|
||||
name: 'openai',
|
||||
displayName: 'Mock OpenAI for F7',
|
||||
contractVersion: '1.0',
|
||||
models: ['gpt-5.5'],
|
||||
auth: { type: 'oauth', storage: 'file', path: '/tmp/fake', refresh: 'auto' },
|
||||
async * spawn(ir, _authContext) {
|
||||
// Snapshot the model field as received
|
||||
capturedIR = { model: ir.model };
|
||||
yield { type: 'delta', role: 'assistant', content: 'mock-openai-response' };
|
||||
yield { type: 'stop', finish_reason: 'stop' };
|
||||
},
|
||||
estimateCost: () => null,
|
||||
quotaStatus: async () => null,
|
||||
healthCheck: async () => ({ ok: true, latencyMs: 0 }),
|
||||
hints: { requiresTTY: false, concurrentSpawnSafe: true, maxConcurrent: 4, cacheable: true },
|
||||
};
|
||||
|
||||
// Mock anthropic provider that always fails immediately with SPAWN_FAILED
|
||||
const mockAnthropic = {
|
||||
name: 'anthropic',
|
||||
displayName: 'Mock Anthropic for F7',
|
||||
contractVersion: '1.0',
|
||||
models: ['claude-sonnet-4-6'],
|
||||
auth: { type: 'oauth', storage: 'file', path: '/tmp/fake', refresh: 'auto' },
|
||||
async * spawn(_ir, _authContext) {
|
||||
throw new ProviderError('mock anthropic always fails (F7 test)', 'SPAWN_FAILED');
|
||||
// eslint-disable-next-line no-unreachable
|
||||
yield { type: 'stop' };
|
||||
},
|
||||
estimateCost: () => null,
|
||||
quotaStatus: async () => null,
|
||||
healthCheck: async () => ({ ok: true, latencyMs: 0 }),
|
||||
hints: { requiresTTY: false, concurrentSpawnSafe: true, maxConcurrent: 4, cacheable: true },
|
||||
};
|
||||
|
||||
// Save + clear loadedProviders, install the mocks
|
||||
const savedProviders = new Map(loadedProviders);
|
||||
loadedProviders.clear();
|
||||
loadedProviders.set('anthropic', mockAnthropic);
|
||||
loadedProviders.set('openai', mockOpenAI);
|
||||
|
||||
// 2-hop chain with DIFFERENT models per hop — the bug only manifests when
|
||||
// hopModel !== requested model on the hop that actually serves.
|
||||
setFC36q({
|
||||
chains: {
|
||||
'claude-sonnet-4-6': [
|
||||
{ provider: 'anthropic', model: 'claude-sonnet-4-6' },
|
||||
{ provider: 'openai', model: 'gpt-5.5' },
|
||||
],
|
||||
},
|
||||
soft_triggers: {},
|
||||
});
|
||||
|
||||
const srv = createOlpServer();
|
||||
let port = 28400 + Math.floor(Math.random() * 100);
|
||||
await new Promise((resolve, reject) => {
|
||||
srv.listen(port, '127.0.0.1', resolve);
|
||||
srv.once('error', e => {
|
||||
if (e.code === 'EADDRINUSE') { port++; srv.listen(port, '127.0.0.1', resolve); srv.once('error', reject); }
|
||||
else reject(e);
|
||||
});
|
||||
});
|
||||
|
||||
try {
|
||||
const r = await fetch({
|
||||
port,
|
||||
method: 'POST',
|
||||
path: '/v1/chat/completions',
|
||||
body: {
|
||||
model: 'claude-sonnet-4-6',
|
||||
messages: [{ role: 'user', content: 'F7 test' }],
|
||||
},
|
||||
});
|
||||
assert.equal(r.status, 200, `expected 200, got ${r.status}: ${r.body.slice(0, 300)}`);
|
||||
assert.equal(r.headers['x-olp-provider-used'], 'openai',
|
||||
'openai must be the serving hop (anthropic failed)');
|
||||
// The F7 assertion: the openai hop received the chain's configured model,
|
||||
// NOT the user's original request model.
|
||||
assert.ok(capturedIR, 'mock openai must have been invoked');
|
||||
assert.equal(capturedIR.model, 'gpt-5.5',
|
||||
`F7: openai hop must receive its chain-configured model 'gpt-5.5', got '${capturedIR.model}'`);
|
||||
} finally {
|
||||
await new Promise(r => srv.close(r));
|
||||
resetFC36q();
|
||||
// Restore original providers
|
||||
loadedProviders.clear();
|
||||
for (const [k, v] of savedProviders) loadedProviders.set(k, v);
|
||||
resetAC36q();
|
||||
}
|
||||
});
|
||||
|
||||
it('36r (F7) — per-hop model override preserved on streaming single-hop path', async () => {
|
||||
// The streaming path (server.mjs ~line 1390-1434) has its own spawn call site
|
||||
// inside sourceFactory that previously also passed `ir` unchanged. F7 fix
|
||||
// wraps it with the same { ...ir, model: streamModel } substitution. This
|
||||
// test exercises a single-hop streaming request and asserts the mock saw
|
||||
// the chain's configured model.
|
||||
|
||||
const { __setAuthConfig: setAC36r, __resetAuthConfig: resetAC36r,
|
||||
__setFallbackConfig: setFC36r, __resetFallbackConfig: resetFC36r,
|
||||
createOlpServer, loadedProviders } = await import('./server.mjs');
|
||||
|
||||
setAC36r({ allow_anonymous: true });
|
||||
|
||||
let capturedModel = null;
|
||||
const mockOpenAIStream = {
|
||||
name: 'openai',
|
||||
displayName: 'Mock Codex for F7 streaming',
|
||||
contractVersion: '1.0',
|
||||
models: ['gpt-5.5'],
|
||||
auth: { type: 'oauth', storage: 'file', path: '/tmp/fake', refresh: 'auto' },
|
||||
async * spawn(ir, _authContext) {
|
||||
capturedModel = ir.model;
|
||||
yield { type: 'delta', role: 'assistant', content: 'streamed-content' };
|
||||
yield { type: 'stop', finish_reason: 'stop' };
|
||||
},
|
||||
estimateCost: () => null,
|
||||
quotaStatus: async () => null,
|
||||
healthCheck: async () => ({ ok: true, latencyMs: 0 }),
|
||||
hints: { requiresTTY: false, concurrentSpawnSafe: true, maxConcurrent: 4, cacheable: true },
|
||||
};
|
||||
|
||||
const savedProviders = new Map(loadedProviders);
|
||||
loadedProviders.clear();
|
||||
loadedProviders.set('openai', mockOpenAIStream);
|
||||
|
||||
// Single-hop chain with a DIFFERENT model than the request — the bug is
|
||||
// visible only when the user requests one model name and the chain remaps it.
|
||||
setFC36r({
|
||||
chains: {
|
||||
'gpt-aliased': [
|
||||
{ provider: 'openai', model: 'gpt-5.5' },
|
||||
],
|
||||
},
|
||||
soft_triggers: {},
|
||||
});
|
||||
|
||||
const srv = createOlpServer();
|
||||
let port = 28500 + Math.floor(Math.random() * 100);
|
||||
await new Promise((resolve, reject) => {
|
||||
srv.listen(port, '127.0.0.1', resolve);
|
||||
srv.once('error', e => {
|
||||
if (e.code === 'EADDRINUSE') { port++; srv.listen(port, '127.0.0.1', resolve); srv.once('error', reject); }
|
||||
else reject(e);
|
||||
});
|
||||
});
|
||||
|
||||
try {
|
||||
const r = await fetch({
|
||||
port,
|
||||
method: 'POST',
|
||||
path: '/v1/chat/completions',
|
||||
body: {
|
||||
model: 'gpt-aliased',
|
||||
stream: true,
|
||||
messages: [{ role: 'user', content: 'F7 streaming test' }],
|
||||
},
|
||||
});
|
||||
assert.equal(r.status, 200, `expected 200, got ${r.status}: ${r.body.slice(0, 200)}`);
|
||||
assert.equal(capturedModel, 'gpt-5.5',
|
||||
`F7 streaming: openai must receive chain-configured model 'gpt-5.5', got '${capturedModel}'`);
|
||||
} finally {
|
||||
await new Promise(r => srv.close(r));
|
||||
resetFC36r();
|
||||
loadedProviders.clear();
|
||||
for (const [k, v] of savedProviders) loadedProviders.set(k, v);
|
||||
resetAC36r();
|
||||
}
|
||||
});
|
||||
|
||||
// ── F5 (D76 v0.4.3): OLP_BIND env honored + safety warn ──────────────────
|
||||
it('36s (F5) — server.mjs reads OLP_BIND with safe default 127.0.0.1', () => {
|
||||
// F5 fix: previously bind was hard-coded `server.listen(PORT, '127.0.0.1', ...)`.
|
||||
// D76 adds `const BIND = process.env.OLP_BIND ?? '127.0.0.1'` and the listen
|
||||
// call uses BIND. Pin via source grep so a future refactor that re-introduces
|
||||
// a hardcoded literal in the listen call fails this test.
|
||||
const serverSrc = _readFileSyncS36(_joinS36(import.meta.dirname ?? process.cwd(), 'server.mjs'), 'utf8');
|
||||
assert.ok(/process\.env\.OLP_BIND\s*\?\?\s*['"]127\.0\.0\.1['"]/.test(serverSrc),
|
||||
'server.mjs must read OLP_BIND env with 127.0.0.1 default');
|
||||
assert.ok(/server\.listen\(PORT,\s*BIND\b/.test(serverSrc),
|
||||
'server.mjs server.listen must use the BIND variable (not a hardcoded address)');
|
||||
assert.ok(!/server\.listen\(PORT,\s*['"]127\.0\.0\.1['"]/.test(serverSrc),
|
||||
'server.mjs must NOT pass a hardcoded 127.0.0.1 literal to server.listen anymore');
|
||||
});
|
||||
|
||||
it('36t (F5) — anonymous_key_advertised_with_lan_bind startup warn wiring', () => {
|
||||
// Per ADR 0011 Deployment configurations amendment: when OLP_BIND is
|
||||
// non-loopback AND advertise_anonymous_key is true, server emits a startup
|
||||
// warn so the operator sees the trust-context overlap. Pin the wiring.
|
||||
const serverSrc = _readFileSyncS36(_joinS36(import.meta.dirname ?? process.cwd(), 'server.mjs'), 'utf8');
|
||||
assert.ok(/anonymous_key_advertised_with_lan_bind/.test(serverSrc),
|
||||
'startup warn event name must be present');
|
||||
// Loopback check must compare BIND against ALL three loopback forms
|
||||
assert.ok(/BIND\s*!==\s*['"]127\.0\.0\.1['"][\s\S]{0,200}BIND\s*!==\s*['"]localhost['"][\s\S]{0,200}BIND\s*!==\s*['"]::1['"]/.test(serverSrc),
|
||||
'startup warn must check BIND against all three loopback forms (127.0.0.1 / localhost / ::1)');
|
||||
});
|
||||
|
||||
it('36u (F5) — ADR 0011 Deployment configurations amendment present', () => {
|
||||
const adrPath = _joinS36(import.meta.dirname ?? process.cwd(), 'docs/adr/0011-anonymous-key-deployment-context.md');
|
||||
const adrSrc = _readFileSyncS36(adrPath, 'utf8');
|
||||
assert.ok(/Deployment configurations \(D76 amendment/.test(adrSrc),
|
||||
'ADR 0011 must carry the D76 amendment heading');
|
||||
assert.ok(/OLP_BIND/.test(adrSrc), 'amendment must document OLP_BIND');
|
||||
assert.ok(/anonymous_key_advertised_with_lan_bind/.test(adrSrc),
|
||||
'amendment must cite the startup-warn event name');
|
||||
});
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user