mirror of
https://github.com/dtzp555-max/ocp.git
synced 2026-07-21 21:15:09 +00:00
feat: v2.0.0 — on-demand spawning, session management, full tool access
Major architectural upgrade: - Replace pool system with on-demand spawning (eliminates crash loops, DEGRADED states, and stdin timeout errors from v1.x) - Add session management with --resume support (reduces token waste on multi-turn conversations) - Expand default allowed tools (Bash, Read, Write, Edit, Glob, Grep, WebSearch, WebFetch, Agent) with configurable CLAUDE_ALLOWED_TOOLS - Add system prompt pass-through (CLAUDE_SYSTEM_PROMPT) - Add MCP config support (CLAUDE_MCP_CONFIG) - Add concurrency control (CLAUDE_MAX_CONCURRENT, default 5) - Add periodic auth health monitoring - Add session API endpoints (GET/DELETE /sessions) - Improve /health with full diagnostics (stats, sessions, errors, config) - Fix timeout race condition (graceful SIGTERM → SIGKILL) - Fix ERR_HTTP_HEADERS_SENT by checking headersSent in all response helpers - Document coexistence with Claude Code interactive mode (Telegram, IDE) - No conflict with CC: different ports, protocols, and process models Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -4,13 +4,27 @@
|
|||||||
|
|
||||||
A lightweight, zero-dependency proxy that lets [OpenClaw](https://github.com/openclaw/openclaw) agents talk to Claude through your existing subscription. One command to set up, one file to run.
|
A lightweight, zero-dependency proxy that lets [OpenClaw](https://github.com/openclaw/openclaw) agents talk to Claude through your existing subscription. One command to set up, one file to run.
|
||||||
|
|
||||||
**Why?**
|
## v2.0.0 — Major Upgrade
|
||||||
- **$0 API cost** — uses your Claude Pro/Max subscription, not pay-per-token API
|
|
||||||
- **Zero dependencies** — single Node.js file, no `npm install`
|
**What's new:**
|
||||||
- **One command setup** — `node setup.mjs` handles everything
|
- **On-demand spawning** — eliminates the pool crash loops, DEGRADED states, and stdin timeout errors from v1.x. Each request spawns a fresh `claude -p` process with stdin written immediately. No more stale workers, no more backoff spirals.
|
||||||
- **OpenAI-compatible** — standard `/v1/chat/completions` endpoint
|
- **Session management** — multi-turn conversations use `--resume` to avoid resending full history. Reduces token waste and enables Claude Code's built-in context compression on long conversations.
|
||||||
- **All Claude models** — Opus 4.6, Sonnet 4.6, Haiku 4
|
- **Full tool access** — expanded default tools (Bash, Read, Write, Edit, Glob, Grep, WebSearch, WebFetch, Agent). Configurable via `CLAUDE_ALLOWED_TOOLS` or bypass all checks with `CLAUDE_SKIP_PERMISSIONS=true`.
|
||||||
- **Streaming support** — real-time SSE responses
|
- **System prompt pass-through** — set `CLAUDE_SYSTEM_PROMPT` to inject context into every request.
|
||||||
|
- **MCP config support** — set `CLAUDE_MCP_CONFIG` to load MCP servers (Telegram, etc.) into claude -p calls.
|
||||||
|
- **Concurrency control** — `CLAUDE_MAX_CONCURRENT` prevents runaway process spawning (default: 5).
|
||||||
|
- **Auth health monitoring** — periodic `claude auth status` checks with status exposed on `/health`.
|
||||||
|
- **Session API** — `GET /sessions` to list, `DELETE /sessions` to clear active sessions.
|
||||||
|
- **Improved diagnostics** — `/health` endpoint shows stats, active sessions, recent errors, auth status, and full config.
|
||||||
|
|
||||||
|
**Coexistence with Claude Code interactive mode:**
|
||||||
|
OCP and Claude Code (interactive/Telegram) run on completely different paths and can coexist on the same machine without conflict:
|
||||||
|
- OCP: `localhost:3456` (HTTP) → spawns `claude -p` processes (per-request, stateless)
|
||||||
|
- CC: MCP protocol (in-process) → persistent interactive session
|
||||||
|
- No shared ports, no shared processes, no shared sessions
|
||||||
|
|
||||||
|
**Daemon advantage over CC:**
|
||||||
|
OCP runs as a system daemon (launchd/systemd) that auto-starts on boot and auto-recovers from crashes. Unlike Claude Code interactive mode, OCP does not require a terminal session to stay open — it survives disconnects, reboots, and SSH drops. Combined with OpenClaw's memory system, this means your agents never lose continuity.
|
||||||
|
|
||||||
## How it works
|
## How it works
|
||||||
|
|
||||||
@@ -22,7 +36,7 @@ The proxy translates OpenAI-compatible `/v1/chat/completions` requests into `cla
|
|||||||
|
|
||||||
## Prerequisites
|
## Prerequisites
|
||||||
|
|
||||||
- **Node.js** ≥ 18
|
- **Node.js** >= 18
|
||||||
- **Claude CLI** installed and authenticated (`claude login`)
|
- **Claude CLI** installed and authenticated (`claude login`)
|
||||||
- **OpenClaw** installed
|
- **OpenClaw** installed
|
||||||
|
|
||||||
@@ -49,6 +63,33 @@ openclaw config set agents.defaults.model.primary "claude-local/claude-opus-4-6"
|
|||||||
openclaw gateway restart
|
openclaw gateway restart
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Session Management (v2.0)
|
||||||
|
|
||||||
|
Multi-turn conversations can use sessions to avoid resending full message history on every request.
|
||||||
|
|
||||||
|
**How to enable:** Include a `session_id` or `conversation_id` field in your request body, or set the `X-Session-Id` / `X-Conversation-Id` header.
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"model": "claude-opus-4-6",
|
||||||
|
"session_id": "conv-abc-123",
|
||||||
|
"messages": [
|
||||||
|
{"role": "user", "content": "Hello"},
|
||||||
|
{"role": "assistant", "content": "Hi there!"},
|
||||||
|
{"role": "user", "content": "What did I just say?"}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**First request** with a new session_id: all messages are sent, session is persisted via `--session-id`.
|
||||||
|
**Subsequent requests** with the same session_id: only the latest user message is sent via `--resume`, reducing token consumption.
|
||||||
|
|
||||||
|
Sessions expire after 1 hour of inactivity (configurable via `CLAUDE_SESSION_TTL`).
|
||||||
|
|
||||||
|
**API endpoints:**
|
||||||
|
- `GET /sessions` — list all active sessions
|
||||||
|
- `DELETE /sessions` — clear all sessions
|
||||||
|
|
||||||
## Security
|
## Security
|
||||||
|
|
||||||
- **Localhost only** — the proxy binds to `127.0.0.1` and is not exposed to the internet or your local network
|
- **Localhost only** — the proxy binds to `127.0.0.1` and is not exposed to the internet or your local network
|
||||||
@@ -123,105 +164,77 @@ openclaw gateway restart
|
|||||||
|
|
||||||
| Model ID | Claude CLI model | Notes |
|
| Model ID | Claude CLI model | Notes |
|
||||||
|----------|-----------------|-------|
|
|----------|-----------------|-------|
|
||||||
| `claude-opus-4-6` | opus | Most capable, slower |
|
| `claude-opus-4-6` | claude-opus-4-6 | Most capable, slower |
|
||||||
| `claude-sonnet-4-6` | sonnet | Good balance of speed/quality |
|
| `claude-sonnet-4-6` | claude-sonnet-4-6 | Good balance of speed/quality |
|
||||||
| `claude-haiku-4` | haiku | Fastest, lightweight |
|
| `claude-haiku-4` | claude-haiku-4-5-20251001 | Fastest, lightweight |
|
||||||
|
|
||||||
## Authentication
|
|
||||||
The proxy supports optional Bearer token authentication via the `PROXY_API_KEY` environment variable.
|
|
||||||
|
|
||||||
**When `PROXY_API_KEY` is set**, all requests (except `GET /health`) must include a valid `Authorization: Bearer <token>` header. Requests with a missing or invalid token receive a `401 Unauthorized` response.
|
|
||||||
|
|
||||||
**When `PROXY_API_KEY` is not set**, authentication is disabled and all requests are accepted (backwards compatible with v1.6.x and earlier).
|
|
||||||
|
|
||||||
For local Claude Max / OAuth usage, avoid exporting `ANTHROPIC_API_KEY` globally into the shell that launches the proxy. `v1.7.1` also sanitizes `ANTHROPIC_*` variables before spawning Claude subprocesses to reduce accidental auth-path pollution.
|
|
||||||
|
|
||||||
|
|
||||||
The proxy supports optional Bearer token authentication via the `PROXY_API_KEY` environment variable.
|
|
||||||
|
|
||||||
**When `PROXY_API_KEY` is set**, all requests (except `GET /health`) must include a valid `Authorization: Bearer <token>` header. Requests with a missing or invalid token receive a `401 Unauthorized` response.
|
|
||||||
|
|
||||||
**When `PROXY_API_KEY` is not set**, authentication is disabled and all requests are accepted (backwards compatible with v1.6.x and earlier).
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Start with auth enabled
|
|
||||||
PROXY_API_KEY=my-secret-token node server.mjs
|
|
||||||
|
|
||||||
# Configure OpenClaw provider with the matching key
|
|
||||||
# In openclaw.json, set apiKey under the claude-local provider:
|
|
||||||
"claude-local": {
|
|
||||||
"baseUrl": "http://127.0.0.1:3456/v1",
|
|
||||||
"api": "openai-completions",
|
|
||||||
"apiKey": "my-secret-token",
|
|
||||||
...
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
The proxy logs auth status on startup: `Auth: enabled (PROXY_API_KEY set)` or `Auth: disabled (no PROXY_API_KEY)`.
|
|
||||||
|
|
||||||
## Environment Variables
|
## Environment Variables
|
||||||
|
|
||||||
| Variable | Default | Description |
|
| Variable | Default | Description |
|
||||||
|----------|---------|-------------|
|
|----------|---------|-------------|
|
||||||
| `CLAUDE_PROXY_PORT` | `3456` | Listen port |
|
| `CLAUDE_PROXY_PORT` | `3456` | Listen port |
|
||||||
| `CLAUDE_BIN` | `claude` | Path to claude binary |
|
| `CLAUDE_BIN` | *(auto-detect)* | Path to claude binary |
|
||||||
| `CLAUDE_TIMEOUT` | `120000` | Request timeout (ms) |
|
| `CLAUDE_TIMEOUT` | `300000` | Request timeout (ms) |
|
||||||
| `PROXY_API_KEY` | *(unset)* | Bearer token for API authentication; when unset, auth is disabled |
|
| `CLAUDE_ALLOWED_TOOLS` | `Bash,Read,...,Agent` | Comma-separated tools to pre-approve |
|
||||||
|
| `CLAUDE_SKIP_PERMISSIONS` | `false` | Set `true` to bypass all permission checks |
|
||||||
|
| `CLAUDE_SYSTEM_PROMPT` | *(empty)* | System prompt appended to all requests |
|
||||||
|
| `CLAUDE_MCP_CONFIG` | *(empty)* | Path to MCP server config JSON file |
|
||||||
|
| `CLAUDE_SESSION_TTL` | `3600000` | Session expiry in ms (default: 1 hour) |
|
||||||
|
| `CLAUDE_MAX_CONCURRENT` | `5` | Max concurrent claude processes |
|
||||||
|
| `PROXY_API_KEY` | *(unset)* | Bearer token for API authentication |
|
||||||
|
|
||||||
## API Endpoints
|
## API Endpoints
|
||||||
|
|
||||||
- `GET /v1/models` — List available models
|
- `GET /v1/models` — List available models
|
||||||
- `POST /v1/chat/completions` — Chat completion (streaming + non-streaming)
|
- `POST /v1/chat/completions` — Chat completion (streaming + non-streaming)
|
||||||
- `GET /health` — Health check
|
- `GET /health` — Comprehensive health check (stats, sessions, auth, config)
|
||||||
|
- `GET /sessions` — List active sessions
|
||||||
|
- `DELETE /sessions` — Clear all sessions
|
||||||
|
|
||||||
## Server / Advanced: Docker
|
## Authentication
|
||||||
|
|
||||||
For server deployments or if you prefer Docker:
|
The proxy supports optional Bearer token authentication via the `PROXY_API_KEY` environment variable.
|
||||||
|
|
||||||
|
**When `PROXY_API_KEY` is set**, all requests (except `GET /health`) must include a valid `Authorization: Bearer <token>` header. Requests with a missing or invalid token receive a `401 Unauthorized` response.
|
||||||
|
|
||||||
|
**When `PROXY_API_KEY` is not set**, authentication is disabled and all requests are accepted.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
git clone https://github.com/dtzp555-max/openclaw-claude-proxy.git
|
# Start with auth enabled
|
||||||
cd openclaw-claude-proxy
|
PROXY_API_KEY=my-secret-token node server.mjs
|
||||||
cp .env.example .env # add your CLAUDE_SESSION_TOKEN / CLAUDE_COOKIES
|
|
||||||
docker compose up -d
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Or as a single command if you already have a `.env` ready:
|
## Architecture: v1 vs v2
|
||||||
|
|
||||||
```bash
|
| | v1.x (pool) | v2.0 (on-demand) |
|
||||||
git clone https://github.com/dtzp555-max/openclaw-claude-proxy.git && cd openclaw-claude-proxy && docker compose up -d
|
|---|---|---|
|
||||||
```
|
| Process lifecycle | Pre-spawn idle workers | Spawn per request |
|
||||||
|
| Crash handling | Backoff → DEGRADED → manual restart | No crash loops (no idle workers) |
|
||||||
|
| Session support | None (stateless) | --resume with session tracking |
|
||||||
|
| Tool access | 6 tools hardcoded | Configurable, expanded defaults |
|
||||||
|
| System prompt | None | CLAUDE_SYSTEM_PROMPT env |
|
||||||
|
| MCP support | None | CLAUDE_MCP_CONFIG env |
|
||||||
|
| Concurrency | Unlimited (dangerous) | CLAUDE_MAX_CONCURRENT limit |
|
||||||
|
| Auth monitoring | None | Periodic health checks |
|
||||||
|
| Diagnostics | Basic /health | Full stats, sessions, errors |
|
||||||
|
|
||||||
Health check: `curl http://localhost:3456/health`
|
## Coexistence with Claude Code
|
||||||
|
|
||||||
|
OCP and Claude Code interactive mode (including Telegram bots) are completely independent:
|
||||||
|
|
||||||
|
| | OCP (this proxy) | CC interactive |
|
||||||
|
|---|---|---|
|
||||||
|
| Protocol | HTTP (localhost:3456) | MCP (in-process) |
|
||||||
|
| Process model | Per-request spawn | Persistent session |
|
||||||
|
| Lifecycle | Daemon (auto-start, auto-recover) | Requires terminal |
|
||||||
|
| Permission model | Pre-approved tools | Interactive prompts |
|
||||||
|
| Use case | Automated agent work | Human-in-the-loop |
|
||||||
|
|
||||||
|
Both can run on the same machine simultaneously. No shared state, no port conflicts.
|
||||||
|
|
||||||
## Recovery after OpenClaw upgrade
|
## Recovery after OpenClaw upgrade
|
||||||
|
|
||||||
OpenClaw upgrades (`npm update -g openclaw`) **do not overwrite** the user config at `~/.openclaw/openclaw.json`. However, if the claude-local models stop working after an upgrade, follow these steps:
|
OpenClaw upgrades (`npm update -g openclaw`) **do not overwrite** the user config at `~/.openclaw/openclaw.json`. However, if the claude-local models stop working after an upgrade:
|
||||||
|
|
||||||
### Quick diagnosis
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# 1. Check if proxy is running
|
|
||||||
curl http://127.0.0.1:3456/health
|
|
||||||
# Expected: {"status":"ok"}
|
|
||||||
|
|
||||||
# 2. Verify Claude CLI works
|
|
||||||
claude -p "hello" --model sonnet --output-format text
|
|
||||||
# Expected: text response
|
|
||||||
|
|
||||||
# 3. Verify OpenClaw config
|
|
||||||
cat ~/.openclaw/openclaw.json | grep -A3 claude-local
|
|
||||||
```
|
|
||||||
|
|
||||||
### Common issues and fixes
|
|
||||||
|
|
||||||
| Symptom | Cause | Fix |
|
|
||||||
|---------|-------|-----|
|
|
||||||
| Agent doesn't reply, no proxy logs | Gateway didn't load claude-local provider | Check `models.providers.claude-local` in `openclaw.json` |
|
|
||||||
| Proxy reports `exit 1` | Claude CLI not logged in or token expired | Run `claude login` to re-authenticate |
|
|
||||||
| `🔑 unknown` in `/status` | Normal — no API key, using OAuth | Does not affect functionality, safe to ignore |
|
|
||||||
| `/status` shows Context 0% | Messages not reaching proxy (SSE format issue) | Ensure proxy is latest version with streaming support |
|
|
||||||
| Gateway reports `invalid api type` | OpenClaw renamed API type in new version | Check `api` field is still valid (e.g., `openai-completions`) |
|
|
||||||
| Proxy startup `EADDRINUSE` | Port 3456 already in use | `lsof -i :3456` to find and kill the old process |
|
|
||||||
|
|
||||||
### One-command recovery
|
### One-command recovery
|
||||||
|
|
||||||
@@ -232,19 +245,12 @@ node setup.mjs # reconfigure OpenClaw + start proxy
|
|||||||
openclaw gateway restart
|
openclaw gateway restart
|
||||||
```
|
```
|
||||||
|
|
||||||
### Pre-upgrade backup (recommended)
|
|
||||||
|
|
||||||
```bash
|
|
||||||
cp ~/.openclaw/openclaw.json ~/.openclaw/openclaw.json.bak
|
|
||||||
```
|
|
||||||
|
|
||||||
## Notes
|
## Notes
|
||||||
|
|
||||||
- Cost shows as $0 because billing goes through your Claude subscription
|
- Cost shows as $0 because billing goes through your Claude subscription
|
||||||
- The `🔑` field in `/status` shows the configured auth key (or `unknown` if auth is disabled) — this is normal
|
- Each request spawns a `claude -p` process; concurrent requests are capped by `CLAUDE_MAX_CONCURRENT`
|
||||||
- Each request spawns a `claude -p` process; concurrent requests are supported
|
|
||||||
- The proxy must run on the same machine as the Claude CLI (uses local OAuth)
|
- The proxy must run on the same machine as the Claude CLI (uses local OAuth)
|
||||||
- The same Claude account can be used on multiple machines (shared usage quota)
|
- Session data is stored by Claude CLI on disk; session map is in-memory (lost on proxy restart)
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
{
|
{
|
||||||
"name": "openclaw-claude-proxy",
|
"name": "openclaw-claude-proxy",
|
||||||
"version": "1.8.0",
|
"version": "2.0.0",
|
||||||
"description": "OpenAI-compatible proxy that routes requests through Claude CLI \u2014 use your Claude Pro/Max subscription as an OpenClaw model provider",
|
"description": "OpenAI-compatible proxy that routes requests through Claude CLI \u2014 use your Claude Pro/Max subscription as an OpenClaw model provider",
|
||||||
"type": "module",
|
"type": "module",
|
||||||
"bin": {
|
"bin": {
|
||||||
|
|||||||
+262
-257
@@ -1,21 +1,29 @@
|
|||||||
#!/usr/bin/env node
|
#!/usr/bin/env node
|
||||||
/**
|
/**
|
||||||
* openclaw-claude-proxy — OpenAI-compatible proxy for Claude CLI
|
* openclaw-claude-proxy v2.0.0 — OpenAI-compatible proxy for Claude CLI
|
||||||
*
|
*
|
||||||
* Translates OpenAI chat/completions requests into `claude -p` CLI calls,
|
* Translates OpenAI chat/completions requests into `claude -p` CLI calls,
|
||||||
* letting you use your Claude Pro/Max subscription as an OpenClaw model provider.
|
* letting you use your Claude Pro/Max subscription as an OpenClaw model provider.
|
||||||
*
|
*
|
||||||
* Features:
|
* v2.0.0 highlights:
|
||||||
* - Process pool: pre-spawns CLI processes to eliminate cold start latency
|
* - On-demand spawning: eliminates pool crash loops from v1.x
|
||||||
* - SSE streaming + non-streaming responses
|
* - Session management: --resume support reduces token waste on multi-turn
|
||||||
* - Concurrent request support
|
* - Full tool access: configurable allowedTools (expanded defaults)
|
||||||
|
* - System prompt & MCP config pass-through
|
||||||
|
* - Concurrency control with queuing
|
||||||
|
* - Coexists safely with Claude Code interactive mode (Telegram, IDE, etc.)
|
||||||
*
|
*
|
||||||
* Env vars:
|
* Env vars:
|
||||||
* CLAUDE_PROXY_PORT — listen port (default: 3456)
|
* CLAUDE_PROXY_PORT — listen port (default: 3456)
|
||||||
* CLAUDE_BIN — path to claude binary (default: "claude")
|
* CLAUDE_BIN — path to claude binary (default: auto-detect)
|
||||||
* CLAUDE_TIMEOUT — per-request timeout in ms (default: 120000)
|
* CLAUDE_TIMEOUT — per-request timeout in ms (default: 300000)
|
||||||
* CLAUDE_POOL_SIZE — warm process pool size per model (default: 1)
|
* CLAUDE_ALLOWED_TOOLS — comma-separated tools to allow (default: expanded set)
|
||||||
* PROXY_API_KEY — Bearer token for API authentication (optional, if unset auth is disabled)
|
* CLAUDE_SKIP_PERMISSIONS — "true" to bypass all permission checks (default: false)
|
||||||
|
* CLAUDE_SYSTEM_PROMPT — system prompt appended to all requests
|
||||||
|
* CLAUDE_MCP_CONFIG — path to MCP server config JSON file
|
||||||
|
* CLAUDE_SESSION_TTL — session TTL in ms (default: 3600000 = 1h)
|
||||||
|
* CLAUDE_MAX_CONCURRENT — max concurrent claude processes (default: 5)
|
||||||
|
* PROXY_API_KEY — Bearer token for API auth (optional)
|
||||||
*/
|
*/
|
||||||
import { createServer } from "node:http";
|
import { createServer } from "node:http";
|
||||||
import { spawn, execFileSync } from "node:child_process";
|
import { spawn, execFileSync } from "node:child_process";
|
||||||
@@ -64,26 +72,29 @@ function resolveClaude() {
|
|||||||
process.exit(1);
|
process.exit(1);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ── Configuration ───────────────────────────────────────────────────────
|
||||||
const PORT = parseInt(process.env.CLAUDE_PROXY_PORT || "3456", 10);
|
const PORT = parseInt(process.env.CLAUDE_PROXY_PORT || "3456", 10);
|
||||||
const CLAUDE = resolveClaude();
|
const CLAUDE = resolveClaude();
|
||||||
const TIMEOUT = parseInt(process.env.CLAUDE_TIMEOUT || "300000", 10);
|
const TIMEOUT = parseInt(process.env.CLAUDE_TIMEOUT || "300000", 10);
|
||||||
const POOL_SIZE = parseInt(process.env.CLAUDE_POOL_SIZE || "1", 10);
|
|
||||||
const POOL_MAX_IDLE = parseInt(process.env.CLAUDE_POOL_MAX_IDLE || "60000", 10); // max idle time before recycle
|
|
||||||
const PROXY_API_KEY = process.env.PROXY_API_KEY || "";
|
const PROXY_API_KEY = process.env.PROXY_API_KEY || "";
|
||||||
|
const SKIP_PERMISSIONS = process.env.CLAUDE_SKIP_PERMISSIONS === "true";
|
||||||
|
const ALLOWED_TOOLS = (process.env.CLAUDE_ALLOWED_TOOLS ||
|
||||||
|
"Bash,Read,Write,Edit,Glob,Grep,WebSearch,WebFetch,Agent"
|
||||||
|
).split(",").map(s => s.trim()).filter(Boolean);
|
||||||
|
const SYSTEM_PROMPT = process.env.CLAUDE_SYSTEM_PROMPT || "";
|
||||||
|
const MCP_CONFIG = process.env.CLAUDE_MCP_CONFIG || "";
|
||||||
|
const SESSION_TTL = parseInt(process.env.CLAUDE_SESSION_TTL || "3600000", 10);
|
||||||
|
const MAX_CONCURRENT = parseInt(process.env.CLAUDE_MAX_CONCURRENT || "5", 10);
|
||||||
|
|
||||||
const VERSION = _pkg.version;
|
const VERSION = _pkg.version;
|
||||||
const START_TIME = Date.now();
|
const START_TIME = Date.now();
|
||||||
|
|
||||||
// Model alias mapping: request model → claude CLI --model arg
|
// ── Model mapping ───────────────────────────────────────────────────────
|
||||||
// Maps both shorthand aliases AND full model IDs to the canonical full model ID
|
// Maps request model IDs and aliases to canonical claude CLI model IDs.
|
||||||
// that the claude CLI accepts. Using short names like "sonnet"/"opus"/"haiku"
|
|
||||||
// causes the CLI to reject the --model arg and crash immediately.
|
|
||||||
const MODEL_MAP = {
|
const MODEL_MAP = {
|
||||||
// Full canonical IDs (pass through as-is)
|
|
||||||
"claude-opus-4-6": "claude-opus-4-6",
|
"claude-opus-4-6": "claude-opus-4-6",
|
||||||
"claude-sonnet-4-6": "claude-sonnet-4-6",
|
"claude-sonnet-4-6": "claude-sonnet-4-6",
|
||||||
"claude-haiku-4-5-20251001": "claude-haiku-4-5-20251001",
|
"claude-haiku-4-5-20251001": "claude-haiku-4-5-20251001",
|
||||||
// Short aliases → full canonical IDs
|
|
||||||
"claude-opus-4": "claude-opus-4-6",
|
"claude-opus-4": "claude-opus-4-6",
|
||||||
"claude-haiku-4": "claude-haiku-4-5-20251001",
|
"claude-haiku-4": "claude-haiku-4-5-20251001",
|
||||||
"claude-haiku-4-5": "claude-haiku-4-5-20251001",
|
"claude-haiku-4-5": "claude-haiku-4-5-20251001",
|
||||||
@@ -98,270 +109,228 @@ const MODELS = [
|
|||||||
{ id: "claude-haiku-4-5-20251001", name: "Claude Haiku 4.5" },
|
{ id: "claude-haiku-4-5-20251001", name: "Claude Haiku 4.5" },
|
||||||
];
|
];
|
||||||
|
|
||||||
// ── Process Pool ──────────────────────────────────────────────────────────
|
// ── Session management ──────────────────────────────────────────────────
|
||||||
// Pre-spawns `claude -p` processes that read prompts from stdin.
|
// Maps conversation IDs (from caller) to Claude CLI session UUIDs.
|
||||||
// When a request arrives, we grab a warm process and pipe the prompt in.
|
// Enables --resume for multi-turn conversations, reducing token waste.
|
||||||
// After the process finishes, a new one is spawned to replace it.
|
const sessions = new Map(); // conversationId → { uuid, messageCount, lastUsed, model }
|
||||||
|
|
||||||
const pool = new Map(); // model → [{ proc, ready }]
|
setInterval(() => {
|
||||||
|
const now = Date.now();
|
||||||
|
for (const [id, s] of sessions) {
|
||||||
|
if (now - s.lastUsed > SESSION_TTL) {
|
||||||
|
sessions.delete(id);
|
||||||
|
console.log(`[session] expired ${id.slice(0, 12)}... (idle ${Math.round((now - s.lastUsed) / 60000)}m)`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}, 60000);
|
||||||
|
|
||||||
// Exponential backoff state per model: tracks consecutive fast failures
|
// ── Stats & diagnostics ─────────────────────────────────────────────────
|
||||||
// to prevent a tight spawn/die loop when workers crash on startup.
|
const stats = {
|
||||||
// Delays: 2s base, doubled each failure, capped at 60s.
|
totalRequests: 0,
|
||||||
// After 5 consecutive fast crashes (each lived < 10s, all within 60s),
|
activeRequests: 0,
|
||||||
// the model is marked "degraded" and respawning stops entirely.
|
errors: 0,
|
||||||
const poolBackoff = new Map(); // model → { failures: number, timer: TimeoutId|null, degraded: boolean, windowStart: number }
|
timeouts: 0,
|
||||||
|
sessionHits: 0,
|
||||||
|
sessionMisses: 0,
|
||||||
|
oneOffRequests: 0,
|
||||||
|
};
|
||||||
|
const recentErrors = []; // last 20 errors
|
||||||
|
|
||||||
function spawnWarm(cliModel) {
|
function trackError(msg) {
|
||||||
|
stats.errors++;
|
||||||
|
recentErrors.push({ time: new Date().toISOString(), message: String(msg).slice(0, 200) });
|
||||||
|
if (recentErrors.length > 20) recentErrors.shift();
|
||||||
|
}
|
||||||
|
|
||||||
|
// ── Auth health check ───────────────────────────────────────────────────
|
||||||
|
let authStatus = { ok: null, lastCheck: 0, message: "" };
|
||||||
|
|
||||||
|
async function checkAuth() {
|
||||||
|
try {
|
||||||
const env = { ...process.env };
|
const env = { ...process.env };
|
||||||
delete env.CLAUDECODE;
|
delete env.CLAUDECODE;
|
||||||
delete env.ANTHROPIC_API_KEY;
|
delete env.ANTHROPIC_API_KEY;
|
||||||
delete env.ANTHROPIC_BASE_URL;
|
delete env.ANTHROPIC_BASE_URL;
|
||||||
delete env.ANTHROPIC_AUTH_TOKEN;
|
delete env.ANTHROPIC_AUTH_TOKEN;
|
||||||
|
execFileSync(CLAUDE, ["auth", "status"], { encoding: "utf8", timeout: 10000, env });
|
||||||
const proc = spawn(CLAUDE, [
|
authStatus = { ok: true, lastCheck: Date.now(), message: "authenticated" };
|
||||||
"-p", "--model", cliModel,
|
} catch (e) {
|
||||||
"--output-format", "text",
|
const msg = (e.stderr || e.message || "").slice(0, 200);
|
||||||
"--no-session-persistence",
|
authStatus = { ok: false, lastCheck: Date.now(), message: msg };
|
||||||
"--allowedTools", "Bash", "Read", "Write", "Edit", "Glob", "Grep",
|
console.error(`[auth] check failed: ${msg}`);
|
||||||
], { env, stdio: ["pipe", "pipe", "pipe"] });
|
|
||||||
|
|
||||||
const entry = { proc, cliModel, ready: true, spawnedAt: Date.now() };
|
|
||||||
|
|
||||||
// Capture stderr from pool workers so crash reasons are visible in logs
|
|
||||||
let stderrBuf = "";
|
|
||||||
proc.stderr.on("data", (d) => {
|
|
||||||
stderrBuf += d;
|
|
||||||
if (stderrBuf.length > 500) stderrBuf = stderrBuf.slice(-500);
|
|
||||||
});
|
|
||||||
|
|
||||||
proc.on("error", (err) => {
|
|
||||||
console.error(`[pool] spawn error model=${cliModel}: ${err.message}`);
|
|
||||||
entry.ready = false;
|
|
||||||
});
|
|
||||||
|
|
||||||
proc.on("exit", (code) => {
|
|
||||||
const livedMs = Date.now() - entry.spawnedAt;
|
|
||||||
entry.ready = false;
|
|
||||||
|
|
||||||
// Log stderr from crashed pool worker (first 500 chars) to aid debugging
|
|
||||||
if (stderrBuf.trim()) {
|
|
||||||
console.error(`[pool] worker stderr model=${cliModel} exit=${code} lived=${livedMs}ms: ${stderrBuf.slice(0, 500)}`);
|
|
||||||
}
|
|
||||||
|
|
||||||
// If the process survived > 10s, it was healthy — reset the backoff counter and window
|
|
||||||
if (livedMs > 10000) {
|
|
||||||
const state = poolBackoff.get(cliModel);
|
|
||||||
if (state && (state.failures > 0 || state.degraded)) {
|
|
||||||
console.log(`[pool] resetting backoff for model=${cliModel} (lived ${livedMs}ms)`);
|
|
||||||
state.failures = 0;
|
|
||||||
state.degraded = false;
|
|
||||||
state.windowStart = Date.now();
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// Remove from pool
|
// Check auth on start and every 10 minutes
|
||||||
const arr = pool.get(cliModel);
|
checkAuth();
|
||||||
if (arr) {
|
setInterval(checkAuth, 600000);
|
||||||
const idx = arr.indexOf(entry);
|
|
||||||
if (idx !== -1) arr.splice(idx, 1);
|
|
||||||
}
|
|
||||||
// Replenish: treat as crash (apply backoff) only if it died fast (< 10s)
|
|
||||||
const isCrash = livedMs <= 10000;
|
|
||||||
replenishPool(cliModel, isCrash);
|
|
||||||
});
|
|
||||||
|
|
||||||
return entry;
|
// ── Build CLI arguments ─────────────────────────────────────────────────
|
||||||
|
function buildCliArgs(cliModel, sessionInfo) {
|
||||||
|
const args = ["-p", "--model", cliModel, "--output-format", "text"];
|
||||||
|
|
||||||
|
// Session handling
|
||||||
|
if (sessionInfo?.resume) {
|
||||||
|
args.push("--resume", sessionInfo.uuid);
|
||||||
|
} else if (sessionInfo?.uuid) {
|
||||||
|
args.push("--session-id", sessionInfo.uuid);
|
||||||
|
} else {
|
||||||
|
args.push("--no-session-persistence");
|
||||||
}
|
}
|
||||||
|
|
||||||
// Recycle idle processes to prevent stale connections
|
// Permissions
|
||||||
function recycleStaleProcesses() {
|
if (SKIP_PERMISSIONS) {
|
||||||
const now = Date.now();
|
args.push("--dangerously-skip-permissions");
|
||||||
for (const [cliModel, arr] of pool) {
|
} else if (ALLOWED_TOOLS.length > 0) {
|
||||||
for (const entry of arr) {
|
args.push("--allowedTools", ...ALLOWED_TOOLS);
|
||||||
if (entry.ready && (now - entry.spawnedAt) > POOL_MAX_IDLE) {
|
|
||||||
console.log(`[pool] recycling stale process model=${cliModel} (idle ${Math.round((now - entry.spawnedAt) / 1000)}s)`);
|
|
||||||
entry.ready = false;
|
|
||||||
entry.proc.kill();
|
|
||||||
// exit handler will replenish
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
setInterval(recycleStaleProcesses, 15000); // check every 15s
|
// System prompt
|
||||||
|
if (SYSTEM_PROMPT) {
|
||||||
const BACKOFF_BASE_MS = 2000; // 2s starting delay
|
args.push("--append-system-prompt", SYSTEM_PROMPT);
|
||||||
const BACKOFF_MAX_MS = 60000; // 60s ceiling
|
|
||||||
const CRASH_LIMIT = 5; // max consecutive fast crashes before degraded
|
|
||||||
const CRASH_WINDOW_MS = 60000; // window for counting consecutive fast crashes (60s)
|
|
||||||
|
|
||||||
// replenishPool(cliModel, isCrash)
|
|
||||||
// isCrash=false → initial or manual fill, no backoff applied
|
|
||||||
// isCrash=true → called from exit handler after a fast crash
|
|
||||||
function replenishPool(cliModel, isCrash = false) {
|
|
||||||
if (!pool.has(cliModel)) pool.set(cliModel, []);
|
|
||||||
if (!poolBackoff.has(cliModel)) poolBackoff.set(cliModel, { failures: 0, timer: null, degraded: false, windowStart: Date.now() });
|
|
||||||
|
|
||||||
const arr = pool.get(cliModel);
|
|
||||||
const state = poolBackoff.get(cliModel);
|
|
||||||
|
|
||||||
// If this model is degraded (too many consecutive fast crashes), stop respawning
|
|
||||||
if (state.degraded) {
|
|
||||||
console.error(`[pool] DEGRADED: model=${cliModel} will not be respawned. Restart the proxy to retry.`);
|
|
||||||
return;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
const alive = arr.filter((e) => e.ready).length;
|
// MCP config
|
||||||
const needed = POOL_SIZE - alive;
|
if (MCP_CONFIG) {
|
||||||
if (needed <= 0) return;
|
args.push("--mcp-config", MCP_CONFIG);
|
||||||
|
|
||||||
// Cancel any pending backoff timer for this model before scheduling a new one
|
|
||||||
if (state.timer) {
|
|
||||||
clearTimeout(state.timer);
|
|
||||||
state.timer = null;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// Only track failures and apply backoff when this is a crash respawn
|
return args;
|
||||||
if (!isCrash) {
|
|
||||||
// Immediate spawn — no backoff on initial fill or manual replenish
|
|
||||||
const currentAlive = arr.filter((e) => e.ready).length;
|
|
||||||
const currentNeeded = POOL_SIZE - currentAlive;
|
|
||||||
for (let i = 0; i < currentNeeded; i++) {
|
|
||||||
const entry = spawnWarm(cliModel);
|
|
||||||
arr.push(entry);
|
|
||||||
console.log(`[pool] pre-spawned model=${cliModel} (pool size: ${arr.filter(e => e.ready).length})`);
|
|
||||||
}
|
|
||||||
return;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// --- Crash path: apply exponential backoff and degraded-state logic ---
|
// ── Format messages to prompt text ──────────────────────────────────────
|
||||||
|
function messagesToPrompt(messages) {
|
||||||
const now = Date.now();
|
return messages.map((m) => {
|
||||||
|
|
||||||
// Reset window if the last crash was outside the rolling window
|
|
||||||
if ((now - state.windowStart) > CRASH_WINDOW_MS) {
|
|
||||||
state.windowStart = now;
|
|
||||||
state.failures = 0;
|
|
||||||
}
|
|
||||||
|
|
||||||
state.failures += 1;
|
|
||||||
|
|
||||||
// Check if we've hit the crash limit within the rolling window
|
|
||||||
if (state.failures >= CRASH_LIMIT) {
|
|
||||||
state.degraded = true;
|
|
||||||
console.error(
|
|
||||||
`[pool] DEGRADED: model=${cliModel} crashed ${state.failures} times in ` +
|
|
||||||
`${Math.round((now - state.windowStart) / 1000)}s. ` +
|
|
||||||
`Stopping respawn to prevent CPU spin. Restart the proxy to retry.`
|
|
||||||
);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Exponential backoff: 2s, 4s, 8s, 16s, 32s … capped at 60s
|
|
||||||
const delayMs = Math.min(BACKOFF_BASE_MS * Math.pow(2, state.failures - 1), BACKOFF_MAX_MS);
|
|
||||||
console.warn(`[pool] backoff model=${cliModel} delay=${delayMs}ms (failures=${state.failures}/${CRASH_LIMIT})`);
|
|
||||||
|
|
||||||
state.timer = setTimeout(() => {
|
|
||||||
state.timer = null;
|
|
||||||
const currentAlive = arr.filter((e) => e.ready).length;
|
|
||||||
const currentNeeded = POOL_SIZE - currentAlive;
|
|
||||||
for (let i = 0; i < currentNeeded; i++) {
|
|
||||||
const entry = spawnWarm(cliModel);
|
|
||||||
arr.push(entry);
|
|
||||||
console.log(`[pool] re-spawned model=${cliModel} (pool size: ${arr.filter(e => e.ready).length}, failures=${state.failures}/${CRASH_LIMIT})`);
|
|
||||||
}
|
|
||||||
}, delayMs);
|
|
||||||
}
|
|
||||||
|
|
||||||
function getWarmProcess(cliModel) {
|
|
||||||
const arr = pool.get(cliModel) || [];
|
|
||||||
const entry = arr.find((e) => e.ready);
|
|
||||||
if (entry) {
|
|
||||||
entry.ready = false; // mark as in-use
|
|
||||||
const warmMs = Date.now() - entry.spawnedAt;
|
|
||||||
console.log(`[pool] using warm process model=${cliModel} (warm for ${warmMs}ms)`);
|
|
||||||
return entry.proc;
|
|
||||||
}
|
|
||||||
return null;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Initialize pool for all models
|
|
||||||
function initPool() {
|
|
||||||
for (const cliModel of new Set(Object.values(MODEL_MAP))) {
|
|
||||||
replenishPool(cliModel);
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// ── Call claude CLI ─────────────────────────────────────────────────────
|
|
||||||
function callClaude(model, messages) {
|
|
||||||
return new Promise((resolve, reject) => {
|
|
||||||
const prompt = messages
|
|
||||||
.map((m) => {
|
|
||||||
const text = typeof m.content === "string" ? m.content : JSON.stringify(m.content);
|
const text = typeof m.content === "string" ? m.content : JSON.stringify(m.content);
|
||||||
if (m.role === "system") return `[System] ${text}`;
|
if (m.role === "system") return `[System] ${text}`;
|
||||||
if (m.role === "assistant") return `[Assistant] ${text}`;
|
if (m.role === "assistant") return `[Assistant] ${text}`;
|
||||||
return text;
|
return text;
|
||||||
})
|
}).join("\n\n");
|
||||||
.join("\n\n");
|
}
|
||||||
|
|
||||||
|
// ── Call claude CLI ─────────────────────────────────────────────────────
|
||||||
|
// On-demand spawning: each request spawns a fresh `claude -p` process.
|
||||||
|
// No pool = no crash loops, no stale workers, no degraded states.
|
||||||
|
// Stdin is written immediately so there's no 3s stdin timeout issue.
|
||||||
|
function callClaude(model, messages, conversationId) {
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
if (stats.activeRequests >= MAX_CONCURRENT) {
|
||||||
|
return reject(new Error(`concurrency limit reached (${stats.activeRequests}/${MAX_CONCURRENT})`));
|
||||||
|
}
|
||||||
|
stats.activeRequests++;
|
||||||
|
stats.totalRequests++;
|
||||||
|
|
||||||
const cliModel = MODEL_MAP[model] || model;
|
const cliModel = MODEL_MAP[model] || model;
|
||||||
|
let sessionInfo = null;
|
||||||
|
let prompt;
|
||||||
|
|
||||||
// Try to use a warm process from the pool
|
// ── Session logic ──
|
||||||
let proc = getWarmProcess(cliModel);
|
if (conversationId && sessions.has(conversationId)) {
|
||||||
let usedPool = !!proc;
|
// Resume existing session: only send the latest user message
|
||||||
|
const session = sessions.get(conversationId);
|
||||||
|
session.lastUsed = Date.now();
|
||||||
|
sessionInfo = { uuid: session.uuid, resume: true };
|
||||||
|
stats.sessionHits++;
|
||||||
|
|
||||||
|
const lastUserMsg = [...messages].reverse().find((m) => m.role === "user");
|
||||||
|
prompt = lastUserMsg
|
||||||
|
? (typeof lastUserMsg.content === "string" ? lastUserMsg.content : JSON.stringify(lastUserMsg.content))
|
||||||
|
: "";
|
||||||
|
session.messageCount = messages.length;
|
||||||
|
|
||||||
|
console.log(`[session] resume conv=${conversationId.slice(0, 12)}... uuid=${session.uuid.slice(0, 8)}... msgs=${messages.length} prompt_chars=${prompt.length}`);
|
||||||
|
|
||||||
|
} else if (conversationId) {
|
||||||
|
// New session: send all messages, persist session for future --resume
|
||||||
|
const uuid = randomUUID();
|
||||||
|
sessions.set(conversationId, { uuid, messageCount: messages.length, lastUsed: Date.now(), model: cliModel });
|
||||||
|
sessionInfo = { uuid, resume: false };
|
||||||
|
stats.sessionMisses++;
|
||||||
|
prompt = messagesToPrompt(messages);
|
||||||
|
|
||||||
|
console.log(`[session] new conv=${conversationId.slice(0, 12)}... uuid=${uuid.slice(0, 8)}... msgs=${messages.length}`);
|
||||||
|
|
||||||
|
} else {
|
||||||
|
// One-off request, no session
|
||||||
|
stats.oneOffRequests++;
|
||||||
|
prompt = messagesToPrompt(messages);
|
||||||
|
}
|
||||||
|
|
||||||
|
const cliArgs = buildCliArgs(cliModel, sessionInfo);
|
||||||
|
|
||||||
if (!proc) {
|
|
||||||
// Cold start fallback: spawn fresh
|
|
||||||
console.log(`[pool] no warm process for model=${cliModel}, cold starting...`);
|
|
||||||
const env = { ...process.env };
|
const env = { ...process.env };
|
||||||
delete env.CLAUDECODE;
|
delete env.CLAUDECODE;
|
||||||
delete env.ANTHROPIC_API_KEY;
|
delete env.ANTHROPIC_API_KEY;
|
||||||
delete env.ANTHROPIC_BASE_URL;
|
delete env.ANTHROPIC_BASE_URL;
|
||||||
delete env.ANTHROPIC_AUTH_TOKEN;
|
delete env.ANTHROPIC_AUTH_TOKEN;
|
||||||
proc = spawn(CLAUDE, [
|
|
||||||
"-p", "--model", cliModel,
|
const proc = spawn(CLAUDE, cliArgs, { env, stdio: ["pipe", "pipe", "pipe"] });
|
||||||
"--output-format", "text",
|
|
||||||
"--no-session-persistence",
|
|
||||||
"--allowedTools", "Bash", "Read", "Write", "Edit", "Glob", "Grep",
|
|
||||||
"--", prompt,
|
|
||||||
], { env, stdio: ["ignore", "pipe", "pipe"] });
|
|
||||||
}
|
|
||||||
|
|
||||||
let stdout = "";
|
let stdout = "";
|
||||||
let stderr = "";
|
let stderr = "";
|
||||||
const t0 = Date.now();
|
const t0 = Date.now();
|
||||||
|
let settled = false;
|
||||||
|
|
||||||
|
function settle(err, result) {
|
||||||
|
if (settled) return;
|
||||||
|
settled = true;
|
||||||
|
clearTimeout(timer);
|
||||||
|
stats.activeRequests--;
|
||||||
|
|
||||||
|
if (err) {
|
||||||
|
trackError(err.message || String(err));
|
||||||
|
|
||||||
|
// If session resume failed, remove session so next request starts fresh
|
||||||
|
if (sessionInfo?.resume && conversationId) {
|
||||||
|
console.warn(`[session] resume failed for ${conversationId.slice(0, 12)}..., removing stale session`);
|
||||||
|
sessions.delete(conversationId);
|
||||||
|
}
|
||||||
|
|
||||||
|
reject(err);
|
||||||
|
} else {
|
||||||
|
resolve(result);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
proc.stdout.on("data", (d) => (stdout += d));
|
proc.stdout.on("data", (d) => (stdout += d));
|
||||||
proc.stderr.on("data", (d) => (stderr += d));
|
proc.stderr.on("data", (d) => (stderr += d));
|
||||||
|
|
||||||
proc.on("close", (code) => {
|
proc.on("close", (code) => {
|
||||||
const elapsed = Date.now() - t0;
|
const elapsed = Date.now() - t0;
|
||||||
if (code !== 0) {
|
if (code !== 0) {
|
||||||
console.error(`[claude] exit=${code} model=${cliModel} elapsed=${elapsed}ms stderr=${stderr.slice(0, 300)}`);
|
console.error(`[claude] exit=${code} model=${cliModel} elapsed=${elapsed}ms stderr=${stderr.slice(0, 500)}`);
|
||||||
reject(new Error(stderr || stdout || `exit ${code}`));
|
settle(new Error(stderr.slice(0, 300) || stdout.slice(0, 300) || `claude exit ${code}`));
|
||||||
} else {
|
} else {
|
||||||
console.log(`[claude] ok model=${cliModel} chars=${stdout.length} elapsed=${elapsed}ms pool=${usedPool}`);
|
console.log(`[claude] ok model=${cliModel} chars=${stdout.length} elapsed=${elapsed}ms session=${conversationId ? conversationId.slice(0, 12) + "..." : "none"}`);
|
||||||
resolve(stdout.trim());
|
settle(null, stdout.trim());
|
||||||
}
|
}
|
||||||
});
|
});
|
||||||
proc.on("error", reject);
|
|
||||||
|
|
||||||
// Log prompt size for debugging
|
proc.on("error", (err) => {
|
||||||
console.log(`[claude] request model=${cliModel} prompt_chars=${prompt.length} pool=${usedPool}`);
|
console.error(`[claude] spawn error: ${err.message}`);
|
||||||
|
settle(err);
|
||||||
|
});
|
||||||
|
|
||||||
// If using pool process, send prompt via stdin
|
// Write prompt to stdin immediately — no idle timeout issue
|
||||||
if (usedPool) {
|
|
||||||
proc.stdin.write(prompt);
|
proc.stdin.write(prompt);
|
||||||
proc.stdin.end();
|
proc.stdin.end();
|
||||||
}
|
|
||||||
|
|
||||||
const timer = setTimeout(() => { proc.kill(); reject(new Error("timeout")); }, TIMEOUT);
|
console.log(`[claude] spawned model=${cliModel} prompt_chars=${prompt.length} session=${conversationId ? conversationId.slice(0, 12) + "..." : "none"}`);
|
||||||
proc.on("close", () => clearTimeout(timer));
|
|
||||||
|
// Timeout with graceful kill
|
||||||
|
const timer = setTimeout(() => {
|
||||||
|
stats.timeouts++;
|
||||||
|
console.error(`[claude] timeout after ${TIMEOUT}ms model=${cliModel}`);
|
||||||
|
proc.kill("SIGTERM");
|
||||||
|
setTimeout(() => { try { proc.kill("SIGKILL"); } catch {} }, 5000);
|
||||||
|
settle(new Error(`timeout after ${TIMEOUT}ms`));
|
||||||
|
}, TIMEOUT);
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
// ── Response helpers ────────────────────────────────────────────────────
|
// ── Response helpers ────────────────────────────────────────────────────
|
||||||
function jsonResponse(res, status, data) {
|
function jsonResponse(res, status, data) {
|
||||||
|
if (res.headersSent || res.writableEnded || res.destroyed) return;
|
||||||
res.writeHead(status, { "Content-Type": "application/json" });
|
res.writeHead(status, { "Content-Type": "application/json" });
|
||||||
res.end(JSON.stringify(data));
|
res.end(JSON.stringify(data));
|
||||||
}
|
}
|
||||||
@@ -371,25 +340,23 @@ function sendSSE(res, data) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
function streamResponse(res, id, model, content) {
|
function streamResponse(res, id, model, content) {
|
||||||
|
if (res.headersSent || res.writableEnded || res.destroyed) return;
|
||||||
res.writeHead(200, {
|
res.writeHead(200, {
|
||||||
"Content-Type": "text/event-stream",
|
"Content-Type": "text/event-stream",
|
||||||
"Cache-Control": "no-cache",
|
"Cache-Control": "no-cache",
|
||||||
"Connection": "keep-alive",
|
"Connection": "keep-alive",
|
||||||
});
|
});
|
||||||
const created = Math.floor(Date.now() / 1000);
|
const created = Math.floor(Date.now() / 1000);
|
||||||
// Role chunk
|
|
||||||
sendSSE(res, {
|
sendSSE(res, {
|
||||||
id, object: "chat.completion.chunk", created, model,
|
id, object: "chat.completion.chunk", created, model,
|
||||||
choices: [{ index: 0, delta: { role: "assistant" }, finish_reason: null }],
|
choices: [{ index: 0, delta: { role: "assistant" }, finish_reason: null }],
|
||||||
});
|
});
|
||||||
// Content chunks (~500 chars each)
|
|
||||||
for (let i = 0; i < content.length; i += 500) {
|
for (let i = 0; i < content.length; i += 500) {
|
||||||
sendSSE(res, {
|
sendSSE(res, {
|
||||||
id, object: "chat.completion.chunk", created, model,
|
id, object: "chat.completion.chunk", created, model,
|
||||||
choices: [{ index: 0, delta: { content: content.slice(i, i + 500) }, finish_reason: null }],
|
choices: [{ index: 0, delta: { content: content.slice(i, i + 500) }, finish_reason: null }],
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
// Finish
|
|
||||||
sendSSE(res, {
|
sendSSE(res, {
|
||||||
id, object: "chat.completion.chunk", created, model,
|
id, object: "chat.completion.chunk", created, model,
|
||||||
choices: [{ index: 0, delta: {}, finish_reason: "stop" }],
|
choices: [{ index: 0, delta: {}, finish_reason: "stop" }],
|
||||||
@@ -420,10 +387,13 @@ async function handleChatCompletions(req, res) {
|
|||||||
const model = parsed.model || "claude-sonnet-4-6";
|
const model = parsed.model || "claude-sonnet-4-6";
|
||||||
const stream = parsed.stream;
|
const stream = parsed.stream;
|
||||||
|
|
||||||
|
// Session ID: from request body, header, or null (one-off)
|
||||||
|
const conversationId = parsed.session_id || parsed.conversation_id || req.headers["x-session-id"] || req.headers["x-conversation-id"] || null;
|
||||||
|
|
||||||
if (!messages?.length) return jsonResponse(res, 400, { error: "messages required" });
|
if (!messages?.length) return jsonResponse(res, 400, { error: "messages required" });
|
||||||
|
|
||||||
try {
|
try {
|
||||||
const content = await callClaude(model, messages);
|
const content = await callClaude(model, messages, conversationId);
|
||||||
const id = `chatcmpl-${randomUUID()}`;
|
const id = `chatcmpl-${randomUUID()}`;
|
||||||
|
|
||||||
if (stream) {
|
if (stream) {
|
||||||
@@ -444,8 +414,8 @@ async function handleChatCompletions(req, res) {
|
|||||||
// ── HTTP server ─────────────────────────────────────────────────────────
|
// ── HTTP server ─────────────────────────────────────────────────────────
|
||||||
const server = createServer(async (req, res) => {
|
const server = createServer(async (req, res) => {
|
||||||
res.setHeader("Access-Control-Allow-Origin", "*");
|
res.setHeader("Access-Control-Allow-Origin", "*");
|
||||||
res.setHeader("Access-Control-Allow-Methods", "GET, POST, OPTIONS");
|
res.setHeader("Access-Control-Allow-Methods", "GET, POST, DELETE, OPTIONS");
|
||||||
res.setHeader("Access-Control-Allow-Headers", "Content-Type, Authorization");
|
res.setHeader("Access-Control-Allow-Headers", "Content-Type, Authorization, X-Session-Id, X-Conversation-Id");
|
||||||
if (req.method === "OPTIONS") { res.writeHead(204); res.end(); return; }
|
if (req.method === "OPTIONS") { res.writeHead(204); res.end(); return; }
|
||||||
|
|
||||||
// Bearer token auth (skip for /health and when PROXY_API_KEY is not set)
|
// Bearer token auth (skip for /health and when PROXY_API_KEY is not set)
|
||||||
@@ -473,49 +443,84 @@ const server = createServer(async (req, res) => {
|
|||||||
return handleChatCompletions(req, res);
|
return handleChatCompletions(req, res);
|
||||||
}
|
}
|
||||||
|
|
||||||
// GET /health — includes pool status, version, uptime
|
// GET /health — comprehensive diagnostics
|
||||||
if (req.url === "/health") {
|
if (req.url === "/health") {
|
||||||
const poolStatus = {};
|
|
||||||
for (const [model, arr] of pool) {
|
|
||||||
const readyCount = arr.filter(e => e.ready).length;
|
|
||||||
const errorCount = arr.filter(e => !e.ready).length;
|
|
||||||
poolStatus[model] = {
|
|
||||||
total: arr.length,
|
|
||||||
ready: readyCount,
|
|
||||||
error: errorCount,
|
|
||||||
status: readyCount > 0 ? "ready" : "error",
|
|
||||||
};
|
|
||||||
}
|
|
||||||
const uptimeMs = Date.now() - START_TIME;
|
|
||||||
let binaryOk = false;
|
let binaryOk = false;
|
||||||
try { accessSync(CLAUDE, constants.X_OK); binaryOk = true; } catch {}
|
try { accessSync(CLAUDE, constants.X_OK); binaryOk = true; } catch {}
|
||||||
return jsonResponse(res, 200, {
|
|
||||||
status: binaryOk ? "ok" : "degraded",
|
const uptimeMs = Date.now() - START_TIME;
|
||||||
version: VERSION,
|
const sessionList = [];
|
||||||
uptime: uptimeMs,
|
for (const [id, s] of sessions) {
|
||||||
uptimeHuman: `${Math.floor(uptimeMs / 3600000)}h ${Math.floor((uptimeMs % 3600000) / 60000)}m ${Math.floor((uptimeMs % 60000) / 1000)}s`,
|
sessionList.push({
|
||||||
claudeBinary: CLAUDE,
|
id: id.slice(0, 12) + "...",
|
||||||
claudeBinaryOk: binaryOk,
|
model: s.model,
|
||||||
pool: poolStatus,
|
messages: s.messageCount,
|
||||||
|
idleMs: Date.now() - s.lastUsed,
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
// Catch-all: try to handle any POST with messages
|
return jsonResponse(res, 200, {
|
||||||
|
status: binaryOk && authStatus.ok !== false ? "ok" : "degraded",
|
||||||
|
version: VERSION,
|
||||||
|
architecture: "on-demand (v2)",
|
||||||
|
uptime: uptimeMs,
|
||||||
|
uptimeHuman: `${Math.floor(uptimeMs / 3600000)}h ${Math.floor((uptimeMs % 3600000) / 60000)}m`,
|
||||||
|
claudeBinary: CLAUDE,
|
||||||
|
claudeBinaryOk: binaryOk,
|
||||||
|
auth: authStatus,
|
||||||
|
config: {
|
||||||
|
timeout: TIMEOUT,
|
||||||
|
maxConcurrent: MAX_CONCURRENT,
|
||||||
|
sessionTTL: SESSION_TTL,
|
||||||
|
allowedTools: SKIP_PERMISSIONS ? "all (skip-permissions)" : ALLOWED_TOOLS,
|
||||||
|
systemPrompt: SYSTEM_PROMPT ? `${SYSTEM_PROMPT.slice(0, 50)}...` : "(none)",
|
||||||
|
mcpConfig: MCP_CONFIG || "(none)",
|
||||||
|
},
|
||||||
|
stats,
|
||||||
|
sessions: sessionList,
|
||||||
|
recentErrors: recentErrors.slice(-5),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
// DELETE /sessions — clear all sessions
|
||||||
|
if (req.url === "/sessions" && req.method === "DELETE") {
|
||||||
|
const count = sessions.size;
|
||||||
|
sessions.clear();
|
||||||
|
return jsonResponse(res, 200, { cleared: count });
|
||||||
|
}
|
||||||
|
|
||||||
|
// GET /sessions — list active sessions
|
||||||
|
if (req.url === "/sessions" && req.method === "GET") {
|
||||||
|
const list = [];
|
||||||
|
for (const [id, s] of sessions) {
|
||||||
|
list.push({ id, uuid: s.uuid, model: s.model, messages: s.messageCount, lastUsed: new Date(s.lastUsed).toISOString() });
|
||||||
|
}
|
||||||
|
return jsonResponse(res, 200, { sessions: list });
|
||||||
|
}
|
||||||
|
|
||||||
|
// Catch-all POST
|
||||||
if (req.method === "POST") {
|
if (req.method === "POST") {
|
||||||
return handleChatCompletions(req, res);
|
return handleChatCompletions(req, res);
|
||||||
}
|
}
|
||||||
|
|
||||||
jsonResponse(res, 404, { error: "Not found. Endpoints: GET /v1/models, POST /v1/chat/completions, GET /health" });
|
jsonResponse(res, 404, { error: "Not found. Endpoints: GET /v1/models, POST /v1/chat/completions, GET /health, GET|DELETE /sessions" });
|
||||||
});
|
});
|
||||||
|
|
||||||
// ── Start ──────────────────────────────────────────────────────────────
|
// ── Start ───────────────────────────────────────────────────────────────
|
||||||
initPool();
|
|
||||||
|
|
||||||
server.listen(PORT, "0.0.0.0", () => {
|
server.listen(PORT, "0.0.0.0", () => {
|
||||||
console.log(`openclaw-claude-proxy v${VERSION} listening on http://0.0.0.0:${PORT}`);
|
console.log(`openclaw-claude-proxy v${VERSION} listening on http://0.0.0.0:${PORT}`);
|
||||||
|
console.log(`Architecture: on-demand spawning (no pool)`);
|
||||||
console.log(`Models: ${MODELS.map((m) => m.id).join(", ")}`);
|
console.log(`Models: ${MODELS.map((m) => m.id).join(", ")}`);
|
||||||
console.log(`Claude binary: ${CLAUDE}`);
|
console.log(`Claude binary: ${CLAUDE}`);
|
||||||
console.log(`Timeout: ${TIMEOUT}ms`);
|
console.log(`Timeout: ${TIMEOUT}ms | Max concurrent: ${MAX_CONCURRENT}`);
|
||||||
console.log(`Pool size: ${POOL_SIZE} per model, max idle: ${POOL_MAX_IDLE / 1000}s`);
|
console.log(`Tools: ${SKIP_PERMISSIONS ? "all (skip-permissions)" : ALLOWED_TOOLS.join(", ")}`);
|
||||||
|
console.log(`Sessions: TTL=${SESSION_TTL / 1000}s`);
|
||||||
|
if (SYSTEM_PROMPT) console.log(`System prompt: "${SYSTEM_PROMPT.slice(0, 80)}..."`);
|
||||||
|
if (MCP_CONFIG) console.log(`MCP config: ${MCP_CONFIG}`);
|
||||||
console.log(`Auth: ${PROXY_API_KEY ? "enabled (PROXY_API_KEY set)" : "disabled (no PROXY_API_KEY)"}`);
|
console.log(`Auth: ${PROXY_API_KEY ? "enabled (PROXY_API_KEY set)" : "disabled (no PROXY_API_KEY)"}`);
|
||||||
|
console.log(`---`);
|
||||||
|
console.log(`Coexistence: This proxy does NOT conflict with Claude Code interactive mode.`);
|
||||||
|
console.log(` OCP uses: localhost:${PORT} (HTTP) → claude -p (per-request process)`);
|
||||||
|
console.log(` CC uses: MCP protocol (in-process) → persistent session`);
|
||||||
|
console.log(` Both can run simultaneously on the same machine.`);
|
||||||
});
|
});
|
||||||
|
|||||||
Reference in New Issue
Block a user