mirror of
https://github.com/dtzp555-max/ocp.git
synced 2026-07-21 21:15:09 +00:00
feat(chat): forward OpenAI image_url parts to Claude (multimodal vision) (#154)
* feat(chat): forward OpenAI image_url parts to Claude (multimodal vision) `POST /v1/chat/completions` previously flattened every message to plain text via contentToText(), replacing image_url parts with "[non-text content omitted]" — so images were silently dropped (issue #110). This adds real multimodal support: OpenAI `image_url` content parts are translated to Anthropic image blocks and fed to the Claude CLI over `--input-format stream-json`, keeping subscription auth (the reason OCP routes through the CLI rather than the API). Class B.1 (OpenAI-compatibility surface), authorized by ADR 0006. Request shape follows OpenAI's published vision / chat-completions spec (https://platform.openai.com/docs/guides/vision and the chat/completions `content` image_url part) — no field is introduced beyond OpenAI's shape. The CLI's `--input-format stream-json` is the transport for this Class B endpoint, not a forwarded cli.js operation, so there is no Class A cli.js citation to make; scope is justified under the ALIGNMENT.md Class B mapping of Rule 2 (no invention beyond the cited OpenAI spec). Mechanism (verified empirically against the installed CLI, v2.1.206): a user message whose `content` is an Anthropic block array including `{type:"image",source:{type:"base64",media_type,data}}` fed to `claude -p --input-format stream-json` is correctly described by the model. Confirmed live end-to-end through this endpoint (a base64 PNG returns the correct color). Design: - New pure module `lib/multimodal.mjs` (mirrors the lib/*.mjs pattern; unit- testable without a live server): hasImageContent, buildImageBlocks, buildStreamJsonInput, MultimodalError. - server.mjs: text path is byte-for-byte unchanged. Only when a request carries an image_url part does spawnClaudeProcess switch stdin to a stream-json user envelope and buildCliArgs add `--input-format stream-json`. Image parsing runs before any stats mutation so a validation failure never leaks counters/slots. - Images bypass the text char budget (CLAUDE_MAX_PROMPT_CHARS) and are bounded by explicit byte/count caps with clear 4xx errors (413 for size/count, 400 for malformed/unsupported/disabled-remote), never a silent drop. Scope decisions (v1): - Base64 data URIs supported by default (image/jpeg,png,gif,webp). - Remote http(s) image URLs OFF by default behind CLAUDE_IMAGE_ALLOW_URL; when enabled they are passed through as an Anthropic url-source (OCP never fetches the URL itself, so no OCP-side SSRF surface). - Audio/file parts deferred: existing placeholder behavior preserved. - Images anywhere in multi-turn history, not just the last message. New env vars (documented in README Environment Variables table): CLAUDE_IMAGE_ALLOW_URL, CLAUDE_MAX_IMAGE_BYTES, CLAUDE_MAX_IMAGES, CLAUDE_MAX_IMAGE_TOTAL_BYTES, CLAUDE_MAX_BODY_SIZE (now configurable; default 5 MB unchanged). Tests: 26 unit tests in test-features.mjs covering data-URI parse, multiple images, text/image ordering, multi-turn history images, malformed/oversized/ too-many handling, remote-URL policy, and text-path parity. `npm test` green (289 passed). `node --check` clean. No alignment-blacklist tokens added. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(chat): address PR #154 review blockers — TUI guard, fail-closed caps, text budget Remediates the three merge-blocking findings from the maintainer's review of the multimodal vision PR. Class B.1 (OpenAI-compat surface): request shape per OpenAI vision spec (image_url content parts), authorized by ADR 0006. No new wire shape and no cli.js surface change — the stream-json image contract is the CLI's native input format (already cited in the base commit); these are correctness fixes on the OCP-owned validation/dispatch layer. F1 — TUI mode silently dropped images and returned 200. callClaudeTui() renders every non-text part as "[non-text content omitted]", so a vision request in CLAUDE_TUI_MODE=true was answered about an image the model never saw. Now handleChatCompletions fails loudly with 400 images_unsupported_in_tui_mode instead of a silent drop (ALIGNMENT.md forbids serving text the model did not mean). Documented in README § Images / Multimodal. F3 — NaN env parsing failed open. CLAUDE_MAX_BODY_SIZE=unlimited -> NaN -> `body.length > NaN` always false -> body cap gone (OOM DoS); =5MB -> 5 bytes -> proxy bricked; same on CLAUDE_MAX_IMAGES / _IMAGE_BYTES / _IMAGE_TOTAL_BYTES. Added lib/env.mjs parsePositiveInt (pure, fail-closed) + a thin parseIntEnv warn wrapper; a malformed cap now keeps the safe default and warns at startup. F2 — images let unbounded text bypass MAX_PROMPT_CHARS. buildImageBlocks only counted textChars and never truncated; the budget was never passed in. Threaded maxTextChars (= MAX_PROMPT_CHARS) into the multimodal transform, which now truncates text tail-first (mirroring messagesToPrompt) while preserving image blocks, and logs prompt_truncated. Tests: +11 in test-features.mjs (all pure-module, per the repo's no-server-import pattern) covering the text-budget enforcement and the fail-closed cap parsing, including the exact F2 (500k chars + 1 image) and F3 (unlimited/5MB/0/20.5) scenarios. 300 passed, 0 failed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(server): address PR #154 review round 2 — MAX_PROMPT_CHARS fail-closed + system-only image guard Closes the two residual gaps from the round-2 review, both traced to the same root as the already-fixed blockers. Class B.1 (OpenAI-compat vision): authorized by ADR 0006; request shape is the OpenAI `image_url` content part. No new wire shape — the Anthropic image block over `--input-format stream-json` is the CLI's native contract (cli.js buildStreamJsonInput path, verified live in round 1). MAX_PROMPT_CHARS is OCP's own truncation guard, not a cli.js operation. Gap (a) — MAX_PROMPT_CHARS was left on the raw parseInt while every other cap moved to the fail-closed helper. `let MAX_PROMPT_CHARS = parseInt(env||"150000",10)` sat five lines above the parseIntEnv helper this PR added, so CLAUDE_MAX_PROMPT_CHARS=unlimited → NaN → enforceTextBudget's `!(NaN > 0)` early-return → 500k chars passed unbounded, truncated:false, silently defeating F2's text-budget guarantee under a plausible operator config. Fix: hoist parseIntEnv above the declaration and derive MAX_PROMPT_CHARS through it (keeps `let` for the settings API). A misconfigured value now keeps the 150k default and warns, like the other caps. Gap (b) — an image present ONLY in a system message silently dropped in non-TUI mode. Detection runs on the full message list, but extraction/spawn filter role==="system" out, so a system-only image was detected as multimodal, survived no filter, fell to the text path, and rendered as "[non-text content omitted]" → 200 with a hallucinated answer — the one silent-drop outcome F1 exists to forbid. Fix: after filtering, if hasImageContent(full) is true but no image survives, return `400 images_unsupported_in_system_messages`. Narrow (OpenAI disallows images in the system role) so no legitimate request is rejected. Documented in README § Images. Tests: +4 (all pure-module) — parsePositiveInt('unlimited') keeps the 150k default (gap a) and a valid override is honored; hasImageContent proves the guard predicate fires for a system-only image (true on full list, false after the system filter) and does NOT fire for a user-message image. 304 passed, 0 failed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: vvlasy-openclaw <vvlasy-openclaw@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: vvlasy-openclaw <vvlasy@gmail.com> Co-authored-by: dtzp555 <dtzp555@gmail.com>
This commit is contained in:
@@ -27,7 +27,7 @@ One proxy. Multiple IDEs. All models. **$0 API cost.**
|
||||
- [How It Works](#how-it-works)
|
||||
- Reference: [Available Models](#available-models) · [API Endpoints](#api-endpoints) · [Environment Variables](#environment-variables)
|
||||
- Modes & operations: [LAN & multi-user](#lan--multi-user) → [`docs/lan-mode.md`](docs/lan-mode.md) · [Subscription-pool (TUI) mode](#subscription-pool-tui-mode) → [`docs/tui-mode.md`](docs/tui-mode.md) · [Upgrading](#upgrading) → [`docs/upgrading.md`](docs/upgrading.md)
|
||||
- [Built-in Usage Monitoring](#built-in-usage-monitoring) · [Response Cache](#response-cache) · [OpenClaw Integration](#openclaw-integration)
|
||||
- [Built-in Usage Monitoring](#built-in-usage-monitoring) · [Response Cache](#response-cache) · [Images / Multimodal](#images--multimodal-vision) · [OpenClaw Integration](#openclaw-integration)
|
||||
- [Troubleshooting](#troubleshooting) → [`docs/troubleshooting.md`](docs/troubleshooting.md)
|
||||
- [Repository Layout](#repository-layout) · [Security](#security) · [Governance](#governance) · [Support OCP](#support-ocp) · [License](#license)
|
||||
|
||||
@@ -222,12 +222,17 @@ The canonical list lives in [`models.json`](./models.json) — the single source
|
||||
| `CLAUDE_MAX_CONCURRENT` | `8` | Max concurrent claude processes (`-p`/stream-json path) |
|
||||
| `CLAUDE_MAX_QUEUE` | `16` | Max requests **waiting** for a `-p` concurrency slot. Beyond `CLAUDE_MAX_CONCURRENT`, requests queue (up to this cap) instead of being rejected; when the queue is **also** full, the request gets `HTTP 429` + `Retry-After` (not an opaque 500). Surfaced on `/health.concurrency` + `/health.stats.queueRejections`. |
|
||||
| `CLAUDE_QUEUE_RETRY_AFTER` | `5` | Seconds advertised in the `Retry-After` header on a `-p` concurrency-overflow `429`. |
|
||||
| `CLAUDE_MAX_PROMPT_CHARS` | *(derived)* | Prompt truncation limit in chars. Default derives from the models.json SPOT: `max(contextWindow) × 3` — currently **600,000** (≈150–200k tokens). Setting this env var (or the runtime settings API) overrides the derivation absolutely. See [ADR 0009](docs/adr/0009-spot-derived-prompt-budget.md). Note: very large prompts burn subscription-window quota quickly and slow TTFT; the TUI-mode paste path is untested beyond ~hundreds of KB. |
|
||||
| `CLAUDE_MAX_PROMPT_CHARS` | *(derived)* | Prompt truncation limit in chars. Default derives from the models.json SPOT: `max(contextWindow) × 3` — currently **600,000** (≈150–200k tokens). Setting this env var (or the runtime settings API) overrides the derivation absolutely. See [ADR 0009](docs/adr/0009-spot-derived-prompt-budget.md). Note: very large prompts burn subscription-window quota quickly and slow TTFT; the TUI-mode paste path is untested beyond ~hundreds of KB. Applies to **text only** — image bytes bypass this budget (see [Images / Multimodal](#images--multimodal-vision)). |
|
||||
| `CLAUDE_SESSION_TTL` | `3600000` | Session expiry (ms, default: 1 hour) |
|
||||
| `CLAUDE_CACHE_TTL` | `0` | Response cache TTL (ms, 0 = disabled). Set to e.g. `300000` for 5-min cache. See [Response Cache](#response-cache). |
|
||||
| `CLAUDE_ALLOWED_TOOLS` | `Bash,Read,...,Agent` | Comma-separated tools to pre-approve |
|
||||
| `CLAUDE_SKIP_PERMISSIONS` | `false` | Bypass all permission checks |
|
||||
| `CLAUDE_MCP_CONFIG` | *(unset)* | Path to an MCP server config JSON, passed to the spawned `claude` as `--mcp-config` (both the `-p` path and TUI `OCP_TUI_FULL_TOOLS` panes) |
|
||||
| `CLAUDE_MAX_BODY_SIZE` | `5242880` | Max request body size (bytes, default 5 MB). Base64 image payloads inflate ~33%; raise this to admit larger multimodal requests. Fail-closed parsing: a garbage value keeps the default. |
|
||||
| `CLAUDE_IMAGE_ALLOW_URL` | `false` | Allow remote `http(s)` image URLs in `image_url` parts. **Off by default** (v1 supports base64 `data:` URIs only). When on, the URL is passed through to Anthropic as a `url` image source — **OCP does not fetch it** (no OCP-side SSRF surface); unreachable/blocked URLs surface as an API error. |
|
||||
| `CLAUDE_MAX_IMAGE_BYTES` | `5242880` | Per-image decoded-byte cap (default 5 MB). Over-cap images get `HTTP 413`. |
|
||||
| `CLAUDE_MAX_IMAGES` | `20` | Max image parts per request. Over-cap gets `HTTP 413`. |
|
||||
| `CLAUDE_MAX_IMAGE_TOTAL_BYTES` | `20971520` | Aggregate decoded-byte cap across all images in a request (default 20 MB). Over-cap gets `HTTP 413`. |
|
||||
| `CLAUDE_SYSTEM_PROMPT` | *(unset)* | Operator-wide system-prompt text appended (last) to every request's composed system prompt on the default `-p` path. TUI-mode panes are unaffected (they keep the interactive CLI's own system prompt). Echoed truncated on `/health.systemPrompt`. Note: changing this value and restarting auto-invalidates the response cache (the key carries a boot-config epoch, #177). |
|
||||
| `CLAUDE_NO_CONTEXT` | `false` | Suppress CLAUDE.md and auto-memory injection (pure API mode) |
|
||||
| `PROXY_API_KEY` | *(unset)* | Bearer token for shared-mode authentication |
|
||||
@@ -382,6 +387,91 @@ ocp settings cacheTTL 0 # disable at runtime
|
||||
|
||||
Cache is **disabled by default** (`CLAUDE_CACHE_TTL=0`). All data is stored locally in `~/.ocp/ocp.db`. **Hash format upgrade in v3.13.0:** legacy `v1` cache rows don't match new `v2`-format lookups; they orphan and are reaped by the TTL cleanup interval within one window — no migration script required.
|
||||
|
||||
## Images / Multimodal (Vision)
|
||||
|
||||
`POST /v1/chat/completions` accepts OpenAI-style multimodal `content` parts, so a
|
||||
message can carry images alongside text and Claude will actually see them. This
|
||||
follows OpenAI's [vision](https://platform.openai.com/docs/guides/vision) /
|
||||
[chat-completions `image_url`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-messages)
|
||||
request shape — no OCP-invented fields. (Class B.1 endpoint; see ADR 0006.)
|
||||
|
||||
Under the hood, when a request carries an image OCP feeds the conversation to the
|
||||
Claude CLI as Anthropic image blocks over `--input-format stream-json`. Text-only
|
||||
requests are completely unaffected (unchanged code path).
|
||||
|
||||
### Supported input
|
||||
|
||||
- **Base64 data URIs** (default, recommended):
|
||||
`data:image/png;base64,<...>`. Media types: `image/jpeg`, `image/png`,
|
||||
`image/gif`, `image/webp`.
|
||||
- **Remote `http(s)` URLs** — **off by default**. Set `CLAUDE_IMAGE_ALLOW_URL=1`
|
||||
to enable; the URL is passed through to Anthropic (OCP never fetches it itself,
|
||||
so there is no OCP-side SSRF surface).
|
||||
- Images may appear in **any** message in the history (multi-turn), not just the
|
||||
last one.
|
||||
- Non-image, non-text parts (audio, files) are **not** yet supported and are
|
||||
replaced with a `[non-text content omitted]` placeholder (deferred to a future
|
||||
version).
|
||||
|
||||
### Example (base64 data URI)
|
||||
|
||||
```bash
|
||||
curl -X POST http://127.0.0.1:3456/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "claude-sonnet-4-6",
|
||||
"messages": [{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{ "type": "text", "text": "What is in this image?" },
|
||||
{ "type": "image_url",
|
||||
"image_url": { "url": "data:image/png;base64,iVBORw0KGgoAAA..." } }
|
||||
]
|
||||
}]
|
||||
}'
|
||||
```
|
||||
|
||||
### Not supported in TUI mode
|
||||
|
||||
Multimodal images require the default `-p` spawn path. In **TUI / subscription-pool
|
||||
mode** (`CLAUDE_TUI_MODE=true`) the CLI is driven interactively and cannot carry
|
||||
image blocks, so a request with an `image_url` part returns **`400
|
||||
images_unsupported_in_tui_mode`** rather than silently dropping the image and
|
||||
answering about something the model never saw. Remove the images, or run OCP
|
||||
without TUI mode, to use vision.
|
||||
|
||||
Images must also live in a **user or assistant** message, not a `system` message
|
||||
(system content is not forwarded to the CLI as image blocks). An `image_url` part
|
||||
present only in a system message returns **`400 images_unsupported_in_system_messages`**
|
||||
for the same reason — fail loudly rather than answer about an unseen image. This matches
|
||||
the OpenAI vision spec, which does not place images in the system role.
|
||||
|
||||
### Limits
|
||||
|
||||
Images bypass the text `CLAUDE_MAX_PROMPT_CHARS` budget and are instead bounded by
|
||||
their own byte/count caps. The **text** in a multimodal request is still subject to
|
||||
`CLAUDE_MAX_PROMPT_CHARS` (older text is truncated exactly as on the text-only
|
||||
path — only the image bytes are exempt). All numeric caps are parsed **fail-closed**:
|
||||
a malformed value (e.g. `CLAUDE_MAX_BODY_SIZE=unlimited` or `=5MB`) is rejected with
|
||||
a startup warning and the safe default is kept — a misconfigured cap can never
|
||||
silently disable the guard. Requests that violate a cap get a clear `4xx` (never a
|
||||
silent drop):
|
||||
|
||||
| Cap | Env var | Default | Error |
|
||||
|-----|---------|---------|-------|
|
||||
| Request body | `CLAUDE_MAX_BODY_SIZE` | 5 MB | `413` request body too large |
|
||||
| Per-image bytes | `CLAUDE_MAX_IMAGE_BYTES` | 5 MB | `413` `image_too_large` |
|
||||
| Total image bytes | `CLAUDE_MAX_IMAGE_TOTAL_BYTES` | 20 MB | `413` `images_too_large` |
|
||||
| Image count | `CLAUDE_MAX_IMAGES` | 20 | `413` `too_many_images` |
|
||||
| Unsupported media type | — | — | `400` `unsupported_image_type` |
|
||||
| Malformed data URI | — | — | `400` `invalid_data_uri` |
|
||||
| Remote URL while disabled | `CLAUDE_IMAGE_ALLOW_URL` | off | `400` `remote_url_disabled` |
|
||||
|
||||
Base64 payloads are large: a 5 MB image is ~6.7 MB as a data URI, so raise
|
||||
`CLAUDE_MAX_BODY_SIZE` (and, if needed, `CLAUDE_MAX_IMAGE_BYTES`) to admit big
|
||||
images. Vision support depends on the target model — request a current
|
||||
vision-capable Claude model.
|
||||
|
||||
## OpenClaw Integration
|
||||
|
||||
OCP was originally built for [OpenClaw](https://github.com/openclaw/openclaw) and includes deep integration:
|
||||
|
||||
Reference in New Issue
Block a user