mirror of
https://github.com/dtzp555-max/ocp.git
synced 2026-07-22 13:35:08 +00:00
* feat(server): honor OpenAI response_format for structured-output clients `/v1/chat/completions` advertises OpenAI compatibility but ignored `response_format`, so clients requiring machine-parseable JSON (Home Assistant AI Tasks, Honcho, OpenAI-SDK scripts) received free-form assistant prose — markdown tables, ```json fences, trailing commentary — that fails JSON.parse. This honors the OpenAI `response_format` contract on the `-p` path: - New `lib/structured-output.mjs` (pure, unit-tested): `detectStructuredOutput` (json_schema / json_object), `structuredSystemInstruction` (strict JSON-only steering, escalated on retry), `extractJsonPayload` (string-aware balanced slice that unwraps fences/prose), and a minimal JSON-Schema `validateJsonSchema` (types, required, enum, const, additionalProperties, nullability, items, min/maxItems). - `server.mjs`: `runStructuredCompletion` retries up to `OCP_STRUCTURED_MAX_ATTEMPTS` (default 3), returns the canonical JSON string as `message.content`, and yields HTTP 422 (`invalid_response_error`) if no valid JSON can be produced. Structured requests take their own path (bypass the cache, which does not key on response_format). Non-structured requests are byte-for-byte unchanged, streaming included. - Nullability precedence: a `null` value is accepted whenever the schema permits null (`type:["x","null"]` / `nullable:true`), even if a bare `enum` omits null — matches OpenAI behaviour and fixes real Home Assistant schemas (`type:["string","null"], enum:["Loxone"]`) that otherwise 422 on null. - README: Structured Outputs section + `OCP_STRUCTURED_MAX_ATTEMPTS` env row. - 18 new unit tests (281 passed, 0 failed). Endpoint class: B.1 (OpenAI-compatibility surface, `/v1/chat/completions`). Specification: OpenAI chat/completions `response_format` (https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format). Authorizing ADR: ADR 0006 — OpenAI shim scope. cli.js does NOT perform this operation (it speaks Anthropic's protocol, not OpenAI's); scope is justified under ADR 0006 Class B.1 (OpenAI spec as protocol authority). Revives closed PR #99. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(server): structured-output caching + json_mode alias Follow-up on the response_format path, closing the two gaps vs closed PR #99: - **Validated caching (improves on #99).** Structured responses now use the OCP cache when CLAUDE_CACHE_TTL>0, on a structured-keyed hash: cacheHash gains an `structured` marker folding the detected response_format/schema into the key, so a JSON reply never collides with the conversational answer to the same prompt and different schemas never share a slot. Only a *validated* result is written back — a 422 is never cached. (#99 cached the fence-stripped but *unvalidated* output; this caches only schema-valid JSON.) The marker is absent for normal requests, so existing cache hashes are byte-identical. - **json_mode alias.** Honor the non-standard top-level `json_mode: true` flag as a json_object alias, matching #99's activation set. Disclosed as non-spec. - README: json_mode shape + caching note. +2 unit tests (283 passed, 0 failed). Endpoint class: B.1 (/v1/chat/completions), ADR 0006. json_mode is a non-OpenAI convenience alias (disclosed); everything else stays within OpenAI's response_format spec. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(server): address PR #153 review — $ref/strict, extraction safety, refusal, singleflight Remediates the maintainer's merge-blocking findings on the structured-output PR, and rebases onto current main (the one-hunk test-features.mjs conflict — both sides appended tests — resolved by keeping both blocks). Class B.1 (OpenAI-compat): spec authority is OpenAI chat/completions `response_format` (https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format), authorized by ADR 0006. No cli.js analogue (claude -p has no native response_format); scope is the B.1 shim, not a Class A forward. Finding 1 (correctness gate) — strict:true + $ref/$defs rejected valid objects 100% of the time. `noExtra = addl === false || (strict && addl === undefined)` treated a nested {$ref:"#/$defs/step"} as an empty-properties object and, under strict, rejected every real key as "additional property not allowed" — exactly the shape the OpenAI SDK emits (zodResponseFormat / client.beta.chat.completions.parse) and OpenAI's own docs example. Fix: validateJsonSchema now resolves same-document $ref against the root $defs/definitions, handles allOf/anyOf/oneOf composition, and only infers additionalProperties:false from strict when the object actually declares its own non-empty properties and is not a composite. Explicit additionalProperties:false is always honoured, so validation is not weakened (tests prove an extra key and a missing required key still fail under strict). Finding 2 (correctness gate) — the extractor served JSON the model did not mean. json_object mode had no validation at all: a refusal like `I can't. The schema is {"type":"object"}` returned the embedded object as the answer. Now json_object requires the WHOLE reply to parse as a single JSON value, and schema mode rejects a reply carrying more than one top-level JSON value (Schema:{}/Answer:{}, Option A/Option B) rather than silently picking the first. The schema-validated value is still returned; nothing unvalidated is served. Finding 3 — replaced the invented `invalid_response_error` 422 with OpenAI's assistant `refusal` field (200, content:null, refusal:<reason>, finish_reason:"stop"), streaming and non-streaming, so SDK clients take their refusal branch instead of throwing an opaque UnprocessableEntityError. Finding 5 — runStructuredCompletion no longer bypasses stampede protection. Identical concurrent one-off structured requests now share one singleflight (independent of cache enablement), so N callers no longer cost N × up-to-3 spawns. Cache read/write still gated on CLAUDE_CACHE_TTL; refusals are never cached. Docs: README structured-output § updated (refusal field, $ref/composition support, whole-reply json_object rule, ambiguous-multi-value rejection) plus a Caching & cost paragraph stating the post-2026-06-15 model, the up-to-N-spawn worst case, the singleflight + validated-cache guards, and the OCP_STRUCTURED_MAX_ATTEMPTS=1 / per-key quota levers. Tests: +11 (all pure-module) — OpenAI's doc $ref/$defs schema under strict:true accepts a conforming reply and still rejects extra/missing keys; anyOf/allOf; unresolvable $ref skipped; json_object refusal-embedded-json rejected; >1-top-level-value rejected. 360 passed, 0 failed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(server): address PR #153 review round 2 — cyclic-$ref guard + NaN attempts guard Remediates the two remaining merge-blocking findings from the round-2 review. Class B.1 (OpenAI-compat): spec authority is OpenAI chat/completions `response_format` (https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format), authorized by ADR 0006. No cli.js analogue (claude -p has no native response_format, and the retry cap OCP_STRUCTURED_MAX_ATTEMPTS is OCP's own coercion loop, not a cli.js operation); scope stays the B.1 shim, not a Class A forward. BLOCKER — cyclic $ref stack-overflowed the validator. resolveRef + the $ref branch of validateJsonSchema had no cycle detection: a pure ref→ref cycle ({$defs:{a:{$ref:b},b:{$ref:a}},$ref:a}) recursed independent of the data and threw RangeError for ANY reply value (even `5`), caught upstream as a 500 but only after 1–3 metered spawns — a request-controlled cost-amplification / grief vector on an authed path. Fix: validateJsonSchema now threads a `refChain` of the $ref pointers resolved on the current path WITHOUT consuming data (a $ref hop, or an allOf/anyOf/oneOf branch — all re-validate the same value); a pointer reappearing on that chain fails closed with a `cyclic $ref detected` error. Data-consuming recursion (properties/items/ additionalProperties) deliberately resets the chain, because a JSON value is a finite tree so those always terminate — a legitimately recursive schema (Node→child:Node) must NOT be flagged. A REF_DEPTH_CAP backstops any threading mistake. MUST-FIX — OCP_STRUCTURED_MAX_ATTEMPTS NaN guard was broken. `Math.max(1, parseInt(env ||"3",10))` === `Math.max(1, NaN)` === NaN for a non-integer value, so the retry loop `attempt < NaN` never ran → 0 spawns, every structured request silently refused (fails closed on cost but bricks the feature and ignores the intended floor). Fix: extracted a pure fail-closed resolveMaxAttempts() into lib/structured-output.mjs — rejects NaN/non-finite/<1, keeps the documented default of 3, and warns at startup. server.mjs now derives STRUCTURED_MAX_ATTEMPTS through it. Tests: +8 (all pure-module) — a→b→a and self (a→a) cyclic $ref fail closed without overflowing the stack; a cycle routed through anyOf; a legitimate recursive Node schema is NOT flagged; resolveMaxAttempts honors valid integers, defaults on unset/empty/null, and fails closed (not NaN, not 0) on abc/0/-1/NaN/Infinity/blank with a startup warn. 368 passed, 0 failed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: vvlasy-openclaw <vvlasy@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: dtzp555 <dtzp555@gmail.com>
489 lines
20 KiB
JavaScript
489 lines
20 KiB
JavaScript
// keys.mjs — API key management and usage tracking for OCP LAN mode
|
|
// Uses Node.js built-in SQLite (node:sqlite) — zero external dependencies.
|
|
import { DatabaseSync } from "node:sqlite";
|
|
import { randomBytes, createHash } from "node:crypto";
|
|
import { join } from "node:path";
|
|
import { mkdirSync, chmodSync } from "node:fs";
|
|
import { homedir } from "node:os";
|
|
|
|
// Resolved LAZILY, on first getDb() — not at module top-level. Two reasons, and the second is
|
|
// the bug this fixes:
|
|
//
|
|
// 1. Merely IMPORTING keys.mjs should not, as a side effect, create directories in the
|
|
// operator's home.
|
|
// 2. OCP_DIR_OVERRIDE exists so the test suite can point the key store at a scratch dir — and
|
|
// because ESM hoists imports, a top-level `const OCP_DIR = ...` here would be evaluated
|
|
// BEFORE an importing module's body could set the env var. Eager resolution made the
|
|
// override unsettable in the one place that needs it. (test-features.mjs carried a comment
|
|
// claiming it could "set env before the first getDb() call" — it could not, because nothing
|
|
// here ever read an env var. So `npm test` wrote real, UNREVOKED api_keys rows into the
|
|
// operator's live ~/.ocp/ocp.db: two per run, unbounded — 737 junk keys against 12 real ones
|
|
// on the maintainer's host — and two concurrent runs raced one file, which is the ~1-in-6
|
|
// flake in `listKeys includes quota fields`.)
|
|
//
|
|
// The override is gated on NODE_ENV === "test", and that gate is the ACTUAL guard. An earlier
|
|
// cut of this fix relied on the variable merely having an awkward name — i.e. a naming convention
|
|
// plus a comment — which is precisely the failure mode this whole change exists to indict (a
|
|
// comment describing an intention that nothing enforces). The two-key gate means NEITHER var
|
|
// alone does anything: a stray OCP_DIR_OVERRIDE with no NODE_ENV is inert, and NODE_ENV=test with
|
|
// no override just resolves the default dir.
|
|
//
|
|
// This gate does NOT, by itself, prove a production daemon can't be redirected — an earlier
|
|
// version of this comment overclaimed that ("a production server runs without NODE_ENV, so it
|
|
// CANNOT honor the override no matter how the variable got in"). That is only true while the
|
|
// daemon's env actually lacks NODE_ENV=test, which is an assumption, not something this file can
|
|
// enforce. What makes it hold in the shipped configuration is defense-in-depth in OCP's launchers:
|
|
// the plist/systemd units strip both vars on every (re)install (scripts/lib/plist-merge.mjs
|
|
// NEVER_PRESERVE), and `ocp` restart's manual nohup fallback strips them (`env -u`). So a server
|
|
// OCP itself started cannot carry the test-only redirection. The one residual path is an operator
|
|
// who hand-launches `node server.mjs` with BOTH vars explicitly exported, bypassing every
|
|
// launcher — a case no library-level gate can catch. The loud getDb() log below ("NOT the default
|
|
// ~/.ocp/ocp.db") is the backstop there: a wrong key store is at least never silent (in
|
|
// AUTH_MODE=multi that would otherwise be a total auth outage with nothing on /health to show it).
|
|
function resolveOcpDir() {
|
|
const override = process.env.NODE_ENV === "test" ? process.env.OCP_DIR_OVERRIDE : null;
|
|
const dir = override || join(homedir(), ".ocp");
|
|
mkdirSync(dir, { recursive: true, mode: 0o700 });
|
|
// Tighten the directory mode in case it already existed with broader permissions.
|
|
try { chmodSync(dir, 0o700); } catch { /* ignore EPERM on pre-existing dirs */ }
|
|
return dir;
|
|
}
|
|
|
|
let db;
|
|
let dbPath; // resolved on first open, alongside the db handle
|
|
|
|
export function getDb() {
|
|
if (!db) {
|
|
dbPath = join(resolveOcpDir(), "ocp.db");
|
|
// Say which store we opened. Silence was the other half of the bug: a server on the wrong
|
|
// key store looks exactly like a server on the right one until every request 401s.
|
|
if (dbPath !== join(homedir(), ".ocp", "ocp.db")) {
|
|
console.error(`[keys] key store: ${dbPath} (NOT the default ~/.ocp/ocp.db)`);
|
|
}
|
|
db = new DatabaseSync(dbPath);
|
|
db.exec("PRAGMA journal_mode = WAL");
|
|
db.exec("PRAGMA foreign_keys = ON");
|
|
initSchema();
|
|
// Tighten mode on the DB file (0600) after creation / first open.
|
|
try { chmodSync(dbPath, 0o600); } catch { /* ignore — same-user access still works */ }
|
|
}
|
|
return db;
|
|
}
|
|
|
|
// Which file the key store actually opened. Exported so a test can ASSERT it is not the
|
|
// operator's real db — the bug this replaced was invisible precisely because nothing checked.
|
|
export function getDbPath() { return dbPath; }
|
|
|
|
function initSchema() {
|
|
db.exec(`
|
|
CREATE TABLE IF NOT EXISTS api_keys (
|
|
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
|
key TEXT UNIQUE NOT NULL,
|
|
name TEXT NOT NULL,
|
|
created_at TEXT NOT NULL DEFAULT (datetime('now')),
|
|
revoked INTEGER NOT NULL DEFAULT 0
|
|
);
|
|
|
|
CREATE TABLE IF NOT EXISTS usage_log (
|
|
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
|
key_id INTEGER,
|
|
key_name TEXT NOT NULL DEFAULT 'anonymous',
|
|
model TEXT NOT NULL,
|
|
prompt_chars INTEGER NOT NULL DEFAULT 0,
|
|
response_chars INTEGER NOT NULL DEFAULT 0,
|
|
elapsed_ms INTEGER NOT NULL DEFAULT 0,
|
|
success INTEGER NOT NULL DEFAULT 1,
|
|
created_at TEXT NOT NULL DEFAULT (datetime('now')),
|
|
FOREIGN KEY (key_id) REFERENCES api_keys(id)
|
|
);
|
|
|
|
CREATE INDEX IF NOT EXISTS idx_usage_created ON usage_log(created_at);
|
|
CREATE INDEX IF NOT EXISTS idx_usage_key ON usage_log(key_id);
|
|
|
|
CREATE TABLE IF NOT EXISTS response_cache (
|
|
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
|
hash TEXT UNIQUE NOT NULL,
|
|
model TEXT NOT NULL,
|
|
response TEXT NOT NULL,
|
|
created_at TEXT NOT NULL DEFAULT (datetime('now')),
|
|
last_hit_at TEXT,
|
|
hits INTEGER NOT NULL DEFAULT 0
|
|
);
|
|
CREATE INDEX IF NOT EXISTS idx_cache_hash ON response_cache(hash);
|
|
CREATE INDEX IF NOT EXISTS idx_cache_created ON response_cache(created_at);
|
|
`);
|
|
|
|
// Idempotent migrations: add quota columns if they don't exist yet.
|
|
for (const col of [
|
|
"ALTER TABLE api_keys ADD COLUMN quota_daily INTEGER DEFAULT NULL",
|
|
"ALTER TABLE api_keys ADD COLUMN quota_weekly INTEGER DEFAULT NULL",
|
|
"ALTER TABLE api_keys ADD COLUMN quota_monthly INTEGER DEFAULT NULL",
|
|
]) {
|
|
try { db.exec(col); } catch (e) {
|
|
// SQLite throws "duplicate column name" if already present — safe to ignore.
|
|
if (!e.message?.includes("duplicate column")) throw e;
|
|
}
|
|
}
|
|
}
|
|
|
|
// ── Key CRUD ──
|
|
|
|
export function createKey(name) {
|
|
const key = "ocp_" + randomBytes(24).toString("base64url");
|
|
const d = getDb();
|
|
const stmt = d.prepare("INSERT INTO api_keys (key, name) VALUES (?, ?)");
|
|
const result = stmt.run(key, name);
|
|
return { id: result.lastInsertRowid, key, name };
|
|
}
|
|
|
|
export function listKeys() {
|
|
const d = getDb();
|
|
return d.prepare(
|
|
"SELECT id, key, name, created_at, revoked, quota_daily, quota_weekly, quota_monthly FROM api_keys ORDER BY created_at DESC"
|
|
).all().map(({ key, ...rest }) => ({
|
|
...rest,
|
|
keyPreview: key.slice(0, 8) + "..." + key.slice(-4),
|
|
}));
|
|
}
|
|
|
|
export function revokeKey(idOrName) {
|
|
const d = getDb();
|
|
const stmt = d.prepare(
|
|
"UPDATE api_keys SET revoked = 1 WHERE (id = ? OR name = ?) AND revoked = 0"
|
|
);
|
|
return stmt.run(idOrName, idOrName).changes > 0;
|
|
}
|
|
|
|
export function validateKey(key) {
|
|
const d = getDb();
|
|
const row = d.prepare(
|
|
"SELECT id, name FROM api_keys WHERE key = ? AND revoked = 0"
|
|
).get(key);
|
|
return row || null;
|
|
}
|
|
|
|
// ── Usage recording ──
|
|
|
|
export function recordUsage({ keyId, keyName, model, promptChars, responseChars, elapsedMs, success }) {
|
|
const d = getDb();
|
|
d.prepare(`
|
|
INSERT INTO usage_log (key_id, key_name, model, prompt_chars, response_chars, elapsed_ms, success)
|
|
VALUES (?, ?, ?, ?, ?, ?, ?)
|
|
`).run(keyId ?? null, keyName || "anonymous", model, promptChars, responseChars, elapsedMs, success ? 1 : 0);
|
|
}
|
|
|
|
// ── Usage queries ──
|
|
|
|
export function getUsageByKey({ since, until } = {}) {
|
|
const d = getDb();
|
|
let where = "WHERE 1=1";
|
|
const params = [];
|
|
if (since) { where += " AND created_at >= ?"; params.push(since); }
|
|
if (until) { where += " AND created_at <= ?"; params.push(until); }
|
|
|
|
return d.prepare(`
|
|
SELECT
|
|
key_name,
|
|
COUNT(*) as requests,
|
|
SUM(CASE WHEN success = 1 THEN 1 ELSE 0 END) as successes,
|
|
SUM(CASE WHEN success = 0 THEN 1 ELSE 0 END) as errors,
|
|
SUM(prompt_chars) as total_prompt_chars,
|
|
SUM(response_chars) as total_response_chars,
|
|
SUM(elapsed_ms) as total_elapsed_ms,
|
|
AVG(elapsed_ms) as avg_elapsed_ms,
|
|
MIN(created_at) as first_request,
|
|
MAX(created_at) as last_request
|
|
FROM usage_log
|
|
${where}
|
|
GROUP BY key_name
|
|
ORDER BY requests DESC
|
|
`).all(...params);
|
|
}
|
|
|
|
export function getUsageTimeline({ keyName, hours = 24 } = {}) {
|
|
const d = getDb();
|
|
const since = new Date(Date.now() - hours * 3600000).toISOString();
|
|
let where = "WHERE created_at >= ?";
|
|
const params = [since];
|
|
if (keyName) { where += " AND key_name = ?"; params.push(keyName); }
|
|
|
|
return d.prepare(`
|
|
SELECT
|
|
strftime('%Y-%m-%dT%H:00:00', created_at) as hour,
|
|
COUNT(*) as requests,
|
|
SUM(prompt_chars) as prompt_chars,
|
|
SUM(response_chars) as response_chars,
|
|
AVG(elapsed_ms) as avg_elapsed_ms
|
|
FROM usage_log
|
|
${where}
|
|
GROUP BY hour
|
|
ORDER BY hour
|
|
`).all(...params);
|
|
}
|
|
|
|
export function getRecentUsage(limit = 50) {
|
|
const d = getDb();
|
|
return d.prepare(`
|
|
SELECT key_name, model, prompt_chars, response_chars, elapsed_ms, success, created_at
|
|
FROM usage_log
|
|
ORDER BY created_at DESC
|
|
LIMIT ?
|
|
`).all(limit);
|
|
}
|
|
|
|
// ── SQLite datetime helper ──
|
|
// SQLite datetime('now') stores as 'YYYY-MM-DD HH:MM:SS' (no T, no Z).
|
|
// JavaScript .toISOString() produces 'YYYY-MM-DDTHH:MM:SS.sssZ'.
|
|
// String comparison between the two breaks for same-day ranges (T > space).
|
|
// This helper formats Date to match SQLite's format for correct comparisons.
|
|
function sqliteDatetime(date) {
|
|
return date.toISOString().replace("T", " ").replace(/\.\d{3}Z$/, "");
|
|
}
|
|
|
|
// ── Quota management ──
|
|
|
|
// Returns { period, limit, used, resetsIn } if a quota is exceeded, null otherwise.
|
|
// Anonymous/admin callers (keyId === null) are never subject to quotas.
|
|
export function checkQuota(keyId, _keyName) {
|
|
if (keyId === null || keyId === undefined) return null;
|
|
|
|
const d = getDb();
|
|
const keyRow = d.prepare(
|
|
"SELECT quota_daily, quota_weekly, quota_monthly FROM api_keys WHERE id = ? AND revoked = 0"
|
|
).get(keyId);
|
|
if (!keyRow) return null;
|
|
|
|
const now = new Date();
|
|
|
|
// UTC period boundaries (SQLite-compatible format)
|
|
const startOfToday = sqliteDatetime(new Date(Date.UTC(now.getUTCFullYear(), now.getUTCMonth(), now.getUTCDate())));
|
|
const sevenDaysAgo = sqliteDatetime(new Date(Date.now() - 7 * 86400000));
|
|
const thirtyDaysAgo = sqliteDatetime(new Date(Date.now() - 30 * 86400000));
|
|
|
|
// Next reset times for human display
|
|
const tomorrowUTC = new Date(Date.UTC(now.getUTCFullYear(), now.getUTCMonth(), now.getUTCDate() + 1));
|
|
function msToHuman(ms) {
|
|
if (ms <= 0) return "now";
|
|
const h = Math.floor(ms / 3600000);
|
|
const m = Math.floor((ms % 3600000) / 60000);
|
|
if (h >= 24) { const d = Math.floor(h / 24); return `${d}d ${h % 24}h`; }
|
|
return h > 0 ? `${h}h ${m}m` : `${m}m`;
|
|
}
|
|
|
|
// Single query for all periods (widest window = monthly)
|
|
const row = d.prepare(`
|
|
SELECT
|
|
SUM(CASE WHEN created_at >= ? THEN 1 ELSE 0 END) as daily_cnt,
|
|
SUM(CASE WHEN created_at >= ? THEN 1 ELSE 0 END) as weekly_cnt,
|
|
COUNT(*) as monthly_cnt
|
|
FROM usage_log
|
|
WHERE key_id = ? AND success = 1 AND created_at >= ?
|
|
`).get(startOfToday, sevenDaysAgo, keyId, thirtyDaysAgo);
|
|
|
|
const checks = [
|
|
{ period: "daily", limit: keyRow.quota_daily, used: row?.daily_cnt ?? 0, resetsIn: msToHuman(tomorrowUTC - now) },
|
|
{ period: "weekly", limit: keyRow.quota_weekly, used: row?.weekly_cnt ?? 0, resetsIn: "rolling 7-day window" },
|
|
{ period: "monthly", limit: keyRow.quota_monthly, used: row?.monthly_cnt ?? 0, resetsIn: "rolling 30-day window" },
|
|
];
|
|
|
|
for (const { period, limit, used, resetsIn } of checks) {
|
|
if (limit === null || limit === undefined) continue;
|
|
if (used >= limit) {
|
|
return { period, limit, used, resetsIn };
|
|
}
|
|
}
|
|
|
|
return null;
|
|
}
|
|
|
|
// Set quota for a key. Only updates fields explicitly present in the input object.
|
|
// Pass null to clear a specific limit. Omit a field to leave it unchanged.
|
|
export function updateKeyQuota(idOrName, updates = {}) {
|
|
const d = getDb();
|
|
const setClauses = [];
|
|
const params = [];
|
|
if ("daily" in updates) { setClauses.push("quota_daily = ?"); params.push(updates.daily ?? null); }
|
|
if ("weekly" in updates) { setClauses.push("quota_weekly = ?"); params.push(updates.weekly ?? null); }
|
|
if ("monthly" in updates){ setClauses.push("quota_monthly = ?");params.push(updates.monthly ?? null); }
|
|
if (setClauses.length === 0) return false;
|
|
params.push(idOrName, idOrName);
|
|
const result = d.prepare(
|
|
`UPDATE api_keys SET ${setClauses.join(", ")} WHERE id = ? OR name = ?`
|
|
).run(...params);
|
|
return result.changes > 0;
|
|
}
|
|
|
|
// Returns { daily: { limit, used }, weekly: { limit, used }, monthly: { limit, used } }
|
|
export function getKeyQuota(keyId) {
|
|
const d = getDb();
|
|
const keyRow = d.prepare(
|
|
"SELECT quota_daily, quota_weekly, quota_monthly FROM api_keys WHERE id = ?"
|
|
).get(keyId);
|
|
if (!keyRow) return null;
|
|
|
|
const now = new Date();
|
|
const startOfToday = sqliteDatetime(new Date(Date.UTC(now.getUTCFullYear(), now.getUTCMonth(), now.getUTCDate())));
|
|
const sevenDaysAgo = sqliteDatetime(new Date(Date.now() - 7 * 86400000));
|
|
const thirtyDaysAgo = sqliteDatetime(new Date(Date.now() - 30 * 86400000));
|
|
|
|
const row = d.prepare(`
|
|
SELECT
|
|
SUM(CASE WHEN created_at >= ? THEN 1 ELSE 0 END) as daily_cnt,
|
|
SUM(CASE WHEN created_at >= ? THEN 1 ELSE 0 END) as weekly_cnt,
|
|
COUNT(*) as monthly_cnt
|
|
FROM usage_log
|
|
WHERE key_id = ? AND success = 1 AND created_at >= ?
|
|
`).get(startOfToday, sevenDaysAgo, keyId, thirtyDaysAgo);
|
|
|
|
return {
|
|
daily: { limit: keyRow.quota_daily ?? null, used: row?.daily_cnt ?? 0 },
|
|
weekly: { limit: keyRow.quota_weekly ?? null, used: row?.weekly_cnt ?? 0 },
|
|
monthly: { limit: keyRow.quota_monthly ?? null, used: row?.monthly_cnt ?? 0 },
|
|
};
|
|
}
|
|
|
|
// ── Response cache ──
|
|
|
|
// Generate a cache key from model + messages + request params that affect output.
|
|
// opts.keyId isolates per-API-key cache pools (v2 hash format).
|
|
// When keyId is absent/null/empty, falls back to "anon" (shared anonymous pool).
|
|
export function cacheHash(model, messages, opts = {}) {
|
|
const keyId = opts.keyId || "anon";
|
|
const h = createHash("sha256");
|
|
h.update(`v2|k:${keyId}|`);
|
|
h.update(model);
|
|
if (opts.temperature != null) h.update(`t:${opts.temperature}`);
|
|
if (opts.max_tokens != null) h.update(`mt:${opts.max_tokens}`);
|
|
if (opts.top_p != null) h.update(`tp:${opts.top_p}`);
|
|
// #176: fold the server's boot-config epoch into the key, so a config change that shapes
|
|
// answers (operator system prompt, wrapper text, allowed tools, NO_CONTEXT) invalidates
|
|
// the persistent cache instead of serving answers composed under the old config. Callers
|
|
// that omit it (older paths, tests) hash byte-identically to before.
|
|
if (opts.configEpoch != null) h.update(`ce:${opts.configEpoch}|`);
|
|
// Structured-output (OpenAI response_format / json_mode) requests must never share a cache slot
|
|
// with the conversational answer to the same prompt, nor with a different schema — the steering
|
|
// instruction and validated JSON payload differ. Keying on the detected descriptor isolates them.
|
|
// Absent for normal requests → hashes are byte-identical to pre-change.
|
|
if (opts.structured != null) h.update(`s:${JSON.stringify(opts.structured)}`);
|
|
for (const m of messages) {
|
|
h.update(m.role || "");
|
|
h.update(typeof m.content === "string" ? m.content : JSON.stringify(m.content));
|
|
}
|
|
return h.digest("hex");
|
|
}
|
|
|
|
// Check whether any message (or content part) carries an Anthropic cache_control field.
|
|
// If true, OCP should skip its own cache to avoid interfering with prompt-caching intent.
|
|
export function hasCacheControl(messages) {
|
|
for (const m of messages || []) {
|
|
if (m && typeof m === "object") {
|
|
if (m.cache_control) return true;
|
|
if (Array.isArray(m.content)) {
|
|
for (const part of m.content) {
|
|
if (part && typeof part === "object" && part.cache_control) return true;
|
|
}
|
|
}
|
|
}
|
|
}
|
|
return false;
|
|
}
|
|
|
|
// Look up a cached response. Returns { response, hits } or null.
|
|
// Also updates last_hit_at and increments hits counter on hit.
|
|
export function getCachedResponse(hash, ttlMs) {
|
|
const d = getDb();
|
|
const cutoff = sqliteDatetime(new Date(Date.now() - ttlMs));
|
|
const row = d.prepare(
|
|
"SELECT id, response, hits FROM response_cache WHERE hash = ? AND created_at >= ?"
|
|
).get(hash, cutoff);
|
|
if (!row) return null;
|
|
// Update hit stats
|
|
d.prepare("UPDATE response_cache SET hits = hits + 1, last_hit_at = datetime('now') WHERE id = ?").run(row.id);
|
|
return { response: row.response, hits: row.hits + 1 };
|
|
}
|
|
|
|
// Store a response in the cache
|
|
export function setCachedResponse(hash, model, response) {
|
|
const d = getDb();
|
|
// Upsert: if hash already exists (race condition), just update
|
|
d.prepare(`
|
|
INSERT INTO response_cache (hash, model, response) VALUES (?, ?, ?)
|
|
ON CONFLICT(hash) DO UPDATE SET response = excluded.response, created_at = datetime('now'), hits = 0
|
|
`).run(hash, model, response);
|
|
}
|
|
|
|
// Clear all cached responses, or expired ones only
|
|
export function clearCache(ttlMs = null) {
|
|
const d = getDb();
|
|
if (ttlMs === null) {
|
|
const result = d.prepare("DELETE FROM response_cache").run();
|
|
return result.changes;
|
|
}
|
|
const cutoff = sqliteDatetime(new Date(Date.now() - ttlMs));
|
|
const result = d.prepare("DELETE FROM response_cache WHERE created_at < ?").run(cutoff);
|
|
return result.changes;
|
|
}
|
|
|
|
// Get cache statistics
|
|
export function getCacheStats() {
|
|
const d = getDb();
|
|
const total = d.prepare("SELECT COUNT(*) as cnt FROM response_cache").get()?.cnt ?? 0;
|
|
const totalHits = d.prepare("SELECT SUM(hits) as total FROM response_cache").get()?.total ?? 0;
|
|
const sizeBytes = d.prepare("SELECT SUM(LENGTH(response)) as size FROM response_cache").get()?.size ?? 0;
|
|
return { entries: total, totalHits, sizeBytes };
|
|
}
|
|
|
|
// ── Singleflight stampede protection ──
|
|
|
|
// In-memory singleflight Map: hash → { promise, requesters }
|
|
// Deduplicates concurrent identical cache-miss flows so only one upstream call runs.
|
|
// Per ADR 0005 / spec D4: in-process scope only (single Node process per host).
|
|
const inflightMap = new Map();
|
|
|
|
// `retryIf` (optional, audit finding M1): a predicate applied on the FOLLOWER path only.
|
|
// When a follower joins an existing flight and the shared promise rejects with an error for
|
|
// which retryIf(err) is true (in practice: the LEADER's client disconnected while queued —
|
|
// an error that is personal to the leader, not a verdict about the upstream), the follower
|
|
// does NOT inherit that rejection. Instead it re-enters singleflight with its OWN fn: it
|
|
// either becomes the new leader (the map entry is already deleted — see the finally below,
|
|
// which runs before any follower's catch because it is attached upstream of the promise the
|
|
// followers await) or joins a flight another retrying follower just created. The leader's
|
|
// own rejection is never retried here — its error belongs to it (leader path returns the
|
|
// bare promise). Callers that pass no retryIf get the exact pre-M1 share-everything behavior.
|
|
export function singleflight(hash, fn, retryIf) {
|
|
const existing = inflightMap.get(hash);
|
|
if (existing) {
|
|
existing.requesters++;
|
|
if (!retryIf) return existing.promise;
|
|
return existing.promise.catch((err) => {
|
|
if (!retryIf(err)) throw err;
|
|
return singleflight(hash, fn, retryIf);
|
|
});
|
|
}
|
|
// Wrap fn() in Promise.resolve().then() so synchronous throws don't escape.
|
|
const promise = Promise.resolve().then(fn).finally(() => {
|
|
inflightMap.delete(hash);
|
|
});
|
|
inflightMap.set(hash, { promise, requesters: 1 });
|
|
return promise;
|
|
}
|
|
|
|
export function getInflightStats() {
|
|
let totalRequesters = 0;
|
|
for (const entry of inflightMap.values()) totalRequesters += entry.requesters;
|
|
return {
|
|
inflight: inflightMap.size,
|
|
requesters: totalRequesters,
|
|
};
|
|
}
|
|
|
|
// Find a key by id or name (returns { id, name } or null)
|
|
export function findKey(idOrName) {
|
|
const d = getDb();
|
|
return d.prepare("SELECT id, name FROM api_keys WHERE id = ? OR name = ?").get(idOrName, idOrName) || null;
|
|
}
|
|
|
|
export function closeDb() {
|
|
if (db) { db.close(); db = null; dbPath = undefined; } // clear both — a path to a closed db is a footgun
|
|
}
|