Files
memory-continuity/memory/2026-03-07.md
T

3.9 KiB

2026-03-07

Memory/Embeddings incident + external API inventory (OpenClaw)

  • Issue: memory_search (semantic memory retrieval) failed with 429 insufficient_quota due to embeddings provider quota exhaustion.
  • Root cause: OpenClaw agents.defaults.memorySearch.provider = openai with remote.apiKey present (OpenAI embeddings key). Not related to Brave web search.

External API / Provider inventory (from ~/.openclaw/openclaw.json)

  • Chat/model routing:
    • Primary: openai-codex/gpt-5.4
    • Fallbacks: openai-codex/gpt-5.2, github-copilot/claude-opus-4.6
  • Memory semantic search:
    • memorySearch.enabled = true
    • memorySearch.provider = openai
    • memorySearch.remote.apiKey present
    • memorySearch.sources = [memory, sessions]
  • Web tools:
    • tools.web.search.enabled = true (Brave Search key present)
    • tools.web.fetch.enabled = true
  • Auth profiles present: anthropic:manual, github-copilot:github, openai:manual, openai-codex:default.

User request

  • Tao requested: archive this inventory info and keep it updated whenever APIs/providers are added/removed in the future.

Comms protocol update (task progress)

  • Tao reported a recurring issue: after assigning a task, the agent may go silent (stuck or working) unless Tao explicitly asks.
  • Preference selected: progress updates at milestones (Option B).
  • Policy: acknowledge within ~30s with plan+ETA; then send updates at each milestone; if ETA slips, proactively report block/reason/next-step.

Memory architecture preference (agents)

  • Tao preference: each Agent should have independent soul + independent memory; do not share memory across agents.
  • Disaster recovery preference: restore automation should be maximized; when restoring, default to safe mode: rename existing ~/.openclaw before restoring (Option B).

Clawkeeper project + gateway safety incident

  • Created and published new open-source repo: https://github.com/dtzp555-max/clawkeeper
    • Purpose: OpenClaw Memory Ops Kit (doctor + embeddings probe + provider switch + backup plan + verify-backup + restore tooling).
    • Added doctor to perform a real OpenAI embeddings probe (minimal request) to detect 429/insufficient_quota early.
    • Added backup-plan --json and verify-backup <tgz> to ensure backups include required paths.
    • Added restore-guide and restore-apply with SAFE default: do not restart gateway unless explicitly requested.
  • Incident: calling openclaw gateway restart during restore testing could leave LaunchAgent in a bad state (not loaded / restart fails), forcing openclaw gateway install to recover; can cause temporary messaging disconnects.
  • New hard rule from Tao: do not harm the primary Mac gateway during development/testing; use a separate test machine.

Test machine ("lobster" 232)

  • Host: 172.16.2.232 (ubuntu-srv), user: administrator.
  • SSH key: ~/.ssh/openclaw_backup_key
  • Tao confirmed key was installed; verified login: ssh -i ~/.ssh/openclaw_backup_key administrator@172.16.2.232 'echo OK'.
  • Policy: perform potentially disruptive OpenClaw install/restart tests on 232 instead of the primary Mac.

Multi-agent dev workflow preference

  • Tao clarified the intended architecture for complex development:
    • main = PM / architect / coordinator / reviewer / reporter.
    • execution agents do the hands-on implementation.
  • Dedicated execution agent for Codex work:
    • Name: codex_worker
    • Workspace: ~/.openclaw/workspaces/codex_worker
    • Agent dir: ~/.openclaw/agents/codex_worker/agent
    • Model target: openai-codex/gpt-5.4
  • Preference: codex_worker should avoid automatic fallback; if GPT-5.4 is unavailable, report to Tao for model decision.
  • Preference: execution agents should get the strongest execution model; main should remain strong for PM/summary work, but not necessarily do the coding itself.
  • This architecture is the current default, but Tao may revise it as needs change.