fleet-memory/hindsight-integrations/openclaw
Nicolò Boschi 576016f5dc
feat: add @vectorize-io/hindsight-all daemon lifecycle package (#949)
* feat: add @vectorize-io/hindsight-embed daemon lifecycle package

Create a new top-level `hindsight-embed-npm/` package that owns the daemon
lifecycle for the Python `hindsight-embed` CLI: spawning via `uvx`, writing
the profile, waiting for `/health`, and shutting down. Nothing more.

Deliberately does not ship an HTTP client — `@vectorize-io/hindsight-client`
already covers retain / recall / reflect / createBank against the Hindsight
API, and the two packages compose: once `manager.start()` returns, consumers
talk to the daemon via `new HindsightClient({ baseUrl: manager.getBaseUrl() })`.

`HindsightEmbedManagerOptions.env` forwards an arbitrary `Record<string,
string>` to both the daemon process and the profile config via `--env K=V`,
and `extraProfileCreateArgs` / `extraDaemonStartArgs` escape hatches cover
any new CLI flag without waiting for a wrapper release.

Refactor `hindsight-integrations/openclaw` to consume both packages:
`HindsightEmbedManager` for daemon lifecycle in local mode, `HindsightClient`
for all HTTP memory operations. Drop the bespoke subprocess/HTTP client that
used to live in openclaw. The retain queue stays local to openclaw (it's a
client-side reliability workaround with a single consumer today — will move
to the client package or server-side when a second consumer needs it).

Wire the new package into the main release pipeline (versioned alongside
the other core packages, published from `v*` tags) and add a CI build job.

* docs: add Embedded Node.js SDK page for @vectorize-io/hindsight-embed

* refactor: rename hindsight-embed-npm to hindsight-all, restructure docs sidebar

The Node package previously named @vectorize-io/hindsight-embed was
semantically misnamed: hindsight-embed (Python) is a CLI tool, while what
this Node package actually provides is the Node equivalent of hindsight-all
— a programmatic lifecycle manager for a local Hindsight daemon. Rename to
match.

Package rename
  - hindsight-embed-npm/ → hindsight-all-npm/ (git mv, history preserved)
  - @vectorize-io/hindsight-embed → @vectorize-io/hindsight-all
  - class HindsightEmbedManager → HindsightServer (matches Python hindsight-all)
  - HindsightEmbedManagerOptions → HindsightServerOptions
  - src/manager.ts → src/server.ts, src/manager.test.ts → src/server.test.ts
  - openclaw (index.ts, backfill.ts, tests) and the claude-code Python port
    updated to reference the new names

Docs restructure
  - Split sdks/python.md: now client-only content. New sdks/hindsight-all.md
    covers the programmatic hindsight-all Python package (HindsightServer and
    HindsightEmbedded).
  - Rename sdks/embed-npm.md → sdks/hindsight-all-npm.md with HindsightServer
    examples.
  - New "Installation" sidebar section, placed after Hosting, containing
    Docker / Kubernetes / Bare Metal (anchor links into developer/installation)
    plus Programmatic API (Python), Programmatic API (Node.js), and Daemon CLI.
  - Add si-docker, si-kubernetes, si-nodedotjs, lu-hard-drive to the sidebar
    ICON_MAP.

Docs dev-server fix
  - docusaurus.config.ts: drop the flaky NODE_ENV sniff for including the
    "Next" version. Use INCLUDE_CURRENT_VERSION exclusively. NODE_ENV was
    unreliable across hot-reload paths and caused the Next version to
    disappear intermittently when editing files.
  - scripts/dev/start-docs.sh: export INCLUDE_CURRENT_VERSION=true so local
    dev always shows Next; production builds leave it unset.

Lockfile cleanup
  - package-lock.json and hindsight-integrations/openclaw/package-lock.json
    had extraneous hindsight-embed-npm blocks left over from the rename.
    Removed manually and verified with npm install.

* ci: fix openclaw jobs by pre-building workspace deps; regenerate docs-skill

The build-openclaw-integration and test-openclaw-integration jobs failed
with "Failed to resolve entry for package @vectorize-io/hindsight-all"
because openclaw depends on two monorepo workspaces via `file:` deps
(@vectorize-io/hindsight-client and @vectorize-io/hindsight-all) whose
`dist/` directories are gitignored and never built before openclaw's npm ci.
Both jobs now install the root workspace and build the two deps first,
mirroring the release-control-plane pattern.

Also regenerate skills/hindsight-docs/references/* via
./scripts/generate-docs-skill.sh:
  - new skill pages for sdks/hindsight-all{.md,-npm.md}
  - updated skill pages for sdks/embed.md and sdks/python.md to match
    the new H1s and split content
  - incidental refreshes to changelog/index.md, developer/models.md,
    openapi.json, and uv.lock that verify-generated-files picked up

* ci: build openclaw before running tests so symlink test can realpath dist
2026-04-10 15:51:44 +02:00
..
src feat: add @vectorize-io/hindsight-all daemon lifecycle package (#949) 2026-04-10 15:51:44 +02:00
tests feat: add @vectorize-io/hindsight-all daemon lifecycle package (#949) 2026-04-10 15:51:44 +02:00
.gitignore feat(openclaw): add config-aware history backfill CLI (#878) 2026-04-09 10:44:19 +02:00
install.sh fix: openclaw improve config setup (#258) 2026-01-30 17:36:49 +01:00
openclaw.plugin.json feat(openclaw): add session pattern filtering for ignore and stateless sessions (#909) 2026-04-09 10:43:35 +02:00
package-lock.json feat: add @vectorize-io/hindsight-all daemon lifecycle package (#949) 2026-04-10 15:51:44 +02:00
package.json feat: add @vectorize-io/hindsight-all daemon lifecycle package (#949) 2026-04-10 15:51:44 +02:00
README.md feat(openclaw): add config-aware history backfill CLI (#878) 2026-04-09 10:44:19 +02:00
tsconfig.json fix: rename openclawd to openclaw (#252) 2026-01-30 13:04:42 +01:00
vitest.config.ts fix: rename openclawd to openclaw (#252) 2026-01-30 13:04:42 +01:00
vitest.integration.config.ts fix: improve openclaw test coverage (#396) 2026-02-18 14:10:33 +01:00

Hindsight Memory Plugin for OpenClaw

Biomimetic long-term memory for OpenClaw using Hindsight. Automatically captures conversations and intelligently recalls relevant context.

Quick Start

# 1. Configure your LLM provider for memory extraction
# Option A: OpenAI
export OPENAI_API_KEY="sk-your-key"

# Option B: Claude Code (no API key needed)
export HINDSIGHT_API_LLM_PROVIDER=claude-code

# Option C: OpenAI Codex (no API key needed)
export HINDSIGHT_API_LLM_PROVIDER=openai-codex

# 2. Install and enable the plugin
openclaw plugins install @vectorize-io/hindsight-openclaw

# 3. Start OpenClaw
openclaw gateway

That's it! The plugin will automatically start capturing and recalling memories.

Features

  • Auto-capture and auto-recall of memories each turn, injected into system prompt space so recalled memories stay out of the visible chat transcript
  • Memory isolation — configurable per agent, channel, user, or provider via dynamicBankGranularity
  • Historical backfill CLI — import prior OpenClaw session history into Hindsight using the active plugin bank-routing config by default
  • Retention controls — choose which message roles to retain, toggle auto-retain on/off, and stamp retained documents with consistent tags/source metadata

Configuration

Optional settings in ~/.openclaw/openclaw.json under plugins.entries.hindsight-openclaw.config:

Option Default Description
apiPort 9077 Port for the local Hindsight daemon
daemonIdleTimeout 0 Seconds before daemon shuts down from inactivity (0 = never)
embedPort 0 Port for hindsight-embed server (0 = auto-assign)
embedVersion "latest" hindsight-embed version
embedPackagePath Local path to hindsight-embed package for development
bankMission Agent identity/purpose stored on the memory bank. Helps the engine understand context for better fact extraction. Set once per bank — not a recall prompt.
llmProvider auto-detect LLM provider override for memory extraction (openai, anthropic, gemini, groq, ollama, openai-codex, claude-code)
llmModel provider default LLM model override used with llmProvider
llmApiKeyEnv provider standard env var Custom env var name for the provider API key
dynamicBankId true Enable per-context memory banks
bankId Static bank ID used when dynamicBankId is false. Can also be set with HINDSIGHT_BANK_ID.
bankIdPrefix Prefix for bank IDs (e.g. "prod")
retainTags [] Tags applied to every retained document, useful for cross-agent/source labeling (e.g. source_system:openclaw, agent:agentname)
retainSource "openclaw" source value written into retained document metadata
dynamicBankGranularity ["agent", "channel", "user"] Fields used to derive bank ID. Options: agent, channel, user, provider
excludeProviders ["heartbeat"] Message providers to skip for recall/retain (e.g. heartbeat, slack, telegram, discord)
autoRecall true Auto-inject memories before each turn. Set to false when the agent has its own recall tool.
autoRetain true Auto-retain conversations after each turn
retainRoles ["user", "assistant"] Which message roles to retain. Options: user, assistant, system, tool
retainEveryNTurns 1 Retain every Nth turn. 1 = every turn (default). Values > 1 enable chunked retention with a sliding window.
retainOverlapTurns 0 Extra prior turns included when chunked retention fires. Window = retainEveryNTurns + retainOverlapTurns. Only applies when retainEveryNTurns > 1.
recallBudget "mid" Recall effort: low, mid, or high. Higher budgets use more retrieval strategies.
recallMaxTokens 1024 Max tokens for recall response. Controls how much memory context is injected per turn.
recallTypes ["world", "experience"] Memory types to recall. Options: world, experience, observation. Excludes verbose observation entries by default.
recallRoles ["user", "assistant"] Roles included when building prior context for recall query composition. Options: user, assistant, system, tool.
recallTopK Max number of memories to inject per turn. Applied after API response as a hard cap.
recallContextTurns 1 Number of user turns to include when composing recall query context. 1 keeps latest-message-only behavior.
recallMaxQueryChars 800 Maximum character length for the composed recall query before calling recall.
recallPromptPreamble built-in string Prompt text placed above recalled memories in the injected <hindsight_memories> system-context block.
hindsightApiUrl External Hindsight API URL (skips local daemon)
hindsightApiToken Auth token for external API
ignoreSessionPatterns [] Session key glob patterns to skip entirely — no recall, no retain (e.g. ["agent:*:cron:**"])
statelessSessionPatterns [] Session key glob patterns for read-only sessions — retain is always skipped; recall is skipped when skipStatelessSessions is true (e.g. ["agent:*:subagent:**", "agent:*:heartbeat:**"])
skipStatelessSessions true When true, sessions matching statelessSessionPatterns also skip recall. Set to false to allow recall but still skip retain.

Session pattern filtering

ignoreSessionPatterns and statelessSessionPatterns accept glob patterns matched against the session key (format: agent:<agentId>:<type>:<uuid>).

Glob syntax:

  • * — matches any characters except : (single segment)
  • ** — matches anything including : (multiple segments)
Pattern Matches
agent:*:cron:** All cron sessions for any agent
agent:*:subagent:** All subagent sessions for any agent
agent:main:** All sessions under the main agent

Difference between the two options:

ignoreSessionPatterns statelessSessionPatterns
Retain Skipped Always skipped
Recall Skipped Skipped only when skipStatelessSessions: true

Example config — exclude cron jobs from memory entirely, allow subagents to read but not write memories:

{
  "ignoreSessionPatterns": ["agent:*:cron:**"],
  "statelessSessionPatterns": ["agent:*:subagent:**"],
  "skipStatelessSessions": false
}

Retention details

Retained documents use stable session-scoped IDs like openclaw:agent:agentname:discord:channel:123:turn:000001 (or ...:window:000002 for chunked retention), and include richer metadata such as session_key, agent_id, provider, channel_id, thread_id, sender_id, turn_index, and retention_scope.

Documentation

For full documentation, configuration options, troubleshooting, and development guide, see:

OpenClaw Integration Documentation

Development

To test local changes to the Hindsight package before publishing:

  1. Add embedPackagePath to your plugin config in ~/.openclaw/openclaw.json:
{
  "plugins": {
    "entries": {
      "hindsight-openclaw": {
        "enabled": true,
        "config": {
          "embedPackagePath": "/path/to/hindsight-wt3/hindsight-embed"
        }
      }
    }
  }
}
  1. The plugin will use uv run --directory <path> hindsight-embed instead of uvx hindsight-embed@latest

  2. To use a specific profile for testing:

# Check daemon status
uvx hindsight-embed@latest -p openclaw daemon status

# View logs
tail -f ~/.hindsight/profiles/openclaw.log

# List profiles
uvx hindsight-embed@latest profile list

Backfilling Existing OpenClaw History

The package includes a config-aware backfill CLI for importing historical OpenClaw sessions into Hindsight.

By default it mirrors the active plugin settings for:

  • dynamicBankId
  • dynamicBankGranularity
  • bankIdPrefix
  • local daemon vs external hindsightApiUrl

Dry-run example:

npx --package @vectorize-io/hindsight-openclaw hindsight-openclaw-backfill \
  --openclaw-root ~/.openclaw \
  --dry-run

Direct invocation from a built checkout:

node dist/backfill.js --openclaw-root ~/.openclaw --dry-run

Migration-oriented overrides are explicit:

node dist/backfill.js \
  --openclaw-root ~/.openclaw \
  --bank-strategy agent \
  --agent proj-run \
  --resume \
  --max-pending-operations 10

Useful options:

  • --agent <id> limit import to selected agents
  • --exclude-archive ignore sessions-archive-from-migration_backup
  • --bank-strategy mirror-config|agent|fixed
  • --resume skip only entries already finalized as completed
  • --checkpoint <path> store progress outside the default location
  • --wait-until-drained block until the touched bank queues have finished and checkpoint state can be finalized

License

MIT