* doc: update cookbook
* fix(cookbook): preserve tag keys during sync, strip local .md links
- Fix extract_tags_from_readme/notebook to return dict[str,str] preserving
sdk/topic keys instead of bare values, preventing topics like
"Customer Service" from being misclassified as SDK
- Add strip_local_md_links() to remove relative .md references that
would cause broken link errors in Docusaurus build
* ci: run test-doc-examples independently without waiting for test-rust-cli
Build the CLI directly in the job instead of downloading the artifact,
so test-doc-examples can start at the beginning in parallel with all other jobs.
* feat: webhook system with task-owned retry, retain.completed event, and UI
- New webhook system: register per-bank webhooks with HMAC signing, configurable
HTTP method/timeout/headers/params (http_config JSONB), and PATCH support
- Webhook deliveries run as async_operations (webhook_delivery type) with
task-owned retry via RetryTaskAt exception and exponential backoff
(60s / 5m / 30m / 2h / 8h, max 6 attempts)
- New retain.completed event fires per-document for both sync and async retain
- Delivery debug info (status code, response body) stored in result_metadata
- Control plane UI: webhooks tab per bank with create/edit/delete and a
deliveries table with cursor pagination and expandable response details
- 28 webhook tests covering HMAC signing, delivery retries, CRUD endpoints,
PATCH update, and retain.completed queuing
- Docs page at developer/api/webhooks documenting event payloads and delivery
- OpenAPI spec and all client SDKs (Python, TypeScript, Rust, Go) regenerated
* fix: update tests for task-owned retry model and guard _webhook_manager attribute
- test_worker.py: test_executor_exception_triggers_retry now raises RetryTaskAt
(plain exceptions are immediate failures in the new system); rename
test_executor_exception_marks_failed_after_max_retries to
test_executor_exception_marks_failed_immediately to reflect new semantics
- test_batch_api.py: remove max_retries kwarg from WorkerPoller constructor
- memory_engine.py: use getattr for _webhook_manager in _fire_retain_webhook
to avoid AttributeError when engine is created without __init__ (tests)
* fix: remove max_retries from benchmark WorkerPoller call
* fix(webhooks): transactional outbox, observations_deleted tracking, sidebar
- Queue webhook delivery rows atomically with the primary operation using the
transactional outbox pattern — prevents lost events on process crash:
- Retain (sync + async): outbox_callback passed into orchestrator.retain_batch
and called inside the DB transaction, replacing the post-commit fire call
- Consolidation: new _mark_operation_completed_and_fire_webhook combines the
status UPDATE and webhook INSERT in one transaction
- Added fire_event_with_conn() to WebhookManager for in-connection delivery
- Track observations_deleted count in consolidation stats and expose it in the
consolidation.completed webhook payload (was always None)
- Add Webhooks page to docs sidebar
- Document at-least-once delivery guarantee with operation_id dedup guidance
* fix(ui): add retain.completed to available webhook event types
* feat(ui): add delete confirmation dialog for webhooks
* fix(webhooks): include operation_id in task_payload so delivery is marked completed
The task_payload JSON was missing the operation_id field, causing execute_task
to see operation_id=None and skip _mark_operation_completed — leaving every
delivery row stuck in 'pending' forever.
Added a test that inserts a real async_operations row and verifies the status
transitions to 'completed' after a successful execute_task call.
* style: fix prettier formatting in webhooks-view
|
||
|---|---|---|
| .. | ||
| common | ||
| consolidation | ||
| locomo | ||
| longmemeval | ||
| perf | ||
| visualizer | ||
| .DS_Store | ||
| __init__.py | ||
| README.md | ||
Hindsight Benchmarks
This directory contains benchmark suites for evaluating Hindsight's memory capabilities.
Prerequisites
-
Set up your environment variables in
.envat the project root:cp .env.example .env # Edit .env with your API keys -
Make sure you have
uvinstalled.
Available Benchmarks
LoComo
Tests conversational memory with multi-turn dialogues.
# Run from project root
./scripts/benchmarks/run-locomo.sh
# With options
./scripts/benchmarks/run-locomo.sh --max-conversations 10
./scripts/benchmarks/run-locomo.sh --skip-ingestion # Reuse existing data
./scripts/benchmarks/run-locomo.sh --use-think # Use think API
./scripts/benchmarks/run-locomo.sh --conversation conv-26 # Single conversation
Options:
--max-conversations N- Limit number of conversations--max-questions N- Limit questions per conversation--skip-ingestion- Skip data ingestion, use existing--use-think- Use think API instead of search + LLM--conversation NAME- Run specific conversation only--api-url URL- Custom API URL (default: local memory)--only-failed- Retry only failed questions--only-invalid- Retry only invalid questions
LongMemEval
Tests long-term memory across different categories.
# Run from project root
./scripts/benchmarks/run-longmemeval.sh
# With options
./scripts/benchmarks/run-longmemeval.sh --max-instances 50
./scripts/benchmarks/run-longmemeval.sh --category single-session-user
./scripts/benchmarks/run-longmemeval.sh --parallel 4 # Faster evaluation
Options:
--max-instances N- Limit total questions--max-instances-per-category N- Limit per category--skip-ingestion- Skip data ingestion--category NAME- Filter by category:single-session-usermulti-sessionsingle-session-preferencetemporal-reasoningknowledge-updatesingle-session-assistant
--parallel N- Parallel instances (default: 1)--only-failed- Retry failed questions--fill- Resume interrupted runs
Consolidation Performance
Tests consolidation throughput and identifies bottlenecks.
./scripts/benchmarks/run-consolidation.sh
# With custom memory count
NUM_MEMORIES=200 ./scripts/benchmarks/run-consolidation.sh
Retain Performance
Measures retain operation performance (throughput and token usage).
Prerequisites: API server must be running (./scripts/dev/start-api.sh)
# Basic usage
./scripts/benchmarks/run-retain-perf.sh \
--document hindsight-dev/benchmarks/perf/test_data/sample_document.txt
# Save results to JSON
./scripts/benchmarks/run-retain-perf.sh \
--document ./my_document.txt \
--bank-id my-test-bank \
--output results/retain_perf.json
Options:
--document PATH- Document file to retain (required)--bank-id ID- Bank ID to use (default: perf-test)--context TEXT- Optional context--api-url URL- API URL (default: http://localhost:8000)--timeout SECONDS- Request timeout (default: 300)--output PATH- Save results to JSON file
See perf/README.md for detailed documentation.
Visualizer
View benchmark results in a web UI:
./scripts/benchmarks/start-visualizer.sh
# Opens at http://localhost:8001
Results
Results are saved in JSON format in each benchmark's results/ directory.