fleet-memory/hindsight-dev/benchmarks/README.md
Nicolò Boschi aefb3fcf4d
fix: improve async batch retain with large payloads (#366)
* fix: improve async batch retain with large payloads

* fix: improve async batch retain with large payloads

* api

* api

* api

* api

* api

* Clean up perf benchmark: keep only Python files

- Remove README.md and PERFORMANCE_FINDINGS.md
- Remove results/ JSON files (gitignored)
- Remove test_data/ directory
- Keep only __init__.py and retain_perf.py

* docs: explain automatic batch optimization for async retain

- Add section explaining Hindsight automatically handles batch sizing
- Users don't need to manually tune batch sizes with async mode
- Hindsight splits large batches (>10k tokens) into optimized sub-batches
- Include example showing best practices

* docs: remove emojis and code example from performance page

* fix: correct OperationDetails type to match API response

- Change optional fields to use | null instead of ?
- Fixes TypeScript compilation error in control plane build

* fix: use discriminated union for OperationDetails type

- Support both success and error states properly
- Fixes TypeScript error when setting error state

* fix: use unique document_ids in batch retain examples

- Each item in a batch must have unique document_id
- Update both Python and JavaScript examples
- Fixes test-doc-examples CI failure

* chore: trigger CI

* fix: test mocking and duplicate document_ids in examples

- Mock _get_pool() in test_async_retain_tags.py to avoid _initialized error
- Set _initialized = True on mocked MemoryEngine instances
- Fix duplicate document_ids in retain.py and retain.mjs examples

* fix: properly mock async pool/connection and fix more duplicate document_ids

- Use AsyncMock for pool.acquire() to fix 'can't be used in await' error
- Fix duplicate document_ids in retain-async examples (retain.py and retain.mjs)
- Remove batch-level document_id parameter that caused duplicates

* ci: collect all doc example failures and show summary

- Run all Python/Node.js/CLI examples regardless of individual failures
- Collect failure list and display summary at the end
- Show pass/fail count and list of failed files
- Exit with failure only after running all examples

* refactor: extract doc example testing to standalone script

- Create scripts/test-doc-examples.sh to run all examples
- Collects logs of failed examples separately
- Shows full error logs only for failures at the end
- Clean summary with pass/fail counts
- Proper exit codes
- Replaces inline bash in CI workflow

* fix: doc examples - duplicate document_ids and error handling

- retain.py: move document_id to item level to avoid duplicates
- documents.mjs: add error handling for getDocument to show clear error message

* fix: update tests for duplicate document_id validation

- test_async_retain_tags: verify operation structure instead of exact UUID
- test_delete_bank: use unique document_ids (team-doc-1, team-doc-2)
2026-02-16 12:51:42 +01:00

3.3 KiB

Hindsight Benchmarks

This directory contains benchmark suites for evaluating Hindsight's memory capabilities.

Prerequisites

  1. Set up your environment variables in .env at the project root:

    cp .env.example .env
    # Edit .env with your API keys
    
  2. Make sure you have uv installed.

Available Benchmarks

LoComo

Tests conversational memory with multi-turn dialogues.

# Run from project root
./scripts/benchmarks/run-locomo.sh

# With options
./scripts/benchmarks/run-locomo.sh --max-conversations 10
./scripts/benchmarks/run-locomo.sh --skip-ingestion  # Reuse existing data
./scripts/benchmarks/run-locomo.sh --use-think       # Use think API
./scripts/benchmarks/run-locomo.sh --conversation conv-26  # Single conversation

Options:

  • --max-conversations N - Limit number of conversations
  • --max-questions N - Limit questions per conversation
  • --skip-ingestion - Skip data ingestion, use existing
  • --use-think - Use think API instead of search + LLM
  • --conversation NAME - Run specific conversation only
  • --api-url URL - Custom API URL (default: local memory)
  • --only-failed - Retry only failed questions
  • --only-invalid - Retry only invalid questions

LongMemEval

Tests long-term memory across different categories.

# Run from project root
./scripts/benchmarks/run-longmemeval.sh

# With options
./scripts/benchmarks/run-longmemeval.sh --max-instances 50
./scripts/benchmarks/run-longmemeval.sh --category single-session-user
./scripts/benchmarks/run-longmemeval.sh --parallel 4  # Faster evaluation

Options:

  • --max-instances N - Limit total questions
  • --max-instances-per-category N - Limit per category
  • --skip-ingestion - Skip data ingestion
  • --category NAME - Filter by category:
    • single-session-user
    • multi-session
    • single-session-preference
    • temporal-reasoning
    • knowledge-update
    • single-session-assistant
  • --parallel N - Parallel instances (default: 1)
  • --only-failed - Retry failed questions
  • --fill - Resume interrupted runs

Consolidation Performance

Tests consolidation throughput and identifies bottlenecks.

./scripts/benchmarks/run-consolidation.sh

# With custom memory count
NUM_MEMORIES=200 ./scripts/benchmarks/run-consolidation.sh

Retain Performance

Measures retain operation performance (throughput and token usage).

Prerequisites: API server must be running (./scripts/dev/start-api.sh)

# Basic usage
./scripts/benchmarks/run-retain-perf.sh \
    --document hindsight-dev/benchmarks/perf/test_data/sample_document.txt

# Save results to JSON
./scripts/benchmarks/run-retain-perf.sh \
    --document ./my_document.txt \
    --bank-id my-test-bank \
    --output results/retain_perf.json

Options:

  • --document PATH - Document file to retain (required)
  • --bank-id ID - Bank ID to use (default: perf-test)
  • --context TEXT - Optional context
  • --api-url URL - API URL (default: http://localhost:8000)
  • --timeout SECONDS - Request timeout (default: 300)
  • --output PATH - Save results to JSON file

See perf/README.md for detailed documentation.

Visualizer

View benchmark results in a web UI:

./scripts/benchmarks/start-visualizer.sh
# Opens at http://localhost:8001

Results

Results are saved in JSON format in each benchmark's results/ directory.