* fix(recall): cap entity fanout in graph expansion to prevent slow queries On large banks, the entity co-occurrence self-join in _expand_combined() produces massive intermediate row counts when seeds reference high-fanout entities (e.g. an entity with 25K+ mentions). This causes recall latency to degrade significantly. Changes: - Replace unbounded entity self-join with LATERAL per-entity cap (graph_per_entity_limit, default 200), reducing intermediate rows from potentially millions to at most num_entities * 200 - Add ORDER BY unit_id DESC in LATERAL subquery for deterministic recency-biased sampling (rides the PK index, no extra sort) - Add timeout fallback (graph_expansion_timeout, default 10s) that drops entity expansion and falls back to semantic+causal only - Add composite index (entity_id, unit_id) on unit_entities for index-only scans in the LATERAL subquery - Merge 3 unmerged migration heads into one - Fix recall_perf.py dotenv override issue Unlike the approach in #895, this does NOT filter out hub entities entirely — all entities are kept but capped equally, preserving retrieval quality for queries about frequently-mentioned entities. Benchmarked on a 67K-unit bank (top entity = 25K mentions): - retrieval_graph: 0.337s → 0.055s (84% faster) - end-to-end recall: 0.912s → 0.519s (43% faster) * fix(tests): fix broken test_combined_scoring and test_reranking_proof_count - test_combined_scoring: replace MagicMock(spec=RetrievalResult) with real dataclass instances — MagicMock attributes returned nested mocks that failed on >= comparisons with int - test_reranking_proof_count: remove deleted `embedding` param from RetrievalResult constructor, use None for occurred_start/end to get neutral recency (datetime.now gave recency=1.0 which boosted scores) * refactor: rename config to link_expansion_ prefix, fix observation fanout - Rename GRAPH_PER_ENTITY_LIMIT → LINK_EXPANSION_PER_ENTITY_LIMIT and GRAPH_EXPANSION_TIMEOUT → LINK_EXPANSION_TIMEOUT to follow the convention that these are specific to the link_expansion graph retriever - Apply the same LATERAL per-entity cap to _expand_observations(), which had the same unbounded self-join through unit_entities * style: fix formatting in config.py |
||
|---|---|---|
| .. | ||
| common | ||
| consolidation | ||
| locomo | ||
| longmemeval | ||
| perf | ||
| visualizer | ||
| .DS_Store | ||
| __init__.py | ||
| README.md | ||
Hindsight Benchmarks
This directory contains benchmark suites for evaluating Hindsight's memory capabilities.
Prerequisites
-
Set up your environment variables in
.envat the project root:cp .env.example .env # Edit .env with your API keys -
Make sure you have
uvinstalled.
Available Benchmarks
LoComo
Tests conversational memory with multi-turn dialogues.
# Run from project root
./scripts/benchmarks/run-locomo.sh
# With options
./scripts/benchmarks/run-locomo.sh --max-conversations 10
./scripts/benchmarks/run-locomo.sh --skip-ingestion # Reuse existing data
./scripts/benchmarks/run-locomo.sh --use-think # Use think API
./scripts/benchmarks/run-locomo.sh --conversation conv-26 # Single conversation
Options:
--max-conversations N- Limit number of conversations--max-questions N- Limit questions per conversation--skip-ingestion- Skip data ingestion, use existing--use-think- Use think API instead of search + LLM--conversation NAME- Run specific conversation only--api-url URL- Custom API URL (default: local memory)--only-failed- Retry only failed questions--only-invalid- Retry only invalid questions
LongMemEval
Tests long-term memory across different categories.
# Run from project root
./scripts/benchmarks/run-longmemeval.sh
# With options
./scripts/benchmarks/run-longmemeval.sh --max-instances 50
./scripts/benchmarks/run-longmemeval.sh --category single-session-user
./scripts/benchmarks/run-longmemeval.sh --parallel 4 # Faster evaluation
Options:
--max-instances N- Limit total questions--max-instances-per-category N- Limit per category--skip-ingestion- Skip data ingestion--category NAME- Filter by category:single-session-usermulti-sessionsingle-session-preferencetemporal-reasoningknowledge-updatesingle-session-assistant
--parallel N- Parallel instances (default: 1)--only-failed- Retry failed questions--fill- Resume interrupted runs
Consolidation Performance
Tests consolidation throughput and identifies bottlenecks.
./scripts/benchmarks/run-consolidation.sh
# With custom memory count
NUM_MEMORIES=200 ./scripts/benchmarks/run-consolidation.sh
Retain Performance
Measures retain operation performance (throughput and token usage).
Prerequisites: API server must be running (./scripts/dev/start-api.sh)
# Basic usage
./scripts/benchmarks/run-retain-perf.sh \
--document hindsight-dev/benchmarks/perf/test_data/sample_document.txt
# Save results to JSON
./scripts/benchmarks/run-retain-perf.sh \
--document ./my_document.txt \
--bank-id my-test-bank \
--output results/retain_perf.json
Options:
--document PATH- Document file to retain (required)--bank-id ID- Bank ID to use (default: perf-test)--context TEXT- Optional context--api-url URL- API URL (default: http://localhost:8000)--timeout SECONDS- Request timeout (default: 300)--output PATH- Save results to JSON file
See perf/README.md for detailed documentation.
Visualizer
View benchmark results in a web UI:
./scripts/benchmarks/start-visualizer.sh
# Opens at http://localhost:8001
Results
Results are saved in JSON format in each benchmark's results/ directory.