* feat: allow chunks only in recall * feat: fetch chunks independently of max_tokens filtering Changes: - Chunks now fetched BEFORE max_tokens filtering (Step 5.5) - Implements batching: (max_chunk_tokens / retain_chunk_size) * 2 - Loop-based fetching until budget exhausted or no more chunks - Handles varying chunk sizes across documents - When max_tokens=0: returns 0 facts but still returns chunks - When max_tokens>0: backward compatible (chunks match filtered facts) Tests: - Added test_recall_chunks_independence.py with 5 comprehensive tests - Tests chunk independence, batching, ordering, and backward compat Docs: - Updated recall.mdx to explain new chunk behavior - Updated memory_engine.py docstrings Fixes chunk-related test failures by reordering chunks to match filtered facts when max_tokens > 0 (backward compatibility). * fix: fetch chunks after token filtering when max_tokens>0 Changes: - When max_tokens=0: fetch chunks BEFORE token filtering (new behavior) - When max_tokens>0: fetch chunks AFTER token filtering (backward compat) - This ensures chunk ordering matches filtered facts for max_tokens>0 - Fixes test failures in test_chunks_and_entities_follow_fact_order, test_chunk_fact_mapping, test_chunk_ordering_preservation, etc. The previous approach tried to reorder prefetched chunks, but that caused issues when the chunk budget was exhausted before all facts were processed. The new approach fetches chunks based on the correct fact set for each scenario. * fix: use ConfigResolver for bank-specific retain_chunk_size Fixes error: Field 'retain_chunk_size' is bank-configurable and cannot be accessed from global config. Changed from: - config.retain_chunk_size (global config, not allowed) To: - bank_config.retain_chunk_size (resolved from ConfigResolver) This ensures the correct chunk size is used for each bank, respecting any bank-specific overrides. * fix: correct Budget import in test_recall_chunks_independence Changed from: - from hindsight_api.engine.interface import Budget (incorrect) To: - from hindsight_api.engine.memory_engine import Budget (correct) This fixes the ImportError that was preventing the tests from running. * fix: prevent infinite loop in chunk fetching and improve test content - Add max(1, ...) to estimated_batch_size to prevent division resulting in 0 - Update test content to use more substantial examples that generate facts - Add request_context parameter to all retain_async and recall_async test calls * refactor: simplify chunk fetching to always use pre-filtering approach Remove backward compatibility code that fetched chunks after token filtering. Now chunks are always fetched from top-scored results before max_tokens filtering, regardless of max_tokens value. This simplifies the code by: - Removing duplicate chunk fetching logic - Eliminating conditional behavior based on max_tokens - Making chunk fetching behavior consistent and predictable Chunks are still fetched in batches and respect max_chunk_tokens limit. |
||
|---|---|---|
| .. | ||
| hindsight_api | ||
| tests | ||
| pyproject.toml | ||
| README.md | ||
Hindsight API
Memory System for AI Agents — Temporal + Semantic + Entity Memory Architecture using PostgreSQL with pgvector.
Hindsight gives AI agents persistent memory that works like human memory: it stores facts, tracks entities and relationships, handles temporal reasoning ("what happened last spring?"), and forms opinions based on configurable disposition traits.
Installation
pip install hindsight-api
Quick Start
Run the Server
# Set your LLM provider
export HINDSIGHT_API_LLM_PROVIDER=openai
export HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx
# Start the server (uses embedded PostgreSQL by default)
hindsight-api
The server starts at http://localhost:8888 with:
- REST API for memory operations
- MCP server at
/mcpfor tool-use integration
Use the Python API
from hindsight_api import MemoryEngine
# Create and initialize the memory engine
memory = MemoryEngine()
await memory.initialize()
# Create a memory bank for your agent
bank = await memory.create_memory_bank(
name="my-assistant",
background="A helpful coding assistant"
)
# Store a memory
await memory.retain(
memory_bank_id=bank.id,
content="The user prefers Python for data science projects"
)
# Recall memories
results = await memory.recall(
memory_bank_id=bank.id,
query="What programming language does the user prefer?"
)
# Reflect with reasoning
response = await memory.reflect(
memory_bank_id=bank.id,
query="Should I recommend Python or R for this ML project?"
)
CLI Options
hindsight-api --help
# Common options
hindsight-api --port 9000 # Custom port (default: 8888)
hindsight-api --host 127.0.0.1 # Bind to localhost only
hindsight-api --workers 4 # Multiple worker processes
hindsight-api --log-level debug # Verbose logging
Configuration
Configure via environment variables:
| Variable | Description | Default |
|---|---|---|
HINDSIGHT_API_DATABASE_URL |
PostgreSQL connection string | pg0 (embedded) |
HINDSIGHT_API_LLM_PROVIDER |
openai, anthropic, gemini, groq, ollama, lmstudio |
openai |
HINDSIGHT_API_LLM_API_KEY |
API key for LLM provider | - |
HINDSIGHT_API_LLM_MODEL |
Model name | gpt-4o-mini |
HINDSIGHT_API_HOST |
Server bind address | 0.0.0.0 |
HINDSIGHT_API_PORT |
Server port | 8888 |
Example with External PostgreSQL
export HINDSIGHT_API_DATABASE_URL=postgresql://user:pass@localhost:5432/hindsight
export HINDSIGHT_API_LLM_PROVIDER=groq
export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx
hindsight-api
Docker
docker run --rm -it -p 8888:8888 \
-e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
-v $HOME/.hindsight-docker:/home/hindsight/.pg0 \
ghcr.io/vectorize-io/hindsight:latest
MCP Server
For local MCP integration without running the full API server:
hindsight-local-mcp
This runs a stdio-based MCP server that can be used directly with MCP-compatible clients.
Key Features
- Multi-Strategy Retrieval (TEMPR) — Semantic, keyword, graph, and temporal search combined with RRF fusion
- Entity Graph — Automatic entity extraction and relationship tracking
- Temporal Reasoning — Native support for time-based queries
- Disposition Traits — Configurable skepticism, literalism, and empathy influence opinion formation
- Three Memory Types — World facts, bank actions, and formed opinions with confidence scores
Documentation
Full documentation: https://hindsight.vectorize.io
License
Apache 2.0