* Add Hindsight as git subtree + BCGU noise filtering tests Adds hindsight server source as a subtree under hindsight-api/ so we can iterate on server-side fixes directly. test_bcgu_noise_filtering.py proves that a well-crafted retain_custom_instructions (BCGU_RETAIN_MISSION) can suppress talking-head noise at fact extraction time — eliminating the need for client-side --filter-vision-noise preprocessing. Tests cover: - Default mode extracts 3 noise facts from talking-head frame (problem documented) - BCGU mission produces 0 noise facts from same talking-head frame - BCGU mission still extracts 2 high-value ChatGPT screen facts correctly - Mixed doc (2 talking-head + 2 screen): 0% noise ratio with BCGU mission - Pure talking-head doc: 0 facts extracted All 5 tests pass in ~32s using gpt-4o-mini. * fix(consolidation): respect mission context over ephemeral-state heuristic Two related fixes for the consolidation engine when a bank mission is configured: 1. **Mission override for ephemeral-state filter** (`prompts.py`): The system prompt previously instructed the LLM to discard any fact that looked like "ephemeral state" (e.g. current position, transient actions). When a mission is active the mission itself defines what is valuable — timestamped screen actions, session events, tool interactions may all be mission-critical even though they look ephemeral. Added a MISSION OVERRIDE block that explicitly tells the LLM the mission takes priority over the generic ephemeral-state guidance. 2. **Remove contradictory durable-knowledge nudge** (`consolidator.py`): The user-prompt builder was injecting "Focus on DURABLE knowledge that serves this mission, not ephemeral state" alongside the mission text. This phrasing contradicted missions that intentionally capture timestamped events. Replaced with a neutral directive that simply signals the mission overrides general rules. 3. **JSON control-character sanitisation** (`consolidator.py`): LLMs occasionally embed literal ASCII control characters (0x00–0x1f) inside JSON string values, causing `json.loads` to raise a JSONDecodeError. Added a try/except that strips control characters and retries the parse before re-raising, preventing spurious failures. * refactor(consolidation): move sanitize_llm_output to llm_wrapper, reuse in consolidator - Add `sanitize_llm_output()` to `llm_wrapper.py` as the single canonical function for stripping characters that break downstream systems (ASCII control chars 0x00-0x08/0x0B-0x0C/0x0E-0x1F/0x7F and Unicode surrogates). Tab, newline, and carriage-return are preserved. - Reduce `_sanitize_text()` in `fact_extraction.py` to a thin wrapper that delegates to `sanitize_llm_output()`. - Update `consolidator.py` to import and call `sanitize_llm_output()` directly instead of reimplementing the logic inline. - Remove test_bcgu_noise_filtering.py (should not have been committed). * fix(consolidation): apply sanitize_llm_output to observation text fields sanitize_llm_output was imported but unused after the old _call_llm_once path was removed. The batch flow uses structured Pydantic output so there's no raw json.loads call — instead, apply sanitization via field_validator on _CreateAction.text and _UpdateAction.text so control characters are stripped before observation text reaches the database. * fix(entity-resolver): correct mention_count for new entities in batch retain When the same entity (e.g. "Bob") appears across N items in a single batch retain, _resolve_entities_batch_impl deduplicates them into one name group before inserting, then queued only ONE _EntityStat regardless of N. The flush therefore always incremented mention_count by 1 beyond the INSERT value — giving 2 for any number of mentions. Two-part fix: - INSERT with mention_count=0 so the post-transaction flush is the single source of truth for the count (avoids an off-by-one for N=1 as well). - Append one _EntityStat per original mention (len(g.indices)) instead of one per unique name, so flush_pending_stats() adds the correct total N. This makes the batch path consistent with the single-entity path, which already accumulates one stat per mention via entities_to_update. |
||
|---|---|---|
| .. | ||
| hindsight_api | ||
| tests | ||
| pyproject.toml | ||
| README.md | ||
Hindsight API
Memory System for AI Agents — Temporal + Semantic + Entity Memory Architecture using PostgreSQL with pgvector.
Hindsight gives AI agents persistent memory that works like human memory: it stores facts, tracks entities and relationships, handles temporal reasoning ("what happened last spring?"), and forms opinions based on configurable disposition traits.
Installation
pip install hindsight-api
Quick Start
Run the Server
# Set your LLM provider
export HINDSIGHT_API_LLM_PROVIDER=openai
export HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx
# Start the server (uses embedded PostgreSQL by default)
hindsight-api
The server starts at http://localhost:8888 with:
- REST API for memory operations
- MCP server at
/mcpfor tool-use integration
Use the Python API
from hindsight_api import MemoryEngine
# Create and initialize the memory engine
memory = MemoryEngine()
await memory.initialize()
# Create a memory bank for your agent
bank = await memory.create_memory_bank(
name="my-assistant",
background="A helpful coding assistant"
)
# Store a memory
await memory.retain(
memory_bank_id=bank.id,
content="The user prefers Python for data science projects"
)
# Recall memories
results = await memory.recall(
memory_bank_id=bank.id,
query="What programming language does the user prefer?"
)
# Reflect with reasoning
response = await memory.reflect(
memory_bank_id=bank.id,
query="Should I recommend Python or R for this ML project?"
)
CLI Options
hindsight-api --help
# Common options
hindsight-api --port 9000 # Custom port (default: 8888)
hindsight-api --host 127.0.0.1 # Bind to localhost only
hindsight-api --workers 4 # Multiple worker processes
hindsight-api --log-level debug # Verbose logging
Configuration
Configure via environment variables:
| Variable | Description | Default |
|---|---|---|
HINDSIGHT_API_DATABASE_URL |
PostgreSQL connection string | pg0 (embedded) |
HINDSIGHT_API_LLM_PROVIDER |
openai, anthropic, gemini, groq, ollama, lmstudio |
openai |
HINDSIGHT_API_LLM_API_KEY |
API key for LLM provider | - |
HINDSIGHT_API_LLM_MODEL |
Model name | gpt-4o-mini |
HINDSIGHT_API_HOST |
Server bind address | 0.0.0.0 |
HINDSIGHT_API_PORT |
Server port | 8888 |
Example with External PostgreSQL
export HINDSIGHT_API_DATABASE_URL=postgresql://user:pass@localhost:5432/hindsight
export HINDSIGHT_API_LLM_PROVIDER=groq
export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx
hindsight-api
Docker
docker run --rm -it -p 8888:8888 \
-e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
-v $HOME/.hindsight-docker:/home/hindsight/.pg0 \
ghcr.io/vectorize-io/hindsight:latest
MCP Server
For local MCP integration without running the full API server:
hindsight-local-mcp
This runs a stdio-based MCP server that can be used directly with MCP-compatible clients.
Key Features
- Multi-Strategy Retrieval (TEMPR) — Semantic, keyword, graph, and temporal search combined with RRF fusion
- Entity Graph — Automatic entity extraction and relationship tracking
- Temporal Reasoning — Native support for time-based queries
- Disposition Traits — Configurable skepticism, literalism, and empathy influence opinion formation
- Three Memory Types — World facts, bank actions, and formed opinions with confidence scores
Documentation
Full documentation: https://hindsight.vectorize.io
License
Apache 2.0