* feat: add comprehensive OpenTelemetry tracing - Add tool execution spans for reflect operations - Add tool call information (names, params) to spans - Change verification scope from 'test' to 'verification' - Add hindsight.reflect_generation span for done() processing - Implement no-op tracer for improved code readability - Update documentation for OTEL configuration - Resolve merge conflicts from rebase * fix: properly serialize Pydantic models in span recording - Add _serialize_for_span() helper to handle Pydantic models - Update all providers to use the helper function - Fixes test failures with 'Object of type X is not JSON serializable' * feat: add Grafana LGTM stack for unified local observability Add Grafana LGTM (Loki, Grafana, Tempo, Mimir) as the recommended local development observability stack. This provides traces, metrics, and logs in a single Docker container instead of separate tools. Changes: - Add scripts/dev/grafana/ with docker-compose and README - Add scripts/dev/start-grafana.sh startup script - Update .env.example to reference Grafana LGTM - Update configuration docs to emphasize Grafana LGTM as primary option - Reorder OTLP backend list to show Grafana LGTM first Benefits: - Single container vs multiple separate tools (Jaeger, SigNoz, etc.) - ~515MB image with full observability stack - Compatible with existing OTLP configuration - Simpler local development setup * chore: remove SigNoz scripts and references Remove SigNoz observability stack in favor of Grafana LGTM as the sole recommended local development tracing solution. Changes: - Delete scripts/dev/signoz/ directory and all SigNoz configurations - Delete scripts/dev/start-signoz.sh startup script - Remove SigNoz references from .env.example - Remove SigNoz from OTLP backends list in configuration docs Grafana LGTM provides the same capabilities (traces, metrics, logs) in a simpler single-container setup. * feat: add consolidation span hierarchy for tracing Add parent-child span structure for consolidation operations: - hindsight.consolidation: Parent span for each memory being processed - hindsight.consolidation_recall: Child span for finding related observations - LLM call span: Automatically created by LLM provider (scope="consolidation") This enables detailed timing breakdown in Grafana Tempo: - Total consolidation time per memory - Time spent in recall - Time spent in LLM call - Time spent executing actions (create/update) All consolidation tests pass (31/31). * feat: add Prometheus metrics and GenAI dashboard to Grafana stack Add comprehensive metrics and dashboarding to the Grafana LGTM stack: Metrics Collection: - Configure Prometheus to scrape Hindsight API /metrics endpoint - Scrape interval: 10 seconds - Targets hindsight-api on host.docker.internal:8888 GenAI Dashboard: - Pre-configured dashboard with 6 panels: - LLM call rate (by provider/model) - LLM call duration (p50/p95 by scope) - Token usage - input tokens/sec by scope - Token usage - output tokens/sec by scope - Operations rate (retain/recall/reflect/consolidation) - Operation duration p95 by operation type Configuration: - Mount prometheus.yml for metrics scraping - Mount dashboards directory for auto-provisioning - Add host.docker.internal mapping for container->host access - Dashboard provisioning with auto-reload every 10s Documentation: - Updated README with metrics viewing instructions - Added PromQL query examples - Documented dashboard access and navigation This provides full observability: traces (Tempo) + metrics (Prometheus/Mimir) + dashboards (Grafana) * refactor: merge Grafana setup into existing monitoring stack Consolidate the separate scripts/dev/grafana/ setup into the existing scripts/dev/monitoring/ stack, using Grafana LGTM (Loki, Grafana, Tempo, Mimir). Changes: - Remove separate scripts/dev/grafana/ directory and start-grafana.sh - Rewrite scripts/dev/monitoring/start.sh to use Docker + Grafana LGTM (was: download native Prometheus/Grafana binaries) - Add docker-compose.yaml for Grafana LGTM container - Add prometheus.yml for scraping Hindsight API metrics - Mount existing dashboards from monitoring/grafana/dashboards/ - Add comprehensive README.md Benefits: - Single unified monitoring command: ./scripts/dev/start-monitoring.sh - Uses existing dashboard files (hindsight-operations, hindsight-llm, hindsight-api-service) - Simpler setup: Docker-based vs downloading/running native binaries - Full observability: traces + metrics + logs + dashboards in one container - Standard ports: Grafana on 3000, OTLP on 4317/4318 Architecture: - Grafana LGTM container (~515MB) provides all components - Dashboards auto-provisioned from monitoring/grafana/dashboards/ - Prometheus scrapes host.docker.internal:8888/metrics - Shared hindsight-network for future service-to-service tracing * fix: run monitoring stack in foreground for easy Ctrl+C stop Change docker-compose from detached (-d) to foreground mode. Users can now stop the stack with Ctrl+C instead of needing to run docker-compose down separately. * fix: remove invalid home dashboard path and obsolete version field - Remove GF_DASHBOARDS_DEFAULT_HOME_DASHBOARD_PATH environment variable (was pointing to wrong path causing 'Failed to load home dashboard' error) - Remove obsolete 'version' field from docker-compose.yaml (docker-compose v2+ doesn't require version field) * fix: load Hindsight dashboards in Grafana LGTM Mount Hindsight dashboard JSON files and custom provisioning config to make dashboards visible in Grafana. Changes: - Mount hindsight-operations.json, hindsight-llm.json, hindsight-api-service.json to /otel-lgtm/ - Create grafana-dashboards.yaml with all dashboard providers (default + Hindsight) - Mount custom provisioning config to override LGTM default All 3 Hindsight dashboards now appear in Grafana UI with metrics from Prometheus scraping the Hindsight API /metrics endpoint. * fix: configure Prometheus to scrape Hindsight API metrics Update prometheus.yml to include both OTLP receiver config (from LGTM) and scrape_configs for pulling metrics from Hindsight API. Changes: - Mount prometheus.yml to /otel-lgtm/prometheus.yaml (where LGTM reads it) - Add scrape_configs section to pull from host.docker.internal:8888/metrics - Keep OTLP receiver configuration for trace metrics - Set scrape_interval to 5s Verified: Prometheus now successfully scrapes hindsight_llm_calls_total and other Hindsight metrics. Dashboards now show live data! * feat: add comprehensive tracing for recall and improve reflect/mental_model_refresh spans - Add recall operation tracing with parent-child span hierarchy - Parent: hindsight.recall with attributes (bank_id, query, fact_types, etc.) - Children: recall_embedding, recall_retrieval, recall_fusion, recall_rerank - Fixed context propagation using start_as_current_span() - Improve reflect tracing spans - Remove reflect_generation spans, use reflect instead - Change done() tool processing to hindsight.reflect_tool_call - Fix mental_model_refresh span nesting - Add _skip_span parameter to reflect_async to avoid duplicate hindsight.reflect spans - Mental model refresh now has clean span hierarchy without nested reflect parent - Add comprehensive tracing verification tests - Test span hierarchy and attributes for all operations - Verify parent-child relationships - 5 passing tests covering recall, reflect, consolidation, and mental_model_refresh * refactor: remove redundant is_tracing_enabled() checks - Remove all is_tracing_enabled() conditional checks before tracing calls - NoOpTracer/NoOpSpan handle disabled tracing automatically - Simplify code by always calling tracer methods directly - Fix NoOpTracer.start_as_current_span() to yield NoOpSpan instead of None Changes: - memory_engine.py: Remove 5 is_tracing_enabled checks in recall spans - agent.py: Remove 2 is_tracing_enabled checks in reflect tool spans - tracing.py: Fix NoOpTracer context manager to yield proper NoOpSpan This eliminates ~50 lines of redundant conditional code while maintaining identical behavior. * docs: simplify distributed tracing section in monitoring.md - Make tracing documentation more concise - Focus on span hierarchy and attributes - Remove verbose troubleshooting and performance sections - Keep configuration.md for env vars only
10 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
Hindsight is an agent memory system that provides long-term memory for AI agents using biomimetic data structures. Memories are organized as:
- World facts: General knowledge ("The sky is blue")
- Experience facts: Personal experiences ("I visited Paris in 2023")
- Mental models: Consolidated knowledge synthesized from facts ("User prefers functional programming patterns")
Development Commands
API Server (Python/FastAPI)
# Start API server (loads .env automatically)
./scripts/dev/start-api.sh
# Run all tests (parallelized with pytest-xdist)
cd hindsight-api && uv run pytest tests/
# Run specific test file
cd hindsight-api && uv run pytest tests/test_http_api_integration.py -v
# Run single test function
cd hindsight-api && uv run pytest tests/test_retain.py::test_retain_simple -v
# Lint and format
cd hindsight-api && uv run ruff check .
cd hindsight-api && uv run ruff format .
# Type checking (uses ty - extremely fast type checker from Astral)
cd hindsight-api && uv run ty check hindsight_api/
Control Plane (Next.js)
./scripts/dev/start-control-plane.sh
# Or manually:
cd hindsight-control-plane && npm run dev
Documentation Site (Docusaurus)
./scripts/dev/start-docs.sh
Generating Clients/OpenAPI
# Regenerate OpenAPI spec after API changes (REQUIRED after changing endpoints)
./scripts/generate-openapi.sh
# Regenerate all client SDKs (Python, TypeScript, Rust)
./scripts/generate-clients.sh
Benchmarks
./scripts/benchmarks/run-longmemeval.sh
./scripts/benchmarks/run-locomo.sh
./scripts/benchmarks/start-visualizer.sh # View results at localhost:8001
Architecture
Monorepo Structure
- hindsight-api/: Core FastAPI server with memory engine (Python, uv)
- hindsight/: Embedded Python bundle (hindsight-all package)
- hindsight-control-plane/: Admin UI (Next.js, npm)
- hindsight-cli/: CLI tool (Rust, cargo, uses progenitor for API client)
- hindsight-clients/: Generated SDK clients (Python, TypeScript, Rust)
- hindsight-docs/: Docusaurus documentation site
- hindsight-integrations/: Framework integrations (LiteLLM, OpenAI)
- hindsight-dev/: Development tools and benchmarks
Core Engine (hindsight-api/hindsight_api/engine/)
memory_engine.py: Main orchestrator (~170KB) for retain/recall/reflect operationsllm_wrapper.py: LLM abstraction supporting OpenAI, Anthropic, Gemini, Groq, Ollama, LM Studioembeddings.py: Embedding generation (local sentence-transformers or TEI)cross_encoder.py: Reranking (local or TEI)entity_resolver.py: Entity extraction and normalizationquery_analyzer.py: Query intent analysis
retain/: Memory ingestion pipeline
orchestrator.py: Coordinates the retain flowfact_extraction.py: LLM-based fact extraction from contentlink_utils.py: Entity link creation and management
search/: Multi-strategy retrieval
retrieval.py: Main retrieval orchestratorgraph_retrieval.py: Entity/relationship graph traversalmpfp_retrieval.py: Multi-Path Fact Propagation retrievalfusion.py: Reciprocal rank fusion for combining resultsreranking.py: Cross-encoder reranking
API Layer (hindsight-api/hindsight_api/api/)
http.py: FastAPI HTTP routers (~80KB) for all REST endpointsmcp.py: Model Context Protocol server implementation
Main operations:
- Retain: Store memories, extracts facts/entities/relationships
- Recall: Retrieve memories via 4 parallel strategies (semantic, BM25, graph, temporal) + reranking
- Reflect: Disposition-aware reasoning using memories and mental models.
Database
PostgreSQL with pgvector. Schema managed via Alembic migrations in hindsight-api/hindsight_api/alembic/. Migrations run automatically on API startup.
Key tables: banks, memory_units, documents, entities, entity_links
Adding Database Migrations
-
Create a new migration file in
hindsight-api/hindsight_api/alembic/versions/:- File name format:
<revision_id>_<description>.py(e.g.,f1a2b3c4d5e6_add_new_index.py) - Use a unique hex revision ID (12 chars)
- Set
down_revisionto the previous migration's revision ID
- File name format:
-
Migration template:
"""Description of the migration Revision ID: f1a2b3c4d5e6 Revises: <previous_revision_id> Create Date: YYYY-MM-DD """ from collections.abc import Sequence from alembic import context, op revision: str = "f1a2b3c4d5e6" down_revision: str | Sequence[str] | None = "<previous_revision_id>" branch_labels: str | Sequence[str] | None = None depends_on: str | Sequence[str] | None = None def _get_schema_prefix() -> str: """Get schema prefix for table names (required for multi-tenant support).""" schema = context.config.get_main_option("target_schema") return f'"{schema}".' if schema else "" def upgrade() -> None: schema = _get_schema_prefix() op.execute(f"CREATE INDEX ... ON {schema}table_name(...)") def downgrade() -> None: schema = _get_schema_prefix() op.execute(f"DROP INDEX IF EXISTS {schema}index_name") -
Run migrations locally:
# Set database URL and run migrations uv run hindsight-admin run-db-migration # Run on a specific tenant schema uv run hindsight-admin run-db-migration --schema tenant_xyz
Key Conventions
Code Quality
Always run the lint script after making Python or TypeScript/Node changes:
./scripts/hooks/lint.sh
This runs the same checks as the pre-commit hook (Ruff for Python, ESLint/Prettier for TypeScript).
Memory Banks
- Each bank is an isolated memory store (like a "brain" for one user/agent)
- Banks have dispositions (skepticism, literalism, empathy traits 1-5) affecting reflect
- Banks can have background context
- Bank isolation is strict - no cross-bank data leakage
API Design
- All endpoints operate on a single bank per request
- Multi-bank queries are client responsibility to orchestrate
- Disposition traits only affect reflect, not recall
Control Plane API Routes
When adding or modifying parameters in the dataplane API (hindsight-api), you must also update the control plane routes that proxy to it:
-
API Routes (
hindsight-control-plane/src/app/api/):recall/route.ts- proxies to/v1/default/banks/{bank_id}/memories/recallreflect/route.ts- proxies to/v1/default/banks/{bank_id}/reflectmemories/retain/route.ts- proxies to/v1/default/banks/{bank_id}/memories/retain- Other routes follow the same pattern
-
Client types (
hindsight-control-plane/src/lib/api.ts):- Update the TypeScript type definitions for
recall(),reflect(),retain()etc.
- Update the TypeScript type definitions for
-
Checklist when adding new API parameters:
- Add parameter extraction in the route handler (destructure from
body) - Pass the parameter to the SDK call
- Update the client type definition in
lib/api.ts - Update any UI components that need to use the new parameter
- Add parameter extraction in the route handler (destructure from
Python Style
- Python 3.11+, type hints required
- Async throughout (asyncpg, async FastAPI)
- Pydantic models for request/response
- Ruff for linting (line-length 120)
- No Python files at project root - maintain clean directory structure
- Never use multi-item tuple return values - prefer dataclass or Pydantic model for structured returns
Type Safety with Pydantic Models
NEVER use raw dict types for structured data. Always use Pydantic models:
- Use Pydantic
BaseModelfor all data structures passed between functions - Add
@field_validatorfor type coercion (e.g., ensuring datetimes are timezone-aware) - Avoid
dict.get()patterns - use typed model attributes instead - Parse external data (JSON, API responses) into Pydantic models at the boundary
- This catches type errors at parse time, not deep in business logic
# BAD - error-prone dict access
def process(data: dict) -> str:
return data.get("name", "") # No validation, silent failures
# GOOD - typed and validated
class UserData(BaseModel):
name: str
created_at: datetime
@field_validator("created_at", mode="before")
@classmethod
def ensure_tz_aware(cls, v):
if isinstance(v, str):
v = datetime.fromisoformat(v.replace("Z", "+00:00"))
if v.tzinfo is None:
return v.replace(tzinfo=timezone.utc)
return v
def process(data: UserData) -> str:
return data.name # Type-safe, validated at construction
TypeScript Style
- Next.js App Router for control plane
- Tailwind CSS with shadcn/ui components
Adding New API Configuration Flags
When adding a new environment variable configuration:
-
config.py (
hindsight-api/hindsight_api/config.py):- Add
ENV_*constant for the environment variable name - Add
DEFAULT_*constant for the default value - Add field to
HindsightConfigdataclass - Add initialization in
from_env()method
- Add
-
main.py (
hindsight-api/hindsight_api/main.py):- Add field to the manual
HindsightConfig()constructor call (search for "CLI override")
- Add field to the manual
-
Use the config in code:
from ...config import get_config config = get_config() value = config.your_new_field -
Documentation (
hindsight-docs/docs/developer/configuration.md):- Add to appropriate section table with Variable, Description, Default
Environment Setup
cp .env.example .env
# Edit .env with LLM API key
# Python deps
uv sync --directory hindsight-api/
# Node deps (uses npm workspaces)
npm install
Required env vars:
HINDSIGHT_API_LLM_PROVIDER: openai, anthropic, gemini, groq, ollama, lmstudioHINDSIGHT_API_LLM_API_KEY: Your API keyHINDSIGHT_API_LLM_MODEL: Model name (e.g., o3-mini, claude-sonnet-4-20250514)
Optional (uses local models by default):
HINDSIGHT_API_EMBEDDINGS_PROVIDER: local (default) or teiHINDSIGHT_API_RERANKER_PROVIDER: local (default) or teiHINDSIGHT_API_DATABASE_URL: External PostgreSQL (uses embedded pg0 by default)