fleet-memory/CLAUDE.md
Nicolò Boschi 69dec8ec34
feat: add otel traceability (#330)
* feat: add comprehensive OpenTelemetry tracing

- Add tool execution spans for reflect operations
- Add tool call information (names, params) to spans
- Change verification scope from 'test' to 'verification'
- Add hindsight.reflect_generation span for done() processing
- Implement no-op tracer for improved code readability
- Update documentation for OTEL configuration
- Resolve merge conflicts from rebase

* fix: properly serialize Pydantic models in span recording

- Add _serialize_for_span() helper to handle Pydantic models
- Update all providers to use the helper function
- Fixes test failures with 'Object of type X is not JSON serializable'

* feat: add Grafana LGTM stack for unified local observability

Add Grafana LGTM (Loki, Grafana, Tempo, Mimir) as the recommended
local development observability stack. This provides traces, metrics,
and logs in a single Docker container instead of separate tools.

Changes:
- Add scripts/dev/grafana/ with docker-compose and README
- Add scripts/dev/start-grafana.sh startup script
- Update .env.example to reference Grafana LGTM
- Update configuration docs to emphasize Grafana LGTM as primary option
- Reorder OTLP backend list to show Grafana LGTM first

Benefits:
- Single container vs multiple separate tools (Jaeger, SigNoz, etc.)
- ~515MB image with full observability stack
- Compatible with existing OTLP configuration
- Simpler local development setup

* chore: remove SigNoz scripts and references

Remove SigNoz observability stack in favor of Grafana LGTM as the
sole recommended local development tracing solution.

Changes:
- Delete scripts/dev/signoz/ directory and all SigNoz configurations
- Delete scripts/dev/start-signoz.sh startup script
- Remove SigNoz references from .env.example
- Remove SigNoz from OTLP backends list in configuration docs

Grafana LGTM provides the same capabilities (traces, metrics, logs)
in a simpler single-container setup.

* feat: add consolidation span hierarchy for tracing

Add parent-child span structure for consolidation operations:
- hindsight.consolidation: Parent span for each memory being processed
- hindsight.consolidation_recall: Child span for finding related observations
- LLM call span: Automatically created by LLM provider (scope="consolidation")

This enables detailed timing breakdown in Grafana Tempo:
- Total consolidation time per memory
- Time spent in recall
- Time spent in LLM call
- Time spent executing actions (create/update)

All consolidation tests pass (31/31).

* feat: add Prometheus metrics and GenAI dashboard to Grafana stack

Add comprehensive metrics and dashboarding to the Grafana LGTM stack:

Metrics Collection:
- Configure Prometheus to scrape Hindsight API /metrics endpoint
- Scrape interval: 10 seconds
- Targets hindsight-api on host.docker.internal:8888

GenAI Dashboard:
- Pre-configured dashboard with 6 panels:
  - LLM call rate (by provider/model)
  - LLM call duration (p50/p95 by scope)
  - Token usage - input tokens/sec by scope
  - Token usage - output tokens/sec by scope
  - Operations rate (retain/recall/reflect/consolidation)
  - Operation duration p95 by operation type

Configuration:
- Mount prometheus.yml for metrics scraping
- Mount dashboards directory for auto-provisioning
- Add host.docker.internal mapping for container->host access
- Dashboard provisioning with auto-reload every 10s

Documentation:
- Updated README with metrics viewing instructions
- Added PromQL query examples
- Documented dashboard access and navigation

This provides full observability: traces (Tempo) + metrics (Prometheus/Mimir) + dashboards (Grafana)

* refactor: merge Grafana setup into existing monitoring stack

Consolidate the separate scripts/dev/grafana/ setup into the existing
scripts/dev/monitoring/ stack, using Grafana LGTM (Loki, Grafana, Tempo, Mimir).

Changes:
- Remove separate scripts/dev/grafana/ directory and start-grafana.sh
- Rewrite scripts/dev/monitoring/start.sh to use Docker + Grafana LGTM
  (was: download native Prometheus/Grafana binaries)
- Add docker-compose.yaml for Grafana LGTM container
- Add prometheus.yml for scraping Hindsight API metrics
- Mount existing dashboards from monitoring/grafana/dashboards/
- Add comprehensive README.md

Benefits:
- Single unified monitoring command: ./scripts/dev/start-monitoring.sh
- Uses existing dashboard files (hindsight-operations, hindsight-llm, hindsight-api-service)
- Simpler setup: Docker-based vs downloading/running native binaries
- Full observability: traces + metrics + logs + dashboards in one container
- Standard ports: Grafana on 3000, OTLP on 4317/4318

Architecture:
- Grafana LGTM container (~515MB) provides all components
- Dashboards auto-provisioned from monitoring/grafana/dashboards/
- Prometheus scrapes host.docker.internal:8888/metrics
- Shared hindsight-network for future service-to-service tracing

* fix: run monitoring stack in foreground for easy Ctrl+C stop

Change docker-compose from detached (-d) to foreground mode.
Users can now stop the stack with Ctrl+C instead of needing
to run docker-compose down separately.

* fix: remove invalid home dashboard path and obsolete version field

- Remove GF_DASHBOARDS_DEFAULT_HOME_DASHBOARD_PATH environment variable
  (was pointing to wrong path causing 'Failed to load home dashboard' error)
- Remove obsolete 'version' field from docker-compose.yaml
  (docker-compose v2+ doesn't require version field)

* fix: load Hindsight dashboards in Grafana LGTM

Mount Hindsight dashboard JSON files and custom provisioning config
to make dashboards visible in Grafana.

Changes:
- Mount hindsight-operations.json, hindsight-llm.json, hindsight-api-service.json to /otel-lgtm/
- Create grafana-dashboards.yaml with all dashboard providers (default + Hindsight)
- Mount custom provisioning config to override LGTM default

All 3 Hindsight dashboards now appear in Grafana UI with metrics
from Prometheus scraping the Hindsight API /metrics endpoint.

* fix: configure Prometheus to scrape Hindsight API metrics

Update prometheus.yml to include both OTLP receiver config (from LGTM)
and scrape_configs for pulling metrics from Hindsight API.

Changes:
- Mount prometheus.yml to /otel-lgtm/prometheus.yaml (where LGTM reads it)
- Add scrape_configs section to pull from host.docker.internal:8888/metrics
- Keep OTLP receiver configuration for trace metrics
- Set scrape_interval to 5s

Verified: Prometheus now successfully scrapes hindsight_llm_calls_total
and other Hindsight metrics. Dashboards now show live data!

* feat: add comprehensive tracing for recall and improve reflect/mental_model_refresh spans

- Add recall operation tracing with parent-child span hierarchy
  - Parent: hindsight.recall with attributes (bank_id, query, fact_types, etc.)
  - Children: recall_embedding, recall_retrieval, recall_fusion, recall_rerank
  - Fixed context propagation using start_as_current_span()

- Improve reflect tracing spans
  - Remove reflect_generation spans, use reflect instead
  - Change done() tool processing to hindsight.reflect_tool_call

- Fix mental_model_refresh span nesting
  - Add _skip_span parameter to reflect_async to avoid duplicate hindsight.reflect spans
  - Mental model refresh now has clean span hierarchy without nested reflect parent

- Add comprehensive tracing verification tests
  - Test span hierarchy and attributes for all operations
  - Verify parent-child relationships
  - 5 passing tests covering recall, reflect, consolidation, and mental_model_refresh

* refactor: remove redundant is_tracing_enabled() checks

- Remove all is_tracing_enabled() conditional checks before tracing calls
- NoOpTracer/NoOpSpan handle disabled tracing automatically
- Simplify code by always calling tracer methods directly
- Fix NoOpTracer.start_as_current_span() to yield NoOpSpan instead of None

Changes:
- memory_engine.py: Remove 5 is_tracing_enabled checks in recall spans
- agent.py: Remove 2 is_tracing_enabled checks in reflect tool spans
- tracing.py: Fix NoOpTracer context manager to yield proper NoOpSpan

This eliminates ~50 lines of redundant conditional code while maintaining
identical behavior.

* docs: simplify distributed tracing section in monitoring.md

- Make tracing documentation more concise
- Focus on span hierarchy and attributes
- Remove verbose troubleshooting and performance sections
- Keep configuration.md for env vars only
2026-02-10 12:20:48 +01:00

10 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project Overview

Hindsight is an agent memory system that provides long-term memory for AI agents using biomimetic data structures. Memories are organized as:

  • World facts: General knowledge ("The sky is blue")
  • Experience facts: Personal experiences ("I visited Paris in 2023")
  • Mental models: Consolidated knowledge synthesized from facts ("User prefers functional programming patterns")

Development Commands

API Server (Python/FastAPI)

# Start API server (loads .env automatically)
./scripts/dev/start-api.sh

# Run all tests (parallelized with pytest-xdist)
cd hindsight-api && uv run pytest tests/

# Run specific test file
cd hindsight-api && uv run pytest tests/test_http_api_integration.py -v

# Run single test function
cd hindsight-api && uv run pytest tests/test_retain.py::test_retain_simple -v

# Lint and format
cd hindsight-api && uv run ruff check .
cd hindsight-api && uv run ruff format .

# Type checking (uses ty - extremely fast type checker from Astral)
cd hindsight-api && uv run ty check hindsight_api/

Control Plane (Next.js)

./scripts/dev/start-control-plane.sh
# Or manually:
cd hindsight-control-plane && npm run dev

Documentation Site (Docusaurus)

./scripts/dev/start-docs.sh

Generating Clients/OpenAPI

# Regenerate OpenAPI spec after API changes (REQUIRED after changing endpoints)
./scripts/generate-openapi.sh

# Regenerate all client SDKs (Python, TypeScript, Rust)
./scripts/generate-clients.sh

Benchmarks

./scripts/benchmarks/run-longmemeval.sh
./scripts/benchmarks/run-locomo.sh
./scripts/benchmarks/start-visualizer.sh  # View results at localhost:8001

Architecture

Monorepo Structure

  • hindsight-api/: Core FastAPI server with memory engine (Python, uv)
  • hindsight/: Embedded Python bundle (hindsight-all package)
  • hindsight-control-plane/: Admin UI (Next.js, npm)
  • hindsight-cli/: CLI tool (Rust, cargo, uses progenitor for API client)
  • hindsight-clients/: Generated SDK clients (Python, TypeScript, Rust)
  • hindsight-docs/: Docusaurus documentation site
  • hindsight-integrations/: Framework integrations (LiteLLM, OpenAI)
  • hindsight-dev/: Development tools and benchmarks

Core Engine (hindsight-api/hindsight_api/engine/)

  • memory_engine.py: Main orchestrator (~170KB) for retain/recall/reflect operations
  • llm_wrapper.py: LLM abstraction supporting OpenAI, Anthropic, Gemini, Groq, Ollama, LM Studio
  • embeddings.py: Embedding generation (local sentence-transformers or TEI)
  • cross_encoder.py: Reranking (local or TEI)
  • entity_resolver.py: Entity extraction and normalization
  • query_analyzer.py: Query intent analysis

retain/: Memory ingestion pipeline

  • orchestrator.py: Coordinates the retain flow
  • fact_extraction.py: LLM-based fact extraction from content
  • link_utils.py: Entity link creation and management

search/: Multi-strategy retrieval

  • retrieval.py: Main retrieval orchestrator
  • graph_retrieval.py: Entity/relationship graph traversal
  • mpfp_retrieval.py: Multi-Path Fact Propagation retrieval
  • fusion.py: Reciprocal rank fusion for combining results
  • reranking.py: Cross-encoder reranking

API Layer (hindsight-api/hindsight_api/api/)

  • http.py: FastAPI HTTP routers (~80KB) for all REST endpoints
  • mcp.py: Model Context Protocol server implementation

Main operations:

  • Retain: Store memories, extracts facts/entities/relationships
  • Recall: Retrieve memories via 4 parallel strategies (semantic, BM25, graph, temporal) + reranking
  • Reflect: Disposition-aware reasoning using memories and mental models.

Database

PostgreSQL with pgvector. Schema managed via Alembic migrations in hindsight-api/hindsight_api/alembic/. Migrations run automatically on API startup.

Key tables: banks, memory_units, documents, entities, entity_links

Adding Database Migrations

  1. Create a new migration file in hindsight-api/hindsight_api/alembic/versions/:

    • File name format: <revision_id>_<description>.py (e.g., f1a2b3c4d5e6_add_new_index.py)
    • Use a unique hex revision ID (12 chars)
    • Set down_revision to the previous migration's revision ID
  2. Migration template:

    """Description of the migration
    
    Revision ID: f1a2b3c4d5e6
    Revises: <previous_revision_id>
    Create Date: YYYY-MM-DD
    """
    from collections.abc import Sequence
    from alembic import context, op
    
    revision: str = "f1a2b3c4d5e6"
    down_revision: str | Sequence[str] | None = "<previous_revision_id>"
    branch_labels: str | Sequence[str] | None = None
    depends_on: str | Sequence[str] | None = None
    
    def _get_schema_prefix() -> str:
        """Get schema prefix for table names (required for multi-tenant support)."""
        schema = context.config.get_main_option("target_schema")
        return f'"{schema}".' if schema else ""
    
    def upgrade() -> None:
        schema = _get_schema_prefix()
        op.execute(f"CREATE INDEX ... ON {schema}table_name(...)")
    
    def downgrade() -> None:
        schema = _get_schema_prefix()
        op.execute(f"DROP INDEX IF EXISTS {schema}index_name")
    
  3. Run migrations locally:

    # Set database URL and run migrations
    uv run hindsight-admin run-db-migration
    
    # Run on a specific tenant schema
    uv run hindsight-admin run-db-migration --schema tenant_xyz
    

Key Conventions

Code Quality

Always run the lint script after making Python or TypeScript/Node changes:

./scripts/hooks/lint.sh

This runs the same checks as the pre-commit hook (Ruff for Python, ESLint/Prettier for TypeScript).

Memory Banks

  • Each bank is an isolated memory store (like a "brain" for one user/agent)
  • Banks have dispositions (skepticism, literalism, empathy traits 1-5) affecting reflect
  • Banks can have background context
  • Bank isolation is strict - no cross-bank data leakage

API Design

  • All endpoints operate on a single bank per request
  • Multi-bank queries are client responsibility to orchestrate
  • Disposition traits only affect reflect, not recall

Control Plane API Routes

When adding or modifying parameters in the dataplane API (hindsight-api), you must also update the control plane routes that proxy to it:

  1. API Routes (hindsight-control-plane/src/app/api/):

    • recall/route.ts - proxies to /v1/default/banks/{bank_id}/memories/recall
    • reflect/route.ts - proxies to /v1/default/banks/{bank_id}/reflect
    • memories/retain/route.ts - proxies to /v1/default/banks/{bank_id}/memories/retain
    • Other routes follow the same pattern
  2. Client types (hindsight-control-plane/src/lib/api.ts):

    • Update the TypeScript type definitions for recall(), reflect(), retain() etc.
  3. Checklist when adding new API parameters:

    • Add parameter extraction in the route handler (destructure from body)
    • Pass the parameter to the SDK call
    • Update the client type definition in lib/api.ts
    • Update any UI components that need to use the new parameter

Python Style

  • Python 3.11+, type hints required
  • Async throughout (asyncpg, async FastAPI)
  • Pydantic models for request/response
  • Ruff for linting (line-length 120)
  • No Python files at project root - maintain clean directory structure
  • Never use multi-item tuple return values - prefer dataclass or Pydantic model for structured returns

Type Safety with Pydantic Models

NEVER use raw dict types for structured data. Always use Pydantic models:

  • Use Pydantic BaseModel for all data structures passed between functions
  • Add @field_validator for type coercion (e.g., ensuring datetimes are timezone-aware)
  • Avoid dict.get() patterns - use typed model attributes instead
  • Parse external data (JSON, API responses) into Pydantic models at the boundary
  • This catches type errors at parse time, not deep in business logic
# BAD - error-prone dict access
def process(data: dict) -> str:
    return data.get("name", "")  # No validation, silent failures

# GOOD - typed and validated
class UserData(BaseModel):
    name: str
    created_at: datetime

    @field_validator("created_at", mode="before")
    @classmethod
    def ensure_tz_aware(cls, v):
        if isinstance(v, str):
            v = datetime.fromisoformat(v.replace("Z", "+00:00"))
        if v.tzinfo is None:
            return v.replace(tzinfo=timezone.utc)
        return v

def process(data: UserData) -> str:
    return data.name  # Type-safe, validated at construction

TypeScript Style

  • Next.js App Router for control plane
  • Tailwind CSS with shadcn/ui components

Adding New API Configuration Flags

When adding a new environment variable configuration:

  1. config.py (hindsight-api/hindsight_api/config.py):

    • Add ENV_* constant for the environment variable name
    • Add DEFAULT_* constant for the default value
    • Add field to HindsightConfig dataclass
    • Add initialization in from_env() method
  2. main.py (hindsight-api/hindsight_api/main.py):

    • Add field to the manual HindsightConfig() constructor call (search for "CLI override")
  3. Use the config in code:

    from ...config import get_config
    config = get_config()
    value = config.your_new_field
    
  4. Documentation (hindsight-docs/docs/developer/configuration.md):

    • Add to appropriate section table with Variable, Description, Default

Environment Setup

cp .env.example .env
# Edit .env with LLM API key

# Python deps
uv sync --directory hindsight-api/

# Node deps (uses npm workspaces)
npm install

Required env vars:

  • HINDSIGHT_API_LLM_PROVIDER: openai, anthropic, gemini, groq, ollama, lmstudio
  • HINDSIGHT_API_LLM_API_KEY: Your API key
  • HINDSIGHT_API_LLM_MODEL: Model name (e.g., o3-mini, claude-sonnet-4-20250514)

Optional (uses local models by default):

  • HINDSIGHT_API_EMBEDDINGS_PROVIDER: local (default) or tei
  • HINDSIGHT_API_RERANKER_PROVIDER: local (default) or tei
  • HINDSIGHT_API_DATABASE_URL: External PostgreSQL (uses embedded pg0 by default)