fleet-memory/CLAUDE.md
Nicolò Boschi 8d731f2e5f
feat: implement hierarchical configuration (system, tenant, bank) (#329)
* feat: implement hierarchical configuration (system, tenant, bank)

* feat: implement hierarchical configuration (system, tenant, bank)

* docs: add instructions for hierarchical config in CLAUDE.md

* feat: add ENABLE_BANK_CONFIG_API flag (disabled by default)

- Add HINDSIGHT_API_ENABLE_BANK_CONFIG_API env var (default: false)
- Return 403 Forbidden from bank config endpoints when disabled
- Update tests to enable the flag
- Update CLAUDE.md documentation

This provides security control over the bank configuration API,
ensuring it's only accessible when explicitly enabled.

* docs: add hierarchical configuration section

* feat(cli): add bank config commands (config, set-config, reset-config)

- Add 'hindsight bank config' to view bank configuration
- Add 'hindsight bank set-config' to update LLM settings per bank
- Add 'hindsight bank reset-config' to reset to defaults
- Implements client API calls to new bank config endpoints

* fix(cli): fix compilation errors in bank config commands

- Fix type signature: use ApiClient instead of api::Client
- Fix confirmation: use ui::prompt_confirmation instead of ui::confirm
- Fix error handling: use anyhow! macro instead of errors::Error
- Fix type conversion: convert HashMap to serde_json::Map for API call

* feat: implement type-safe hierarchical config with bank overrides

Implements a production-ready hierarchical configuration system that prevents
accidentally using global defaults when bank-specific overrides exist.

- Created StaticConfigProxy that wraps HindsightConfig
- get_config() now returns proxy that blocks access to bank-configurable fields
- Raises ConfigFieldAccessError with clear message when accessing configurable fields
- Added _get_raw_config() for internal use only
- Forces developers to use resolve_full_config(bank_id, context) for bank settings

- Added resolve_full_config() method that returns complete HindsightConfig
- Resolves hierarchy: Global (env) → Tenant → Bank
- No caching to support multi-server deployments (always fresh from DB)
- LLM provider pooling handles expensive operations separately

- Updated entire retain pipeline to pass resolved config through call chain
- memory_engine.py: Resolves config at top level where bank_id/context available
- orchestrator.py: Accepts and passes config to fact_extraction
- fact_extraction.py: Uses passed config instead of get_config()
- utils.py: Added optional config param for backward compatibility

- consolidator.py: Uses resolve_full_config() for enable_observations check
- memory_engine.py: Resolves config before triggering consolidation

- Renamed "Memory Bank" to "Bank Configuration" with tabs
- Combined Stats and Operations into "General" tab
- Consolidated Profile and Configuration into "Configuration" tab
- Moved Actions dropdown to page level (outside tabs)

- Created new component for managing bank-specific config
- Displays configurable fields: retain_chunk_size, retain_extraction_mode, etc.
- Edit via dialog with form validation
- Reset to defaults via AlertDialog confirmation
- Shows field IDs in monospace for clarity
- Visual separation with borders and hover effects

- Removed inline edit mode, switched to dialog-based editing
- Separate dialogs for Disposition and Mission editing
- Read-only display with clear edit buttons
- Removed duplicate stats cards and operations

- bank-stats-view.tsx: Overview statistics (memories, links, documents, pending ops)
- bank-operations-view.tsx: Background operations table with filtering

**Problem**: Consolidation always used global enable_observations, ignoring bank overrides
**Root Cause**: consolidator.py called get_config() instead of resolving bank-specific config
**Solution**: Pass resolved config through the entire pipeline

**Problem**: asyncpg returning JSONB as JSON string instead of parsed dict
**Solution**: Explicit JSON parsing in config_resolver.py with type checking

- All 19 API integration tests pass
- All 10 hierarchical config tests pass
- Retain operations work correctly with bank-specific config
- Consolidation respects bank-specific enable_observations setting

- Updated developer/configuration.md with type-safe config access pattern
- Added examples showing correct usage patterns
- Documented ConfigFieldAccessError and resolution methods

- get_config() now returns StaticConfigProxy (blocks configurable field access)
- Code accessing bank-configurable fields must use resolve_full_config()
- Clear migration path with helpful error messages

Fixes hierarchical configuration to be production-ready with proper type safety.

* refactor: remove LLM client pool and simplify config resolver

Since LLM config (provider, model, api_key) is now static and not
bank-configurable, the LLMClientPool is no longer needed.

Changes:
- Remove hindsight_api/llm_client_pool.py (no longer needed)
- Remove memory_engine._get_bank_llm_config() (dead code, never called)
- Simplify config_resolver.py by eliminating duplication between
  resolve_full_config() and get_bank_config()
- get_bank_config() now calls resolve_full_config() and filters results
- Remove outdated "LLM provider pooling" comments from docstrings

All tests pass (10 hierarchical config tests, 19 API integration tests)

* fix: update tests to use _get_raw_config() for configurable fields

Fixed test fixtures that were accessing configurable fields (like
enable_observations) from get_config(), which now raises
ConfigFieldAccessError due to type-safe config access.

Changes:
- test_consolidation.py: Changed enable_observations fixture to use
  _get_raw_config() instead of get_config()
- test_consolidation.py: Updated test_consolidation_returns_disabled_status
  to set bank config instead of mocking get_config()
- test_link_expansion_retrieval.py: Changed fixture to use _get_raw_config()
- test_observations.py: Changed disable_observations fixture to use
  _get_raw_config()
- Regenerated OpenAPI spec and clients

All 39 previously failing tests now pass.

* fix: add missing config parameter to test calls of extract_facts_from_text()

Fixed 45 test failures where tests were calling extract_facts_from_text()
without the new required config parameter.

Changes:
- Added config=_get_raw_config() to all extract_facts_from_text() calls
- Fixed test_main_module.py to patch _get_raw_config instead of get_config
- Updated 6 test files with 37 function call sites

All tests should now pass.

* fix: add missing config parameter to test_skip_podcast_meta_commentary

One more test was missing the config parameter for extract_facts_from_text().
2026-02-12 13:14:57 +01:00

12 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project Overview

Hindsight is an agent memory system that provides long-term memory for AI agents using biomimetic data structures. Memories are organized as:

  • World facts: General knowledge ("The sky is blue")
  • Experience facts: Personal experiences ("I visited Paris in 2023")
  • Mental models: Consolidated knowledge synthesized from facts ("User prefers functional programming patterns")

Development Commands

API Server (Python/FastAPI)

# Start API server (loads .env automatically)
./scripts/dev/start-api.sh

# Run all tests (parallelized with pytest-xdist)
cd hindsight-api && uv run pytest tests/

# Run specific test file
cd hindsight-api && uv run pytest tests/test_http_api_integration.py -v

# Run single test function
cd hindsight-api && uv run pytest tests/test_retain.py::test_retain_simple -v

# Lint and format
cd hindsight-api && uv run ruff check .
cd hindsight-api && uv run ruff format .

# Type checking (uses ty - extremely fast type checker from Astral)
cd hindsight-api && uv run ty check hindsight_api/

Control Plane (Next.js)

./scripts/dev/start-control-plane.sh
# Or manually:
cd hindsight-control-plane && npm run dev

Documentation Site (Docusaurus)

./scripts/dev/start-docs.sh

Generating Clients/OpenAPI

# Regenerate OpenAPI spec after API changes (REQUIRED after changing endpoints)
./scripts/generate-openapi.sh

# Regenerate all client SDKs (Python, TypeScript, Rust)
./scripts/generate-clients.sh

Benchmarks

./scripts/benchmarks/run-longmemeval.sh
./scripts/benchmarks/run-locomo.sh
./scripts/benchmarks/start-visualizer.sh  # View results at localhost:8001

Architecture

Monorepo Structure

  • hindsight-api/: Core FastAPI server with memory engine (Python, uv)
  • hindsight/: Embedded Python bundle (hindsight-all package)
  • hindsight-control-plane/: Admin UI (Next.js, npm)
  • hindsight-cli/: CLI tool (Rust, cargo, uses progenitor for API client)
  • hindsight-clients/: Generated SDK clients (Python, TypeScript, Rust)
  • hindsight-docs/: Docusaurus documentation site
  • hindsight-integrations/: Framework integrations (LiteLLM, OpenAI)
  • hindsight-dev/: Development tools and benchmarks

Core Engine (hindsight-api/hindsight_api/engine/)

  • memory_engine.py: Main orchestrator (~170KB) for retain/recall/reflect operations
  • llm_wrapper.py: LLM abstraction supporting OpenAI, Anthropic, Gemini, Groq, Ollama, LM Studio
  • embeddings.py: Embedding generation (local sentence-transformers or TEI)
  • cross_encoder.py: Reranking (local or TEI)
  • entity_resolver.py: Entity extraction and normalization
  • query_analyzer.py: Query intent analysis

retain/: Memory ingestion pipeline

  • orchestrator.py: Coordinates the retain flow
  • fact_extraction.py: LLM-based fact extraction from content
  • link_utils.py: Entity link creation and management

search/: Multi-strategy retrieval

  • retrieval.py: Main retrieval orchestrator
  • graph_retrieval.py: Entity/relationship graph traversal
  • mpfp_retrieval.py: Multi-Path Fact Propagation retrieval
  • fusion.py: Reciprocal rank fusion for combining results
  • reranking.py: Cross-encoder reranking

API Layer (hindsight-api/hindsight_api/api/)

  • http.py: FastAPI HTTP routers (~80KB) for all REST endpoints
  • mcp.py: Model Context Protocol server implementation

Main operations:

  • Retain: Store memories, extracts facts/entities/relationships
  • Recall: Retrieve memories via 4 parallel strategies (semantic, BM25, graph, temporal) + reranking
  • Reflect: Disposition-aware reasoning using memories and mental models.

Database

PostgreSQL with pgvector. Schema managed via Alembic migrations in hindsight-api/hindsight_api/alembic/. Migrations run automatically on API startup.

Key tables: banks, memory_units, documents, entities, entity_links

Adding Database Migrations

  1. Create a new migration file in hindsight-api/hindsight_api/alembic/versions/:

    • File name format: <revision_id>_<description>.py (e.g., f1a2b3c4d5e6_add_new_index.py)
    • Use a unique hex revision ID (12 chars)
    • Set down_revision to the previous migration's revision ID
  2. Migration template:

    """Description of the migration
    
    Revision ID: f1a2b3c4d5e6
    Revises: <previous_revision_id>
    Create Date: YYYY-MM-DD
    """
    from collections.abc import Sequence
    from alembic import context, op
    
    revision: str = "f1a2b3c4d5e6"
    down_revision: str | Sequence[str] | None = "<previous_revision_id>"
    branch_labels: str | Sequence[str] | None = None
    depends_on: str | Sequence[str] | None = None
    
    def _get_schema_prefix() -> str:
        """Get schema prefix for table names (required for multi-tenant support)."""
        schema = context.config.get_main_option("target_schema")
        return f'"{schema}".' if schema else ""
    
    def upgrade() -> None:
        schema = _get_schema_prefix()
        op.execute(f"CREATE INDEX ... ON {schema}table_name(...)")
    
    def downgrade() -> None:
        schema = _get_schema_prefix()
        op.execute(f"DROP INDEX IF EXISTS {schema}index_name")
    
  3. Run migrations locally:

    # Set database URL and run migrations
    uv run hindsight-admin run-db-migration
    
    # Run on a specific tenant schema
    uv run hindsight-admin run-db-migration --schema tenant_xyz
    

Key Conventions

Code Quality

Always run the lint script after making Python or TypeScript/Node changes:

./scripts/hooks/lint.sh

This runs the same checks as the pre-commit hook (Ruff for Python, ESLint/Prettier for TypeScript).

Memory Banks

  • Each bank is an isolated memory store (like a "brain" for one user/agent)
  • Banks have dispositions (skepticism, literalism, empathy traits 1-5) affecting reflect
  • Banks can have background context
  • Bank isolation is strict - no cross-bank data leakage

API Design

  • All endpoints operate on a single bank per request
  • Multi-bank queries are client responsibility to orchestrate
  • Disposition traits only affect reflect, not recall

Control Plane API Routes

When adding or modifying parameters in the dataplane API (hindsight-api), you must also update the control plane routes that proxy to it:

  1. API Routes (hindsight-control-plane/src/app/api/):

    • recall/route.ts - proxies to /v1/default/banks/{bank_id}/memories/recall
    • reflect/route.ts - proxies to /v1/default/banks/{bank_id}/reflect
    • memories/retain/route.ts - proxies to /v1/default/banks/{bank_id}/memories/retain
    • Other routes follow the same pattern
  2. Client types (hindsight-control-plane/src/lib/api.ts):

    • Update the TypeScript type definitions for recall(), reflect(), retain() etc.
  3. Checklist when adding new API parameters:

    • Add parameter extraction in the route handler (destructure from body)
    • Pass the parameter to the SDK call
    • Update the client type definition in lib/api.ts
    • Update any UI components that need to use the new parameter

Python Style

  • Python 3.11+, type hints required
  • Async throughout (asyncpg, async FastAPI)
  • Pydantic models for request/response
  • Ruff for linting (line-length 120)
  • No Python files at project root - maintain clean directory structure
  • Never use multi-item tuple return values - prefer dataclass or Pydantic model for structured returns

Type Safety with Pydantic Models

NEVER use raw dict types for structured data. Always use Pydantic models:

  • Use Pydantic BaseModel for all data structures passed between functions
  • Add @field_validator for type coercion (e.g., ensuring datetimes are timezone-aware)
  • Avoid dict.get() patterns - use typed model attributes instead
  • Parse external data (JSON, API responses) into Pydantic models at the boundary
  • This catches type errors at parse time, not deep in business logic
# BAD - error-prone dict access
def process(data: dict) -> str:
    return data.get("name", "")  # No validation, silent failures

# GOOD - typed and validated
class UserData(BaseModel):
    name: str
    created_at: datetime

    @field_validator("created_at", mode="before")
    @classmethod
    def ensure_tz_aware(cls, v):
        if isinstance(v, str):
            v = datetime.fromisoformat(v.replace("Z", "+00:00"))
        if v.tzinfo is None:
            return v.replace(tzinfo=timezone.utc)
        return v

def process(data: UserData) -> str:
    return data.name  # Type-safe, validated at construction

TypeScript Style

  • Next.js App Router for control plane
  • Tailwind CSS with shadcn/ui components

Adding New API Configuration Flags

Configuration follows a hierarchical system: Global (env vars) → Tenant (via extension) → Bank (database).

Fields must be categorized as either hierarchical (can be overridden per-tenant/bank) or static (server-level only).

Adding a New Configuration Field

  1. config.py (hindsight-api/hindsight_api/config.py):

    • Add ENV_* constant for the environment variable name (e.g., ENV_MY_SETTING = "HINDSIGHT_API_MY_SETTING")
    • Add DEFAULT_* constant for the default value
    • Add field to HindsightConfig dataclass with type annotation
    • Mark as hierarchical or static by adding to _HIERARCHICAL_FIELDS set (hierarchical) or leaving it out (static)
    • Add initialization in from_env() method
    # Hierarchical field (can be overridden per-bank)
    _HIERARCHICAL_FIELDS = {
        ...,
        "my_setting",  # Add here for hierarchical
    }
    
    # Static field - just don't add to _HIERARCHICAL_FIELDS
    
  2. main.py (hindsight-api/hindsight_api/main.py):

    • Add field to the manual HindsightConfig() constructor call (search for "CLI override")
  3. Use hierarchical config in MemoryEngine:

    # Config is resolved automatically per bank via ConfigResolver
    config_dict = await self._config_resolver.get_bank_config(bank_id, context)
    value = config_dict["my_setting"]
    
  4. Use static config (non-hierarchical):

    from ...config import get_config
    config = get_config()
    value = config.my_static_field
    
  5. Documentation (hindsight-docs/docs/developer/configuration.md):

    • Add to appropriate section table with Variable, Description, Default
    • Mark if it's hierarchical (can be overridden per-bank)

Hierarchical vs Static Guidelines

Hierarchical (per-bank overridable):

  • LLM settings (provider, model, API key, base URL)
  • Operation-specific settings (retain mode, chunk size, etc.)
  • Feature flags that vary by customer/bank

Static (server-level only):

  • Infrastructure settings (database URL, port, host)
  • Global limits (max concurrent operations)
  • System-wide feature flags

Environment Setup

cp .env.example .env
# Edit .env with LLM API key

# Python deps
uv sync --directory hindsight-api/

# Node deps (uses npm workspaces)
npm install

Required env vars:

  • HINDSIGHT_API_LLM_PROVIDER: openai, anthropic, gemini, groq, ollama, lmstudio
  • HINDSIGHT_API_LLM_API_KEY: Your API key
  • HINDSIGHT_API_LLM_MODEL: Model name (e.g., o3-mini, claude-sonnet-4-20250514)

Optional (uses local models by default):

  • HINDSIGHT_API_EMBEDDINGS_PROVIDER: local (default) or tei
  • HINDSIGHT_API_RERANKER_PROVIDER: local (default) or tei
  • HINDSIGHT_API_DATABASE_URL: External PostgreSQL (uses embedded pg0 by default)
  • HINDSIGHT_API_ENABLE_BANK_CONFIG_API: Enable per-bank config API (default: false, disabled for security)