fleet-memory/hindsight-docs/docs/developer/server.md
2025-11-25 19:28:26 +01:00

14 KiB

sidebar_position
6

Server Deployment

Guide to deploying the Hindsight server in production.

Deployment Options

The simplest way to deploy Hindsight with all dependencies:

# Clone the repository
git clone https://github.com/vectorize-io/hindsight.git
cd hindsight

# Create environment file
cp .env.example .env
# Edit .env with your LLM API key

# Start all services
cd docker
./start.sh

Services will be available at:

Docker Compose Commands

# Start all services
cd docker && ./start.sh

# Stop services
cd docker && ./stop.sh

# Clean all data
cd docker && ./clean.sh

# View logs
docker-compose logs -f api
docker-compose logs -f control-plane

Helm Chart (Kubernetes)

For Kubernetes deployments, use the Hindsight Helm chart:

helm repo add hindsight https://vectorize-io.github.io/hindsight
helm install hindsight hindsight/hindsight \
  --set api.llm.provider=groq \
  --set api.llm.apiKey=gsk_xxxxxxxxxxxx

See the Helm chart documentation for configuration options.

pip install

For custom deployments, install the all-in-one package:

pip install hindsight-all

Run the server:

# Configure via environment variables
export HINDSIGHT_API_LLM_PROVIDER=groq
export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx

# Start the server (Ctrl+C to stop)
hindsight-api

By default, it uses pg0 (embedded PostgreSQL) so you can run it without any external dependencies.

CLI Options

hindsight-api --help

# Common options
hindsight-api --port 9000           # Custom port (default: 8888)
hindsight-api --host 127.0.0.1      # Bind to localhost only
hindsight-api --mcp                 # Enable MCP server at /mcp
hindsight-api --log-level debug     # Verbose logging

Using External PostgreSQL

To use an external PostgreSQL database instead of embedded pg0:

export HINDSIGHT_API_DATABASE_URL=postgresql://user:pass@localhost:5432/hindsight
hindsight-api

You'll need PostgreSQL with the pgvector extension enabled.

Architecture Overview

┌─────────────────────────────────────────────────────────────────────────────┐
│                              CLIENTS                                         │
├─────────────────┬─────────────────┬─────────────────┬───────────────────────┤
│  Python Client  │  Node.js Client │      CLI        │   AI Assistants       │
│  hindsight-client  │  @hindsight/client │   hindsight-cli    │  (Claude, etc.)       │
└────────┬────────┴────────┬────────┴────────┬────────┴───────────┬───────────┘
         │                 │                 │                     │
         ▼                 ▼                 ▼                     ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│                           HINDSIGHT API SERVER                               │
│                          (localhost:8888)                                    │
├─────────────────────────────────┬───────────────────────────────────────────┤
│          HTTP API               │              MCP API                       │
│    /api/memories/*              │         MCP Server (stdio)                 │
│    /api/agents/*                │      hindsight_search, hindsight_think,    │
│    /api/search, /api/think      │      hindsight_store, hindsight_agents     │
└─────────────────────────────────┴───────────────────────────────────────────┘
         │                                           │
         ▼                                           ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│                          PROCESSING PIPELINE                                 │
├─────────────────────────────────────────────────────────────────────────────┤
│  ┌──────────────┐    ┌──────────────┐    ┌──────────────┐                   │
│  │   INGESTION  │    │  RETRIEVAL   │    │  REASONING   │                   │
│  │              │    │   (TEMPR)    │    │   (CARA)     │                   │
│  │  LLM Extract │    │              │    │              │                   │
│  │  Entity Res. │    │  4-way Search│    │  Personality │                   │
│  │  Graph Build │    │  RRF Fusion  │    │  Opinion Gen │                   │
│  └──────────────┘    └──────────────┘    └──────────────┘                   │
└─────────────────────────────────────────────────────────────────────────────┘
         │                      │                      │
         ▼                      ▼                      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│                              ML MODELS                                       │
├───────────────────┬───────────────────┬─────────────────────────────────────┤
│    Embeddings     │   Cross-Encoder   │         LLM Provider                 │
│  all-MiniLM-L6-v2 │ ms-marco-MiniLM   │   OpenAI / Groq / Ollama            │
│    (384-dim)      │   (reranking)     │   (extraction, reasoning)           │
└───────────────────┴───────────────────┴─────────────────────────────────────┘
         │                      │                      │
         ▼                      ▼                      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│                         POSTGRESQL + PGVECTOR                                │
│                          (localhost:5432)                                    │
├─────────────────────────────────────────────────────────────────────────────┤
│  • Memory Units (facts, opinions)    • HNSW Vector Index                    │
│  • Entity Graph (nodes, edges)       • GIN Full-Text Index                  │
│  • Agent Profiles                    • Temporal Indexes                     │
└─────────────────────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────────────────────┐
│                      CONTROL PLANE (Optional)                                │
│                         (localhost:3000)                                     │
├─────────────────────────────────────────────────────────────────────────────┤
│  • Web UI for administration         • Agent management                     │
│  • Memory visualization              • Graph explorer                       │
│  • Connects to API Server            • Monitoring dashboard                 │
└─────────────────────────────────────────────────────────────────────────────┘

Environment Variables

API Server (HINDSIGHT_API_*)

Variable Description Default
HINDSIGHT_API_DATABASE_URL PostgreSQL connection string Required
HINDSIGHT_API_LLM_PROVIDER LLM provider: openai, groq, ollama groq
HINDSIGHT_API_LLM_API_KEY API key for LLM provider Required (except ollama)
HINDSIGHT_API_LLM_MODEL Model name llama-3.1-70b-versatile
HINDSIGHT_API_LLM_BASE_URL Custom LLM endpoint Provider default
HINDSIGHT_API_HOST Server bind address 0.0.0.0
HINDSIGHT_API_PORT Server port 8888
HINDSIGHT_API_MCP_ENABLED Enable MCP server true

Control Plane (HINDSIGHT_CP_*)

Variable Description Default
HINDSIGHT_CP_DATAPLANE_API_URL API server URL http://localhost:8888
HINDSIGHT_CP_HOSTNAME Server bind address 0.0.0.0
HINDSIGHT_CP_PORT Server port 3000

Example Configuration

# .env file

# Database
HINDSIGHT_API_DATABASE_URL=postgresql://hindsight:hindsight_dev@localhost:5432/hindsight

# LLM - Using Groq (fast inference)
HINDSIGHT_API_LLM_PROVIDER=groq
HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx
HINDSIGHT_API_LLM_MODEL=llama-3.1-70b-versatile

# LLM - Using OpenAI
# HINDSIGHT_API_LLM_PROVIDER=openai
# HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx
# HINDSIGHT_API_LLM_MODEL=gpt-4o

# LLM - Using Ollama (local, no API key)
# HINDSIGHT_API_LLM_PROVIDER=ollama
# HINDSIGHT_API_LLM_BASE_URL=http://localhost:11434/v1
# HINDSIGHT_API_LLM_MODEL=llama3.1

# Control Plane
HINDSIGHT_CP_DATAPLANE_API_URL=http://localhost:8888

ML Models

Hindsight uses several ML models that are downloaded automatically on first run:

Embedding Model

Model Dimensions Purpose
all-MiniLM-L6-v2 384 Semantic vector embeddings

Used for memory vectorization, query embedding, and semantic similarity search.

Cross-Encoder (Reranking)

Model Purpose
cross-encoder/ms-marco-MiniLM-L-6-v2 Neural reranking

Reranks search results for precision after initial retrieval.

Temporal Parser

Model Purpose
t5-small Temporal expression parsing

Parses natural language time expressions like "last spring" or "in June 2024".

LLM (Configurable)

Used for fact extraction, entity resolution, opinion generation, and think responses.

Supported providers:

  • OpenAI: GPT-4, GPT-4o, GPT-3.5-turbo
  • Groq: Llama 3.1, Mixtral (fast inference)
  • Ollama: Any local model

Database Schema

PostgreSQL with pgvector extension:

-- Memory units (facts and opinions)
CREATE TABLE memory_units (
    id UUID PRIMARY KEY,
    agent_id VARCHAR NOT NULL,
    text TEXT NOT NULL,
    fact_type VARCHAR NOT NULL,  -- 'world', 'agent', 'opinion'
    confidence_score FLOAT,
    embedding VECTOR(384),
    occurred_start DATE,
    occurred_end DATE,
    mentioned_at TIMESTAMP,
    context VARCHAR,
    document_id VARCHAR
);

-- Entity graph
CREATE TABLE entities (
    id UUID PRIMARY KEY,
    agent_id VARCHAR NOT NULL,
    name VARCHAR NOT NULL,
    entity_type VARCHAR NOT NULL,
    canonical_name VARCHAR
);

CREATE TABLE entity_links (
    memory_id UUID REFERENCES memory_units(id),
    entity_id UUID REFERENCES entities(id),
    PRIMARY KEY (memory_id, entity_id)
);

-- Agent profiles
CREATE TABLE agent_profiles (
    agent_id VARCHAR PRIMARY KEY,
    background TEXT,
    openness FLOAT DEFAULT 0.5,
    conscientiousness FLOAT DEFAULT 0.5,
    extraversion FLOAT DEFAULT 0.5,
    agreeableness FLOAT DEFAULT 0.5,
    neuroticism FLOAT DEFAULT 0.5,
    bias_strength FLOAT DEFAULT 0.5
);

Indexes

  • HNSW Index: Fast approximate nearest neighbor for vector search
  • GIN Index: Full-text search with BM25 ranking
  • B-tree Indexes: Agent ID, timestamps, entity lookups

Health Checks

API Server

curl http://localhost:8888/api/v1/agents

Control Plane

curl http://localhost:3000/

Production Considerations

For production deployments:

  1. Use managed PostgreSQL with pgvector extension (AWS RDS, Google Cloud SQL, Supabase)
  2. Set proper secrets via environment variables or secrets manager
  3. Configure resource limits for ML model inference
  4. Set up monitoring for API latency and error rates
  5. Use HTTPS with proper TLS certificates
  6. Configure rate limiting at load balancer level

Resource Requirements

Component CPU Memory Notes
API Server 2+ cores 4GB+ ML models loaded in memory
PostgreSQL 2+ cores 4GB+ Depends on data size
Control Plane 1 core 512MB Lightweight Next.js app