diff --git a/hindsight-docs/docs/developer/installation.md b/hindsight-docs/docs/developer/installation.md index d5ce0c3c..796f599d 100644 --- a/hindsight-docs/docs/developer/installation.md +++ b/hindsight-docs/docs/developer/installation.md @@ -246,5 +246,5 @@ See the [Python SDK](../sdks/python.md) for the full API reference. ## Next Steps - [Configuration](./configuration.md) — Environment variables and settings -- [Models](./models.md) — ML models and providers +- [Models](./models.mdx) — ML models and providers - [Monitoring](./monitoring.md) — Metrics and observability diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/api/quickstart.mdx b/hindsight-docs/versioned_docs/version-0.4/developer/api/quickstart.mdx index 65bacef2..2f6bab98 100644 --- a/hindsight-docs/versioned_docs/version-0.4/developer/api/quickstart.mdx +++ b/hindsight-docs/versioned_docs/version-0.4/developer/api/quickstart.mdx @@ -9,12 +9,17 @@ Get up and running with Hindsight in 60 seconds. import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; import CodeSnippet from '@site/src/components/CodeSnippet'; +import {ClientsGrid, IntegrationsGrid} from '@site/src/components/SupportedGrids'; {/* Import raw source files */} import quickstartPy from '!!raw-loader!@site/examples/api/quickstart.py'; import quickstartMjs from '!!raw-loader!@site/examples/api/quickstart.mjs'; import quickstartSh from '!!raw-loader!@site/examples/api/quickstart.sh'; +## Clients + + + ## Start the API Server @@ -100,6 +105,10 @@ curl -fsSL https://hindsight.vectorize.io/get-cli | bash --- +## Integrations + + + ## Next Steps - [**Retain**](./retain) — Advanced options for storing memories diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/api/recall.mdx b/hindsight-docs/versioned_docs/version-0.4/developer/api/recall.mdx index a6ef4f4b..8fee9e41 100644 --- a/hindsight-docs/versioned_docs/version-0.4/developer/api/recall.mdx +++ b/hindsight-docs/versioned_docs/version-0.4/developer/api/recall.mdx @@ -184,6 +184,66 @@ Use this for strict scope enforcement where a memory must explicitly belong to * A memory with tags `["user:alice", "team", "project:x"]` will still match a filter of `["user:alice", "team"]` under `all_strict` — extra tags on the memory are not a problem. The filter only requires the memory to contain **at least** the specified tags. ::: +### tag_groups + +`tag_groups` is a list of compound boolean tag filters. The groups in the list are AND-ed together at the top level. Each group is a recursive boolean expression: a **leaf** node `{tags, match}`, or a **compound** node `{and: [...]}`, `{or: [...]}`, or `{not: ...}`. + +`tag_groups` and `tags` / `tags_match` can be used simultaneously — they are AND-ed together. + +#### Leaf node + +```json +{ "tags": ["step:5", "step:8"], "match": "any_strict" } +``` + +`match` accepts the same values as `tags_match`: `any`, `all`, `any_strict`, `all_strict`. Defaults to `any_strict`. + +#### Compound nodes + +```json +{ "and": [ , , ... ] } +{ "or": [ , , ... ] } +{ "not": } +``` + +#### Examples + +**Step filter AND user scope** — two top-level groups AND-ed: + +```json +{ + "tag_groups": [ + { "tags": ["step:5", "step:8", "step:12"], "match": "any_strict" }, + { "tags": ["user:ep_42"], "match": "all_strict" } + ] +} +``` + +**Nested OR inside AND** — user must match, plus either step OR priority: + +```json +{ + "tag_groups": [ + { "tags": ["user:alice"], "match": "all_strict" }, + { "or": [ + { "tags": ["step:5"], "match": "any_strict" }, + { "tags": ["priority:high"], "match": "all_strict" } + ]} + ] +} +``` + +**Exclusion** — user must match, but archived memories are excluded: + +```json +{ + "tag_groups": [ + { "tags": ["user:alice"], "match": "all_strict" }, + { "not": { "tags": ["archived"], "match": "any_strict" } } + ] +} +``` + ### trace When set to `true`, the response includes a detailed debug trace covering the query embedding, entry points, per-strategy retrieval results, RRF fusion candidates, reranked results, temporal constraints detected, and per-phase timings. Has no effect on the retrieval logic itself. Useful for understanding why specific memories were or were not returned. diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/configuration.md b/hindsight-docs/versioned_docs/version-0.4/developer/configuration.md index 43f3368c..012746b7 100644 --- a/hindsight-docs/versioned_docs/version-0.4/developer/configuration.md +++ b/hindsight-docs/versioned_docs/version-0.4/developer/configuration.md @@ -160,7 +160,7 @@ To switch between backends: | Variable | Description | Default | |----------|-------------|---------| -| `HINDSIGHT_API_LLM_PROVIDER` | Provider: `openai`, `openai-codex`, `claude-code`, `anthropic`, `gemini`, `groq`, `ollama`, `lmstudio`, `vertexai` | `openai` | +| `HINDSIGHT_API_LLM_PROVIDER` | Provider: `openai`, `openai-codex`, `claude-code`, `anthropic`, `gemini`, `groq`, `minimax`, `ollama`, `lmstudio`, `vertexai` | `openai` | | `HINDSIGHT_API_LLM_API_KEY` | API key for LLM provider | - | | `HINDSIGHT_API_LLM_MODEL` | Model name | `gpt-5-mini` | | `HINDSIGHT_API_LLM_BASE_URL` | Custom LLM endpoint | Provider default | @@ -410,7 +410,7 @@ Supported OpenAI embedding dimensions: | Variable | Description | Default | |----------|-------------|---------| -| `HINDSIGHT_API_RERANKER_PROVIDER` | Provider: `local`, `tei`, `cohere`, `zeroentropy`, `flashrank`, `litellm`, `litellm-sdk`, or `rrf` | `local` | +| `HINDSIGHT_API_RERANKER_PROVIDER` | Provider: `local`, `tei`, `cohere`, `zeroentropy`, `flashrank`, `litellm`, `litellm-sdk`, `jina-mlx`, or `rrf` | `local` | | `HINDSIGHT_API_RERANKER_LOCAL_MODEL` | Model for local provider | `cross-encoder/ms-marco-MiniLM-L-6-v2` | | `HINDSIGHT_API_RERANKER_LOCAL_MAX_CONCURRENT` | Max concurrent local reranking (prevents CPU thrashing under load) | `4` | | `HINDSIGHT_API_RERANKER_LOCAL_TRUST_REMOTE_CODE` | Allow loading models with custom code (security risk, disabled by default) | `false` | @@ -426,10 +426,12 @@ Supported OpenAI embedding dimensions: | `HINDSIGHT_API_RERANKER_LITELLM_SDK_API_KEY` | LiteLLM **SDK** API key for direct reranking (no proxy needed) | - | | `HINDSIGHT_API_RERANKER_LITELLM_SDK_MODEL` | LiteLLM SDK rerank model (e.g., `deepinfra/Qwen3-reranker-8B`) | `cohere/rerank-english-v3.0` | | `HINDSIGHT_API_RERANKER_LITELLM_SDK_API_BASE` | Custom API base URL for LiteLLM SDK (optional) | - | +| `HINDSIGHT_API_RERANKER_LITELLM_MAX_TOKENS_PER_DOC` | Truncate documents to this many tokens before sending to the reranker (applies to both `litellm` and `litellm-sdk`). Use for models with small context windows (e.g. set to `900` for a 1024-token limit model). Unset by default (no truncation). | - | | `HINDSIGHT_API_RERANKER_ZEROENTROPY_API_KEY` | ZeroEntropy API key for reranking | - | | `HINDSIGHT_API_RERANKER_ZEROENTROPY_MODEL` | ZeroEntropy rerank model (`zerank-2`, `zerank-2-small`) | `zerank-2` | | `HINDSIGHT_API_RERANKER_FLASHRANK_MODEL` | FlashRank model for fast CPU-based reranking | `ms-marco-MiniLM-L-12-v2` | | `HINDSIGHT_API_RERANKER_FLASHRANK_CACHE_DIR` | Cache directory for FlashRank models | System default | +| `HINDSIGHT_API_RERANKER_JINA_MLX_MODEL_PATH` | Local path to downloaded `jina-reranker-v3-mlx` model (auto-downloads from HuggingFace if unset) | - | ```bash # Local (default) - uses SentenceTransformers CrossEncoder @@ -472,6 +474,10 @@ export HINDSIGHT_API_RERANKER_LITELLM_MODEL=cohere/rerank-english-v3.0 # or voy export HINDSIGHT_API_RERANKER_PROVIDER=litellm-sdk export HINDSIGHT_API_RERANKER_LITELLM_SDK_API_KEY=your-deepinfra-api-key export HINDSIGHT_API_RERANKER_LITELLM_SDK_MODEL=deepinfra/Qwen3-reranker-8B # or cohere/rerank-english-v3.0, etc. + +# Jina MLX - Apple Silicon native reranking (no GPU/cloud required) +# Model (~1.2 GB) is downloaded automatically from HuggingFace Hub on first use. +export HINDSIGHT_API_RERANKER_PROVIDER=jina-mlx ``` #### LiteLLM Proxy vs SDK @@ -488,6 +494,14 @@ Both support the same providers: - **Jina AI** (`jina_ai/jina-reranker-v2`) - **AWS Bedrock** (`bedrock/...`) +#### Jina MLX (Apple Silicon) + +The `jina-mlx` provider uses [`jinaai/jina-reranker-v3-mlx`](https://huggingface.co/jinaai/jina-reranker-v3-mlx), optimized for Apple Silicon. The model (~1.2 GB) is downloaded from HuggingFace Hub automatically on first startup and cached locally. + +:::note License +`jina-reranker-v3-mlx` is licensed under CC BY-NC 4.0. Contact Jina AI for commercial usage. +::: + ### Authentication By default, Hindsight runs without authentication. For production deployments, enable API key authentication using the built-in tenant extension: @@ -530,6 +544,7 @@ For advanced authentication (JWT, OAuth, multi-tenant schemas), implement a cust | `HINDSIGHT_API_GRAPH_RETRIEVER` | Graph retrieval algorithm: `link_expansion`, `mpfp`, or `bfs` | `link_expansion` | | `HINDSIGHT_API_RECALL_MAX_CONCURRENT` | Max concurrent recall operations per worker (backpressure) | `32` | | `HINDSIGHT_API_RECALL_CONNECTION_BUDGET` | Max concurrent DB connections per recall operation | `4` | +| `HINDSIGHT_API_RECALL_MAX_QUERY_TOKENS` | Maximum token length of a recall query; requests exceeding this limit are rejected with HTTP 400 | `500` | | `HINDSIGHT_API_RERANKER_MAX_CANDIDATES` | Max candidates to rerank per recall (RRF pre-filters the rest) | `300` | | `HINDSIGHT_API_MPFP_TOP_K_NEIGHBORS` | Fan-out limit per node in MPFP graph traversal | `20` | | `HINDSIGHT_API_MENTAL_MODEL_REFRESH_CONCURRENCY` | Max concurrent mental model refreshes | `8` | diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/index.md b/hindsight-docs/versioned_docs/version-0.4/developer/index.mdx similarity index 97% rename from hindsight-docs/versioned_docs/version-0.4/developer/index.md rename to hindsight-docs/versioned_docs/version-0.4/developer/index.mdx index 8c25e95b..22a712d2 100644 --- a/hindsight-docs/versioned_docs/version-0.4/developer/index.md +++ b/hindsight-docs/versioned_docs/version-0.4/developer/index.mdx @@ -3,6 +3,8 @@ sidebar_position: 1 slug: / --- +import {ClientsGrid, IntegrationsGrid} from '@site/src/components/SupportedGrids'; + # Overview ## Why Hindsight? @@ -114,6 +116,14 @@ The **mission** tells Hindsight what knowledge to prioritize and provides contex These settings only affect the `reflect` operation, not `recall`. +## Clients & Languages + + + +## Integrations + + + ## Next Steps ### Getting Started diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/installation.md b/hindsight-docs/versioned_docs/version-0.4/developer/installation.md index e30ba8be..28a9e5c1 100644 --- a/hindsight-docs/versioned_docs/version-0.4/developer/installation.md +++ b/hindsight-docs/versioned_docs/version-0.4/developer/installation.md @@ -8,28 +8,28 @@ Hindsight can be deployed in several ways depending on your infrastructure and r ## Prerequisites -### PostgreSQL with pgvector +### PostgreSQL -Hindsight requires PostgreSQL with the **pgvector** extension for vector similarity search. +Hindsight requires PostgreSQL 14+ with a vector extension for similarity search. The supported extensions are: + +- **pgvector** (default) +- **pgvectorscale** +- **vchord** + +Configure which one to use with `HINDSIGHT_API_VECTOR_EXTENSION`. See [Configuration](./configuration) for details. **By default**, Hindsight uses **pg0** — an embedded PostgreSQL that runs locally on your machine. This is convenient for development but **not recommended for production**. -**For production**, use an external PostgreSQL with pgvector: +**For production**, use an external PostgreSQL with one of the supported vector extensions: - **Supabase** — Managed PostgreSQL with pgvector built-in - **Neon** — Serverless PostgreSQL with pgvector -- **Azure Database for PostgreSQL** — With pgvector and pg_diskann (DiskANN) support +- **Azure Database for PostgreSQL** — With pgvector and pgvectorscale support - **AWS RDS** / **Cloud SQL** — With pgvector extension enabled -- **Self-hosted** — PostgreSQL 14+ with pgvector installed +- **Self-hosted** — PostgreSQL 14+ with your preferred vector extension ### LLM Provider -You need an LLM API key for fact extraction, entity resolution, and answer generation: - -- **Groq** (recommended): Fast inference with `gpt-oss-20b` -- **OpenAI**: GPT-4o, GPT-4o-mini -- **Ollama**: Run models locally - -See [Models](./models) for detailed comparison and configuration. +You need an LLM API key for fact extraction, entity resolution, and answer generation. See [Models](./models) for supported providers, model recommendations, and configuration. --- @@ -53,61 +53,12 @@ docker run --rm -it --pull always -p 8888:8888 -p 9999:9999 \ ### Docker Image Variants -Hindsight provides two image variants with different size/capability tradeoffs: +| Variant | Size (AMD64) | Size (ARM64) | When to use | +|---------|--------------|--------------|-------------| +| **Full** (`latest`) | ~9 GB | ~3.7 GB | Default. Works out of the box with no external services except the LLM. | +| **Slim** (`slim`) | ~500 MB | ~500 MB | Use when you already rely on external services for embeddings and reranking (OpenAI, Cohere, TEI). Significantly smaller image, faster deploys. Requires [external providers](./configuration#embeddings). | -| Variant | Size (AMD64) | Size (ARM64) | Use Case | -|---------|--------------|--------------|----------| -| **Full** (`latest`) | ~9 GB | ~3.7 GB | Includes local ML models (embeddings, reranking) | -| **Slim** (`slim`) | ~500 MB | ~500 MB | Requires external embedding/reranking providers | - -**Full image** (default): -```bash -docker run --rm -it -p 8888:8888 \ - -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \ - ghcr.io/vectorize-io/hindsight:latest -``` -- ✅ Works out of the box with local ML models -- ✅ No additional services needed -- ❌ Larger image size (AMD64 includes CUDA libraries for GPU support) - -**Slim image**: -```bash -docker run --rm -it -p 8888:8888 \ - -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \ - -e HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai \ - -e HINDSIGHT_API_RERANKER_PROVIDER=cohere \ - -e HINDSIGHT_API_COHERE_API_KEY=$COHERE_API_KEY \ - ghcr.io/vectorize-io/hindsight:latest-slim -``` -- ✅ Dramatically smaller image (~95% reduction on AMD64) -- ✅ Faster pull/deploy times -- ✅ Lower memory footprint -- ❌ Requires external embedding/reranking services (OpenAI, Cohere, TEI) - -**When to use slim:** -- Cloud deployments where image size matters -- Using managed embedding services (OpenAI, Cohere) -- Running on Text Embeddings Inference (TEI) infrastructure -- Kubernetes environments with fast pull requirements - -:::warning Slim Image Requires External Providers -If you run the slim image **without** setting external embedding providers, you'll see this error: - -``` -ImportError: sentence-transformers is required for LocalSTEmbeddings. -Install it with: pip install sentence-transformers -``` - -**Fix:** Always set embedding and reranking providers when using slim images: -```bash --e HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai --e HINDSIGHT_API_EMBEDDINGS_OPENAI_API_KEY=sk-xxx --e HINDSIGHT_API_RERANKER_PROVIDER=cohere --e HINDSIGHT_API_COHERE_API_KEY=xxx -``` -::: - -See [Configuration](./configuration#embeddings) for all embedding provider options. +The slim image corresponds to the [`hindsight-api-slim`](#bare-metal-pip) pip package. See [Configuration](./configuration#embeddings) for external provider options. ### Available Tags @@ -120,7 +71,7 @@ ghcr.io/vectorize-io/hindsight:0.4.9-slim # Slim, specific version # API only ghcr.io/vectorize-io/hindsight-api:latest -ghcr.io/vectorize-io/hindsight-api:slim +ghcr.io/vectorize-io/hindsight-api:latest-slim # Control Plane only ghcr.io/vectorize-io/hindsight-control-plane:latest @@ -175,14 +126,17 @@ See the [Helm chart values.yaml](https://github.com/vectorize-io/hindsight/tree/ ## Bare Metal (pip) -**Best for**: Custom deployments, integration into existing Python applications +**Best for**: Running Hindsight as a standalone service on a host machine. ### Install ```bash -pip install hindsight-all +pip install hindsight-api # Full — works out of the box +pip install hindsight-api-slim # Slim — requires external services for embeddings, reranking, and the database ``` +When using `hindsight-api-slim`, you must configure external providers for all model operations. See [Configuration](./configuration#embeddings) for details. + ### Run with Embedded Database For development and testing, Hindsight can run with an embedded PostgreSQL (pg0): @@ -253,8 +207,44 @@ PORT=80 HINDSIGHT_CP_DATAPLANE_API_URL=https://api.hindsight.io npx @vectorize-i --- +## Embedded in a Python Application + +**Best for**: Using Hindsight programmatically from Python without running a separate server process. + +```bash +pip install hindsight-all # Full — works out of the box +pip install hindsight-all-slim # Slim — requires external services for embeddings, reranking, and the database +``` + +`hindsight-all` supports two modes of embedding: + +**In-process** (`HindsightServer`): the server runs in a background thread inside your application. Best when you want the tightest integration and are already managing your own process lifecycle. + +```python +from hindsight import HindsightServer, HindsightClient + +with HindsightServer(llm_provider="openai", llm_api_key="sk-xxx") as server: + client = HindsightClient(base_url=server.url) + client.retain(bank_id="alice", content="Alice prefers concise answers.") + results = client.recall(bank_id="alice", query="How should I respond to Alice?") +``` + +**Managed subprocess** (`HindsightEmbedded`): the server runs as a background daemon process, shared across multiple Python processes or sessions. The daemon starts on first use and shuts down automatically after an idle timeout. + +```python +from hindsight import HindsightEmbedded + +client = HindsightEmbedded(llm_provider="openai", llm_api_key="sk-xxx") +client.retain(bank_id="alice", content="Alice prefers concise answers.") +results = client.recall(bank_id="alice", query="How should I respond to Alice?") +``` + +See the [Python SDK](../sdks/python.md) for the full API reference. + +--- + ## Next Steps - [Configuration](./configuration.md) — Environment variables and settings -- [Models](./models.md) — ML models and providers +- [Models](./models.mdx) — ML models and providers - [Monitoring](./monitoring.md) — Metrics and observability diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/models.md b/hindsight-docs/versioned_docs/version-0.4/developer/models.mdx similarity index 95% rename from hindsight-docs/versioned_docs/version-0.4/developer/models.md rename to hindsight-docs/versioned_docs/version-0.4/developer/models.mdx index 7b2c817d..22bb8c95 100644 --- a/hindsight-docs/versioned_docs/version-0.4/developer/models.md +++ b/hindsight-docs/versioned_docs/version-0.4/developer/models.mdx @@ -1,16 +1,16 @@ +import {LLMProvidersGrid} from '@site/src/components/SupportedGrids'; + # Models Hindsight uses several machine learning models for different tasks. ## Overview -| Model Type | Purpose | Default | Configurable | -|------------|---------|---------|--------------| -| **LLM** | Fact extraction, reasoning, generation | Provider-specific | Yes | -| **Embedding** | Vector representations for semantic search | `BAAI/bge-small-en-v1.5` | Yes | -| **Cross-Encoder** | Reranking search results | `cross-encoder/ms-marco-MiniLM-L-6-v2` | Yes | +- **LLM** — Fact extraction, reasoning, and generation. Provider-specific, fully configurable. +- **Embedding** — Vector representations for semantic search. Default: `BAAI/bge-small-en-v1.5`. +- **Cross-Encoder** — Reranking search results. Default: `cross-encoder/ms-marco-MiniLM-L-6-v2`. -All local models (embedding, cross-encoder) are automatically downloaded from HuggingFace on first run. +Embedding and cross-encoder models are downloaded automatically from HuggingFace on first run. --- @@ -18,7 +18,11 @@ All local models (embedding, cross-encoder) are automatically downloaded from Hu Used for fact extraction, entity resolution, mental model consolidation, and answer synthesis. -**Supported providers:** OpenAI, Anthropic, Gemini, Groq, Ollama, LM Studio, and **any OpenAI-compatible API** +**Supported providers:** + + + +Also supports **any OpenAI-compatible API** (e.g., Azure OpenAI, Together AI, Fireworks). :::tip OpenAI-Compatible Providers Hindsight works with any provider that exposes an OpenAI-compatible API (e.g., Azure OpenAI). Simply set `HINDSIGHT_API_LLM_PROVIDER=openai` and configure `HINDSIGHT_API_LLM_BASE_URL` to point to your provider's endpoint. @@ -63,6 +67,7 @@ Each provider has a recommended default model that's used when `HINDSIGHT_API_LL | `anthropic` | `claude-haiku-4-5-20251001` | | `gemini` | `gemini-2.5-flash` | | `groq` | `openai/gpt-oss-120b` | +| `minimax` | `MiniMax-M2.5` | | `ollama` | `gemma3:12b` | | `lmstudio` | `local-model` | | `vertexai` | `gemini-2.0-flash-001` | @@ -144,6 +149,11 @@ export HINDSIGHT_API_LLM_PROVIDER=lmstudio export HINDSIGHT_API_LLM_BASE_URL=http://localhost:1234/v1 export HINDSIGHT_API_LLM_MODEL=your-local-model +# MiniMax (204K context window) +export HINDSIGHT_API_LLM_PROVIDER=minimax +export HINDSIGHT_API_LLM_API_KEY=your-minimax-api-key +export HINDSIGHT_API_LLM_MODEL=MiniMax-M2.5 + # Vertex AI (Google Cloud) export HINDSIGHT_API_LLM_PROVIDER=vertexai export HINDSIGHT_API_LLM_MODEL=gemini-2.0-flash-001 diff --git a/hindsight-docs/versioned_docs/version-0.4/sdks/embed.md b/hindsight-docs/versioned_docs/version-0.4/sdks/embed.md index f94dd8e1..7b1ba402 100644 --- a/hindsight-docs/versioned_docs/version-0.4/sdks/embed.md +++ b/hindsight-docs/versioned_docs/version-0.4/sdks/embed.md @@ -88,7 +88,7 @@ The daemon starts automatically on first use! | Variable | Description | Default | |----------|-------------|---------| | `HINDSIGHT_EMBED_LLM_API_KEY` | **Required**. API key for LLM provider | - | -| `HINDSIGHT_EMBED_LLM_PROVIDER` | LLM provider: `openai`, `anthropic`, `gemini`, `groq`, `ollama` | `openai` | +| `HINDSIGHT_EMBED_LLM_PROVIDER` | LLM provider: `openai`, `anthropic`, `gemini`, `groq`, `minimax`, `ollama` | `openai` | | `HINDSIGHT_EMBED_LLM_MODEL` | Model name | `gpt-4o-mini` | | `HINDSIGHT_EMBED_BANK_ID` | Default memory bank ID | `default` | | `HINDSIGHT_EMBED_DAEMON_IDLE_TIMEOUT` | Seconds before daemon auto-exits when idle (0 = never) | `300` | diff --git a/hindsight-docs/versioned_sidebars/version-0.4-sidebars.json b/hindsight-docs/versioned_sidebars/version-0.4-sidebars.json index 1d74ec9c..6329cd25 100644 --- a/hindsight-docs/versioned_sidebars/version-0.4-sidebars.json +++ b/hindsight-docs/versioned_sidebars/version-0.4-sidebars.json @@ -5,15 +5,78 @@ "label": "Architecture", "collapsible": false, "items": [ - { "type": "doc", "id": "developer/index", "label": "Overview", "customProps": { "icon": "lu-book" } }, - { "type": "doc", "id": "developer/retain", "label": "Retain", "customProps": { "icon": "lu-brain" } }, - { "type": "doc", "id": "developer/retrieval", "label": "Recall", "customProps": { "icon": "lu-search" } }, - { "type": "doc", "id": "developer/reflect", "label": "Reflect", "customProps": { "icon": "lu-message" } }, - { "type": "doc", "id": "developer/observations", "label": "Observations", "customProps": { "icon": "lu-activity" } }, - { "type": "doc", "id": "developer/multilingual", "label": "Multilingual", "customProps": { "icon": "lu-languages" } }, - { "type": "doc", "id": "developer/performance", "label": "Performance", "customProps": { "icon": "lu-zap" } }, - { "type": "doc", "id": "developer/storage", "label": "Storage", "customProps": { "icon": "lu-database" } }, - { "type": "doc", "id": "developer/rag-vs-hindsight", "label": "RAG vs Memory", "customProps": { "icon": "lu-compare" } } + { + "type": "doc", + "id": "developer/index", + "label": "Overview", + "customProps": { + "icon": "lu-book" + } + }, + { + "type": "doc", + "id": "developer/retain", + "label": "Retain", + "customProps": { + "icon": "lu-brain" + } + }, + { + "type": "doc", + "id": "developer/retrieval", + "label": "Recall", + "customProps": { + "icon": "lu-search" + } + }, + { + "type": "doc", + "id": "developer/reflect", + "label": "Reflect", + "customProps": { + "icon": "lu-message" + } + }, + { + "type": "doc", + "id": "developer/observations", + "label": "Observations", + "customProps": { + "icon": "lu-activity" + } + }, + { + "type": "doc", + "id": "developer/multilingual", + "label": "Multilingual", + "customProps": { + "icon": "lu-languages" + } + }, + { + "type": "doc", + "id": "developer/performance", + "label": "Performance", + "customProps": { + "icon": "lu-zap" + } + }, + { + "type": "doc", + "id": "developer/storage", + "label": "Storage", + "customProps": { + "icon": "lu-database" + } + }, + { + "type": "doc", + "id": "developer/rag-vs-hindsight", + "label": "RAG vs Memory", + "customProps": { + "icon": "lu-compare" + } + } ] }, { @@ -21,15 +84,87 @@ "label": "API", "collapsible": false, "items": [ - { "type": "doc", "id": "developer/api/quickstart", "label": "Quick Start", "customProps": { "icon": "lu-rocket" } }, - { "type": "doc", "id": "developer/api/retain", "label": "Retain", "customProps": { "icon": "lu-brain" } }, - { "type": "doc", "id": "developer/api/recall", "label": "Recall", "customProps": { "icon": "lu-search" } }, - { "type": "doc", "id": "developer/api/reflect", "label": "Reflect", "customProps": { "icon": "lu-message" } }, - { "type": "doc", "id": "developer/api/mental-models", "label": "Mental Models", "customProps": { "icon": "lu-layers" } }, - { "type": "doc", "id": "developer/api/memory-banks", "label": "Memory Banks", "customProps": { "icon": "lu-memory" } }, - { "type": "doc", "id": "developer/api/documents", "label": "Documents", "customProps": { "icon": "lu-file" } }, - { "type": "doc", "id": "developer/api/operations", "label": "Operations", "customProps": { "icon": "lu-cpu" } }, - { "type": "doc", "id": "developer/api/webhooks", "label": "Webhooks", "customProps": { "icon": "lu-webhook" } } + { + "type": "doc", + "id": "developer/api/quickstart", + "label": "Quick Start", + "customProps": { + "icon": "lu-rocket" + } + }, + { + "type": "doc", + "id": "developer/api/retain", + "label": "Retain", + "customProps": { + "icon": "lu-brain" + } + }, + { + "type": "doc", + "id": "developer/api/recall", + "label": "Recall", + "customProps": { + "icon": "lu-search" + } + }, + { + "type": "doc", + "id": "developer/api/reflect", + "label": "Reflect", + "customProps": { + "icon": "lu-message" + } + }, + { + "type": "doc", + "id": "developer/api/mental-models", + "label": "Mental Models", + "customProps": { + "icon": "lu-layers" + } + }, + { + "type": "doc", + "id": "developer/api/memory-banks", + "label": "Memory Banks", + "customProps": { + "icon": "lu-memory" + } + }, + { + "type": "doc", + "id": "developer/api/documents", + "label": "Documents", + "customProps": { + "icon": "lu-file" + } + }, + { + "type": "doc", + "id": "developer/api/operations", + "label": "Operations", + "customProps": { + "icon": "lu-cpu" + } + }, + { + "type": "doc", + "id": "developer/api/webhooks", + "label": "Webhooks", + "customProps": { + "icon": "lu-webhook" + } + }, + { + "type": "link", + "href": "/api-reference", + "label": "API Reference", + "customProps": { + "icon": "lu-book-open", + "iconAfter": "lu-arrow-up-right" + } + } ] }, { @@ -37,11 +172,46 @@ "label": "Clients", "collapsible": false, "items": [ - { "type": "doc", "id": "sdks/python", "label": "Python", "customProps": { "icon": "si-python" } }, - { "type": "doc", "id": "sdks/nodejs", "label": "TypeScript", "customProps": { "icon": "/img/icons/typescript.png" } }, - { "type": "doc", "id": "sdks/go", "label": "Go", "customProps": { "icon": "si-go" } }, - { "type": "doc", "id": "sdks/cli", "label": "CLI", "customProps": { "icon": "lu-terminal" } }, - { "type": "doc", "id": "sdks/embed", "label": "Embedded Python", "customProps": { "icon": "/img/icons/package.svg" } } + { + "type": "doc", + "id": "sdks/python", + "label": "Python", + "customProps": { + "icon": "si-python" + } + }, + { + "type": "doc", + "id": "sdks/nodejs", + "label": "TypeScript", + "customProps": { + "icon": "/img/icons/typescript.png" + } + }, + { + "type": "doc", + "id": "sdks/go", + "label": "Go", + "customProps": { + "icon": "si-go" + } + }, + { + "type": "doc", + "id": "sdks/cli", + "label": "CLI", + "customProps": { + "icon": "lu-terminal" + } + }, + { + "type": "doc", + "id": "sdks/embed", + "label": "Embedded Python", + "customProps": { + "icon": "/img/icons/package.svg" + } + } ] }, { @@ -49,14 +219,70 @@ "label": "Integrations", "collapsible": false, "items": [ - { "type": "doc", "id": "sdks/integrations/local-mcp", "label": "Local MCP Server", "customProps": { "icon": "/img/icons/mcp.png" } }, - { "type": "doc", "id": "sdks/integrations/litellm", "label": "LiteLLM", "customProps": { "icon": "/img/icons/litellm.png" } }, - { "type": "doc", "id": "sdks/integrations/openclaw", "label": "OpenClaw", "customProps": { "icon": "/img/icons/openclaw.png" } }, - { "type": "doc", "id": "sdks/integrations/ai-sdk", "label": "Vercel AI SDK", "customProps": { "icon": "/img/icons/vercel.png" } }, - { "type": "doc", "id": "sdks/integrations/chat", "label": "Vercel Chat SDK", "customProps": { "icon": "/img/icons/vercel.png" } }, - { "type": "doc", "id": "sdks/integrations/crewai", "label": "CrewAI", "customProps": { "icon": "/img/icons/crewai.png" } }, - { "type": "doc", "id": "sdks/integrations/pydantic-ai", "label": "Pydantic AI", "customProps": { "icon": "/img/icons/pydanticai.png" } }, - { "type": "doc", "id": "sdks/integrations/skills", "label": "Skills", "customProps": { "icon": "/img/icons/skills.png" } } + { + "type": "doc", + "id": "sdks/integrations/local-mcp", + "label": "Local MCP Server", + "customProps": { + "icon": "/img/icons/mcp.png" + } + }, + { + "type": "doc", + "id": "sdks/integrations/litellm", + "label": "LiteLLM", + "customProps": { + "icon": "/img/icons/litellm.png" + } + }, + { + "type": "doc", + "id": "sdks/integrations/openclaw", + "label": "OpenClaw", + "customProps": { + "icon": "/img/icons/openclaw.png" + } + }, + { + "type": "doc", + "id": "sdks/integrations/ai-sdk", + "label": "Vercel AI SDK", + "customProps": { + "icon": "/img/icons/vercel.png" + } + }, + { + "type": "doc", + "id": "sdks/integrations/chat", + "label": "Vercel Chat SDK", + "customProps": { + "icon": "/img/icons/vercel.png" + } + }, + { + "type": "doc", + "id": "sdks/integrations/crewai", + "label": "CrewAI", + "customProps": { + "icon": "/img/icons/crewai.png" + } + }, + { + "type": "doc", + "id": "sdks/integrations/pydantic-ai", + "label": "Pydantic AI", + "customProps": { + "icon": "/img/icons/pydanticai.png" + } + }, + { + "type": "doc", + "id": "sdks/integrations/skills", + "label": "Skills", + "customProps": { + "icon": "/img/icons/skills.png" + } + } ] }, { @@ -64,14 +290,149 @@ "label": "Hosting", "collapsible": false, "items": [ - { "type": "doc", "id": "developer/installation", "label": "Installation", "customProps": { "icon": "lu-package" } }, - { "type": "doc", "id": "developer/services", "label": "Services", "customProps": { "icon": "lu-server" } }, - { "type": "doc", "id": "developer/configuration", "label": "Configuration", "customProps": { "icon": "lu-settings" } }, - { "type": "doc", "id": "developer/admin-cli", "label": "Admin CLI", "customProps": { "icon": "lu-terminal" } }, - { "type": "doc", "id": "developer/extensions", "label": "Extensions", "customProps": { "icon": "lu-plug" } }, - { "type": "doc", "id": "developer/models", "label": "Models", "customProps": { "icon": "lu-cpu" } }, - { "type": "doc", "id": "developer/monitoring", "label": "Monitoring", "customProps": { "icon": "lu-activity" } }, - { "type": "doc", "id": "developer/mcp-server", "label": "MCP Server", "customProps": { "icon": "lu-network" } } + { + "type": "link", + "href": "https://ui.hindsight.vectorize.io/signup", + "label": "Cloud", + "customProps": { + "icon": "lu-cloud", + "iconAfter": "lu-arrow-up-right" + } + }, + { + "type": "doc", + "id": "developer/installation", + "label": "Installation", + "customProps": { + "icon": "lu-package" + } + }, + { + "type": "doc", + "id": "developer/services", + "label": "Services", + "customProps": { + "icon": "lu-server" + } + }, + { + "type": "doc", + "id": "developer/configuration", + "label": "Configuration", + "customProps": { + "icon": "lu-settings" + } + }, + { + "type": "doc", + "id": "developer/admin-cli", + "label": "Admin CLI", + "customProps": { + "icon": "lu-terminal" + } + }, + { + "type": "doc", + "id": "developer/extensions", + "label": "Extensions", + "customProps": { + "icon": "lu-plug" + } + }, + { + "type": "doc", + "id": "developer/models", + "label": "Models", + "customProps": { + "icon": "lu-cpu" + } + }, + { + "type": "doc", + "id": "developer/monitoring", + "label": "Monitoring", + "customProps": { + "icon": "lu-activity" + } + }, + { + "type": "doc", + "id": "developer/mcp-server", + "label": "MCP Server", + "customProps": { + "icon": "lu-network" + } + } + ] + }, + { + "type": "category", + "label": "More", + "collapsible": false, + "items": [ + { + "type": "link", + "href": "/cookbook", + "label": "Cookbook", + "customProps": { + "icon": "lu-book", + "iconAfter": "lu-arrow-up-right" + } + }, + { + "type": "link", + "href": "/blog", + "label": "Blog", + "customProps": { + "icon": "lu-rss", + "iconAfter": "lu-arrow-up-right" + } + }, + { + "type": "link", + "href": "https://join.slack.com/t/hindsight-space/shared_invite/zt-3nhbm4w29-LeSJ5Ixi6j8PdiYOCPlOgg", + "label": "Community", + "customProps": { + "icon": "si-slack", + "iconAfter": "lu-arrow-up-right" + } + }, + { + "type": "link", + "href": "https://github.com/vectorize-io/hindsight", + "label": "GitHub", + "customProps": { + "icon": "si-github", + "iconAfter": "lu-arrow-up-right" + } + }, + { + "type": "link", + "href": "https://benchmarks.hindsight.vectorize.io/", + "label": "Benchmarks", + "customProps": { + "icon": "lu-chart-bar", + "iconAfter": "lu-arrow-up-right" + } + }, + { + "type": "link", + "href": "https://benchmarks.hindsight.vectorize.io/", + "label": "Which Model Should I Use?", + "customProps": { + "icon": "lu-cpu", + "iconAfter": "lu-arrow-up-right" + } + }, + { + "type": "link", + "href": "https://arxiv.org/abs/2512.12818", + "label": "Paper", + "customProps": { + "icon": "lu-file-text", + "iconAfter": "lu-arrow-up-right" + } + } ] } ] diff --git a/skills/hindsight-docs/references/developer/api/quickstart.md b/skills/hindsight-docs/references/developer/api/quickstart.md index 97d9e5d2..dea9e3c8 100644 --- a/skills/hindsight-docs/references/developer/api/quickstart.md +++ b/skills/hindsight-docs/references/developer/api/quickstart.md @@ -5,6 +5,10 @@ Get up and running with Hindsight in 60 seconds. {/* Import raw source files */} +## Clients + + + ## Start the API Server ### pip (API only) @@ -113,6 +117,10 @@ hindsight memory reflect my-bank "Tell me about Alice" --- +## Integrations + + + ## Next Steps - [**Retain**](./retain) — Advanced options for storing memories diff --git a/skills/hindsight-docs/references/developer/api/recall.md b/skills/hindsight-docs/references/developer/api/recall.md index 91a5edc4..d7ec1027 100644 --- a/skills/hindsight-docs/references/developer/api/recall.md +++ b/skills/hindsight-docs/references/developer/api/recall.md @@ -333,6 +333,66 @@ Use this for strict scope enforcement where a memory must explicitly belong to * > **💡 Extra tags are fine** > A memory with tags `["user:alice", "team", "project:x"]` will still match a filter of `["user:alice", "team"]` under `all_strict` — extra tags on the memory are not a problem. The filter only requires the memory to contain **at least** the specified tags. +### tag_groups + +`tag_groups` is a list of compound boolean tag filters. The groups in the list are AND-ed together at the top level. Each group is a recursive boolean expression: a **leaf** node `{tags, match}`, or a **compound** node `{and: [...]}`, `{or: [...]}`, or `{not: ...}`. + +`tag_groups` and `tags` / `tags_match` can be used simultaneously — they are AND-ed together. + +#### Leaf node + +```json +{ "tags": ["step:5", "step:8"], "match": "any_strict" } +``` + +`match` accepts the same values as `tags_match`: `any`, `all`, `any_strict`, `all_strict`. Defaults to `any_strict`. + +#### Compound nodes + +```json +{ "and": [ , , ... ] } +{ "or": [ , , ... ] } +{ "not": } +``` + +#### Examples + +**Step filter AND user scope** — two top-level groups AND-ed: + +```json +{ + "tag_groups": [ + { "tags": ["step:5", "step:8", "step:12"], "match": "any_strict" }, + { "tags": ["user:ep_42"], "match": "all_strict" } + ] +} +``` + +**Nested OR inside AND** — user must match, plus either step OR priority: + +```json +{ + "tag_groups": [ + { "tags": ["user:alice"], "match": "all_strict" }, + { "or": [ + { "tags": ["step:5"], "match": "any_strict" }, + { "tags": ["priority:high"], "match": "all_strict" } + ]} + ] +} +``` + +**Exclusion** — user must match, but archived memories are excluded: + +```json +{ + "tag_groups": [ + { "tags": ["user:alice"], "match": "all_strict" }, + { "not": { "tags": ["archived"], "match": "any_strict" } } + ] +} +``` + ### trace When set to `true`, the response includes a detailed debug trace covering the query embedding, entry points, per-strategy retrieval results, RRF fusion candidates, reranked results, temporal constraints detected, and per-phase timings. Has no effect on the retrieval logic itself. Useful for understanding why specific memories were or were not returned. diff --git a/skills/hindsight-docs/references/developer/configuration.md b/skills/hindsight-docs/references/developer/configuration.md index 43f3368c..012746b7 100644 --- a/skills/hindsight-docs/references/developer/configuration.md +++ b/skills/hindsight-docs/references/developer/configuration.md @@ -160,7 +160,7 @@ To switch between backends: | Variable | Description | Default | |----------|-------------|---------| -| `HINDSIGHT_API_LLM_PROVIDER` | Provider: `openai`, `openai-codex`, `claude-code`, `anthropic`, `gemini`, `groq`, `ollama`, `lmstudio`, `vertexai` | `openai` | +| `HINDSIGHT_API_LLM_PROVIDER` | Provider: `openai`, `openai-codex`, `claude-code`, `anthropic`, `gemini`, `groq`, `minimax`, `ollama`, `lmstudio`, `vertexai` | `openai` | | `HINDSIGHT_API_LLM_API_KEY` | API key for LLM provider | - | | `HINDSIGHT_API_LLM_MODEL` | Model name | `gpt-5-mini` | | `HINDSIGHT_API_LLM_BASE_URL` | Custom LLM endpoint | Provider default | @@ -410,7 +410,7 @@ Supported OpenAI embedding dimensions: | Variable | Description | Default | |----------|-------------|---------| -| `HINDSIGHT_API_RERANKER_PROVIDER` | Provider: `local`, `tei`, `cohere`, `zeroentropy`, `flashrank`, `litellm`, `litellm-sdk`, or `rrf` | `local` | +| `HINDSIGHT_API_RERANKER_PROVIDER` | Provider: `local`, `tei`, `cohere`, `zeroentropy`, `flashrank`, `litellm`, `litellm-sdk`, `jina-mlx`, or `rrf` | `local` | | `HINDSIGHT_API_RERANKER_LOCAL_MODEL` | Model for local provider | `cross-encoder/ms-marco-MiniLM-L-6-v2` | | `HINDSIGHT_API_RERANKER_LOCAL_MAX_CONCURRENT` | Max concurrent local reranking (prevents CPU thrashing under load) | `4` | | `HINDSIGHT_API_RERANKER_LOCAL_TRUST_REMOTE_CODE` | Allow loading models with custom code (security risk, disabled by default) | `false` | @@ -426,10 +426,12 @@ Supported OpenAI embedding dimensions: | `HINDSIGHT_API_RERANKER_LITELLM_SDK_API_KEY` | LiteLLM **SDK** API key for direct reranking (no proxy needed) | - | | `HINDSIGHT_API_RERANKER_LITELLM_SDK_MODEL` | LiteLLM SDK rerank model (e.g., `deepinfra/Qwen3-reranker-8B`) | `cohere/rerank-english-v3.0` | | `HINDSIGHT_API_RERANKER_LITELLM_SDK_API_BASE` | Custom API base URL for LiteLLM SDK (optional) | - | +| `HINDSIGHT_API_RERANKER_LITELLM_MAX_TOKENS_PER_DOC` | Truncate documents to this many tokens before sending to the reranker (applies to both `litellm` and `litellm-sdk`). Use for models with small context windows (e.g. set to `900` for a 1024-token limit model). Unset by default (no truncation). | - | | `HINDSIGHT_API_RERANKER_ZEROENTROPY_API_KEY` | ZeroEntropy API key for reranking | - | | `HINDSIGHT_API_RERANKER_ZEROENTROPY_MODEL` | ZeroEntropy rerank model (`zerank-2`, `zerank-2-small`) | `zerank-2` | | `HINDSIGHT_API_RERANKER_FLASHRANK_MODEL` | FlashRank model for fast CPU-based reranking | `ms-marco-MiniLM-L-12-v2` | | `HINDSIGHT_API_RERANKER_FLASHRANK_CACHE_DIR` | Cache directory for FlashRank models | System default | +| `HINDSIGHT_API_RERANKER_JINA_MLX_MODEL_PATH` | Local path to downloaded `jina-reranker-v3-mlx` model (auto-downloads from HuggingFace if unset) | - | ```bash # Local (default) - uses SentenceTransformers CrossEncoder @@ -472,6 +474,10 @@ export HINDSIGHT_API_RERANKER_LITELLM_MODEL=cohere/rerank-english-v3.0 # or voy export HINDSIGHT_API_RERANKER_PROVIDER=litellm-sdk export HINDSIGHT_API_RERANKER_LITELLM_SDK_API_KEY=your-deepinfra-api-key export HINDSIGHT_API_RERANKER_LITELLM_SDK_MODEL=deepinfra/Qwen3-reranker-8B # or cohere/rerank-english-v3.0, etc. + +# Jina MLX - Apple Silicon native reranking (no GPU/cloud required) +# Model (~1.2 GB) is downloaded automatically from HuggingFace Hub on first use. +export HINDSIGHT_API_RERANKER_PROVIDER=jina-mlx ``` #### LiteLLM Proxy vs SDK @@ -488,6 +494,14 @@ Both support the same providers: - **Jina AI** (`jina_ai/jina-reranker-v2`) - **AWS Bedrock** (`bedrock/...`) +#### Jina MLX (Apple Silicon) + +The `jina-mlx` provider uses [`jinaai/jina-reranker-v3-mlx`](https://huggingface.co/jinaai/jina-reranker-v3-mlx), optimized for Apple Silicon. The model (~1.2 GB) is downloaded from HuggingFace Hub automatically on first startup and cached locally. + +:::note License +`jina-reranker-v3-mlx` is licensed under CC BY-NC 4.0. Contact Jina AI for commercial usage. +::: + ### Authentication By default, Hindsight runs without authentication. For production deployments, enable API key authentication using the built-in tenant extension: @@ -530,6 +544,7 @@ For advanced authentication (JWT, OAuth, multi-tenant schemas), implement a cust | `HINDSIGHT_API_GRAPH_RETRIEVER` | Graph retrieval algorithm: `link_expansion`, `mpfp`, or `bfs` | `link_expansion` | | `HINDSIGHT_API_RECALL_MAX_CONCURRENT` | Max concurrent recall operations per worker (backpressure) | `32` | | `HINDSIGHT_API_RECALL_CONNECTION_BUDGET` | Max concurrent DB connections per recall operation | `4` | +| `HINDSIGHT_API_RECALL_MAX_QUERY_TOKENS` | Maximum token length of a recall query; requests exceeding this limit are rejected with HTTP 400 | `500` | | `HINDSIGHT_API_RERANKER_MAX_CANDIDATES` | Max candidates to rerank per recall (RRF pre-filters the rest) | `300` | | `HINDSIGHT_API_MPFP_TOP_K_NEIGHBORS` | Fan-out limit per node in MPFP graph traversal | `20` | | `HINDSIGHT_API_MENTAL_MODEL_REFRESH_CONCURRENCY` | Max concurrent mental model refreshes | `8` | diff --git a/skills/hindsight-docs/references/developer/index.md b/skills/hindsight-docs/references/developer/index.md index 8c25e95b..9191df56 100644 --- a/skills/hindsight-docs/references/developer/index.md +++ b/skills/hindsight-docs/references/developer/index.md @@ -1,7 +1,4 @@ ---- -sidebar_position: 1 -slug: / ---- + # Overview @@ -114,6 +111,14 @@ The **mission** tells Hindsight what knowledge to prioritize and provides contex These settings only affect the `reflect` operation, not `recall`. +## Clients & Languages + + + +## Integrations + + + ## Next Steps ### Getting Started diff --git a/skills/hindsight-docs/references/developer/installation.md b/skills/hindsight-docs/references/developer/installation.md index e30ba8be..d5ce0c3c 100644 --- a/skills/hindsight-docs/references/developer/installation.md +++ b/skills/hindsight-docs/references/developer/installation.md @@ -8,28 +8,28 @@ Hindsight can be deployed in several ways depending on your infrastructure and r ## Prerequisites -### PostgreSQL with pgvector +### PostgreSQL -Hindsight requires PostgreSQL with the **pgvector** extension for vector similarity search. +Hindsight requires PostgreSQL 14+ with a vector extension for similarity search. The supported extensions are: + +- **pgvector** (default) +- **pgvectorscale** +- **vchord** + +Configure which one to use with `HINDSIGHT_API_VECTOR_EXTENSION`. See [Configuration](./configuration) for details. **By default**, Hindsight uses **pg0** — an embedded PostgreSQL that runs locally on your machine. This is convenient for development but **not recommended for production**. -**For production**, use an external PostgreSQL with pgvector: +**For production**, use an external PostgreSQL with one of the supported vector extensions: - **Supabase** — Managed PostgreSQL with pgvector built-in - **Neon** — Serverless PostgreSQL with pgvector -- **Azure Database for PostgreSQL** — With pgvector and pg_diskann (DiskANN) support +- **Azure Database for PostgreSQL** — With pgvector and pgvectorscale support - **AWS RDS** / **Cloud SQL** — With pgvector extension enabled -- **Self-hosted** — PostgreSQL 14+ with pgvector installed +- **Self-hosted** — PostgreSQL 14+ with your preferred vector extension ### LLM Provider -You need an LLM API key for fact extraction, entity resolution, and answer generation: - -- **Groq** (recommended): Fast inference with `gpt-oss-20b` -- **OpenAI**: GPT-4o, GPT-4o-mini -- **Ollama**: Run models locally - -See [Models](./models) for detailed comparison and configuration. +You need an LLM API key for fact extraction, entity resolution, and answer generation. See [Models](./models) for supported providers, model recommendations, and configuration. --- @@ -53,61 +53,12 @@ docker run --rm -it --pull always -p 8888:8888 -p 9999:9999 \ ### Docker Image Variants -Hindsight provides two image variants with different size/capability tradeoffs: +| Variant | Size (AMD64) | Size (ARM64) | When to use | +|---------|--------------|--------------|-------------| +| **Full** (`latest`) | ~9 GB | ~3.7 GB | Default. Works out of the box with no external services except the LLM. | +| **Slim** (`slim`) | ~500 MB | ~500 MB | Use when you already rely on external services for embeddings and reranking (OpenAI, Cohere, TEI). Significantly smaller image, faster deploys. Requires [external providers](./configuration#embeddings). | -| Variant | Size (AMD64) | Size (ARM64) | Use Case | -|---------|--------------|--------------|----------| -| **Full** (`latest`) | ~9 GB | ~3.7 GB | Includes local ML models (embeddings, reranking) | -| **Slim** (`slim`) | ~500 MB | ~500 MB | Requires external embedding/reranking providers | - -**Full image** (default): -```bash -docker run --rm -it -p 8888:8888 \ - -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \ - ghcr.io/vectorize-io/hindsight:latest -``` -- ✅ Works out of the box with local ML models -- ✅ No additional services needed -- ❌ Larger image size (AMD64 includes CUDA libraries for GPU support) - -**Slim image**: -```bash -docker run --rm -it -p 8888:8888 \ - -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \ - -e HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai \ - -e HINDSIGHT_API_RERANKER_PROVIDER=cohere \ - -e HINDSIGHT_API_COHERE_API_KEY=$COHERE_API_KEY \ - ghcr.io/vectorize-io/hindsight:latest-slim -``` -- ✅ Dramatically smaller image (~95% reduction on AMD64) -- ✅ Faster pull/deploy times -- ✅ Lower memory footprint -- ❌ Requires external embedding/reranking services (OpenAI, Cohere, TEI) - -**When to use slim:** -- Cloud deployments where image size matters -- Using managed embedding services (OpenAI, Cohere) -- Running on Text Embeddings Inference (TEI) infrastructure -- Kubernetes environments with fast pull requirements - -:::warning Slim Image Requires External Providers -If you run the slim image **without** setting external embedding providers, you'll see this error: - -``` -ImportError: sentence-transformers is required for LocalSTEmbeddings. -Install it with: pip install sentence-transformers -``` - -**Fix:** Always set embedding and reranking providers when using slim images: -```bash --e HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai --e HINDSIGHT_API_EMBEDDINGS_OPENAI_API_KEY=sk-xxx --e HINDSIGHT_API_RERANKER_PROVIDER=cohere --e HINDSIGHT_API_COHERE_API_KEY=xxx -``` -::: - -See [Configuration](./configuration#embeddings) for all embedding provider options. +The slim image corresponds to the [`hindsight-api-slim`](#package-variants) pip package. See [Configuration](./configuration#embeddings) for external provider options. ### Available Tags @@ -120,7 +71,7 @@ ghcr.io/vectorize-io/hindsight:0.4.9-slim # Slim, specific version # API only ghcr.io/vectorize-io/hindsight-api:latest -ghcr.io/vectorize-io/hindsight-api:slim +ghcr.io/vectorize-io/hindsight-api:latest-slim # Control Plane only ghcr.io/vectorize-io/hindsight-control-plane:latest @@ -175,14 +126,17 @@ See the [Helm chart values.yaml](https://github.com/vectorize-io/hindsight/tree/ ## Bare Metal (pip) -**Best for**: Custom deployments, integration into existing Python applications +**Best for**: Running Hindsight as a standalone service on a host machine. ### Install ```bash -pip install hindsight-all +pip install hindsight-api # Full — works out of the box +pip install hindsight-api-slim # Slim — requires external services for embeddings, reranking, and the database ``` +When using `hindsight-api-slim`, you must configure external providers for all model operations. See [Configuration](./configuration#embeddings) for details. + ### Run with Embedded Database For development and testing, Hindsight can run with an embedded PostgreSQL (pg0): @@ -253,6 +207,42 @@ PORT=80 HINDSIGHT_CP_DATAPLANE_API_URL=https://api.hindsight.io npx @vectorize-i --- +## Embedded in a Python Application + +**Best for**: Using Hindsight programmatically from Python without running a separate server process. + +```bash +pip install hindsight-all # Full — works out of the box +pip install hindsight-all-slim # Slim — requires external services for embeddings, reranking, and the database +``` + +`hindsight-all` supports two modes of embedding: + +**In-process** (`HindsightServer`): the server runs in a background thread inside your application. Best when you want the tightest integration and are already managing your own process lifecycle. + +```python +from hindsight import HindsightServer, HindsightClient + +with HindsightServer(llm_provider="openai", llm_api_key="sk-xxx") as server: + client = HindsightClient(base_url=server.url) + client.retain(bank_id="alice", content="Alice prefers concise answers.") + results = client.recall(bank_id="alice", query="How should I respond to Alice?") +``` + +**Managed subprocess** (`HindsightEmbedded`): the server runs as a background daemon process, shared across multiple Python processes or sessions. The daemon starts on first use and shuts down automatically after an idle timeout. + +```python +from hindsight import HindsightEmbedded + +client = HindsightEmbedded(llm_provider="openai", llm_api_key="sk-xxx") +client.retain(bank_id="alice", content="Alice prefers concise answers.") +results = client.recall(bank_id="alice", query="How should I respond to Alice?") +``` + +See the [Python SDK](../sdks/python.md) for the full API reference. + +--- + ## Next Steps - [Configuration](./configuration.md) — Environment variables and settings diff --git a/skills/hindsight-docs/references/developer/models.md b/skills/hindsight-docs/references/developer/models.md index 7b2c817d..8680de45 100644 --- a/skills/hindsight-docs/references/developer/models.md +++ b/skills/hindsight-docs/references/developer/models.md @@ -1,16 +1,15 @@ + # Models Hindsight uses several machine learning models for different tasks. ## Overview -| Model Type | Purpose | Default | Configurable | -|------------|---------|---------|--------------| -| **LLM** | Fact extraction, reasoning, generation | Provider-specific | Yes | -| **Embedding** | Vector representations for semantic search | `BAAI/bge-small-en-v1.5` | Yes | -| **Cross-Encoder** | Reranking search results | `cross-encoder/ms-marco-MiniLM-L-6-v2` | Yes | +- **LLM** — Fact extraction, reasoning, and generation. Provider-specific, fully configurable. +- **Embedding** — Vector representations for semantic search. Default: `BAAI/bge-small-en-v1.5`. +- **Cross-Encoder** — Reranking search results. Default: `cross-encoder/ms-marco-MiniLM-L-6-v2`. -All local models (embedding, cross-encoder) are automatically downloaded from HuggingFace on first run. +Embedding and cross-encoder models are downloaded automatically from HuggingFace on first run. --- @@ -18,14 +17,17 @@ All local models (embedding, cross-encoder) are automatically downloaded from Hu Used for fact extraction, entity resolution, mental model consolidation, and answer synthesis. -**Supported providers:** OpenAI, Anthropic, Gemini, Groq, Ollama, LM Studio, and **any OpenAI-compatible API** +**Supported providers:** -:::tip OpenAI-Compatible Providers + + +Also supports **any OpenAI-compatible API** (e.g., Azure OpenAI, Together AI, Fireworks). + +> **💡 OpenAI-Compatible Providers** +> Hindsight works with any provider that exposes an OpenAI-compatible API (e.g., Azure OpenAI). Simply set `HINDSIGHT_API_LLM_PROVIDER=openai` and configure `HINDSIGHT_API_LLM_BASE_URL` to point to your provider's endpoint. See [Configuration](./configuration#llm-provider) for setup examples. -::: - ### Benchmarks Not sure which model to use? The **[Model Leaderboard](https://benchmarks.hindsight.vectorize.io/)** benchmarks models across accuracy, speed, cost, and reliability for retain, reflect, and observation consolidation so you can pick the right trade-off for your use case. @@ -63,6 +65,7 @@ Each provider has a recommended default model that's used when `HINDSIGHT_API_LL | `anthropic` | `claude-haiku-4-5-20251001` | | `gemini` | `gemini-2.5-flash` | | `groq` | `openai/gpt-oss-120b` | +| `minimax` | `MiniMax-M2.5` | | `ollama` | `gemma3:12b` | | `lmstudio` | `local-model` | | `vertexai` | `gemini-2.0-flash-001` | @@ -97,7 +100,8 @@ export HINDSIGHT_API_RETAIN_LLM_PROVIDER=anthropic Other LLM models not listed above may work with Hindsight, but they must support **at least 65,000 output tokens** to ensure reliable fact extraction. If you need support for a specific model that doesn't meet this requirement, please [open an issue](https://github.com/hindsight-ai/hindsight/issues) to request an exception. -:::tip Models with Limited Output Tokens +> **💡 Models with Limited Output Tokens** +> If your model only supports 32k or fewer output tokens (e.g., some older models), you can reduce the retain completion token limit: ```bash @@ -109,8 +113,6 @@ export HINDSIGHT_API_RETAIN_MAX_COMPLETION_TOKENS=16000 ``` **Important:** `HINDSIGHT_API_RETAIN_MAX_COMPLETION_TOKENS` must be greater than `HINDSIGHT_API_RETAIN_CHUNK_SIZE` (default: 3000). The system will validate this on startup and provide an error message if the configuration is invalid. -::: - ### Configuration ```bash @@ -144,6 +146,11 @@ export HINDSIGHT_API_LLM_PROVIDER=lmstudio export HINDSIGHT_API_LLM_BASE_URL=http://localhost:1234/v1 export HINDSIGHT_API_LLM_MODEL=your-local-model +# MiniMax (204K context window) +export HINDSIGHT_API_LLM_PROVIDER=minimax +export HINDSIGHT_API_LLM_API_KEY=your-minimax-api-key +export HINDSIGHT_API_LLM_MODEL=MiniMax-M2.5 + # Vertex AI (Google Cloud) export HINDSIGHT_API_LLM_PROVIDER=vertexai export HINDSIGHT_API_LLM_MODEL=gemini-2.0-flash-001 @@ -210,8 +217,8 @@ You can use any model supported by OpenAI Codex CLI Use your Claude Pro or Max subscription for Hindsight without separate Anthropic API costs. - -:::warning Terms of Service Notice +> **⚠️ Terms of Service Notice** +> This integration uses the Claude Agent SDK with your personal Claude Pro/Max subscription credentials. You must be logged into Claude Code on your own machine before using this provider. @@ -235,9 +242,6 @@ credentials. You must be logged into Claude Code on your own machine before usin For production or team use, we recommend using `HINDSIGHT_API_LLM_PROVIDER=anthropic` with an API key from the [Anthropic Console](https://console.anthropic.com/). -::: - - **Prerequisites:** - Active Claude Pro or Max subscription - Claude Code CLI installed @@ -282,7 +286,6 @@ You can use any model supported by Claude Code CLI. - Usage billed to your Claude subscription (not separate API costs) - For personal development use only (see Claude Terms of Service) - --- ### Vertex AI Setup (Google Cloud) @@ -376,10 +379,9 @@ Converts text into dense vector representations for semantic similarity search. | `embed-english-v3.0` | 1024 | English text | | `embed-multilingual-v3.0` | 1024 | 100+ languages | -:::warning Embedding Dimensions +> **⚠️ Embedding Dimensions** +> Hindsight automatically detects the embedding dimension at startup and adjusts the database schema. Once memories are stored, you cannot change dimensions without losing data. -::: - **Configuration Examples:** ```bash diff --git a/skills/hindsight-docs/references/sdks/embed.md b/skills/hindsight-docs/references/sdks/embed.md index f94dd8e1..7b1ba402 100644 --- a/skills/hindsight-docs/references/sdks/embed.md +++ b/skills/hindsight-docs/references/sdks/embed.md @@ -88,7 +88,7 @@ The daemon starts automatically on first use! | Variable | Description | Default | |----------|-------------|---------| | `HINDSIGHT_EMBED_LLM_API_KEY` | **Required**. API key for LLM provider | - | -| `HINDSIGHT_EMBED_LLM_PROVIDER` | LLM provider: `openai`, `anthropic`, `gemini`, `groq`, `ollama` | `openai` | +| `HINDSIGHT_EMBED_LLM_PROVIDER` | LLM provider: `openai`, `anthropic`, `gemini`, `groq`, `minimax`, `ollama` | `openai` | | `HINDSIGHT_EMBED_LLM_MODEL` | Model name | `gpt-4o-mini` | | `HINDSIGHT_EMBED_BANK_ID` | Default memory bank ID | `default` | | `HINDSIGHT_EMBED_DAEMON_IDLE_TIMEOUT` | Seconds before daemon auto-exits when idle (0 = never) | `300` | diff --git a/skills/hindsight-docs/references/sdks/nodejs.md b/skills/hindsight-docs/references/sdks/nodejs.md index 38d87132..767125eb 100644 --- a/skills/hindsight-docs/references/sdks/nodejs.md +++ b/skills/hindsight-docs/references/sdks/nodejs.md @@ -2,7 +2,7 @@ sidebar_position: 2 --- -# Node.js Client +# TypeScript Client Official TypeScript/JavaScript client for the Hindsight API.