diff --git a/hindsight-docs/docs/developer/installation.md b/hindsight-docs/docs/developer/installation.md
index d5ce0c3c..796f599d 100644
--- a/hindsight-docs/docs/developer/installation.md
+++ b/hindsight-docs/docs/developer/installation.md
@@ -246,5 +246,5 @@ See the [Python SDK](../sdks/python.md) for the full API reference.
## Next Steps
- [Configuration](./configuration.md) — Environment variables and settings
-- [Models](./models.md) — ML models and providers
+- [Models](./models.mdx) — ML models and providers
- [Monitoring](./monitoring.md) — Metrics and observability
diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/api/quickstart.mdx b/hindsight-docs/versioned_docs/version-0.4/developer/api/quickstart.mdx
index 65bacef2..2f6bab98 100644
--- a/hindsight-docs/versioned_docs/version-0.4/developer/api/quickstart.mdx
+++ b/hindsight-docs/versioned_docs/version-0.4/developer/api/quickstart.mdx
@@ -9,12 +9,17 @@ Get up and running with Hindsight in 60 seconds.
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
import CodeSnippet from '@site/src/components/CodeSnippet';
+import {ClientsGrid, IntegrationsGrid} from '@site/src/components/SupportedGrids';
{/* Import raw source files */}
import quickstartPy from '!!raw-loader!@site/examples/api/quickstart.py';
import quickstartMjs from '!!raw-loader!@site/examples/api/quickstart.mjs';
import quickstartSh from '!!raw-loader!@site/examples/api/quickstart.sh';
+## Clients
+
+
+
## Start the API Server
@@ -100,6 +105,10 @@ curl -fsSL https://hindsight.vectorize.io/get-cli | bash
---
+## Integrations
+
+
+
## Next Steps
- [**Retain**](./retain) — Advanced options for storing memories
diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/api/recall.mdx b/hindsight-docs/versioned_docs/version-0.4/developer/api/recall.mdx
index a6ef4f4b..8fee9e41 100644
--- a/hindsight-docs/versioned_docs/version-0.4/developer/api/recall.mdx
+++ b/hindsight-docs/versioned_docs/version-0.4/developer/api/recall.mdx
@@ -184,6 +184,66 @@ Use this for strict scope enforcement where a memory must explicitly belong to *
A memory with tags `["user:alice", "team", "project:x"]` will still match a filter of `["user:alice", "team"]` under `all_strict` — extra tags on the memory are not a problem. The filter only requires the memory to contain **at least** the specified tags.
:::
+### tag_groups
+
+`tag_groups` is a list of compound boolean tag filters. The groups in the list are AND-ed together at the top level. Each group is a recursive boolean expression: a **leaf** node `{tags, match}`, or a **compound** node `{and: [...]}`, `{or: [...]}`, or `{not: ...}`.
+
+`tag_groups` and `tags` / `tags_match` can be used simultaneously — they are AND-ed together.
+
+#### Leaf node
+
+```json
+{ "tags": ["step:5", "step:8"], "match": "any_strict" }
+```
+
+`match` accepts the same values as `tags_match`: `any`, `all`, `any_strict`, `all_strict`. Defaults to `any_strict`.
+
+#### Compound nodes
+
+```json
+{ "and": [ , , ... ] }
+{ "or": [ , , ... ] }
+{ "not": }
+```
+
+#### Examples
+
+**Step filter AND user scope** — two top-level groups AND-ed:
+
+```json
+{
+ "tag_groups": [
+ { "tags": ["step:5", "step:8", "step:12"], "match": "any_strict" },
+ { "tags": ["user:ep_42"], "match": "all_strict" }
+ ]
+}
+```
+
+**Nested OR inside AND** — user must match, plus either step OR priority:
+
+```json
+{
+ "tag_groups": [
+ { "tags": ["user:alice"], "match": "all_strict" },
+ { "or": [
+ { "tags": ["step:5"], "match": "any_strict" },
+ { "tags": ["priority:high"], "match": "all_strict" }
+ ]}
+ ]
+}
+```
+
+**Exclusion** — user must match, but archived memories are excluded:
+
+```json
+{
+ "tag_groups": [
+ { "tags": ["user:alice"], "match": "all_strict" },
+ { "not": { "tags": ["archived"], "match": "any_strict" } }
+ ]
+}
+```
+
### trace
When set to `true`, the response includes a detailed debug trace covering the query embedding, entry points, per-strategy retrieval results, RRF fusion candidates, reranked results, temporal constraints detected, and per-phase timings. Has no effect on the retrieval logic itself. Useful for understanding why specific memories were or were not returned.
diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/configuration.md b/hindsight-docs/versioned_docs/version-0.4/developer/configuration.md
index 43f3368c..012746b7 100644
--- a/hindsight-docs/versioned_docs/version-0.4/developer/configuration.md
+++ b/hindsight-docs/versioned_docs/version-0.4/developer/configuration.md
@@ -160,7 +160,7 @@ To switch between backends:
| Variable | Description | Default |
|----------|-------------|---------|
-| `HINDSIGHT_API_LLM_PROVIDER` | Provider: `openai`, `openai-codex`, `claude-code`, `anthropic`, `gemini`, `groq`, `ollama`, `lmstudio`, `vertexai` | `openai` |
+| `HINDSIGHT_API_LLM_PROVIDER` | Provider: `openai`, `openai-codex`, `claude-code`, `anthropic`, `gemini`, `groq`, `minimax`, `ollama`, `lmstudio`, `vertexai` | `openai` |
| `HINDSIGHT_API_LLM_API_KEY` | API key for LLM provider | - |
| `HINDSIGHT_API_LLM_MODEL` | Model name | `gpt-5-mini` |
| `HINDSIGHT_API_LLM_BASE_URL` | Custom LLM endpoint | Provider default |
@@ -410,7 +410,7 @@ Supported OpenAI embedding dimensions:
| Variable | Description | Default |
|----------|-------------|---------|
-| `HINDSIGHT_API_RERANKER_PROVIDER` | Provider: `local`, `tei`, `cohere`, `zeroentropy`, `flashrank`, `litellm`, `litellm-sdk`, or `rrf` | `local` |
+| `HINDSIGHT_API_RERANKER_PROVIDER` | Provider: `local`, `tei`, `cohere`, `zeroentropy`, `flashrank`, `litellm`, `litellm-sdk`, `jina-mlx`, or `rrf` | `local` |
| `HINDSIGHT_API_RERANKER_LOCAL_MODEL` | Model for local provider | `cross-encoder/ms-marco-MiniLM-L-6-v2` |
| `HINDSIGHT_API_RERANKER_LOCAL_MAX_CONCURRENT` | Max concurrent local reranking (prevents CPU thrashing under load) | `4` |
| `HINDSIGHT_API_RERANKER_LOCAL_TRUST_REMOTE_CODE` | Allow loading models with custom code (security risk, disabled by default) | `false` |
@@ -426,10 +426,12 @@ Supported OpenAI embedding dimensions:
| `HINDSIGHT_API_RERANKER_LITELLM_SDK_API_KEY` | LiteLLM **SDK** API key for direct reranking (no proxy needed) | - |
| `HINDSIGHT_API_RERANKER_LITELLM_SDK_MODEL` | LiteLLM SDK rerank model (e.g., `deepinfra/Qwen3-reranker-8B`) | `cohere/rerank-english-v3.0` |
| `HINDSIGHT_API_RERANKER_LITELLM_SDK_API_BASE` | Custom API base URL for LiteLLM SDK (optional) | - |
+| `HINDSIGHT_API_RERANKER_LITELLM_MAX_TOKENS_PER_DOC` | Truncate documents to this many tokens before sending to the reranker (applies to both `litellm` and `litellm-sdk`). Use for models with small context windows (e.g. set to `900` for a 1024-token limit model). Unset by default (no truncation). | - |
| `HINDSIGHT_API_RERANKER_ZEROENTROPY_API_KEY` | ZeroEntropy API key for reranking | - |
| `HINDSIGHT_API_RERANKER_ZEROENTROPY_MODEL` | ZeroEntropy rerank model (`zerank-2`, `zerank-2-small`) | `zerank-2` |
| `HINDSIGHT_API_RERANKER_FLASHRANK_MODEL` | FlashRank model for fast CPU-based reranking | `ms-marco-MiniLM-L-12-v2` |
| `HINDSIGHT_API_RERANKER_FLASHRANK_CACHE_DIR` | Cache directory for FlashRank models | System default |
+| `HINDSIGHT_API_RERANKER_JINA_MLX_MODEL_PATH` | Local path to downloaded `jina-reranker-v3-mlx` model (auto-downloads from HuggingFace if unset) | - |
```bash
# Local (default) - uses SentenceTransformers CrossEncoder
@@ -472,6 +474,10 @@ export HINDSIGHT_API_RERANKER_LITELLM_MODEL=cohere/rerank-english-v3.0 # or voy
export HINDSIGHT_API_RERANKER_PROVIDER=litellm-sdk
export HINDSIGHT_API_RERANKER_LITELLM_SDK_API_KEY=your-deepinfra-api-key
export HINDSIGHT_API_RERANKER_LITELLM_SDK_MODEL=deepinfra/Qwen3-reranker-8B # or cohere/rerank-english-v3.0, etc.
+
+# Jina MLX - Apple Silicon native reranking (no GPU/cloud required)
+# Model (~1.2 GB) is downloaded automatically from HuggingFace Hub on first use.
+export HINDSIGHT_API_RERANKER_PROVIDER=jina-mlx
```
#### LiteLLM Proxy vs SDK
@@ -488,6 +494,14 @@ Both support the same providers:
- **Jina AI** (`jina_ai/jina-reranker-v2`)
- **AWS Bedrock** (`bedrock/...`)
+#### Jina MLX (Apple Silicon)
+
+The `jina-mlx` provider uses [`jinaai/jina-reranker-v3-mlx`](https://huggingface.co/jinaai/jina-reranker-v3-mlx), optimized for Apple Silicon. The model (~1.2 GB) is downloaded from HuggingFace Hub automatically on first startup and cached locally.
+
+:::note License
+`jina-reranker-v3-mlx` is licensed under CC BY-NC 4.0. Contact Jina AI for commercial usage.
+:::
+
### Authentication
By default, Hindsight runs without authentication. For production deployments, enable API key authentication using the built-in tenant extension:
@@ -530,6 +544,7 @@ For advanced authentication (JWT, OAuth, multi-tenant schemas), implement a cust
| `HINDSIGHT_API_GRAPH_RETRIEVER` | Graph retrieval algorithm: `link_expansion`, `mpfp`, or `bfs` | `link_expansion` |
| `HINDSIGHT_API_RECALL_MAX_CONCURRENT` | Max concurrent recall operations per worker (backpressure) | `32` |
| `HINDSIGHT_API_RECALL_CONNECTION_BUDGET` | Max concurrent DB connections per recall operation | `4` |
+| `HINDSIGHT_API_RECALL_MAX_QUERY_TOKENS` | Maximum token length of a recall query; requests exceeding this limit are rejected with HTTP 400 | `500` |
| `HINDSIGHT_API_RERANKER_MAX_CANDIDATES` | Max candidates to rerank per recall (RRF pre-filters the rest) | `300` |
| `HINDSIGHT_API_MPFP_TOP_K_NEIGHBORS` | Fan-out limit per node in MPFP graph traversal | `20` |
| `HINDSIGHT_API_MENTAL_MODEL_REFRESH_CONCURRENCY` | Max concurrent mental model refreshes | `8` |
diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/index.md b/hindsight-docs/versioned_docs/version-0.4/developer/index.mdx
similarity index 97%
rename from hindsight-docs/versioned_docs/version-0.4/developer/index.md
rename to hindsight-docs/versioned_docs/version-0.4/developer/index.mdx
index 8c25e95b..22a712d2 100644
--- a/hindsight-docs/versioned_docs/version-0.4/developer/index.md
+++ b/hindsight-docs/versioned_docs/version-0.4/developer/index.mdx
@@ -3,6 +3,8 @@ sidebar_position: 1
slug: /
---
+import {ClientsGrid, IntegrationsGrid} from '@site/src/components/SupportedGrids';
+
# Overview
## Why Hindsight?
@@ -114,6 +116,14 @@ The **mission** tells Hindsight what knowledge to prioritize and provides contex
These settings only affect the `reflect` operation, not `recall`.
+## Clients & Languages
+
+
+
+## Integrations
+
+
+
## Next Steps
### Getting Started
diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/installation.md b/hindsight-docs/versioned_docs/version-0.4/developer/installation.md
index e30ba8be..28a9e5c1 100644
--- a/hindsight-docs/versioned_docs/version-0.4/developer/installation.md
+++ b/hindsight-docs/versioned_docs/version-0.4/developer/installation.md
@@ -8,28 +8,28 @@ Hindsight can be deployed in several ways depending on your infrastructure and r
## Prerequisites
-### PostgreSQL with pgvector
+### PostgreSQL
-Hindsight requires PostgreSQL with the **pgvector** extension for vector similarity search.
+Hindsight requires PostgreSQL 14+ with a vector extension for similarity search. The supported extensions are:
+
+- **pgvector** (default)
+- **pgvectorscale**
+- **vchord**
+
+Configure which one to use with `HINDSIGHT_API_VECTOR_EXTENSION`. See [Configuration](./configuration) for details.
**By default**, Hindsight uses **pg0** — an embedded PostgreSQL that runs locally on your machine. This is convenient for development but **not recommended for production**.
-**For production**, use an external PostgreSQL with pgvector:
+**For production**, use an external PostgreSQL with one of the supported vector extensions:
- **Supabase** — Managed PostgreSQL with pgvector built-in
- **Neon** — Serverless PostgreSQL with pgvector
-- **Azure Database for PostgreSQL** — With pgvector and pg_diskann (DiskANN) support
+- **Azure Database for PostgreSQL** — With pgvector and pgvectorscale support
- **AWS RDS** / **Cloud SQL** — With pgvector extension enabled
-- **Self-hosted** — PostgreSQL 14+ with pgvector installed
+- **Self-hosted** — PostgreSQL 14+ with your preferred vector extension
### LLM Provider
-You need an LLM API key for fact extraction, entity resolution, and answer generation:
-
-- **Groq** (recommended): Fast inference with `gpt-oss-20b`
-- **OpenAI**: GPT-4o, GPT-4o-mini
-- **Ollama**: Run models locally
-
-See [Models](./models) for detailed comparison and configuration.
+You need an LLM API key for fact extraction, entity resolution, and answer generation. See [Models](./models) for supported providers, model recommendations, and configuration.
---
@@ -53,61 +53,12 @@ docker run --rm -it --pull always -p 8888:8888 -p 9999:9999 \
### Docker Image Variants
-Hindsight provides two image variants with different size/capability tradeoffs:
+| Variant | Size (AMD64) | Size (ARM64) | When to use |
+|---------|--------------|--------------|-------------|
+| **Full** (`latest`) | ~9 GB | ~3.7 GB | Default. Works out of the box with no external services except the LLM. |
+| **Slim** (`slim`) | ~500 MB | ~500 MB | Use when you already rely on external services for embeddings and reranking (OpenAI, Cohere, TEI). Significantly smaller image, faster deploys. Requires [external providers](./configuration#embeddings). |
-| Variant | Size (AMD64) | Size (ARM64) | Use Case |
-|---------|--------------|--------------|----------|
-| **Full** (`latest`) | ~9 GB | ~3.7 GB | Includes local ML models (embeddings, reranking) |
-| **Slim** (`slim`) | ~500 MB | ~500 MB | Requires external embedding/reranking providers |
-
-**Full image** (default):
-```bash
-docker run --rm -it -p 8888:8888 \
- -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
- ghcr.io/vectorize-io/hindsight:latest
-```
-- ✅ Works out of the box with local ML models
-- ✅ No additional services needed
-- ❌ Larger image size (AMD64 includes CUDA libraries for GPU support)
-
-**Slim image**:
-```bash
-docker run --rm -it -p 8888:8888 \
- -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
- -e HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai \
- -e HINDSIGHT_API_RERANKER_PROVIDER=cohere \
- -e HINDSIGHT_API_COHERE_API_KEY=$COHERE_API_KEY \
- ghcr.io/vectorize-io/hindsight:latest-slim
-```
-- ✅ Dramatically smaller image (~95% reduction on AMD64)
-- ✅ Faster pull/deploy times
-- ✅ Lower memory footprint
-- ❌ Requires external embedding/reranking services (OpenAI, Cohere, TEI)
-
-**When to use slim:**
-- Cloud deployments where image size matters
-- Using managed embedding services (OpenAI, Cohere)
-- Running on Text Embeddings Inference (TEI) infrastructure
-- Kubernetes environments with fast pull requirements
-
-:::warning Slim Image Requires External Providers
-If you run the slim image **without** setting external embedding providers, you'll see this error:
-
-```
-ImportError: sentence-transformers is required for LocalSTEmbeddings.
-Install it with: pip install sentence-transformers
-```
-
-**Fix:** Always set embedding and reranking providers when using slim images:
-```bash
--e HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai
--e HINDSIGHT_API_EMBEDDINGS_OPENAI_API_KEY=sk-xxx
--e HINDSIGHT_API_RERANKER_PROVIDER=cohere
--e HINDSIGHT_API_COHERE_API_KEY=xxx
-```
-:::
-
-See [Configuration](./configuration#embeddings) for all embedding provider options.
+The slim image corresponds to the [`hindsight-api-slim`](#bare-metal-pip) pip package. See [Configuration](./configuration#embeddings) for external provider options.
### Available Tags
@@ -120,7 +71,7 @@ ghcr.io/vectorize-io/hindsight:0.4.9-slim # Slim, specific version
# API only
ghcr.io/vectorize-io/hindsight-api:latest
-ghcr.io/vectorize-io/hindsight-api:slim
+ghcr.io/vectorize-io/hindsight-api:latest-slim
# Control Plane only
ghcr.io/vectorize-io/hindsight-control-plane:latest
@@ -175,14 +126,17 @@ See the [Helm chart values.yaml](https://github.com/vectorize-io/hindsight/tree/
## Bare Metal (pip)
-**Best for**: Custom deployments, integration into existing Python applications
+**Best for**: Running Hindsight as a standalone service on a host machine.
### Install
```bash
-pip install hindsight-all
+pip install hindsight-api # Full — works out of the box
+pip install hindsight-api-slim # Slim — requires external services for embeddings, reranking, and the database
```
+When using `hindsight-api-slim`, you must configure external providers for all model operations. See [Configuration](./configuration#embeddings) for details.
+
### Run with Embedded Database
For development and testing, Hindsight can run with an embedded PostgreSQL (pg0):
@@ -253,8 +207,44 @@ PORT=80 HINDSIGHT_CP_DATAPLANE_API_URL=https://api.hindsight.io npx @vectorize-i
---
+## Embedded in a Python Application
+
+**Best for**: Using Hindsight programmatically from Python without running a separate server process.
+
+```bash
+pip install hindsight-all # Full — works out of the box
+pip install hindsight-all-slim # Slim — requires external services for embeddings, reranking, and the database
+```
+
+`hindsight-all` supports two modes of embedding:
+
+**In-process** (`HindsightServer`): the server runs in a background thread inside your application. Best when you want the tightest integration and are already managing your own process lifecycle.
+
+```python
+from hindsight import HindsightServer, HindsightClient
+
+with HindsightServer(llm_provider="openai", llm_api_key="sk-xxx") as server:
+ client = HindsightClient(base_url=server.url)
+ client.retain(bank_id="alice", content="Alice prefers concise answers.")
+ results = client.recall(bank_id="alice", query="How should I respond to Alice?")
+```
+
+**Managed subprocess** (`HindsightEmbedded`): the server runs as a background daemon process, shared across multiple Python processes or sessions. The daemon starts on first use and shuts down automatically after an idle timeout.
+
+```python
+from hindsight import HindsightEmbedded
+
+client = HindsightEmbedded(llm_provider="openai", llm_api_key="sk-xxx")
+client.retain(bank_id="alice", content="Alice prefers concise answers.")
+results = client.recall(bank_id="alice", query="How should I respond to Alice?")
+```
+
+See the [Python SDK](../sdks/python.md) for the full API reference.
+
+---
+
## Next Steps
- [Configuration](./configuration.md) — Environment variables and settings
-- [Models](./models.md) — ML models and providers
+- [Models](./models.mdx) — ML models and providers
- [Monitoring](./monitoring.md) — Metrics and observability
diff --git a/hindsight-docs/versioned_docs/version-0.4/developer/models.md b/hindsight-docs/versioned_docs/version-0.4/developer/models.mdx
similarity index 95%
rename from hindsight-docs/versioned_docs/version-0.4/developer/models.md
rename to hindsight-docs/versioned_docs/version-0.4/developer/models.mdx
index 7b2c817d..22bb8c95 100644
--- a/hindsight-docs/versioned_docs/version-0.4/developer/models.md
+++ b/hindsight-docs/versioned_docs/version-0.4/developer/models.mdx
@@ -1,16 +1,16 @@
+import {LLMProvidersGrid} from '@site/src/components/SupportedGrids';
+
# Models
Hindsight uses several machine learning models for different tasks.
## Overview
-| Model Type | Purpose | Default | Configurable |
-|------------|---------|---------|--------------|
-| **LLM** | Fact extraction, reasoning, generation | Provider-specific | Yes |
-| **Embedding** | Vector representations for semantic search | `BAAI/bge-small-en-v1.5` | Yes |
-| **Cross-Encoder** | Reranking search results | `cross-encoder/ms-marco-MiniLM-L-6-v2` | Yes |
+- **LLM** — Fact extraction, reasoning, and generation. Provider-specific, fully configurable.
+- **Embedding** — Vector representations for semantic search. Default: `BAAI/bge-small-en-v1.5`.
+- **Cross-Encoder** — Reranking search results. Default: `cross-encoder/ms-marco-MiniLM-L-6-v2`.
-All local models (embedding, cross-encoder) are automatically downloaded from HuggingFace on first run.
+Embedding and cross-encoder models are downloaded automatically from HuggingFace on first run.
---
@@ -18,7 +18,11 @@ All local models (embedding, cross-encoder) are automatically downloaded from Hu
Used for fact extraction, entity resolution, mental model consolidation, and answer synthesis.
-**Supported providers:** OpenAI, Anthropic, Gemini, Groq, Ollama, LM Studio, and **any OpenAI-compatible API**
+**Supported providers:**
+
+
+
+Also supports **any OpenAI-compatible API** (e.g., Azure OpenAI, Together AI, Fireworks).
:::tip OpenAI-Compatible Providers
Hindsight works with any provider that exposes an OpenAI-compatible API (e.g., Azure OpenAI). Simply set `HINDSIGHT_API_LLM_PROVIDER=openai` and configure `HINDSIGHT_API_LLM_BASE_URL` to point to your provider's endpoint.
@@ -63,6 +67,7 @@ Each provider has a recommended default model that's used when `HINDSIGHT_API_LL
| `anthropic` | `claude-haiku-4-5-20251001` |
| `gemini` | `gemini-2.5-flash` |
| `groq` | `openai/gpt-oss-120b` |
+| `minimax` | `MiniMax-M2.5` |
| `ollama` | `gemma3:12b` |
| `lmstudio` | `local-model` |
| `vertexai` | `gemini-2.0-flash-001` |
@@ -144,6 +149,11 @@ export HINDSIGHT_API_LLM_PROVIDER=lmstudio
export HINDSIGHT_API_LLM_BASE_URL=http://localhost:1234/v1
export HINDSIGHT_API_LLM_MODEL=your-local-model
+# MiniMax (204K context window)
+export HINDSIGHT_API_LLM_PROVIDER=minimax
+export HINDSIGHT_API_LLM_API_KEY=your-minimax-api-key
+export HINDSIGHT_API_LLM_MODEL=MiniMax-M2.5
+
# Vertex AI (Google Cloud)
export HINDSIGHT_API_LLM_PROVIDER=vertexai
export HINDSIGHT_API_LLM_MODEL=gemini-2.0-flash-001
diff --git a/hindsight-docs/versioned_docs/version-0.4/sdks/embed.md b/hindsight-docs/versioned_docs/version-0.4/sdks/embed.md
index f94dd8e1..7b1ba402 100644
--- a/hindsight-docs/versioned_docs/version-0.4/sdks/embed.md
+++ b/hindsight-docs/versioned_docs/version-0.4/sdks/embed.md
@@ -88,7 +88,7 @@ The daemon starts automatically on first use!
| Variable | Description | Default |
|----------|-------------|---------|
| `HINDSIGHT_EMBED_LLM_API_KEY` | **Required**. API key for LLM provider | - |
-| `HINDSIGHT_EMBED_LLM_PROVIDER` | LLM provider: `openai`, `anthropic`, `gemini`, `groq`, `ollama` | `openai` |
+| `HINDSIGHT_EMBED_LLM_PROVIDER` | LLM provider: `openai`, `anthropic`, `gemini`, `groq`, `minimax`, `ollama` | `openai` |
| `HINDSIGHT_EMBED_LLM_MODEL` | Model name | `gpt-4o-mini` |
| `HINDSIGHT_EMBED_BANK_ID` | Default memory bank ID | `default` |
| `HINDSIGHT_EMBED_DAEMON_IDLE_TIMEOUT` | Seconds before daemon auto-exits when idle (0 = never) | `300` |
diff --git a/hindsight-docs/versioned_sidebars/version-0.4-sidebars.json b/hindsight-docs/versioned_sidebars/version-0.4-sidebars.json
index 1d74ec9c..6329cd25 100644
--- a/hindsight-docs/versioned_sidebars/version-0.4-sidebars.json
+++ b/hindsight-docs/versioned_sidebars/version-0.4-sidebars.json
@@ -5,15 +5,78 @@
"label": "Architecture",
"collapsible": false,
"items": [
- { "type": "doc", "id": "developer/index", "label": "Overview", "customProps": { "icon": "lu-book" } },
- { "type": "doc", "id": "developer/retain", "label": "Retain", "customProps": { "icon": "lu-brain" } },
- { "type": "doc", "id": "developer/retrieval", "label": "Recall", "customProps": { "icon": "lu-search" } },
- { "type": "doc", "id": "developer/reflect", "label": "Reflect", "customProps": { "icon": "lu-message" } },
- { "type": "doc", "id": "developer/observations", "label": "Observations", "customProps": { "icon": "lu-activity" } },
- { "type": "doc", "id": "developer/multilingual", "label": "Multilingual", "customProps": { "icon": "lu-languages" } },
- { "type": "doc", "id": "developer/performance", "label": "Performance", "customProps": { "icon": "lu-zap" } },
- { "type": "doc", "id": "developer/storage", "label": "Storage", "customProps": { "icon": "lu-database" } },
- { "type": "doc", "id": "developer/rag-vs-hindsight", "label": "RAG vs Memory", "customProps": { "icon": "lu-compare" } }
+ {
+ "type": "doc",
+ "id": "developer/index",
+ "label": "Overview",
+ "customProps": {
+ "icon": "lu-book"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/retain",
+ "label": "Retain",
+ "customProps": {
+ "icon": "lu-brain"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/retrieval",
+ "label": "Recall",
+ "customProps": {
+ "icon": "lu-search"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/reflect",
+ "label": "Reflect",
+ "customProps": {
+ "icon": "lu-message"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/observations",
+ "label": "Observations",
+ "customProps": {
+ "icon": "lu-activity"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/multilingual",
+ "label": "Multilingual",
+ "customProps": {
+ "icon": "lu-languages"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/performance",
+ "label": "Performance",
+ "customProps": {
+ "icon": "lu-zap"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/storage",
+ "label": "Storage",
+ "customProps": {
+ "icon": "lu-database"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/rag-vs-hindsight",
+ "label": "RAG vs Memory",
+ "customProps": {
+ "icon": "lu-compare"
+ }
+ }
]
},
{
@@ -21,15 +84,87 @@
"label": "API",
"collapsible": false,
"items": [
- { "type": "doc", "id": "developer/api/quickstart", "label": "Quick Start", "customProps": { "icon": "lu-rocket" } },
- { "type": "doc", "id": "developer/api/retain", "label": "Retain", "customProps": { "icon": "lu-brain" } },
- { "type": "doc", "id": "developer/api/recall", "label": "Recall", "customProps": { "icon": "lu-search" } },
- { "type": "doc", "id": "developer/api/reflect", "label": "Reflect", "customProps": { "icon": "lu-message" } },
- { "type": "doc", "id": "developer/api/mental-models", "label": "Mental Models", "customProps": { "icon": "lu-layers" } },
- { "type": "doc", "id": "developer/api/memory-banks", "label": "Memory Banks", "customProps": { "icon": "lu-memory" } },
- { "type": "doc", "id": "developer/api/documents", "label": "Documents", "customProps": { "icon": "lu-file" } },
- { "type": "doc", "id": "developer/api/operations", "label": "Operations", "customProps": { "icon": "lu-cpu" } },
- { "type": "doc", "id": "developer/api/webhooks", "label": "Webhooks", "customProps": { "icon": "lu-webhook" } }
+ {
+ "type": "doc",
+ "id": "developer/api/quickstart",
+ "label": "Quick Start",
+ "customProps": {
+ "icon": "lu-rocket"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/api/retain",
+ "label": "Retain",
+ "customProps": {
+ "icon": "lu-brain"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/api/recall",
+ "label": "Recall",
+ "customProps": {
+ "icon": "lu-search"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/api/reflect",
+ "label": "Reflect",
+ "customProps": {
+ "icon": "lu-message"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/api/mental-models",
+ "label": "Mental Models",
+ "customProps": {
+ "icon": "lu-layers"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/api/memory-banks",
+ "label": "Memory Banks",
+ "customProps": {
+ "icon": "lu-memory"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/api/documents",
+ "label": "Documents",
+ "customProps": {
+ "icon": "lu-file"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/api/operations",
+ "label": "Operations",
+ "customProps": {
+ "icon": "lu-cpu"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/api/webhooks",
+ "label": "Webhooks",
+ "customProps": {
+ "icon": "lu-webhook"
+ }
+ },
+ {
+ "type": "link",
+ "href": "/api-reference",
+ "label": "API Reference",
+ "customProps": {
+ "icon": "lu-book-open",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ }
]
},
{
@@ -37,11 +172,46 @@
"label": "Clients",
"collapsible": false,
"items": [
- { "type": "doc", "id": "sdks/python", "label": "Python", "customProps": { "icon": "si-python" } },
- { "type": "doc", "id": "sdks/nodejs", "label": "TypeScript", "customProps": { "icon": "/img/icons/typescript.png" } },
- { "type": "doc", "id": "sdks/go", "label": "Go", "customProps": { "icon": "si-go" } },
- { "type": "doc", "id": "sdks/cli", "label": "CLI", "customProps": { "icon": "lu-terminal" } },
- { "type": "doc", "id": "sdks/embed", "label": "Embedded Python", "customProps": { "icon": "/img/icons/package.svg" } }
+ {
+ "type": "doc",
+ "id": "sdks/python",
+ "label": "Python",
+ "customProps": {
+ "icon": "si-python"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/nodejs",
+ "label": "TypeScript",
+ "customProps": {
+ "icon": "/img/icons/typescript.png"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/go",
+ "label": "Go",
+ "customProps": {
+ "icon": "si-go"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/cli",
+ "label": "CLI",
+ "customProps": {
+ "icon": "lu-terminal"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/embed",
+ "label": "Embedded Python",
+ "customProps": {
+ "icon": "/img/icons/package.svg"
+ }
+ }
]
},
{
@@ -49,14 +219,70 @@
"label": "Integrations",
"collapsible": false,
"items": [
- { "type": "doc", "id": "sdks/integrations/local-mcp", "label": "Local MCP Server", "customProps": { "icon": "/img/icons/mcp.png" } },
- { "type": "doc", "id": "sdks/integrations/litellm", "label": "LiteLLM", "customProps": { "icon": "/img/icons/litellm.png" } },
- { "type": "doc", "id": "sdks/integrations/openclaw", "label": "OpenClaw", "customProps": { "icon": "/img/icons/openclaw.png" } },
- { "type": "doc", "id": "sdks/integrations/ai-sdk", "label": "Vercel AI SDK", "customProps": { "icon": "/img/icons/vercel.png" } },
- { "type": "doc", "id": "sdks/integrations/chat", "label": "Vercel Chat SDK", "customProps": { "icon": "/img/icons/vercel.png" } },
- { "type": "doc", "id": "sdks/integrations/crewai", "label": "CrewAI", "customProps": { "icon": "/img/icons/crewai.png" } },
- { "type": "doc", "id": "sdks/integrations/pydantic-ai", "label": "Pydantic AI", "customProps": { "icon": "/img/icons/pydanticai.png" } },
- { "type": "doc", "id": "sdks/integrations/skills", "label": "Skills", "customProps": { "icon": "/img/icons/skills.png" } }
+ {
+ "type": "doc",
+ "id": "sdks/integrations/local-mcp",
+ "label": "Local MCP Server",
+ "customProps": {
+ "icon": "/img/icons/mcp.png"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/integrations/litellm",
+ "label": "LiteLLM",
+ "customProps": {
+ "icon": "/img/icons/litellm.png"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/integrations/openclaw",
+ "label": "OpenClaw",
+ "customProps": {
+ "icon": "/img/icons/openclaw.png"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/integrations/ai-sdk",
+ "label": "Vercel AI SDK",
+ "customProps": {
+ "icon": "/img/icons/vercel.png"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/integrations/chat",
+ "label": "Vercel Chat SDK",
+ "customProps": {
+ "icon": "/img/icons/vercel.png"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/integrations/crewai",
+ "label": "CrewAI",
+ "customProps": {
+ "icon": "/img/icons/crewai.png"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/integrations/pydantic-ai",
+ "label": "Pydantic AI",
+ "customProps": {
+ "icon": "/img/icons/pydanticai.png"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "sdks/integrations/skills",
+ "label": "Skills",
+ "customProps": {
+ "icon": "/img/icons/skills.png"
+ }
+ }
]
},
{
@@ -64,14 +290,149 @@
"label": "Hosting",
"collapsible": false,
"items": [
- { "type": "doc", "id": "developer/installation", "label": "Installation", "customProps": { "icon": "lu-package" } },
- { "type": "doc", "id": "developer/services", "label": "Services", "customProps": { "icon": "lu-server" } },
- { "type": "doc", "id": "developer/configuration", "label": "Configuration", "customProps": { "icon": "lu-settings" } },
- { "type": "doc", "id": "developer/admin-cli", "label": "Admin CLI", "customProps": { "icon": "lu-terminal" } },
- { "type": "doc", "id": "developer/extensions", "label": "Extensions", "customProps": { "icon": "lu-plug" } },
- { "type": "doc", "id": "developer/models", "label": "Models", "customProps": { "icon": "lu-cpu" } },
- { "type": "doc", "id": "developer/monitoring", "label": "Monitoring", "customProps": { "icon": "lu-activity" } },
- { "type": "doc", "id": "developer/mcp-server", "label": "MCP Server", "customProps": { "icon": "lu-network" } }
+ {
+ "type": "link",
+ "href": "https://ui.hindsight.vectorize.io/signup",
+ "label": "Cloud",
+ "customProps": {
+ "icon": "lu-cloud",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/installation",
+ "label": "Installation",
+ "customProps": {
+ "icon": "lu-package"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/services",
+ "label": "Services",
+ "customProps": {
+ "icon": "lu-server"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/configuration",
+ "label": "Configuration",
+ "customProps": {
+ "icon": "lu-settings"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/admin-cli",
+ "label": "Admin CLI",
+ "customProps": {
+ "icon": "lu-terminal"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/extensions",
+ "label": "Extensions",
+ "customProps": {
+ "icon": "lu-plug"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/models",
+ "label": "Models",
+ "customProps": {
+ "icon": "lu-cpu"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/monitoring",
+ "label": "Monitoring",
+ "customProps": {
+ "icon": "lu-activity"
+ }
+ },
+ {
+ "type": "doc",
+ "id": "developer/mcp-server",
+ "label": "MCP Server",
+ "customProps": {
+ "icon": "lu-network"
+ }
+ }
+ ]
+ },
+ {
+ "type": "category",
+ "label": "More",
+ "collapsible": false,
+ "items": [
+ {
+ "type": "link",
+ "href": "/cookbook",
+ "label": "Cookbook",
+ "customProps": {
+ "icon": "lu-book",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ },
+ {
+ "type": "link",
+ "href": "/blog",
+ "label": "Blog",
+ "customProps": {
+ "icon": "lu-rss",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ },
+ {
+ "type": "link",
+ "href": "https://join.slack.com/t/hindsight-space/shared_invite/zt-3nhbm4w29-LeSJ5Ixi6j8PdiYOCPlOgg",
+ "label": "Community",
+ "customProps": {
+ "icon": "si-slack",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ },
+ {
+ "type": "link",
+ "href": "https://github.com/vectorize-io/hindsight",
+ "label": "GitHub",
+ "customProps": {
+ "icon": "si-github",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ },
+ {
+ "type": "link",
+ "href": "https://benchmarks.hindsight.vectorize.io/",
+ "label": "Benchmarks",
+ "customProps": {
+ "icon": "lu-chart-bar",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ },
+ {
+ "type": "link",
+ "href": "https://benchmarks.hindsight.vectorize.io/",
+ "label": "Which Model Should I Use?",
+ "customProps": {
+ "icon": "lu-cpu",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ },
+ {
+ "type": "link",
+ "href": "https://arxiv.org/abs/2512.12818",
+ "label": "Paper",
+ "customProps": {
+ "icon": "lu-file-text",
+ "iconAfter": "lu-arrow-up-right"
+ }
+ }
]
}
]
diff --git a/skills/hindsight-docs/references/developer/api/quickstart.md b/skills/hindsight-docs/references/developer/api/quickstart.md
index 97d9e5d2..dea9e3c8 100644
--- a/skills/hindsight-docs/references/developer/api/quickstart.md
+++ b/skills/hindsight-docs/references/developer/api/quickstart.md
@@ -5,6 +5,10 @@ Get up and running with Hindsight in 60 seconds.
{/* Import raw source files */}
+## Clients
+
+
+
## Start the API Server
### pip (API only)
@@ -113,6 +117,10 @@ hindsight memory reflect my-bank "Tell me about Alice"
---
+## Integrations
+
+
+
## Next Steps
- [**Retain**](./retain) — Advanced options for storing memories
diff --git a/skills/hindsight-docs/references/developer/api/recall.md b/skills/hindsight-docs/references/developer/api/recall.md
index 91a5edc4..d7ec1027 100644
--- a/skills/hindsight-docs/references/developer/api/recall.md
+++ b/skills/hindsight-docs/references/developer/api/recall.md
@@ -333,6 +333,66 @@ Use this for strict scope enforcement where a memory must explicitly belong to *
> **💡 Extra tags are fine**
>
A memory with tags `["user:alice", "team", "project:x"]` will still match a filter of `["user:alice", "team"]` under `all_strict` — extra tags on the memory are not a problem. The filter only requires the memory to contain **at least** the specified tags.
+### tag_groups
+
+`tag_groups` is a list of compound boolean tag filters. The groups in the list are AND-ed together at the top level. Each group is a recursive boolean expression: a **leaf** node `{tags, match}`, or a **compound** node `{and: [...]}`, `{or: [...]}`, or `{not: ...}`.
+
+`tag_groups` and `tags` / `tags_match` can be used simultaneously — they are AND-ed together.
+
+#### Leaf node
+
+```json
+{ "tags": ["step:5", "step:8"], "match": "any_strict" }
+```
+
+`match` accepts the same values as `tags_match`: `any`, `all`, `any_strict`, `all_strict`. Defaults to `any_strict`.
+
+#### Compound nodes
+
+```json
+{ "and": [ , , ... ] }
+{ "or": [ , , ... ] }
+{ "not": }
+```
+
+#### Examples
+
+**Step filter AND user scope** — two top-level groups AND-ed:
+
+```json
+{
+ "tag_groups": [
+ { "tags": ["step:5", "step:8", "step:12"], "match": "any_strict" },
+ { "tags": ["user:ep_42"], "match": "all_strict" }
+ ]
+}
+```
+
+**Nested OR inside AND** — user must match, plus either step OR priority:
+
+```json
+{
+ "tag_groups": [
+ { "tags": ["user:alice"], "match": "all_strict" },
+ { "or": [
+ { "tags": ["step:5"], "match": "any_strict" },
+ { "tags": ["priority:high"], "match": "all_strict" }
+ ]}
+ ]
+}
+```
+
+**Exclusion** — user must match, but archived memories are excluded:
+
+```json
+{
+ "tag_groups": [
+ { "tags": ["user:alice"], "match": "all_strict" },
+ { "not": { "tags": ["archived"], "match": "any_strict" } }
+ ]
+}
+```
+
### trace
When set to `true`, the response includes a detailed debug trace covering the query embedding, entry points, per-strategy retrieval results, RRF fusion candidates, reranked results, temporal constraints detected, and per-phase timings. Has no effect on the retrieval logic itself. Useful for understanding why specific memories were or were not returned.
diff --git a/skills/hindsight-docs/references/developer/configuration.md b/skills/hindsight-docs/references/developer/configuration.md
index 43f3368c..012746b7 100644
--- a/skills/hindsight-docs/references/developer/configuration.md
+++ b/skills/hindsight-docs/references/developer/configuration.md
@@ -160,7 +160,7 @@ To switch between backends:
| Variable | Description | Default |
|----------|-------------|---------|
-| `HINDSIGHT_API_LLM_PROVIDER` | Provider: `openai`, `openai-codex`, `claude-code`, `anthropic`, `gemini`, `groq`, `ollama`, `lmstudio`, `vertexai` | `openai` |
+| `HINDSIGHT_API_LLM_PROVIDER` | Provider: `openai`, `openai-codex`, `claude-code`, `anthropic`, `gemini`, `groq`, `minimax`, `ollama`, `lmstudio`, `vertexai` | `openai` |
| `HINDSIGHT_API_LLM_API_KEY` | API key for LLM provider | - |
| `HINDSIGHT_API_LLM_MODEL` | Model name | `gpt-5-mini` |
| `HINDSIGHT_API_LLM_BASE_URL` | Custom LLM endpoint | Provider default |
@@ -410,7 +410,7 @@ Supported OpenAI embedding dimensions:
| Variable | Description | Default |
|----------|-------------|---------|
-| `HINDSIGHT_API_RERANKER_PROVIDER` | Provider: `local`, `tei`, `cohere`, `zeroentropy`, `flashrank`, `litellm`, `litellm-sdk`, or `rrf` | `local` |
+| `HINDSIGHT_API_RERANKER_PROVIDER` | Provider: `local`, `tei`, `cohere`, `zeroentropy`, `flashrank`, `litellm`, `litellm-sdk`, `jina-mlx`, or `rrf` | `local` |
| `HINDSIGHT_API_RERANKER_LOCAL_MODEL` | Model for local provider | `cross-encoder/ms-marco-MiniLM-L-6-v2` |
| `HINDSIGHT_API_RERANKER_LOCAL_MAX_CONCURRENT` | Max concurrent local reranking (prevents CPU thrashing under load) | `4` |
| `HINDSIGHT_API_RERANKER_LOCAL_TRUST_REMOTE_CODE` | Allow loading models with custom code (security risk, disabled by default) | `false` |
@@ -426,10 +426,12 @@ Supported OpenAI embedding dimensions:
| `HINDSIGHT_API_RERANKER_LITELLM_SDK_API_KEY` | LiteLLM **SDK** API key for direct reranking (no proxy needed) | - |
| `HINDSIGHT_API_RERANKER_LITELLM_SDK_MODEL` | LiteLLM SDK rerank model (e.g., `deepinfra/Qwen3-reranker-8B`) | `cohere/rerank-english-v3.0` |
| `HINDSIGHT_API_RERANKER_LITELLM_SDK_API_BASE` | Custom API base URL for LiteLLM SDK (optional) | - |
+| `HINDSIGHT_API_RERANKER_LITELLM_MAX_TOKENS_PER_DOC` | Truncate documents to this many tokens before sending to the reranker (applies to both `litellm` and `litellm-sdk`). Use for models with small context windows (e.g. set to `900` for a 1024-token limit model). Unset by default (no truncation). | - |
| `HINDSIGHT_API_RERANKER_ZEROENTROPY_API_KEY` | ZeroEntropy API key for reranking | - |
| `HINDSIGHT_API_RERANKER_ZEROENTROPY_MODEL` | ZeroEntropy rerank model (`zerank-2`, `zerank-2-small`) | `zerank-2` |
| `HINDSIGHT_API_RERANKER_FLASHRANK_MODEL` | FlashRank model for fast CPU-based reranking | `ms-marco-MiniLM-L-12-v2` |
| `HINDSIGHT_API_RERANKER_FLASHRANK_CACHE_DIR` | Cache directory for FlashRank models | System default |
+| `HINDSIGHT_API_RERANKER_JINA_MLX_MODEL_PATH` | Local path to downloaded `jina-reranker-v3-mlx` model (auto-downloads from HuggingFace if unset) | - |
```bash
# Local (default) - uses SentenceTransformers CrossEncoder
@@ -472,6 +474,10 @@ export HINDSIGHT_API_RERANKER_LITELLM_MODEL=cohere/rerank-english-v3.0 # or voy
export HINDSIGHT_API_RERANKER_PROVIDER=litellm-sdk
export HINDSIGHT_API_RERANKER_LITELLM_SDK_API_KEY=your-deepinfra-api-key
export HINDSIGHT_API_RERANKER_LITELLM_SDK_MODEL=deepinfra/Qwen3-reranker-8B # or cohere/rerank-english-v3.0, etc.
+
+# Jina MLX - Apple Silicon native reranking (no GPU/cloud required)
+# Model (~1.2 GB) is downloaded automatically from HuggingFace Hub on first use.
+export HINDSIGHT_API_RERANKER_PROVIDER=jina-mlx
```
#### LiteLLM Proxy vs SDK
@@ -488,6 +494,14 @@ Both support the same providers:
- **Jina AI** (`jina_ai/jina-reranker-v2`)
- **AWS Bedrock** (`bedrock/...`)
+#### Jina MLX (Apple Silicon)
+
+The `jina-mlx` provider uses [`jinaai/jina-reranker-v3-mlx`](https://huggingface.co/jinaai/jina-reranker-v3-mlx), optimized for Apple Silicon. The model (~1.2 GB) is downloaded from HuggingFace Hub automatically on first startup and cached locally.
+
+:::note License
+`jina-reranker-v3-mlx` is licensed under CC BY-NC 4.0. Contact Jina AI for commercial usage.
+:::
+
### Authentication
By default, Hindsight runs without authentication. For production deployments, enable API key authentication using the built-in tenant extension:
@@ -530,6 +544,7 @@ For advanced authentication (JWT, OAuth, multi-tenant schemas), implement a cust
| `HINDSIGHT_API_GRAPH_RETRIEVER` | Graph retrieval algorithm: `link_expansion`, `mpfp`, or `bfs` | `link_expansion` |
| `HINDSIGHT_API_RECALL_MAX_CONCURRENT` | Max concurrent recall operations per worker (backpressure) | `32` |
| `HINDSIGHT_API_RECALL_CONNECTION_BUDGET` | Max concurrent DB connections per recall operation | `4` |
+| `HINDSIGHT_API_RECALL_MAX_QUERY_TOKENS` | Maximum token length of a recall query; requests exceeding this limit are rejected with HTTP 400 | `500` |
| `HINDSIGHT_API_RERANKER_MAX_CANDIDATES` | Max candidates to rerank per recall (RRF pre-filters the rest) | `300` |
| `HINDSIGHT_API_MPFP_TOP_K_NEIGHBORS` | Fan-out limit per node in MPFP graph traversal | `20` |
| `HINDSIGHT_API_MENTAL_MODEL_REFRESH_CONCURRENCY` | Max concurrent mental model refreshes | `8` |
diff --git a/skills/hindsight-docs/references/developer/index.md b/skills/hindsight-docs/references/developer/index.md
index 8c25e95b..9191df56 100644
--- a/skills/hindsight-docs/references/developer/index.md
+++ b/skills/hindsight-docs/references/developer/index.md
@@ -1,7 +1,4 @@
----
-sidebar_position: 1
-slug: /
----
+
# Overview
@@ -114,6 +111,14 @@ The **mission** tells Hindsight what knowledge to prioritize and provides contex
These settings only affect the `reflect` operation, not `recall`.
+## Clients & Languages
+
+
+
+## Integrations
+
+
+
## Next Steps
### Getting Started
diff --git a/skills/hindsight-docs/references/developer/installation.md b/skills/hindsight-docs/references/developer/installation.md
index e30ba8be..d5ce0c3c 100644
--- a/skills/hindsight-docs/references/developer/installation.md
+++ b/skills/hindsight-docs/references/developer/installation.md
@@ -8,28 +8,28 @@ Hindsight can be deployed in several ways depending on your infrastructure and r
## Prerequisites
-### PostgreSQL with pgvector
+### PostgreSQL
-Hindsight requires PostgreSQL with the **pgvector** extension for vector similarity search.
+Hindsight requires PostgreSQL 14+ with a vector extension for similarity search. The supported extensions are:
+
+- **pgvector** (default)
+- **pgvectorscale**
+- **vchord**
+
+Configure which one to use with `HINDSIGHT_API_VECTOR_EXTENSION`. See [Configuration](./configuration) for details.
**By default**, Hindsight uses **pg0** — an embedded PostgreSQL that runs locally on your machine. This is convenient for development but **not recommended for production**.
-**For production**, use an external PostgreSQL with pgvector:
+**For production**, use an external PostgreSQL with one of the supported vector extensions:
- **Supabase** — Managed PostgreSQL with pgvector built-in
- **Neon** — Serverless PostgreSQL with pgvector
-- **Azure Database for PostgreSQL** — With pgvector and pg_diskann (DiskANN) support
+- **Azure Database for PostgreSQL** — With pgvector and pgvectorscale support
- **AWS RDS** / **Cloud SQL** — With pgvector extension enabled
-- **Self-hosted** — PostgreSQL 14+ with pgvector installed
+- **Self-hosted** — PostgreSQL 14+ with your preferred vector extension
### LLM Provider
-You need an LLM API key for fact extraction, entity resolution, and answer generation:
-
-- **Groq** (recommended): Fast inference with `gpt-oss-20b`
-- **OpenAI**: GPT-4o, GPT-4o-mini
-- **Ollama**: Run models locally
-
-See [Models](./models) for detailed comparison and configuration.
+You need an LLM API key for fact extraction, entity resolution, and answer generation. See [Models](./models) for supported providers, model recommendations, and configuration.
---
@@ -53,61 +53,12 @@ docker run --rm -it --pull always -p 8888:8888 -p 9999:9999 \
### Docker Image Variants
-Hindsight provides two image variants with different size/capability tradeoffs:
+| Variant | Size (AMD64) | Size (ARM64) | When to use |
+|---------|--------------|--------------|-------------|
+| **Full** (`latest`) | ~9 GB | ~3.7 GB | Default. Works out of the box with no external services except the LLM. |
+| **Slim** (`slim`) | ~500 MB | ~500 MB | Use when you already rely on external services for embeddings and reranking (OpenAI, Cohere, TEI). Significantly smaller image, faster deploys. Requires [external providers](./configuration#embeddings). |
-| Variant | Size (AMD64) | Size (ARM64) | Use Case |
-|---------|--------------|--------------|----------|
-| **Full** (`latest`) | ~9 GB | ~3.7 GB | Includes local ML models (embeddings, reranking) |
-| **Slim** (`slim`) | ~500 MB | ~500 MB | Requires external embedding/reranking providers |
-
-**Full image** (default):
-```bash
-docker run --rm -it -p 8888:8888 \
- -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
- ghcr.io/vectorize-io/hindsight:latest
-```
-- ✅ Works out of the box with local ML models
-- ✅ No additional services needed
-- ❌ Larger image size (AMD64 includes CUDA libraries for GPU support)
-
-**Slim image**:
-```bash
-docker run --rm -it -p 8888:8888 \
- -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
- -e HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai \
- -e HINDSIGHT_API_RERANKER_PROVIDER=cohere \
- -e HINDSIGHT_API_COHERE_API_KEY=$COHERE_API_KEY \
- ghcr.io/vectorize-io/hindsight:latest-slim
-```
-- ✅ Dramatically smaller image (~95% reduction on AMD64)
-- ✅ Faster pull/deploy times
-- ✅ Lower memory footprint
-- ❌ Requires external embedding/reranking services (OpenAI, Cohere, TEI)
-
-**When to use slim:**
-- Cloud deployments where image size matters
-- Using managed embedding services (OpenAI, Cohere)
-- Running on Text Embeddings Inference (TEI) infrastructure
-- Kubernetes environments with fast pull requirements
-
-:::warning Slim Image Requires External Providers
-If you run the slim image **without** setting external embedding providers, you'll see this error:
-
-```
-ImportError: sentence-transformers is required for LocalSTEmbeddings.
-Install it with: pip install sentence-transformers
-```
-
-**Fix:** Always set embedding and reranking providers when using slim images:
-```bash
--e HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai
--e HINDSIGHT_API_EMBEDDINGS_OPENAI_API_KEY=sk-xxx
--e HINDSIGHT_API_RERANKER_PROVIDER=cohere
--e HINDSIGHT_API_COHERE_API_KEY=xxx
-```
-:::
-
-See [Configuration](./configuration#embeddings) for all embedding provider options.
+The slim image corresponds to the [`hindsight-api-slim`](#package-variants) pip package. See [Configuration](./configuration#embeddings) for external provider options.
### Available Tags
@@ -120,7 +71,7 @@ ghcr.io/vectorize-io/hindsight:0.4.9-slim # Slim, specific version
# API only
ghcr.io/vectorize-io/hindsight-api:latest
-ghcr.io/vectorize-io/hindsight-api:slim
+ghcr.io/vectorize-io/hindsight-api:latest-slim
# Control Plane only
ghcr.io/vectorize-io/hindsight-control-plane:latest
@@ -175,14 +126,17 @@ See the [Helm chart values.yaml](https://github.com/vectorize-io/hindsight/tree/
## Bare Metal (pip)
-**Best for**: Custom deployments, integration into existing Python applications
+**Best for**: Running Hindsight as a standalone service on a host machine.
### Install
```bash
-pip install hindsight-all
+pip install hindsight-api # Full — works out of the box
+pip install hindsight-api-slim # Slim — requires external services for embeddings, reranking, and the database
```
+When using `hindsight-api-slim`, you must configure external providers for all model operations. See [Configuration](./configuration#embeddings) for details.
+
### Run with Embedded Database
For development and testing, Hindsight can run with an embedded PostgreSQL (pg0):
@@ -253,6 +207,42 @@ PORT=80 HINDSIGHT_CP_DATAPLANE_API_URL=https://api.hindsight.io npx @vectorize-i
---
+## Embedded in a Python Application
+
+**Best for**: Using Hindsight programmatically from Python without running a separate server process.
+
+```bash
+pip install hindsight-all # Full — works out of the box
+pip install hindsight-all-slim # Slim — requires external services for embeddings, reranking, and the database
+```
+
+`hindsight-all` supports two modes of embedding:
+
+**In-process** (`HindsightServer`): the server runs in a background thread inside your application. Best when you want the tightest integration and are already managing your own process lifecycle.
+
+```python
+from hindsight import HindsightServer, HindsightClient
+
+with HindsightServer(llm_provider="openai", llm_api_key="sk-xxx") as server:
+ client = HindsightClient(base_url=server.url)
+ client.retain(bank_id="alice", content="Alice prefers concise answers.")
+ results = client.recall(bank_id="alice", query="How should I respond to Alice?")
+```
+
+**Managed subprocess** (`HindsightEmbedded`): the server runs as a background daemon process, shared across multiple Python processes or sessions. The daemon starts on first use and shuts down automatically after an idle timeout.
+
+```python
+from hindsight import HindsightEmbedded
+
+client = HindsightEmbedded(llm_provider="openai", llm_api_key="sk-xxx")
+client.retain(bank_id="alice", content="Alice prefers concise answers.")
+results = client.recall(bank_id="alice", query="How should I respond to Alice?")
+```
+
+See the [Python SDK](../sdks/python.md) for the full API reference.
+
+---
+
## Next Steps
- [Configuration](./configuration.md) — Environment variables and settings
diff --git a/skills/hindsight-docs/references/developer/models.md b/skills/hindsight-docs/references/developer/models.md
index 7b2c817d..8680de45 100644
--- a/skills/hindsight-docs/references/developer/models.md
+++ b/skills/hindsight-docs/references/developer/models.md
@@ -1,16 +1,15 @@
+
# Models
Hindsight uses several machine learning models for different tasks.
## Overview
-| Model Type | Purpose | Default | Configurable |
-|------------|---------|---------|--------------|
-| **LLM** | Fact extraction, reasoning, generation | Provider-specific | Yes |
-| **Embedding** | Vector representations for semantic search | `BAAI/bge-small-en-v1.5` | Yes |
-| **Cross-Encoder** | Reranking search results | `cross-encoder/ms-marco-MiniLM-L-6-v2` | Yes |
+- **LLM** — Fact extraction, reasoning, and generation. Provider-specific, fully configurable.
+- **Embedding** — Vector representations for semantic search. Default: `BAAI/bge-small-en-v1.5`.
+- **Cross-Encoder** — Reranking search results. Default: `cross-encoder/ms-marco-MiniLM-L-6-v2`.
-All local models (embedding, cross-encoder) are automatically downloaded from HuggingFace on first run.
+Embedding and cross-encoder models are downloaded automatically from HuggingFace on first run.
---
@@ -18,14 +17,17 @@ All local models (embedding, cross-encoder) are automatically downloaded from Hu
Used for fact extraction, entity resolution, mental model consolidation, and answer synthesis.
-**Supported providers:** OpenAI, Anthropic, Gemini, Groq, Ollama, LM Studio, and **any OpenAI-compatible API**
+**Supported providers:**
-:::tip OpenAI-Compatible Providers
+
+
+Also supports **any OpenAI-compatible API** (e.g., Azure OpenAI, Together AI, Fireworks).
+
+> **💡 OpenAI-Compatible Providers**
+>
Hindsight works with any provider that exposes an OpenAI-compatible API (e.g., Azure OpenAI). Simply set `HINDSIGHT_API_LLM_PROVIDER=openai` and configure `HINDSIGHT_API_LLM_BASE_URL` to point to your provider's endpoint.
See [Configuration](./configuration#llm-provider) for setup examples.
-:::
-
### Benchmarks
Not sure which model to use? The **[Model Leaderboard](https://benchmarks.hindsight.vectorize.io/)** benchmarks models across accuracy, speed, cost, and reliability for retain, reflect, and observation consolidation so you can pick the right trade-off for your use case.
@@ -63,6 +65,7 @@ Each provider has a recommended default model that's used when `HINDSIGHT_API_LL
| `anthropic` | `claude-haiku-4-5-20251001` |
| `gemini` | `gemini-2.5-flash` |
| `groq` | `openai/gpt-oss-120b` |
+| `minimax` | `MiniMax-M2.5` |
| `ollama` | `gemma3:12b` |
| `lmstudio` | `local-model` |
| `vertexai` | `gemini-2.0-flash-001` |
@@ -97,7 +100,8 @@ export HINDSIGHT_API_RETAIN_LLM_PROVIDER=anthropic
Other LLM models not listed above may work with Hindsight, but they must support **at least 65,000 output tokens** to ensure reliable fact extraction. If you need support for a specific model that doesn't meet this requirement, please [open an issue](https://github.com/hindsight-ai/hindsight/issues) to request an exception.
-:::tip Models with Limited Output Tokens
+> **💡 Models with Limited Output Tokens**
+>
If your model only supports 32k or fewer output tokens (e.g., some older models), you can reduce the retain completion token limit:
```bash
@@ -109,8 +113,6 @@ export HINDSIGHT_API_RETAIN_MAX_COMPLETION_TOKENS=16000
```
**Important:** `HINDSIGHT_API_RETAIN_MAX_COMPLETION_TOKENS` must be greater than `HINDSIGHT_API_RETAIN_CHUNK_SIZE` (default: 3000). The system will validate this on startup and provide an error message if the configuration is invalid.
-:::
-
### Configuration
```bash
@@ -144,6 +146,11 @@ export HINDSIGHT_API_LLM_PROVIDER=lmstudio
export HINDSIGHT_API_LLM_BASE_URL=http://localhost:1234/v1
export HINDSIGHT_API_LLM_MODEL=your-local-model
+# MiniMax (204K context window)
+export HINDSIGHT_API_LLM_PROVIDER=minimax
+export HINDSIGHT_API_LLM_API_KEY=your-minimax-api-key
+export HINDSIGHT_API_LLM_MODEL=MiniMax-M2.5
+
# Vertex AI (Google Cloud)
export HINDSIGHT_API_LLM_PROVIDER=vertexai
export HINDSIGHT_API_LLM_MODEL=gemini-2.0-flash-001
@@ -210,8 +217,8 @@ You can use any model supported by OpenAI Codex CLI
Use your Claude Pro or Max subscription for Hindsight without separate Anthropic API costs.
-
-:::warning Terms of Service Notice
+> **⚠️ Terms of Service Notice**
+>
This integration uses the Claude Agent SDK with your personal Claude Pro/Max subscription
credentials. You must be logged into Claude Code on your own machine before using this provider.
@@ -235,9 +242,6 @@ credentials. You must be logged into Claude Code on your own machine before usin
For production or team use, we recommend using `HINDSIGHT_API_LLM_PROVIDER=anthropic` with
an API key from the [Anthropic Console](https://console.anthropic.com/).
-:::
-
-
**Prerequisites:**
- Active Claude Pro or Max subscription
- Claude Code CLI installed
@@ -282,7 +286,6 @@ You can use any model supported by Claude Code CLI.
- Usage billed to your Claude subscription (not separate API costs)
- For personal development use only (see Claude Terms of Service)
-
---
### Vertex AI Setup (Google Cloud)
@@ -376,10 +379,9 @@ Converts text into dense vector representations for semantic similarity search.
| `embed-english-v3.0` | 1024 | English text |
| `embed-multilingual-v3.0` | 1024 | 100+ languages |
-:::warning Embedding Dimensions
+> **⚠️ Embedding Dimensions**
+>
Hindsight automatically detects the embedding dimension at startup and adjusts the database schema. Once memories are stored, you cannot change dimensions without losing data.
-:::
-
**Configuration Examples:**
```bash
diff --git a/skills/hindsight-docs/references/sdks/embed.md b/skills/hindsight-docs/references/sdks/embed.md
index f94dd8e1..7b1ba402 100644
--- a/skills/hindsight-docs/references/sdks/embed.md
+++ b/skills/hindsight-docs/references/sdks/embed.md
@@ -88,7 +88,7 @@ The daemon starts automatically on first use!
| Variable | Description | Default |
|----------|-------------|---------|
| `HINDSIGHT_EMBED_LLM_API_KEY` | **Required**. API key for LLM provider | - |
-| `HINDSIGHT_EMBED_LLM_PROVIDER` | LLM provider: `openai`, `anthropic`, `gemini`, `groq`, `ollama` | `openai` |
+| `HINDSIGHT_EMBED_LLM_PROVIDER` | LLM provider: `openai`, `anthropic`, `gemini`, `groq`, `minimax`, `ollama` | `openai` |
| `HINDSIGHT_EMBED_LLM_MODEL` | Model name | `gpt-4o-mini` |
| `HINDSIGHT_EMBED_BANK_ID` | Default memory bank ID | `default` |
| `HINDSIGHT_EMBED_DAEMON_IDLE_TIMEOUT` | Seconds before daemon auto-exits when idle (0 = never) | `300` |
diff --git a/skills/hindsight-docs/references/sdks/nodejs.md b/skills/hindsight-docs/references/sdks/nodejs.md
index 38d87132..767125eb 100644
--- a/skills/hindsight-docs/references/sdks/nodejs.md
+++ b/skills/hindsight-docs/references/sdks/nodejs.md
@@ -2,7 +2,7 @@
sidebar_position: 2
---
-# Node.js Client
+# TypeScript Client
Official TypeScript/JavaScript client for the Hindsight API.