diff --git a/hindsight-docs/docs/developer/models.md b/hindsight-docs/docs/developer/models.md index 85e24ff0..889cd77f 100644 --- a/hindsight-docs/docs/developer/models.md +++ b/hindsight-docs/docs/developer/models.md @@ -6,14 +6,70 @@ Hindsight uses several machine learning models for different tasks. | Model Type | Purpose | Default | Configurable | |------------|---------|---------|--------------| +| **LLM** | Fact extraction, reasoning, generation | Provider-specific | Yes | | **Embedding** | Vector representations for semantic search | `BAAI/bge-small-en-v1.5` | Yes | | **Cross-Encoder** | Reranking search results | `cross-encoder/ms-marco-MiniLM-L-6-v2` | Yes | -| **LLM** | Fact extraction, reasoning, generation | Provider-specific | Yes | All local models (embedding, cross-encoder) are automatically downloaded from HuggingFace on first run. --- +## LLM + +Used for fact extraction, entity resolution, opinion generation, and answer synthesis. + +**Supported providers:** OpenAI, Gemini, Groq, Ollama + +### Tested Models + +The following models have been tested and verified to work correctly with Hindsight: + +| Provider | Model | +|----------|-------| +| **OpenAI** | `gpt-5` | +| **OpenAI** | `gpt-5-mini` | +| **OpenAI** | `gpt-5-nano` | +| **OpenAI** | `gpt-4.1-mini` | +| **OpenAI** | `gpt-4.1-nano` | +| **OpenAI** | `gpt-4o-mini` | +| **Gemini** | `gemini-2.5-flash` | +| **Gemini** | `gemini-2.5-flash-lite` | +| **Groq** | `openai/gpt-oss-120b` | +| **Groq** | `openai/gpt-oss-20b` | +| **Groq** | `llama-3.3-70b-versatile` | + +### Using Other Models + +Other LLM models not listed above may work with Hindsight, but they must support **at least 65,000 output tokens** to ensure reliable fact extraction. If you need support for a specific model that doesn't meet this requirement, please [open an issue](https://github.com/hindsight-ai/hindsight/issues) to request an exception. + +### Configuration + +```bash +# Groq (recommended) +export HINDSIGHT_API_LLM_PROVIDER=groq +export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx +export HINDSIGHT_API_LLM_MODEL=openai/gpt-oss-20b + +# OpenAI +export HINDSIGHT_API_LLM_PROVIDER=openai +export HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx +export HINDSIGHT_API_LLM_MODEL=gpt-4o + +# Gemini +export HINDSIGHT_API_LLM_PROVIDER=gemini +export HINDSIGHT_API_LLM_API_KEY=xxxxxxxxxxxx +export HINDSIGHT_API_LLM_MODEL=gemini-2.0-flash + +# Ollama (local) +export HINDSIGHT_API_LLM_PROVIDER=ollama +export HINDSIGHT_API_LLM_BASE_URL=http://localhost:11434/v1 +export HINDSIGHT_API_LLM_MODEL=llama3.1 +``` + +**Note:** The LLM is the primary bottleneck for retain operations. See [Performance](./performance) for optimization strategies. + +--- + ## Embedding Model Converts text into dense vector representations for semantic similarity search. @@ -22,14 +78,13 @@ Converts text into dense vector representations for semantic similarity search. **Alternatives:** -| Model | Dimensions | Use Case | -|-------|------------|----------| -| `BAAI/bge-small-en-v1.5` | 384 | Default, fast, good quality | -| `BAAI/bge-base-en-v1.5` | 768 | Higher accuracy, slower | -| `sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2` | 384 | Multilingual (50+ languages) | +| Model | Use Case | +|-------|----------| +| `BAAI/bge-small-en-v1.5` | Default, fast, good quality | +| `sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2` | Multilingual (50+ languages) | :::warning -All embedding models must produce 384-dimensional vectors to match the database schema. +All embedding models must produce **384-dimensional vectors** to match the database schema. ::: **Configuration:** @@ -71,44 +126,3 @@ export HINDSIGHT_API_RERANKER_LOCAL_MODEL=cross-encoder/ms-marco-MiniLM-L-6-v2 export HINDSIGHT_API_RERANKER_PROVIDER=tei export HINDSIGHT_API_RERANKER_TEI_URL=http://localhost:8081 ``` - ---- - -## LLM - -Used for fact extraction, entity resolution, opinion generation, and answer synthesis. - -**Supported providers:** Groq, OpenAI, Gemini, Ollama - -| Provider | Recommended Model | Best For | -|----------|------------------|----------| -| **Groq** | `openai/gpt-oss-20b` | Fast inference, high throughput (recommended) | -| **OpenAI** | `gpt-4o` | Good quality | -| **Gemini** | `gemini-2.0-flash` | Good quality, cost effective | -| **Ollama** | `llama3.1` | Local deployment, privacy | - -**Configuration:** - -```bash -# Groq (recommended) -export HINDSIGHT_API_LLM_PROVIDER=groq -export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx -export HINDSIGHT_API_LLM_MODEL=openai/gpt-oss-20b - -# OpenAI -export HINDSIGHT_API_LLM_PROVIDER=openai -export HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx -export HINDSIGHT_API_LLM_MODEL=gpt-4o - -# Gemini -export HINDSIGHT_API_LLM_PROVIDER=gemini -export HINDSIGHT_API_LLM_API_KEY=xxxxxxxxxxxx -export HINDSIGHT_API_LLM_MODEL=gemini-2.0-flash - -# Ollama (local) -export HINDSIGHT_API_LLM_PROVIDER=ollama -export HINDSIGHT_API_LLM_BASE_URL=http://localhost:11434/v1 -export HINDSIGHT_API_LLM_MODEL=llama3.1 -``` - -**Note:** The LLM is the primary bottleneck for retain operations. See [Performance](./performance) for optimization strategies. diff --git a/hindsight-docs/static/llms-full.txt b/hindsight-docs/static/llms-full.txt index aff377c9..3a62968b 100644 --- a/hindsight-docs/static/llms-full.txt +++ b/hindsight-docs/static/llms-full.txt @@ -3,7 +3,7 @@ > Agent Memory that Works Like Human Memory This file contains the complete Hindsight documentation for LLM consumption. -Generated: 2025-12-15T09:47:27.854Z +Generated: 2025-12-15T10:18:45.836Z --- @@ -2750,14 +2750,70 @@ Hindsight uses several machine learning models for different tasks. | Model Type | Purpose | Default | Configurable | |------------|---------|---------|--------------| +| **LLM** | Fact extraction, reasoning, generation | Provider-specific | Yes | | **Embedding** | Vector representations for semantic search | `BAAI/bge-small-en-v1.5` | Yes | | **Cross-Encoder** | Reranking search results | `cross-encoder/ms-marco-MiniLM-L-6-v2` | Yes | -| **LLM** | Fact extraction, reasoning, generation | Provider-specific | Yes | All local models (embedding, cross-encoder) are automatically downloaded from HuggingFace on first run. --- +## LLM + +Used for fact extraction, entity resolution, opinion generation, and answer synthesis. + +**Supported providers:** OpenAI, Gemini, Groq, Ollama + +### Tested Models + +The following models have been tested and verified to work correctly with Hindsight: + +| Provider | Model | +|----------|-------| +| **OpenAI** | `gpt-5` | +| **OpenAI** | `gpt-5-mini` | +| **OpenAI** | `gpt-5-nano` | +| **OpenAI** | `gpt-4.1-mini` | +| **OpenAI** | `gpt-4.1-nano` | +| **OpenAI** | `gpt-4o-mini` | +| **Gemini** | `gemini-2.5-flash` | +| **Gemini** | `gemini-2.5-flash-lite` | +| **Groq** | `openai/gpt-oss-120b` | +| **Groq** | `openai/gpt-oss-20b` | +| **Groq** | `llama-3.3-70b-versatile` | + +### Using Other Models + +Other LLM models not listed above may work with Hindsight, but they must support **at least 65,000 output tokens** to ensure reliable fact extraction. If you need support for a specific model that doesn't meet this requirement, please [open an issue](https://github.com/hindsight-ai/hindsight/issues) to request an exception. + +### Configuration + +```bash +# Groq (recommended) +export HINDSIGHT_API_LLM_PROVIDER=groq +export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx +export HINDSIGHT_API_LLM_MODEL=openai/gpt-oss-20b + +# OpenAI +export HINDSIGHT_API_LLM_PROVIDER=openai +export HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx +export HINDSIGHT_API_LLM_MODEL=gpt-4o + +# Gemini +export HINDSIGHT_API_LLM_PROVIDER=gemini +export HINDSIGHT_API_LLM_API_KEY=xxxxxxxxxxxx +export HINDSIGHT_API_LLM_MODEL=gemini-2.0-flash + +# Ollama (local) +export HINDSIGHT_API_LLM_PROVIDER=ollama +export HINDSIGHT_API_LLM_BASE_URL=http://localhost:11434/v1 +export HINDSIGHT_API_LLM_MODEL=llama3.1 +``` + +**Note:** The LLM is the primary bottleneck for retain operations. See [Performance](./performance) for optimization strategies. + +--- + ## Embedding Model Converts text into dense vector representations for semantic similarity search. @@ -2766,14 +2822,13 @@ Converts text into dense vector representations for semantic similarity search. **Alternatives:** -| Model | Dimensions | Use Case | -|-------|------------|----------| -| `BAAI/bge-small-en-v1.5` | 384 | Default, fast, good quality | -| `BAAI/bge-base-en-v1.5` | 768 | Higher accuracy, slower | -| `sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2` | 384 | Multilingual (50+ languages) | +| Model | Use Case | +|-------|----------| +| `BAAI/bge-small-en-v1.5` | Default, fast, good quality | +| `sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2` | Multilingual (50+ languages) | :::warning -All embedding models must produce 384-dimensional vectors to match the database schema. +All embedding models must produce **384-dimensional vectors** to match the database schema. ::: **Configuration:** @@ -2816,47 +2871,6 @@ export HINDSIGHT_API_RERANKER_PROVIDER=tei export HINDSIGHT_API_RERANKER_TEI_URL=http://localhost:8081 ``` ---- - -## LLM - -Used for fact extraction, entity resolution, opinion generation, and answer synthesis. - -**Supported providers:** Groq, OpenAI, Gemini, Ollama - -| Provider | Recommended Model | Best For | -|----------|------------------|----------| -| **Groq** | `openai/gpt-oss-20b` | Fast inference, high throughput (recommended) | -| **OpenAI** | `gpt-4o` | Good quality | -| **Gemini** | `gemini-2.0-flash` | Good quality, cost effective | -| **Ollama** | `llama3.1` | Local deployment, privacy | - -**Configuration:** - -```bash -# Groq (recommended) -export HINDSIGHT_API_LLM_PROVIDER=groq -export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx -export HINDSIGHT_API_LLM_MODEL=openai/gpt-oss-20b - -# OpenAI -export HINDSIGHT_API_LLM_PROVIDER=openai -export HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx -export HINDSIGHT_API_LLM_MODEL=gpt-4o - -# Gemini -export HINDSIGHT_API_LLM_PROVIDER=gemini -export HINDSIGHT_API_LLM_API_KEY=xxxxxxxxxxxx -export HINDSIGHT_API_LLM_MODEL=gemini-2.0-flash - -# Ollama (local) -export HINDSIGHT_API_LLM_PROVIDER=ollama -export HINDSIGHT_API_LLM_BASE_URL=http://localhost:11434/v1 -export HINDSIGHT_API_LLM_MODEL=llama3.1 -``` - -**Note:** The LLM is the primary bottleneck for retain operations. See [Performance](./performance) for optimization strategies. - ---