* feat: introduce hindsight-api-slim and hindsight-all-slim packages
Closes#552
- Move all source code from hindsight-api/ to new hindsight-api-slim/
- hindsight-api-slim has heavy ML deps (torch, sentence-transformers,
transformers, einops, flashrank, mlx, mlx-lm, safetensors) and
pg0-embedded as optional extras: [local-ml], [embedded-db], [all]
- hindsight-api becomes a zero-code meta-package depending on
hindsight-api-slim[all] for full backward compatibility
- Add hindsight-all-slim meta-package: hindsight-api-slim + client + embed
- hindsight-all updated to depend on hindsight-api-slim[all]
- pg0.py: lazy-import pg0 with clear ImportError pointing to [embedded-db]
- Dockerfile: replace sed hack with proper uv sync --extra flags
- Update release.yml, test.yml, lint.sh, release.sh, CLAUDE.md and
all path references throughout the repo
* refactor: rename hindsight/ directory to hindsight-all/
* docs: document hindsight-api-slim and hindsight-all-slim package variants
Add package variants table and extras explanation to installation.md
* docs: remove emojis from installation.md, use professional tone
* docs: link Docker slim variant to pip package variants section
* docs: consolidate Docker image variants into single table
* ci: fix working-directory paths after package restructure
- Replace all hindsight-api → hindsight-api-slim in test.yml
- Replace hindsight → hindsight-all in test.yml
- Add --extra embedded-db to test-embed API install step
* ci: add local-ml and embedded-db extras to API sync steps
These extras were previously implicit in the old hindsight-api package
(which bundled everything). Now that hindsight-api-slim uses optional
extras, we must explicitly request local-ml and embedded-db in CI.
* ci: add API install step with embedded-db to test-embed smoke test
The smoke test starts hindsight-api as a daemon, which requires pg0-embedded.
Add a dedicated install step for hindsight-api-slim with embedded-db extra
so the daemon can start successfully.
* ci: remove --no-install-project when using optional extras
When --no-install-project is combined with --extra, the optional deps
are not installed because extras require the project to be active.
Remove --no-install-project from steps that need local-ml or embedded-db.
* ci: fix ordering of uv sync steps to preserve optional extras
When uv sync runs for a different workspace member, it removes optional
extras installed for other members. Fix by always running extra-requiring
API sync last, after other workspace member syncs.
Also remove --no-install-project from embedded-db sync in test-embed,
as --no-install-project prevents optional extras from being active.
* ci: add local-ml extra to test-embed API install for smoke test
The smoke test starts the full API server which needs sentence-transformers
for local embeddings (default provider). Add local-ml extra to the install.
* ci: simplify extras with --all-extras and add slim pip smoke test
- Replace explicit --extra local-ml --extra embedded-db with --all-extras
for cleaner, more maintainable sync steps
- Add test-pip-slim job: tests hindsight-api-slim[embedded-db] without
local ML models, using Cohere for embeddings/reranking (mirrors Docker
slim smoke test approach)
* ci: simplify slim smoke test to health check only (mirrors Docker test)
* feat: add OAuth extension hooks for MCP authentication
Add extension points in core that allow cloud extensions to support
OAuth 2.1 (RFC 9728 / RFC 7591) for MCP server authentication:
- HttpExtension.get_root_router() for well-known endpoint mounting
- AuthenticationError.headers for WWW-Authenticate propagation
- MCP middleware forwards auth error headers to clients
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* docs: document get_root_router and AuthenticationError.headers
Add documentation for the new extension points introduced in the
OAuth extension hooks commit.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* Remove OAuth-specific wording from extension docs
Make the AuthenticationError headers example generic instead of
OAuth-specific, since these are general-purpose extension hooks.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* feat: add bank-scoped validation to engine methods and HTTP handlers
Add validate_bank_read/validate_bank_write hooks to all bank-scoped
engine methods so the operation validator can enforce per-bank API key
restrictions. Add OperationValidationError handling to HTTP handlers
and MCP tools to return proper 403 responses. Add allowed_bank_ids
field to RequestContext.
* Add OperationValidationError handling to mental model GET and DELETE endpoints
Add directives, memory browsing, documents, operations, tags, and bank
management tools to the MCP server. Expose previously hardcoded parameters
(budget, types, tags, response_schema, trigger) on retain, recall, reflect,
and mental model tools. Update docs for all new tools and parameters.
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* Fix MCP extra args rejection and bank ID resolution priority
Two fixes to the MCP middleware:
1. Strip unknown tool arguments: LLMs frequently add extra fields
like "explanation" to tool calls. FastMCP's Pydantic TypeAdapter
rejects these with "Unexpected keyword argument". The middleware
now intercepts tools/call requests and removes unknown fields
before they reach validation.
2. Bank ID resolution priority: Path now takes priority over header.
Previously X-Bank-Id header was checked first, meaning /mcp/my-bank/
with X-Bank-Id: other-bank would silently use other-bank in multi-bank
mode. Now the URL path is authoritative — single-bank mode connections
cannot be overridden by headers.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* docs: update MCP server docs with mental model tools and fixes
- Add all mental model tools (create, list, get, update, delete, refresh)
- Add list_banks and create_bank tool docs
- Document single-bank vs multi-bank modes
- Fix bank selection priority: path > header > default
- Add Accept header to curl example
- Add timestamp param to retain, max_tokens to recall
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* fix: move mental model usage metering into engine for MCP support
Mental model validation hooks (validate_mental_model_get, validate_mental_model_refresh)
were only called in REST HTTP handlers, not in the engine. MCP tools call engine methods
directly, so usage metering was skipped entirely for MCP mental model operations.
Moved pre-validation and post-completion hooks into memory_engine.py (matching the
retain/recall/reflect pattern) and removed the duplicate code from http.py.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: remove double validation from create_mental_model and add internal checks
- Remove pre-validation from create_mental_model since callers always call
submit_async_refresh_mental_model next (which validates), preventing
double credit checks
- Add is_internal checks to mental model metering validators (matching
the existing pattern for recall/reflect) so background worker tasks
skip billing
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: prevent 307 redirect on /mcp that breaks MCP tool discovery
Starlette's Mount class redirects /mcp to /mcp/ with a 307 Temporary
Redirect. Many MCP clients don't follow POST redirects, which causes
tool discovery to fail (0 tools discovered despite successful auth).
Add _MCPPathRewriteMiddleware that rewrites /mcp to /mcp/ at the ASGI
level before routing, preventing the redirect entirely. Both /mcp and
/mcp/ now work identically.
Add regression test test_mcp_no_trailing_slash_works to verify URLs
with and without trailing slashes discover tools correctly.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* harden MCP server for real-world usage
- Remove MCP_ENDPOINTS blocklist so banks named "sse"/"messages" route correctly
- Scope SSE body rewriting to text/event-stream responses only to prevent data corruption
- Add _validate_mental_model_inputs for name, source_query, max_tokens validation in MCP tools
- Improve "not found" error messages to include bank_id context
- Fix fragile tool count assertions (exact → minimum bounds)
- Add integration tests: tool execution, input validation, edge-case bank names
- Add unit tests for validation helper and tool-level validation
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor: replace Mount + rewrite middleware with wrapping middleware
Starlette's Mount class redirects /mcp -> /mcp/ with 307, which MCP clients
don't follow. Previously we patched this with _MCPPathRewriteMiddleware.
Now MCPMiddleware wraps the FastAPI app directly via add_middleware, intercepting
/mcp* requests before they reach Starlette's router. No Mount means no redirect.
- Remove _MCPPathRewriteMiddleware (no longer needed)
- Remove app.mount() call
- Add prefix parameter to MCPMiddleware
- Use app.add_middleware() for proper Starlette integration
- Simplify path stripping (just remove prefix, no mount/root_path handling)
- Update routing test to match current behavior (no MCP_ENDPOINTS blocklist)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: update stale docstring referencing removed _MCPPathRewriteMiddleware
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* Add mental model CRUD tools to MCP server
Expose mental models (pinned reflections) as 6 new MCP tools:
- list_mental_models: List with optional tag filtering
- get_mental_model: Get by ID
- create_mental_model: Create with async content generation
- update_mental_model: Update name/source_query/tags
- delete_mental_model: Delete by ID
- refresh_mental_model: Re-run source query to update content
Both multi-bank (bank_id param) and single-bank modes supported,
following the same patterns as existing retain/recall/reflect tools.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: include mental model tools in single-bank MCP mode and update tests
The single-bank mode tool set was hardcoded to only retain/recall/reflect,
excluding the new mental model tools. Updated all 3 test layers (unit,
routing, HTTP integration) to assert mental model tool exposure.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: update extension test tool count for mental model tools
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: move mental model usage metering into engine for MCP support
Mental model validation hooks (validate_mental_model_get, validate_mental_model_refresh)
were only called in REST HTTP handlers, not in the engine. MCP tools call engine methods
directly, so usage metering was skipped entirely for MCP mental model operations.
Moved pre-validation and post-completion hooks into memory_engine.py (matching the
retain/recall/reflect pattern) and removed the duplicate code from http.py.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: remove double validation from create_mental_model and add internal checks
- Remove pre-validation from create_mental_model since callers always call
submit_async_refresh_mental_model next (which validates), preventing
double credit checks
- Add is_internal checks to mental model metering validators (matching
the existing pattern for recall/reflect) so background worker tasks
skip billing
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
MCP middleware was discarding tenant_id and api_key_id after authentication.
The authenticate_mcp() call mutated a RequestContext with these fields, but
tools later created a fresh RequestContext without them. This caused
UsageMeteringValidator to see tenant_id="unknown" and skip billing entirely.
Propagate tenant_id and api_key_id via ContextVars (same pattern as bank_id
and api_key) so the RequestContext passed to the memory engine has the full
auth context needed for usage tracking.
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* feat: improve mcp tools based on endpoint
* feat: improve mcp tools based on endpoint
* test: add integration test for MCP endpoint routing
- Add test_mcp_endpoint_routing.py to verify single-bank vs multi-bank tool exposure
- Verifies /mcp/ exposes all tools with bank_id parameters
- Verifies /mcp/{bank_id}/ only exposes scoped tools without bank_id parameters
- Regression test for issue #317
Related: #317, #318
* test: use StreamableHTTP client for MCP endpoint routing test
Replace httpx AsyncClient SSE parsing with proper MCP StreamableHTTP
client. This correctly tests the MCP server using the actual protocol
that clients will use.
Fixes#317
* feat: add TenantExtension auth to MCP endpoint
Replace static MCP_AUTH_TOKEN check with TenantExtension authentication,
making MCP use the same auth path as REST API.
- MCPMiddleware now calls tenant_extension.authenticate()
- Sets _current_schema from TenantContext for multi-tenant isolation
- Returns 401 on AuthenticationError (same as REST API)
- DefaultTenantExtension: no auth (local dev)
- ApiKeyTenantExtension: validates against env var
- CloudTenantExtension: HMAC + DB lookup (production)
Adds tests for middleware auth rejection, acceptance, and schema routing.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Address PR review: backwards compatibility for MCP auth
- Keep MCP_AUTH_TOKEN env var for legacy MCP servers
- Add authenticate_mcp() method to TenantExtension base class
- Default implementation calls authenticate()
- Extensions can override to opt-out of MCP auth
- Add mcp_auth_disabled config option to ApiKeyTenantExtension
- Set HINDSIGHT_API_TENANT_MCP_AUTH_DISABLED=true to skip MCP auth
- Remove CloudTenantExtension from public docstring
- Add tests for legacy auth token and mcp_auth_disabled flag
- Update MCP docs with new auth configuration
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Add search_docs MCP tool for documentation search
Implements a new MCP tool that searches Hindsight documentation using
Vectorize RAG pipelines. The tool supports:
- Searching core (OSS) docs, cloud docs, or both
- Configurable number of results (1-10)
- Returns ranked results with URLs, similarity scores, and text snippets
New environment variables:
- HINDSIGHT_API_VECTORIZE_ORG_ID
- HINDSIGHT_API_VECTORIZE_API_TOKEN
- HINDSIGHT_API_VECTORIZE_CORE_PIPELINE_ID
- HINDSIGHT_API_VECTORIZE_CLOUD_PIPELINE_ID
- HINDSIGHT_API_VECTORIZE_API_BASE_URL
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Add documentation for search_docs MCP tool
- Add Vectorize environment variables to configuration.md
- Add search_docs tool to MCP server available tools
- Add reflect tool documentation (was missing)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Add tests for search_docs MCP tool
Tests cover:
- DocsSource enum values and parsing
- _clean_text HTML stripping helper
- _search_vectorize_pipeline with mocked httpx
- Tool registration and function execution
- Source filtering (core/cloud/all)
- Result sorting by similarity
- Error handling for pipeline failures
- HTML cleaning in results
- Invalid source defaulting to 'all'
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Move search_docs to hindsight-cloud, add MCPExtension pattern
- Add MCPExtension base class for registering additional MCP tools
- Load MCPExtension in create_mcp_server when configured
- Remove search_docs tool (moved to hindsight-cloud CloudMCPExtension)
- Remove Vectorize config from hindsight-core
- Add tests for MCPExtension pattern
- Update docs to remove search_docs references
The MCPExtension pattern allows cloud (or any extension package) to
register additional MCP tools via:
HINDSIGHT_API_MCP_EXTENSION=package.module:ExtensionClass
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Address PR review feedback
- Remove CloudTenantExtension mention from MCPMiddleware docstring
- Fix docs: clarify that ApiKeyTenantExtension must be explicitly enabled
- Revert changes to versioned docs (0.3 and 0.4) - synced automatically on release
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Format mcp.py line length
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
* feat(mcp): add Bearer token authentication support
Add HINDSIGHT_API_MCP_AUTH_TOKEN environment variable to enable
authentication for MCP endpoint. When set, all requests must include
a valid Authorization header (Bearer token or direct token).
If not set, MCP endpoint remains open for backwards compatibility
with local development environments.
* fix: propagate Bearer token from MCP middleware to tools for tenant auth
MCP tools were creating RequestContext() without api_key, causing
"Invalid API key" errors when tenant extension validates requests.
Now the Bearer token is extracted in middleware, stored in a context
variable, and passed through to all MCP tool RequestContext instances.
* feat(mcp): add async_processing parameter to retain tool
Add async_processing parameter (default: True) to the MCP retain tool
to allow non-blocking memory storage. When True, memories are queued
for background processing and the tool returns immediately. When False,
the tool waits for completion before returning.
This matches the async behavior available in the HTTP API.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat(mcp): add list_memories and reflect tools
Add two missing MCP tools to achieve feature parity with HTTP API:
- list_memories: browse memories with pagination and full-text search
(equivalent to GET /memories/list)
- reflect: LLM-based reasoning over memories with disposition awareness
(equivalent to POST /reflect)
Both tools follow the existing pattern with JSON string responses
and proper error handling.
* docs: improve CLAUDE.md with detailed architecture info
- Add memory types explanation (world, experience, opinion, observation)
- Document retain/ and search/ submodule structure
- Add commands for single test run, ruff format, ty type checking
- Note MCP server implementation in API layer
- Add optional environment variables section
- Clarify conventions (no Python files at root, npm workspaces)
* chore: add .mcp.json and .osgrep to gitignore
These are user-specific development tool configs that should not be committed.
* changes
* refactor(mcp): remove list_memories tool
The list_memories endpoint is for debugging/exploration, not agent use.
Agents should use recall for semantic search instead.
Feedback from maintainer: "this tool is misleading for the agent,
it should use recall, the list method is mostly for debugging and
exploration, not for real usage"
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* refactor(mcp): remove list_banks and create_bank tools
These admin/orchestration tools are not needed for typical agent usage.
Agents work with a single configured bank via X-Bank-Id header.
MCP now exposes only core memory operations:
- retain: store memories
- recall: semantic search
- reflect: LLM reasoning over memories
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
---------
Co-authored-by: Anton Evseev <a.evseev@xsolla.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
* misc: add mcp integration tests and increase test coverage
* misc: add mcp integration tests and increase test coverage
* misc: add mcp integration tests and increase test coverage
* feat(mcp): Add multi-bank access and new MCP tools
Enables orchestrator agents to access multiple memory banks from a
single MCP connection, with new tools for bank management.
## New MCP Tools
- `reflect` - Thoughtful analysis using bank's personality and memories
- `list_banks` - Discover all available memory banks
- `create_bank` - Create new banks programmatically
## Multi-Bank Access
- Added optional `bank_id` parameter to `retain`, `recall`, `reflect`
- Allows cross-bank operations from a single MCP session
- Defaults to session bank if not specified
## Claude Code Compatibility
- Enabled `stateless_http=True` for proper Claude Code integration
- Responses now include `bank_id` for transparency
## Documentation
- Added docker-compose.example.yml with env var substitution
- Added HINDSIGHT-DOCKER.md setup guide with volume persistence docs
- Updated .gitignore to exclude local docker-compose.yml
## Use Case
Orchestrator agents can now:
- Maintain a private meta-orchestration bank
- Access shared project knowledge banks
- Query across banks for cross-context insights
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Address PR review feedback: remove docker files, improve reflect description
- Remove HINDSIGHT-DOCKER.md and docker-compose.example.yml per reviewer request
- Improve reflect tool description with clearer guidance for AI agents:
- Added "WHEN TO USE THIS TOOL" section
- Added "EXAMPLES OF GOOD QUERIES" with concrete use cases
- Added "HOW IT DIFFERS FROM RECALL" to clarify when to use each tool
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>