fleet-memory/hindsight-api/hindsight_api/engine
Nicolò Boschi eeb938fc65
fix: truncate documents exceeding LiteLLM reranker context limit (#549)
* fix: register embedded profiles in CLI metadata on daemon start

When HindsightEmbedded(profile="myapp") starts a daemon, the profile
was never written to metadata.json or given a .env file, making it
invisible to `hindsight-embed profile list` and other CLI commands.

Add _register_profile() to DaemonEmbedManager which saves HINDSIGHT_API_*
config to ~/.hindsight/profiles/{name}.env and registers the port in
metadata.json. Called after a successful new daemon start and when the
daemon is already running, so orphaned profiles also get registered on
next use.

* fix: truncate documents exceeding LiteLLM reranker context limit

Add HINDSIGHT_API_RERANKER_LITELLM_MAX_TOKENS_PER_DOC env var for both
litellm and litellm-sdk reranker providers. When set, documents are
truncated to the configured token limit using tiktoken (cl100k_base)
before being sent to the reranker, preventing BadRequestError for
models with small context windows (e.g. 1024-token limit).

* refactor: use shared _tiktoken_encoder for doc truncation in LiteLLM reranker

* refactor: use _get_tiktoken_encoding() consistently, remove eager module-level encoder instance

* doc: add HINDSIGHT_API_RERANKER_LITELLM_MAX_TOKENS_PER_DOC to configuration reference
2026-03-13 10:18:18 +01:00
..
consolidation fix: cancel in-flight async ops when bank is deleted (#545) 2026-03-12 09:39:31 +01:00
directives feat: revisit mental models, directives and reflections (#179) 2026-01-22 17:13:16 +01:00
parsers Fix Iris parser httpx read timeout for file uploads (#524) 2026-03-09 11:22:44 +01:00
providers feat: add MiniMax LLM provider support (#550) 2026-03-13 10:17:55 +01:00
reflect feat: entity labels — optional, free_values, multi_value, UI polish (#450) 2026-03-02 13:05:25 +01:00
retain perf: replace window-function retrieval with UNION ALL + per-bank HNSW indexes (#541) 2026-03-11 12:09:50 +01:00
search perf: replace window-function retrieval with UNION ALL + per-bank HNSW indexes (#541) 2026-03-11 12:09:50 +01:00
storage Fix GCS auth for Workload Identity Federation credentials (#518) 2026-03-07 08:59:51 +01:00
__init__.py feat: extensions (#54) 2025-12-22 11:05:23 +01:00
cross_encoder.py fix: truncate documents exceeding LiteLLM reranker context limit (#549) 2026-03-13 10:18:18 +01:00
db_budget.py fix: batch queries on recall (#149) 2026-01-13 13:20:22 +01:00
db_utils.py fix: strip null bytes from parsed file content before retain (#535) 2026-03-10 16:24:11 +01:00
embeddings.py fix: pass encoding_format="float" in LiteLLM embedding calls (#434) 2026-02-25 10:32:05 +01:00
entity_resolver.py fix(consolidation): respect bank mission over ephemeral-state heuristic (#525) 2026-03-09 15:04:36 +01:00
interface.py feat: add tags filtering and q description fix for list documents API (#468) 2026-03-02 17:03:16 +01:00
jina_mlx_reranker.py feat: add jina-mlx reranker provider for Apple Silicon (#542) 2026-03-11 15:15:58 +01:00
llm_interface.py feat: support Batch API for retain (openai/groq) (#365) 2026-02-16 13:31:50 +01:00
llm_wrapper.py feat: add MiniMax LLM provider support (#550) 2026-03-13 10:17:55 +01:00
memory_engine.py fix: truncate documents exceeding LiteLLM reranker context limit (#549) 2026-03-13 10:18:18 +01:00
operation_metadata.py fix: improve async batch retain with large payloads (#366) 2026-02-16 12:51:42 +01:00
query_analyzer.py fix(performance): improve recall and retain performance on large banks (#469) 2026-03-03 13:35:22 +01:00
response_models.py feat: include source facts in observation recall (#404) 2026-02-19 14:54:19 +01:00
task_backend.py fix: retain async fails if timestamp is set (#251) 2026-01-30 13:04:33 +01:00
utils.py feat: implement hierarchical configuration (system, tenant, bank) (#329) 2026-02-12 13:14:57 +01:00