fleet-memory/hindsight-api/hindsight_api
Nicolò Boschi 77defd96e9
fix(reflect): prevent context_length_exceeded on large memory banks (#462)
* fix(reflect): prevent context_length_exceeded on large memory banks (#457)

The reflect agent's agentic loop accumulated tool-call messages across
iterations with no upper bound on token count, causing
context_length_exceeded errors on banks with 19K+ nodes.

Changes:
- Add proactive token-budget guard: before each call_with_tools, count
  accumulated message tokens via tiktoken; if >= max_context_tokens and
  evidence has been gathered, immediately synthesize from what was found
- Detect context-overflow errors specifically (_is_context_overflow_error)
  and skip the retry path — retrying after overflow only makes it worse
- Truncate context_history in build_final_prompt to a 60K-token budget
  so the fallback synthesis prompt itself cannot overflow
- Add HINDSIGHT_API_REFLECT_MAX_CONTEXT_TOKENS config (default 100000)
  wired through config.py → main.py → memory_engine → run_reflect_agent
- Tests: unit tests for helpers + mock-LLM behavior tests + an
  end-to-end integration test using a real LLM with max_context_tokens=1

* fix(reflect): derive final prompt context budget from max_context_tokens

Replace the hardcoded _FINAL_PROMPT_CONTEXT_BUDGET (60K tokens) with
a fraction of max_context_tokens (80%), so the fallback synthesis prompt
automatically scales with whatever context window is configured.
2026-03-02 12:03:12 +01:00
..
admin feat: new 'worker' service (#176) 2026-01-20 10:17:56 +01:00
alembic feat: observation_scopes field to drive observations granularity (#447) 2026-02-28 10:46:21 +01:00
api Add bank-scoped validation to engine and HTTP handlers (#454) 2026-03-02 09:55:21 +01:00
engine fix(reflect): prevent context_length_exceeded on large memory banks (#462) 2026-03-02 12:03:12 +01:00
extensions Add bank-scoped validation to engine and HTTP handlers (#454) 2026-03-02 09:55:21 +01:00
worker fix: resolve consolidation deadlock caused by zombie processing tasks on retry (#463) 2026-03-02 11:45:41 +01:00
__init__.py Release v0.4.14 2026-02-26 18:03:14 +01:00
banner.py feat: support for pgvectorscale (DiskANN) (#378) 2026-02-16 14:19:56 +01:00
config.py fix(reflect): prevent context_length_exceeded on large memory banks (#462) 2026-03-02 12:03:12 +01:00
config_resolver.py Fix bank config API for multi-tenant schema isolation (#417) 2026-02-20 23:43:52 +01:00
daemon.py feat: improve openclaw and hindisght-embed params (#279) 2026-02-03 09:39:04 +01:00
main.py fix(reflect): prevent context_length_exceeded on large memory banks (#462) 2026-03-02 12:03:12 +01:00
mcp_local.py fix(mcp): unify hindsight-mcp-local and server mcp (#407) 2026-02-19 17:57:17 +01:00
mcp_tools.py Add bank-scoped validation to engine and HTTP handlers (#454) 2026-03-02 09:55:21 +01:00
metrics.py chore: remove dead code (#245) 2026-01-30 09:16:32 +01:00
migrations.py fix: raise error when embedding dimensions exceed pgvector HNSW limit (#361) 2026-02-26 10:24:17 +01:00
models.py Add bank-scoped validation to engine and HTTP handlers (#454) 2026-03-02 09:55:21 +01:00
pg0.py feat: support vertex as llm provider (#233) 2026-01-29 16:13:57 -05:00
server.py Fix: Load extensions in server.py for multi-worker deployments (#155) 2026-01-13 17:55:33 +01:00
tracing.py feat: add otel traceability (#330) 2026-02-10 12:20:48 +01:00
utils.py fix(helm): improve appVersion usage (#326) 2026-02-09 11:35:08 +01:00