* fix(reflect): prevent context_length_exceeded on large memory banks (#457) The reflect agent's agentic loop accumulated tool-call messages across iterations with no upper bound on token count, causing context_length_exceeded errors on banks with 19K+ nodes. Changes: - Add proactive token-budget guard: before each call_with_tools, count accumulated message tokens via tiktoken; if >= max_context_tokens and evidence has been gathered, immediately synthesize from what was found - Detect context-overflow errors specifically (_is_context_overflow_error) and skip the retry path — retrying after overflow only makes it worse - Truncate context_history in build_final_prompt to a 60K-token budget so the fallback synthesis prompt itself cannot overflow - Add HINDSIGHT_API_REFLECT_MAX_CONTEXT_TOKENS config (default 100000) wired through config.py → main.py → memory_engine → run_reflect_agent - Tests: unit tests for helpers + mock-LLM behavior tests + an end-to-end integration test using a real LLM with max_context_tokens=1 * fix(reflect): derive final prompt context budget from max_context_tokens Replace the hardcoded _FINAL_PROMPT_CONTEXT_BUDGET (60K tokens) with a fraction of max_context_tokens (80%), so the fallback synthesis prompt automatically scales with whatever context window is configured. |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| agent.py | ||
| models.py | ||
| observations.py | ||
| prompts.py | ||
| tools.py | ||
| tools_schema.py | ||