fleet-memory/hindsight-api/hindsight_api/engine/retain
Nicolò Boschi 00ccf0b218
fix(consolidation): respect bank mission over ephemeral-state heuristic (#525)
* Add Hindsight as git subtree + BCGU noise filtering tests

Adds hindsight server source as a subtree under hindsight-api/ so we
can iterate on server-side fixes directly.

test_bcgu_noise_filtering.py proves that a well-crafted
retain_custom_instructions (BCGU_RETAIN_MISSION) can suppress
talking-head noise at fact extraction time — eliminating the need for
client-side --filter-vision-noise preprocessing.

Tests cover:
- Default mode extracts 3 noise facts from talking-head frame (problem documented)
- BCGU mission produces 0 noise facts from same talking-head frame
- BCGU mission still extracts 2 high-value ChatGPT screen facts correctly
- Mixed doc (2 talking-head + 2 screen): 0% noise ratio with BCGU mission
- Pure talking-head doc: 0 facts extracted

All 5 tests pass in ~32s using gpt-4o-mini.

* fix(consolidation): respect mission context over ephemeral-state heuristic

Two related fixes for the consolidation engine when a bank mission is
configured:

1. **Mission override for ephemeral-state filter** (`prompts.py`):
   The system prompt previously instructed the LLM to discard any fact
   that looked like "ephemeral state" (e.g. current position, transient
   actions).  When a mission is active the mission itself defines what is
   valuable — timestamped screen actions, session events, tool interactions
   may all be mission-critical even though they look ephemeral.  Added a
   MISSION OVERRIDE block that explicitly tells the LLM the mission takes
   priority over the generic ephemeral-state guidance.

2. **Remove contradictory durable-knowledge nudge** (`consolidator.py`):
   The user-prompt builder was injecting "Focus on DURABLE knowledge that
   serves this mission, not ephemeral state" alongside the mission text.
   This phrasing contradicted missions that intentionally capture
   timestamped events.  Replaced with a neutral directive that simply
   signals the mission overrides general rules.

3. **JSON control-character sanitisation** (`consolidator.py`):
   LLMs occasionally embed literal ASCII control characters (0x00–0x1f)
   inside JSON string values, causing `json.loads` to raise a
   JSONDecodeError.  Added a try/except that strips control characters
   and retries the parse before re-raising, preventing spurious failures.

* refactor(consolidation): move sanitize_llm_output to llm_wrapper, reuse in consolidator

- Add `sanitize_llm_output()` to `llm_wrapper.py` as the single canonical
  function for stripping characters that break downstream systems
  (ASCII control chars 0x00-0x08/0x0B-0x0C/0x0E-0x1F/0x7F and Unicode
  surrogates). Tab, newline, and carriage-return are preserved.
- Reduce `_sanitize_text()` in `fact_extraction.py` to a thin wrapper
  that delegates to `sanitize_llm_output()`.
- Update `consolidator.py` to import and call `sanitize_llm_output()`
  directly instead of reimplementing the logic inline.
- Remove test_bcgu_noise_filtering.py (should not have been committed).

* fix(consolidation): apply sanitize_llm_output to observation text fields

sanitize_llm_output was imported but unused after the old _call_llm_once
path was removed. The batch flow uses structured Pydantic output so
there's no raw json.loads call — instead, apply sanitization via
field_validator on _CreateAction.text and _UpdateAction.text so control
characters are stripped before observation text reaches the database.

* fix(entity-resolver): correct mention_count for new entities in batch retain

When the same entity (e.g. "Bob") appears across N items in a single batch
retain, _resolve_entities_batch_impl deduplicates them into one name group
before inserting, then queued only ONE _EntityStat regardless of N. The
flush therefore always incremented mention_count by 1 beyond the INSERT
value — giving 2 for any number of mentions.

Two-part fix:
- INSERT with mention_count=0 so the post-transaction flush is the single
  source of truth for the count (avoids an off-by-one for N=1 as well).
- Append one _EntityStat per original mention (len(g.indices)) instead of
  one per unique name, so flush_pending_stats() adds the correct total N.

This makes the batch path consistent with the single-entity path, which
already accumulates one stat per mention via entities_to_update.
2026-03-09 15:04:36 +01:00
..
__init__.py fix(performance): improve recall and retain performance on large banks (#469) 2026-03-03 13:35:22 +01:00
bank_utils.py feat: introduce mental models (#132) 2026-01-16 11:16:41 +01:00
chunk_storage.py feat: extensions (#54) 2025-12-22 11:05:23 +01:00
embedding_processing.py feat: entity labels — optional, free_values, multi_value, UI polish (#450) 2026-03-02 13:05:25 +01:00
embedding_utils.py fix(performance): improve recall and retain performance on large banks (#469) 2026-03-03 13:35:22 +01:00
entity_labels.py feat: entity labels — optional, free_values, multi_value, UI polish (#450) 2026-03-02 13:05:25 +01:00
entity_processing.py feat: entity labels — optional, free_values, multi_value, UI polish (#450) 2026-03-02 13:05:25 +01:00
fact_extraction.py fix(consolidation): respect bank mission over ephemeral-state heuristic (#525) 2026-03-09 15:04:36 +01:00
fact_storage.py feat: entity labels — optional, free_values, multi_value, UI polish (#450) 2026-03-02 13:05:25 +01:00
link_creation.py fix: doc build and lint files (#34) 2025-12-16 13:49:09 +01:00
link_utils.py fix(performance): improve recall and retain performance on large banks (#469) 2026-03-03 13:35:22 +01:00
orchestrator.py feat: webhook system with retain.completed event, UI, and docs (#487) 2026-03-04 14:17:01 +01:00
types.py feat: support timestamp="unset" to retain content without a date (#465) 2026-03-02 12:03:23 +01:00