* Add Hindsight as git subtree + BCGU noise filtering tests Adds hindsight server source as a subtree under hindsight-api/ so we can iterate on server-side fixes directly. test_bcgu_noise_filtering.py proves that a well-crafted retain_custom_instructions (BCGU_RETAIN_MISSION) can suppress talking-head noise at fact extraction time — eliminating the need for client-side --filter-vision-noise preprocessing. Tests cover: - Default mode extracts 3 noise facts from talking-head frame (problem documented) - BCGU mission produces 0 noise facts from same talking-head frame - BCGU mission still extracts 2 high-value ChatGPT screen facts correctly - Mixed doc (2 talking-head + 2 screen): 0% noise ratio with BCGU mission - Pure talking-head doc: 0 facts extracted All 5 tests pass in ~32s using gpt-4o-mini. * fix(consolidation): respect mission context over ephemeral-state heuristic Two related fixes for the consolidation engine when a bank mission is configured: 1. **Mission override for ephemeral-state filter** (`prompts.py`): The system prompt previously instructed the LLM to discard any fact that looked like "ephemeral state" (e.g. current position, transient actions). When a mission is active the mission itself defines what is valuable — timestamped screen actions, session events, tool interactions may all be mission-critical even though they look ephemeral. Added a MISSION OVERRIDE block that explicitly tells the LLM the mission takes priority over the generic ephemeral-state guidance. 2. **Remove contradictory durable-knowledge nudge** (`consolidator.py`): The user-prompt builder was injecting "Focus on DURABLE knowledge that serves this mission, not ephemeral state" alongside the mission text. This phrasing contradicted missions that intentionally capture timestamped events. Replaced with a neutral directive that simply signals the mission overrides general rules. 3. **JSON control-character sanitisation** (`consolidator.py`): LLMs occasionally embed literal ASCII control characters (0x00–0x1f) inside JSON string values, causing `json.loads` to raise a JSONDecodeError. Added a try/except that strips control characters and retries the parse before re-raising, preventing spurious failures. * refactor(consolidation): move sanitize_llm_output to llm_wrapper, reuse in consolidator - Add `sanitize_llm_output()` to `llm_wrapper.py` as the single canonical function for stripping characters that break downstream systems (ASCII control chars 0x00-0x08/0x0B-0x0C/0x0E-0x1F/0x7F and Unicode surrogates). Tab, newline, and carriage-return are preserved. - Reduce `_sanitize_text()` in `fact_extraction.py` to a thin wrapper that delegates to `sanitize_llm_output()`. - Update `consolidator.py` to import and call `sanitize_llm_output()` directly instead of reimplementing the logic inline. - Remove test_bcgu_noise_filtering.py (should not have been committed). * fix(consolidation): apply sanitize_llm_output to observation text fields sanitize_llm_output was imported but unused after the old _call_llm_once path was removed. The batch flow uses structured Pydantic output so there's no raw json.loads call — instead, apply sanitization via field_validator on _CreateAction.text and _UpdateAction.text so control characters are stripped before observation text reaches the database. * fix(entity-resolver): correct mention_count for new entities in batch retain When the same entity (e.g. "Bob") appears across N items in a single batch retain, _resolve_entities_batch_impl deduplicates them into one name group before inserting, then queued only ONE _EntityStat regardless of N. The flush therefore always incremented mention_count by 1 beyond the INSERT value — giving 2 for any number of mentions. Two-part fix: - INSERT with mention_count=0 so the post-transaction flush is the single source of truth for the count (avoids an off-by-one for N=1 as well). - Append one _EntityStat per original mention (len(g.indices)) instead of one per unique name, so flush_pending_stats() adds the correct total N. This makes the batch path consistent with the single-entity path, which already accumulates one stat per mention via entities_to_update.
83 lines
5 KiB
Python
83 lines
5 KiB
Python
"""Prompts for the consolidation engine."""
|
|
|
|
# Default mission when no bank-specific mission is set
|
|
_DEFAULT_MISSION = "Track every detail: names, numbers, dates, places, and relationships. Prefer specifics over abstractions, never generalise."
|
|
|
|
# Processing rules — always present regardless of mission
|
|
_PROCESSING_RULES = """Processing rules (always apply):
|
|
- REDUNDANT: same info worded differently → UPDATE the existing observation.
|
|
- CONTRADICTION/UPDATE: capture both states with temporal markers ("used to X, now Y").
|
|
- RESOLVE REFERENCES: when a new fact provides a concrete value resolving a vague placeholder in an existing observation (e.g. "home country", "hometown", "birthplace", "native language", "her ex", "that city"), UPDATE the observation to embed the resolved value explicitly. Example: new fact says "grandma in Sweden" + existing observation says "moved from her home country" → update to "home country is Sweden".
|
|
- NEVER merge observations about different people or unrelated topics."""
|
|
|
|
# Data section — format placeholders {facts_text} and {observations_text} are substituted at call time
|
|
_BATCH_DATA_SECTION = """
|
|
NEW FACTS:
|
|
{facts_text}
|
|
|
|
EXISTING OBSERVATIONS (JSON array, pooled from recalls across all facts above):
|
|
{observations_text}
|
|
|
|
Each observation includes:
|
|
- id: unique identifier for updating
|
|
- text: the observation content
|
|
- proof_count: number of supporting memories
|
|
- occurred_start/occurred_end: temporal range of source facts
|
|
- source_memories: array of supporting facts with their text and dates
|
|
|
|
Compare the facts against existing observations:
|
|
- Same topic as an existing observation → UPDATE it (observation_id + source_fact_ids)
|
|
- New topic with durable knowledge → CREATE a new observation (source_fact_ids)
|
|
- Cross-reference facts within the batch: a later fact may resolve a vague reference in an earlier one
|
|
- Purely ephemeral facts → omit them unless the MISSION above explicitly targets such data (e.g. timestamped events, session state, screen content)"""
|
|
|
|
# Output format — JSON braces escaped as {{ }} so .format() leaves them literal
|
|
_BATCH_OUTPUT_FORMAT = """
|
|
Output a JSON object with three arrays.
|
|
|
|
## EXAMPLE
|
|
|
|
Input facts:
|
|
[a1b2c3d4-e5f6-7890-abcd-ef1234567890] Alice mentioned she works long hours, often past midnight | Involving: Alice (occurred_start=2024-01-15, mentioned_at=2024-01-15)
|
|
[b2c3d4e5-f6a7-8901-bcde-f12345678901] Alice said she's exhausted from the project deadlines | Involving: Alice (occurred_start=2024-01-20, mentioned_at=2024-01-20)
|
|
|
|
Good observation text — clean prose, no metadata, each fact tracked distinctly:
|
|
"Alice works long hours, often past midnight."
|
|
"Alice feels exhausted from project deadlines."
|
|
|
|
Bad observation text — NEVER do this (verbatim copy of fact text with metadata):
|
|
"Alice mentioned she works long hours, often past midnight | Involving: Alice (occurred_start=2024-01-15, mentioned_at=2024-01-15)"
|
|
|
|
Observation text rules:
|
|
- Write clean prose — NEVER copy raw fact lines or their metadata (temporal fields, "Involving:", "When:" labels, UUIDs).
|
|
- Parenthesized metadata like (occurred_start=...) and pipe-separated labels like "| Involving: ..." are fact formatting — strip them entirely from observation text.
|
|
- How many observations to create and how much to aggregate is driven by the MISSION above.
|
|
|
|
{{"creates": [{{"text": "Alice works long hours, often past midnight.", "source_fact_ids": ["a1b2c3d4-e5f6-7890-abcd-ef1234567890"]}}, {{"text": "Alice feels exhausted from project deadlines.", "source_fact_ids": ["b2c3d4e5-f6a7-8901-bcde-f12345678901"]}}],
|
|
"updates": [{{"text": "Alice works at Acme Corp as a senior engineer", "observation_id": "c3d4e5f6-a7b8-9012-cdef-123456789012", "source_fact_ids": ["d4e5f6a7-b8c9-0123-defa-234567890123"]}}],
|
|
"deletes": [{{"observation_id": "e5f6a7b8-c9d0-1234-efab-345678901234"}}]}}
|
|
|
|
Rules:
|
|
- "source_fact_ids": copy the EXACT UUID strings shown in brackets [uuid] from NEW FACTS — never use integers or positions.
|
|
- "observation_id": copy the EXACT "id" UUID string from EXISTING OBSERVATIONS.
|
|
- One create/update may reference multiple facts when they jointly support the observation.
|
|
- "deletes": only when an observation is directly superseded or contradicted by new facts.
|
|
- Do NOT include "tags" — handled automatically.
|
|
- Return {{"creates": [], "updates": [], "deletes": []}} if nothing durable is found."""
|
|
|
|
|
|
def build_batch_consolidation_prompt(observations_mission: str | None = None) -> str:
|
|
"""
|
|
Build the consolidation prompt for batch mode (multiple facts per LLM call).
|
|
|
|
The mission defines *what* to track (customisable per bank).
|
|
Processing rules and output format are always present regardless of mission.
|
|
"""
|
|
mission = observations_mission or _DEFAULT_MISSION
|
|
|
|
return (
|
|
"You are a memory consolidation system. Synthesize facts into observations "
|
|
"and merge with existing observations when appropriate.\n\n"
|
|
f"## MISSION\n{mission}\n\n"
|
|
f"{_PROCESSING_RULES}" + _BATCH_DATA_SECTION + _BATCH_OUTPUT_FORMAT
|
|
)
|