* feat: support timestamp="unset" to retain content without a date When callers retain timeless content (e.g. fictional documents, static reference material), passing timestamp="unset" now skips the utcnow() default so mentioned_at is stored as NULL instead of an artificial date. - HTTP: validate_timestamp recognises "unset" sentinel and threads it through api_retain as event_date=None (key present, value None), which the orchestrator distinguishes from key-absent (still defaults to now) - Orchestrator: new branching logic separates "key absent" → utcnow() from "key present but None" → no date - types.py: RetainContent.event_date and ProcessedFact.mentioned_at are now datetime | None; removed the unused _now_utc factory - fact_extraction.py: all event_date params accept datetime | None; _build_user_message emits "Event Date: Unknown" when None; removed mentioned_at from the Fact LLM response model (LLM never sets it) - embedding_processing: skip date suffix when fact_date is None - entity_resolver: COALESCE(event_date, now()) for first_seen/last_seen so entities table NOT NULL constraint is preserved - link_utils: skip temporal linking for units without event_date - Migration aa2b3c4d5e6f: DROP NOT NULL on memory_units.event_date - Tests: test_retain_no_timestamp and test_retain_omit_timestamp_defaults_to_now - Docs + OpenAPI + TypeScript client updated * refactor: replace _TIMESTAMP_UNKNOWN sentinel with plain string comparison The sentinel object() was only needed to distinguish "unset" from None at the boundary — but since the field type is datetime | str | None, "unset" can pass through the validator unchanged and be compared directly. * chore: regenerate OpenAPI spec and clients after timestamp type change timestamp field is now datetime | str | None to accept the "unset" sentinel value.
58 lines
1.7 KiB
Python
58 lines
1.7 KiB
Python
"""
|
|
Embedding processing for retain pipeline.
|
|
|
|
Handles augmenting fact texts with temporal information and generating embeddings.
|
|
"""
|
|
|
|
import logging
|
|
|
|
from . import embedding_utils
|
|
from .types import ExtractedFact
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
def augment_texts_with_dates(facts: list[ExtractedFact], format_date_fn) -> list[str]:
|
|
"""
|
|
Augment fact texts with readable dates for better temporal matching.
|
|
|
|
This allows queries like "camping in June" to match facts that happened in June.
|
|
|
|
Args:
|
|
facts: List of ExtractedFact objects
|
|
format_date_fn: Function to format datetime to readable string
|
|
|
|
Returns:
|
|
List of augmented text strings (same length as facts)
|
|
"""
|
|
augmented_texts = []
|
|
for fact in facts:
|
|
# Use occurred_start as the representative date
|
|
fact_date = fact.occurred_start or fact.mentioned_at
|
|
if fact_date is not None:
|
|
readable_date = format_date_fn(fact_date)
|
|
# Augment text with date for embedding (but store original text in DB)
|
|
augmented_text = f"{fact.fact_text} (happened in {readable_date})"
|
|
else:
|
|
augmented_text = fact.fact_text
|
|
augmented_texts.append(augmented_text)
|
|
return augmented_texts
|
|
|
|
|
|
async def generate_embeddings_batch(embeddings_model, texts: list[str]) -> list[list[float]]:
|
|
"""
|
|
Generate embeddings for a batch of texts.
|
|
|
|
Args:
|
|
embeddings_model: Embeddings model instance
|
|
texts: List of text strings to embed
|
|
|
|
Returns:
|
|
List of embedding vectors (same length as texts)
|
|
"""
|
|
if not texts:
|
|
return []
|
|
|
|
embeddings = await embedding_utils.generate_embeddings_batch(embeddings_model, texts)
|
|
|
|
return embeddings
|