* fix: sanitize null bytes from text fields before PostgreSQL insertion Fixes 'invalid byte sequence for encoding UTF8: 0x00' error during batch retain Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: consolidate _sanitize_text into fact_extraction module Address review feedback: reuse existing _sanitize_text from fact_extraction instead of duplicating in fact_storage. The consolidated function now handles both: - Null bytes (\x00) for PostgreSQL compatibility - Unicode surrogates (U+D800-U+DFFF) for UTF-8/LLM API compatibility Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| bank_utils.py | ||
| chunk_storage.py | ||
| deduplication.py | ||
| embedding_processing.py | ||
| embedding_utils.py | ||
| entity_processing.py | ||
| fact_extraction.py | ||
| fact_storage.py | ||
| link_creation.py | ||
| link_utils.py | ||
| orchestrator.py | ||
| types.py | ||