* feat: add @vectorize-io/hindsight-embed daemon lifecycle package
Create a new top-level `hindsight-embed-npm/` package that owns the daemon
lifecycle for the Python `hindsight-embed` CLI: spawning via `uvx`, writing
the profile, waiting for `/health`, and shutting down. Nothing more.
Deliberately does not ship an HTTP client — `@vectorize-io/hindsight-client`
already covers retain / recall / reflect / createBank against the Hindsight
API, and the two packages compose: once `manager.start()` returns, consumers
talk to the daemon via `new HindsightClient({ baseUrl: manager.getBaseUrl() })`.
`HindsightEmbedManagerOptions.env` forwards an arbitrary `Record<string,
string>` to both the daemon process and the profile config via `--env K=V`,
and `extraProfileCreateArgs` / `extraDaemonStartArgs` escape hatches cover
any new CLI flag without waiting for a wrapper release.
Refactor `hindsight-integrations/openclaw` to consume both packages:
`HindsightEmbedManager` for daemon lifecycle in local mode, `HindsightClient`
for all HTTP memory operations. Drop the bespoke subprocess/HTTP client that
used to live in openclaw. The retain queue stays local to openclaw (it's a
client-side reliability workaround with a single consumer today — will move
to the client package or server-side when a second consumer needs it).
Wire the new package into the main release pipeline (versioned alongside
the other core packages, published from `v*` tags) and add a CI build job.
* docs: add Embedded Node.js SDK page for @vectorize-io/hindsight-embed
* refactor: rename hindsight-embed-npm to hindsight-all, restructure docs sidebar
The Node package previously named @vectorize-io/hindsight-embed was
semantically misnamed: hindsight-embed (Python) is a CLI tool, while what
this Node package actually provides is the Node equivalent of hindsight-all
— a programmatic lifecycle manager for a local Hindsight daemon. Rename to
match.
Package rename
- hindsight-embed-npm/ → hindsight-all-npm/ (git mv, history preserved)
- @vectorize-io/hindsight-embed → @vectorize-io/hindsight-all
- class HindsightEmbedManager → HindsightServer (matches Python hindsight-all)
- HindsightEmbedManagerOptions → HindsightServerOptions
- src/manager.ts → src/server.ts, src/manager.test.ts → src/server.test.ts
- openclaw (index.ts, backfill.ts, tests) and the claude-code Python port
updated to reference the new names
Docs restructure
- Split sdks/python.md: now client-only content. New sdks/hindsight-all.md
covers the programmatic hindsight-all Python package (HindsightServer and
HindsightEmbedded).
- Rename sdks/embed-npm.md → sdks/hindsight-all-npm.md with HindsightServer
examples.
- New "Installation" sidebar section, placed after Hosting, containing
Docker / Kubernetes / Bare Metal (anchor links into developer/installation)
plus Programmatic API (Python), Programmatic API (Node.js), and Daemon CLI.
- Add si-docker, si-kubernetes, si-nodedotjs, lu-hard-drive to the sidebar
ICON_MAP.
Docs dev-server fix
- docusaurus.config.ts: drop the flaky NODE_ENV sniff for including the
"Next" version. Use INCLUDE_CURRENT_VERSION exclusively. NODE_ENV was
unreliable across hot-reload paths and caused the Next version to
disappear intermittently when editing files.
- scripts/dev/start-docs.sh: export INCLUDE_CURRENT_VERSION=true so local
dev always shows Next; production builds leave it unset.
Lockfile cleanup
- package-lock.json and hindsight-integrations/openclaw/package-lock.json
had extraneous hindsight-embed-npm blocks left over from the rename.
Removed manually and verified with npm install.
* ci: fix openclaw jobs by pre-building workspace deps; regenerate docs-skill
The build-openclaw-integration and test-openclaw-integration jobs failed
with "Failed to resolve entry for package @vectorize-io/hindsight-all"
because openclaw depends on two monorepo workspaces via `file:` deps
(@vectorize-io/hindsight-client and @vectorize-io/hindsight-all) whose
`dist/` directories are gitignored and never built before openclaw's npm ci.
Both jobs now install the root workspace and build the two deps first,
mirroring the release-control-plane pattern.
Also regenerate skills/hindsight-docs/references/* via
./scripts/generate-docs-skill.sh:
- new skill pages for sdks/hindsight-all{.md,-npm.md}
- updated skill pages for sdks/embed.md and sdks/python.md to match
the new H1s and split content
- incidental refreshes to changelog/index.md, developer/models.md,
openapi.json, and uv.lock that verify-generated-files picked up
* ci: build openclaw before running tests so symlink test can realpath dist
PR #932 added update_mode (replace/append) to retain items but
did not update the docs. Add a section explaining the parameter,
when to use append mode, and a JSON example.
Closes#957
* feat(openclaw): add session pattern filtering for ignore and stateless sessions
Adds three new config options to the OpenClaw plugin that allow filtering
sessions by key pattern before recall and retain operations fire:
- `ignoreSessionPatterns`: glob patterns for sessions to skip entirely
(no recall, no retain). Useful for cron/scheduled agent sessions.
- `statelessSessionPatterns`: glob patterns for read-only sessions —
retain is always skipped; recall is also skipped when
`skipStatelessSessions` is true (default).
- `skipStatelessSessions`: boolean (default: true). When false, sessions
matching statelessSessionPatterns can still recall but never retain.
Pattern syntax mirrors lossless-claw: `*` matches non-colon characters,
`**` matches anything including colons. Session keys follow the OpenClaw
format `agent:<agentId>:<type>:<uuid>`.
Example config:
ignoreSessionPatterns: ["agent:*:cron:**"]
statelessSessionPatterns: ["agent:*:subagent:**", "agent:*💓**"]
skipStatelessSessions: true
Implementation:
- New `session-patterns.ts` module with compile/match utilities
- Session filter applied in `before_prompt_build` and `agent_end` hooks
immediately after the existing `excludeProviders` check
- New fields wired through `getPluginConfig`
- Schema added to `openclaw.plugin.json` (additionalProperties: false
was already set, causing config validation errors without this)
- 11 unit tests in `session-patterns.test.ts`
- 5 integration tests added to `hooks.integration.test.ts`
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(openclaw): support HINDSIGHT_API_TOKEN in integration tests
Pass HINDSIGHT_API_TOKEN env var through to HindsightClient and plugin
config in integration tests so tests work against authenticated APIs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(openclaw): document session pattern filtering options
Add ignoreSessionPatterns, statelessSessionPatterns, and skipStatelessSessions
to the README config table with glob syntax reference and usage examples.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Marco Rutsch <marco@rutimka.de>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: add 0.5.0 release notes and changelog
* docs: include all commits since v0.4.22 and add recall perf to blog
* docs: include all commits since v0.4.22 and add recall perf to blog
* docs: add openrouter default model to provider table
* docs: reorder blog sections, fix code snippets, remove paperclip
* docs: add hermes integration docs link
* docs: fix broken anchor in blog post TOC
greenlet 3.4.0 lacks manylinux_2_41_aarch64 wheels. Use a UV_CONSTRAINT
file instead of the workspace lock file (which doesn't work in the
single-package Docker context).
Without the lock file, uv sync resolves fresh and picks up greenlet
3.4.0 which lacks arm64 wheels for manylinux_2_41, breaking the
multi-arch Docker build.
* fix: exclude local-llm from [all] extra to avoid heavy llama-cpp-python dep
local-llm (llama-cpp-python) requires C++ compilation and is only needed
for the built-in llamacpp provider. Keep it as a separate opt-in:
pip install 'hindsight-api-slim[local-llm]'
* feat: add local-llm optional extra to hindsight-all
Allows: pip install 'hindsight-all[local-llm]' to get built-in llamacpp support.
* chore: regenerate uv.lock from workspace root
* feat: add built-in llama.cpp LLM provider for fully local inference
Add `llamacpp` as a new LLM provider that manages a llama-cpp-python server
subprocess. Auto-downloads Gemma 4 E2B Q4_K_M (~3.5 GB) on first use and
runs inference locally via Metal/CUDA with no external services needed.
- New provider: `HINDSIGHT_API_LLM_PROVIDER=llamacpp`
- Singleton server shared across retain/reflect/consolidation
- Configurable: model path, GPU layers, context size, grammar enforcement
- User-extensible via `HINDSIGHT_API_LLAMACPP_EXTRA_ARGS`
- Flash attention + prompt caching enabled by default
- LLM provider cleanup on shutdown (stops subprocess)
- hindsight-embed: `--ui` flag on `daemon start`, removed FORCE_CPU on macOS
- Docs: configuration.md, models.mdx, providers grid updated
* chore: regenerate docs skill and update lockfile for local-llm dep
* feat: add update_mode='append' for retain to concatenate content to existing documents
When retaining with update_mode='append' and a document_id that already exists,
the new content is appended to the existing document text and the full document
is reprocessed. Delta retain automatically skips unchanged chunks, so only the
new content triggers LLM extraction.
- Add update_mode field to MemoryItem (API), RetainContentDict (internal), MCP tools
- Validate that update_mode='append' requires a document_id
- Fetch existing document content and prepend before processing in orchestrator
- Update Python, TypeScript, Go generated clients and top-level client wrappers
- Add tests for append, multiple appends, no-existing-doc, validation, and default replace
* fix: add update_mode field to Rust CLI and client MemoryItem initializers
* chore: regenerate docs skill references for update_mode
* docs: add best practice for filtering recall by memory shape (#856)
Add guidance on using entity labels with `tag: true` to deterministically
filter recall results when a bank contains different memory shapes
(e.g., concise rules vs. detailed procedures).
* feat: add OpenRouter support for LLM, embeddings, and reranking
OpenRouter is OpenAI-compatible for chat/embeddings and Cohere-compatible
for reranking, so no new provider classes are needed.
- LLM: added as OpenAICompatibleLLM provider (default model: qwen/qwen3.5-9b)
- Embeddings: reuses OpenAIEmbeddings with OpenRouter base URL (default: perplexity/pplx-embed-v1-0.6b)
- Reranker: reuses CohereCrossEncoder with OpenRouter rerank endpoint (default: cohere/rerank-v3.5)
- API key fallback chain: dedicated key → shared OPENROUTER_API_KEY → LLM_API_KEY
* chore: regenerate docs skill references and fix formatting
Extend format_facts_for_prompt() to include occurred_end and mentioned_at
temporal fields (when non-null), matching the MemoryFact model. Also add
RecallResponse.to_prompt_string() to Python and TypeScript client SDKs so
users can serialize recall results (with chunks and entity summaries) into
LLM-ready prompt strings.
Closes#924
* fix: make LiteLLM SDK embeddings encoding_format configurable (#925)
The hardcoded encoding_format='float' breaks providers like Voyage AI
(only accepts 'base64') and Gemini (doesn't support the parameter at all).
Add HINDSIGHT_API_EMBEDDINGS_LITELLM_SDK_ENCODING_FORMAT config option
that defaults to 'float' for backwards compatibility. Set to empty string
to omit the parameter for incompatible providers.
* chore: regenerate docs skill after configuration change
* security: bump lodash, lodash-es, and defu in root lockfile
Fixes Dependabot alerts in the root npm workspace lockfile:
- GHSA-r5fr-rjxr-66jc (high) lodash <4.18.1 (alert #338)
- GHSA-r5fr-rjxr-66jc (high) lodash-es <4.18.1 (alert #335)
- GHSA-737v-mqg7-c878 (high) defu <6.1.7 (alert #343)
defu (6.1.4 -> 6.1.7) and lodash (4.17.23 -> 4.18.1) were bumped via
targeted `npm update`. lodash-es was pinned exactly to 4.17.23 by
@chevrotain packages (transitive dep of mermaid in hindsight-docs),
so a `lodash-es` override (>=4.18.1) is added to the root package.json
to force resolution to the patched 4.18.1.
Verified: `npm ci` succeeds with 0 vulnerabilities. Mermaid/chevrotain
consumers all dedupe to lodash-es 4.18.1. lodash-es 4.x is semver-
compatible.
* chore: regenerate hindsight-docs skill
Picks up FAQ and best-practice sections added in #905 that were not
regenerated at merge time, so that `verify-generated-files` passes
for this branch.
* security: bump vite across integrations to patched versions
Fixes Dependabot alerts for vite transitive dev dependency:
- GHSA-v2wj-q39q-566r (high): server.fs.deny bypass with queries
- GHSA-p9ff-h696-f583 (high): related vite server vulnerability
Adds a `vite` entry to the npm `overrides` in each integration's
package.json to force the patched version (>=8.0.5). To make this
possible in ai-sdk, chat, and openclaw — which pinned vitest ^4.0.18
whose vite peer is `^6.0.0 || ^7.0.0` — the minor-compatible bump
vitest ^4.0.18 -> ^4.1.2 is also included. vitest 4.1.x supports
vite 8.x (peer: ^6 || ^7 || ^8), so all six integrations converge on
vite 8.x consistently.
paperclip had no overrides block; one was added.
Verified locally: `npm ci && npx vitest run` passes in all six
integrations (ai-sdk 23, chat 28, openclaw 66, opencode 89, paperclip 27,
nemoclaw 36 tests).
* chore: regenerate hindsight-docs skill
Picks up FAQ and best-practice sections added in #905 that were not
regenerated at merge time, so that `verify-generated-files` passes
for this branch.
Some LLM providers (e.g. Anthropic Haiku) return 1-indexed
content_index values. When only one content item is provided,
this causes KeyError: 1 since the dict only has key 0.
Clamp content_index to the valid range instead of crashing.
Fixes#873
Co-authored-by: easonysliu <easonysliu@tencent.com>
* fix(recall): cap entity fanout in graph expansion to prevent slow queries
On large banks, the entity co-occurrence self-join in _expand_combined()
produces massive intermediate row counts when seeds reference high-fanout
entities (e.g. an entity with 25K+ mentions). This causes recall latency
to degrade significantly.
Changes:
- Replace unbounded entity self-join with LATERAL per-entity cap
(graph_per_entity_limit, default 200), reducing intermediate rows
from potentially millions to at most num_entities * 200
- Add ORDER BY unit_id DESC in LATERAL subquery for deterministic
recency-biased sampling (rides the PK index, no extra sort)
- Add timeout fallback (graph_expansion_timeout, default 10s) that
drops entity expansion and falls back to semantic+causal only
- Add composite index (entity_id, unit_id) on unit_entities for
index-only scans in the LATERAL subquery
- Merge 3 unmerged migration heads into one
- Fix recall_perf.py dotenv override issue
Unlike the approach in #895, this does NOT filter out hub entities
entirely — all entities are kept but capped equally, preserving
retrieval quality for queries about frequently-mentioned entities.
Benchmarked on a 67K-unit bank (top entity = 25K mentions):
- retrieval_graph: 0.337s → 0.055s (84% faster)
- end-to-end recall: 0.912s → 0.519s (43% faster)
* fix(tests): fix broken test_combined_scoring and test_reranking_proof_count
- test_combined_scoring: replace MagicMock(spec=RetrievalResult) with real
dataclass instances — MagicMock attributes returned nested mocks that
failed on >= comparisons with int
- test_reranking_proof_count: remove deleted `embedding` param from
RetrievalResult constructor, use None for occurred_start/end to get
neutral recency (datetime.now gave recency=1.0 which boosted scores)
* refactor: rename config to link_expansion_ prefix, fix observation fanout
- Rename GRAPH_PER_ENTITY_LIMIT → LINK_EXPANSION_PER_ENTITY_LIMIT and
GRAPH_EXPANSION_TIMEOUT → LINK_EXPANSION_TIMEOUT to follow the
convention that these are specific to the link_expansion graph retriever
- Apply the same LATERAL per-entity cap to _expand_observations(), which
had the same unbounded self-join through unit_entities
* style: fix formatting in config.py
* feat(openclaw): support exact static bank ids
* test(openclaw): use generic static bank id example
* feat(openclaw): support bankId static bank configuration
---------
Co-authored-by: Aldous the Orchestrator <Aldoustheorchestrator@users.noreply.github.com>
Fixes Dependabot alerts:
- GHSA-jjhc-v7c2-5hh6 (critical): Authentication bypass via OIDC userinfo
cache key collision (CVE-2026-35030)
- GHSA-53mr-6c8q-9789 (high): related litellm vulnerability
Updates both hindsight-api-slim and hindsight-integrations/litellm to
require litellm >=1.83.0. The previous upper cap (<=1.82.6) was set due
to the 1.82.7/1.82.8 supply chain compromise, which has since been yanked
from PyPI; 1.83.0 was published from the new secure CI/CD v2 pipeline
and is safe.
The uv.lock diffs are large because the current uv version (0.9.11)
upgrades the lockfile format (adds revision=3 and upload-time fields);
only litellm itself changes version (1.81.10/1.80.10 -> 1.83.0).
All 68 tests in hindsight-integrations/litellm pass against 1.83.0.
* fix(mcp): validate UUID inputs at engine level and add sync_retain tool (#888)
- Add UUID validation in memory_engine for get_memory_unit, delete_memory_unit,
get_mental_model, delete_mental_model, get_mental_model_history (raises ValueError)
- Catch ValueError → 400 in HTTP route handlers
- Add sync_retain MCP tool that calls retain_batch_async directly for immediate
availability (no polling needed)
- Register sync_retain in _ALL_TOOLS, _SINGLE_BANK_TOOLS, UI MCP_TOOL_GROUPS
- Add code-review check for MCP tool registration completeness
* fix: remove UUID validation for mental model IDs (column is TEXT, not UUID)
Mental model IDs are TEXT columns that accept arbitrary string IDs
(e.g., 'team-communication-preferences'). UUID validation was incorrectly
added to get_mental_model, delete_mental_model, and get_mental_model_history.
* test: add regression tests for #874 and #894
Add tests for None event_date in fact extraction (AttributeError fix)
and for _register_profile skipping .env overwrite with short config keys.
* fix(config): validate entity_labels structure on PATCH (#891)
Config PATCH accepted bare strings in entity_labels values without
validation, causing silent failures at retain time. Now validates
via parse_entity_labels() before writing to DB, and fixes the
BankTemplateConfig type from list[str] to list[dict[str, Any]].
* fix(scripts): handle Python client generator README crash gracefully
The openapi-generator sometimes crashes writing README_onlypackage.mustache.
Allow the failure with || true since all API/model files are generated
before that step, and add a verification check for api_client.py.
* chore: regenerate docs skill openapi.json
Add guidance on using entity labels with `tag: true` to deterministically
filter recall results when a bank contains different memory shapes
(e.g., concise rules vs. detailed procedures).
* feat: add OpenCode persistent memory plugin
Add hindsight-opencode integration with:
- Three custom tools: hindsight_retain, hindsight_recall, hindsight_reflect
- Auto-retain on session.idle with document_id deduplication
- Memory injection on session start via system transform hook
- Memory preservation during context window compaction
- Sliding window retain with retainOverlapTurns support
- 4-level config hierarchy (defaults, user file, plugin options, env vars)
- Dynamic bank ID derivation (agent, project, channel, user dimensions)
- CI job, release script entry, docs page
79 tests across 6 test files.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: address review findings for opencode integration
1. Pre-compaction retain now uses shared retainSession() helper,
respecting retainMode, documentId, and session_id metadata
consistently with idle-retain (was bypassing retention policy).
2. System transform recall is only consumed after successful injection.
If Hindsight is briefly unavailable, the plugin retries on the next
LLM call instead of permanently skipping recall for the session.
3. Config validation for retainMode and recallBudget — typos like
"full_session" or "maximum" now log a warning and fall back to
the default instead of silently changing retention semantics.
85 tests (6 new covering compaction documentId, recall retry, and
config validation).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: docs/tools findings from second review round
1. Remove "session" from supported dynamic bank fields in docs —
the implementation can't vary bank ID per session since it's
derived once at plugin startup.
2. Explicit tools (retain, reflect) now call ensureBankMission()
before API calls, so bankMission/retainMission are applied even
when the agent uses tools exclusively without triggering hooks.
3. Added tests for mission setup via tools path.
88 tests pass.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: recall retry semantics and README bank scoping clarity
1. recallForContext now returns { context, ok } to distinguish
"no results" (ok=true) from "API error" (ok=false). System
transform consumes the session on ok=true even with 0 results,
so empty banks don't cause repeated queries. Only transient API
failures preserve retry.
2. README clarifies that channel/user bank dimensions are process-
scoped (set via env vars before launch), not per-session dynamic
within a running OpenCode process.
89 tests pass.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: review fixes for opencode integration
- Rename CI job from build-opencode-integration to test-opencode-integration
to match naming convention for integrations that run tests
- Fix tsconfig module resolution to Node16 (consistent with other integrations)
- Extract shared makeConfig test helper to avoid duplication across 3 test files
* fix: remove unused PluginState import from tools.ts
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nicolò Boschi <boschi1997@gmail.com>
* fix: make bank_id metric label opt-in to prevent OTel memory leak
bank_id as an OTel metric attribute creates unbounded histogram growth
since each unique bank_id produces never-evicted time series. Default
to excluding it; opt in with HINDSIGHT_API_METRICS_INCLUDE_BANK_ID=true
for deployments with few banks.
Closes#850
* refactor: use config.py for metrics_include_bank_id setting
Move HINDSIGHT_API_METRICS_INCLUDE_BANK_ID from direct os.getenv in
metrics.py to the standard HindsightConfig path. Add configuration
documentation.
LLM agents frequently serialize list/dict tool arguments as JSON strings
instead of native types (e.g., tags='["a","b"]' instead of tags=["a","b"]),
causing Pydantic validation failures. This extends _make_tools_tolerant to
detect array/object parameters from the JSON Schema and auto-coerce string
values via json.loads before validation.
Also fixes _make_tools_tolerant compatibility with FastMCP 3.x by adding
a _get_mcp_tools helper that supports both 2.x and 3.x internal APIs.
* feat(recall): add proof_count boost to combined scoring
Observations with more supporting evidence now rank slightly higher
in recall results. proof_count is threaded through the retrieval
pipeline and applied as a multiplicative boost in reranking:
- types.py: add proof_count field to RetrievalResult
- retrieval.py: include proof_count in SELECT columns
- reranking.py: add log1p-normalized proof_count boost (alpha=0.1)
The boost uses the same multiplicative pattern as recency and temporal
signals. proof_count=1 is neutral, proof_count=50 gives ~+5% boost.
Non-observation fact types are unaffected (neutral 0.5).
* fix(retrieval): Apply proof_count boost to graph and temporal retrieval, normalize scaling
* fix(retrieval): correct proof_norm math to zero-center at count 1
* fix(retrieval): Apply proof_count boost to link_expansion retrieval
* fix: remove BFS zombie, clamp proof_norm to [0,1], fix test comment (log1p->math.log)
- Add CI job for paperclip integration tests with change detection
- Add paperclip to valid release integrations
- Validate hindsightApiUrl is set in loadConfig()
- Log warnings on recall/retain failures instead of silently swallowing
- Remove hardcoded timeout from reflect call
- Fix tsconfig module resolution to Node16
- Update tests to pass required hindsightApiUrl
When the daemon is already running, ensure_running() calls _register_profile()
with a config dict using short keys (llm_api_key, llm_provider, etc.) that do
not match the HINDSIGHT_API_* prefix filter. This caused api_config to always
be empty, and create_profile() would overwrite the existing .env with an empty
file on every CLI command.
Add an early return guard so _register_profile() skips the create_profile()
call when api_config is empty, preserving any existing profile configuration.
Fixes#894
DateparserQueryAnalyzer.analyze() called dateparser.search.search_dates()
without any error handling, so internal bugs in the third-party library
propagated all the way up the search/consolidation pipeline and failed
the calling task.
Observed traceback:
File ".../engine/query_analyzer.py", line 140, in analyze
results = self._search_dates(query, settings=settings)
File ".../dateparser/search/search.py", line 294, in search_dates
"Dates": self.search.search_parse(...)
File ".../dateparser/search/search.py", line 168, in search_parse
translated, original = self.search(shortname, text, settings)
File ".../dateparser/languages/locale.py", line 224, in translate_search
[original_tokens[i], original_tokens[i + 1]],
IndexError: list index out of range
Wrap the call in a try/except so any parser failure is treated as
"no temporal constraint found" — the caller can then fall back to
non-temporal retrieval instead of erroring out the whole task. The
failure is logged at WARNING level so we still notice it.
Add a regression test that monkey-patches _search_dates to raise an
IndexError and asserts the analyzer returns an empty constraint and
emits a warning log.
* Fix AttributeError when event_date is None in fact_extraction
`_extract_facts_from_chunk` crashes with `'NoneType' object has no
attribute 'isoformat'` when retaining documents without a timestamp.
Two locations fixed:
- Line 1058: debug log called `event_date.isoformat()` without a None
check
- Line 921: `parse_datetime_flexible()` can return None, so re-check
before calling `.strftime()` / `.isoformat()`
Fixes#874
* Revert unnecessary None guard on line 921
The original `if event_date is not None:` already guards that block.
Only line 1058 needed the fix.
- Add cross-platform file locking support
- Use fcntl on Unix-like systems, msvcrt on Windows
- Add detailed documentation explaining why we don't use external libraries
- Fixes issue where module couldn't be imported on Windows due to missing fcntl
Co-authored-by: yishun.eason <yishun.eason@bytedance.com>
* feat(helm): add persistent volume for local model cache
When using local reranker (e.g., BAAI/bge-reranker-v2-m3) or local
embedding models, the models are downloaded to /home/hindsight/.cache
on every pod restart, causing slow startup and unnecessary bandwidth.
Add optional persistent volume support:
- api: PVC mounted at /home/hindsight/.cache
- worker: volumeClaimTemplate (StatefulSet) at same path
Disabled by default. Enable via:
api.persistence.modelCache.enabled: true
worker.persistence.modelCache.enabled: true
Closes#860
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(helm): add extraVolumes and extraVolumeMounts for api and worker
Allow users to mount arbitrary volumes (configMaps, secrets, emptyDir,
etc.) into api and worker pods via values, following common helm chart
library conventions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Mistral (and several other providers) reject 'max_completion_tokens' with a 422
because they haven't adopted the newer OpenAI parameter name. When the openai
provider is configured with a custom base_url (e.g. Mistral, Together AI),
fall back to the widely-supported 'max_tokens' parameter.
Native OpenAI (no custom base_url) and Groq still use 'max_completion_tokens'.
Fixes#852