fleet-memory/hindsight-api/hindsight_api
Nicolò Boschi 1caf5ec9ee
feat: add jina-mlx reranker provider for Apple Silicon (#542)
* feat: add JinaMLXCrossEncoder for native Apple Silicon reranking

Adds a new `jina-mlx` reranker provider backed by jinaai/jina-reranker-v3-mlx,
a 0.6B multilingual listwise reranker running via the MLX framework on Apple Silicon.
The model is downloaded automatically from HuggingFace Hub on first use.

Benchmarked latencies (Apple Silicon): 1 doc→32ms, 5→45ms, 10→60ms, 20→94ms.
Sub-linear scaling because all docs are ranked in a single forward pass.

- Embeds the MLX reranker implementation (_MLXReranker / _MLPProjector) directly
  in cross_encoder.py with no transformers/PyTorch dependency
- Adds `mlx`, `mlx-lm`, `safetensors` to pyproject.toml optional deps (uv add)
- Updates configuration.md with provider docs and benchmark table

* refactor: import MLXReranker from repo rerank.py instead of duplicating code

Use importlib to load MLXReranker directly from the model repo's own rerank.py
(downloaded via snapshot_download). Also pin exact minimum versions for
mlx>=0.31.0, mlx-lm>=0.31.1, safetensors>=0.6.2 (verified against installed versions).

* refactor: move MLX reranker impl to dedicated jina_mlx_reranker.py

Replaces the importlib hack with a proper module. jina_mlx_reranker.py is
adapted from jinaai/jina-reranker-v3-mlx/rerank.py (CC BY-NC 4.0) with the
source clearly documented at the top of the file.

* docs: simplify jina-mlx reranker docs

* fix: disable GIN fastupdate on source_memory_ids index to prevent deadlocks

GIN fastupdate buffers inserts in a pending list and flushes it with
AccessExclusiveLock when full. Under concurrent test load (8 xdist workers
all running retain_async), two workers can trigger a flush simultaneously
and deadlock. Recreating the index with fastupdate=off eliminates the
flush/lock cycle at the cost of slightly slower individual inserts.

* fix: drop per-bank HNSW indexes after transaction to avoid AccessExclusiveLock deadlock

When deleting a bank, the previous code dropped HNSW indexes inside the
same transaction as the DELETE FROM memory_units. Since DROP INDEX needs
AccessExclusiveLock on the parent table and DELETE holds RowExclusiveLock,
two concurrent bank deletions deadlocked on the same table lock.

Fix: capture internal_id inside the transaction, commit, then drop the
indexes outside the transaction so no row-level locks are held.
2026-03-11 15:15:58 +01:00
..
admin Fix run-db-migration for all-tenant upgrades (#530) 2026-03-10 10:10:03 +01:00
alembic feat: add jina-mlx reranker provider for Apple Silicon (#542) 2026-03-11 15:15:58 +01:00
api feat: make recall max query tokens configurable via env var (#544) 2026-03-11 14:56:59 +01:00
engine feat: add jina-mlx reranker provider for Apple Silicon (#542) 2026-03-11 15:15:58 +01:00
extensions Add on_file_convert_complete extension hook after file-to-markdown conversion (#507) 2026-03-06 09:56:53 +01:00
webhooks fix: use correct schema name in webhook outbox callback to prevent silent transaction rollback (#499) 2026-03-05 17:07:48 +01:00
worker feat: webhook system with retain.completed event, UI, and docs (#487) 2026-03-04 14:17:01 +01:00
__init__.py Release v0.4.17 2026-03-10 17:18:35 +01:00
banner.py feat: support for pgvectorscale (DiskANN) (#378) 2026-02-16 14:19:56 +01:00
config.py feat: make recall max query tokens configurable via env var (#544) 2026-03-11 14:56:59 +01:00
config_resolver.py Fix bank config API for multi-tenant schema isolation (#417) 2026-02-20 23:43:52 +01:00
daemon.py feat: improve openclaw and hindisght-embed params (#279) 2026-02-03 09:39:04 +01:00
main.py feat: make recall max query tokens configurable via env var (#544) 2026-03-11 14:56:59 +01:00
mcp_local.py fix(mcp): unify hindsight-mcp-local and server mcp (#407) 2026-02-19 17:57:17 +01:00
mcp_tools.py Fix bank-level MCP tool filtering for FastMCP 3.x (#491) 2026-03-04 10:29:29 -05:00
metrics.py chore: remove dead code (#245) 2026-01-30 09:16:32 +01:00
migrations.py perf: replace window-function retrieval with UNION ALL + per-bank HNSW indexes (#541) 2026-03-11 12:09:50 +01:00
models.py Add bank-scoped validation to engine and HTTP handlers (#454) 2026-03-02 09:55:21 +01:00
pg0.py feat: support vertex as llm provider (#233) 2026-01-29 16:13:57 -05:00
server.py Fix: Load extensions in server.py for multi-worker deployments (#155) 2026-01-13 17:55:33 +01:00
tracing.py feat: add otel traceability (#330) 2026-02-10 12:20:48 +01:00
utils.py fix(helm): improve appVersion usage (#326) 2026-02-09 11:35:08 +01:00