* feat: add Pydantic AI integration to CI, release pipeline, and docs
- Add test-pydantic-ai-integration job to CI (test.yml)
- Add build, publish, and artifact steps to release workflow (release.yml)
- Add hindsight-integrations/pydantic-ai to release.sh version bumping
- Add Pydantic AI documentation page (sdks/integrations/pydantic-ai.md)
- Add Pydantic AI entry to sidebar with icon
* docs: remove Requirements section from pydantic-ai integration page
* feat: add tags filtering and fix offset pagination docs for list documents API
- Add `tags` and `tags_match` query params to GET /banks/{bank_id}/documents
- Supports any, all, any_strict, all_strict matching modes (default: any_strict)
- Fix `q` param description — it's a case-insensitive substring match on document ID only
- Add tests for offset pagination and all tags_match modes
- Regenerate OpenAPI spec and Python/TypeScript/Go clients
- Document the new filtering options in docs/developer/api/documents.mdx
* fix(cli): pass new tags/tags_match args to list_documents
196 lines
5.6 KiB
Text
196 lines
5.6 KiB
Text
---
|
|
sidebar_position: 8
|
|
---
|
|
|
|
# Documents
|
|
|
|
Track and manage document sources in your memory bank. Documents provide traceability — knowing where memories came from.
|
|
|
|
import Tabs from '@theme/Tabs';
|
|
import TabItem from '@theme/TabItem';
|
|
import CodeSnippet from '@site/src/components/CodeSnippet';
|
|
|
|
{/* Import raw source files */}
|
|
import documentsPy from '!!raw-loader!@site/examples/api/documents.py';
|
|
import documentsMjs from '!!raw-loader!@site/examples/api/documents.mjs';
|
|
|
|
:::tip Prerequisites
|
|
Make sure you've completed the [Quick Start](./quickstart) and understand [how retain works](./retain).
|
|
:::
|
|
|
|
## What Are Documents?
|
|
|
|
Documents are containers for retained content. They help you:
|
|
|
|
- **Track sources** — Know which PDF, conversation, or file a memory came from
|
|
- **Update content** — Re-retain a document to update its facts
|
|
- **Delete in bulk** — Remove all memories from a document at once
|
|
- **Organize memories** — Group related facts by source
|
|
|
|
## Chunks
|
|
|
|
When you retain content, Hindsight splits it into chunks before extracting facts. These chunks are stored alongside the extracted memories, preserving the original text segments.
|
|
|
|
**Why chunks matter:**
|
|
- **Context preservation** — Chunks contain the raw text that generated facts, useful when you need the exact wording
|
|
- **Richer recall** — Including chunks in recall provides surrounding context for matched facts
|
|
|
|
:::tip Include Chunks in Recall
|
|
Use `include_chunks=True` in your recall calls to get the original text chunks alongside fact results. See [Recall](./recall) for details.
|
|
:::
|
|
|
|
## Retain with Document ID
|
|
|
|
Associate retained content with a document:
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-retain" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-retain" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
# Retain content with document ID
|
|
hindsight memory retain my-bank "Meeting notes content..." --doc-id notes-2024-03-15
|
|
|
|
# Batch retain from files
|
|
hindsight memory retain-files my-bank docs/
|
|
```
|
|
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Update Documents
|
|
|
|
Re-retaining with the same document_id **replaces** the old content:
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-update" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-update" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
# Original
|
|
hindsight memory retain my-bank "Project deadline: March 31" --doc-id project-plan
|
|
|
|
# Update
|
|
hindsight memory retain my-bank "Project deadline: April 15 (extended)" --doc-id project-plan
|
|
```
|
|
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Get Document
|
|
|
|
Retrieve a document's original text and metadata. This is useful for expanding document context after a recall operation returns memories with document references.
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-get" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-get" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
hindsight document get my-bank meeting-2024-03-15
|
|
```
|
|
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Delete Document
|
|
|
|
Remove a document and all its associated memories:
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-delete" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-delete" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
hindsight document delete my-bank meeting-2024-03-15
|
|
```
|
|
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
:::warning
|
|
Deleting a document permanently removes all memories extracted from it. This action cannot be undone.
|
|
:::
|
|
|
|
## List Documents
|
|
|
|
List documents in a bank with optional filtering by ID and tags.
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={documentsPy} section="document-list" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={documentsMjs} section="document-list" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
```bash
|
|
# List all documents
|
|
hindsight document list my-bank
|
|
|
|
# Filter by ID substring
|
|
hindsight document list my-bank --q report
|
|
|
|
# Filter by tags
|
|
hindsight document list my-bank --tags team-a --tags team-b
|
|
```
|
|
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
### Filtering Options
|
|
|
|
| Parameter | Description |
|
|
|---|---|
|
|
| `q` | Case-insensitive substring match on document ID. `report` matches `report-2024`, `annual-report`, etc. |
|
|
| `tags` | Filter by document tags. Accepts multiple values. |
|
|
| `tags_match` | How to match tags (default: `any_strict`). See below. |
|
|
| `limit` / `offset` | Pagination. Default limit is 100. |
|
|
|
|
**`tags_match` modes:**
|
|
|
|
| Mode | Behaviour |
|
|
|---|---|
|
|
| `any_strict` *(default)* | Document must have **at least one** of the specified tags. Untagged docs excluded. |
|
|
| `any` | Same as `any_strict` but also includes untagged documents. |
|
|
| `all_strict` | Document must have **all** specified tags. Untagged docs excluded. |
|
|
| `all` | Same as `all_strict` but also includes untagged documents. |
|
|
|
|
## Document Response Format
|
|
|
|
```json
|
|
{
|
|
"id": "meeting-2024-03-15",
|
|
"bank_id": "my-bank",
|
|
"original_text": "Alice presented the Q4 roadmap...",
|
|
"content_hash": "abc123def456",
|
|
"memory_unit_count": 12,
|
|
"created_at": "2024-03-15T14:00:00Z",
|
|
"updated_at": "2024-03-15T14:00:00Z"
|
|
}
|
|
```
|
|
|
|
## Next Steps
|
|
|
|
- [**Operations**](./operations) — Monitor background tasks
|
|
- [**Memory Banks**](./memory-banks) — Configure bank settings
|