* feat: Record LLM token metrics via Prometheus Wire up the existing token metrics infrastructure to actually record token usage from LLM calls. The MetricsCollector already had record_tokens() method and Prometheus counters (hindsight.tokens.input, hindsight.tokens.output), but they were never being populated. Changes: - Import get_metrics_collector in llm_wrapper.py - Call record_tokens() after successful LLM calls for: - OpenAI/Groq (using response.usage.prompt_tokens, completion_tokens) - Anthropic (using response.usage.input_tokens, output_tokens) - Gemini (using response.usage_metadata.prompt_token_count, candidates_token_count) - Add test file to verify token metrics are recorded Note: Ollama's native API doesn't return token usage, so metrics are not recorded for that provider. The token metrics will now be available via /metrics endpoint: - hindsight_tokens_input_total - hindsight_tokens_output_total * feat: add per-request token usage tracking to retain and reflect endpoints - Add TokenUsage model with input_tokens, output_tokens, total_tokens - Return usage metrics in retain response (sync operations only) - Return usage metrics in reflect response - Update Python, TypeScript, and Rust clients - Add API documentation for usage fields - Add changelog entry
249 lines
8.9 KiB
Text
249 lines
8.9 KiB
Text
---
|
|
sidebar_position: 3
|
|
---
|
|
|
|
# Reflect
|
|
|
|
Generate disposition-aware responses using retrieved memories.
|
|
|
|
When you call **reflect**, Hindsight performs a multi-step reasoning process:
|
|
1. **Recalls** relevant memories from the bank based on your query
|
|
2. **Applies** the bank's disposition traits to shape the reasoning style
|
|
3. **Generates** a contextual answer grounded in the retrieved facts
|
|
4. **Forms opinions** in the background based on the reasoning (available in subsequent calls)
|
|
|
|
The response includes the generated answer along with the facts that were used, providing full transparency into how the answer was derived.
|
|
|
|
import Tabs from '@theme/Tabs';
|
|
import TabItem from '@theme/TabItem';
|
|
import CodeSnippet from '@site/src/components/CodeSnippet';
|
|
|
|
{/* Import raw source files */}
|
|
import reflectPy from '!!raw-loader!@site/examples/api/reflect.py';
|
|
import reflectMjs from '!!raw-loader!@site/examples/api/reflect.mjs';
|
|
import reflectSh from '!!raw-loader!@site/examples/api/reflect.sh';
|
|
|
|
:::info How Reflect Works
|
|
Learn about disposition-driven reasoning and opinion formation in the [Reflect Architecture](/developer/reflect) guide.
|
|
:::
|
|
|
|
:::tip Prerequisites
|
|
Make sure you've completed the [Quick Start](./quickstart) to install the client and start the server.
|
|
:::
|
|
|
|
## Basic Usage
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-basic" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-basic" language="javascript" />
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
<CodeSnippet code={reflectSh} section="reflect-basic" language="bash" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Parameters
|
|
|
|
| Parameter | Type | Default | Description |
|
|
|-----------|------|---------|-------------|
|
|
| `query` | string | required | Question or prompt |
|
|
| `budget` | string | "low" | Budget level: "low", "mid", "high" |
|
|
| `context` | string | None | Additional context for the query |
|
|
| `max_tokens` | int | 4096 | Maximum tokens for the response |
|
|
| `response_schema` | object | None | JSON Schema for [structured output](#structured-output) |
|
|
|
|
### Response Fields
|
|
|
|
| Field | Type | Description |
|
|
|-------|------|-------------|
|
|
| `text` | string | The generated answer text |
|
|
| `based_on` | array | Facts used to generate the response |
|
|
| `structured_output` | object | Parsed structured output (when `response_schema` provided) |
|
|
| `usage` | TokenUsage | Token usage metrics for the LLM call |
|
|
|
|
The `usage` field contains:
|
|
- `input_tokens`: Number of input/prompt tokens consumed
|
|
- `output_tokens`: Number of output/completion tokens generated
|
|
- `total_tokens`: Sum of input and output tokens
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-with-params" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-with-params" language="javascript" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## The Role of Context
|
|
|
|
The `context` parameter steers how the reflection is performed without impacting the memory recall. It provides situational information that helps shape the reasoning and response.
|
|
|
|
**How context is used:**
|
|
- **Shapes reasoning**: Helps understand the situation when formulating an answer
|
|
- **Disambiguates intent**: Clarifies what aspect of the query matters most
|
|
- **Does not affect recall**: The same memories are retrieved regardless of context
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-with-context" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-with-context" language="javascript" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Opinion Formation
|
|
|
|
When reflect reasons about a question, it may form new **opinions** based on the evidence in the memory bank. These opinions are created in the background and become available in subsequent `reflect` and `recall` calls.
|
|
|
|
**Why opinions matter:**
|
|
- **Consistent thinking**: Opinions ensure the memory bank maintains a coherent perspective over time
|
|
- **Evolving viewpoints**: As more information is retained, opinions can be refined or updated
|
|
- **Grounded reasoning**: Opinions are always derived from factual evidence in the memory bank
|
|
|
|
Opinions are stored as a special memory type and are automatically retrieved when relevant to future queries. This creates a natural evolution of the bank's perspective, similar to how humans form and refine their views based on accumulated experience.
|
|
|
|
## Disposition Influence
|
|
|
|
The bank's disposition affects reflect responses:
|
|
|
|
| Trait | Low (1) | High (5) |
|
|
|-------|---------|----------|
|
|
| **Skepticism** | Trusting, accepts claims | Questions and doubts claims |
|
|
| **Literalism** | Flexible interpretation | Exact, literal interpretation |
|
|
| **Empathy** | Detached, fact-focused | Considers emotional context |
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-disposition" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-disposition" language="javascript" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
## Using Sources
|
|
|
|
The `based_on` field shows which memories informed the response:
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
<CodeSnippet code={reflectPy} section="reflect-sources" language="python" />
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
<CodeSnippet code={reflectMjs} section="reflect-sources" language="javascript" />
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
This enables:
|
|
- **Transparency** — users see why the bank said something
|
|
- **Verification** — check if the response is grounded in facts
|
|
- **Debugging** — understand retrieval quality
|
|
|
|
## Structured Output
|
|
|
|
For applications that need to process responses programmatically, you can request structured output by providing a JSON Schema via `response_schema`. When provided, the response includes a `structured_output` field with the LLM response parsed according to the schema. The `text` field will be empty since only a single LLM call is made for efficiency.
|
|
|
|
The easiest way to define a schema is using **Pydantic models**:
|
|
|
|
<Tabs>
|
|
<TabItem value="python" label="Python">
|
|
|
|
```python
|
|
from pydantic import BaseModel
|
|
from hindsight_client import Hindsight
|
|
|
|
# Define your response structure with Pydantic
|
|
class HiringRecommendation(BaseModel):
|
|
recommendation: str
|
|
confidence: str # "low", "medium", "high"
|
|
key_factors: list[str]
|
|
risks: list[str] = []
|
|
|
|
with Hindsight() as client:
|
|
response = client.reflect(
|
|
bank_id="hiring-team",
|
|
query="Should we hire Alice for the ML team lead position?",
|
|
response_schema=HiringRecommendation.model_json_schema(),
|
|
)
|
|
|
|
# Parse structured output into Pydantic model
|
|
result = HiringRecommendation.model_validate(response.structured_output)
|
|
print(f"Recommendation: {result.recommendation}")
|
|
print(f"Confidence: {result.confidence}")
|
|
print(f"Key factors: {result.key_factors}")
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="node" label="Node.js">
|
|
|
|
```javascript
|
|
import { Hindsight } from "@anthropic-ai/hindsight";
|
|
|
|
const client = new Hindsight();
|
|
|
|
// Define JSON schema directly
|
|
const responseSchema = {
|
|
type: "object",
|
|
properties: {
|
|
recommendation: { type: "string" },
|
|
confidence: { type: "string", enum: ["low", "medium", "high"] },
|
|
key_factors: { type: "array", items: { type: "string" } },
|
|
risks: { type: "array", items: { type: "string" } },
|
|
},
|
|
required: ["recommendation", "confidence", "key_factors"],
|
|
};
|
|
|
|
const response = await client.reflect({
|
|
bankId: "hiring-team",
|
|
query: "Should we hire Alice for the ML team lead position?",
|
|
responseSchema: responseSchema,
|
|
});
|
|
|
|
// Structured output
|
|
console.log(response.structuredOutput.recommendation);
|
|
console.log(response.structuredOutput.keyFactors);
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="cli" label="CLI">
|
|
|
|
First, create a JSON schema file `schema.json`:
|
|
```json
|
|
{
|
|
"type": "object",
|
|
"properties": {
|
|
"recommendation": {"type": "string"},
|
|
"confidence": {"type": "string", "enum": ["low", "medium", "high"]},
|
|
"key_factors": {"type": "array", "items": {"type": "string"}}
|
|
},
|
|
"required": ["recommendation", "confidence", "key_factors"]
|
|
}
|
|
```
|
|
|
|
Then use the `--schema` flag:
|
|
```bash
|
|
hindsight memory reflect hiring-team \
|
|
"Should we hire Alice for the ML team lead position?" \
|
|
--schema schema.json
|
|
```
|
|
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
| Use Case | Why Structured Output Helps |
|
|
|----------|----------------------------|
|
|
| **Decision pipelines** | Parse recommendations into workflow systems |
|
|
| **Dashboards** | Extract confidence scores, risk factors for visualization |
|
|
| **Multi-agent systems** | Pass structured data between agents |
|
|
| **Auditing** | Log structured decisions with clear reasoning |
|
|
|
|
**Tips:**
|
|
- Use Pydantic's `model_json_schema()` for type-safe schema generation
|
|
- Use `model_validate()` to parse the response back into your Pydantic model
|
|
- Keep schemas focused — extract only what you need
|
|
- Use `Optional` fields for data that may not always be available
|