* feat: introduce hindsight-api-slim and hindsight-all-slim packages Closes #552 - Move all source code from hindsight-api/ to new hindsight-api-slim/ - hindsight-api-slim has heavy ML deps (torch, sentence-transformers, transformers, einops, flashrank, mlx, mlx-lm, safetensors) and pg0-embedded as optional extras: [local-ml], [embedded-db], [all] - hindsight-api becomes a zero-code meta-package depending on hindsight-api-slim[all] for full backward compatibility - Add hindsight-all-slim meta-package: hindsight-api-slim + client + embed - hindsight-all updated to depend on hindsight-api-slim[all] - pg0.py: lazy-import pg0 with clear ImportError pointing to [embedded-db] - Dockerfile: replace sed hack with proper uv sync --extra flags - Update release.yml, test.yml, lint.sh, release.sh, CLAUDE.md and all path references throughout the repo * refactor: rename hindsight/ directory to hindsight-all/ * docs: document hindsight-api-slim and hindsight-all-slim package variants Add package variants table and extras explanation to installation.md * docs: remove emojis from installation.md, use professional tone * docs: link Docker slim variant to pip package variants section * docs: consolidate Docker image variants into single table * ci: fix working-directory paths after package restructure - Replace all hindsight-api → hindsight-api-slim in test.yml - Replace hindsight → hindsight-all in test.yml - Add --extra embedded-db to test-embed API install step * ci: add local-ml and embedded-db extras to API sync steps These extras were previously implicit in the old hindsight-api package (which bundled everything). Now that hindsight-api-slim uses optional extras, we must explicitly request local-ml and embedded-db in CI. * ci: add API install step with embedded-db to test-embed smoke test The smoke test starts hindsight-api as a daemon, which requires pg0-embedded. Add a dedicated install step for hindsight-api-slim with embedded-db extra so the daemon can start successfully. * ci: remove --no-install-project when using optional extras When --no-install-project is combined with --extra, the optional deps are not installed because extras require the project to be active. Remove --no-install-project from steps that need local-ml or embedded-db. * ci: fix ordering of uv sync steps to preserve optional extras When uv sync runs for a different workspace member, it removes optional extras installed for other members. Fix by always running extra-requiring API sync last, after other workspace member syncs. Also remove --no-install-project from embedded-db sync in test-embed, as --no-install-project prevents optional extras from being active. * ci: add local-ml extra to test-embed API install for smoke test The smoke test starts the full API server which needs sentence-transformers for local embeddings (default provider). Add local-ml extra to the install. * ci: simplify extras with --all-extras and add slim pip smoke test - Replace explicit --extra local-ml --extra embedded-db with --all-extras for cleaner, more maintainable sync steps - Add test-pip-slim job: tests hindsight-api-slim[embedded-db] without local ML models, using Cohere for embeddings/reranking (mirrors Docker slim smoke test approach) * ci: simplify slim smoke test to health check only (mirrors Docker test)
103 lines
3.5 KiB
Python
103 lines
3.5 KiB
Python
"""
|
|
Test to analyze fact extraction token usage and identify optimization opportunities.
|
|
"""
|
|
import asyncio
|
|
import logging
|
|
import time
|
|
from datetime import datetime
|
|
|
|
import pytest
|
|
|
|
from hindsight_api.config import get_config, clear_config_cache, _get_raw_config
|
|
from hindsight_api.engine.llm_wrapper import LLMConfig
|
|
from hindsight_api.engine.retain.fact_extraction import extract_facts_from_text
|
|
|
|
logging.basicConfig(level=logging.INFO)
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
@pytest.fixture
|
|
def llm_config():
|
|
"""Create LLM config from environment."""
|
|
clear_config_cache()
|
|
config = get_config()
|
|
return LLMConfig(
|
|
provider=config.retain_llm_provider or config.llm_provider,
|
|
api_key=config.retain_llm_api_key or config.llm_api_key,
|
|
model=config.retain_llm_model or config.llm_model,
|
|
base_url=config.retain_llm_base_url or config.llm_base_url,
|
|
)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_fact_extraction_basic_analysis(llm_config):
|
|
"""
|
|
Test fact extraction and analyze token usage with sample content.
|
|
|
|
This test helps identify:
|
|
1. How many facts are extracted
|
|
2. Token usage (input/output ratio)
|
|
3. Types of facts being extracted
|
|
"""
|
|
content = """
|
|
Alice is a senior software engineer at TechCorp with 8 years of experience.
|
|
She has a Kubernetes certification (CKA) and leads the platform team.
|
|
Bob is her colleague who works on the frontend. He's been at the company for 3 years.
|
|
They're working on a new microservices migration project together.
|
|
The deadline for the first milestone is end of Q2.
|
|
Alice prefers to use Go for backend services while Bob advocates for TypeScript.
|
|
"""
|
|
|
|
logger.info(f"Content length: {len(content)} chars (~{len(content) // 4} tokens)")
|
|
|
|
start_time = time.time()
|
|
|
|
facts, chunks, usage = await extract_facts_from_text(
|
|
text=content,
|
|
event_date=datetime.now(),
|
|
llm_config=llm_config,
|
|
agent_name="test-agent",
|
|
context="Friday Standup meeting",
|
|
config=_get_raw_config(),
|
|
)
|
|
|
|
duration = time.time() - start_time
|
|
|
|
logger.info(f"\n{'='*60}")
|
|
logger.info(f"EXTRACTION RESULTS")
|
|
logger.info(f"{'='*60}")
|
|
logger.info(f"Duration: {duration:.2f}s")
|
|
logger.info(f"Chunks: {len(chunks)}")
|
|
logger.info(f"Facts extracted: {len(facts)}")
|
|
logger.info(f"Input tokens: {usage.input_tokens}")
|
|
logger.info(f"Output tokens: {usage.output_tokens}")
|
|
logger.info(f"Token ratio (out/in): {usage.output_tokens / max(1, usage.input_tokens):.2f}")
|
|
|
|
# Analyze facts by type
|
|
fact_types = {}
|
|
for fact in facts:
|
|
ft = fact.fact_type
|
|
fact_types[ft] = fact_types.get(ft, 0) + 1
|
|
|
|
logger.info(f"\nFacts by type:")
|
|
for ft, count in sorted(fact_types.items()):
|
|
logger.info(f" {ft}: {count}")
|
|
|
|
# Show sample facts
|
|
logger.info(f"\nSample facts (first 10):")
|
|
for i, fact in enumerate(facts[:10]):
|
|
logger.info(f"\n [{i+1}] {fact.fact_type}: {fact.fact[:150]}...")
|
|
|
|
# Show facts containing key terms
|
|
key_terms = ["kubernetes", "k8s", "CKA", "certification", "Alice"]
|
|
logger.info(f"\n{'='*60}")
|
|
logger.info(f"FACTS CONTAINING KEY TERMS")
|
|
logger.info(f"{'='*60}")
|
|
|
|
for term in key_terms:
|
|
matching = [f for f in facts if term.lower() in f.fact.lower()]
|
|
logger.info(f"\n'{term}' ({len(matching)} facts):")
|
|
for fact in matching[:3]:
|
|
logger.info(f" - {fact.fact[:200]}...")
|
|
|
|
assert len(facts) > 0, "Should extract at least one fact"
|