Files
vectorize-io__hindsight/hindsight-api-slim
Nicolò Boschi 179938a655 fix(retain): stop chunk ids colliding across banks (#4257)
* fix(retain): stop chunk ids colliding across banks

Chunk ids flattened (bank_id, document_id, chunk_index) with a plain
underscore join, so ('a', 'b_c') and ('a_b', 'c') both produced 'a_b_c_0'.
Both are arbitrary caller-supplied strings, and 'chunks' is keyed on the id
alone, so the second bank's retain overwrote the first bank's chunk row.

Build the id through engine/chunk_ids.py, which escapes the separator inside
each component and so is injective. Ids are unchanged where neither component
contains a separator; parsing still reads the ambiguous ids written before
this. Ids that already collide in a deployed database stay possible, so the
upsert now refuses a conflicting row owned by another bank instead of
overwriting it, and the delta delete is scoped to the bank when one is given.

Fixes #4244

* test(retain): cover legacy chunk ids on read and on update

A document stored before the id fix keeps its unescaped chunk ids — there is
no migration — so both reading it and re-retaining over it have to stay
correct. Ages a freshly retained document's rows back to the legacy shape and
asserts the addressed chunk route still resolves them, and that editing one
section replaces exactly the chunks covering it (escaped id) while every other
chunk keeps the legacy id it was stored under, one row per chunk_index.

* refactor(retain): require bank_id on the chunk delete, drop the duplicate id helper

Review follow-ups on the #4244 fix. `delete_chunks_by_ids` took bank_id as an
optional argument with a None default, so its bank predicate was applied under a
conditional — the latent shape the isolation rules warn about, even though every
caller passes one. Make it required and let both statements carry it
unconditionally.

`chunk_index_in` and `resolve_chunk_id_in` resolved the same id the same way for
every input; the caller that had the document id can compare it against the
resolved one instead, so only the latter remains.
2026-09-09 14:54:55 +02:00
..

Hindsight API

Memory System for AI Agents — Temporal + Semantic + Entity Memory Architecture using PostgreSQL with pgvector.

Hindsight gives AI agents persistent memory that works like human memory: it stores facts, tracks entities and relationships, handles temporal reasoning ("what happened last spring?"), and forms opinions based on configurable disposition traits.

Installation

pip install hindsight-api

Quick Start

Run the Server

# Set your LLM provider
export HINDSIGHT_API_LLM_PROVIDER=openai
export HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx

# Start the server (uses embedded PostgreSQL by default)
hindsight-api

The server starts at http://localhost:8888 with:

  • REST API for memory operations
  • MCP server at /mcp for tool-use integration

Use the Python API

from hindsight_api import MemoryEngine

# Create and initialize the memory engine
memory = MemoryEngine()
await memory.initialize()

# Create a memory bank for your agent
bank = await memory.create_memory_bank(
    name="my-assistant",
    background="A helpful coding assistant"
)

# Store a memory
await memory.retain(
    memory_bank_id=bank.id,
    content="The user prefers Python for data science projects"
)

# Recall memories
results = await memory.recall(
    memory_bank_id=bank.id,
    query="What programming language does the user prefer?"
)

# Reflect with reasoning
response = await memory.reflect(
    memory_bank_id=bank.id,
    query="Should I recommend Python or R for this ML project?"
)

CLI Options

hindsight-api --help

# Common options
hindsight-api --port 9000          # Custom port (default: 8888)
hindsight-api --host 127.0.0.1     # Bind to localhost only
hindsight-api --workers 4          # Multiple worker processes
hindsight-api --log-level debug    # Verbose logging

Configuration

Configure via environment variables:

Variable Description Default
HINDSIGHT_API_DATABASE_URL PostgreSQL connection string pg0 (embedded)
HINDSIGHT_API_LLM_PROVIDER LLM provider, including openai, anthropic, gemini, groq, ollama, lmstudio, and github-copilot openai
HINDSIGHT_API_LLM_API_KEY API key for LLM provider -
HINDSIGHT_API_LLM_MODEL Model name gpt-4o-mini
HINDSIGHT_API_HOST Server bind address 0.0.0.0
HINDSIGHT_API_PORT Server port 8888

Example with External PostgreSQL

export HINDSIGHT_API_DATABASE_URL=postgresql://user:pass@localhost:5432/hindsight
export HINDSIGHT_API_LLM_PROVIDER=groq
export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx

hindsight-api

Docker

docker run -it --name hindsight --restart unless-stopped -p 8888:8888 \
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
  -v $HOME/.hindsight-docker:/home/hindsight/.pg0 \
  ghcr.io/vectorize-io/hindsight:latest

MCP Server

For local MCP integration without running the full API server:

hindsight-local-mcp

This runs a stdio-based MCP server that can be used directly with MCP-compatible clients.

Key Features

  • Multi-Strategy Retrieval (TEMPR) — Semantic, keyword, graph, and temporal search combined with RRF fusion
  • Entity Graph — Automatic entity extraction and relationship tracking
  • Temporal Reasoning — Native support for time-based queries
  • Disposition Traits — Configurable skepticism, literalism, and empathy influence opinion formation
  • Three Memory Types — World facts, experience facts (the bank's own actions), and observations

Documentation

Full documentation: https://hindsight.vectorize.io

License

Apache 2.0