Files
vectorize-io__hindsight/hindsight-api-slim
Nicolò Boschi b05546b57d fix(recall): a store-owned bank's attachment lookup does not read memory_units (#4306)
Recall resolves the attachments behind each fact by reading memory_units.attachment_ids, on
every recall. For a bank whose memories store owns its rows that table holds none of them, so
the read can only return nothing -- and it is not the cheap empty read it looks like.

memory_units carries partial vector indexes per bank, and the planner opens and locks every
index on a table to plan any statement against it. In a tenant with 5,286 banks that is 15,877
indexes: planning took 434-484 ms and 15,880 locks (15,863 on the slow path) for a statement
that executed in 0.04 ms, against ~2 ms and 22 locks in a tenant with a few banks. Twenty
concurrent recalls saturated LWLock:LockManager for the whole database: a 2 vCPU API pod fell
from 180 to ~13 recalls/s, and unrelated pods' readiness checks slowed from 6 ms to ~870 ms.

The guard uses the same store_owned_for check as the other store-owned paths. Store-owned banks
report no attachments on recall until the store can carry the ids itself.

The test asserts the lookup returns before it reads the bank profile or takes a connection,
not merely that it returns {}: an empty result is what the expensive read produced as well.
2026-09-11 09:13:29 +02:00
..

Hindsight API

Memory System for AI Agents — Temporal + Semantic + Entity Memory Architecture using PostgreSQL with pgvector.

Hindsight gives AI agents persistent memory that works like human memory: it stores facts, tracks entities and relationships, handles temporal reasoning ("what happened last spring?"), and forms opinions based on configurable disposition traits.

Installation

pip install hindsight-api

Quick Start

Run the Server

# Set your LLM provider
export HINDSIGHT_API_LLM_PROVIDER=openai
export HINDSIGHT_API_LLM_API_KEY=sk-xxxxxxxxxxxx

# Start the server (uses embedded PostgreSQL by default)
hindsight-api

The server starts at http://localhost:8888 with:

  • REST API for memory operations
  • MCP server at /mcp for tool-use integration

Use the Python API

from hindsight_api import MemoryEngine

# Create and initialize the memory engine
memory = MemoryEngine()
await memory.initialize()

# Create a memory bank for your agent
bank = await memory.create_memory_bank(
    name="my-assistant",
    background="A helpful coding assistant"
)

# Store a memory
await memory.retain(
    memory_bank_id=bank.id,
    content="The user prefers Python for data science projects"
)

# Recall memories
results = await memory.recall(
    memory_bank_id=bank.id,
    query="What programming language does the user prefer?"
)

# Reflect with reasoning
response = await memory.reflect(
    memory_bank_id=bank.id,
    query="Should I recommend Python or R for this ML project?"
)

CLI Options

hindsight-api --help

# Common options
hindsight-api --port 9000          # Custom port (default: 8888)
hindsight-api --host 127.0.0.1     # Bind to localhost only
hindsight-api --workers 4          # Multiple worker processes
hindsight-api --log-level debug    # Verbose logging

Configuration

Configure via environment variables:

Variable Description Default
HINDSIGHT_API_DATABASE_URL PostgreSQL connection string pg0 (embedded)
HINDSIGHT_API_LLM_PROVIDER LLM provider, including openai, anthropic, gemini, groq, ollama, lmstudio, and github-copilot openai
HINDSIGHT_API_LLM_API_KEY API key for LLM provider -
HINDSIGHT_API_LLM_MODEL Model name gpt-4o-mini
HINDSIGHT_API_HOST Server bind address 0.0.0.0
HINDSIGHT_API_PORT Server port 8888

Example with External PostgreSQL

export HINDSIGHT_API_DATABASE_URL=postgresql://user:pass@localhost:5432/hindsight
export HINDSIGHT_API_LLM_PROVIDER=groq
export HINDSIGHT_API_LLM_API_KEY=gsk_xxxxxxxxxxxx

hindsight-api

Docker

docker run -it --name hindsight --restart unless-stopped -p 8888:8888 \
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
  -v $HOME/.hindsight-docker:/home/hindsight/.pg0 \
  ghcr.io/vectorize-io/hindsight:latest

MCP Server

For local MCP integration without running the full API server:

hindsight-local-mcp

This runs a stdio-based MCP server that can be used directly with MCP-compatible clients.

Key Features

  • Multi-Strategy Retrieval (TEMPR) — Semantic, keyword, graph, and temporal search combined with RRF fusion
  • Entity Graph — Automatic entity extraction and relationship tracking
  • Temporal Reasoning — Native support for time-based queries
  • Disposition Traits — Configurable skepticism, literalism, and empathy influence opinion formation
  • Three Memory Types — World facts, experience facts (the bank's own actions), and observations

Documentation

Full documentation: https://hindsight.vectorize.io

License

Apache 2.0