Open source · MIT · Phase 0 spike

Your agents deserve
a memory that survives the session.

agent-memory-es is a self-owned, Elasticsearch-native long-term memory service for AI agents — three memory tiers, hybrid recall, and strict private/team/common visibility. Run it on your cluster; no memory SaaS between your agents and their past.

3memory tiers — episodic · semantic · procedural
2→1BM25 + kNN arms, RRF-fused into one ranking
0visibility a client can widen — owner comes from the API key

Why

Context windows are whiteboards.
They get erased.

Every session starts from zero: re-explained preferences, re-discovered decisions, re-learned mistakes. A memory bank fixes that — with the same discipline you'd expect from any multi-user data store.

🧠

Three kinds of memory

episodic raw events in the user's words, semantic distilled facts, procedural playbooks. Each searched together or apart.

🔐

Tenancy is not a filter arg

Owner identity comes from the API key. The recall filter owner_id == me OR visibility ∈ {team, common} is applied server-side, on both retrieval arms.

🛡️

Guarded promotion

private → team/common passes a deterministic sensitive-marker guard. "Forget this" tombstones block re-learning forgotten values.

🔎

Hybrid recall

BM25 over text + entities, dense kNN over embeddings, fused with RRF, cut to a token budget. No embeddings backend? Degrades to BM25-only.

🧹

Memory that maintains itself

A consolidation worker proposes and applies dedup/supersession across semantic facts. Superseded docs stay as audit history.

🔌

Any agent can plug in

Hermes provider plugin, MCP endpoint (Claude Code, Cursor, Codex), or plain curl. Same five operations everywhere.

Visibility model

Shared knowledge, private banks.

One store, three visibility levels, one server-side rule. A team playbook and a personal API key never mix by accident.

private

Only the owner

Default for everything. Recall filter pins owner_id to the API key's owner — never a request argument.

team

Team-shared

Explicit promotion only, through the deterministic guard. Sensitive-marker text stays private, same text ⇒ same decision.

common

Everyone

Fleet-wide playbooks and procedural knowledge. Rejection tombstones keep forgotten values from sneaking back.

Architecture

Small surface, boring parts, real guarantees.

FastAPI in front of Elasticsearch, a consolidation worker behind it. Every box below is clickable in the interactive version — with links to the exact source lines.

System topology archify · source-linked · open full diagram ↗

Recall pipeline

Two arms, one ranking, zero score tuning.

01

Visibility filter

Applied to both arms before anything scores.

02

BM25 arm

Match over text + extracted entities.

03

kNN arm

Dense-vector top-k similarity.

04

RRF fuse

Rank fusion across arms — no cross-arm score tuning.

05

Budget cut

Top fused hits until the token budget is spent.

Recall pipeline (static) open SVG ↗
Recall pipeline: BM25 arm and kNN arm, both visibility-filtered, fused with RRF, cut at a token budget

Quick start

Running in three commands.

# 1. boot the stack (backend :8123 + worker + Elasticsearch)
cd backend && export AMES_ADMIN_TOKEN=change-me && docker compose up -d

# 2. mint a key — owner identity binds to it
curl -s -X POST "localhost:8123/admin/keys?owner_id=alice" -H "X-Admin-Token: $AMES_ADMIN_TOKEN"

# 3. remember, then recall
curl -s -X POST localhost:8123/memory/retain -H "X-API-Key: $KEY" \
  -d '{"kind":"semantic","text":"Deploys go through OmniRoute catchup branches"}'
curl -s -X POST localhost:8123/memory/recall -H "X-API-Key: $KEY" \
  -d '{"query":"how do deploys work"}'

Hermes Agent user? Install the standalone plugin ↗ and these become native agent tools.