Pipeline walkthrough
Every mechanism below mirrors real code paths — source file named per step. ← landing page
Retain — POST /memory/retain
A memory lands in one of three indices with a full tenancy envelope.
- Validate — kind ∈ episodic/semantic/procedural, visibility ∈ private/team/common
memory.py:retain
- Tombstone check — if the user rejected this exact value before, retain raises
tombstone.py:is_rejected
- Embed — dense vector via ES inference or OpenAI-compatible endpoint; on failure → BM25-only, never a rejected write
embeddings.py
- Index — doc with owner_id (from API key), visibility, entities, occurred_at
store.py MAPPINGS
Source: backend/app/memory.py · store.py
Recall — POST /memory/recall
Two retrieval arms over the visibility-filtered set, fused by rank.
- Visibility filter — owner_id == caller OR visibility ∈ {team, common}; identity is the API key, never an argument
memory.py:_visibility_filter
- BM25 arm — match over text + entities
- kNN arm — top-k by dense_vector similarity
- RRF fuse — rank-based fusion, then token-budget cut
Source: backend/app/memory.py · diagram: docs/diagrams/recall-pipeline.svg
Reflect — POST /memory/reflect
Recall as evidence, then an LLM answers with citations.
- Gather evidence — same recall pipeline, top hits as context
memory.py
- LLM call — any OpenAI-compatible endpoint
llm.py (AMES_LLM_BASE)
- Cited answer — response references memory ids it used
main.py:reflect
Q: how do deploys work?
A: Deploys go through OmniRoute catchup branches [ames-semantic:Zu3f...]
— 1 source, owner-scoped evidence only
Source: backend/app/llm.py · main.py
Promote & the guard — POST /memory/promote
private → team/common is an explicit, guarded operation.
- Deterministic guard — regex over sensitive markers (api key, token, secret, credential, ssh key…); a hit blocks shared promotion
memory.py:guard_promotion
- Same text, same decision — no classifier, no LLM in the path
- Rejection tombstones — "forget this" survives as a block on re-learning
tombstone.py
Source: backend/app/memory.py (guard_promotion, _PRIVATE_MARKERS) · tests: backend/tests/test_promotion_guard.py
Consolidate — worker loop
Semantic facts rot; the worker proposes and then applies supersessions.
- Every interval — default 600 s
worker.py (AMES_CONSOLIDATE_INTERVAL)
- Propose pass — dry-run: duplicates / supersessions surfaced for visibility
memory.consolidate(dry_run=True)
- Apply pass — supersede old facts; keep_both cases are surfaced, never auto-merged
- Audit — superseded docs flagged inactive, not deleted
14:02:11 {'alice': {'proposals': 3, 'superseded': 2}}
Source: backend/app/worker.py · tests: backend/tests/test_worker.py