Skip to content
Memory runtimeOpen-source · Apache 2.0 · self-hosted

Letta lets the model manage memoryStatewave manages memory for the model

Letta hands memory to the model and lets it decide. Statewave assembles a deterministic, token-bounded context bundle and can return an integrity-hashed receipt of what the agent saw.

S

ContextAssembler

deterministic · token-bounded

budget 1,180/1,500
ranked context bundle
profile_factprefers succulents, low-waterrank 1
procedurerefund flow for orders < 30drank 2
episodeold address on fileexpired · dropped
receipt01J9Z4··· · sha256:a3f9c1 · bundle:7c21 (enforce)
0.905LoCoMo, n=1,540
4ranking signals, deterministic
0.967LongMemEval, n=30
0API keys to run
Postgres-only, runs offline

The model manages memory, or a runtime does

Letta gives the agent tools to edit its own memory blocks, tracked as commits in a git-backed context tree (MemFS): memory is the model's job, decided turn by turn and paid for in tokens. Statewave ingests each event as an immutable episode, compiles it into typed memories, and assembles a ranked bundle the same way every call, with no model in the loop.

laLetta · MemFS, agent-managed
contexteverything the agent sees, one git-tracked tree
blocksmemory files inside that tree, model-edited
historyevery edit is a commit, diffable and revertable
agent tool call →the model decides what to read and write

Retrieval rides on the model

There's no ranking or expiry step: the agent reads and edits its own context tree directly, and whatever it last wrote is what it sees next time. Every operation spends inference tokens, and an unwritten fact is gone.

SStatewave · Record → Compile → Context → Govern
01Recordimmutable episodes
02Compiletyped memories
03Contextranked bundle
04Governpolicy + receipt
typed memory kinds · ranking priority
profile_fact
10
procedure
8
episode_summary
5
raw_episode
3

Assembly is deterministic and inspectable

Given the same subject, task, token budget, and point in time, the assembler returns the identical bundle every run. Four signals set the order: kind priority, recency, task relevance, and temporal validity.

How order is decidedscore = priority + recency + relevance + validity
KIND PRIORITY
3–10
typed profile facts outrank raw episodes
RECENCY
0–5
linear by age, newest scores highest
TASK RELEVANCE
0–8
word overlap (0-5) or cosine similarity (0-8)
TEMPORAL VALIDITY
−4…+3
valid facts gain +3, expired ones lose 4

Where each capability lives

Both give agents memory. They diverge on who does the work (the model, or the runtime) and on what you can prove once it has.

RETRIEVAL & RANKING

How context is selected

la
The agent reads and edits its own context tree with tool calls
S
Deterministic assembly, ranked to a token budget

What ranks results

la
The model’s judgment, turn by turn
S
Kind priority, recency, relevance, validity

Cost of retrieval

la
Inference tokens on every memory operation
S
Mechanical; no LLM on the read path

Same query, same result

la
Varies with the model and its tool choices
S
Byte-identical bundle every run
GOVERNANCE & PROVENANCE

Proof of what the agent saw

la
None; reconstruct it from message logs
S
Immutable, ULID-addressable receipt with an integrity hash

Policy on the read path

la
Implement it in your app or tools
S
Declarative bundles: deny or redact by label and caller

Reliability of capture

la
If the model does not save it, it is gone
S
Every event recorded as an immutable episode

Subject deletion (GDPR)

la
Remove or rewrite files in the context tree directly
S
One call clears episodes, memories, and receipts
OPERATIONS & LICENSING

Scope

la
A full stateful-agent platform: you adopt the runtime
S
A memory layer that drops into your existing stack

Storage

la
Self-hosted App Server (Docker Compose, Railway, or Fly.io); memory lives in git-tracked MemFS, no separate archival store
S
Postgres and pgvector, nothing else to run

Interface

la
CLI (letta-code), desktop app, chat.letta.com, Slack/Telegram/Discord
S
REST, Python and TypeScript SDKs, MCP server, connectors

License

la
Apache 2.0, with managed Letta Cloud
S
Apache 2.0 throughout, runs fully offline

Letta (formerly MemGPT) manages memory through the agent's own tool calls against a git-tracked context tree (MemFS), so retrieval depends on the model driving it. Rows reflect each product's public docs and source as of August 2026.

laReach for Letta when

You're building agent-first, want the model to manage its own context, and value adaptive, self-editing memory over reproducibility or an audit trail.

SReach for Statewave when

Agents run in production across many sessions and you need deterministic context, source provenance, policy on the read path, and an optional auditable receipt for every decision.

One returning customer, two runtimes

A support agent resumes a customer thread three weeks later. In between, the customer moved house and pasted a card number into an earlier message. The same episode history runs through each system.

laLetta · agent reads its own context tree
# the agent's memory is its working tree
› agent.send_message("where do I ship it")
  ↳ cat context/cust_5521.md
# returns whatever's currently in the file
old address · Elm Ststale, still stored
card 4242 4242 ····pii, no policy gate
There's no ranking or expiry step: whatever the model last wrote to the file is what it reads next time, stale entries and all. Redaction and retention stay the agent's job.
SStatewave · assemble + govern
# assemble a ranked, bounded bundle
› get_context(subject="cust_5521",
  task="where do I ship it", max_tokens=1500)
new address · Oak Avevalid +3 · ranked #1
old address · Elm St−4 expired · dropped
card ●●●● ●●●● ····label:pii · redacted
The runtime decides: the superseded address scores out, the card is redacted by its policy label, and the receipt stores an integrity hash of what was delivered.

Every call can leave a receipt

Letta leaves auditability to your application and its message logs. Statewave governs assembly on the read path and emits an immutable receipt, all in the Apache 2.0 core.

state-assembly receiptimmutable · ULID-addressable
receipt_id01J9Z4RT8K···
integrity_hashsha256:a3f9c1e0···
policy_bundlebundle:7c21 (enforce)
included · 3 facts, 2 episodes · 1,180/1,500 tokens
profile_factconf 0.92 · valid[ep_4, ep_9]
procedureconf 0.88 · valid[ep_2]
episode_summarysupersededdropped
1 memory redacted · label:pii
{}
Policy engine
Content-hashed YAML or JSON bundles. Deny or redact by sensitivity label and caller identity; log_only audits a policy before you enforce it.
#
Sensitivity labels
Per-memory pii, financial, and secret tags in a GIN-indexed array, so policy filters run inside the query.
←
Full provenance
Every compiled memory keeps the source episode ids, confidence score, and validity window it was derived from.
⌫
Subject deletion
One GDPR-style call erases every episode, memory, and receipt for a subject. No orphaned rows.

The benchmarks Statewave has tested

Statewave's scores are fixed: no model sits on the read path, so the same subject and budget return the same answer every run. Letta has published a LoCoMo figure of its own (74.0%, GPT-4o mini, above Mem0's 68.5%), but not one produced under the same conditions as this harness. What it does publish continuously is a leaderboard that scores the driving LLM, and the same runtime swings 56 points and 21× in cost by model.

SStatewave: fixed benchmark scoresMODEL-INDEPENDENT
0.905accuracy
LoCoMo
n=1,540 · robust figure
0.967accuracy
LongMemEval
n=30 · directional
Hybrid retrieval lift · v10
LoCoMo +2.1LongMemEval +16.0
laLetta: a range, not a numberVARIES BY MODEL

Letta's leaderboard runs the identical runtime across 15 LLMs. The score is the model's, not the memory's: pick a different model and it moves. leaderboard.letta.com

Context-Bench · filesystem suite · 15 models
30%100%
37%
minimax/minimax-m2.5 · lowest
93%
openai/gpt-5.2-codex-xhigh · best
56 pts
score spread across 15 models
21×
cost spread, $27.36 → $567.66

Context-Bench measures an agent's context engineering, not a memory layer in isolation, so it is not comparable to the LoCoMo and LongMemEval figures above. It is shown instead to make the opposite point: the number moves with the model. leaderboard.letta.com, last updated 13 March 2026.

reproduce it yourself · statewave-memory-benchmarks
$ git clone https://github.com/smaramwbc/statewave-memory-benchmarks.git
$ cd statewave-memory-benchmarks && pip install -r requirements.txt
$ python -m benchmarks.locomo.run --backend statewave \
  --answerer-model gpt-4o --judge-model gpt-4o
# same gpt-4.1 extraction, judge & scoring code untouched from upstream

Add Statewave to your Letta stack

Letta is the whole agent runtime; Statewave is only the memory layer. Keep your agent loop and point its memory calls at Statewave: each write lands as an immutable episode.

one command · connects any MCP client
$ npx @statewavedev/statewave
→ API + admin console + Postgres up via Docker · healthy in under 2 min · no account
LETTA
STATEWAVE
memory_insert("context/cust_5521.md", "…")
create_episode(subject="cust_5521", event="…")
Ingested as an immutable episode; compilers extract typed memories.
cat context/cust_5521.md
get_context(subject="cust_5521", task="…", max_tokens=600)
MemFS has no search or ranking step, only whatever the file currently holds; Statewave returns a ranked, token-bounded bundle plus an optional receipt.
memory_replace("context/cust_5521.md", "…")
create_episode(subject="cust_5521", event="…")
Facts are compiled from episodes, not hand-edited files.
git rm context/cust_5521.md
delete_subject("cust_5521")
Removes every episode, memory, and receipt for the subject in one call.

Frequently asked

How is Statewave different from Letta?

In Letta the agent manages its own memory: it edits memory blocks in a git-tracked context tree (MemFS) with tool calls, so retrieval is the model’s job and costs tokens every turn. Statewave compiles episodes into typed memories, ranks them to a token budget, applies policy on the read path, and can return an integrity-hashed receipt, with no model in the loop.

Do I have to replace my agent framework?

No. Letta is a whole agent runtime; Statewave is only the memory layer. Keep your existing agent or framework and point its memory reads and writes at Statewave over REST, the SDKs, or MCP.

What makes retrieval deterministic?

A fixed scoring model applied to a hybrid lexical and vector candidate set: kind priority (3–10), recency (0–5), task relevance (0–8), and temporal validity (−4 to +3). The same subject, task, budget, and point in time produce the same bundle every time.

What is a state-assembly receipt?

An immutable, ULID-addressable record of one context call. It carries a byte-level integrity hash of what was delivered and references the policy bundle hash, so ‘what did the agent see, under which policy’ is answerable forever.

Does it work with Claude, Cursor, or Codex?

Yes. One command (npx @statewavedev/statewave) boots the runtime, and its shipped MCP server connects any MCP-compatible client: Claude, Cursor, Copilot, and agent runtimes.

Can I run it fully offline?

Yes. Storage is Postgres-only and self-hosted. The heuristic compiler keeps everything on your network; nothing leaves unless you configure an LLM compiler or hosted embeddings.

Give your agent context it can prove

Self-host the Apache 2.0 runtime, wire it to your MCP client, and every context call can return a receipt.