Skip to content

Why Statewave

We built Statewave because we believe AI should remember.

Not a chat history. Memory — the kind that holds across days, makes promises stick, and turns yesterday's lesson into today's better answer.

Every system we admire works because someone remembered. The line of code that didn't need rewriting. The customer's name. The mistake we agreed never to repeat.

So we built it carefully. Self-hosted, because what your agents remember is yours. Provenance-first, because any answer worth giving is worth tracing. Open source, because infrastructure this important shouldn't live behind someone else's login.

In the end, only memories matter.

Built with care, in the open.
THE PROBLEM

The infrastructure gap

AI support agents forget. Every session starts from zero. Returning customers re-explain who they are, what plan they're on, what they asked last time. Agents make the same mistakes they made before. This isn't a capability gap in the LLM — it's an infrastructure gap. Most AI applications have no memory layer.

Session 01

Context is created.

Session 02

Context is lost.

Result

The agent starts again.

How it works

From question to memory-aware answer

Statewave assembles a token-bounded, ranked context bundle so every reply is grounded in the subject's real history.

  1. 1

    A returning user asks

    "Following up on the billing issue from last week — any update?"

    The question implicitly references a prior episode the agent has no built-in way to recall.

  2. 2

    Statewave assembles context

    Search memories

    • • preferences · procedures
    • • past decisions

    Pull relevant episodes

    • • recent interactions
    • • support events

    Rank & filter

    • • ranked, deterministic, token-bounded
  3. 3

    The LLM answers with memory

    Personalized, grounded, and aware of prior history.

    "Yes — the refund posted Friday and your account is active again."

COMPARISON

Statewave vs alternatives

Deterministic

Prompt stuffing
✗
Naive RAG
✗
Statewave
✓

Token-bounded

Prompt stuffing
✗
Naive RAG
Truncation
Statewave
✓ Ranked packing

Provenance

Prompt stuffing
✗
Naive RAG
✗
Statewave
✓ Episode-level

Structured extraction

Prompt stuffing
✗
Naive RAG
✗
Statewave
✓ Typed memories

Temporal reasoning

Prompt stuffing
✗
Naive RAG
✗
Statewave
✓ Validity windows

Confidence scoring

Prompt stuffing
✗
Naive RAG
✗
Statewave
✓

Idempotent

Prompt stuffing
N/A
Naive RAG
N/A
Statewave
✓

Subject lifecycle

Prompt stuffing
✗
Naive RAG
✗
Statewave
✓ Full CRUD + delete

Cost at scale

Prompt stuffing
Linear growth
Naive RAG
Index bloat
Statewave
Bounded by budget
TECHNICAL FOUNDATION

Key technical properties

Event sources flowing into the stateful workflows stack

Deterministic

Same subject + task + budget → same context bundle. No non-determinism from vector-only retrieval.

Token-bounded

Context assembly respects a configurable token budget. Items are packed by ranked score, not truncated arbitrarily.

Provenance-traced

Every memory traces to its source episode IDs. Every context bundle reports which facts and episodes were included.

Idempotent

Recompiling the same subject produces no duplicate memories. Safe to run on schedule or on-demand.

Subject-centric

Everything organized around subjects. Full lifecycle: ingest → compile → retrieve → inspect → delete.

Self-hosted storage

Postgres-only. Episodes and compiled memories stay in your infrastructure. Whether prompt content leaves depends on your compiler and embedding choice — heuristic mode is fully local.

WHO IT IS FOR

Who this is for

Good fit

  • Teams building AI support agents with returning customers
  • Engineering leads who want measurable context quality
  • Teams that need provenance — "why did the agent say X?"
  • Self-hosted storage requirements — episodes and memories stay on your infrastructure (heuristic compiler keeps everything local; LLM compiler or hosted embeddings will send content to the chosen provider)
  • Small capable teams using AI coding tools

Not yet a fit

  • —Need a hosted SaaS (Statewave is self-hosted infrastructure)
  • —Just need a vector database (use pgvector/Pinecone directly)
  • —Building chatbots with no multi-session requirement
  • —Need verified high-throughput scale today (multi-replica API is supported, but not load-tested beyond 10k subjects; single Postgres, no cross-region clustering)
  • —Looking for a complete agent framework

FAQ

Frequently asked questions

Why does an agent need a memory runtime instead of a bigger context window?

A context window is per-call state — re-sent, re-priced, and discarded every turn. A memory runtime keeps durable facts outside the prompt and assembles only what the current task needs, packed to a token budget rather than truncated at the end. Statewave ranks candidates on four fixed signals: kind priority (profile_fact 10, procedure 8, episode_summary 5, raw_episode 3), recency, task relevance, and temporal validity (+3 while valid, −4 once expired). The same subject, task, and budget return the same bundle every run.

What does provenance-first mean in practice?

Every compiled memory carries the IDs of the episodes it was derived from, a confidence score, and a validity window, so any fact an agent used traces back to the raw event that produced it. Context assembly can also emit a state-assembly receipt: an immutable, ULID-addressable record carrying a SHA-256 hash of the exact bytes handed to the model, plus the content hash of the policy bundle in force. "Why did the agent say that, and under which policy?" stays answerable months later.

Does anything leave my infrastructure?

Storage does not: episodes, compiled memories, and their embeddings live in your own Postgres with pgvector, and ranking is a single local query. What leaves depends on two choices you make. The heuristic compiler runs entirely on your network. An LLM compiler routes through LiteLLM, which supports 100+ providers including self-hosted Ollama and vLLM, so fully local is a configuration you can verify rather than a promise you have to trust.

Who is Statewave not a good fit for today?

Teams who want a hosted SaaS — Statewave is self-hosted infrastructure with no managed cloud. Teams who only need nearest-neighbor search, where pgvector or a vector database on its own is simpler. Chatbots with no multi-session requirement, which have nothing to remember. And workloads needing verified high-throughput scale today: the multi-replica API is supported but has not been load-tested beyond 10,000 subjects, on a single Postgres with no cross-region clustering.

Answers last checked against the Statewave docs and repositories on .