Skip to content

Your model forgets everythingOne header fixes it

A drop-in, OpenAI-compatible HTTP proxy that gives OpenRouter calls persistent memory. Point your existing client at it, add one header, and every request arrives with the context of the ones before it.

36 unit testsPython 3.11–3.133 memory-aware endpointsApache 2.0
POST/v1/chat/completions2001.2s
prompt

Has this customer hit this before?

no subject header

I have no previous context for this customer.

X-Statewave-Subject: acct:88413 episodes · 412 tokens

Twice in the last 30 days, both SSO timeouts following the 14:00 deploy.

One header, same model, same call.

Why route through Statewave

One header, zero rewrite

Change the base URL, add one header. Everything else in your integration stays exactly as it is.

Context that survives the call

Every request arrives with the memory of the ones before it, not just the current turn.

Fails open, not closed

If Statewave is unreachable, the call still goes through — just without memory for that turn.

Self-hosted, Apache 2.0

Runs next to your own infrastructure. No managed service, no vendor lock-in.

A base URL and one header

Two lines, with header trust enabled on the proxy; in JWT mode, send X-Statewave-Token instead. Everything else in your integration stays exactly as it is. A request with no subject gets no memory, so memory is opt-in per request rather than a global mode.

2 lines changed
client = OpenAI(base_url="http://localhost:8080/v1", api_key="sk-or-...") client.chat.completions.create(    model="openai/gpt-4o",    messages=[{"role": "user", "content": "What coffee do I like?"}],    extra_headers={"X-Statewave-Subject": "user:42"},)

What just happened

  • A memory bundle for user:42 was injected as a system message.
  • The call went to OpenRouter unchanged otherwise.
  • The turn was written back as an episode after the response was sent.
beforeOpenAI Client→OpenRouter
afterOpenAI Client→Statewave Proxy→OpenRouter

Memory in, completion out, episode written after

01
Client calls the proxy
02
Context fetched
03
Forwarded to OpenRouter
04
Reply relayed back
async · off the critical path
Episode written back
One path through the proxy. The write happens after the reply, not before it.

What the proxy touches

  • Adds the memory bundle, reads the reply text back out.
  • Strips Statewave headers and the statewave_subject body field before the call goes upstream.
  • Leaves model, temperature, tools and every other parameter untouched.

What it leaves alone

  • No subject: no context fetched, no episode written.
  • Non-completion paths such as /v1/models forward as-is.
  • Your OpenRouter key stays yours; the proxy forwards it, never replaces it.
No added write latency.

The episode write is fire-and-forget. It happens after the reply is already on its way to the client.

Streaming is not a special case.

SSE chunks relay byte-for-byte as they arrive. The reply is reassembled line by line, so a long stream costs the reply text, not a second copy of the body. The episode is written when the stream closes.

Empty replies write nothing.

A turn with no answer in it is noise in the subject's memory, not history.

Where the bundle goes, per endpoint

EndpointBundle goesReply read from
POST /v1/chat/completionsa system message, first in messageschoices[].message.content
POST /v1/completionsahead of promptchoices[].text
POST /v1/responsesahead of instructions; input untouchedoutput[].content[].text

Every other path (/v1/models, /v1/credits, the rest) is proxied straight through, so this is a drop-in base URL replacement.

Who the memory belongs to

A subject is the unit of memory. A session narrows it to one run of turns.

subject: user:42
turn 1
oat lattes
no session
turn 2
no sugar
session: sess_abc
turn 3
decaf after 4pm
session: sess_abc
compiled memory bundle for user:42
A session scopes a run of turns inside a subject; the subject keeps the memory.
subject
Who the memory belongs to: user:42, team:acme. This is the unit memory accumulates against.
session
Optional. Scopes a run of turns inside a subject.
set via
The X-Statewave-Subject header, or a statewave_subject body field for clients that cannot set headers. The header wins, and body fields are stripped before the request reaches OpenRouter.
id format
1–256 characters of letters, digits, underscore, dot, dash or colon. Anything else is rejected with 400 before any upstream call.
subject and session, over plain HTTP
curl http://localhost:8080/v1/chat/completions \
  -H "X-Statewave-Subject: user:42" \
  -H "X-Statewave-Session: sess_abc" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-4o","messages":[{"role":"user","content":"What coffee do I like?"}]}'

One decision to get right before you deploy

Whoever can name a subject can read and write that subject's memory.

So the proxy will not take a subject id on faith. Until you tell it which clients are trustworthy, a request that names one is refused outright rather than quietly reading somebody else's memory.

How the proxy decides whether to trust a subjectA request naming a subject reaches the trust gate. With nothing configured it is rejected with 400. With the trust flag set the header is taken as sent. With a JWT secret set the subject is taken from the token's sub claim.
request
names a subject
trust gate
400 rejected
nothing configured
header as sent
trusted network
sub claim
signed token
One gate, three outcomes. Pick a setting below to follow its branch.

An enhancement, never a hard dependency

If context assembly or the episode write fails, whether the server is down, the key is wrong, or the call times out, it is logged and the completion still goes through, just without memory for that turn. A Statewave outage degrades your app's memory quality. It does not take it down.

Fails-open comparisonOne request forks at the context step. When Statewave is healthy the bundle is assembled; when it is unreachable the failure is logged rather than raised. Both paths rejoin and return 200 OK, and only the memory in the reply differs.
request
context ok
bundle assembled
context failed
logged, not raised
completion
200 OK
One step differs. Hover a branch to trace it — both reply 200 OK, only the memory in the reply changes.

On shutdown, in-flight episode writes are drained before the HTTP client closes, since that is the only place a turn exists before Statewave has it.

Context read fails

The call still goes out

No bundle is injected, the error is logged, and the completion goes through without memory. The turn is still written back afterwards. The client sees a normal reply with no memory in it.

Episode write fails

The client never notices

The write happens after the reply has been sent, so a failure there cannot affect the response. That turn is missing from the subject’s history and the next one carries on from what is stored.

OpenRouter fails

The upstream error reaches you

Upstream status codes and error bodies are relayed rather than rewritten, so your existing error handling keeps working. Nothing is written to memory for a turn that produced no answer.

The settings that decide how it behaves

Everything is read from the environment, so the same image runs on a laptop and behind a gateway with no code change. OPENROUTER_API_KEY (unless clients send their own key), STATEWAVE_URL and one trust setting (STATEWAVE_TRUST_CLIENT_SUBJECT or PROXY_JWT_SECRET) are required for memory to work at all; the rest decide who is allowed to name a subject and which tenant it's written to.

VariableWhat it doesWhen
OPENROUTER_API_KEYThe key used for upstream calls when the client does not send its own.When clients do not send their own key
STATEWAVE_URLWhere the memory runtime lives. Defaults to http://localhost:8000.Always
STATEWAVE_API_KEYThe key Statewave itself requires, when the server enforces one. The proxy fails open, so a missing key means a normal reply with no memory and no error.Whenever the Statewave server requires an API key
STATEWAVE_TRUST_CLIENT_SUBJECTTakes the subject header at face value.Local, private network, or behind an authenticating gateway
PROXY_JWT_SECRETRequires a signed token on every route but /health, and takes the subject from its sub claim.Anything reachable by clients you do not control
STATEWAVE_CALLER_ID / STATEWAVE_CALLER_TYPEIdentifies the caller to Statewave.Optional; overrides the default caller openrouter-gateway
STATEWAVE_TENANT_IDPins which tenant the turn is written to. Without it, JWT mode takes the tenant from the token's tenant claim; otherwise the caller's X-Tenant-ID header decides.Any multi-tenant deployment

The shipped .env.example lists the remaining optional settings with their defaults. Note that the process does not read .env on its own, so pass --env-file .env when you start it.

Running in five commands

git clone https://github.com/smaramwbc/statewave-openrouter
cd statewave-openrouter
python3 -m venv .venv && . .venv/bin/activate && pip install .
cp .env.example .env     # set OPENROUTER_API_KEY, STATEWAVE_URL and STATEWAVE_TRUST_CLIENT_SUBJECT=1
uvicorn statewave_openrouter:app --env-file .env --port 8080
verify
curl http://localhost:8080/health
# {"status":"ok"}
client.ts
const client = new OpenAI({
  baseURL: "http://localhost:8080/v1",
  apiKey: process.env.OPENROUTER_API_KEY,
});

await client.chat.completions.create(
  { model: "openai/gpt-4o", messages },
  { headers: { "X-Statewave-Subject": "user:42" } },
);

Before you put it in front of users

  • Decide how subjects are trusted: a trusted header on a private network, signed tokens for anything public.
  • Pick subject ids stable for the user's lifetime, not per install or device.
  • Point your health check at /health; context and episode failures never surface as a request error.
  • Pin STATEWAVE_TENANT_ID before any multi-tenant deployment. Without it, a caller's X-Tenant-ID header decides which tenant the turn is written to, unless the proxy runs in JWT mode, where the token's tenant claim decides.
  • Talking to Statewave directly? The Python and TypeScript SDKs cover that path.

FAQ

Frequently asked questions

What is statewave-openrouter?

An OpenAI-compatible HTTP proxy that gives OpenRouter calls persistent memory. It assembles a bundle for the subject before each call and writes the turn back as an episode after the reply.

How much do I have to change in my code?

Two lines: point your existing OpenAI client at the proxy base URL and add an X-Statewave-Subject header to the request, with header trust enabled on the proxy. In JWT mode, send X-Statewave-Token instead. Everything else in the integration stays the same, and a request with no subject (no header, and in JWT mode no sub claim) gets no memory.

Does the proxy add latency to completions?

The episode write is fire-and-forget, so it costs nothing. The context fetch is one blocking read ahead of the upstream call.

What happens if Statewave is down?

It fails open. The failure is logged and the completion still goes through, just without memory for that turn.

Does it work with streaming?

Yes. SSE chunks relay byte for byte and the episode is written once the stream closes.

Which endpoints support memory?

Three endpoints are memory-aware: /v1/chat/completions, where the bundle becomes the first system message; /v1/completions, where it is prefixed onto the prompt; and /v1/responses, where it is prepended to top-level instructions. Every other path is proxied verbatim.

Do I need the Statewave SDK to use the proxy?

No. The proxy speaks the OpenAI HTTP API, so any OpenAI-compatible client in any language works: the Python SDK, the TypeScript SDK, plain curl, or an HTTP library you already use. The SDKs are for talking to Statewave directly, not for going through the proxy.

Can one proxy serve several subjects at once?

Yes. The subject is read per request, and two subjects never share a bundle.

What happens to the statewave_subject body field?

It is read by the proxy and then stripped from the payload before the request is forwarded, so OpenRouter never sees a field it does not recognise. If both the header and the body field are present, the header wins.

Is it tied to OpenRouter models only?

The upstream is OpenRouter, so any model OpenRouter routes to is available, and the model string passes through untouched. Switching models is a change to your request, not to the proxy.

How do I run it in production?

Build the container or install the package from the repository and run it under uvicorn behind whatever ingress you already use, set PROXY_JWT_SECRET if untrusted clients can reach it, and point your load balancer's health check at /health. Shutdown drains in-flight episode writes before the HTTP client closes.

Answers last checked against the Statewave docs and repositories on .

Give OpenRouter calls persistent memory

Self-host the Apache 2.0 proxy, point your existing client at it, and every call ships with the context of the ones before it.