Agent memory
What an agent remembers across conversations: how memories are written, ranked, deduplicated and labelled by provenance, and how to read, seed and delete them.
A conversation has a history (Sessions); an agent has a memory. Memories are short notes (a decision, a preference, a fact worth keeping) that survive from one session to the next and are recalled into the prompt of later turns. They are what lets an agent say "as you told me last week, the client closes on 30 June" in a brand-new session.
Memories are private to the agent that wrote them. They are not a knowledge base: documents belong in Knowledge sources, where they are retrieved, cited and signed into a proof. Memory is for what the agent learned from the conversation.
How memories are written
There are two ways in.
- The agent writes them itself. At the end of a text answer the agent may emit up to three memory blocks (a short key, a short value, an optional type). They are stripped from what you receive, then stored. The ones saved during a turn come back in
memoriesSavedin the chat response (Agent chat). This happens on every chat call unless memory is switched off. Structured JSON answers (responseFormat.type: "json") do not carry memory instructions. - You write them through the API.
POST /v1/agents/:id/memorystores a fact you want the agent to know, without waiting for a conversation to produce it. See the endpoints below for how this path differs.
Every agent-written memory is checked before it is stored. Nothing is refused, but the memory is labelled with where its content came from.
Provenance
A model that invents "revenue grew 50%" in turn three must not be able to recall it in turn nine as if it had read it in a document. So at write time Subsidia looks for the memory's checkable content (every figure, or enough of its content words) in the evidence that turn actually had, strongest source first:
| Provenance | Found in | Starting salience | How it is recalled |
|---|---|---|---|
knowledge | Passages retrieved from the knowledge base this turn | 0.65 | Labelled "sourced: knowledge base". |
data | Structured data injected through the context array | 0.65 | Labelled "sourced: workspace data". |
tool | A tool result or a colleague consultation of the turn | 0.65 | Labelled "sourced: tool result". |
stated | What the user said in the conversation | 0.50 | Labelled "said in conversation". |
unverified | Nowhere: the agent's own claim | 0.35 | Labelled UNVERIFIED, with an instruction never to present it as established. |
A figure is the part of a memory that can be wrong in a costly way, so when a memory contains numbers, every number must appear in the source for it to be credited to that source. Unverified memories rank lower and fade sooner, but they are kept: an agent's own judgment is worth remembering as long as it is not mistaken for a fact. The provenance is stored in the memory's metadata.provenance.
Deduplication
Models re-label the same fact on every turn (client_year_end, then fiscal_closing_date). Two protections keep the memory from filling with copies:
- A memory with the same key replaces the previous value.
- A memory with a different key but the same content is folded into the existing record: among the agent's 60 most recent memories, one whose embedding is more than 0.92 cosine-similar to the new value, or whose content words overlap by at least 80%, is treated as a twin. The longer value is kept, the new key is recorded as an alias, and the salience goes up slightly (repetition is evidence of importance, but a claim asserted without a source stays unverified).
These protections apply to memories the agent writes. Memories you create with POST /v1/agents/:id/memory are inserted as they are.
How memories are recalled
Before each turn, the agent's memories are ranked against the new message and the best twelve are placed in its prompt. The score is a weighted sum: 50% semantic similarity to the message, 20% recency (a 14-day half-life, counted from the last time the memory was recalled), 20% salience, and 10% congruence with the agent's current mood. Memories without an embedding, such as those created through the API, get a neutral semantic score: they can still be recalled, but they win less often against memories that match the question closely. If the embedding service is unavailable, the agent falls back to its most recent memories.
Global memories (a fact worth sharing with every agent of the workspace, flagged by the agent when it writes it) are added on top, up to eight. They only exist for agents with sharing enabled in the console; for other agents the flag is ignored and the memory stays private.
Turning memory off
Memory is on by default for chat. To call an agent without it being read or written, use the Gateway address flag +nomemory, for example agent:claire+nomemory as the model name of a chat completion (Model addressing). The agent then neither recalls nor stores anything for that call. This is the right setting for tests, evaluations and any caller that keeps its own conversation state.
The /v1/agents/:id/chat endpoints have no per-call memory switch: they always recall and may always write. Postes never read or write memory (Postes).
Endpoints
All three routes need the agents scope on restricted keys, and the two write routes are refused for readonly keys. See Agents.
GET/v1/agents/:id/memory
List the memories of an agent.
agentsReturns the agent's own memories (those not flagged global), as stored rows. Rows include internal fields such as the embedding vector, a list of floating-point numbers that you can ignore; drop it before logging. Useful to audit what an agent has learned, to find an UNVERIFIED claim before it spreads, or to export memory.
Path parameters
idstringrequiredThe agent id.
Request examples
curl https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/memory \ -H "Authorization: Bearer $SUBSIDIA_API_KEY"Responses
An object with a memory array, in no guaranteed order. The list is not paginated.
{ "memory": [{ "id": "cm2kb7q4w0007qz0fm3n8c1xa", "agentId": "cm2k8x1ab0001qz0f7h3d9t4e", "workspaceId": null, "key": "client_fiscal_year_end", "value": "The client closes its books on 30 June.", "metadata": { "type": "fact", "provenance": "knowledge" }, "isGlobal": false, "type": "fact", "connectedTo": [], "valence": null, "salience": 0.65, "embedding": [0.0121, -0.0443, "... 768 numbers"], "lastRecalledAt": "2026-10-09T09:02:11.402Z", "recallCount": 3, "createdAt": "2026-10-02T14:20:05.771Z", "updatedAt": "2026-10-09T09:02:11.402Z"}] }Errors
- 404Unknown agent id, or the agent belongs to another workspace.
- 403
scope_deniedA restricted key without theagentsscope.
POST/v1/agents/:id/memory
Add a memory to an agent.
agentsStores a fact the agent should know from its next turn on. The memory is created exactly as sent: no provenance check, no deduplication against existing memories, and no embedding at creation time (so it is ranked with a neutral semantic score). It starts with the default salience of 0.5. If you call it twice with the same key you get two records.
Keep values short and factual, one idea per memory, like a note on a card. For documents, use Knowledge sources instead.
Path parameters
idstringrequiredThe agent id.
Request body
keystringrequiredShort label in snake_case, for exampleclient_fiscal_year_end.valuestringrequiredThe thing to remember, one or two sentences.
Request examples
curl https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/memory \ -H "Authorization: Bearer $SUBSIDIA_API_KEY" \ -H "Content-Type: application/json" \ -d '{"key": "client_fiscal_year_end", "value": "The client closes its books on 30 June."}'Responses
The stored memory.
{ "memory": { "id": "cm2kb7q4w0007qz0fm3n8c1xa", "agentId": "cm2k8x1ab0001qz0f7h3d9t4e", "workspaceId": null, "key": "client_fiscal_year_end", "value": "The client closes its books on 30 June.", "metadata": null, "isGlobal": false, "type": null, "connectedTo": [], "valence": null, "salience": 0.5, "embedding": [], "lastRecalledAt": null, "recallCount": 0, "createdAt": "2026-10-02T14:20:05.771Z", "updatedAt": "2026-10-09T09:02:11.402Z"} }Errors
- 400
keyorvaluemissing, empty or not a string. - 404Unknown agent id.
- 403
read_only_keyThe key carries thereadonlyscope.
DELETE/v1/agents/:id/memory/:memoryId
Delete one memory.
agentsRemoves a single memory. Use it to retract a wrong or UNVERIFIED claim. The memory must belong to the agent in the path.
Path parameters
idstringrequiredThe agent id.memoryIdstringrequiredTheidof the memory row, from the list endpoint.
Request examples
curl -X DELETE https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/memory/$MEMORY_ID \ -H "Authorization: Bearer $SUBSIDIA_API_KEY"Responses
Deleted. No body.
Errors
- 404The agent or the memory does not exist, or the memory belongs to another agent.
- 403
read_only_keyThe key carries thereadonlyscope.
Frequently asked
Can an agent remember something a user told it in confidence?
Memory is scoped to the agent and workspace and is never shown to other workspaces. But anything an agent stores is recalled into later prompts, for any user of that agent. Do not use a shared agent as a private notebook, and delete a memory that should not be kept.
Why did my memory not show up in answers?
Only the twelve best-ranked memories enter each prompt. A memory created through the API has no embedding, so it competes on recency and salience alone. Phrase it the way a user will ask about it, or put the document in a knowledge source if it must be found reliably.
Does memory count as a question?
No. Reading and writing memory is part of the chat turn that triggers it; the CRUD routes above call no model and consume nothing.