Synapse
Synapse is the engine behind every call: it masks personal data, picks the model by task, meters usage in questions and spreads work over machines.
Synapse is the layer every Subsidia call passes through, whether it arrives on the Gateway, through an agent, or from one of the product modules. It makes four decisions on your behalf, and a fifth thing, the certificate, records them:
- Is there personal data in this request, and may it leave? (the PII shield)
- Which provider and model should answer? (routing by task and sensitivity)
- What does this cost the workspace? (metering in questions)
- Which machine runs it? (multi-node scheduling, for installations with several)
As an API consumer you mostly never see these decisions, which is the point. This page explains how they work so that you can predict them, steer them with the model address, and read the few places where they are exposed.
- Request
- Address parsing
- PII scan + mask
- Route: provider + model
- Machine (node)
- Answer
- Restore values
- Meter + sign
Routing
Routing turns two inputs, a task and a sensitivity, into a provider and a model. You do not pick the model by name: whatever you put in model is recorded but not obeyed (apart from the address prefixes and flags, which change the task and the sensitivity).
How the task is decided
The task is resolved in this order. The first rule that applies wins.
- A sensitive call (see below) goes to the local lane.
- An agent task goes to the full-quality model. A call is agentic when it uses
pulse-agentor+agent, when the request definestools, when the system prompt is longer than 8 000 characters (agent harnesses ship very long operating instructions), or when tool calls and tool results already appear in the conversation. Small models derail tool loops, so these never go to a small one. - A simple or extraction task goes to the fast tier.
- A complex or code task goes to the full tier.
- Otherwise Synapse auto-classifies: JSON mode (
response_format: json_object) counts as complex, an estimated prompt of 500 tokens or more counts as complex, anything shorter is simple. The estimate is the character count divided by four, the same heuristicPOST /v1/messages/count_tokensreturns.
The call classifier is deliberately cheap and deterministic. There is no extra model call on the Gateway path, so routing adds no latency and no cost.
| Tier | Chosen for | Model |
|---|---|---|
| Fast | Short, simple asks and extraction. | The cheapest fast model among the providers allowed for the workspace. If no distinct fast model is configured, the full model is used. |
| Full | Complex, code and agent work, JSON mode, long prompts. | The installation's configured full-quality model, or the best of the allowed providers when an administrator has enabled several. |
| Local | Sensitive calls and the local preference. | The locally hosted model (Ollama). Nothing leaves the machine. |
| You send | Resolved task | Tier |
|---|---|---|
pulse-auto, 20-word question | simple | Fast |
pulse-auto, 3 000-word contract | complex | Full |
pulse-auto, response_format: json_object | complex | Full |
Any model, request with tools | agent | Full |
pulse-agent | agent | Full |
pulse-sensitive, or +sensitive | sensitive (on an installation with a local lane) | Local |
gpt-4o or any unknown name | As pulse-auto | As pulse-auto |
Providers and administrator policy
An installation can be configured with several providers (for example a European hosted provider, a US provider and a local model). Two rules bound what Synapse may choose:
- Workspace allow-list. An administrator can restrict which providers the workspace may use. Synapse never picks outside it, and never silently replaces a refused choice with an external provider: a request whose preference cannot be honoured is refused.
- The local lane is never restricted. A model that runs on the installation's own machines cannot leak anything, so it is always available.
Until an administrator writes an allow-list, routing stays on the installation's active provider. The x-pulse-provider header and the certificate tell you which one answered.
Sensitive calls
A call is treated as sensitive in two ways:
- Explicitly, with
pulse-sensitiveor the+sensitiveflag. This always wins. - Automatically, when the shield finds personal data in the request. This auto-escalation is skipped for agentic traffic (tool loops and harness prompts are full of incidental identifiers such as emails in code, and would otherwise land on a small local model every turn).
What "sensitive" does depends on the installation:
| Installation | Sensitive call |
|---|---|
Local or on-premise (APP_MODE=local) |
Served by the local model, whatever provider preference was configured. The text reaches the model unmasked, since nothing leaves the machine. |
| Hosted cloud without a local lane | There is no local machine to route to. The call goes through normal routing, but the shield masks personal data before any external provider sees it. |
The certificate records the outcome: egress.provider is ollama when the call stayed local, and routing.taskType is sensitive.
PII shield
The shield finds personal values in the system prompt and every message, replaces each with a token before the text is sent, and swaps the originals back into the answer before you see it. The model reasons over [EMAIL_k3j2h], you read jean.dupont@example.com.
It only masks what leaves
Masking exists to protect data from third parties, so Synapse applies it only when the request can reach an external provider. A request that can only land on a local model is sent as is: full fidelity, no tokens, no restoration step. That is why local answers are often better on names and figures.
- The shield is a no-op for local providers and when the mode is
off. - When a call routes to a local model after the scan, the original text is sent, even if PII was found.
- The count of masked values is reported as
pii_masked, and is0for a local call.
Modes
The mode is a server setting chosen by the installation administrator (it can be changed from the admin interface, with a deployment-level default). It is not a per-request parameter.
| Mode | What it masks | Use it for |
|---|---|---|
standard (default) | Values with a rigid syntactic shape: EMAIL, PHONE (and fax), CREDIT_CARD (Luhn-validated), IBAN, SSN, IP_ADDRESS, MAC_ADDRESS. Free text such as names and dates is left untouched, so ordinary prose is not mangled. | Accounting, legal and general business traffic. |
medical | Everything in standard, plus HIPAA Safe Harbor identifiers: title-anchored and bare person names (NAME), DATE (the year is kept), AGE for ages 90 and over, ZIP, URL, record and account numbers (MRN), national health identifiers (NIR for France, NHS_NUMBER for the UK), VIN. Clinical content itself (diagnoses, medications) is preserved so the model can still reason about the case. | Health data. |
off | Nothing. | Installations that only use local models, or tests. |
Consistent pseudonymization
Tokens are deterministic per entity. The suffix is a hash of the normalised value (accents, case and leading titles removed), so:
- The same entity gets the same token everywhere: in every turn of a conversation, across separate requests, and across documents. "Dr Jean Dupuis" and "jean dupuis" become the same alias.
- Distinct entities stay distinct. Two different people never collapse to one token, which would let a model confuse them.
- Restoration works inside a single call from a vault kept for that call only.
This is why a model can follow "the second email from [EMAIL_k3j2h]" across a long conversation without ever seeing the address.
Reply to jean.dupont@example.com and confirm the transfer to FR7630006000011234567890189.Call +33 6 12 34 56 78 if there is any issue.The exact token suffixes depend on the value; do not parse them. The categories that were masked, never the values, are written to the certificate's egress.pii.categories.
What the API tells you
Routing decisions are exposed in three places, from lightest to most complete. Everything below is also true of the Anthropic dialect except where noted.
| Where | Field | Meaning |
|---|---|---|
| Response header (non-streamed) | x-pulse-provider | The provider that answered: ollama, openai, anthropic, mistral... |
| Response header (non-streamed) | x-pulse-pii-masked | Number of values masked in the request. 0 for a local call. |
| Response header | x-pulse-proof | The certificate id (also on streamed responses). |
| Body, OpenAI dialect | model | On a non-streamed response, the model that actually answered, not the string you sent. On streamed chunks, the string you sent. |
| Body, OpenAI dialect | pulse.provider, pulse.task, pulse.pii_masked | Provider, resolved task type (simple, complex, agent, sensitive...) and masked count. Non-streamed responses. |
| Body, Anthropic dialect | model | The model that answered. The pulse block does not exist in this dialect; use the headers. |
| Certificate | routing.taskType, routing.reason | The resolved task and a sentence such as task=complex -> configured model (openai/gpt-4o) or sensitive=true -> local Ollama (data never leaves machine). |
| Certificate | egress.provider, egress.model, egress.pii | Who answered and what was masked. |
curl -s -D - -o /dev/null https://api.subsidia.protypa.fr/v1/chat/completions \ -H "Authorization: Bearer $SUBSIDIA_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"pulse-auto","messages":[{"role":"user","content":"Hi"}]}' \ | grep -i "^x-pulse"Metering and cost
Customers buy questions, and Synapse is where usage becomes questions. For an API consumer this comes down to four facts:
- Preflight check. Before any model is called, the workspace allowance is checked. An exhausted allowance is refused up front (
402 insufficient_quotaon the OpenAI dialect,429 rate_limit_erroron the Anthropic dialect), so a refused call consumes nothing. - Charged on the model that answered.
pulse-automay route one call to a small model and the next to a large one. A larger model counts for more questions per answer, so routing a simple ask to the fast tier is also the cheapest outcome for you. Local models are the lightest class. - Charged after the fact. The deduction happens once the completion has been served; an accounting failure is logged on the server and never fails your call.
- Local installations are not metered against a question allowance. They keep a usage ledger for the dashboard.
A key can also carry its own monthly ceiling (API_KEY_BUDGET_EXCEEDED, a 402). See Rate limits and quotas.
The usage object in every response (prompt_tokens, completion_tokens, total_tokens) is the real token count reported by the model, useful for your own cost tracking. It is not a bill.
Behind the scenes Synapse also estimates a dollar cost per call from the model's public price (local models cost zero) and compares it with what the configured full model would have cost. Those figures feed the operator dashboard of the installation (cost, saved cost, latency, time to first token); they are not returned by the Gateway.
Multiple machines
On an installation with more than one machine running local models, Synapse spreads calls across them. It is entirely transparent to API clients: there is nothing to configure on your side and no field to set.
- One whole request per machine. A single request is never split across machines; sharding one model over a local network is slower than running it on one box.
- Least in-flight first. A new request goes to the healthy machine with the fewest requests already running.
- Failover. An unhealthy machine is skipped, and a request that fails on one machine is retried on another.
- Discovery is automatic, admission is a decision. A new machine on the network appears as a candidate and joins routing once an administrator admits it (or a shared cluster key is configured).
The visible effect is capacity and resilience: more concurrent calls without queueing, and no outage when one box is down. The egress.provider in the certificate is still ollama.
Putting it together
- 1Start with pulse-auto
It gives you masking when needed, a small model for small asks, and the full model for real work. This is the right default for almost everything.
- 2Declare intent with the address when it matters
pulse-agentfor tool loops,pulse-sensitivewhen the data must stay local,agent:<slug>when a configured agent should answer. See Model addressing. - 3Observe
Log
x-pulse-proof,x-pulse-providerandx-pulse-pii-maskedwith each call. They cost nothing and answer most support questions. - 4Keep the evidence
Store the certificate with the record the answer fed. It states the provider, the masking and the routing reason in a form a third party can verify.
Frequently asked
Can I choose the model by name?
No. The model name you send is recorded, but Synapse chooses the model from the task and the installation's configuration. What you control is the intent: pulse-agent for the full model, pulse-sensitive for local, an agent address for an agent. This keeps routing, cost and the data-residency guarantee in one place instead of in every client.
Why is my answer slower or lower quality than usual?
Check pulse.task and x-pulse-provider. A short prompt resolves to the fast tier, and a call forced local may use a smaller model than the hosted one. Send pulse-agent to force the full model for that call, or add context so the call is classified as complex.
The shield masked something that was not personal, or missed something.
Standard mode is built on syntactic patterns and validators (a card number must pass the Luhn check), so it is conservative and does not detect free-text names. If you need names and dates masked, the installation must run in medical mode, or the call must be local. Report false positives with the proof id; the certificate shows the categories involved.
Does masking change my prompt in the certificate?
The certificate hashes your original request (request.sha256) and, separately, the masked text that left toward the provider (egress.sha256). Both are hashes, so neither shows content.
Are streamed responses masked and restored too?
Yes, in the same way, and the hash in the certificate covers the restored text you received. Streamed responses carry x-pulse-proof but not x-pulse-provider or x-pulse-pii-masked, because the headers are sent before routing completes; read the certificate instead.
Is Synapse a separate API?
No. Synapse is the engine inside the endpoints you already call. Its decisions surface in headers, response fields and certificates, as described above.