Skip to content

Synapse

Synapse is the engine behind every call: it masks personal data, picks the model by task, meters usage in questions and spreads work over machines.

Synapse is the layer every Subsidia call passes through, whether it arrives on the Gateway, through an agent, or from one of the product modules. It makes four decisions on your behalf, and a fifth thing, the certificate, records them:

  1. Is there personal data in this request, and may it leave? (the PII shield)
  2. Which provider and model should answer? (routing by task and sensitivity)
  3. What does this cost the workspace? (metering in questions)
  4. Which machine runs it? (multi-node scheduling, for installations with several)

As an API consumer you mostly never see these decisions, which is the point. This page explains how they work so that you can predict them, steer them with the model address, and read the few places where they are exposed.

  1. Request
  2. Address parsing
  3. PII scan + mask
  4. Route: provider + model
  5. Machine (node)
  6. Answer
  7. Restore values
  8. Meter + sign
The path of one call. Masking happens before routing is applied to the text, restoration after the model answers.

Routing

Routing turns two inputs, a task and a sensitivity, into a provider and a model. You do not pick the model by name: whatever you put in model is recorded but not obeyed (apart from the address prefixes and flags, which change the task and the sensitivity).

How the task is decided

The task is resolved in this order. The first rule that applies wins.

  1. A sensitive call (see below) goes to the local lane.
  2. An agent task goes to the full-quality model. A call is agentic when it uses pulse-agent or +agent, when the request defines tools, when the system prompt is longer than 8 000 characters (agent harnesses ship very long operating instructions), or when tool calls and tool results already appear in the conversation. Small models derail tool loops, so these never go to a small one.
  3. A simple or extraction task goes to the fast tier.
  4. A complex or code task goes to the full tier.
  5. Otherwise Synapse auto-classifies: JSON mode (response_format: json_object) counts as complex, an estimated prompt of 500 tokens or more counts as complex, anything shorter is simple. The estimate is the character count divided by four, the same heuristic POST /v1/messages/count_tokens returns.

The call classifier is deliberately cheap and deterministic. There is no extra model call on the Gateway path, so routing adds no latency and no cost.

TierChosen forModel
FastShort, simple asks and extraction.The cheapest fast model among the providers allowed for the workspace. If no distinct fast model is configured, the full model is used.
FullComplex, code and agent work, JSON mode, long prompts.The installation's configured full-quality model, or the best of the allowed providers when an administrator has enabled several.
LocalSensitive calls and the local preference.The locally hosted model (Ollama). Nothing leaves the machine.
You sendResolved taskTier
pulse-auto, 20-word questionsimpleFast
pulse-auto, 3 000-word contractcomplexFull
pulse-auto, response_format: json_objectcomplexFull
Any model, request with toolsagentFull
pulse-agentagentFull
pulse-sensitive, or +sensitivesensitive (on an installation with a local lane)Local
gpt-4o or any unknown nameAs pulse-autoAs pulse-auto

Providers and administrator policy

An installation can be configured with several providers (for example a European hosted provider, a US provider and a local model). Two rules bound what Synapse may choose:

  • Workspace allow-list. An administrator can restrict which providers the workspace may use. Synapse never picks outside it, and never silently replaces a refused choice with an external provider: a request whose preference cannot be honoured is refused.
  • The local lane is never restricted. A model that runs on the installation's own machines cannot leak anything, so it is always available.

Until an administrator writes an allow-list, routing stays on the installation's active provider. The x-pulse-provider header and the certificate tell you which one answered.

Sensitive calls

A call is treated as sensitive in two ways:

  • Explicitly, with pulse-sensitive or the +sensitive flag. This always wins.
  • Automatically, when the shield finds personal data in the request. This auto-escalation is skipped for agentic traffic (tool loops and harness prompts are full of incidental identifiers such as emails in code, and would otherwise land on a small local model every turn).

What "sensitive" does depends on the installation:

Installation Sensitive call
Local or on-premise (APP_MODE=local) Served by the local model, whatever provider preference was configured. The text reaches the model unmasked, since nothing leaves the machine.
Hosted cloud without a local lane There is no local machine to route to. The call goes through normal routing, but the shield masks personal data before any external provider sees it.

The certificate records the outcome: egress.provider is ollama when the call stayed local, and routing.taskType is sensitive.

PII shield

The shield finds personal values in the system prompt and every message, replaces each with a token before the text is sent, and swaps the originals back into the answer before you see it. The model reasons over [EMAIL_k3j2h], you read jean.dupont@example.com.

It only masks what leaves

Masking exists to protect data from third parties, so Synapse applies it only when the request can reach an external provider. A request that can only land on a local model is sent as is: full fidelity, no tokens, no restoration step. That is why local answers are often better on names and figures.

  • The shield is a no-op for local providers and when the mode is off.
  • When a call routes to a local model after the scan, the original text is sent, even if PII was found.
  • The count of masked values is reported as pii_masked, and is 0 for a local call.

Modes

The mode is a server setting chosen by the installation administrator (it can be changed from the admin interface, with a deployment-level default). It is not a per-request parameter.

ModeWhat it masksUse it for
standard (default)Values with a rigid syntactic shape: EMAIL, PHONE (and fax), CREDIT_CARD (Luhn-validated), IBAN, SSN, IP_ADDRESS, MAC_ADDRESS. Free text such as names and dates is left untouched, so ordinary prose is not mangled.Accounting, legal and general business traffic.
medicalEverything in standard, plus HIPAA Safe Harbor identifiers: title-anchored and bare person names (NAME), DATE (the year is kept), AGE for ages 90 and over, ZIP, URL, record and account numbers (MRN), national health identifiers (NIR for France, NHS_NUMBER for the UK), VIN. Clinical content itself (diagnoses, medications) is preserved so the model can still reason about the case.Health data.
offNothing.Installations that only use local models, or tests.

Consistent pseudonymization

Tokens are deterministic per entity. The suffix is a hash of the normalised value (accents, case and leading titles removed), so:

  • The same entity gets the same token everywhere: in every turn of a conversation, across separate requests, and across documents. "Dr Jean Dupuis" and "jean dupuis" become the same alias.
  • Distinct entities stay distinct. Two different people never collapse to one token, which would let a model confuse them.
  • Restoration works inside a single call from a vault kept for that call only.

This is why a model can follow "the second email from [EMAIL_k3j2h]" across a long conversation without ever seeing the address.

What the provider receives (illustration)
Reply to jean.dupont@example.com and confirm the transfer to FR7630006000011234567890189.
Call +33 6 12 34 56 78 if there is any issue.

The exact token suffixes depend on the value; do not parse them. The categories that were masked, never the values, are written to the certificate's egress.pii.categories.

What the API tells you

Routing decisions are exposed in three places, from lightest to most complete. Everything below is also true of the Anthropic dialect except where noted.

WhereFieldMeaning
Response header (non-streamed)x-pulse-providerThe provider that answered: ollama, openai, anthropic, mistral...
Response header (non-streamed)x-pulse-pii-maskedNumber of values masked in the request. 0 for a local call.
Response headerx-pulse-proofThe certificate id (also on streamed responses).
Body, OpenAI dialectmodelOn a non-streamed response, the model that actually answered, not the string you sent. On streamed chunks, the string you sent.
Body, OpenAI dialectpulse.provider, pulse.task, pulse.pii_maskedProvider, resolved task type (simple, complex, agent, sensitive...) and masked count. Non-streamed responses.
Body, Anthropic dialectmodelThe model that answered. The pulse block does not exist in this dialect; use the headers.
Certificaterouting.taskType, routing.reasonThe resolved task and a sentence such as task=complex -> configured model (openai/gpt-4o) or sensitive=true -> local Ollama (data never leaves machine).
Certificateegress.provider, egress.model, egress.piiWho answered and what was masked.
Reading the routing decision from a call
curl -s -D - -o /dev/null https://api.subsidia.protypa.fr/v1/chat/completions \
-H "Authorization: Bearer $SUBSIDIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"pulse-auto","messages":[{"role":"user","content":"Hi"}]}' \
| grep -i "^x-pulse"

Metering and cost

Customers buy questions, and Synapse is where usage becomes questions. For an API consumer this comes down to four facts:

  • Preflight check. Before any model is called, the workspace allowance is checked. An exhausted allowance is refused up front (402 insufficient_quota on the OpenAI dialect, 429 rate_limit_error on the Anthropic dialect), so a refused call consumes nothing.
  • Charged on the model that answered. pulse-auto may route one call to a small model and the next to a large one. A larger model counts for more questions per answer, so routing a simple ask to the fast tier is also the cheapest outcome for you. Local models are the lightest class.
  • Charged after the fact. The deduction happens once the completion has been served; an accounting failure is logged on the server and never fails your call.
  • Local installations are not metered against a question allowance. They keep a usage ledger for the dashboard.

A key can also carry its own monthly ceiling (API_KEY_BUDGET_EXCEEDED, a 402). See Rate limits and quotas.

The usage object in every response (prompt_tokens, completion_tokens, total_tokens) is the real token count reported by the model, useful for your own cost tracking. It is not a bill.

Behind the scenes Synapse also estimates a dollar cost per call from the model's public price (local models cost zero) and compares it with what the configured full model would have cost. Those figures feed the operator dashboard of the installation (cost, saved cost, latency, time to first token); they are not returned by the Gateway.

Multiple machines

On an installation with more than one machine running local models, Synapse spreads calls across them. It is entirely transparent to API clients: there is nothing to configure on your side and no field to set.

  • One whole request per machine. A single request is never split across machines; sharding one model over a local network is slower than running it on one box.
  • Least in-flight first. A new request goes to the healthy machine with the fewest requests already running.
  • Failover. An unhealthy machine is skipped, and a request that fails on one machine is retried on another.
  • Discovery is automatic, admission is a decision. A new machine on the network appears as a candidate and joins routing once an administrator admits it (or a shared cluster key is configured).

The visible effect is capacity and resilience: more concurrent calls without queueing, and no outage when one box is down. The egress.provider in the certificate is still ollama.

Putting it together

  1. 1
    Start with pulse-auto

    It gives you masking when needed, a small model for small asks, and the full model for real work. This is the right default for almost everything.

  2. 2
    Declare intent with the address when it matters

    pulse-agent for tool loops, pulse-sensitive when the data must stay local, agent:<slug> when a configured agent should answer. See Model addressing.

  3. 3
    Observe

    Log x-pulse-proof, x-pulse-provider and x-pulse-pii-masked with each call. They cost nothing and answer most support questions.

  4. 4
    Keep the evidence

    Store the certificate with the record the answer fed. It states the provider, the masking and the routing reason in a form a third party can verify.

Frequently asked

Can I choose the model by name?

No. The model name you send is recorded, but Synapse chooses the model from the task and the installation's configuration. What you control is the intent: pulse-agent for the full model, pulse-sensitive for local, an agent address for an agent. This keeps routing, cost and the data-residency guarantee in one place instead of in every client.

Why is my answer slower or lower quality than usual?

Check pulse.task and x-pulse-provider. A short prompt resolves to the fast tier, and a call forced local may use a smaller model than the hosted one. Send pulse-agent to force the full model for that call, or add context so the call is classified as complex.

The shield masked something that was not personal, or missed something.

Standard mode is built on syntactic patterns and validators (a card number must pass the Luhn check), so it is conservative and does not detect free-text names. If you need names and dates masked, the installation must run in medical mode, or the call must be local. Report false positives with the proof id; the certificate shows the categories involved.

Does masking change my prompt in the certificate?

The certificate hashes your original request (request.sha256) and, separately, the masked text that left toward the provider (egress.sha256). Both are hashes, so neither shows content.

Are streamed responses masked and restored too?

Yes, in the same way, and the hash in the certificate covers the restored text you received. Streamed responses carry x-pulse-proof but not x-pulse-provider or x-pulse-pii-masked, because the headers are sent before routing completes; read the certificate instead.

Is Synapse a separate API?

No. Synapse is the engine inside the endpoints you already call. Its decisions surface in headers, response fields and certificates, as described above.

Related