# Synapse

> Synapse is the engine behind every call: it masks personal data, picks the model by task, meters usage in questions and spreads work over machines.

Synapse is the layer every Subsidia call passes through, whether it arrives on the [Gateway](https://dev.subsidia.protypa.fr/docs/gateway.md), through an [agent](https://dev.subsidia.protypa.fr/docs/agent-chat.md), or from one of the product modules. It makes four decisions on your behalf, and a fifth thing, the certificate, records them:

1. **Is there personal data in this request, and may it leave?** (the PII shield)
2. **Which provider and model should answer?** (routing by task and sensitivity)
3. **What does this cost the workspace?** (metering in questions)
4. **Which machine runs it?** (multi-node scheduling, for installations with several)

As an API consumer you mostly never see these decisions, which is the point. This page explains how they work so that you can predict them, steer them with the [model address](https://dev.subsidia.protypa.fr/docs/model-addressing.md), and read the few places where they are exposed.

**The path of one call. Masking happens before routing is applied to the text, restoration after the model answers.**

Flow: Request -> Address parsing -> PII scan + mask -> Route: provider + model -> Machine (node) -> Answer -> Restore values -> Meter + sign

## Routing

Routing turns two inputs, a **task** and a **sensitivity**, into a provider and a model. You do not pick the model by name: whatever you put in `model` is recorded but not obeyed (apart from the address prefixes and flags, which change the task and the sensitivity).

**How the task is decided**

The task is resolved in this order. The first rule that applies wins.

1. A **sensitive** call (see below) goes to the local lane.
2. An **agent** task goes to the full-quality model. A call is agentic when it uses `pulse-agent` or `+agent`, when the request defines `tools`, when the system prompt is longer than 8 000 characters (agent harnesses ship very long operating instructions), or when tool calls and tool results already appear in the conversation. Small models derail tool loops, so these never go to a small one.
3. A **simple** or **extraction** task goes to the fast tier.
4. A **complex** or **code** task goes to the full tier.
5. Otherwise Synapse **auto-classifies**: JSON mode (`response_format: json_object`) counts as complex, an estimated prompt of 500 tokens or more counts as complex, anything shorter is simple. The estimate is the character count divided by four, the same heuristic `POST /v1/messages/count_tokens` returns.

The call classifier is deliberately cheap and deterministic. There is no extra model call on the Gateway path, so routing adds no latency and no cost.

| Tier | Chosen for | Model |
| --- | --- | --- |
| Fast | Short, simple asks and extraction. | The cheapest fast model among the providers allowed for the workspace. If no distinct fast model is configured, the full model is used. |
| Full | Complex, code and agent work, JSON mode, long prompts. | The installation's configured full-quality model, or the best of the allowed providers when an administrator has enabled several. |
| Local | Sensitive calls and the `local` preference. | The locally hosted model (Ollama). Nothing leaves the machine. |

| You send | Resolved task | Tier |
| --- | --- | --- |
| `pulse-auto`, 20-word question | `simple` | Fast |
| `pulse-auto`, 3 000-word contract | `complex` | Full |
| `pulse-auto`, `response_format: json_object` | `complex` | Full |
| Any model, request with `tools` | `agent` | Full |
| `pulse-agent` | `agent` | Full |
| `pulse-sensitive`, or `+sensitive` | `sensitive` (on an installation with a local lane) | Local |
| `gpt-4o` or any unknown name | As `pulse-auto` | As `pulse-auto` |

**Providers and administrator policy**

An installation can be configured with several providers (for example a European hosted provider, a US provider and a local model). Two rules bound what Synapse may choose:

- **Workspace allow-list.** An administrator can restrict which providers the workspace may use. Synapse never picks outside it, and never silently replaces a refused choice with an external provider: a request whose preference cannot be honoured is refused.
- **The local lane is never restricted.** A model that runs on the installation's own machines cannot leak anything, so it is always available.

Until an administrator writes an allow-list, routing stays on the installation's active provider. The `x-pulse-provider` header and the certificate tell you which one answered.

### Sensitive calls

A call is treated as sensitive in two ways:

- **Explicitly**, with `pulse-sensitive` or the `+sensitive` flag. This always wins.
- **Automatically**, when the shield finds personal data in the request. This auto-escalation is skipped for agentic traffic (tool loops and harness prompts are full of incidental identifiers such as emails in code, and would otherwise land on a small local model every turn).

What "sensitive" does depends on the installation:

| Installation | Sensitive call |
|---|---|
| Local or on-premise (`APP_MODE=local`) | Served by the local model, whatever provider preference was configured. The text reaches the model **unmasked**, since nothing leaves the machine. |
| Hosted cloud without a local lane | There is no local machine to route to. The call goes through normal routing, but the shield masks personal data before any external provider sees it. |

The certificate records the outcome: `egress.provider` is `ollama` when the call stayed local, and `routing.taskType` is `sensitive`.

> **WARNING: Verify, do not assume, for hard data-residency requirements**
> If a contract says a category of data must never reach an external provider, check `x-pulse-provider` (or `payload.egress.provider` in the [certificate](https://dev.subsidia.protypa.fr/docs/proofs.md)) in your integration tests against the target installation, and fail the build when it is not local. The address expresses intent; the certificate records what happened.

## PII shield

The shield finds personal values in the system prompt and every message, replaces each with a token before the text is sent, and swaps the originals back into the answer before you see it. The model reasons over `[EMAIL_k3j2h]`, you read `jean.dupont@example.com`.

**It only masks what leaves**

Masking exists to protect data from third parties, so Synapse applies it **only when the request can reach an external provider**. A request that can only land on a local model is sent as is: full fidelity, no tokens, no restoration step. That is why local answers are often better on names and figures.

- The shield is a no-op for local providers and when the mode is `off`.
- When a call routes to a local model after the scan, the original text is sent, even if PII was found.
- The count of masked values is reported as `pii_masked`, and is `0` for a local call.

**Modes**

The mode is a server setting chosen by the installation administrator (it can be changed from the admin interface, with a deployment-level default). It is not a per-request parameter.

| Mode | What it masks | Use it for |
| --- | --- | --- |
| `standard` (default) | Values with a rigid syntactic shape: `EMAIL`, `PHONE` (and fax), `CREDIT_CARD` (Luhn-validated), `IBAN`, `SSN`, `IP_ADDRESS`, `MAC_ADDRESS`. Free text such as names and dates is left untouched, so ordinary prose is not mangled. | Accounting, legal and general business traffic. |
| `medical` | Everything in `standard`, plus HIPAA Safe Harbor identifiers: title-anchored and bare person names (`NAME`), `DATE` (the year is kept), `AGE` for ages 90 and over, `ZIP`, `URL`, record and account numbers (`MRN`), national health identifiers (`NIR` for France, `NHS_NUMBER` for the UK), `VIN`. Clinical content itself (diagnoses, medications) is preserved so the model can still reason about the case. | Health data. |
| `off` | Nothing. | Installations that only use local models, or tests. |

> **INFO: Medical mode is not the right default for professional firms**
> Medical mode strips dates and bare names, which removes exactly the facts an accounting or legal question depends on. Use `standard` unless the traffic is clinical.

### Consistent pseudonymization

Tokens are **deterministic per entity**. The suffix is a hash of the normalised value (accents, case and leading titles removed), so:

- The same entity gets the same token everywhere: in every turn of a conversation, across separate requests, and across documents. "Dr Jean Dupuis" and "jean dupuis" become the same alias.
- Distinct entities stay distinct. Two different people never collapse to one token, which would let a model confuse them.
- Restoration works inside a single call from a vault kept for that call only.

This is why a model can follow "the second email from `[EMAIL_k3j2h]`" across a long conversation without ever seeing the address.

**What the provider receives (illustration)**

_You send_

```text
Reply to jean.dupont@example.com and confirm the transfer to FR7630006000011234567890189.
Call +33 6 12 34 56 78 if there is any issue.
```

_External provider receives_

```text
Reply to [EMAIL_k3j2h] and confirm the transfer to [IBAN_9f2x1a].
Call [PHONE_1zq8d0] if there is any issue.
```

_You read in the answer_

```text
Dear Mr Dupont, the transfer to FR7630006000011234567890189 is confirmed.
(values restored before the answer reaches you)
```

The exact token suffixes depend on the value; do not parse them. The categories that were masked, never the values, are written to the certificate's `egress.pii.categories`.

## What the API tells you

Routing decisions are exposed in three places, from lightest to most complete. Everything below is also true of the Anthropic dialect except where noted.

| Where | Field | Meaning |
| --- | --- | --- |
| Response header (non-streamed) | `x-pulse-provider` | The provider that answered: `ollama`, `openai`, `anthropic`, `mistral`... |
| Response header (non-streamed) | `x-pulse-pii-masked` | Number of values masked in the request. `0` for a local call. |
| Response header | `x-pulse-proof` | The certificate id (also on streamed responses). |
| Body, OpenAI dialect | `model` | On a non-streamed response, the model that actually answered, not the string you sent. On streamed chunks, the string you sent. |
| Body, OpenAI dialect | `pulse.provider`, `pulse.task`, `pulse.pii_masked` | Provider, resolved task type (`simple`, `complex`, `agent`, `sensitive`...) and masked count. Non-streamed responses. |
| Body, Anthropic dialect | `model` | The model that answered. The `pulse` block does not exist in this dialect; use the headers. |
| Certificate | `routing.taskType`, `routing.reason` | The resolved task and a sentence such as `task=complex -> configured model (openai/gpt-4o)` or `sensitive=true -> local Ollama (data never leaves machine)`. |
| Certificate | `egress.provider`, `egress.model`, `egress.pii` | Who answered and what was masked. |

**Reading the routing decision from a call**

_curl_

```bash
curl -s -D - -o /dev/null https://api.subsidia.protypa.fr/v1/chat/completions \
  -H "Authorization: Bearer $SUBSIDIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"pulse-auto","messages":[{"role":"user","content":"Hi"}]}' \
  | grep -i "^x-pulse"
```

_TypeScript_

```typescript
const res = await fetch('https://api.subsidia.protypa.fr/v1/chat/completions', {
  method: 'POST',
  headers: {
    Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    model: 'pulse-auto',
    messages: [{ role: 'user', content: 'Hi' }],
  }),
})

console.log({
  provider: res.headers.get('x-pulse-provider'),
  masked: Number(res.headers.get('x-pulse-pii-masked')),
  proof: res.headers.get('x-pulse-proof'),
})

const body = await res.json()
console.log(body.model, body.pulse.task) // model that answered, resolved task
```

_Python_

```python
import os, requests

res = requests.post(
    "https://api.subsidia.protypa.fr/v1/chat/completions",
    headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
    json={"model": "pulse-auto", "messages": [{"role": "user", "content": "Hi"}]},
)

print(res.headers["x-pulse-provider"], res.headers["x-pulse-pii-masked"], res.headers["x-pulse-proof"])
body = res.json()
print(body["model"], body["pulse"]["task"])
```

> **TIP: Guard a residency rule in a test**
> In a CI job, send a representative prompt to `pulse-sensitive` and assert that `x-pulse-provider` is `ollama`, then fetch the certificate and assert the same on `payload.egress.provider`. The first catches a misconfigured installation, the second proves the evidence matches.

## Metering and cost

Customers buy **questions**, and Synapse is where usage becomes questions. For an API consumer this comes down to four facts:

- **Preflight check.** Before any model is called, the workspace allowance is checked. An exhausted allowance is refused up front (`402 insufficient_quota` on the OpenAI dialect, `429 rate_limit_error` on the Anthropic dialect), so a refused call consumes nothing.
- **Charged on the model that answered.** `pulse-auto` may route one call to a small model and the next to a large one. A larger model counts for more questions per answer, so routing a simple ask to the fast tier is also the cheapest outcome for you. Local models are the lightest class.
- **Charged after the fact.** The deduction happens once the completion has been served; an accounting failure is logged on the server and never fails your call.
- **Local installations are not metered** against a question allowance. They keep a usage ledger for the dashboard.

A key can also carry its own monthly ceiling (`API_KEY_BUDGET_EXCEEDED`, a 402). See [Rate limits and quotas](https://dev.subsidia.protypa.fr/docs/rate-limits.md).

The `usage` object in every response (`prompt_tokens`, `completion_tokens`, `total_tokens`) is the real token count reported by the model, useful for your own cost tracking. It is not a bill.

Behind the scenes Synapse also estimates a dollar cost per call from the model's public price (local models cost zero) and compares it with what the configured full model would have cost. Those figures feed the operator dashboard of the installation (cost, saved cost, latency, time to first token); they are not returned by the Gateway.

## Multiple machines

On an installation with more than one machine running local models, Synapse spreads calls across them. It is entirely transparent to API clients: there is nothing to configure on your side and no field to set.

- **One whole request per machine.** A single request is never split across machines; sharding one model over a local network is slower than running it on one box.
- **Least in-flight first.** A new request goes to the healthy machine with the fewest requests already running.
- **Failover.** An unhealthy machine is skipped, and a request that fails on one machine is retried on another.
- **Discovery is automatic, admission is a decision.** A new machine on the network appears as a candidate and joins routing once an administrator admits it (or a shared cluster key is configured).

The visible effect is capacity and resilience: more concurrent calls without queueing, and no outage when one box is down. The `egress.provider` in the certificate is still `ollama`.

## Putting it together

1. **Start with pulse-auto**

   It gives you masking when needed, a small model for small asks, and the full model for real work. This is the right default for almost everything.

2. **Declare intent with the address when it matters**

   `pulse-agent` for tool loops, `pulse-sensitive` when the data must stay local, `agent:<slug>` when a configured agent should answer. See [Model addressing](https://dev.subsidia.protypa.fr/docs/model-addressing.md).

3. **Observe**

   Log `x-pulse-proof`, `x-pulse-provider` and `x-pulse-pii-masked` with each call. They cost nothing and answer most support questions.

4. **Keep the evidence**

   Store the [certificate](https://dev.subsidia.protypa.fr/docs/proofs.md) with the record the answer fed. It states the provider, the masking and the routing reason in a form a third party can verify.

## Frequently asked

**Can I choose the model by name?**

No. The model name you send is recorded, but Synapse chooses the model from the task and the installation's configuration. What you control is the intent: `pulse-agent` for the full model, `pulse-sensitive` for local, an agent address for an agent. This keeps routing, cost and the data-residency guarantee in one place instead of in every client.

**Why is my answer slower or lower quality than usual?**

Check `pulse.task` and `x-pulse-provider`. A short prompt resolves to the fast tier, and a call forced local may use a smaller model than the hosted one. Send `pulse-agent` to force the full model for that call, or add context so the call is classified as complex.

**The shield masked something that was not personal, or missed something.**

Standard mode is built on syntactic patterns and validators (a card number must pass the Luhn check), so it is conservative and does not detect free-text names. If you need names and dates masked, the installation must run in medical mode, or the call must be local. Report false positives with the proof id; the certificate shows the categories involved.

**Does masking change my prompt in the certificate?**

The certificate hashes your original request (`request.sha256`) and, separately, the masked text that left toward the provider (`egress.sha256`). Both are hashes, so neither shows content.

**Are streamed responses masked and restored too?**

Yes, in the same way, and the hash in the certificate covers the restored text you received. Streamed responses carry `x-pulse-proof` but not `x-pulse-provider` or `x-pulse-pii-masked`, because the headers are sent before routing completes; read the certificate instead.

**Is Synapse a separate API?**

No. Synapse is the engine inside the endpoints you already call. Its decisions surface in headers, response fields and certificates, as described above.
