# Knowledge sources

> Ingest text, URLs and files (PDF, Office, email, images) into Cortex, track ingestion, and attach sources to agents.

A **knowledge source** is one document in your workspace: a file, a web page or a block of text. When you add one, Cortex parses it, cuts it into passages, embeds them, extracts atomic facts and prepares health-check questions. From then on its passages can be retrieved by [`/v1/knowledge/query`](https://dev.subsidia.protypa.fr/docs/knowledge-query.md), by agents linked to it, and by Gateway calls that carry the `+cortex` flag.

Ingestion is asynchronous. The API answers immediately with the source id and a status; you poll the source (or listen to a [webhook](https://dev.subsidia.protypa.fr/docs/webhooks.md)) until it is ready.

**Ingestion lifecycle. The `progressStage` field of a source walks through these stages while `status` is `processing`.**

Flow: Request (text, url or file) -> Duplicate check -> Parsing (+ OCR) -> Chunks -> Contextualizing -> Embedding -> Facts + contradictions -> Health QA -> ready

## Concepts

| Field | Values | Meaning |
| --- | --- | --- |
| `status` | `processing`, `ready`, `failed` | Lifecycle of the ingestion. Only `ready` sources are searchable. A failed source keeps the reason in `error`. |
| `progressStage` | `parsing`, `chunks`, `contextualizing`, `embedding`, `facts`, `qa`, `done` | Where a `processing` source currently is. `progressPercent` (0-100) and `etaSeconds` accompany it. |
| `sourceType` | `text`, `file`, `url` | How the source entered the workspace. |
| `collectionId` | string or null | The collection (a "client file" or dossier in the app) the source is filed under. Retrieval results echo it back as `collectionId` and `dossierName`. |

> **INFO: Identical content is ingested once**
> Every text or file is hashed (SHA-256 of the raw bytes). If the workspace already holds a non-failed source with the same hash, the API returns `status: "duplicate"` with the `id` of the existing source and `duplicateOf` set to its name, and nothing is re-ingested. A `failed` source does not block a retry. URL sources are exempt, because their content is only known after the fetch.

> **INFO: Ingestion consumes no questions**
> Ingestion is background work and does not count against your question allowance. What your plan limits is the **volume of documents**: when the workspace is full the request fails with `402` and `code: "DOCUMENT_LIMIT"`. See [Rate limits and quotas](https://dev.subsidia.protypa.fr/docs/rate-limits.md).

## Supported formats

| Format | Extensions | Notes |
| --- | --- | --- |
| PDF | `pdf` | Text layer first. A scanned PDF falls back to OCR (French and English by default, 60 pages by default on the server). Passages carry a `page` number when known. |
| Word | `docx` | Headings are kept as sections. |
| Spreadsheets | `xlsx`, `xlsm`, `csv`, `fec` | Tables are linearised into readable rows, capped at 5000 rows per sheet (a truncation is stated in the text). Legacy `.xls` is refused with a message asking for `.xlsx` or `.csv`. |
| Presentations | `pptx` | Slide text. |
| Email | `eml`, `msg` | Headers and body. |
| Images | `png`, `jpg`, `jpeg`, `webp`, `bmp` | Read by OCR. |
| Web and data | `html`, `htm`, `xml`, `json` | HTML is converted to Markdown. |
| Text | `md`, `markdown`, `txt`, `text` | Used as is. |

Any other type is rejected at parsing time: the source ends in `failed` with the message `Unsupported file type ...`. Passages are cut on the document structure (headings, paragraphs, table rows) and never exceed 2000 characters.

## Limits

| Limit | Value |
| --- | --- |
| File size (`/upload`) | 25 MB. Larger files return `413`. |
| Request rate | 30 requests per minute for `POST /v1/knowledge-sources` and `POST /v1/knowledge-sources/upload`; 429 beyond that. |
| URL ingestion | Public `http(s)` addresses only. Private, local and non-http addresses are refused with `400`. |
| Documents per workspace | Set by the plan. `402 DOCUMENT_LIMIT` when reached. |

## Endpoints

> **WARNING: Scopes**
> The source endpoints need the `knowledge` scope on a restricted key. Linking a source to an agent goes through `/v1/agents/...` and needs the `agents` scope. A missing scope is a `403 scope_denied`. See [Authentication](https://dev.subsidia.protypa.fr/docs/authentication.md).

### POST /v1/knowledge-sources

Ingest raw text or a public URL.

Creates a source from text you send or from a page Cortex fetches. Send **either** `content` **or** `url`. The response comes back at once with `status: "processing"`; poll `GET /v1/knowledge-sources/:id` until `ready`.

Use this route when you already hold the text (a CRM note, an export). For files, use the upload route: you do not need to extract the text yourself, and confidential documents never have to be published to a URL.

- **Authentication:** API key (Bearer or x-api-key)
- **Scopes:** `knowledge`

#### Request body

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `name` | `string` | no |  | Display name. Required with `content`; defaults to the URL for URL ingestion. |
| `description` | `string` | no |  | Free description. |
| `content` | `string` | no |  | Raw text to ingest. Must not be empty. |
| `url` | `string` | no |  | Public http(s) page or document to fetch and ingest. Alternative to `content`. |
| `agentIds` | `string[]` | no |  | Agent ids to link the source to at creation, so those agents can retrieve from it. |
| `collectionId` | `string` | no |  | Collection to file the source under. |
| `contextualize` | `boolean` | no | `true` | Generate a one-line context for each passage before embedding. Improves retrieval, costs ingestion time. Set `false` to skip. |
| `extractFacts` | `boolean` | no | `true` | Extract atomic facts and detect contradictions with other sources. See [Facts, contradictions and health](https://dev.subsidia.protypa.fr/docs/knowledge-insights.md). |
| `generateQa` | `boolean` | no | `true` | Generate question/answer pairs used by the health replay. |

#### Request examples

_curl_

```bash
curl https://api.subsidia.protypa.fr/v1/knowledge-sources \
  -H "Authorization: Bearer $SUBSIDIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Termination policy",
    "content": "Either party may terminate the agreement with 30 days written notice."
  }'
```

_TypeScript_

```typescript
const res = await fetch('https://api.subsidia.protypa.fr/v1/knowledge-sources', {
  method: 'POST',
  headers: {
    Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    name: 'Termination policy',
    content: 'Either party may terminate the agreement with 30 days written notice.',
  }),
})
const { id, status } = await res.json() // status: "processing"
```

_Python_

```python
import os, requests

res = requests.post(
    "https://api.subsidia.protypa.fr/v1/knowledge-sources",
    headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
    json={
        "name": "Termination policy",
        "content": "Either party may terminate the agreement with 30 days written notice.",
    },
)
source = res.json()  # {"id": "...", "status": "processing"}
```

#### Responses

**201**: The source was created and is being processed. A duplicate returns `status: "duplicate"` and `duplicateOf`.

```json
{ "id": "clx9k2m4p0001abcd", "status": "processing" }
```

**400**: Neither `content` nor `url`; `content` without `name`; a URL that is invalid, private or not http(s).

```json
{ "error": "content or url is required" }
```

#### Errors

| Status | Code | When |
| --- | --- | --- |
| 400 |  | Missing `content`/`url`, missing `name` with `content`, or an unsafe URL. |
| 401 |  | Missing or invalid key. |
| 402 | `DOCUMENT_LIMIT` | The workspace reached the document volume of its plan. |
| 403 | `scope_denied` | The key lacks the `knowledge` scope, or is read-only (`read_only_key`). |
| 429 | `rate_limited` | More than 30 requests per minute. |
| 500 |  | Unexpected failure while creating the source. |

### POST /v1/knowledge-sources/upload

Upload a file (multipart) and ingest it.

Sends the file itself, as `multipart/form-data`. Cortex parses it server-side (OCR included), so a script can push PDFs, Word and Excel files, emails and scans without any text extraction on your side.

The `file` part carries the document; every other part is an optional text field. Keep field values as plain strings (JSON text for `agentIds` and `metadata`).

- **Authentication:** API key (Bearer or x-api-key)
- **Scopes:** `knowledge`

#### Request body

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `file` | `file` | yes |  | The document. 25 MB maximum. See the formats table above. |
| `name` | `string` | no |  | Display name. Defaults to the file name. |
| `description` | `string` | no |  | Free description. |
| `collectionId` | `string` | no |  | Existing collection to file the source under. An unknown id is a `400 Collection not found`. |
| `collectionName` | `string` | no |  | Collection by name. Created if missing, reused if it already exists. Ignored when `collectionId` is set. |
| `metadata` | `string (JSON object)` | no |  | Provenance you want stored with the source, for example `{"origin":"nightly-sync"}`. Arrays and malformed JSON are ignored. |
| `agentIds` | `string (JSON array)` | no |  | Agent ids to link, for example `["agent_1","agent_2"]`. Malformed JSON is ignored. |
| `contextualize` | `string` | no |  | Send `"false"` to skip the per-passage context generation. |

#### Request examples

_curl_

```bash
curl https://api.subsidia.protypa.fr/v1/knowledge-sources/upload \
  -H "Authorization: Bearer $SUBSIDIA_API_KEY" \
  -F "file=@./contrat-bail-2025.pdf" \
  -F "collectionName=Dupont SARL" \
  -F 'metadata={"origin":"nightly-sync"}'
```

_TypeScript_

```typescript
import { readFile } from 'node:fs/promises'

const form = new FormData()
form.append('file', new Blob([await readFile('./contrat-bail-2025.pdf')], { type: 'application/pdf' }), 'contrat-bail-2025.pdf')
form.append('collectionName', 'Dupont SARL')
form.append('metadata', JSON.stringify({ origin: 'nightly-sync' }))

const res = await fetch('https://api.subsidia.protypa.fr/v1/knowledge-sources/upload', {
  method: 'POST',
  headers: { Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY }, // no Content-Type: fetch sets the boundary
  body: form,
})
const source = await res.json() // { id, status }
```

_Python_

```python
import os, requests

with open("contrat-bail-2025.pdf", "rb") as f:
    res = requests.post(
        "https://api.subsidia.protypa.fr/v1/knowledge-sources/upload",
        headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
        files={"file": ("contrat-bail-2025.pdf", f, "application/pdf")},
        data={"collectionName": "Dupont SARL", "metadata": '{"origin": "nightly-sync"}'},
    )
print(res.json())  # {"id": "...", "status": "processing"}
```

#### Responses

**201**: Accepted. `status` is `processing`, or `duplicate` (with `duplicateOf`) when the same bytes are already in the workspace.

```json
{ "id": "clx9k2m4p0002abcd", "status": "processing" }
```

**400**: No `file` part, or `collectionId` does not exist in this workspace.

**413**: File larger than 25 MB.

```json
{ "error": "File too large (25 MB max)" }
```

#### Errors

| Status | Code | When |
| --- | --- | --- |
| 400 |  | No file part, or unknown `collectionId`. |
| 402 | `DOCUMENT_LIMIT` | Document volume of the plan reached. |
| 413 |  | File over 25 MB. |
| 429 | `rate_limited` | More than 30 requests per minute. |

#### Notes

An unsupported file type is **not** rejected at upload: the request succeeds and the source ends in `failed` with the parsing error. Check the source after processing.

### GET /v1/knowledge-sources

List the sources of the workspace.

Newest first. Each entry carries its ingestion state, `_count.chunks` (number of passages) and the agents it is linked to (`agents: [{ id, name }]`). Not paginated.

- **Authentication:** API key
- **Scopes:** `knowledge`

#### Request examples

_curl_

```bash
curl https://api.subsidia.protypa.fr/v1/knowledge-sources -H "Authorization: Bearer $SUBSIDIA_API_KEY"
```

_TypeScript_

```typescript
const { sources } = await fetch('https://api.subsidia.protypa.fr/v1/knowledge-sources', {
  headers: { Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY },
}).then((r) => r.json())
const failed = sources.filter((s: any) => s.status === 'failed')
```

_Python_

```python
import os, requests

sources = requests.get(
    "https://api.subsidia.protypa.fr/v1/knowledge-sources",
    headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
).json()["sources"]
failed = [s for s in sources if s["status"] == "failed"]
```

#### Responses

**200**: The list.

```json
{
  "sources": [
    {
      "id": "clx9k2m4p0002abcd",
      "name": "contrat-bail-2025.pdf",
      "sourceType": "file",
      "fileName": "contrat-bail-2025.pdf",
      "mimeType": "application/pdf",
      "status": "ready",
      "error": null,
      "collectionId": "col_123",
      "progressStage": "done",
      "progressPercent": 100,
      "createdAt": "2026-10-09T08:12:44.000Z",
      "_count": { "chunks": 42 },
      "agents": [{ "id": "agent_1", "name": "Accueil" }]
    }
  ]
}
```

### GET /v1/knowledge-sources/:id

Get one source and its ingestion state.

Poll this route after an ingestion call. `status` moves from `processing` to `ready` or `failed`; while it is `processing`, `progressStage` and `progressPercent` tell you how far it is. On `failed`, `error` explains why.

- **Authentication:** API key
- **Scopes:** `knowledge`

#### Path parameters

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `id` | `string` | yes |  | Source id. |

#### Request examples

_curl_

```bash
curl https://api.subsidia.protypa.fr/v1/knowledge-sources/$SOURCE_ID -H "Authorization: Bearer $SUBSIDIA_API_KEY"
```

_TypeScript_

```typescript
async function waitUntilReady(id: string) {
  for (;;) {
    const { source } = await fetch('https://api.subsidia.protypa.fr/v1/knowledge-sources/' + id, {
      headers: { Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY },
    }).then((r) => r.json())
    if (source.status !== 'processing') return source
    await new Promise((r) => setTimeout(r, 3000))
  }
}
```

_Python_

```python
import os, time, requests

def wait_until_ready(source_id):
    while True:
        source = requests.get(
            "https://api.subsidia.protypa.fr/v1/knowledge-sources/" + source_id,
            headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
        ).json()["source"]
        if source["status"] != "processing":
            return source
        time.sleep(3)
```

#### Responses

**200**: The source.

```json
{
  "source": {
    "id": "clx9k2m4p0002abcd",
    "name": "contrat-bail-2025.pdf",
    "status": "processing",
    "progressStage": "embedding",
    "progressPercent": 62,
    "etaSeconds": 18,
    "error": null,
    "agents": [],
    "_count": { "chunks": 42 }
  }
}
```

**404**: Unknown id, or the source belongs to another workspace.

```json
{ "error": "Knowledge source not found" }
```

#### Errors

| Status | Code | When |
| --- | --- | --- |
| 404 |  | No such source in this workspace. |

### DELETE /v1/knowledge-sources/:id

Delete a source.

Removes the source from the workspace so its passages can no longer be retrieved. Certificates of answers that already cited it are unaffected. Irreversible.

- **Authentication:** API key
- **Scopes:** `knowledge`

#### Path parameters

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `id` | `string` | yes |  | Source id. |

#### Request examples

_curl_

```bash
curl -X DELETE https://api.subsidia.protypa.fr/v1/knowledge-sources/$SOURCE_ID -H "Authorization: Bearer $SUBSIDIA_API_KEY"
```

_Python_

```python
import os, requests

res = requests.delete(
    "https://api.subsidia.protypa.fr/v1/knowledge-sources/" + source_id,
    headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
)
assert res.status_code == 204
```

#### Responses

**204**: Deleted. No body.

**404**: Unknown id.

```json
{ "error": "Knowledge source not found" }
```

#### Errors

| Status | Code | When |
| --- | --- | --- |
| 403 | `read_only_key` | The key is read-only. |

## Attach a source to an agent

An [agent](https://dev.subsidia.protypa.fr/docs/agents.md) answers from the sources linked to it. Link at creation with `agentIds`, or afterwards with the two routes below. Linking is many-to-many: one source can serve several agents, and unlinking never deletes the source. Both routes use the `agents` scope.

### POST /v1/agents/:id/knowledge-sources/:sourceId

Link a source to an agent.

Linking twice is harmless. Returns the agent with the list of its linked sources.

- **Authentication:** API key
- **Scopes:** `agents`

#### Path parameters

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `id` | `string` | yes |  | Agent id. |
| `sourceId` | `string` | yes |  | Source id. |

#### Request examples

_curl_

```bash
curl -X POST https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/knowledge-sources/$SOURCE_ID \
  -H "Authorization: Bearer $SUBSIDIA_API_KEY"
```

_TypeScript_

```typescript
const { agent } = await fetch(
  'https://api.subsidia.protypa.fr/v1/agents/' + agentId + '/knowledge-sources/' + sourceId,
  { method: 'POST', headers: { Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY } },
).then((r) => r.json())
console.log(agent.knowledgeSources)
```

_Python_

```python
import os, requests

agent = requests.post(
    "https://api.subsidia.protypa.fr/v1/agents/" + agent_id + "/knowledge-sources/" + source_id,
    headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
).json()["agent"]
```

#### Responses

**200**: The updated agent.

```json
{ "agent": { "id": "agent_1", "name": "Accueil", "knowledgeSources": [{ "id": "clx9k2m4p0002abcd", "name": "contrat-bail-2025.pdf" }] } }
```

**404**: `Agent not found` or `Knowledge source not found` in this workspace.

#### Errors

| Status | Code | When |
| --- | --- | --- |
| 404 |  | The agent or the source does not exist in this workspace. |

### DELETE /v1/agents/:id/knowledge-sources/:sourceId

Unlink a source from an agent.

The agent stops retrieving from the source. The source stays in the workspace.

- **Authentication:** API key
- **Scopes:** `agents`

#### Path parameters

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `id` | `string` | yes |  | Agent id. |
| `sourceId` | `string` | yes |  | Source id. |

#### Request examples

_curl_

```bash
curl -X DELETE https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/knowledge-sources/$SOURCE_ID \
  -H "Authorization: Bearer $SUBSIDIA_API_KEY"
```

_Python_

```python
import os, requests

requests.delete(
    "https://api.subsidia.protypa.fr/v1/agents/" + agent_id + "/knowledge-sources/" + source_id,
    headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
)
```

#### Responses

**204**: Unlinked. No body.

**404**: `Agent not found`.

## Know when a source is ready

Polling works for scripts. For pipelines, subscribe to the `source.ingested` webhook (payload: `sourceId`, `name`, `chunkCount`, `factCount`, `contradictionCount`, `qaCount`) and `source.failed` (payload: `sourceId`, `name`, `error`). See [Webhooks](https://dev.subsidia.protypa.fr/docs/webhooks.md).

## Frequently asked

**Can I update a document?**

Not in place. Delete the old source and upload the new version; a changed file has a different hash, so it is ingested as new content.

**A source stays in `processing` for a long time.**

`progressStage` and `etaSeconds` show the current step. Large PDFs, scans (OCR) and contextualization take time, especially on a small local machine. If the server restarts mid-ingestion, sources that got past parsing resume where they stopped; the others are marked `failed` with a clear message and can be uploaded again.

**Does ingestion send my documents to an external model?**

Ingestion uses the model configured for the workspace. With a local model nothing leaves the machine; with an external provider the PII shield applies as for any other call. See [Synapse](https://dev.subsidia.protypa.fr/docs/synapse.md).
