Knowledge sources
Ingest text, URLs and files (PDF, Office, email, images) into Cortex, track ingestion, and attach sources to agents.
A knowledge source is one document in your workspace: a file, a web page or a block of text. When you add one, Cortex parses it, cuts it into passages, embeds them, extracts atomic facts and prepares health-check questions. From then on its passages can be retrieved by /v1/knowledge/query, by agents linked to it, and by Gateway calls that carry the +cortex flag.
Ingestion is asynchronous. The API answers immediately with the source id and a status; you poll the source (or listen to a webhook) until it is ready.
- Request (text, url or file)
- Duplicate check
- Parsing (+ OCR)
- Chunks
- Contextualizing
- Embedding
- Facts + contradictions
- Health QA
- ready
Concepts
| Field | Values | Meaning |
|---|---|---|
status | processing, ready, failed | Lifecycle of the ingestion. Only ready sources are searchable. A failed source keeps the reason in error. |
progressStage | parsing, chunks, contextualizing, embedding, facts, qa, done | Where a processing source currently is. progressPercent (0-100) and etaSeconds accompany it. |
sourceType | text, file, url | How the source entered the workspace. |
collectionId | string or null | The collection (a "client file" or dossier in the app) the source is filed under. Retrieval results echo it back as collectionId and dossierName. |
Supported formats
| Format | Extensions | Notes |
|---|---|---|
pdf | Text layer first. A scanned PDF falls back to OCR (French and English by default, 60 pages by default on the server). Passages carry a page number when known. | |
| Word | docx | Headings are kept as sections. |
| Spreadsheets | xlsx, xlsm, csv, fec | Tables are linearised into readable rows, capped at 5000 rows per sheet (a truncation is stated in the text). Legacy .xls is refused with a message asking for .xlsx or .csv. |
| Presentations | pptx | Slide text. |
eml, msg | Headers and body. | |
| Images | png, jpg, jpeg, webp, bmp | Read by OCR. |
| Web and data | html, htm, xml, json | HTML is converted to Markdown. |
| Text | md, markdown, txt, text | Used as is. |
Any other type is rejected at parsing time: the source ends in failed with the message Unsupported file type .... Passages are cut on the document structure (headings, paragraphs, table rows) and never exceed 2000 characters.
Limits
| Limit | Value |
|---|---|
File size (/upload) | 25 MB. Larger files return 413. |
| Request rate | 30 requests per minute for POST /v1/knowledge-sources and POST /v1/knowledge-sources/upload; 429 beyond that. |
| URL ingestion | Public http(s) addresses only. Private, local and non-http addresses are refused with 400. |
| Documents per workspace | Set by the plan. 402 DOCUMENT_LIMIT when reached. |
Endpoints
POST/v1/knowledge-sources
Ingest raw text or a public URL.
knowledgeCreates a source from text you send or from a page Cortex fetches. Send either content or url. The response comes back at once with status: "processing"; poll GET /v1/knowledge-sources/:id until ready.
Use this route when you already hold the text (a CRM note, an export). For files, use the upload route: you do not need to extract the text yourself, and confidential documents never have to be published to a URL.
Request body
namestringDisplay name. Required withcontent; defaults to the URL for URL ingestion.descriptionstringFree description.contentstringRaw text to ingest. Must not be empty.urlstringPublic http(s) page or document to fetch and ingest. Alternative tocontent.agentIdsstring[]Agent ids to link the source to at creation, so those agents can retrieve from it.collectionIdstringCollection to file the source under.contextualizebooleandefaulttrueGenerate a one-line context for each passage before embedding. Improves retrieval, costs ingestion time. Setfalseto skip.extractFactsbooleandefaulttrueExtract atomic facts and detect contradictions with other sources. See Facts, contradictions and health.generateQabooleandefaulttrueGenerate question/answer pairs used by the health replay.
Request examples
curl https://api.subsidia.protypa.fr/v1/knowledge-sources \ -H "Authorization: Bearer $SUBSIDIA_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "name": "Termination policy", "content": "Either party may terminate the agreement with 30 days written notice." }'Responses
The source was created and is being processed. A duplicate returns status: "duplicate" and duplicateOf.
{ "id": "clx9k2m4p0001abcd", "status": "processing" }Errors
- 400Missing
content/url, missingnamewithcontent, or an unsafe URL. - 401Missing or invalid key.
- 402
DOCUMENT_LIMITThe workspace reached the document volume of its plan. - 403
scope_deniedThe key lacks theknowledgescope, or is read-only (read_only_key). - 429
rate_limitedMore than 30 requests per minute. - 500Unexpected failure while creating the source.
POST/v1/knowledge-sources/upload
Upload a file (multipart) and ingest it.
knowledgeSends the file itself, as multipart/form-data. Cortex parses it server-side (OCR included), so a script can push PDFs, Word and Excel files, emails and scans without any text extraction on your side.
The file part carries the document; every other part is an optional text field. Keep field values as plain strings (JSON text for agentIds and metadata).
Request body
filefilerequiredThe document. 25 MB maximum. See the formats table above.namestringDisplay name. Defaults to the file name.descriptionstringFree description.collectionIdstringExisting collection to file the source under. An unknown id is a400 Collection not found.collectionNamestringCollection by name. Created if missing, reused if it already exists. Ignored whencollectionIdis set.metadatastring (JSON object)Provenance you want stored with the source, for example{"origin":"nightly-sync"}. Arrays and malformed JSON are ignored.agentIdsstring (JSON array)Agent ids to link, for example["agent_1","agent_2"]. Malformed JSON is ignored.contextualizestringSend"false"to skip the per-passage context generation.
Request examples
curl https://api.subsidia.protypa.fr/v1/knowledge-sources/upload \ -H "Authorization: Bearer $SUBSIDIA_API_KEY" \ -F "file=@./contrat-bail-2025.pdf" \ -F "collectionName=Dupont SARL" \ -F 'metadata={"origin":"nightly-sync"}'Responses
Accepted. status is processing, or duplicate (with duplicateOf) when the same bytes are already in the workspace.
{ "id": "clx9k2m4p0002abcd", "status": "processing" }Errors
- 400No file part, or unknown
collectionId. - 402
DOCUMENT_LIMITDocument volume of the plan reached. - 413File over 25 MB.
- 429
rate_limitedMore than 30 requests per minute.
Notes
An unsupported file type is not rejected at upload: the request succeeds and the source ends in failed with the parsing error. Check the source after processing.
GET/v1/knowledge-sources
List the sources of the workspace.
knowledgeNewest first. Each entry carries its ingestion state, _count.chunks (number of passages) and the agents it is linked to (agents: [{ id, name }]). Not paginated.
Request examples
curl https://api.subsidia.protypa.fr/v1/knowledge-sources -H "Authorization: Bearer $SUBSIDIA_API_KEY"Responses
The list.
{ "sources": [ { "id": "clx9k2m4p0002abcd", "name": "contrat-bail-2025.pdf", "sourceType": "file", "fileName": "contrat-bail-2025.pdf", "mimeType": "application/pdf", "status": "ready", "error": null, "collectionId": "col_123", "progressStage": "done", "progressPercent": 100, "createdAt": "2026-10-09T08:12:44.000Z", "_count": { "chunks": 42 }, "agents": [{ "id": "agent_1", "name": "Accueil" }] } ]}GET/v1/knowledge-sources/:id
Get one source and its ingestion state.
knowledgePoll this route after an ingestion call. status moves from processing to ready or failed; while it is processing, progressStage and progressPercent tell you how far it is. On failed, error explains why.
Path parameters
idstringrequiredSource id.
Request examples
curl https://api.subsidia.protypa.fr/v1/knowledge-sources/$SOURCE_ID -H "Authorization: Bearer $SUBSIDIA_API_KEY"Responses
The source.
{ "source": { "id": "clx9k2m4p0002abcd", "name": "contrat-bail-2025.pdf", "status": "processing", "progressStage": "embedding", "progressPercent": 62, "etaSeconds": 18, "error": null, "agents": [], "_count": { "chunks": 42 } }}Errors
- 404No such source in this workspace.
DELETE/v1/knowledge-sources/:id
Delete a source.
knowledgeRemoves the source from the workspace so its passages can no longer be retrieved. Certificates of answers that already cited it are unaffected. Irreversible.
Path parameters
idstringrequiredSource id.
Request examples
curl -X DELETE https://api.subsidia.protypa.fr/v1/knowledge-sources/$SOURCE_ID -H "Authorization: Bearer $SUBSIDIA_API_KEY"Responses
Deleted. No body.
Errors
- 403
read_only_keyThe key is read-only.
Attach a source to an agent
An agent answers from the sources linked to it. Link at creation with agentIds, or afterwards with the two routes below. Linking is many-to-many: one source can serve several agents, and unlinking never deletes the source. Both routes use the agents scope.
POST/v1/agents/:id/knowledge-sources/:sourceId
Link a source to an agent.
agentsLinking twice is harmless. Returns the agent with the list of its linked sources.
Path parameters
idstringrequiredAgent id.sourceIdstringrequiredSource id.
Request examples
curl -X POST https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/knowledge-sources/$SOURCE_ID \ -H "Authorization: Bearer $SUBSIDIA_API_KEY"Responses
The updated agent.
{ "agent": { "id": "agent_1", "name": "Accueil", "knowledgeSources": [{ "id": "clx9k2m4p0002abcd", "name": "contrat-bail-2025.pdf" }] } }Errors
- 404The agent or the source does not exist in this workspace.
DELETE/v1/agents/:id/knowledge-sources/:sourceId
Unlink a source from an agent.
agentsThe agent stops retrieving from the source. The source stays in the workspace.
Path parameters
idstringrequiredAgent id.sourceIdstringrequiredSource id.
Request examples
curl -X DELETE https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/knowledge-sources/$SOURCE_ID \ -H "Authorization: Bearer $SUBSIDIA_API_KEY"Responses
Unlinked. No body.
Know when a source is ready
Polling works for scripts. For pipelines, subscribe to the source.ingested webhook (payload: sourceId, name, chunkCount, factCount, contradictionCount, qaCount) and source.failed (payload: sourceId, name, error). See Webhooks.
Frequently asked
Can I update a document?
Not in place. Delete the old source and upload the new version; a changed file has a different hash, so it is ingested as new content.
A source stays in `processing` for a long time.
progressStage and etaSeconds show the current step. Large PDFs, scans (OCR) and contextualization take time, especially on a small local machine. If the server restarts mid-ingestion, sources that got past parsing resume where they stopped; the others are marked failed with a clear message and can be uploaded again.
Does ingestion send my documents to an external model?
Ingestion uses the model configured for the workspace. With a local model nothing leaves the machine; with an external provider the PII shield applies as for any other call. See Synapse.