Skip to content

Knowledge sources

Ingest text, URLs and files (PDF, Office, email, images) into Cortex, track ingestion, and attach sources to agents.

A knowledge source is one document in your workspace: a file, a web page or a block of text. When you add one, Cortex parses it, cuts it into passages, embeds them, extracts atomic facts and prepares health-check questions. From then on its passages can be retrieved by /v1/knowledge/query, by agents linked to it, and by Gateway calls that carry the +cortex flag.

Ingestion is asynchronous. The API answers immediately with the source id and a status; you poll the source (or listen to a webhook) until it is ready.

  1. Request (text, url or file)
  2. Duplicate check
  3. Parsing (+ OCR)
  4. Chunks
  5. Contextualizing
  6. Embedding
  7. Facts + contradictions
  8. Health QA
  9. ready
Ingestion lifecycle. The `progressStage` field of a source walks through these stages while `status` is `processing`.

Concepts

FieldValuesMeaning
statusprocessing, ready, failedLifecycle of the ingestion. Only ready sources are searchable. A failed source keeps the reason in error.
progressStageparsing, chunks, contextualizing, embedding, facts, qa, doneWhere a processing source currently is. progressPercent (0-100) and etaSeconds accompany it.
sourceTypetext, file, urlHow the source entered the workspace.
collectionIdstring or nullThe collection (a "client file" or dossier in the app) the source is filed under. Retrieval results echo it back as collectionId and dossierName.

Supported formats

FormatExtensionsNotes
PDFpdfText layer first. A scanned PDF falls back to OCR (French and English by default, 60 pages by default on the server). Passages carry a page number when known.
WorddocxHeadings are kept as sections.
Spreadsheetsxlsx, xlsm, csv, fecTables are linearised into readable rows, capped at 5000 rows per sheet (a truncation is stated in the text). Legacy .xls is refused with a message asking for .xlsx or .csv.
PresentationspptxSlide text.
Emaileml, msgHeaders and body.
Imagespng, jpg, jpeg, webp, bmpRead by OCR.
Web and datahtml, htm, xml, jsonHTML is converted to Markdown.
Textmd, markdown, txt, textUsed as is.

Any other type is rejected at parsing time: the source ends in failed with the message Unsupported file type .... Passages are cut on the document structure (headings, paragraphs, table rows) and never exceed 2000 characters.

Limits

LimitValue
File size (/upload)25 MB. Larger files return 413.
Request rate30 requests per minute for POST /v1/knowledge-sources and POST /v1/knowledge-sources/upload; 429 beyond that.
URL ingestionPublic http(s) addresses only. Private, local and non-http addresses are refused with 400.
Documents per workspaceSet by the plan. 402 DOCUMENT_LIMIT when reached.

Endpoints

POST/v1/knowledge-sources

Ingest raw text or a public URL.

API key (Bearer or x-api-key)knowledge

Creates a source from text you send or from a page Cortex fetches. Send either content or url. The response comes back at once with status: "processing"; poll GET /v1/knowledge-sources/:id until ready.

Use this route when you already hold the text (a CRM note, an export). For files, use the upload route: you do not need to extract the text yourself, and confidential documents never have to be published to a URL.

Request body

  • namestring
    Display name. Required with content; defaults to the URL for URL ingestion.
  • descriptionstring
    Free description.
  • contentstring
    Raw text to ingest. Must not be empty.
  • urlstring
    Public http(s) page or document to fetch and ingest. Alternative to content.
  • agentIdsstring[]
    Agent ids to link the source to at creation, so those agents can retrieve from it.
  • collectionIdstring
    Collection to file the source under.
  • contextualizebooleandefault true
    Generate a one-line context for each passage before embedding. Improves retrieval, costs ingestion time. Set false to skip.
  • extractFactsbooleandefault true
    Extract atomic facts and detect contradictions with other sources. See Facts, contradictions and health.
  • generateQabooleandefault true
    Generate question/answer pairs used by the health replay.

Request examples

curl https://api.subsidia.protypa.fr/v1/knowledge-sources \
-H "Authorization: Bearer $SUBSIDIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Termination policy",
"content": "Either party may terminate the agreement with 30 days written notice."
}'

Responses

The source was created and is being processed. A duplicate returns status: "duplicate" and duplicateOf.

Example response
{ "id": "clx9k2m4p0001abcd", "status": "processing" }

Errors

  • 400Missing content/url, missing name with content, or an unsafe URL.
  • 401Missing or invalid key.
  • 402DOCUMENT_LIMITThe workspace reached the document volume of its plan.
  • 403scope_deniedThe key lacks the knowledge scope, or is read-only (read_only_key).
  • 429rate_limitedMore than 30 requests per minute.
  • 500Unexpected failure while creating the source.

POST/v1/knowledge-sources/upload

Upload a file (multipart) and ingest it.

API key (Bearer or x-api-key)knowledge

Sends the file itself, as multipart/form-data. Cortex parses it server-side (OCR included), so a script can push PDFs, Word and Excel files, emails and scans without any text extraction on your side.

The file part carries the document; every other part is an optional text field. Keep field values as plain strings (JSON text for agentIds and metadata).

Request body

  • filefilerequired
    The document. 25 MB maximum. See the formats table above.
  • namestring
    Display name. Defaults to the file name.
  • descriptionstring
    Free description.
  • collectionIdstring
    Existing collection to file the source under. An unknown id is a 400 Collection not found.
  • collectionNamestring
    Collection by name. Created if missing, reused if it already exists. Ignored when collectionId is set.
  • metadatastring (JSON object)
    Provenance you want stored with the source, for example {"origin":"nightly-sync"}. Arrays and malformed JSON are ignored.
  • agentIdsstring (JSON array)
    Agent ids to link, for example ["agent_1","agent_2"]. Malformed JSON is ignored.
  • contextualizestring
    Send "false" to skip the per-passage context generation.

Request examples

curl https://api.subsidia.protypa.fr/v1/knowledge-sources/upload \
-H "Authorization: Bearer $SUBSIDIA_API_KEY" \
-F "file=@./contrat-bail-2025.pdf" \
-F "collectionName=Dupont SARL" \
-F 'metadata={"origin":"nightly-sync"}'

Responses

Accepted. status is processing, or duplicate (with duplicateOf) when the same bytes are already in the workspace.

Example response
{ "id": "clx9k2m4p0002abcd", "status": "processing" }

Errors

  • 400No file part, or unknown collectionId.
  • 402DOCUMENT_LIMITDocument volume of the plan reached.
  • 413File over 25 MB.
  • 429rate_limitedMore than 30 requests per minute.

Notes

An unsupported file type is not rejected at upload: the request succeeds and the source ends in failed with the parsing error. Check the source after processing.

GET/v1/knowledge-sources

List the sources of the workspace.

API keyknowledge

Newest first. Each entry carries its ingestion state, _count.chunks (number of passages) and the agents it is linked to (agents: [{ id, name }]). Not paginated.

Request examples

curl https://api.subsidia.protypa.fr/v1/knowledge-sources -H "Authorization: Bearer $SUBSIDIA_API_KEY"

Responses

The list.

Example response
{
"sources": [
{
"id": "clx9k2m4p0002abcd",
"name": "contrat-bail-2025.pdf",
"sourceType": "file",
"fileName": "contrat-bail-2025.pdf",
"mimeType": "application/pdf",
"status": "ready",
"error": null,
"collectionId": "col_123",
"progressStage": "done",
"progressPercent": 100,
"createdAt": "2026-10-09T08:12:44.000Z",
"_count": { "chunks": 42 },
"agents": [{ "id": "agent_1", "name": "Accueil" }]
}
]
}

GET/v1/knowledge-sources/:id

Get one source and its ingestion state.

API keyknowledge

Poll this route after an ingestion call. status moves from processing to ready or failed; while it is processing, progressStage and progressPercent tell you how far it is. On failed, error explains why.

Path parameters

  • idstringrequired
    Source id.

Request examples

curl https://api.subsidia.protypa.fr/v1/knowledge-sources/$SOURCE_ID -H "Authorization: Bearer $SUBSIDIA_API_KEY"

Responses

The source.

Example response
{
"source": {
"id": "clx9k2m4p0002abcd",
"name": "contrat-bail-2025.pdf",
"status": "processing",
"progressStage": "embedding",
"progressPercent": 62,
"etaSeconds": 18,
"error": null,
"agents": [],
"_count": { "chunks": 42 }
}
}

Errors

  • 404No such source in this workspace.

DELETE/v1/knowledge-sources/:id

Delete a source.

API keyknowledge

Removes the source from the workspace so its passages can no longer be retrieved. Certificates of answers that already cited it are unaffected. Irreversible.

Path parameters

  • idstringrequired
    Source id.

Request examples

curl -X DELETE https://api.subsidia.protypa.fr/v1/knowledge-sources/$SOURCE_ID -H "Authorization: Bearer $SUBSIDIA_API_KEY"

Responses

Deleted. No body.

Errors

  • 403read_only_keyThe key is read-only.

Attach a source to an agent

An agent answers from the sources linked to it. Link at creation with agentIds, or afterwards with the two routes below. Linking is many-to-many: one source can serve several agents, and unlinking never deletes the source. Both routes use the agents scope.

POST/v1/agents/:id/knowledge-sources/:sourceId

Link a source to an agent.

API keyagents

Linking twice is harmless. Returns the agent with the list of its linked sources.

Path parameters

  • idstringrequired
    Agent id.
  • sourceIdstringrequired
    Source id.

Request examples

curl -X POST https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/knowledge-sources/$SOURCE_ID \
-H "Authorization: Bearer $SUBSIDIA_API_KEY"

Responses

The updated agent.

Example response
{ "agent": { "id": "agent_1", "name": "Accueil", "knowledgeSources": [{ "id": "clx9k2m4p0002abcd", "name": "contrat-bail-2025.pdf" }] } }

Errors

  • 404The agent or the source does not exist in this workspace.

DELETE/v1/agents/:id/knowledge-sources/:sourceId

Unlink a source from an agent.

API keyagents

The agent stops retrieving from the source. The source stays in the workspace.

Path parameters

  • idstringrequired
    Agent id.
  • sourceIdstringrequired
    Source id.

Request examples

curl -X DELETE https://api.subsidia.protypa.fr/v1/agents/$AGENT_ID/knowledge-sources/$SOURCE_ID \
-H "Authorization: Bearer $SUBSIDIA_API_KEY"

Responses

Unlinked. No body.

Know when a source is ready

Polling works for scripts. For pipelines, subscribe to the source.ingested webhook (payload: sourceId, name, chunkCount, factCount, contradictionCount, qaCount) and source.failed (payload: sourceId, name, error). See Webhooks.

Frequently asked

Can I update a document?

Not in place. Delete the old source and upload the new version; a changed file has a different hash, so it is ingested as new content.

A source stays in `processing` for a long time.

progressStage and etaSeconds show the current step. Large PDFs, scans (OCR) and contextualization take time, especially on a small local machine. If the server restarts mid-ingestion, sources that got past parsing resume where they stopped; the others are marked failed with a clear message and can be uploaded again.

Does ingestion send my documents to an external model?

Ingestion uses the model configured for the workspace. With a local model nothing leaves the machine; with an external provider the PII shield applies as for any other call. See Synapse.

Related