Querying the knowledge base
POST /v1/knowledge/query runs hybrid retrieval with reranking and returns cited passages plus a trace of why they matched.
POST /v1/knowledge/query returns the passages of your documents that best answer a question. It does not write an answer: you get the evidence and decide what to do with it, which is what you want when you build your own grounded prompt, a citation panel or a verification step.
For a ready-made answer with citations or an honest refusal, see Light. For retrieval inside a normal chat call, add +cortex to the model address in the Gateway.
The retrieval pipeline
- Question
- Query expansion
- Vector search + full-text search
- RRF fusion
- Rerank (guarded)
- topK passages + trace
- 1Expand
A compound question ("What is the notice period and who signs?") is decomposed into sub-queries, so each part can find its own passages. Disable with
expand: false. - 2Search twice per sub-query
A semantic (embedding) search and a full-text search run in parallel. Meaning finds paraphrases; the full-text channel finds exact terms, numbers and names that embeddings blur. A passage found by both ranks higher.
- 3Fuse
The rankings are merged with reciprocal rank fusion (RRF), which rewards agreement between channels without needing comparable scores.
- 4Rerank, with guards
A small pool of the best candidates is re-scored by the model for relevance. If the judge finds nothing relevant, or contradicts the consensus ranking wildly, its verdict is discarded and a deterministic order is used (passages sharing words with the question first). A bad judge therefore cannot corrupt the result. If the pool is no larger than
topK, reranking is skipped since it could not change anything. - 5Return
The
topKbest passages come back with their source, section, page and scores, and atracethat tells you what happened.
POST/v1/knowledge/query
Retrieve the passages that best match a question.
knowledgeReturns up to topK passages from the sources of the workspace, best first, and a retrieval trace. An empty matches array is a normal answer meaning the documents contain nothing relevant; treat it as an honest "not found" and do not let a model fill the gap.
By default the whole workspace is searched. Pass agentId to restrict the search to the sources linked to one agent (see Knowledge sources).
Request body
querystringrequiredThe question or search phrase. Must not be empty. A full question works better than keywords.topKnumberdefault5How many passages to return, 1 to 20.agentIdstringRestrict retrieval to the sources linked to this agent.expandbooleandefaulttrueDecompose compound questions into sub-queries. Setfalsefor short, single-fact lookups to save a model call and latency.rerankbooleandefaulttrueLet the model re-score the fused candidates. Setfalsefor the fastest, cheapest path (embeddings and full-text only).
Request examples
curl https://api.subsidia.protypa.fr/v1/knowledge/query \ -H "Authorization: Bearer $SUBSIDIA_API_KEY" \ -H "Content-Type: application/json" \ -d '{"query": "What is the notice period to terminate the lease?", "topK": 5}'Responses
The matches and the trace. matches may be empty.
{ "matches": [ { "sourceId": "clx9k2m4p0002abcd", "sourceName": "contrat-bail-2025.pdf", "collectionId": "col_123", "dossierName": "Dupont SARL", "chunkId": "ck_8f31a0c2", "section": "Article 12 - Termination", "page": 6, "content": "The tenant may terminate the lease at the end of each three-year period by giving six months' notice by registered letter.", "similarity": 0.81, "score": 0.0328, "matchedBy": ["vector", "fts"], "rerankScore": 9 } ], "trace": { "queries": ["What is the notice period to terminate the lease?"], "vectorBackend": "pgvector", "candidateCount": 14, "rerankApplied": true, "timings": { "expandMs": 0, "searchMs": 41, "rerankMs": 612 }, "tokensUsed": 380 }}Errors
- 400
querymissing or empty,topKout of range. - 401Missing or invalid key.
- 403
scope_deniedThe key lacks theknowledgescope. - 429
rate_limitedMore than 60 requests per minute.
Notes
page is null when the source has no pages (text, Word, spreadsheets). rerankScore is null when reranking did not apply. trace.tokensUsed is the internal cost of expansion and reranking; it is not a question count.
Reading the response
| Field | Meaning |
|---|---|
matches[].content | The passage text, at most about 2000 characters. Quote it, do not paraphrase it, when you cite. |
matches[].sourceId, sourceName | The document it comes from. Use sourceId to link back to GET /v1/knowledge-sources/:id. |
matches[].section, page | Where in the document: nearest heading, and 1-based page for PDFs and scans. |
matches[].collectionId, dossierName | The collection (client file) of the source, or null when unfiled. |
matches[].chunkId | Stable id of the passage. This is what proof certificates commit to. |
matches[].similarity | Best cosine similarity of the passage, 0 to 1. A raw signal, not a probability: do not set a universal threshold on it. |
matches[].score | The fused RRF score. Only meaningful for comparing passages inside this response. |
matches[].matchedBy | vector, fts, or both. A passage matched by both is usually the safest evidence. |
matches[].rerankScore | The model relevance score when reranking applied, else null. |
trace.queries | The sub-queries actually searched (one when expansion is off or the question is simple). |
trace.vectorBackend | pgvector or memory; the engine that served the semantic search. |
trace.candidateCount | Number of candidates before the final cut. |
trace.rerankApplied | Whether the model order was used. false also covers the case where its verdict was discarded by the guards. |
trace.timings | Milliseconds spent expanding, searching and reranking. |
Build a grounded answer
The pattern: retrieve, number the passages, instruct the model to answer only from them and to say so when they do not cover the question, then map the [n] markers in the answer back to sources. This is the recipe Light applies server-side.
If you send the final call through the Gateway, you get a signed certificate for it. If you would rather not assemble the prompt at all, use model: "pulse-auto+cortex" and the Gateway does the retrieval and records the passages in the certificate.
const API = 'https://api.subsidia.protypa.fr'const headers = { Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY, 'Content-Type': 'application/json',}
async function answer(question: string) { const { matches } = await fetch(API + '/v1/knowledge/query', { method: 'POST', headers, body: JSON.stringify({ query: question, topK: 6 }), }).then((r) => r.json())
if (matches.length === 0) return { answer: 'The documents do not cover this.', sources: [] }
const passages = matches .map((m: any, i: number) => '[' + (i + 1) + '] ' + m.sourceName + (m.page ? ' p.' + m.page : '') + '\n' + m.content) .join('\n\n')
const res = await fetch(API + '/v1/chat/completions', { method: 'POST', headers, body: JSON.stringify({ model: 'pulse-auto', messages: [ { role: 'system', content: 'Answer using ONLY the numbered passages below. Cite each claim as [n]. ' + 'If the passages do not contain the answer, reply exactly: The documents do not cover this.\n\n' + passages, }, { role: 'user', content: question }, ], }), }) const proofId = res.headers.get('x-pulse-proof') const completion = await res.json() return { answer: completion.choices[0].message.content, sources: matches, proofId }}Tips
| Situation | Do |
|---|---|
| Single-fact lookup (a date, an amount, a name) | expand: false and rerank: false for the lowest latency. The full-text channel handles exact values well. |
| Compound or vague question | Keep expand: true; ask the whole question rather than keywords. |
| Lists ("all deadlines", "every client") | Raise topK (10 to 20). The query detects enumeration questions and widens its candidate pool, and tries to surface the table that lists the answer. |
| One client or one team | Link the sources to an agent and pass its agentId, or file them in a collection and use Light with dossierId. |
Empty matches | Say so to the user. Do not fall back to a model answer without telling them it is not from their documents. |
| Fresh uploads not found | Check the source is ready. Passages of a processing source are not searchable yet. |
Frequently asked
Does this endpoint consume questions?
It returns passages, not an answer. Expansion and reranking use the workspace model internally (see trace.tokensUsed); setting expand: false and rerank: false means no generative model is called (only the embedding of your question).
Is there a similarity threshold to drop weak matches?
Not in the API. similarity depends on the language and the embedding model, so a fixed threshold is fragile. Prefer the combination of matchedBy containing both channels and a high rerankScore, or let Light decide with its strict grounded prompt.
How do I cite a page?
Use sourceName, page (when not null) and section. Link sourceId back to your own record of the document.