Skip to content

Querying the knowledge base

POST /v1/knowledge/query runs hybrid retrieval with reranking and returns cited passages plus a trace of why they matched.

POST /v1/knowledge/query returns the passages of your documents that best answer a question. It does not write an answer: you get the evidence and decide what to do with it, which is what you want when you build your own grounded prompt, a citation panel or a verification step.

For a ready-made answer with citations or an honest refusal, see Light. For retrieval inside a normal chat call, add +cortex to the model address in the Gateway.

The retrieval pipeline

  1. Question
  2. Query expansion
  3. Vector search + full-text search
  4. RRF fusion
  5. Rerank (guarded)
  6. topK passages + trace
Every query goes through the same stages. Query expansion and reranking use the workspace model and can be switched off per call.
  1. 1
    Expand

    A compound question ("What is the notice period and who signs?") is decomposed into sub-queries, so each part can find its own passages. Disable with expand: false.

  2. 2
    Search twice per sub-query

    A semantic (embedding) search and a full-text search run in parallel. Meaning finds paraphrases; the full-text channel finds exact terms, numbers and names that embeddings blur. A passage found by both ranks higher.

  3. 3
    Fuse

    The rankings are merged with reciprocal rank fusion (RRF), which rewards agreement between channels without needing comparable scores.

  4. 4
    Rerank, with guards

    A small pool of the best candidates is re-scored by the model for relevance. If the judge finds nothing relevant, or contradicts the consensus ranking wildly, its verdict is discarded and a deterministic order is used (passages sharing words with the question first). A bad judge therefore cannot corrupt the result. If the pool is no larger than topK, reranking is skipped since it could not change anything.

  5. 5
    Return

    The topK best passages come back with their source, section, page and scores, and a trace that tells you what happened.

POST/v1/knowledge/query

Retrieve the passages that best match a question.

API key (Bearer or x-api-key)knowledge

Returns up to topK passages from the sources of the workspace, best first, and a retrieval trace. An empty matches array is a normal answer meaning the documents contain nothing relevant; treat it as an honest "not found" and do not let a model fill the gap.

By default the whole workspace is searched. Pass agentId to restrict the search to the sources linked to one agent (see Knowledge sources).

Request body

  • querystringrequired
    The question or search phrase. Must not be empty. A full question works better than keywords.
  • topKnumberdefault 5
    How many passages to return, 1 to 20.
  • agentIdstring
    Restrict retrieval to the sources linked to this agent.
  • expandbooleandefault true
    Decompose compound questions into sub-queries. Set false for short, single-fact lookups to save a model call and latency.
  • rerankbooleandefault true
    Let the model re-score the fused candidates. Set false for the fastest, cheapest path (embeddings and full-text only).

Request examples

curl https://api.subsidia.protypa.fr/v1/knowledge/query \
-H "Authorization: Bearer $SUBSIDIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query": "What is the notice period to terminate the lease?", "topK": 5}'

Responses

The matches and the trace. matches may be empty.

Example response
{
"matches": [
{
"sourceId": "clx9k2m4p0002abcd",
"sourceName": "contrat-bail-2025.pdf",
"collectionId": "col_123",
"dossierName": "Dupont SARL",
"chunkId": "ck_8f31a0c2",
"section": "Article 12 - Termination",
"page": 6,
"content": "The tenant may terminate the lease at the end of each three-year period by giving six months' notice by registered letter.",
"similarity": 0.81,
"score": 0.0328,
"matchedBy": ["vector", "fts"],
"rerankScore": 9
}
],
"trace": {
"queries": ["What is the notice period to terminate the lease?"],
"vectorBackend": "pgvector",
"candidateCount": 14,
"rerankApplied": true,
"timings": { "expandMs": 0, "searchMs": 41, "rerankMs": 612 },
"tokensUsed": 380
}
}

Errors

  • 400query missing or empty, topK out of range.
  • 401Missing or invalid key.
  • 403scope_deniedThe key lacks the knowledge scope.
  • 429rate_limitedMore than 60 requests per minute.

Notes

page is null when the source has no pages (text, Word, spreadsheets). rerankScore is null when reranking did not apply. trace.tokensUsed is the internal cost of expansion and reranking; it is not a question count.

Reading the response

FieldMeaning
matches[].contentThe passage text, at most about 2000 characters. Quote it, do not paraphrase it, when you cite.
matches[].sourceId, sourceNameThe document it comes from. Use sourceId to link back to GET /v1/knowledge-sources/:id.
matches[].section, pageWhere in the document: nearest heading, and 1-based page for PDFs and scans.
matches[].collectionId, dossierNameThe collection (client file) of the source, or null when unfiled.
matches[].chunkIdStable id of the passage. This is what proof certificates commit to.
matches[].similarityBest cosine similarity of the passage, 0 to 1. A raw signal, not a probability: do not set a universal threshold on it.
matches[].scoreThe fused RRF score. Only meaningful for comparing passages inside this response.
matches[].matchedByvector, fts, or both. A passage matched by both is usually the safest evidence.
matches[].rerankScoreThe model relevance score when reranking applied, else null.
trace.queriesThe sub-queries actually searched (one when expansion is off or the question is simple).
trace.vectorBackendpgvector or memory; the engine that served the semantic search.
trace.candidateCountNumber of candidates before the final cut.
trace.rerankAppliedWhether the model order was used. false also covers the case where its verdict was discarded by the guards.
trace.timingsMilliseconds spent expanding, searching and reranking.

Build a grounded answer

The pattern: retrieve, number the passages, instruct the model to answer only from them and to say so when they do not cover the question, then map the [n] markers in the answer back to sources. This is the recipe Light applies server-side.

If you send the final call through the Gateway, you get a signed certificate for it. If you would rather not assemble the prompt at all, use model: "pulse-auto+cortex" and the Gateway does the retrieval and records the passages in the certificate.

Retrieve, then answer strictly from the passages
const API = 'https://api.subsidia.protypa.fr'
const headers = {
Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY,
'Content-Type': 'application/json',
}
async function answer(question: string) {
const { matches } = await fetch(API + '/v1/knowledge/query', {
method: 'POST',
headers,
body: JSON.stringify({ query: question, topK: 6 }),
}).then((r) => r.json())
if (matches.length === 0) return { answer: 'The documents do not cover this.', sources: [] }
const passages = matches
.map((m: any, i: number) => '[' + (i + 1) + '] ' + m.sourceName + (m.page ? ' p.' + m.page : '') + '\n' + m.content)
.join('\n\n')
const res = await fetch(API + '/v1/chat/completions', {
method: 'POST',
headers,
body: JSON.stringify({
model: 'pulse-auto',
messages: [
{
role: 'system',
content:
'Answer using ONLY the numbered passages below. Cite each claim as [n]. ' +
'If the passages do not contain the answer, reply exactly: The documents do not cover this.\n\n' + passages,
},
{ role: 'user', content: question },
],
}),
})
const proofId = res.headers.get('x-pulse-proof')
const completion = await res.json()
return { answer: completion.choices[0].message.content, sources: matches, proofId }
}

Tips

SituationDo
Single-fact lookup (a date, an amount, a name)expand: false and rerank: false for the lowest latency. The full-text channel handles exact values well.
Compound or vague questionKeep expand: true; ask the whole question rather than keywords.
Lists ("all deadlines", "every client")Raise topK (10 to 20). The query detects enumeration questions and widens its candidate pool, and tries to surface the table that lists the answer.
One client or one teamLink the sources to an agent and pass its agentId, or file them in a collection and use Light with dossierId.
Empty matchesSay so to the user. Do not fall back to a model answer without telling them it is not from their documents.
Fresh uploads not foundCheck the source is ready. Passages of a processing source are not searchable yet.

Frequently asked

Does this endpoint consume questions?

It returns passages, not an answer. Expansion and reranking use the workspace model internally (see trace.tokensUsed); setting expand: false and rerank: false means no generative model is called (only the embedding of your question).

Is there a similarity threshold to drop weak matches?

Not in the API. similarity depends on the language and the embedding model, so a fixed threshold is fragile. Prefer the combination of matchedBy containing both channels and a high rerankScore, or let Light decide with its strict grounded prompt.

How do I cite a page?

Use sourceName, page (when not null) and section. Link sourceId back to your own record of the document.

Related