# Querying the knowledge base

> POST /v1/knowledge/query runs hybrid retrieval with reranking and returns cited passages plus a trace of why they matched.

`POST /v1/knowledge/query` returns the passages of your documents that best answer a question. It does **not** write an answer: you get the evidence and decide what to do with it, which is what you want when you build your own grounded prompt, a citation panel or a verification step.

For a ready-made answer with citations or an honest refusal, see [Light](https://dev.subsidia.protypa.fr/docs/light.md). For retrieval inside a normal chat call, add `+cortex` to the model address in the [Gateway](https://dev.subsidia.protypa.fr/docs/gateway.md).

## The retrieval pipeline

**Every query goes through the same stages. Query expansion and reranking use the workspace model and can be switched off per call.**

Flow: Question -> Query expansion -> Vector search + full-text search -> RRF fusion -> Rerank (guarded) -> topK passages + trace

1. **Expand**

   A compound question ("What is the notice period and who signs?") is decomposed into sub-queries, so each part can find its own passages. Disable with `expand: false`.

2. **Search twice per sub-query**

   A semantic (embedding) search and a full-text search run in parallel. Meaning finds paraphrases; the full-text channel finds exact terms, numbers and names that embeddings blur. A passage found by both ranks higher.

3. **Fuse**

   The rankings are merged with reciprocal rank fusion (RRF), which rewards agreement between channels without needing comparable scores.

4. **Rerank, with guards**

   A small pool of the best candidates is re-scored by the model for relevance. If the judge finds nothing relevant, or contradicts the consensus ranking wildly, its verdict is discarded and a deterministic order is used (passages sharing words with the question first). A bad judge therefore cannot corrupt the result. If the pool is no larger than `topK`, reranking is skipped since it could not change anything.

5. **Return**

   The `topK` best passages come back with their source, section, page and scores, and a `trace` that tells you what happened.

> **INFO: What the query respects**
> Results are always limited to your workspace. Collections that a workspace administrator has closed to you are excluded from the search itself, not filtered afterwards.

### POST /v1/knowledge/query

Retrieve the passages that best match a question.

Returns up to `topK` passages from the sources of the workspace, best first, and a retrieval trace. An empty `matches` array is a normal answer meaning the documents contain nothing relevant; treat it as an honest "not found" and do not let a model fill the gap.

By default the whole workspace is searched. Pass `agentId` to restrict the search to the sources linked to one agent (see [Knowledge sources](https://dev.subsidia.protypa.fr/docs/knowledge-sources.md)).

- **Authentication:** API key (Bearer or x-api-key)
- **Scopes:** `knowledge`

#### Request body

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `query` | `string` | yes |  | The question or search phrase. Must not be empty. A full question works better than keywords. |
| `topK` | `number` | no | `5` | How many passages to return, 1 to 20. |
| `agentId` | `string` | no |  | Restrict retrieval to the sources linked to this agent. |
| `expand` | `boolean` | no | `true` | Decompose compound questions into sub-queries. Set `false` for short, single-fact lookups to save a model call and latency. |
| `rerank` | `boolean` | no | `true` | Let the model re-score the fused candidates. Set `false` for the fastest, cheapest path (embeddings and full-text only). |

#### Request examples

_curl_

```bash
curl https://api.subsidia.protypa.fr/v1/knowledge/query \
  -H "Authorization: Bearer $SUBSIDIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query": "What is the notice period to terminate the lease?", "topK": 5}'
```

_TypeScript_

```typescript
const res = await fetch('https://api.subsidia.protypa.fr/v1/knowledge/query', {
  method: 'POST',
  headers: {
    Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({ query: 'What is the notice period to terminate the lease?', topK: 5 }),
})
const { matches, trace } = await res.json()
for (const m of matches) console.log(m.sourceName, m.page, m.content.slice(0, 80))
```

_Python_

```python
import os, requests

res = requests.post(
    "https://api.subsidia.protypa.fr/v1/knowledge/query",
    headers={"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]},
    json={"query": "What is the notice period to terminate the lease?", "topK": 5},
)
data = res.json()
for m in data["matches"]:
    print(m["sourceName"], m["page"], m["content"][:80])
```

#### Responses

**200**: The matches and the trace. `matches` may be empty.

```json
{
  "matches": [
    {
      "sourceId": "clx9k2m4p0002abcd",
      "sourceName": "contrat-bail-2025.pdf",
      "collectionId": "col_123",
      "dossierName": "Dupont SARL",
      "chunkId": "ck_8f31a0c2",
      "section": "Article 12 - Termination",
      "page": 6,
      "content": "The tenant may terminate the lease at the end of each three-year period by giving six months' notice by registered letter.",
      "similarity": 0.81,
      "score": 0.0328,
      "matchedBy": ["vector", "fts"],
      "rerankScore": 9
    }
  ],
  "trace": {
    "queries": ["What is the notice period to terminate the lease?"],
    "vectorBackend": "pgvector",
    "candidateCount": 14,
    "rerankApplied": true,
    "timings": { "expandMs": 0, "searchMs": 41, "rerankMs": 612 },
    "tokensUsed": 380
  }
}
```

**400**: Missing or empty `query`, or a `topK` outside 1-20 (schema validation).

#### Errors

| Status | Code | When |
| --- | --- | --- |
| 400 |  | `query` missing or empty, `topK` out of range. |
| 401 |  | Missing or invalid key. |
| 403 | `scope_denied` | The key lacks the `knowledge` scope. |
| 429 | `rate_limited` | More than 60 requests per minute. |

#### Notes

`page` is `null` when the source has no pages (text, Word, spreadsheets). `rerankScore` is `null` when reranking did not apply. `trace.tokensUsed` is the internal cost of expansion and reranking; it is not a question count.

## Reading the response

| Field | Meaning |
| --- | --- |
| `matches[].content` | The passage text, at most about 2000 characters. Quote it, do not paraphrase it, when you cite. |
| `matches[].sourceId`, `sourceName` | The document it comes from. Use `sourceId` to link back to [`GET /v1/knowledge-sources/:id`](https://dev.subsidia.protypa.fr/docs/knowledge-sources.md). |
| `matches[].section`, `page` | Where in the document: nearest heading, and 1-based page for PDFs and scans. |
| `matches[].collectionId`, `dossierName` | The collection (client file) of the source, or `null` when unfiled. |
| `matches[].chunkId` | Stable id of the passage. This is what [proof certificates](https://dev.subsidia.protypa.fr/docs/proofs.md) commit to. |
| `matches[].similarity` | Best cosine similarity of the passage, 0 to 1. A raw signal, not a probability: do not set a universal threshold on it. |
| `matches[].score` | The fused RRF score. Only meaningful for comparing passages inside this response. |
| `matches[].matchedBy` | `vector`, `fts`, or both. A passage matched by both is usually the safest evidence. |
| `matches[].rerankScore` | The model relevance score when reranking applied, else `null`. |
| `trace.queries` | The sub-queries actually searched (one when expansion is off or the question is simple). |
| `trace.vectorBackend` | `pgvector` or `memory`; the engine that served the semantic search. |
| `trace.candidateCount` | Number of candidates before the final cut. |
| `trace.rerankApplied` | Whether the model order was used. `false` also covers the case where its verdict was discarded by the guards. |
| `trace.timings` | Milliseconds spent expanding, searching and reranking. |

## Build a grounded answer

The pattern: retrieve, number the passages, instruct the model to answer only from them and to say so when they do not cover the question, then map the `[n]` markers in the answer back to sources. This is the recipe Light applies server-side.

If you send the final call through the [Gateway](https://dev.subsidia.protypa.fr/docs/gateway.md), you get a signed certificate for it. If you would rather not assemble the prompt at all, use `model: "pulse-auto+cortex"` and the Gateway does the retrieval and records the passages in the certificate.

**Retrieve, then answer strictly from the passages**

_TypeScript_

```typescript
const API = 'https://api.subsidia.protypa.fr'
const headers = {
  Authorization: 'Bearer ' + process.env.SUBSIDIA_API_KEY,
  'Content-Type': 'application/json',
}

async function answer(question: string) {
  const { matches } = await fetch(API + '/v1/knowledge/query', {
    method: 'POST',
    headers,
    body: JSON.stringify({ query: question, topK: 6 }),
  }).then((r) => r.json())

  if (matches.length === 0) return { answer: 'The documents do not cover this.', sources: [] }

  const passages = matches
    .map((m: any, i: number) => '[' + (i + 1) + '] ' + m.sourceName + (m.page ? ' p.' + m.page : '') + '\n' + m.content)
    .join('\n\n')

  const res = await fetch(API + '/v1/chat/completions', {
    method: 'POST',
    headers,
    body: JSON.stringify({
      model: 'pulse-auto',
      messages: [
        {
          role: 'system',
          content:
            'Answer using ONLY the numbered passages below. Cite each claim as [n]. ' +
            'If the passages do not contain the answer, reply exactly: The documents do not cover this.\n\n' + passages,
        },
        { role: 'user', content: question },
      ],
    }),
  })
  const proofId = res.headers.get('x-pulse-proof')
  const completion = await res.json()
  return { answer: completion.choices[0].message.content, sources: matches, proofId }
}
```

_Python_

```python
import os, requests

API = "https://api.subsidia.protypa.fr"
HEADERS = {"Authorization": "Bearer " + os.environ["SUBSIDIA_API_KEY"]}

def answer(question):
    matches = requests.post(
        API + "/v1/knowledge/query", headers=HEADERS, json={"query": question, "topK": 6}
    ).json()["matches"]
    if not matches:
        return {"answer": "The documents do not cover this.", "sources": []}

    passages = "\n\n".join(
        "[%d] %s%s\n%s" % (i + 1, m["sourceName"], " p.%s" % m["page"] if m["page"] else "", m["content"])
        for i, m in enumerate(matches)
    )
    res = requests.post(
        API + "/v1/chat/completions",
        headers=HEADERS,
        json={
            "model": "pulse-auto",
            "messages": [
                {
                    "role": "system",
                    "content": "Answer using ONLY the numbered passages below. Cite each claim as [n]. "
                    "If the passages do not contain the answer, reply exactly: The documents do not cover this.\n\n" + passages,
                },
                {"role": "user", "content": question},
            ],
        },
    )
    return {
        "answer": res.json()["choices"][0]["message"]["content"],
        "sources": matches,
        "proof": res.headers.get("x-pulse-proof"),
    }
```

> **TIP: Check the citations**
> A model can cite a passage number that does not exist, or attach a claim to the wrong passage. Reject any `[n]` outside `1..matches.length`, and run the final text through [Le Vérificateur](https://dev.subsidia.protypa.fr/docs/verify.md) when the answer matters: it gives every claim a verdict against the same documents, with no model involved.

## Tips

| Situation | Do |
| --- | --- |
| Single-fact lookup (a date, an amount, a name) | `expand: false` and `rerank: false` for the lowest latency. The full-text channel handles exact values well. |
| Compound or vague question | Keep `expand: true`; ask the whole question rather than keywords. |
| Lists ("all deadlines", "every client") | Raise `topK` (10 to 20). The query detects enumeration questions and widens its candidate pool, and tries to surface the table that lists the answer. |
| One client or one team | Link the sources to an agent and pass its `agentId`, or file them in a collection and use [Light](https://dev.subsidia.protypa.fr/docs/light.md) with `dossierId`. |
| Empty `matches` | Say so to the user. Do not fall back to a model answer without telling them it is not from their documents. |
| Fresh uploads not found | Check the source is `ready`. Passages of a `processing` source are not searchable yet. |

## Frequently asked

**Does this endpoint consume questions?**

It returns passages, not an answer. Expansion and reranking use the workspace model internally (see `trace.tokensUsed`); setting `expand: false` and `rerank: false` means no generative model is called (only the embedding of your question).

**Is there a similarity threshold to drop weak matches?**

Not in the API. `similarity` depends on the language and the embedding model, so a fixed threshold is fragile. Prefer the combination of `matchedBy` containing both channels and a high `rerankScore`, or let [Light](https://dev.subsidia.protypa.fr/docs/light.md) decide with its strict grounded prompt.

**How do I cite a page?**

Use `sourceName`, `page` (when not null) and `section`. Link `sourceId` back to your own record of the document.
