> ## Documentation Index
> Fetch the complete documentation index at: https://cortex-e852fafe-t3code-rewrite-docs-declutter.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Knowledge

> Give your agent the documents and app records it answers from.

export const Field = ({name, type, required, recommended}) => {
  const label = required ? 'required' : recommended ? 'recommended' : null;
  const typeLabel = typeof type === 'string' ? type : null;
  const ariaParts = [name, typeLabel && `${typeLabel}`, label].filter(Boolean);
  return <span aria-label={ariaParts.join(', ')} className={label ? 'field-wrap has-field-tip' : 'field-wrap'} style={{
    position: 'relative',
    cursor: label ? 'default' : undefined
  }} tabIndex={label ? 0 : undefined}>
      <span className="field-name-row">
        <code>{name}</code>
        {required && <span className="field-req"> *</span>}
        {recommended && <span className="field-rec"> ●</span>}
      </span>
      {type && <span className="field-type">{type}</span>}
      {label && <span className="field-tip" role="tooltip">
          {label}
        </span>}
    </span>;
};

<a id="1-what-it-is" />

Knowledge is the material your agent searches to answer questions: product docs, policy PDFs, runbooks, wiki pages, Slack threads, tickets, and emails. You send it to [`POST /context/ingest`](/api-reference/v2/endpoint/ingest-context) with `type=knowledge` in one of two forms: **files** that HydraDB reads for you, or **app sources**, records your application already parsed from an app. One uploaded file or app record is a **source**. HydraDB splits its text into searchable passages called **chunks** and returns them through [`POST /query`](/api-reference/v2/endpoint/query).

What one user prefers or did is a [Memory](/essentials/v2/memories), not knowledge. Knowledge can still be private: store it in a collection you choose and restrict it with [Access Control](/essentials/v2/access-control).

This guide builds up in order:

1. [Upload files](#1-upload-files): one request, shown with two PDFs.
2. [Ingest app sources](#2-ingest-app-sources): records from Slack, Jira, Notion, Gmail, and other apps.
3. [Wait for indexing](#3-wait-for-indexing): polling until a source is searchable.
4. [Retrieve knowledge](#4-retrieve-knowledge): querying what you ingested.
5. [Add metadata for filtering](#5-add-metadata-for-filtering).
6. [Link documents that belong together](#6-link-documents-that-belong-together).
7. [Field reference](#field-reference) and [common mistakes](#common-mistakes).

<a id="path-a-files" />

## 1. Upload files

The example below uploads two PDFs. HydraDB parses each file (PDF, DOCX, Markdown, CSV, or TXT), splits it into chunks, and indexes it.

The request is **multipart/form-data** with these fields:

| Form field | What to send |
| - | - |
| `type` | `knowledge` |
| `database` | An existing [database](/api-reference/v2/endpoint/create-tenant) |
| `collection` | Optional. Leave it out to use the database's default collection. A query that names another collection does not read the default one. |
| `documents` | One or more files |
| `document_metadata` | One item per file, in the same order, as a JSON string. Optional, but this is where a file's `id` and `metadata` go. |
| `upsert` | Leave it out. It defaults to `true`, so a file sent again with the same `id` replaces the saved one. |

<Accordion title="Upload files: Python, TypeScript, or cURL">
  Set up your [SDK client](/api-reference/v2/sdks#client-setup) before running the Python or TypeScript example.

  <CodeGroup>
    ```python Python SDK theme={"dark"}
    import json

    with open("/path/to/policy.pdf", "rb") as f1, open("/path/to/runbook.pdf", "rb") as f2:
        result = client.context.ingest(
            type="knowledge",
            database="acme_corp",
            documents=[
                ("policy.pdf", f1, "application/pdf"),
                ("runbook.pdf", f2, "application/pdf"),
            ],
            document_metadata=json.dumps([
                {"id": "policy_v2", "metadata": {"document_type": "policy", "status": "approved"}},
                {"id": "runbook_deploy", "metadata": {"document_type": "runbook"}},
            ]),
        )
    ```

    ```typescript TypeScript SDK theme={"dark"}
    const result = await client.context.ingest({
      type: "knowledge",
      database: "acme_corp",
      documents: [
        { path: "/path/to/policy.pdf", filename: "policy.pdf", contentType: "application/pdf" },
        { path: "/path/to/runbook.pdf", filename: "runbook.pdf", contentType: "application/pdf" },
      ],
      documentMetadata: JSON.stringify([
        { id: "policy_v2", metadata: { document_type: "policy", status: "approved" } },
        { id: "runbook_deploy", metadata: { document_type: "runbook" } },
      ]),
    });
    ```

    ```bash cURL theme={"dark"}
    curl -X POST 'https://api.hydradb.com/context/ingest' \
      -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
      -H "API-Version: 2" \
      -F "type=knowledge" \
      -F "database=acme_corp" \
      -F "documents=@/path/to/policy.pdf" \
      -F "documents=@/path/to/runbook.pdf" \
      -F 'document_metadata=[
        { "id": "policy_v2", "metadata": { "document_type": "policy", "status": "approved" } },
        { "id": "runbook_deploy", "metadata": { "document_type": "runbook" } }
      ]'
    ```
  </CodeGroup>

  The response lists each file in `data.results` with its `id` and a `status` of `queued`. The payload is wrapped in the standard envelope, so read it from `response.data`:

  ```json theme={"dark"}
  {
    "success": true,
    "data": {
      "success": true,
      "message": "Knowledge uploaded successfully",
      "results": [
        { "id": "policy_v2", "filename": "policy.pdf", "status": "queued", "error": null }
      ],
      "success_count": 1,
      "failed_count": 0
    },
    "error": null,
    "meta": {
      "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d",
      "latency_ms": 12.3
    }
  }
  ```
</Accordion>

### What each property in `document_metadata` means

| Property | What it is |
| - | - |
| `id` | Your stable ID for the file. Send the same `id` later to replace the file, or use it to check status and delete. Leave it out to have one generated; the response returns it either way. |
| `metadata` | Values you filter on at query time, such as `document_type`. Declare the keys in the [database metadata schema](/essentials/v2/metadata). Up to 16 KiB as compact JSON. |
| `additional_metadata` | Free-form values for display and bookkeeping. Up to 1 KiB as compact JSON. |
| `relations` | Other sources this file belongs with, as `{ "ids": [...] }`. [Section 6](#6-link-documents-that-belong-together) explains it. |

The key list is closed. Any other key, including `title`, `type`, `url`, and `timestamp`, returns `400`.

## 2. Ingest app sources

When your application already has an item's text and fields, such as a Slack message, a Jira ticket, a Notion page, or a CRM record, send it as an **app source** instead of a file. HydraDB keeps the app's structure: the author, the thread, the ticket's status, and the links between items. [App Sources](/essentials/v2/app-sources) is the full guide, with the classes and a ready-to-adapt pattern for each app.

The request is the same endpoint, with `app_knowledge` in place of `documents`. You can send both in one request.

<Accordion title="Ingest a Slack message: Python, TypeScript, or cURL">
  <CodeGroup>
    ```python Python SDK theme={"dark"}
    import json

    result = client.context.ingest(
        type="knowledge",
        database="acme_corp",
        app_knowledge=json.dumps([
            {
                "id": "slack_C01_1716213600_000100",
                "title": "Pricing discussion - Slack #product",
                "kind": "message",
                "provider": "slack",
                "external_id": "1716213600.000100",
                "fields": {
                    "kind": "message",
                    "body": "We agreed on three tiers: Starter at $29, Pro at $79, Enterprise at $199.",
                    "author": "alice",
                    "thread_id": "1716213600.000100",
                    "created_at": "2026-05-20T10:00:00Z",
                },
                "metadata": {"channel": "product", "status": "finalized"},
                "additional_metadata": {"slack_ts": "1716213600.000100"},
            }
        ]),
    )
    ```

    ```typescript TypeScript SDK theme={"dark"}
    const result = await client.context.ingest({
      type: "knowledge",
      database: "acme_corp",
      appKnowledge: JSON.stringify([
        {
          id: "slack_C01_1716213600_000100",
          title: "Pricing discussion - Slack #product",
          kind: "message",
          provider: "slack",
          external_id: "1716213600.000100",
          fields: {
            kind: "message",
            body: "We agreed on three tiers: Starter at $29, Pro at $79, Enterprise at $199.",
            author: "alice",
            thread_id: "1716213600.000100",
            created_at: "2026-05-20T10:00:00Z",
          },
          metadata: { channel: "product", status: "finalized" },
          additional_metadata: { slack_ts: "1716213600.000100" },
        },
      ]),
    });
    ```

    ```bash cURL theme={"dark"}
    curl -X POST 'https://api.hydradb.com/context/ingest' \
      -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
      -H "API-Version: 2" \
      -F "type=knowledge" \
      -F "database=acme_corp" \
      -F 'app_knowledge=[
        {
          "id": "slack_C01_1716213600_000100",
          "title": "Pricing discussion - Slack #product",
          "kind": "message",
          "provider": "slack",
          "external_id": "1716213600.000100",
          "fields": {
            "kind": "message",
            "body": "We agreed on three tiers: Starter at $29, Pro at $79, Enterprise at $199.",
            "author": "alice",
            "thread_id": "1716213600.000100",
            "created_at": "2026-05-20T10:00:00Z"
          },
          "metadata": { "channel": "product", "status": "finalized" },
          "additional_metadata": { "slack_ts": "1716213600.000100" }
        }
      ]'
    ```
  </CodeGroup>

  The response has the same shape as a file upload, with one entry per item in `data.results`.
</Accordion>

<a id="5-verifying-processing" />

## 3. Wait for indexing

Ingestion runs in the background: HydraDB parses each source, splits it into chunks, finds the people, teams, and topics it mentions and how they relate (see [Context Graphs](/essentials/v2/context-graphs)), and builds the search indexes. The response returns `202 Accepted` with each source's `id`. Poll [`GET /context/status`](/api-reference/v2/endpoint/source-status) with those ids and the same `collection` until every source reaches `completed` or `errored`.

<Accordion title="Poll until indexed: Python, TypeScript, or cURL">
  <CodeGroup>
    ```python Python SDK theme={"dark"}
    import time

    ids = ["policy_v2", "runbook_deploy"]
    terminal = {"completed", "errored"}

    while True:
        status = client.context.status(
            database="acme_corp",
            collection="default",
            ids=ids,
        )
        if all(s.indexing_status in terminal for s in status.data.statuses):
            break
        time.sleep(5)
    ```

    ```typescript TypeScript SDK theme={"dark"}
    const ids = ["policy_v2", "runbook_deploy"];
    const terminal = new Set(["completed", "errored"]);

    let status;
    while (true) {
      status = await client.context.status({
        database: "acme_corp",
        collection: "default",
        ids: ids,
      });

      const allDone = status.data.statuses.every((s) => terminal.has(s.indexingStatus));
      if (allDone) break;

      await new Promise((r) => setTimeout(r, 5_000));
    }
    ```

    ```bash cURL theme={"dark"}
    while true; do
      response=$(curl -s -G 'https://api.hydradb.com/context/status' \
        -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
        -H "API-Version: 2" \
        --data-urlencode "database=acme_corp" \
        --data-urlencode "collection=default" \
        --data-urlencode "ids=policy_v2" \
        --data-urlencode "ids=runbook_deploy")

      statuses=$(echo "$response" | jq -r '.data.statuses[].indexing_status')
      if echo "$statuses" | grep -qv '^\(completed\|errored\)$'; then
        sleep 5
        continue
      fi
      echo "$response" | jq
      break
    done
    ```
  </CodeGroup>

  Each status in the `data` payload looks like this:

  ```json theme={"dark"}
  {
    "statuses": [
      {
        "id": "policy_v2",
        "indexing_status": "completed",
        "success": true
      }
    ]
  }
  ```
</Accordion>

Status values move forward through `queued`, `processing`, `graph_creation`, and `completed`, or land in `errored` on failure. A source is already searchable at `graph_creation`; wait for `completed` only when you need full graph traversal (`graph_context: true`). Poll every few seconds; most documents complete in 1 to 5 minutes. The full table is on [Ingestion Status](/api-reference/v2/endpoint/source-status). If you would rather not poll, [register a webhook](/essentials/v2/webhooks).

## 4. Retrieve knowledge

Search with [`POST /query`](/api-reference/v2/endpoint/query), using `type: "knowledge"`. Use `type: "all"` to search a user's memories in the same call; both stores are read from the same collection, so if the documents live in their own collection, list it next to the user's in `collections`. App sources are searched too, because `query_apps` defaults to `true`.

<Accordion title="Query knowledge: Python, TypeScript, or cURL">
  <CodeGroup>
    ```python Python SDK theme={"dark"}
    result = client.query(
        database="acme_corp",
        query="How do I rotate API keys?",
        type="knowledge",
    )

    for chunk in result.data.chunks:
        print(chunk.source_title, chunk.chunk_content)
    ```

    ```typescript TypeScript SDK theme={"dark"}
    const result = await client.query({
      database: "acme_corp",
      query: "How do I rotate API keys?",
      type: "knowledge",
    });

    for (const chunk of result.data.chunks) {
      console.log(chunk.sourceTitle, chunk.chunkContent);
    }
    ```

    ```bash cURL theme={"dark"}
    curl -X POST 'https://api.hydradb.com/query' \
      -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
      -H "API-Version: 2" \
      -H "Content-Type: application/json" \
      -d '{
        "database": "acme_corp",
        "query": "How do I rotate API keys?",
        "type": "knowledge"
      }'
    ```
  </CodeGroup>

  Matching chunks come back in `data.chunks`, ranked by relevance, each with its `source_title` and text. [How to Use API Results](/essentials/v2/api-results) shows how to turn the response into a prompt. [Query](/essentials/v2/query) covers `mode`, `alpha`, and `recency_bias`.
</Accordion>

<a id="8-metadata-on-knowledge" />

## 5. Add metadata for filtering

[Metadata](/essentials/v2/metadata) is how you bridge structured filters and semantic search. Attach `metadata` at ingestion, and narrow a query to matching sources with `metadata_filters`.

<Accordion title="Filter by metadata: ingest, then query">
  On the `document_metadata` item or app source:

  ```json theme={"dark"}
  {
    "metadata": { "document_type": "runbook", "environment": "production" }
  }
  ```

  Then at query time:

  ```json theme={"dark"}
  {
    "query": "How do I rotate API keys?",
    "metadata_filters": { "document_type": "runbook", "environment": "production" }
  }
  ```
</Accordion>

Declare the `metadata` keys you filter on in the database metadata schema. In a database with a schema, ingest rejects undeclared keys, so filtering on one returns nothing. `additional_metadata` is stored alongside the source for display and bookkeeping; to filter on it, nest the keys under `additional_metadata` in `metadata_filters`, for example `{"metadata_filters": {"additional_metadata": {"author": "Alice"}}}`. `document_metadata` is accepted only as a legacy alias for that nested filter namespace.

See [Metadata](/essentials/v2/metadata) for schema design, filter examples, and the `metadata` vs `additional_metadata` decision rule. See [Create Database](/api-reference/v2/endpoint/create-tenant) for the schema field configuration (`enable_dense_embedding`, `enable_sparse_embedding`).

<a id="7-forceful-relations" />

## 6. Link documents that belong together

Forceful relations let you **declare** links between Knowledge sources at ingestion time. When a query matches a source that has forceful relations attached, HydraDB pulls the linked sources into the `additional_context` field of the query response alongside the primary chunks.

This is separate from the entity and relationship [context graph](/essentials/v2/context-graphs) HydraDB builds automatically. Graph extraction is *derived* from content; forceful relations are *declared* by you, which is useful when you know two documents belong together (a contract and its addendum, a runbook and its troubleshooting guide) but the text does not make the connection explicit. Forceful relations link whole **sources**; to declare the full entity and relation graph *within* a single document (replacing extraction for it), use [Bring Your Own Graph](/essentials/v2/bring-your-own-graph).

On a file, set `relations.ids` on its `document_metadata` item. On an app source, set `relations.ids` the same way, and name the relationship in `app_relations` so the link also carries a predicate. `relations.ids` is what makes the linked text come back with query results; a named relationship alone connects the items without requesting their text. [App Sources](/essentials/v2/app-sources#4-relate-messages-tickets-and-records) covers typed targets in full.

<Accordion title="Link a runbook to its troubleshooting guide: Python, TypeScript, or cURL">
  The Notion guide below links to a troubleshooting page you already ingested with the HydraDB `id` `auth-troubleshooting`. Ingest that page first and wait for `completed`.

  <CodeGroup>
    ```python Python SDK theme={"dark"}
    client.context.ingest(
        type="knowledge",
        database="acme_corp",
        app_knowledge=json.dumps([
            {
                "id": "auth-guide",
                "title": "Authentication Guide",
                "kind": "knowledge_base",
                "provider": "notion",
                "external_id": "AUTH-GUIDE",
                "fields": {"kind": "knowledge_base", "title": "Authentication Guide", "body": "..."},
                "relations": {"ids": ["auth-troubleshooting"]},
                "app_relations": [
                    {"predicate": "linked_to", "target": {"source_id": "auth-troubleshooting"}},
                ],
            }
        ]),
    )
    ```

    ```typescript TypeScript SDK theme={"dark"}
    await client.context.ingest({
      type: "knowledge",
      database: "acme_corp",
      appKnowledge: JSON.stringify([
        {
          id: "auth-guide",
          title: "Authentication Guide",
          kind: "knowledge_base",
          provider: "notion",
          external_id: "AUTH-GUIDE",
          fields: { kind: "knowledge_base", title: "Authentication Guide", body: "..." },
          relations: { ids: ["auth-troubleshooting"] },
          app_relations: [
            { predicate: "linked_to", target: { source_id: "auth-troubleshooting" } },
          ],
        },
      ]),
    });
    ```

    ```bash cURL theme={"dark"}
    curl -X POST 'https://api.hydradb.com/context/ingest' \
      -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
      -H "API-Version: 2" \
      -F "type=knowledge" \
      -F "database=acme_corp" \
      -F 'app_knowledge=[
        {
          "id": "auth-guide",
          "title": "Authentication Guide",
          "kind": "knowledge_base",
          "provider": "notion",
          "external_id": "AUTH-GUIDE",
          "fields": { "kind": "knowledge_base", "title": "Authentication Guide", "body": "..." },
          "relations": { "ids": ["auth-troubleshooting"] },
          "app_relations": [
            { "predicate": "linked_to", "target": { "source_id": "auth-troubleshooting" } }
          ]
        }
      ]'
    ```
  </CodeGroup>
</Accordion>

At query time, linked sources are returned in the `additional_context` field of the [query response](/api-reference/v2/endpoint/query). `query_forceful_relations` defaults to `true`, so you only need to set it explicitly if you want to disable forceful-relation expansion.

<Warning>
  **Knowledge items, thinking mode only.** Forceful relations work within the same store, so pointing a Knowledge item at a Memory ID silently returns nothing. Linked context is fetched only when `mode: "thinking"` is set on the query, or when `mode: "auto"` routes to thinking. In `fast` mode, `query_forceful_relations` is ignored even if set to `true`.
</Warning>

## Field reference

<AccordionGroup>
  <Accordion title="Request fields">
    | Field | Default | Description |
    | - | - | - |
    | <Field name="type" type="knowledge" required /> | None | Selects the Knowledge store. |
    | <Field name="database" type="string" required /> | None | Target database. |
    | <Field name="collection" type="string" /> | default collection | Scope this knowledge to a collection. If omitted, knowledge goes to the database's default collection, which a query naming another `collection` does not read. |
    | <Field name="documents" type="file[]" /> | None | Files for HydraDB to parse: PDF, DOCX, Markdown, CSV, TXT. |
    | <Field name="document_metadata" type="string (JSON array)" /> | None | One item per file in `documents`, in the same order. The item count must match the file count. |
    | <Field name="app_knowledge" type="string (JSON array)" /> | None | App source items. See the [App Sources field reference](/essentials/v2/app-sources#field-reference). |
    | <Field name="upsert" type="boolean" /> | `true` | Replace existing items with the same `id`. |
  </Accordion>

  <Accordion title="document_metadata item fields">
    | Field | Description |
    | - | - |
    | <Field name="id" type="string" /> | Optional stable identifier. Used as upsert key when `upsert: true`. |
    | <Field name="metadata" type="object" /> | Filterable fields. Keys must match the database metadata schema. Up to 16 KiB as compact JSON. |
    | <Field name="additional_metadata" type="object" /> | Free-form fields for display and bookkeeping. Filterable only when nested under `additional_metadata` in `metadata_filters`. Up to 1 KiB as compact JSON. |
    | <Field name="relations" type="object" /> | Forcefully connect this source to others with `{ "ids": [...] }`. |

    Any other key, including `title`, `type`, `url`, and `timestamp`, returns `400`. App sources accept `title`, `url`, and `timestamp`.
  </Accordion>
</AccordionGroup>

## Common mistakes

| Mistake | What goes wrong | Fix |
| - | - | - |
| Querying Knowledge with `type: "memory"` | Nothing returned: wrong store | Use `type: "knowledge"` (or `"all"`) on [`POST /query`](/api-reference/v2/endpoint/query). |
| Using `additional_metadata` for hot filterable fields | Filters work only with nested syntax and are less efficient than schema-backed metadata filters | Put hot filter keys in `metadata` and declare them in the [database schema](/api-reference/v2/endpoint/create-tenant), or filter under `metadata_filters: {"additional_metadata": {...}}` for occasional/free-form fields. See [Metadata](/essentials/v2/metadata). |
| Not polling [`/context/status`](/api-reference/v2/endpoint/source-status) before querying | Chunks aren't indexed yet, so query results are empty | Wait for `indexing_status: completed` (or `graph_creation` for partial availability). |
| Sending `document_metadata` as a JSON object instead of a JSON string | `multipart/form-data` validation fails | Stringify `document_metadata` before sending in the `-F` field. |
| Omitting `database` | `400` error (`database is required`) | Required on every call. |

## Related

* [App Sources](/essentials/v2/app-sources): the full guide to ingesting records from apps
* [Memories](/essentials/v2/memories): the per-user counterpart to Knowledge
* [Query](/essentials/v2/query): how [`POST /query`](/api-reference/v2/endpoint/query) retrieves Knowledge
* [Metadata](/essentials/v2/metadata): designing filterable schemas
* [Context Graphs](/essentials/v2/context-graphs): how the graph layer enriches retrieval
* [Ingest Content: API Reference](/api-reference/v2/endpoint/ingest-context)
* [Ingestion Status: API Reference](/api-reference/v2/endpoint/source-status)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.