> ## Documentation Index
> Fetch the complete documentation index at: https://cortex-e852fafe-t3code-rewrite-docs-declutter.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Semantic Search

> Find passages that mean the same thing as the question, and still match exact codes and names.

Semantic search matches by meaning: a question about "time off" finds a document about "annual leave." Keyword search matches the words themselves, which is what you need for an error code such as `E_AUTH_429`. [`POST /query`](/api-reference/v2/endpoint/query) runs both together by default, a combination called **hybrid search**, so one request handles paraphrases and exact identifiers. BM25 is the keyword-ranking method it uses.

This page explains how the two signals combine and when to change the blend. [Query](/essentials/v2/query) covers the request itself and every other parameter.

This guide builds up in order:

1. [Choose the matching method](#1-choose-the-matching-method): `query_by`, and the three signals a query can use.
2. [Balance meaning and keywords](#2-balance-meaning-and-keywords-with-alpha): the `alpha` value.
3. [Where meaning alone falls short](#3-where-meaning-alone-falls-short): why HydraDB has no pure vector mode.

## 1. Choose the matching method

`query_by` picks how the query text is matched:

* **`"hybrid"`** (the default): matching by meaning and BM25 keyword matching, blended by `alpha`. Use it unless you have a reason not to.
* **`"text"`**: BM25 only. Use it when the literal wording is the point, such as a legal clause, a SKU, or a function name. `operator` says how the terms must match: `"or"` (the default), `"and"`, or `"phrase"`. `"and"` and `"phrase"` are valid only with `query_by: "text"`.

A hybrid query can also draw on a third signal, the [context graph](/essentials/v2/context-graphs), which is returned with the results when `graph_context` is on (the default):

| Signal | Finds | Best for |
| - | - | - |
| Semantic | Text with similar meaning | Natural-language questions, paraphrases, conceptual lookup |
| BM25 keyword | Text with matching tokens | Error codes, identifiers, names, SKUs, exact phrases |
| Graph | Related entities and relationships | Multi-hop questions, dependencies, ownership, project context |

<Accordion title="Hybrid query and exact phrase query: Python, TypeScript, or cURL">
  The first request is a hybrid search kept inside the `phoenix` project with a metadata filter. The second returns only passages that contain the exact phrase "mechanical engineer".

  <CodeGroup>
    ```python Python SDK theme={"dark"}
    result = client.query(
        database="acme",
        collection="team-mobile",
        query="How do we rotate API keys?",
        type="knowledge",
        query_by="hybrid",
        max_results=8,
        alpha=0.8,
        recency_bias=0.2,
        graph_context=True,
        metadata_filters={"project": "phoenix"},
    )
    ```

    ```typescript TypeScript SDK theme={"dark"}
    const result = await client.query({
      database: "acme",
      collection: "team-mobile",
      query: "How do we rotate API keys?",
      type: "knowledge",
      queryBy: "hybrid",
      maxResults: 8,
      alpha: 0.8,
      recencyBias: 0.2,
      graphContext: true,
      metadataFilters: { project: "phoenix" },
    });
    ```

    ```bash cURL theme={"dark"}
    curl -X POST 'https://api.hydradb.com/query' \
      -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
      -H "API-Version: 2" \
      -H "Content-Type: application/json" \
      -d '{
        "database": "acme",
        "collection": "team-mobile",
        "query": "How do we rotate API keys?",
        "type": "knowledge",
        "query_by": "hybrid",
        "max_results": 8,
        "alpha": 0.8,
        "recency_bias": 0.2,
        "graph_context": true,
        "metadata_filters": {
          "project": "phoenix"
        }
      }'
    ```
  </CodeGroup>

  `metadata_filters` are exact constraints that run before ranking and are re-checked after the matching passages are loaded, so use them whenever the query has a scope that must not be violated. Top-level keys match `metadata` and support `equals`, `contains`, and `contains_any` (no range or fuzzy match); nest under `additional_metadata` to filter free-form per-document fields. See [Metadata](/essentials/v2/metadata).

  <CodeGroup>
    ```python Python SDK theme={"dark"}
    result = client.query(
        database="acme",
        query="mechanical engineer",
        type="knowledge",
        query_by="text",
        operator="phrase",
        max_results=10,
    )
    ```

    ```typescript TypeScript SDK theme={"dark"}
    const result = await client.query({
      database: "acme",
      query: "mechanical engineer",
      type: "knowledge",
      queryBy: "text",
      operator: "phrase",
      maxResults: 10,
    });
    ```

    ```bash cURL theme={"dark"}
    curl -X POST 'https://api.hydradb.com/query' \
      -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
      -H "API-Version: 2" \
      -H "Content-Type: application/json" \
      -d '{
        "database": "acme",
        "query": "mechanical engineer",
        "type": "knowledge",
        "query_by": "text",
        "operator": "phrase",
        "max_results": 10
      }'
    ```
  </CodeGroup>
</Accordion>

## 2. Balance meaning and keywords with `alpha`

`alpha` sets the blend in a hybrid query. Higher values lean on meaning; lower values lean on keywords.

| `alpha` | Behavior | Use When |
| - | - | - |
| `1.0` | Semantic-heavy | Conceptual questions and paraphrases |
| `0.8` | Default semantic-leaning hybrid | Most agent query workflows |
| `0.5` | Balanced | Mixed natural language and exact terms |
| `0.2` | Keyword-leaning | Queries with product names, error strings, or IDs |
| `0.0` | Keyword only | Debugging exact-match behavior |

Start with the default (`0.8`; `"auto"` also resolves to `0.8`) and tune from observed results. If users query for exact IDs and get loosely related content, lower `alpha`. If they ask broad conceptual questions and get sparse results, raise it. `alpha` applies only to `query_by: "hybrid"`; it is ignored for `"text"`.

<Accordion title="Technical lookup with a low alpha: Python, TypeScript, or cURL">
  Lower `alpha` when names, IDs, and literal strings matter but the query still has a natural-language part.

  <CodeGroup>
    ```python Python SDK theme={"dark"}
    result = client.query(
        database="acme",
        query="TimeoutError in payments-worker v4.2.1",
        type="knowledge",
        query_by="hybrid",
        max_results=5,
        alpha=0.3,
        recency_bias=0.4,
    )
    ```

    ```typescript TypeScript SDK theme={"dark"}
    const result = await client.query({
      database: "acme",
      query: "TimeoutError in payments-worker v4.2.1",
      type: "knowledge",
      queryBy: "hybrid",
      maxResults: 5,
      alpha: 0.3,
      recencyBias: 0.4,
    });
    ```

    ```bash cURL theme={"dark"}
    curl -X POST 'https://api.hydradb.com/query' \
      -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
      -H "API-Version: 2" \
      -H "Content-Type: application/json" \
      -d '{
        "database": "acme",
        "query": "TimeoutError in payments-worker v4.2.1",
        "type": "knowledge",
        "query_by": "hybrid",
        "max_results": 5,
        "alpha": 0.3,
        "recency_bias": 0.4
      }'
    ```
  </CodeGroup>
</Accordion>

## 3. Where meaning alone falls short

Pure vector search misses things that matter in production:

* Exact identifiers such as `E_AUTH_429` or `payments-worker-v4` may be generalized away.
* A project name can collide with a normal word, like `strawberry` the project and strawberry the fruit.
* Old and new documents can look equally relevant without recency or metadata signals.
* Different users can need different context for the same query.
* Relationship questions need graph context, not only similar text chunks.

That is why HydraDB exposes semantic retrieval through `query_by: "hybrid"` inside the one `/query` endpoint rather than as a separate pure-vector mode. The keyword, recency, metadata, scope, and graph signals ride along with it.

## Related

* [Query](/essentials/v2/query): the request, the recipes, and the full parameter reference
* [Context Graphs](/essentials/v2/context-graphs): the third signal, and how to read it
* [Metadata](/essentials/v2/metadata): designing filterable schemas
* [How to Use API Results](/essentials/v2/api-results): turning the response into a model prompt


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.