Skip to main content
POST /query retrieves Knowledge (documents and app sources), Memories (user preferences, conversation history, inferred content), or both. The response contains matching passages (chunks) and their source details, which your application can put in a model prompt. Search combines meaning, keywords, and relationships between the content. It does not generate the answer. For the full request and response schema, see Query: API Reference.

1. Use-case recipes

Pick the row that matches your goal and use the parameters as a starting point:

2. Parameter reference

When to send database and collection: Multi-tenant. The older names tenant_id and sub_tenant_id still work as deprecated aliases.

Retrieval

Scope

Graph

Shaping results

Access control


3. Tuning heuristics

Most of the time the defaults are right. When they aren’t, here’s where to start:
  • mode: Pick "fast" or "thinking" explicitly when you need predictable retrieval behavior instead of automatic routing.
  • alpha: Start at 0.8. Lower toward 0.3 to 0.5 when the query contains literal tokens (error codes, SKUs, product names). Raise toward 0.9 for conceptual questions.
  • max_results: Start at 10. Drop to 5 for tight context windows; raise to 20 if you rerank downstream.
  • additional_context: Use it when the query alone is ambiguous. Keep it short and factual.
  • graph_context: Keep it on (the default) when answers depend on entity relationships (multi-hop questions, “how does X relate to Y”). Pair with mode: "thinking", because in "fast" mode the graph slice is shallow.
  • query_apps: Leave it on for app data (Slack, Gmail, Jira, and so on), and pair it with mode: "thinking" so relations and threads expand.

4. Minimal working example

A personalized-answer flow takes a single call: POST /query with type: "all" returns merged knowledge and per-user memory in one ranked result set. Both stores are read from the same scope, so if shared knowledge lives in its own collection, list it next to the user’s in collections.

Setup

One call: Knowledge and Memories together

Merge into the LLM prompt

The response is a single RetrievalResult containing chunks[], sources[], and, when applicable, graph_context and additional_context. Chunks from Knowledge and Memories are already interleaved and ranked by relevance, so no manual merging is required. Pass the result through the helper in How to Use API Results to turn it into a context string for your prompt:
If you need to query several users, teams, or workspaces at once, pass collections. A list gives every scope equal normalized weight; an object applies relative ranking weights with at most one decimal place before the final merged ranking. When max_results is set, it caps the final merged response across all selected collections:
If you need to keep Knowledge and Memories formatted differently in the prompt, call /query twice in parallel with type: "knowledge" and type: "memory", then merge client-side.

Production checklist

  • Set per-call timeouts: Generous for thinking (3 to 5 s), tight for fast (500 ms or less). For mode: "auto", size the timeout for the thinking case, since it can resolve to either pipeline and defaults toward thinking when the routing signal is inconclusive.
  • Pass additional_context with known session state (page, feature, role). It sharpens retrieval without extra calls.

5. Common mistakes


6. Advanced patterns

Hybrid + text in two parallel calls. When a query mixes a literal token (error code, SKU, function name) with natural-language intent, run query_by: "hybrid" and query_by: "text" in parallel, dedupe by chunk ID, and treat text hits as a “must include” floor. Recall-then-rerank. Ask for more chunks than you actually need (max_results: 20) and apply your own reranker (recency windows, compliance filters, business rules) before picking the final top-k for the prompt. Cache the prompt context: Include the database, collection scopes, caller ACL, and all query parameters in the cache key. Expire cached results when content or permissions change.