text/event-stream.
Request body
| Field | Type | Required | Description |
|---|---|---|---|
query | string | Yes | The question. Max 4,000 characters |
thread_id | string | No | Continue a conversation. Use the thread_id from an earlier answer |
top_k | number | No | How many passages to retrieve. Default 10, capped at 100 |
group_ids | string[] | No | Answer only from these source groups |
user | { id?, email? } | No | Who is asking, for analytics. id is used when both are sent |
widget_id | string | No | A saved widget’s id. The answer uses that widget’s voice. The request must also pass that widget’s Enabled switch and Allowed origins |
Response stream
Each event is onedata: line holding a JSON object. There are no event: names.
| Frame | Fields | What to do |
|---|---|---|
| First | citations, done: false | The pages found for this question. Show them while the answer streams. Absent when nothing relevant was found |
| Middle | delta, done: false | The next piece of answer text. Append it |
| Last | done: true, thread_id, citations, message_id, answer | See below |
citationsis the list the answer actually cited. It replaces the first list. It can be empty.thread_idcontinues the conversation. Send it back asthread_id.message_idis what you rate.message_idandthread_idare both absent when the workspace uses Zero data retention.answeris only sent when the[n]markers were renumbered. When it is there, show it instead of the text you built fromdelta. See Renumbered answers.
index, title, url, document_id, chunk_id, heading_path (section breadcrumb) and excerpt. An [n] in the answer points at the citation whose index is n.
Renumbered answers
The answer usually cites only some of the pages it was shown. Kelu then renumbers the cited sources from 1 and rewrites the markers to match. The text has already streamed by then, so the fixed text comes on the last frame asanswer.
Reading the stream
Continue a conversation
Sendthread_id in the body, or post to the thread:
user is ignored. widget_id works as on the first question: send it again with each follow-up to keep the widget’s voice and rules.
- A thread from another knowledge base, or one that no longer exists, returns
404. - A
thread_idthat is not a valid id is ignored, and a new conversation starts. - With Zero data retention, conversations are not stored, so no
thread_idis returned. Each question starts a new conversation.
Errors
Refused before the answer starts: you get a normal JSON error with a4xx status. For example 403 for a bad group_ids or a missing CAPTCHA token.
With widget_id, the widget’s own rules can refuse the question too:
| Status | code | Meaning |
|---|---|---|
403 | WIDGET_DISABLED | The widget’s Enabled switch is off |
403 | ORIGIN_NOT_ALLOWED | The page’s origin is not in the widget’s Allowed origins |
404 | NOT_FOUND | No such widget, or it belongs to another knowledge base |
400 | BAD_REQUEST | widget_id is not a valid id |
200, so you get one data: frame with the error, and then the stream ends. No done frame follows.
error_code | Meaning | Retry? |
|---|---|---|
quota_exceeded | The workspace used its monthly question allowance, or the AI provider account is out of credit | No |
misconfigured | The AI provider is not set up correctly | No |
rate_limited | The AI provider is busy | Yes |
timeout, unavailable, internal | Temporary failure | Yes |
content_filtered, truncated, cancelled | The answer was blocked, cut short, or cancelled | Yes |
error is written for the reader, so you can show it as is.
Good to know
- Restricted documents are never used.
- When the knowledge base does not cover the question, the answer says so instead of guessing.
- The WebSocket transport streams the same answers over one connection. The SDK uses it by default.