You send a message. The UI sits still until the model is done. That wait feels broken.
Stream the reply. Tokens show up as the model writes them.
Call chat(). Then wrap the result with toServerSentEventsResponse:
import {
chat,
chatParamsFromRequest,
toServerSentEventsResponse,
} from "@tanstack/ai";
import { openaiText } from "@tanstack/ai-openai";
export async function POST(request: Request) {
const { messages, threadId, runId } = await chatParamsFromRequest(request);
const stream = chat({
adapter: openaiText("gpt-5.6"),
messages,
threadId,
runId,
});
return toServerSentEventsResponse(stream);
}chatParamsFromRequest reads the AG-UI body that useChat sends. If the body is invalid, it throws a Response with status 400. If your framework does not map a thrown Response to HTTP 400, catch it and return it.
The response stream pauses chat() when its output queue is full. It reads more chunks when the client consumes the queued data. This also applies to toHttpResponse with NDJSON. A response with durability does not pause, so its run keeps writing to the log.
import { useState } from "react";
import { useChat, fetchServerSentEvents } from "@tanstack/ai-react";
export function Chat() {
const [input, setInput] = useState("");
const { messages, sendMessage, isLoading, stop } = useChat({
connection: fetchServerSentEvents("/api/chat"),
});
return (
<>
{messages.map((message) => (
<div key={message.id}>
{message.parts.map((part, index) =>
part.type === "text" ? <p key={index}>{part.content}</p> : null,
)}
</div>
))}
<form
onSubmit={(event) => {
event.preventDefault();
if (input.trim() === "") {
return;
}
sendMessage(input);
setInput("");
}}
>
<input
value={input}
onChange={(event) => setInput(event.target.value)}
/>
{isLoading ? (
<button type="button" onClick={stop}>
Stop
</button>
) : (
<button type="submit">Send</button>
)}
</form>
</>
);
}messages updates as chunks arrive. isLoading is true while the run is in flight.
The shared ChatClient processes ready chunks in order without inserting a task between each chunk. It yields after bounded processing work to keep the main thread responsive.
The same pattern works in every UI framework. See Quick Start.
If SSE is blocked, pick another transport on Connection Adapters.
Call stop(). The client aborts the fetch.
Pass the same AbortController to chat() and toServerSentEventsResponse so the server stops the model too:
import {
chat,
chatParamsFromRequest,
toServerSentEventsResponse,
} from "@tanstack/ai";
import { openaiText } from "@tanstack/ai-openai";
export async function POST(request: Request) {
const { messages, threadId, runId } = await chatParamsFromRequest(request);
const abortController = new AbortController();
const stream = chat({
adapter: openaiText("gpt-5.6"),
messages,
threadId,
runId,
abortController,
});
return toServerSentEventsResponse(stream, { abortController });
}AbortError from stop() is expected. Pending client-tool work for that turn does not resume. A later addToolResult() for that turn is ignored.
A dropped connection mid-line throws StreamTruncatedError. The client then moves to error. See Connection Adapters.
For OpenAI and OpenRouter Responses, a stream that ends without response.completed emits RUN_ERROR with code incomplete-stream. This applies to chat and to structured output. Text received before the error remains available. The run does not call onFinish.
Chat Completions adapters built on @tanstack/openai-base also check completion in chat(). A started stream needs a non-null finish_reason or a final usage-only chunk (choices: [] with usage). Otherwise, the adapter closes open lifecycles and emits RUN_ERROR with code incomplete-stream. Partial text remains available. The run calls onError, skips pending server tools, and does not call onFinish.
The usage-only exception preserves provider compatibility. Usage on an earlier chunk or alongside a choice does not satisfy it. The OpenAI SDK hides [DONE], so this check cannot verify that marker. A stream with only [DONE] after content still needs a finish reason or usage-only tail to succeed.
Send a message. Text grows in the UI as tokens arrive.