Streaming
Server-Sent Events in the standard chat-completion-chunk format, with timeouts and idle aborts.
Pass \"stream\": true in the chat-completion request body to receive Server-Sent Events in the standard chunk format. The gateway translates upstream chunks transparently \u2014 whether the upstream speaks its own chunk format or Anthropic's event: … envelopes, the client sees identical chunks.
Request
-N (--no-buffer) disables curl's output buffering so chunks appear as they arrive.
Response format
Each event is a data: line followed by JSON, separated by a blank line. The stream ends with data: [DONE].
The role-establishing first chunk
The gateway always emits a first chunk with delta.role = "assistant" and empty content, even when the upstream doesn't. Strict clients (some CLI tools) drop messages whose first delta doesn't establish the role, so the gateway injects it before relaying anything from the upstream.
Usage is last-seen, not guaranteed
When the upstream reports usage, the gateway keeps the latest counts seen in a standard streaming usage field or Anthropic stop / message_delta event. A normal stream may therefore include usage in a trailing chunk, but clients must not assume that every stream has an authoritative final usage value. If the upstream fails before any usage event arrives, the internal request log records zero tokens for that stream.
Failures before and after the stream opens
Failure handling depends on whether the HTTP response has already become an SSE stream:
- Before streaming starts: an unavailable route or open upstream breaker returns a regular JSON error with HTTP 503 and
error.code = "upstream_unavailable". A transport failure before response headers returns the same code with HTTP 502. - After HTTP 200 and SSE output have started: the status can no longer change. An upstream read, parse, buffer, or idle failure produces a synthetic chunk with
finish_reason: "error", followed bydata: [DONE]. The gateway records the request internally with status 502 and the last usage values it saw.
Timeouts
Three layers protect against a stuck upstream or client connection:
| Timer | Default | Behaviour |
|---|---|---|
| Upstream response deadline | 240 s (UPSTREAM_TIMEOUT_MS) | Aborts while waiting for upstream response headers |
| Idle abort | 120 s (STREAM_IDLE_TIMEOUT_MS; production: 90 s) | Aborts the open SSE stream if no upstream chunk arrives for the configured interval |
| Bun socket idle timeout | 255 s | Keeps the client-bound socket open while the gateway is writing the stream |
UPSTREAM_TIMEOUT_MS measures time until upstream response headers, not time to connect. That distinction decides how you size it: a non-streaming call receives no headers until the model has finished generating, so the value must cover the whole generation. A budget shorter than the model's generation time aborts every slow call and returns HTTP 504 with error.code = "upstream_timeout" and a message naming the gateway's own budget; the gateway neither retries it nor counts it against the circuit breaker. Accepted range is 60–240 s, capped by the 255 s socket timeout above. A non-streaming call that needs longer than the budget cannot be served at all — the answer is to stream, which is what the 504 message tells the caller.
For streaming calls the deadline is disarmed as soon as successful response headers arrive, and STREAM_IDLE_TIMEOUT_MS governs stalls between chunks from then on — which is why a stream may legitimately run for minutes while a non-streaming call on the same model cannot. An idle timeout happens mid-stream: it emits finish_reason: "error" and [DONE], then records internal status 502. If you need longer than 240 s, stream.
Backpressure and cancellation
If the client disconnects mid-stream, the gateway cancels the upstream reader and finalises the internal request log with status 499 and the last-seen usage. There is no error response or synthetic error chunk to send because the client connection is already gone.
SDK examples
TypeScript
Python
When not to stream
Streaming is the wrong default if you need the full response before doing anything (e.g. the response is JSON you need to parse, or you're computing a single embedding). Non-stream requests benefit from the full upstream-validation pass (upstream_invalid_response errors catch malformed JSON early) and avoid the per-chunk parsing overhead in the gateway.