Streaming

Server-Sent Events in the standard chat-completion-chunk format, with timeouts and idle aborts.

Learn

Pass \"stream\": true in the chat-completion request body to receive Server-Sent Events in the standard chunk format. The gateway translates upstream chunks transparently \u2014 whether the upstream speaks its own chunk format or Anthropic's event: … envelopes, the client sees identical chunks.

Request

bash
curl -N -X POST https://api.toygate.store/v1/chat/completions \-H 'Authorization: Bearer gw_<your-key>' \-H 'Content-Type: application/json' \-d '{ "model": "gpt-5.5", "stream": true, "messages": [{ "role": "user", "content": "Stream me a sentence." }] }'

-N (--no-buffer) disables curl's output buffering so chunks appear as they arrive.

Response format

Each event is a data: line followed by JSON, separated by a blank line. The stream ends with data: [DONE].

data: {"id":"chatcmpl-1","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-1","choices":[{"index":0,"delta":{"content":"Hi"},"finish_reason":null}]}
data: {"id":"chatcmpl-1","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]}
data: {"id":"chatcmpl-1","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":4,"completion_tokens":2,"total_tokens":6}}
data: [DONE]

The role-establishing first chunk

The gateway always emits a first chunk with delta.role = "assistant" and empty content, even when the upstream doesn't. Strict clients (some CLI tools) drop messages whose first delta doesn't establish the role, so the gateway injects it before relaying anything from the upstream.

Usage is last-seen, not guaranteed

When the upstream reports usage, the gateway keeps the latest counts seen in a standard streaming usage field or Anthropic stop / message_delta event. A normal stream may therefore include usage in a trailing chunk, but clients must not assume that every stream has an authoritative final usage value. If the upstream fails before any usage event arrives, the internal request log records zero tokens for that stream.

Failures before and after the stream opens

Failure handling depends on whether the HTTP response has already become an SSE stream:

  • Before streaming starts: an unavailable route or open upstream breaker returns a regular JSON error with HTTP 503 and error.code = "upstream_unavailable". A transport failure before response headers returns the same code with HTTP 502.
  • After HTTP 200 and SSE output have started: the status can no longer change. An upstream read, parse, buffer, or idle failure produces a synthetic chunk with finish_reason: "error", followed by data: [DONE]. The gateway records the request internally with status 502 and the last usage values it saw.

Timeouts

Three layers protect against a stuck upstream or client connection:

TimerDefaultBehaviour
Upstream response deadline240 s (UPSTREAM_TIMEOUT_MS)Aborts while waiting for upstream response headers
Idle abort120 s (STREAM_IDLE_TIMEOUT_MS; production: 90 s)Aborts the open SSE stream if no upstream chunk arrives for the configured interval
Bun socket idle timeout255 sKeeps the client-bound socket open while the gateway is writing the stream

UPSTREAM_TIMEOUT_MS measures time until upstream response headers, not time to connect. That distinction decides how you size it: a non-streaming call receives no headers until the model has finished generating, so the value must cover the whole generation. A budget shorter than the model's generation time aborts every slow call and returns HTTP 504 with error.code = "upstream_timeout" and a message naming the gateway's own budget; the gateway neither retries it nor counts it against the circuit breaker. Accepted range is 60–240 s, capped by the 255 s socket timeout above. A non-streaming call that needs longer than the budget cannot be served at all — the answer is to stream, which is what the 504 message tells the caller.

For streaming calls the deadline is disarmed as soon as successful response headers arrive, and STREAM_IDLE_TIMEOUT_MS governs stalls between chunks from then on — which is why a stream may legitimately run for minutes while a non-streaming call on the same model cannot. An idle timeout happens mid-stream: it emits finish_reason: "error" and [DONE], then records internal status 502. If you need longer than 240 s, stream.

Backpressure and cancellation

If the client disconnects mid-stream, the gateway cancels the upstream reader and finalises the internal request log with status 499 and the last-seen usage. There is no error response or synthetic error chunk to send because the client connection is already gone.

SDK examples

TypeScript

ts
import { OpenAI as StandardClient } from "openai";// Or import from any other standard-compatible SDKconst client = new StandardClient({apiKey: process.env.GATEWAY_KEY,baseURL: "https://api.toygate.store/v1",});

Python

python
from openai import OpenAI as StandardClient# Or import from any other standard-compatible SDKclient = StandardClient(  api_key=os.environ["GATEWAY_KEY"],  base_url="https://api.toygate.store/v1",)

When not to stream

Streaming is the wrong default if you need the full response before doing anything (e.g. the response is JSON you need to parse, or you're computing a single embedding). Non-stream requests benefit from the full upstream-validation pass (upstream_invalid_response errors catch malformed JSON early) and avoid the per-chunk parsing overhead in the gateway.