Key concepts
Model slugs, the standard chat-completion surface, and what the gateway abstracts away.
A small mental model goes a long way. ToyGate looks like a standard API to your code, but under the hood it translates and proxies your request to whichever upstream provider runs that model. Understanding the four ideas below means the rest of the docs become obvious.
Model slugs route everything
A request's model field is a slug — a short, human-readable id the gateway owns. Examples: gpt-5.5, claude-opus-4-8. You'll see the full list on the models page.
Your client only ever talks about slugs. The gateway is responsible for resolving each slug to a real upstream model. If the gateway operator swaps the upstream behind gpt-5.5 — moves it to a different provider, upgrades it to a newer snapshot — your code doesn't change a line.
Standard API in, standard API out
The public API speaks the standard chat/completions protocol. Non-streaming responses, streaming responses, function calling, and vision requests for vision-capable models work against the gateway with only the baseURL and API key changed.
Even when the upstream provider speaks a different shape (e.g. Anthropic's /v1/messages), the gateway translates in both directions so the wire format you see is always the standard format.
The upstream is hidden by design
Nothing in the response tells you which provider actually ran the request:
- The
modelfield echoes the slug you asked for, not the upstream's internal id. - The
owned_byfield on the catalog is set by the gateway operator (e.g. "direct request" for self-hosted, or a partner's name). - Errors use stable codes chosen by the gateway. Raw upstream errors never reach your client.
This isolation isn't an accident — it lets the gateway operator swap providers without breaking your integration. Don't depend on inferring the upstream from response shape; it can change tomorrow.
What the gateway does not do
A few things are explicitly out of scope:
- Cache your responses. Every request reaches the upstream. If you want caching, do it in your client (or a CDN in front of the gateway for public read-only traffic).
- Mix providers per request. One slug → one upstream. If you need provider fallback for a single request, build it in your client across multiple slugs.
- Accept structured-output requests yet. Sending
response_formatreturns HTTP 400 witherror.code = "response_format_not_supported". Helpers such as Vercel AI SDKgenerateObjectuse that unsupported field. - Silently drop images.
image_urlparts are forwarded when the selected model advertises vision support. A non-vision model returns HTTP 400 witherror.code = "model_does_not_support_images".
If something feels surprising, one of these four ideas is usually the reason.