Quickstart
Send your first chat-completion through the gateway in under five minutes.
The shortest route from your key to a working chat completion. About five minutes if you already know curl.
Before you start
You need:
- A gateway base URL — your gateway's address. For local development it's
<GatewayUrl/>. - An API key starting with
gw_— create one in the API keys section. Treat it like any secret (don't commit it, don't paste it into chat). - A model slug — see the models page.
gpt-5.5is a safe default.
Every request you send debits AI credits from your account balance. Before doing volume work, review all three controls on the rate limits page.
Send a request
The gateway speaks the standard chat/completions shape on /v1. Drop your existing client in front of it; the only change is the baseURL and the key.
cURL
TypeScript
Python
What you should see
A standard chat-completion response. The model field echoes the slug you asked for, choices[0] contains the assistant turn, and usage reports prompt + completion tokens.
The gateway hid the underlying provider from you — your code never had to care which upstream actually ran the request.
Where to go next
- Key concepts — the model slug, what the gateway hides, the bits worth knowing.
- Authentication — bearer tokens, rotation, the public catalog.
- Streaming — Server-Sent Events, idle timeouts, cancellation.
- Errors — the error envelope and stable codes.
- Rate limits & quotas — request rate, concurrency, AI credit balance, and what to do when a limit trips.
Self-host
ToyGate is open source, and you can run it yourself. The production compose file is docker-compose.prod.yml at the repo root. Once you're up, the ops runbooks under docs/how-to/ops/ cover backup-restore, alerting, certificate renewal, and incident response.