Rate limits & quotas
Three independent controls — per-key request rate, per-key concurrency, and AI credit balance.
The gateway applies three independent controls: a per-key request rate-limit, a per-key concurrency limit, and a per-account AI credit balance. They cover different failure modes and use different error details and recovery behavior.
At a glance
| Control | Counts | Scope | 429 error | Retry-After | Releases or resets |
|---|---|---|---|---|---|
| Request rate-limit | requests in the configured window | one key | type rate_limit_exceeded, code null | 60 seconds | As the window advances |
| Concurrency limit | requests currently in flight | one key | type insufficient_quota, code concurrency_limit | 1 second | When an in-flight request or stream finishes |
| AI credit balance | weighted AI credits (4 buckets) | your account (across all keys) | type insufficient_quota, code null | 60 seconds | Only when you purchase a credit package |
Another 429 — upstream_rate_limited — has nothing to do with these controls. It's the upstream
provider throttling the gateway itself. See Errors for the
distinction.
Per-key request rate-limit
Every API key carries a rate-limit configured for your account. When you exceed it:
What to do
- Back off with exponential jitter (see the snippet in Errors).
- If you regularly trip the limit, ask your gateway contact to raise it, or split the workload across multiple keys with distinct purposes.
- Sharing one key across many processes? Either give each process its own key, or batch requests through a single broker in front of the gateway.
Per-key concurrency limit
Each API key has a cap on requests that are in flight at the same time. A streaming request
holds its slot until the stream finishes. At the cap, OpenAI-compatible generation endpoints
return HTTP 429 with type: "insufficient_quota", code: "concurrency_limit", and
Retry-After: 1.
Queue work locally or lower parallelism before retrying. This limit clears as active requests finish; it does not consume or reset your AI credit balance.
Per-account AI credit balance
Your account has an AI credit balance. Every successful response debits weighted AI credits, calculated from four usage buckets: regular input, cache read, cache write, and output. The cost depends on the model, token type, and caching. When you run out of credits:
What to do
Purchase a credit package to top up your balance — blind retries won't help. The balance
is persistent across days, keys, and processes. The current runtime still returns
Retry-After: 60; treat that as a minimum delay, not a promise that your balance will
automatically refill.
Use authenticated GET /v1/quota to read used, quota, remaining, and the credit status for the account linked to your key.
Why three controls
The request rate-limit protects against single-process bursts (a runaway script hammering one key). The concurrency limit bounds simultaneous upstream work and long-lived streams. The AI credit balance protects against long-tail cost growth (a user happily sending small requests for months until they've burned through their allocation).
All three are independent:
- One massive request can blow through your AI credit balance without ever tripping the request rate-limit or concurrency limit (it's only one request).
- A few slow streams can trip the concurrency limit while staying below the request-rate window.
- Fast, short requests can trip the request rate-limit while concurrency remains low.
What does not count against these key/account controls
- Public model discovery (
/v1/modelsand/v1/catalog*). These endpoints have their own IP-based limiters (60 req/min) but don't count against your key or account credits. - The
/healthendpoint. Always answers, no auth, no limit.
Headers
The current OpenAI-compatible runtime sends Retry-After: 60 for the per-key request rate,
Retry-After: 1 for per-key concurrency, and Retry-After: 60 for AI credit balance exhaustion.
Honour the delay before retrying. For the credit balance, a delayed retry can still fail until
you purchase a credit package.