Rate limits & quotas

Three independent controls — per-key request rate, per-key concurrency, and AI credit balance.

Learn

The gateway applies three independent controls: a per-key request rate-limit, a per-key concurrency limit, and a per-account AI credit balance. They cover different failure modes and use different error details and recovery behavior.

At a glance

ControlCountsScope429 errorRetry-AfterReleases or resets
Request rate-limitrequests in the configured windowone keytype rate_limit_exceeded, code null60 secondsAs the window advances
Concurrency limitrequests currently in flightone keytype insufficient_quota, code concurrency_limit1 secondWhen an in-flight request or stream finishes
AI credit balanceweighted AI credits (4 buckets)your account (across all keys)type insufficient_quota, code null60 secondsOnly when you purchase a credit package

Another 429 — upstream_rate_limited — has nothing to do with these controls. It's the upstream provider throttling the gateway itself. See Errors for the distinction.

Per-key request rate-limit

Every API key carries a rate-limit configured for your account. When you exceed it:

HTTP/1.1 429 Too Many Requests
Content-Type: application/json

{
  "error": { "message": "Rate limit exceeded", "type": "rate_limit_exceeded", "code": null }
}

What to do

  • Back off with exponential jitter (see the snippet in Errors).
  • If you regularly trip the limit, ask your gateway contact to raise it, or split the workload across multiple keys with distinct purposes.
  • Sharing one key across many processes? Either give each process its own key, or batch requests through a single broker in front of the gateway.

Per-key concurrency limit

Each API key has a cap on requests that are in flight at the same time. A streaming request holds its slot until the stream finishes. At the cap, OpenAI-compatible generation endpoints return HTTP 429 with type: "insufficient_quota", code: "concurrency_limit", and Retry-After: 1.

Queue work locally or lower parallelism before retrying. This limit clears as active requests finish; it does not consume or reset your AI credit balance.

Per-account AI credit balance

Your account has an AI credit balance. Every successful response debits weighted AI credits, calculated from four usage buckets: regular input, cache read, cache write, and output. The cost depends on the model, token type, and caching. When you run out of credits:

{
  "error": { "message": "AI credit quota exhausted", "type": "insufficient_quota", "code": null }
}

What to do

Purchase a credit package to top up your balance — blind retries won't help. The balance is persistent across days, keys, and processes. The current runtime still returns Retry-After: 60; treat that as a minimum delay, not a promise that your balance will automatically refill.

Use authenticated GET /v1/quota to read used, quota, remaining, and the credit status for the account linked to your key.

Why three controls

The request rate-limit protects against single-process bursts (a runaway script hammering one key). The concurrency limit bounds simultaneous upstream work and long-lived streams. The AI credit balance protects against long-tail cost growth (a user happily sending small requests for months until they've burned through their allocation).

All three are independent:

  • One massive request can blow through your AI credit balance without ever tripping the request rate-limit or concurrency limit (it's only one request).
  • A few slow streams can trip the concurrency limit while staying below the request-rate window.
  • Fast, short requests can trip the request rate-limit while concurrency remains low.

What does not count against these key/account controls

  • Public model discovery (/v1/models and /v1/catalog*). These endpoints have their own IP-based limiters (60 req/min) but don't count against your key or account credits.
  • The /health endpoint. Always answers, no auth, no limit.

Headers

The current OpenAI-compatible runtime sends Retry-After: 60 for the per-key request rate, Retry-After: 1 for per-key concurrency, and Retry-After: 60 for AI credit balance exhaustion. Honour the delay before retrying. For the credit balance, a delayed retry can still fail until you purchase a credit package.