Authentication
Bearer tokens, key rotation, and the unauthenticated catalog.
The gateway uses bearer tokens for the API and leaves a small public surface open for browsing.
Bearer tokens
The public model-discovery routes /v1/models, /v1/catalog, and /v1/catalog/* do not require an API key. Protected routes such as /v1/chat/completions, /v1/embeddings, and /v1/quota require a bearer token prefixed with gw_. Create the token in your dashboard and put it in the Authorization header.
Failing to include the header, or sending a malformed/rejected key, returns a 401 with the standard error envelope:
Authentication failures currently have no specific error.code. Branch on the HTTP 401 status or type: "authentication_error", not on invalid_api_key.
Storing your key
Treat the key as a secret:
- Don't commit it to a repo. Use environment variables (
process.env.GATEWAY_KEY) or your secrets manager. - Don't paste it into chat. If you suspect it's leaked, revoke it in the signed-in client portal or ask your gateway contact to do so.
- Don't share it across teams. One key per consumer means the rate-limit and credit balance apply to one workload, not many.
A CI script in this repo greps the public docs and admin code for real-looking gw_… keys and fails the build if it finds any. If you're authoring docs that include keys, use a placeholder — never a real value.
Rotation
Keys rotate by issue-and-revoke:
- Create a new personal key in the signed-in client portal, or ask your gateway contact to issue one.
- Deploy it (env var, secrets manager, however you ship config).
- Confirm traffic is using the new key.
- Revoke the old key in the client portal, or ask your gateway contact to revoke it.
There's no expiry by default. Rotation cadence is up to you — quarterly is typical for production keys, more frequently for sensitive use cases.
Per-key request rate-limit
Each key carries its own request rate-limit configured by your gateway contact. When you exceed it, the API returns HTTP 429 with type: "rate_limit_exceeded" and code: null:
Branch on the HTTP 429 status or type, not on error.code. Back off with exponential jitter; if you regularly hit the limit, ask your gateway contact to raise it (or use a separate key for the noisy workload). The full retry strategy lives on Rate limits & quotas.
Public model discovery
The /v1/models, /v1/catalog, and /v1/catalog/* endpoints are anonymous. Anyone can call them without an Authorization header to see which models the gateway exposes — useful for building model pickers, dashboards, or docs. Protected inference and quota routes still require the gw_ bearer token.
The discovery endpoints are IP-rate-limited as a noise filter (60 req/minute, recovers at 1/sec) and responses are heavily cacheable, so they're cheap to embed.
Self-service key management
The signed-in client portal can create and revoke the current user's own keys. It uses the session-protected POST /client/keys and DELETE /client/keys/:id routes. These are cookie-authenticated, self-scoped portal endpoints — not public OpenAI-compatible endpoints, and not APIs that accept a gw_ bearer key.
What you cannot do
- Use a
gw_key to reach/admin/*or/client/*. Those routes use session-cookie authentication for their respective SPAs. - Use a session cookie to call
/v1/*from a server. Use agw_key. - Treat
/client/keysas a public key-management API. Its create and revoke operations require a signed-in user session and are limited to that user's own keys.