Authentication

Bearer tokens, key rotation, and the unauthenticated catalog.

Learn

The gateway uses bearer tokens for the API and leaves a small public surface open for browsing.

Bearer tokens

The public model-discovery routes /v1/models, /v1/catalog, and /v1/catalog/* do not require an API key. Protected routes such as /v1/chat/completions, /v1/embeddings, and /v1/quota require a bearer token prefixed with gw_. Create the token in your dashboard and put it in the Authorization header.

Authorization: Bearer gw_<your-key>

Failing to include the header, or sending a malformed/rejected key, returns a 401 with the standard error envelope:

{
  "error": {
    "message": "Invalid API key",
    "type": "authentication_error",
    "code": null
  }
}

Authentication failures currently have no specific error.code. Branch on the HTTP 401 status or type: "authentication_error", not on invalid_api_key.

Storing your key

Treat the key as a secret:

  • Don't commit it to a repo. Use environment variables (process.env.GATEWAY_KEY) or your secrets manager.
  • Don't paste it into chat. If you suspect it's leaked, revoke it in the signed-in client portal or ask your gateway contact to do so.
  • Don't share it across teams. One key per consumer means the rate-limit and credit balance apply to one workload, not many.

A CI script in this repo greps the public docs and admin code for real-looking gw_… keys and fails the build if it finds any. If you're authoring docs that include keys, use a placeholder — never a real value.

Rotation

Keys rotate by issue-and-revoke:

  1. Create a new personal key in the signed-in client portal, or ask your gateway contact to issue one.
  2. Deploy it (env var, secrets manager, however you ship config).
  3. Confirm traffic is using the new key.
  4. Revoke the old key in the client portal, or ask your gateway contact to revoke it.

There's no expiry by default. Rotation cadence is up to you — quarterly is typical for production keys, more frequently for sensitive use cases.

Per-key request rate-limit

Each key carries its own request rate-limit configured by your gateway contact. When you exceed it, the API returns HTTP 429 with type: "rate_limit_exceeded" and code: null:

{
  "error": {
    "message": "Rate limit exceeded",
    "type": "rate_limit_exceeded",
    "code": null
  }
}

Branch on the HTTP 429 status or type, not on error.code. Back off with exponential jitter; if you regularly hit the limit, ask your gateway contact to raise it (or use a separate key for the noisy workload). The full retry strategy lives on Rate limits & quotas.

Public model discovery

The /v1/models, /v1/catalog, and /v1/catalog/* endpoints are anonymous. Anyone can call them without an Authorization header to see which models the gateway exposes — useful for building model pickers, dashboards, or docs. Protected inference and quota routes still require the gw_ bearer token.

bash
curl https://api.toygate.store/v1/catalog
bash
curl https://api.toygate.store/v1/catalog/gpt-5.5

The discovery endpoints are IP-rate-limited as a noise filter (60 req/minute, recovers at 1/sec) and responses are heavily cacheable, so they're cheap to embed.

Self-service key management

The signed-in client portal can create and revoke the current user's own keys. It uses the session-protected POST /client/keys and DELETE /client/keys/:id routes. These are cookie-authenticated, self-scoped portal endpoints — not public OpenAI-compatible endpoints, and not APIs that accept a gw_ bearer key.

What you cannot do

  • Use a gw_ key to reach /admin/* or /client/*. Those routes use session-cookie authentication for their respective SPAs.
  • Use a session cookie to call /v1/* from a server. Use a gw_ key.
  • Treat /client/keys as a public key-management API. Its create and revoke operations require a signed-in user session and are limited to that user's own keys.