AI Gateway Gateway

Gateway

The model-agnostic chokepoint: the model catalog, reliability routing, fallbacks and response caching, on one page. No client reaches a provider directly — every call is routed, cached, budget-gated, metered and audited here.

This page is read-only visibility into a policy your platform administrator manages — there is no edit control here.

Call the gateway

The gateway exposes an OpenAI-compatible endpoint. Point any OpenAI client at it and pass an sk_... key — no provider credential ever leaves your code.

curl -X POST https://api.forgebench.ai/v1/chat/completions \
-H "Authorization: Bearer sk_..." \
-H "Content-Type: application/json" \
-d '{
  "model": "mock-gpt",
  "messages": [{"role": "user", "content": "hello"}]
}'

Response is OpenAI-shaped plus one addition — trace_id — see Quickstart for the full body.

The model catalog

A flat, searchable list of everything a call can use: label, provider, whether it has a fallback chain, availability, and price per 1K tokens (or per image, for image models). Local and unpriced models read "free"; a model needing a key you haven't configured shows as unavailable, with a link to add one.

Two summary figures sit above the list: the total model count, and how many of them have a fallback chain configured — a quick answer to "how protected is this deployment" without reading every row.

Searching the catalog for a model — vendor filters and availability narrow it as you type
Searching the catalog for a model — vendor filters and availability narrow it as you type

Sort by name, price, context window or tier. Filter by vendor as the catalog grows — the chips are generated from whatever the deployment actually serves, not a fixed list.

Available models are usable in Playground and in an agent's own model choice, within whatever the calling key's allowlist permits.

Reliability routing and fallbacks

Every call is load-balanced across providers, retried on failure, and escalated through a fallback chain automatically — none of it is client-side retry logic you have to write.

SettingWhat it controls
Load balancingHow traffic is spread across providers serving the same model.
RetriesAttempts per call before failing over to a fallback.
TimeoutHow long a single attempt is given before it's treated as failed.
Circuit breakerFailures before a provider is temporarily taken out of rotation, and how long it stays out.

A model marked protected in the catalog has a configured fallback chain. There are two distinct triggers a chain can fire on — worth knowing apart, since they mean different things went wrong:

  • On failure — the primary model errored or is unavailable.
  • Over context limit — the request is too large for the primary model's context window.

Either way, the call escalates to the next model in the chain automatically.

Response cache

A cache hit skips the provider call entirely — but the request is still budgeted, metered and audited exactly as if it had gone through. Caching never becomes a hole in the record.

Isolation is per tenant, always — a cache entry is never shared across tenants. Retention is a fixed TTL set centrally.

Gateway vs Routing

Both pages sit at the same chokepoint and split one job in half. Gateway owns reliability routing — which provider instance serves a model, retried and failed-over automatically, invisible to the caller. Routing owns capability routing — which model a request should use in the first place, based on what the request needs. Configure a model's fallback chain here; configure which requests reach which model in Routing.

Next