AI Gateway Playground

Playground

Send a real governed call from the console and see exactly what it cost, how long it took, and the record it left — without writing any code.

The Playground is not a sandbox or a simulator — it is the same chokepoint your production traffic goes through. Auth, the budget gate, the gateway, metering and audit all run exactly as they will for your application.

Send a call

  1. Pick a model

    The Model field takes any model your workspace can reach. mock-gpt is a built-in test model priced at almost nothing — the right choice while you are proving a path works rather than evaluating output quality.

  2. Load an agent, optionally

    Load from agent pulls in an existing agent's model and system prompt, so you are testing the thing you actually deployed rather than an approximation of it.

  3. Write the prompt

    A system message is optional; the user message is the call.

  4. Generate

    Leave streaming on to watch tokens arrive. The budget gate still fires before the first token either way — streaming changes how you see the response, not when the call is authorised.

A governed call: the response streams in, then the cost receipt lands under it

Read the receipt

Every completion is followed by a line of numbers. This is the whole point of the page — the same call through a raw provider SDK tells you almost none of this:

FieldWhat it tells you
est. costWhat this call cost, priced from the token counts. Multiply by your expected volume before you ship.
modelWhich model actually answered — not necessarily the one you asked for, if routing or a fallback intervened.
prompt / completion / totalToken counts. A prompt that is larger than you expected is the usual cause of a bill that is larger than you expected.
TTFTTime to first token — what a user perceives as responsiveness.
latencyTotal wall-clock for the call.
Audit · TraceLinks straight to this call's audit entry and its trace. One click from "what did that do" to the record.

Compare two models

The Compare tab runs the same prompt against two models side by side, with a receipt under each. This is the cheapest way to answer "is the expensive model actually better for this prompt" — often it is not, and the comparison pays for itself immediately in per-call cost.

Both calls are governed and both are billed. Comparing is not free.

Watching the budget

The Monthly budget meter in the top right updates as you spend. Cross a threshold and the alert you configured fires here too — the Playground is not exempt from anything.

Push past the cap and you get the same 402 your application would, with the same body. That is worth doing once deliberately: see what happens when it is refused.

Next