Playground
Send a real governed call from the console and see exactly what it cost, how long it took, and the record it left — without writing any code.
The Playground is not a sandbox or a simulator — it is the same chokepoint your production traffic goes through. Auth, the budget gate, the gateway, metering and audit all run exactly as they will for your application.
Send a call
Pick a model
The Model field takes any model your workspace can reach.
mock-gptis a built-in test model priced at almost nothing — the right choice while you are proving a path works rather than evaluating output quality.Load an agent, optionally
Load from agent pulls in an existing agent's model and system prompt, so you are testing the thing you actually deployed rather than an approximation of it.
Write the prompt
A system message is optional; the user message is the call.
Generate
Leave streaming on to watch tokens arrive. The budget gate still fires before the first token either way — streaming changes how you see the response, not when the call is authorised.
Read the receipt
Every completion is followed by a line of numbers. This is the whole point of the page — the same call through a raw provider SDK tells you almost none of this:
| Field | What it tells you |
|---|---|
| est. cost | What this call cost, priced from the token counts. Multiply by your expected volume before you ship. |
| model | Which model actually answered — not necessarily the one you asked for, if routing or a fallback intervened. |
| prompt / completion / total | Token counts. A prompt that is larger than you expected is the usual cause of a bill that is larger than you expected. |
| TTFT | Time to first token — what a user perceives as responsiveness. |
| latency | Total wall-clock for the call. |
| Audit · Trace | Links straight to this call's audit entry and its trace. One click from "what did that do" to the record. |
Compare two models
The Compare tab runs the same prompt against two models side by side, with a receipt under each. This is the cheapest way to answer "is the expensive model actually better for this prompt" — often it is not, and the comparison pays for itself immediately in per-call cost.
Both calls are governed and both are billed. Comparing is not free.
Watching the budget
The Monthly budget meter in the top right updates as you spend. Cross a threshold and the alert you configured fires here too — the Playground is not exempt from anything.
Push past the cap and you get the same 402 your application would, with the
same body. That is worth doing once deliberately: see
what happens when it is refused.

