Python SDK
A thin client for the governed chokepoint. If you can call it with requests,
you can call it with this: the SDK exists to save you the boilerplate of
retries, typed errors and streaming, not to hide the API from you.
What this actually is, in plain terms
Forgebench sits in front of real model providers (OpenAI, Anthropic,
Gemini, ...). Instead of your code holding an OpenAI key and calling OpenAI
directly, it holds a Forgebench key (sk_...) and calls Forgebench, which
forwards the request to whichever model you asked for. The client object
below is your one doorway into that: every method on it is a different thing
you can ask Forgebench to do on your behalf.
Doing it this way, instead of just calling OpenAI yourself, is what buys you the governed chokepoint: budget checks before money is spent, a tamper-evident audit row for every call, and a provider credential your code never has to hold. See Core concepts for the vocabulary (Agent, Budget, Policy, etc.) this whole SDK is built around.
Install
pip install "git+https://github.com/seedlinglabs/forgebench-sdk.git#subdirectory=sdk-py"Requires Python 3.9+. The only runtime dependency is
httpx.
Connect to the gateway
Every call, chat, agents, runs, tools, rides through one control-plane
chokepoint: auth → tenant resolution (Postgres RLS) → budget pre-gate →
per-tenant provider key injection → the gateway call → one metering event +
one hash-chained audit row. You never send a provider credential
(api_key/api_base for OpenAI, Anthropic, etc.); the SDK never has one to
send, and the chokepoint strips a client-supplied one anyway.
api_key defaults to $FORGEBENCH_API_KEY; base_url defaults to
$FORGEBENCH_BASE_URL, falling back to http://localhost:8000; pass the real
API URL explicitly against prod:
from forgebench import Forgebench
client = Forgebench(api_key="sk_...", base_url="https://api.forgebench.ai")| Argument | Default |
|---|---|
api_key | $FORGEBENCH_API_KEY |
base_url | $FORGEBENCH_BASE_URL, then http://localhost:8000 |
timeout | 60 seconds, per request |
max_retries | 2, bounded retries on 429/5xx/network errors, with backoff |
default_headers | merged into every request |
http_client | bring your own httpx.Client / httpx.AsyncClient |
An async client is available as AsyncForgebench, same constructor, same
resources, awaited:
import asyncio
from forgebench import AsyncForgebench
async def main():
async with AsyncForgebench(api_key="sk_...", base_url="https://api.forgebench.ai") as client:
resp = await client.chat.completions.create(
model="mock-gpt",
messages=[{"role": "user", "content": "hello"}],
)
print(resp.choices[0].message.content)
asyncio.run(main())whoami
Confirms the key is valid and tells you what it resolves to: the same Principal every governed call is authorized against. Call this once at startup to fail fast on a bad credential, rather than discovering it on your first real request:
me = client.whoami()
print(me.tenant_id, me.roles, me.scopes)Chat completions
resp = client.chat.completions.create(
model="mock-gpt",
messages=[{"role": "user", "content": "Prove the governed path works."}],
)
print(resp.choices[0].message.content)
print(resp.usage.total_tokens)
print(resp.trace_id)
print(resp.call_id)trace_id ties this one response to its audit entry, its metering record and
its trace. call_id is this call's own row on the ledger; pass it as
parent_call_id on the calls it causes (see Call lineage
below). See Make your first call for the full response
shape.
Streaming
for chunk in client.chat.completions.create(
model="mock-gpt",
messages=[{"role": "user", "content": "Stream me a sentence."}],
stream=True,
):
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)The budget gate still runs before the stream opens, so a 402 raises immediately, never mid-stream. Streamed requests are never auto-retried.
Call lineage
Every governed response carries a call_id. Pass it as parent_call_id on
the calls it causes, a sub-agent's turn, the next step of a tool loop,
and the ledger records them as children: cost rolls up to the root, and the
console draws the run as a tree instead of a flat set of rows sharing a
trace. trace_id groups; parent_call_id structures. It's correlation
only; it never affects whether a call is allowed or what it costs.
plan = client.chat.completions.create(model="mock-gpt", messages=[...])
step = client.chat.completions.create(
model="mock-gpt",
messages=[...],
parent_call_id=plan.call_id,
)Tools
There are two entirely separate gates for "may this agent call this tool,"
and which one applies depends on the tool's name in the model's tool_calls
response. Sending the right shape matters more than anything else here: get
the naming wrong and the gate that should authorize the call never sees it as
one of its own.
Sending a tool schema
create() doesn't have first-class tools/tool_choice parameters; pass
them through extra_body, which the control plane forwards untouched to the
provider (extra=ignore server-side, so passthrough is safe). Below, search
is a Tool Registry entry (bare name) and mcp_time_get_current_time is an MCP
tool (mcp_ prefix):
resp = client.chat.completions.create(
model="gpt-4o",
messages=messages,
extra_body={
"tools": [
{
"type": "function",
"function": {
"name": "search",
"description": "Search the KB",
"parameters": {"type": "object", "properties": {"q": {"type": "string"}}},
},
},
{
"type": "function",
"function": {
"name": "mcp_time_get_current_time",
"description": "Current time in a timezone",
"parameters": {"type": "object", "properties": {"tz": {"type": "string"}}},
},
},
],
"tool_choice": "auto",
},
)
for call in resp.choices[0].message.tool_calls or []:
if call["function"]["name"].startswith("mcp_"):
result = your_mcp_client.call(call["function"]["name"], call["function"]["arguments"])
else:
result = your_tool_dispatch(call["function"]["name"], call["function"]["arguments"])tool_calls on the response is a list of plain dicts, OpenAI-shaped
({"id", "type", "function": {"name", "arguments"}}); arguments is a JSON
string, same as the raw API, not pre-parsed.
A response naming a tool this agent isn't bound to never reaches you as-is:
the whole response is refused with 403 tool_not_authorized (Tool Registry)
or the equivalent MCP refusal; see Errors below. Bind the tool on
the Tool Registry or MCP page first; no
amount of retrying the identical call fixes an unbound tool.
Other fields only reachable through extra_body
tools/tool_choice aren't the only fields missing a named parameter on
create(): the control plane's request model accepts several more that only
reach it through this same extra_body passthrough:
| Field | What it's for |
|---|---|
response_format | JSON mode, e.g. {"type": "json_object"} for structured output. |
metadata | The dimensions tagging; see Metadata. |
modalities | ["text", "image"], for image-generating models. |
reasoning_effort | A hint for reasoning-capable models (low/medium/high). |
request_id | Your own idempotency/correlation id. |
source | A free-text tag for where the call came from. |
The agent_tools helper (agent credentials only)
If the API key this client was built with was issued to a registered agent
(not a human/developer key), client.agent_tools gives you the agent's live
allowlist and a ready-made dispatch loop: the API key resolves the tool list
server-side, so a binding revoked centrally drops out of the very next turn
without you hardcoding or re-deploying anything. openai_schema() builds the
model's tool list from the live allowlist (GET /v1/agent-tools); dispatch
then runs every tool_call the gate let through, reports each outcome, and
returns the role: "tool" messages to append before the next turn:
tools = client.agent_tools.openai_schema()
turn = client.chat.completions.create(
model="gpt-4o",
messages=messages,
extra_body={"tools": tools},
)
messages += client.agent_tools.dispatch(
turn.choices[0].message.tool_calls,
execute=lambda name, args: your_mcp_client.call(name, args),
call_id=turn.call_id,
)dispatch never aborts the loop on a failing tool: an exception from
execute is reported as that tool's error and surfaced to the model as
{"error": ...}, so the model gets to decide what to do about a failed call,
same as any other tool result. Pass typed argument schemas per tool with
openai_schema(parameters={"search": {...json schema...}}); the control
plane only stores a tool's name and description, not its input shape.
agent_tools.report(...) is the lower-level primitive dispatch calls
internally; reach for it directly only if you're executing tools outside a
simple execute() callback (e.g. a pre-existing async tool runner).
Agents
Agents execute through the runs subsystem, which routes back through the same governed chokepoint: budget re-gated per hop, metered, audited:
agents = client.agents.list()
for a in agents:
print(a.id, a.name, a.model)
agent = client.agents.create(
name="Support bot",
model="mock-gpt",
system_prompt="You are a helpful support agent.",
)
run = client.agents.run(agent.id, input={"messages": [
{"role": "user", "content": "Help me reset my password."},
]}, wait=True, timeout=30)
print(run.status, run.output)Agent → agent (A2A tasks)
An agent bound to another by an operator can open a task on it through the door; the callee runs wherever it runs, and its own governed calls hang under the task in the ledger:
task = client.agents.call(
"qp-research",
text="10 MCQs on photosynthesis",
data={"grade": 9},
parent_call_id=plan.call_id,
)
if task.state == "input_required":
task = client.agents.call("qp-research", text="CBSE", context_id=task.context_id)
notes = task.artifact_text()To BE a callee with no inbound port (pull delivery: this agent's own
credential long-polls the door). Pass parent_call_id=task.call_id on any
governed call made inside the handler, so it nests under the task, and
client.agents.serve(handle) blocks, long-polling the door for this agent's
tasks:
def handle(task):
answer = client.chat.completions.create(
model="mock-gpt",
messages=[{"role": "user", "content": task.text}],
parent_call_id=task.call_id,
)
return task.done(text=answer.choices[0].message.content)
client.agents.serve(handle)handle may return a TaskReply (task.done(...), task.ask(...),
task.fail(...)), or just a plain str / dict / list of artifacts, which
is coerced into a completed reply automatically. An exception fails the task
with its message rather than crashing the loop.
Runs
run = client.runs.create(model="mock-gpt", input={"prompt": "hello"})
run = client.runs.get(run.id)
run = client.runs.wait(run.id, timeout=30)
print(run.status)wait polls until the run reaches succeeded or failed.
Errors
Every SDK-raised error is a forgebench.ForgebenchError. Network failures raise
ForgebenchConnectionError directly; anything that got an HTTP response raises a
subclass of APIError, one per status code, each carrying .status_code,
.code, .message, .body and .request_id:
| Exception | HTTP | Meaning | Retry? |
|---|---|---|---|
AuthenticationError | 401 | Missing, invalid or revoked API key | No: fix the key |
BudgetExceededError | 402 | The pre-call budget gate fired: no tokens spent | No: degrade deliberately |
PermissionDeniedError | 403 | Authenticated but lacking the required role/scope, model allowlist, or tool binding | No: fix config, don't retry |
NotFoundError | 404 | Resource not found, or not visible to this tenant | No |
ConflictError | 409 | Conflicting state, e.g. deleting something that must be revoked first | No: fix the resource's state first |
ValidationError | 422 | Request body rejected by the control plane | No: fix the request |
RateLimitError | 429 | Too many requests, or a tool call-volume ceiling / repeat-offender circuit breaker | Yes: auto-retried up to max_retries |
ServerError | 5xx | Control-plane error | Yes: auto-retried |
ForgebenchConnectionError | N/A | Network failure reaching the control plane | Depends: not auto-retried |
RateLimitError and ServerError are already retried automatically with
backoff up to max_retries (default 2) before the SDK ever raises them to
your code; by the time you see one, the SDK has already given up. Every
other error means "this exact call will fail again unchanged"; retrying it
in a loop just produces the identical refusal, and for a tool call
specifically, three repeated unauthorized attempts inside ten minutes
escalate to their own 429 (tool_repeatedly_unauthorized) precisely to
stop that pattern before it burns another provider call. PermissionDeniedError
carries which specific cause fired, but not always in the same place: tool
refusals populate .code directly (e.g. .code == "tool_not_authorized"),
while model_not_allowed nests its identifier under .body["detail"]["error"]
instead; .code is None for that one. Check .body when .code comes
back empty:
from forgebench import (
Forgebench,
AuthenticationError,
BudgetExceededError,
PermissionDeniedError,
NotFoundError,
ConflictError,
ValidationError,
RateLimitError,
ServerError,
ForgebenchConnectionError,
)
try:
resp = client.chat.completions.create(model="gpt-4o", messages=msgs, extra_body={"tools": tools})
except BudgetExceededError as e:
log.warning("budget gate refused the call: %s", e)
except PermissionDeniedError as e:
log.error("not authorized: %s", e.body)
except AuthenticationError as e:
log.error("auth failed: %s", e)
except ValidationError as e:
log.error("request rejected: %s", e.body)
except NotFoundError as e:
log.error("not found: %s", e)
except ConflictError as e:
log.error("conflicting state: %s", e)
except RateLimitError as e:
log.warning("rate limited after retries: %s", e)
except ServerError as e:
log.error("control-plane error after retries: %s", e)
except ForgebenchConnectionError as e:
log.error("could not reach the control plane: %s", e)Metering and budget state
s = client.metering_summary()
print(f"spent ${s.spent_usd} / limit ${s.monthly_limit_usd} "
f"(remaining ${s.remaining_usd})")
print(s.by_model)
