Tool Registry
Response-gating for tool calls: no endpoint, no execution, just authorize-or-block. It decides which tools an agent's own model may propose calling: checked on every response, before the agent ever sees a tool it isn't bound to.
Register a tool by the exact string the LLM must emit in
tool_calls[].function.name, bind it to the agents allowed to call it, and
every governed response is checked against that binding before the agent
gets to act on it. Unlike MCP, there is no server behind a Tool Registry
entry, it never dials anything, never holds a credential, never executes a
thing. It is purely the question "was this agent allowed to propose this."
Register a tool
Name it exactly
Must match the string the LLM emits, character for character. Cannot be changed later, so get it right the first time.
Describe it
What a colleague reads when deciding whether an agent should be bound to this.
Confirm the payload location
The standard OpenAI-style shape (
tool_calls[].function.name/.arguments) is detected automatically. If your response puts tool calls somewhere non-standard, set the three paths explicitly.
As you type the name, a live preview shows how the row will read in the registry before you submit. Switching Payload location to Custom paths reveals three fields: the path to the call list, to the name inside one call, and to its arguments, each defaulting to the standard OpenAI shape; a blank field falls back to that default rather than failing.
Bulk registration works from a CSV with columns name, description,
tool_calls_path, name_path, arguments_path; only name is required;
the path columns default to the standard shape. A duplicate or invalid row
is skipped, not fatal.
Opening a registered tool gives its own Overview, Bindings, and Activity tabs, plus a Disable everywhere action, the same platform-wide kill switch described below, available from the tool's own page as well as its row in the list.
Handle it in your code
Register and bind the tool here first. Your calling code doesn't change to
add the gate itself, only to handle the refusal it can now return. A blocked
response raises before your usual tool_calls handling ever runs.
tools/tool_choice aren't first-class parameters on either SDK's chat
call: the control plane accepts and forwards them untouched (extra=ignore),
so send them through Python's extra_body passthrough, or a widened type on
TypeScript (tools isn't declared on ChatCompletionRequest yet). On Python,
tool_calls comes back as a list of plain dicts (OpenAI-shaped on the wire),
not attribute-access objects, so index into it, don't dot into it. A
PermissionDeniedError here means tool_not_authorized: the model proposed
a tool this agent isn't bound to, so bind it rather than retry; a
RateLimitError means either the tool's call-volume ceiling or the
repeat-offender circuit breaker (three unauthorized attempts in ten minutes):
curl -X POST https://api.forgebench.ai/v1/chat/completions \
-H "Authorization: Bearer sk_..." \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Find invoices from last quarter."}],
"tools": [{
"type": "function",
"function": {
"name": "search",
"description": "Search the KB",
"parameters": {"type": "object", "properties": {"q": {"type": "string"}}}
}
}]
}'from forgebench import Forgebench, PermissionDeniedError, RateLimitError
client = Forgebench(api_key="sk_...", base_url="https://api.forgebench.ai")
messages = [{"role": "user", "content": "Find last quarter's invoices."}]
try:
resp = client.chat.completions.create(
model="gpt-4o",
messages=messages,
extra_body={"tools": [...]}, # "search", "send_email", etc.
)
except PermissionDeniedError as e:
log.error("tool not authorized: %s", e)
except RateLimitError as e:
log.warning("tool call rate-limited: %s", e)
else:
for call in resp.choices[0].message.tool_calls or []:
fn = call["function"]
result = your_tool_dispatch(fn["name"], fn["arguments"])import { Forgebench, PermissionDeniedError, RateLimitError, type ChatCompletionRequest } from "@seedlinglabs/forgebench-sdk";
const forgebench = new Forgebench({ apiKey: "sk_...", baseUrl: "https://api.forgebench.ai" });
const messages = [{ role: "user" as const, content: "Find last quarter's invoices." }];
try {
const res = await forgebench.chat.create({
model: "gpt-4o",
messages,
tools: [...], // your tool schema
} as ChatCompletionRequest & Record<string, unknown>);
for (const call of res.choices[0]?.message.tool_calls ?? []) {
const result = await yourToolDispatch(call.function.name, call.function.arguments);
}
} catch (e) {
if (e instanceof PermissionDeniedError) {
console.error("tool not authorized", e);
} else if (e instanceof RateLimitError) {
console.warn("tool call rate-limited", e);
} else {
throw e;
}
}Against curl, the same refusal is just a 403/429 with a JSON body, no
typed exception to catch, see
What a blocked call looks like to the agent
below for its exact shape.
Which agents may call it
Binding is the authorization decision itself: an agent with no binding for a tool cannot call it, full stop. Bindings are made from either side, and take effect on the agent's very next call; bindings are read fresh, never cached.
From the tool: open it, then its Bindings tab
Click Bind agent, choose from the dropdown of agents not yet bound. Remove one later with Unbind on its row.
Or from the agent: open its Tools tab, Registry sub-tab
Click Bind tools and tick as many catalog entries as this agent needs from the checkbox list, separate from MCP allowlist next to it, which is the MCP allowlist instead.
Disabling a tool (enabled: false) is the platform-wide kill switch: it
blocks every agent bound to it at once, without touching a single binding.
Set a call-volume ceiling
Binding decides whether an agent may call a tool at all; a ceiling decides how often: a separate, optional cap on top. Nothing prices a tool call in dollars, so this caps calls per window, not spend.
Open the tool, then its Ceilings tab
Lists every ceiling already set on this tool, with current usage against each.
Add ceiling
Choose a window (per minute, hour, day, or month; a minute window is a rate limit, a month window is a quota, same mechanism either way) and a max call count. Leave the agent field blank to cap every agent calling this tool combined; pick one to scope the ceiling to just that agent.
Save
Takes effect immediately, checked on every future call against this tool.
Tripping a ceiling is a distinct outcome from an unauthorized call, see
ceiling_exceeded below, and refuses with 429, not 403.
What actually gets checked
Authorization runs on the response, after the provider round-trip, because
there is nothing to check until the model has actually proposed a call. That
means a blocked tool call still cost you the completion; the block only
stops the agent from acting on it. Extraction looks at every entry in
tool_calls, matches each against the agent's current bindings, and, if
any entry in the batch is unauthorized, blocks the whole response, not
just that one call. A sibling call that individually would have passed is
recorded withheld, not authorized: it never reached the agent either.
| Outcome | Meaning |
|---|---|
authorized | Every call in the batch was bound. All proceed to the agent. |
unauthorized | This call named a tool the agent has no binding for. The whole response is blocked. |
withheld | This call was individually fine, but a sibling in the same response wasn't; it never reaches the agent either. |
ceiling_exceeded | This call is authorized, but tripped a configured call-volume ceiling for this agent/tool. |
Every attempt is recorded, authorized and unauthorized alike; an agent repeatedly proposing a tool it isn't bound to is itself a signal worth keeping, visible on the tool's own activity log.
What a blocked call looks like to the agent
A blocked response comes back as an error, not a silently-stripped
tool_calls array:
- Unauthorized (
403): this agent is not bound to the tool the model tried to call. - Repeatedly unauthorized (
429): the same agent/tool pair has been blocked three or more times in the last ten minutes. This trips before the provider is even called, so a caller in a retry loop stops paying for a call that is statistically certain to fail the same way again. Binding the tool clears it immediately. - Ceiling reached (
429): the call was authorized, but this agent/tool pair has hit its configured volume cap for the current window.

