Agents Agent-to-agent (A2A) Serving tasks

Serving tasks

An agent becomes a callee the moment another agent has an approved binding to call it. What that callee actually does with an incoming task is entirely up to its own process — Forgebench hands the task over, either by letting the callee pull it or by pushing it to a URL, and waits for a reply in the same shape either way.

Pull: the callee claims its own work

If the callee has no endpoint_url set, it's pull-only: its own process claims work by long-polling GET /v1/agents/me/tasks/next with its own agent credential, then posts the result back. Nothing has to be reachable from the outside — no inbound port, no public IP, no firewall rule to open. That's the same shape a Temporal worker or a CI runner already uses to pull jobs instead of accepting them.

The claim query uses FOR UPDATE SKIP LOCKED, so running several replicas of the same callee polling at once is safe: each replica gets a different task, and none is ever handed out twice.

import os
from forgebench import Forgebench

client = Forgebench(api_key=os.environ["CALLEE_AGENT_KEY"],
               base_url="https://api.forgebench.ai")

def handle(task):
  invoice_id = task.data.get("invoice_id")
  if invoice_id is None:
      return task.ask("which invoice_id should I classify?")
  try:
      category = classify(invoice_id, parent_call_id=task.call_id)
  except ClassificationError as exc:
      return task.fail(str(exc))
  return task.done(data={"category": category})

# blocks; returns the number of tasks handled
client.agents.serve(handle, max_tasks=None)

serve() claims a task, hands it to your handler, and posts the reply, looping until max_tasks/maxTasks is reached or stop()/an aborted signal says to quit. The handler can return:

  • a TaskReply built from task.done(...) / task.ask(...) / task.fail(...) (Python) or the reply builder (TypeScript)
  • a plain string, dict/object, or list of artifacts — converted into a completed reply automatically, so a handler that just wants to answer doesn't have to build the envelope by hand
  • nothing at all — raise/throw, and the exception's message fails the task instead of crashing the whole loop

Finishing a claimed task manually

Without serve(), claim and reply are two separate calls — reach for this if you already have your own polling loop or worker framework and serve() would just be fighting it:

task = client.agents.next_task(wait=20)
if task is not None:
  reply = task.done(text="done")   # or task.ask(...) / task.fail(...)
  client.agents.reply(task.id, reply)

A reply is only accepted while the task is working, and only from the callee that claimed it — if the caller cancels the task in the meantime, the reply is refused with 409 rather than silently applied over a canceled task.

Push: a webhook instead of polling

Setting an endpoint_url on the callee switches delivery from pull to push: the control plane POSTs every task to that URL instead of the callee polling for one. This trades "no inbound port" for "no polling loop to run" — worth it once the callee already has a always-on HTTP service anyway (a FastAPI app, a Lambda behind a URL, and so on).

Configure it from the callee's own Calls tab, or the same call via the API:

The Delivery section on an agent's Calls tab — pull by default, switches to push the moment a URL is saved.
The Delivery section on an agent's Calls tab — pull by default, switches to push the moment a URL is saved.
curl -X PUT https://api.forgebench.ai/v1/agents/$CALLEE_ID/endpoint \
-H "Authorization: Bearer $ADMIN_KEY" \
-H "Content-Type: application/json" \
-d '{
  "url": "https://callee.example.com/a2a",
  "auth": {"type": "bearer", "secret": "s3cret"}
}'

Notes on the endpoint config:

  • auth.secret is sealed in the vault and never returned by any API response or shown again in the console. Omit auth on a later PUT to keep the current credential, or send "clear_auth": true to drop it.
  • "url": null reverts the agent to pull-only.
  • A private-range URL (RFC 1918, etc.) requires "allow_private_network": true plus admin privilege to set. Loopback and cloud-metadata addresses are always blocked, both when the URL is saved and on every dial after that — a hostname that gets re-pointed at a private address later doesn't get a free pass just because it passed the check once.

The control plane POSTs the task to the configured URL:

POST https://callee.example.com/a2a
Content-Type: application/json
X-Forgebench-Task: <task_id>
X-Parent-Call: <task's own call_id>
Authorization: Bearer s3cret          # or your configured header

{"task_id": "...", "call_id": "...", "message": {"role": "user", "parts": [...]}}

Your handler answers in the same shape a pull reply uses (TaskResultIn):

from fastapi import FastAPI, Request

app = FastAPI()

@app.post("/a2a")
async def handle_task(request: Request):
  body = await request.json()
  parts = body["message"]["parts"]
  text = next((p["text"] for p in parts if "text" in p), "")

  # pass this to parent_call_id on this agent's own governed calls
  parent_call_id = body["call_id"]

  return {
      "state": "completed",
      "artifacts": [{"name": "response", "parts": [{"text": f"handled: {text}"}]}],
  }

The response body doesn't have to match TaskResultIn exactly — this is what makes a plain webhook with no A2A awareness at all still work:

  • a JSON object without a state/artifacts key is wrapped as a data artifact holding the whole body
  • a non-JSON (plain text) response becomes a text artifact
  • {"state": "working"} tells the control plane this callee will finish asynchronously; it later posts the real result to POST /v1/agents/me/tasks/{id}/result with its own agent credential, the same route a pull callee uses to finish up

Lineage: the call tree

Every governed call returns call_id in its response body (and X-Call-Id on the streaming path). Send X-Parent-Call: <call_id> on a call that another call caused, and the control plane stamps parent_call_id/root_call_id on it. This is bookkeeping, not policy — no gate reads it, and nothing is refused because of it.

A task participates in the same lineage:

  • the caller's parent_call_id/root_call_id are stamped on the task row when it's opened
  • the task's own call_id becomes the parent for every governed call the callee makes while working on it

Chain that through a few hops — agent A calls agent B, which calls a model and then agent C — and the whole workflow becomes one tree on the ledger, with no orchestrator anywhere having to declare the relationship up front. One query gets its total cost:

SELECT sum(cost_usd) FROM metering_events WHERE root_call_id = :root;

Pass parent_call_id (Python) / parentCallId (TypeScript) to agents.call(...) with the governed call that led to opening the task. A Task object exposes its own call_id/parent_call_id/root_call_id so a serve() handler can continue the chain downward.

Next