Skip to content
Scan to share

A2A Inside: How Prime Agent Runs Multi-Agent Workflows (and What It Shares with PRIME-RL)

https://chatek.co/posts/prime-agent-a2a

A2A Inside: How Prime Agent Runs Multi-Agent Workflows (and What It Shares with PRIME-RL)

Chance Jiang
Published date:

Draft / working notes. Everything below was verified against a live Prime Agent session on a local Mac (macOS, Sept 2026), against the verifiers 0.3.1 source (read from its PyPI sdist), and against the official prime-rl docs. One caveat on method: github.com was unreachable from the machine I worked on, so prime-rl itself is covered via its docs and the Multi-Agent Systems in PRIME-RL post, not its source tree.

I run several Prime Agents on one Mac — one per project directory. They kept feeling like colleagues in separate rooms. So I spent a sprint pulling the multi-agent machinery apart: when does an agent spawn sub-agents, how are those sub-agents scoped, how do agents talk to each other — and how does any of it relate to Prime Intellect’s multi-agent RL stack?

1. When does Prime Agent spawn sub-agents?

The first thing to internalize: spawning is always agent-initiated, never automatic. A Prime Agent creates a child only when it explicitly calls

handle = await rlm('sub-task', name='api-reviewer')

and admission returns immediately with a handle (rlm_child_id, name, session_dir, model) — never the answer. The runtime’s own guidance for when this is worth it:

And when not to: a single known lookup, one edit, one command. Those run inline. If you request an unavailable model or an unsupported thinking level, the spawn fails loudly rather than silently degrading.

The return-at-admission contract is the interesting part. Because rlm() never blocks, an orchestrating agent keeps its own turn moving while children work; results come back later as messages or files (more on that below).

2. How are sub-agents scoped?

There is no separate “role config” object. Scoping is compositional:

The task prompt is the role. Each spawn passes a prompt that defines the child’s whole job. Reusable roles get persisted as subagent specs in the agent’s continual harness and composed into future prompts — a delegation pattern becomes a named, replayable thing.

Capability is scoped by model and thinking level. Children inherit the parent’s model by default. await rlm.find_models("...") resolves an exact selector to pick a different one, and thinking overrides the inherited level.

Access is scoped structurally — the “nuclear family”. A child can message or observe only its parent, siblings (the parent’s other children), and its own direct children. Roots are siblings to each other. Grandchildren and cousins are unreachable directly; communication relays through the intermediate agent. Observation (agent-observe) is strictly read-only and bounded (1–50 messages, 80–2000 chars each).

The daemon enforces identity and transport. Sender identity is derived from the session and cannot be spoofed from Python; message size, rate, and pending-queue limits are enforced before delivery is accepted.

Lifecycle is parent-owned. await rlm.delete_subagent(child) removes a child; a child exists only as long as its parent session does.

Results come back exactly two ways: an explicit reply (agent_message.send(..., receiver_role='parent')) or files the child writes. No hidden shared state — which makes every data path auditable.

3. How A2A communication actually works

All of it goes through a local daemon, exposed as two skills.

Discovery — agent_message.list_agents() returns the family roster: relationship, name, id, depth, status for parent, siblings, and children, including inactive ones. It deliberately does not expose a global session list.

Sending — agent_message.send(message, receiver_role="parent"|"sibling"|"child", receiver_name=...). Identity is daemon-derived; receiver_name is required for siblings and children.

The delivery model is the clever bit. Messages always use steering delivery: a busy target sees them during its active run. Receipts carry a deliveryStatus:

send never blocks waiting for the target. send("all", message) broadcasts to the family roster only, returning per-target receipts; one failed delivery doesn’t reject the others.

Observation — agent_observe complements messaging with read-only list_agents(), get_agent(), and recent_messages() previews, so a parent can check on a child without interrupting it. It cannot prompt, steer, kill, or rename anything.

4. Multiple Prime Agents on one Mac, different projects

This was the practical question that started the sprint. Answer: nothing to enable — it already works. Every Prime Agent on the same host attaches to the same local daemon, and the family model is session lineage, not working directory. Two root agents in different project dirs are siblings by definition.

I verified it live from my own session: the roster showed this session plus 8 sibling roots — my other Prime Agent sessions on this machine, statuses inactive / idle / running, all at depth 0. (I deliberately did not ping the live ones; that would inject turns into real work.)

The recipe:

# in agent A
roster = await agent_message.list_agents()
# find the sibling you want by name/status, then:
await agent_message.send("PR is ready for your review",
                         receiver_role="sibling",
                         receiver_name="01a062e5-...")

# in agent B — the reply comes back the same way
await agent_message.send("On it — found two nits, pushing fixes.",
                         receiver_role="sibling", receiver_name="01a0626c-...")

Boundaries worth knowing:

5. And PRIME-RL? Same philosophy, different layer

The same week, Prime Intellect published Multi-Agent Systems in PRIME-RL. It would be easy to assume Prime Agent is a front-end for it. It isn’t — they’re different layers of the stack: the blog is about multi-agent RL training (Agent.run(task) -> Trace, Env.run(task, agents) programming the interaction, credit assignment via Hierarchical GRPO and RAE); Prime Agent is an inference-time orchestration runtime — no rollouts, rewards, or gradients anywhere.

But reading the verifiers 0.3.1 source made the kinship concrete:

Mapping it onto Prime Agent:

verifiers / prime-rlPrime Agent
Agent.run(task) -> Trace (frozen value)Child session via await rlm(prompt)
AgentConfig typed role knobsTask prompt + model/thinking selector + daemon limits
Roles as AgentConfig fields on EnvConfigPer-spawn prompts / persisted subagent specs
interaction().turn() segmentsSteering-delivered daemon messages; follow-up turns
best_of_n TaskGroup fan-outN parallel spawned children, fan-in via files
Judge ↔ solver via /tmp/verdict.jsonReviewer child + file-based verdicts
finalize() rewards; RAE / Hierarchical GRPO— (no reward, no learning)
Trace / Episode as unified recordSession transcripts via agent-observe

Two convergences stand out. First, both use files as the cross-agent contract — verifiers’ judge writes its verdict to /tmp/verdict.json; Prime Agent’s children fan results in through files. Independent teams, same answer. Second, roles are configuration in both — a seat is something you declare and swap, not something hard-coded into the flow.

And the key structural difference the source makes clear: verifiers centralizes control flow — one Env program awaits each agent’s run and owns all credit assignment. Prime Agent decentralizes — every agent is autonomous, coordination is async message passing through a daemon with family-scoped reach, and “finalize” is just the parent reading replies and files. A single program over agents, versus a distributed program made of agents.

One more connection: the blog’s “Agents beyond RL” section — synthetic data pipelines where every agent’s trace is an auditable artifact — describes exactly what Prime Agent fan-out produces. Role-scoped child sessions with inspectable transcripts and file outputs are the raw material such pipelines consume. Prime Agent could plausibly serve as the orchestration layer whose “episodes” feed a prime-rl training run. Today they’re separate products with no direct integration — but the seams line up.

Takeaways

  1. Spawning is a deliberate act with a clean contract: admit now, answer later, by message or file.
  2. Scoping is prompt + model + family topology — no config zoo, and the daemon enforces the edges.
  3. A2A is steering delivery plus receipts: ping a busy agent mid-run, or wake an idle one into a persistent context.
  4. Root agents on one host are already peers — cross-project orchestration works today, on one daemon.
  5. The training stack and the orchestration runtime converged on the same primitives from opposite directions. That’s usually the sign of a good abstraction.
Next
The 3 Phases of AI Revolution