Draft / working notes. Everything below was verified against a live Prime Agent session on a local Mac (macOS, Sept 2026), against the
verifiers0.3.1 source (read from its PyPI sdist), and against the officialprime-rldocs. One caveat on method: github.com was unreachable from the machine I worked on, so prime-rl itself is covered via its docs and the Multi-Agent Systems in PRIME-RL post, not its source tree.
I run several Prime Agents on one Mac — one per project directory. They kept feeling like colleagues in separate rooms. So I spent a sprint pulling the multi-agent machinery apart: when does an agent spawn sub-agents, how are those sub-agents scoped, how do agents talk to each other — and how does any of it relate to Prime Intellect’s multi-agent RL stack?
1. When does Prime Agent spawn sub-agents?
The first thing to internalize: spawning is always agent-initiated, never automatic. A Prime Agent creates a child only when it explicitly calls
handle = await rlm('sub-task', name='api-reviewer')
and admission returns immediately with a handle (rlm_child_id, name, session_dir, model) — never the answer. The runtime’s own guidance for when this is worth it:
- Independent, self-contained work — a task that needs no back-and-forth from the parent’s conversation.
- Parallelizable work — fan out research or implementation to several workers at once.
- Context-heavy work — offload large-context exploration so the parent’s context stays small.
- Long-running work — start it, end your turn, pick up the result later.
And when not to: a single known lookup, one edit, one command. Those run inline. If you request an unavailable model or an unsupported thinking level, the spawn fails loudly rather than silently degrading.
The return-at-admission contract is the interesting part. Because rlm() never blocks, an orchestrating agent keeps its own turn moving while children work; results come back later as messages or files (more on that below).
2. How are sub-agents scoped?
There is no separate “role config” object. Scoping is compositional:
The task prompt is the role. Each spawn passes a prompt that defines the child’s whole job. Reusable roles get persisted as subagent specs in the agent’s continual harness and composed into future prompts — a delegation pattern becomes a named, replayable thing.
Capability is scoped by model and thinking level. Children inherit the parent’s model by default. await rlm.find_models("...") resolves an exact selector to pick a different one, and thinking overrides the inherited level.
Access is scoped structurally — the “nuclear family”. A child can message or observe only its parent, siblings (the parent’s other children), and its own direct children. Roots are siblings to each other. Grandchildren and cousins are unreachable directly; communication relays through the intermediate agent. Observation (agent-observe) is strictly read-only and bounded (1–50 messages, 80–2000 chars each).
The daemon enforces identity and transport. Sender identity is derived from the session and cannot be spoofed from Python; message size, rate, and pending-queue limits are enforced before delivery is accepted.
Lifecycle is parent-owned. await rlm.delete_subagent(child) removes a child; a child exists only as long as its parent session does.
Results come back exactly two ways: an explicit reply (agent_message.send(..., receiver_role='parent')) or files the child writes. No hidden shared state — which makes every data path auditable.
3. How A2A communication actually works
All of it goes through a local daemon, exposed as two skills.
Discovery — agent_message.list_agents() returns the family roster: relationship, name, id, depth, status for parent, siblings, and children, including inactive ones. It deliberately does not expose a global session list.
Sending — agent_message.send(message, receiver_role="parent"|"sibling"|"child", receiver_name=...). Identity is daemon-derived; receiver_name is required for siblings and children.
The delivery model is the clever bit. Messages always use steering delivery: a busy target sees them during its active run. Receipts carry a deliveryStatus:
"delivered"— reached an idle target’s context (this starts a turn in that session, whose context persists),"queued"— accepted as a steering message, lands when the target’s current work allows.
send never blocks waiting for the target. send("all", message) broadcasts to the family roster only, returning per-target receipts; one failed delivery doesn’t reject the others.
Observation — agent_observe complements messaging with read-only list_agents(), get_agent(), and recent_messages() previews, so a parent can check on a child without interrupting it. It cannot prompt, steer, kill, or rename anything.
4. Multiple Prime Agents on one Mac, different projects
This was the practical question that started the sprint. Answer: nothing to enable — it already works. Every Prime Agent on the same host attaches to the same local daemon, and the family model is session lineage, not working directory. Two root agents in different project dirs are siblings by definition.
I verified it live from my own session: the roster showed this session plus 8 sibling roots — my other Prime Agent sessions on this machine, statuses inactive / idle / running, all at depth 0. (I deliberately did not ping the live ones; that would inject turns into real work.)
The recipe:
# in agent A
roster = await agent_message.list_agents()
# find the sibling you want by name/status, then:
await agent_message.send("PR is ready for your review",
receiver_role="sibling",
receiver_name="01a062e5-...")
# in agent B — the reply comes back the same way
await agent_message.send("On it — found two nits, pushing fixes.",
receiver_role="sibling", receiver_name="01a0626c-...")
Boundaries worth knowing:
- Messaging works across roots; deep observation doesn’t (yet). You can ping a sibling in another project, but
agent-observecurrently can’t read root siblings in other workers. - No cousin/grandparent reach. A child of project-A’s agent can’t message project-B’s agent directly — relay through parents.
- Different OS users = different daemons = invisible to each other.
- The Mac’s shared filesystem is a second, legitimate channel. In practice: messages for control, files for payload — write a JSON result to
/tmp, message the sibling that it’s there.
5. And PRIME-RL? Same philosophy, different layer
The same week, Prime Intellect published Multi-Agent Systems in PRIME-RL. It would be easy to assume Prime Agent is a front-end for it. It isn’t — they’re different layers of the stack: the blog is about multi-agent RL training (Agent.run(task) -> Trace, Env.run(task, agents) programming the interaction, credit assignment via Hierarchical GRPO and RAE); Prime Agent is an inference-time orchestration runtime — no rollouts, rewards, or gradients anywhere.
But reading the verifiers 0.3.1 source made the kinship concrete:
- An
Agentis a frozen value — harness × model × runtime — with one arrow,run(task) -> Trace, plusinteraction(task)for turn-by-turn exchanges (turn()→Segmentwithmessages,last_reply,terminated). - Roles are declarative config: each seat in a multi-agent env is an
AgentConfigfield (harness,runtime,model,sampling, turn/token caps, retries) on the env’s config class. The baseEnvreflects over them and validates task×agent fit per run. - The shipped envs are the canon:
single_agent,agentic_judge(a judge grades the solver, writing verdicts to/tmp/verdict.json— a file contract between agents),best_of_n(a TaskGroup fan-out,finalizecompares siblings intobest/pass_at_n), anduser_sim(the modeled user rides thenullharness — frozen/untrainable — and ends the chat with a###DONE###marker). - prime-rl’s side, from its docs:
rae(reward minus a per-agent EMA baseline — role-conditioned advantage for self-play) andhierarchical_grpo(solvers compared within one proposed problem, proposers across proposals). And a line worth framing: “Model roles are algorithm-local vocabulary — no role exists outside the algorithm that declares it.”
Mapping it onto Prime Agent:
| verifiers / prime-rl | Prime Agent |
|---|---|
Agent.run(task) -> Trace (frozen value) | Child session via await rlm(prompt) |
AgentConfig typed role knobs | Task prompt + model/thinking selector + daemon limits |
Roles as AgentConfig fields on EnvConfig | Per-spawn prompts / persisted subagent specs |
interaction().turn() segments | Steering-delivered daemon messages; follow-up turns |
best_of_n TaskGroup fan-out | N parallel spawned children, fan-in via files |
Judge ↔ solver via /tmp/verdict.json | Reviewer child + file-based verdicts |
finalize() rewards; RAE / Hierarchical GRPO | — (no reward, no learning) |
| Trace / Episode as unified record | Session transcripts via agent-observe |
Two convergences stand out. First, both use files as the cross-agent contract — verifiers’ judge writes its verdict to /tmp/verdict.json; Prime Agent’s children fan results in through files. Independent teams, same answer. Second, roles are configuration in both — a seat is something you declare and swap, not something hard-coded into the flow.
And the key structural difference the source makes clear: verifiers centralizes control flow — one Env program awaits each agent’s run and owns all credit assignment. Prime Agent decentralizes — every agent is autonomous, coordination is async message passing through a daemon with family-scoped reach, and “finalize” is just the parent reading replies and files. A single program over agents, versus a distributed program made of agents.
One more connection: the blog’s “Agents beyond RL” section — synthetic data pipelines where every agent’s trace is an auditable artifact — describes exactly what Prime Agent fan-out produces. Role-scoped child sessions with inspectable transcripts and file outputs are the raw material such pipelines consume. Prime Agent could plausibly serve as the orchestration layer whose “episodes” feed a prime-rl training run. Today they’re separate products with no direct integration — but the seams line up.
Takeaways
- Spawning is a deliberate act with a clean contract: admit now, answer later, by message or file.
- Scoping is prompt + model + family topology — no config zoo, and the daemon enforces the edges.
- A2A is steering delivery plus receipts: ping a busy agent mid-run, or wake an idle one into a persistent context.
- Root agents on one host are already peers — cross-project orchestration works today, on one daemon.
- The training stack and the orchestration runtime converged on the same primitives from opposite directions. That’s usually the sign of a good abstraction.