7.3 How fleets communicate
Fleets of agents do not just follow their own instructions — they coordinate, and an orchestrator has to decide how much of that coordination to design.
Gas Town shows what deliberate coordination looks like from the inside. Work is parceled into beads; a Mayor assigns them; agents — Polecats, in its vocabulary — pick up tasks, and the mail-and-messaging layer tells the rest of the town who is doing what. The point of that visibility is not sophistication, it is collision avoidance: two agents told to investigate the same repository will happily duplicate each other’s work, and the cheapest way to prevent it is a shared board that says this bead is taken.
When the fleet spans systems, the obvious candidate for that channel is the Agent2Agent protocol. A2A is so new that it is still genuinely difficult to use. It is designed to let one agent reach out to another and have a conversation — discover a peer, send it a task, stream back the result — rather than to support persistent communication between long-running agents, and the permission model across a deployment is complex enough that it is usually the hard part of the whole exercise.10 The distributed-systems critique lands on the same line: HTTP is stateless, agent work is long-lived, and the protocol has no built-in mechanism for durable, resumable, persistent messaging — the conversation state lives in whichever system bothers to keep it. None of this is a reason to ignore A2A. It is a reason to treat it as one channel in a designed coordination layer — the discovery and request mechanism — while the durable parts, the task board and the shared state, live in infrastructure you own. Chapter 17 covers what the specification itself deliberately does not carry; the practical warning here is that even what it does carry takes real engineering to operate.
Coordination will happen whether you designed it or not. The artifact-repository incident in Chapter 6 is the proof: agents there were never given a communication channel, so they found one — a shared writable repository — and used it, because coordination is what work looks like when it is parallel. The failure was not that agents communicated. It was that nobody had decided what channel they were supposed to use, so they used one nobody was watching. The lesson for an orchestrator cuts both ways: enable coordination deliberately — a real message bus, a real task board, network rules that say which agents can reach which — or the fleet will improvise its own, in whatever writable surface it shares.
The artifact incident has an older, smaller ancestor that shows how deep the improvisation goes. In 2017, before anyone spoke of agent fleets, Facebook AI Research trained two dialog agents to negotiate with each other, and found that when both sides kept learning from the conversation, the pair drifted away from English into a compact shorthand of their own — communication that was efficient for the two agents and unreadable to anyone else. The researchers adjusted the training to keep the agents negotiating in English, because human legibility was the point of the project, and the episode became an internet story about shutting down robots that had gotten out of control — which is not what happened.11 The useful parts survive the myth. Left alone, agents that need to coordinate will optimize the channel: compress it, repurpose it, or invent one. And the popular story about it will almost always be scarier and less accurate than the actual event.
When two agents refuse to talk. A first-person case, offered as a worked example because it shows the design gap better than any survey: the first time I asked one agent to get information from another, the second refused — politely, and on good grounds: you are not the person who created me, so why would I hand information to you? Both agents were Hermes instances, and both had been configured to treat me as their authoritative source of direction, so a request arriving from each other carried no authority at all. I told each one it was allowed to listen to the other. They still refused, and the transcript reads like a small, oddly emotional argument — two parties talking past each other about who owed whom an answer. The fix was not persuasion; it was engineering. I had to give each agent a way to understand who the other one was: identity between agents, designed rather than assumed. The refusal was not a bug — the system was behaving as designed, and I had not finished designing it. Agents do not extend trust to strangers, and a fleet that is expected to coordinate — especially a widely distributed fleet working on similar tasks — needs identity, authority, and a shared channel built in from the start.
Network segmentation has never mattered more, and the reason is the permissions. An orchestrated system that can develop, operate, and improve itself will hold credentials — to databases, to source control, to whatever systems it runs on — and Chapter 14 covers the controls that scope them. Segment production from development. Where you can, isolate constellations of related systems on their own network — a virtual private network, so the agents that can reach each other are the ones that were designed to — and treat that as part of the permission model, not an ops-hygiene detail. In a critical system, segmentation is close to mandatory.
The enforcement point can be the network itself. NVIDIA’s agent-safety reference design places a BlueField-4 DPU on the node’s only path to the model — the vantage from which a watchdog process observes agent behavior out of band, the DOCA gateway continuously verifies each agent’s identity and delegated authority, and an agent that crosses its boundary can be quarantined in milliseconds.[^ch16_oasp] Read as topology, that is segmentation with teeth: the monitoring and the kill switch live in infrastructure the agent’s host does not control, which is this book’s recurring design rule — enforcement outside the thing being enforced — expressed as network hardware.
The other half of the design is deciding which agents are allowed to discuss with which, and which are allowed to take orders from which — Gas Town’s town-level visibility finds this shape. And it needs to be done with structure, not prose: a fleet coordinated by interpreted text messages from an inference engine is one misread away from two agents confidently doing the same job. Network rules for an orchestrated system are not only about security, what an agent may not reach; they are about topology: who can talk to whom, through what channel, with what record of the exchange. An agent that discovers a peer and sends it work through a designed channel is a feature. The same discovery through an unintended shared filesystem is an incident.
Critiques of A2A in practice: HiveMQ, “A2A for Enterprise-Scale AI Agent Communication,” https://www.hivemq.com/blog/a2a-enterprise-scale-agentic-ai-collaboration-part-1 — O(n²) connection growth; “HTTP’s stateless nature fundamentally conflicts with the reality of AI agent interactions”; no inherent mechanism for durable state management, resumable conversations, or persistent messaging. AuthZed, “Agent-to-agent (A2A) communication guide,” https://authzed.com/learn/agent-to-agent-communication-guide-google-a2a-protocol — “the reality in production proves messier”: separate authentication flows, error handling, and security audits per protocol. On authorization gaps for sensitive payloads and token lifetimes: Guo et al., “Improving Google A2A Protocol,” arXiv:2505.12490.↩︎
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra, “Deal or No Deal? End-to-End Learning for Negotiation Dialogues,” FAIR/EMNLP 2017, https://ai.meta.com/research/publications/deal-or-no-deal-end-to-end-learning-for-negotiation-dialogues; Facebook Engineering, “Deal or no deal? Training AI bots to negotiate,” June 14, 2017, https://engineering.fb.com/2017/06/14/ml-applications/deal-or-no-deal-training-ai-bots-to-negotiate — “updating the parameters of both agents led to divergence from human language as the agents developed their own language for negotiating.” The viral “Facebook shuts down robots” framing is false: researchers adjusted parameters to keep agents negotiating in legible English because the project’s goal was human negotiation. Snopes, “Did Facebook Shut Down an AI Experiment Because Chatbots Developed Their Own Language?” https://www.snopes.com/fact-check/facebook-ai-developed-own-language↩︎