AC

Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation

Multi-agentEMNLP 2026, accepted, to appear · Amazon
Accepted to EMNLP 2026. The proceedings have not been published yet, so this page cites the arXiv version. arXiv comment field, verified 2026-09-15: "Accepted in EMNLP 2026".

PDF arXiv BibTeX

In one sentence

An extra classifier call picks per task whether one AI agent sends work to another as a dependency graph or as prose, gaining 12.7 points on τ-retail and preventing a 14.6-point loss on AppWorld.

Abstract

Multi-agent LLM systems coordinate through natural-language messages that consume 40–60% of their token budget. Replacing these with structured graphs reduces cost but fails on tasks requiring adaptive reasoning. We propose Routed Graph Handoff, where a lightweight LLM router (155 tokens, 0.15% overhead) selects between a typed dependency graph and natural language for each delegation. On four benchmarks (1,050+ trajectories), the routed system matches or exceeds NL-only on every task: +12.7 pp on τ-retail at 3.2× compression (p<0.01), +8.7 pp on BrowseComp at 2.2× compression (p<0.05), and parity on BFCL and AppWorld. Without the router, graph-only delegation regresses 14.6 pp on AppWorld; the router eliminates this at near-zero cost. A graph-aware executor prompt is required: the same schema without interpretation guidance yields no gain. An oracle analysis reveals 8.6 pp of additional headroom, motivating execution-time adaptive routing as future work.

Why it matters

Most multi-agent LLM systems pass work between agents as free prose, which is where most of their failures come from: 76% of failures in the authors' analysis were the receiving agent misreading order or dropping prerequisites. Swapping prose for a typed graph fixes ordering but breaks tasks that need improvisation. The practical takeaway is that one cheap extra classification call, 155 tokens, recovers the better format per task, so you get the accuracy gain without the regression.

Key results

Every number above appears in the paper. Results that were null or went the wrong way are included and marked.

A 155-token router call sends each task to either a 357-token typed graph handoff or a 576-token prose handoff.
Figure 1: System overview. The router (one LLM call) selects Graph (typed DAG, 2× compression) or NL (prose fallback) per task. Taken from the paper.

Topics

This paper is about multi-agent llm systems, agent handoff format, llm routing, prompt compression, typed graph schema, constrained decoding, tau-bench, browsecomp, appworld, bfcl function calling, llm orchestration, and agent delegation.

What this paper does not show

This was accepted to EMNLP 2026, but the proceedings are not out yet, so this page cites the arXiv version, and the code has not been released. The 76 percent failure figure comes from an automated taxonomy, not human coding, spot-checked on 50 cases. The method wins on 2 of 4 benchmarks: BFCL is a tie, 75.4 against 75.3, and the AppWorld result is parity only because the router sends 89 percent of those tasks back to prose. Graph-only loses 14.6 points there on its own. Realized routing is coarse enough that it behaves close to a per-benchmark rule, and the oracle shows 8.6 points still unclaimed.

How to cite

Pratyay Banerjee and Ankit Chadha. 2026. Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation. arXiv:2608.25277.

@misc{banerjee2026routed,
  title         = {Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation},
  author        = {Banerjee, Pratyay and Chadha, Ankit},
  year          = {2026},
  eprint        = {2608.25277},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  note          = {arXiv:2608.25277. Accepted to EMNLP 2026, proceedings not yet published},
  url           = {https://arxiv.org/abs/2608.25277}
}

Related work of mine

arXiv version.

← All publications