AI Agents · LLM Reasoning · Efficient AI
Multi-Agent LLM Systems: Orchestration vs Delegation vs Streaming
Three June-2026 papers each fix one axis of a multi-agent LLM system: Orchestra-o1 decides who does what, SearchSwarm trains one model to delegate, StreamMA decides when agents exchange context.
Three papers, three design questions
The usual multi-agent debate is “one strong agent or many weak ones.” These three papers, all posted on arXiv in June 2026, skip that argument entirely. Each assumes multiple agents and picks a different design axis to optimize, with a measured result attached.
Orchestra-o1 asks who does what. Given a task that mixes text, images, audio and video, how should a main agent decompose it, scope each sub-agent, and pick backends? Its answer is a dependency graph over modality-masked sub-goals plus a trained orchestrator, and it reports 72.8% on OmniGAIA with GPT-5 as the brain.
SearchSwarm asks whether delegation is a trainable skill. Its answer is yes, if you synthesize the training data: a 30B-A3B model fine-tuned on harness-generated delegation trajectories jumps from 43.4 to 68.1 on BrowseComp.
StreamMA asks when agents should talk. Instead of agent A finishing a full chain-of-thought before handing it to agent B, A streams each reasoning step the moment it is written. The result is counter-intuitive: +7.3 percentage points of average accuracy and up to 26.9x faster execution.
Together they map the design space: structure, skill, protocol. This page compares the three on their own numbers and is explicit about where those numbers cannot be lined up.
Key numbers
| Design question | Orchestra-o1 | SearchSwarm | StreamMA | Same harness? |
|---|---|---|---|---|
| Axis being optimized | Task structure (who does what) | Delegation skill (whether one model splits work) | Communication protocol (when context moves) | n/a, different axes |
| Anything trained? | DA-GRPO rubric reward on a Qwen3-8B backbone | SFT on filtered harness-generated delegation trajectories | Nothing, inference-time streaming only | n/a |
| Headline accuracy | 72.8% OmniGAIA with GPT-5; the trained 8B scores 30.0% | 68.1 BrowseComp vs 43.4 for its own base (+24.7) | +7.3 pp average over both serial multi-agent and single-agent baselines | No, three different benchmarks |
| Best controlled jump | +32.8 over AOrchestra running the same GPT-5 (40.0 → 72.8) | +24.7 BrowseComp over its own base model | +22.4 pp on HMMT 2026 with Claude Opus 4.6-high | Only SearchSwarm is a same-model pair |
| Ablation isolating the mechanism | ReAct 12.5 → framework 26.3 → SFT 28.6 → vanilla GRPO 27.7 → DA-GRPO 30.0 | Bare delegation tool +2.3; full harness +10.0 on a 200-question BrowseComp subset | Streaming reaches about 83% of the theoretical pipeline speedup bound | Each within its own paper |
| Efficiency accounting | Cost 341.6 vs AOrchestra’s 565.7 at higher accuracy, in the paper’s own cost units | Delegation keeps the main agent’s context small; no token or dollar cost reported | Stream x4 ≈ $2.75 vs Serial x16 ≈ $5.46; 26.9x wall-clock at 64 agents × 64 steps | No |
Read the table for what it does not contain. There is no shared benchmark, no shared model and no shared cost unit across the three columns, so there is no honest way to say which system “wins.” What can be compared is the shape of the claims: each paper reports its intervention beating its own baseline by roughly 10 to 33 points, and each paper makes an efficiency argument, not just an accuracy one. That pattern is the actual signal for anyone building with these systems.
Structure: Orchestra-o1 and the dependency graph
Orchestra-o1’s premise is that an omnimodal task, say watching a clip and reading a chart to answer a question that needs both, overwhelms a single agent’s context and tool set. Its main agent therefore induces a dependency graph over sub-goals, and each sub-goal carries two masks: a modality mask (text, image, audio, video) and a tool mask. A sub-goal that only needs audio analysis never touches the video pipeline, which is the mechanism that keeps each sub-agent’s context small.
Three pieces sit on top. Modality-aware decomposition builds the graph. A cost-aware matching score assigns each sub-task a backend, trading capability against latency and price. Parallel execution runs all independent sub-tasks in one delegation round, with a proved bound that round-level speedup sits between 1x and the number of tasks. The main agent holds no tools at all; its context stays at 24,576 tokens while sub-agents do up to 30 steps each.
The training contribution is DA-GRPO. Standard GRPO rewards only the final answer, which is a noisy teacher for orchestration: good delegation can still miss the answer, and a lucky answer can excuse bad delegation. DA-GRPO replaces final-answer reward with a rubric scored by an external model (Claude Haiku 4.5): format correctness 0.1, action validity 0.1, tool reasonableness 0.2, decision quality 0.6. The 0.6 weight on decision quality is the design bet: the gradient chases whether the decomposition and delegation were right.
The ablation keeps the two claims separate. The framework alone lifts a ReAct baseline from 12.5% to 26.3% on the 8B setting. SFT adds to 28.6%, vanilla GRPO lands below SFT at 27.7%, and only DA-GRPO reaches 30.0%. So the structure contributes +13.8 points and the RL objective +1.4 over SFT; most of the gain is architecture, not training. And the headline 72.8% is a framework-plus-GPT-5 number, not a trained model: the model the authors actually built scores 30.0%, which still leads open omnimodal agents (OmniAtlas-Qwen3-30B manages 20.8%). The efficiency claim is the more defensible one: against the prior orchestration framework AOrchestra on the same GPT-5 backend, Orchestra-o1 reports higher accuracy (72.8% vs 40.0%) at roughly 60% of the cost (341.6 vs 565.7, in the paper’s own units).
Skill: SearchSwarm and delegation as trainable behavior
SearchSwarm starts from an observation about deep research: a long search task fills the main agent’s window with raw results and page dumps, and the usual fixes are passive, summarizing history once it grows or dropping old tool outputs by rule. SearchSwarm’s alternative is active delegation: the main agent decomposes the task, dispatches bounded subtasks through a call_sub_agent tool, and receives only a short cited report from each sub-agent, which runs in its own context. Delegation itself acts as context compression.
The problem is that natural text almost never contains explicit multi-agent coordination, so there is no scrapeable training data for the skill the authors call delegation intelligence: when to split, how to scope, how to brief, how to fold results back. Their move is to generate it. A harness with inference-time rules (require a brief that states the task and why it matters, force sub-agents to return citations) produces trajectories; the team filters them down to ones that encode correct delegation decisions; and those become SFT data for Tongyi DeepResearch-30B-A3B.
Two results make the paper’s case better than its headline. First, the harness alone does nothing: run it on the untrained base and the model never invokes call_sub_agent at all, scoring exactly like the plain base (43.4 on BrowseComp). Delegation is a behavior that has to be trained in, not prompted in. Second, the ablation on a 200-question BrowseComp subset with DeepSeek V3.2 shows where the harness’s value lives: adding only the bare delegation tool lifts the base framework from 47.7 to 50.0 (+2.3), while the full harness with briefing and citation rules reaches 57.7 (+10.0). The rules carry most of the lift, and the fine-tuning is what makes a model obey them spontaneously.
The fine-tuned result: 68.1 on BrowseComp against 43.4 for its own base, a +24.7 jump, plus 73.3 on BrowseComp-ZH, 82.5 on GAIA and 80.8 on xbench-DeepSearch, the best of the 30B-A3B class on all four. The skill also transfers: with the sub-agent tool disabled, the same single model still beats its base on a 200-question subset (52.0 vs 43.5), meaning the decomposition habit survives without the delegation tool. What the paper does not report is any token, wall-clock or dollar cost for running parallel subagents, so the efficiency of delegation versus one long context is asserted, not measured.
Protocol: StreamMA and the timing of communication
StreamMA changes one variable: the when of inter-agent communication. The standard pattern is generate-then-transfer: agent A finishes its entire chain-of-thought, then hands the whole thing to agent B. StreamMA pipes each reasoning step to B the instant A writes it, so B works on A’s early steps while A is still thinking. The latency win is obvious and large: up to 26.9x wall-clock speedup at the most parallel configuration tested (64 agents, 64 steps each), about 83% of the theoretical pipeline bound.
The accuracy win is the surprise, and the paper’s explanation is sharp: multi-step reasoning quality is non-uniform. Early steps in a chain are more reliable; later steps accumulate drift and over-thinking. A serial downstream agent that waits for the complete chain inherits the error-prone tail and gets misled by it. A streaming agent builds momentum on the reliable head, so the late mistakes arrive diluted rather than copied. Speed and quality come from the same mechanism, which means streaming is not trading accuracy for latency.
Two more results round out the protocol story. In the paper’s cost accounting, a Stream x4 configuration runs about $2.75 against $5.46 for a Serial x16 setup at higher accuracy, roughly half the price. And the authors report a step-level scaling law: increasing reasoning steps per agent while holding agent count fixed improves both effectiveness and efficiency, a dial they describe as orthogonal and composable with adding agents. The claims rest on eight math, science and code benchmarks with two frontier models and three topologies, but the accuracy story is regime-specific: on tasks where the payoff lives in the last step (long planning, multi-hop synthesis), feeding downstream agents the early head could throw away exactly the part that matters, and the paper does not cover that case.
What the three papers agree on
Lined up, the three systems converge on three shared lessons. First, decomposition pays, but through different mechanisms: Orchestra-o1 scopes sub-agents so nothing sees more than it needs, SearchSwarm uses delegation as context compression, and StreamMA shows that when context arrives matters as much as how much. Second, training data is the bottleneck for agent behavior, and synthetic generation is the workaround: Orchestra-o1 grades orchestration decisions with a rubric model, SearchSwarm filters harness trajectories into SFT data, and even StreamMA, which trains nothing, derives its protocol from a closed-form analysis of step reliability. Third, every paper makes an efficiency argument alongside the accuracy one: cheaper at higher accuracy (Orchestra-o1’s 341.6 vs 565.7), cheaper and more accurate (StreamMA’s $2.75 vs $5.46), or at minimum an explicit efficiency gap left open (SearchSwarm reports no cost at all, which is itself a finding).
When to use which
- Cross-modal compound tasks with frontier API access: Orchestra-o1’s structure. Modality-masked sub-agents and cost-aware backend matching are the reusable artifacts; quote its 72.8% only with the GPT-5 backend attached.
- One open model that must delegate reliably: the SearchSwarm route. Build a rule-based harness, generate trajectories, filter for correct delegation decisions, then SFT. Do not expect prompt rules alone to produce delegation: on an untrained base, the harness never triggered the tool once.
- Agents that pass long chains of reasoning and latency matters: StreamMA’s streaming protocol. The gain is tied to long multi-step chains; short single-shot workloads can skip it, as the paper itself concedes.
- You have not fixed the structure yet: none of the three. All three papers measure their intervention against their own already-decomposed design; none of them answers whether your task gains from being multi-agent at all.
Limits and open questions
The fundamental limit is comparability: three benchmarks (OmniGAIA, BrowseComp and a math/science/code suite), three model classes (GPT-5, a 30B-A3B MoE, and Claude Opus 4.6 / GPT-5.4), and three different cost units, with no overlap. Anyone quoting a single table ranking these systems is manufacturing a comparison that no paper made.
Within each paper the gaps are specific. Orchestra-o1 reuses OmniGAIA rather than auditing it, its cost figure has no token or wall-clock breakdown, and the 42.8-point gap between the GPT-5 framework result (72.8%) and the trained 8B (30.0%) means the framework’s value today is mostly as a wrapper for frontier proprietary models. SearchSwarm’s headline sits next to models an order of magnitude larger, but it reports no delegation cost, and its strongest transfer evidence (66.5 on a Qwen3-30B-A3B swap) is measured on a 200-question subset, not the full 1,266-question BrowseComp. StreamMA’s accuracy claim leans on Claude Opus 4.6-high and on a regime where late-step drift is the known failure mode; whether +22.4 points on HMMT survives cheaper models or synthesis-heavy tasks is open. And all three evaluate on English-centric, short-answer agent benchmarks, so none of them says much about long-horizon autonomy.
FAQ
Which multi-agent LLM system is best, Orchestra-o1, SearchSwarm or StreamMA?
There is no measured ranking, because the three systems answer different questions on different benchmarks. Orchestra-o1 reports 72.8% on OmniGAIA with GPT-5 (30.0% for its own trained 8B), SearchSwarm reports 68.1 on BrowseComp versus 43.4 for its base, and StreamMA reports +7.3 pp average across eight reasoning benchmarks. Comparing 72.8, 68.1 and +7.3 as if they were one league is meaningless; pick by the design axis you are missing: task structure, delegation skill, or communication protocol.
Does multi-agent delegation need training, or is prompting enough?
SearchSwarm’s evidence says the delegation reflex specifically needs training: its inference-time harness never triggered call_sub_agent once on the untrained base model, which scored exactly like the plain base. Orchestra-o1 shows the opposite for structure: its orchestration framework alone lifts a ReAct baseline from 12.5% to 26.3% with no RL at all. The reconciliation is that decomposition rules are promptable, but the habit of delegating rather than doing the work yourself is not, at least not for a 30B-class model.
Is Orchestra-o1’s 72.8% OmniGAIA result a model capability?
No. The 72.8% comes from running the orchestration framework on a GPT-5 backend, so it measures the harness plus a frontier model. The model the authors actually trained, an 8B orchestrator with DA-GRPO, scores 30.0%, which leads open omnimodal agents (OmniAtlas-Qwen3-30B reaches 20.8%) but is far from 72.8%. The fair framework comparison is AOrchestra on the same GPT-5: 72.8% versus 40.0%, at roughly 60% of the cost.
Does streaming partial reasoning between agents hurt accuracy?
It improves it in StreamMA’s setup. Across eight benchmarks, streaming beats both the serial multi-agent and single-agent baselines by +7.3 pp on average, peaking at +22.4 pp on HMMT 2026 with Claude Opus 4.6-high. The proposed mechanism is that early reasoning steps are more reliable than late ones, so a downstream agent that starts on the early head avoids inheriting the error-prone tail. The caveat is regime: on tasks where the conclusion only emerges in the final step, streaming the head may discard what matters.
What does multi-agent delegation cost compared with a single agent?
Only two of the three papers give numbers. StreamMA reports Stream x4 at about $2.75 versus $5.46 for a comparable Serial x16 setup, roughly half the cost at higher accuracy, plus up to 26.9x wall-clock speedup at 64 agents by 64 steps. Orchestra-o1 reports its framework at cost 341.6 against 565.7 for the prior orchestration framework on the same GPT-5, in its own cost units. SearchSwarm, despite being the purest delegation system of the three, reports no token, wall-clock or dollar cost for running parallel subagents, so its efficiency claim relative to a single long context remains unmeasured.