AGENTENCODE / RESEARCH & ORCHESTRATION

Multi-Agent Benchmarks: What Research Really Measures

Externally measured findings, primary sources and reproducibility limitations.

Editorial: AgentenCodeReviewed: 9 October 2026Primary sources & external studies
External research is not presented as our own measurement. Unknown remains unknown.

What do multi-agent benchmarks actually measure?

Multi-agent benchmarks evaluate how multiple agents cooperate or compete under specified tasks, models, coordination protocols and scoring procedures. A result is meaningful only inside its study setup. Single-agent capability, tool-using agents and multi-agent coordination are different research targets. The findings presented here were measured externally and curated by AgentenCode. AgentenCode has not run MARBLE or AgentBench as its own benchmark campaign.

Source What was evaluated Appropriate use Not an appropriate substitute for
MultiAgentBench / MARBLE (Zhu et al., ACL 2025) Cooperation, competition and coordination among multiple LLM agents Understanding coordination research and milestones Universal commercial-agent rankings
AgentBench (ICLR 2024) Language models as agents in different environments Context for single-agent task competence Direct Handoff-versus-Graph performance test
Microsoft Agent Framework documentation Available workflow/orchestration patterns Evidence of documented implementation features Independent empirical performance study
AgentenCode agent profiles Source-backed product and governance fields Evidence of particular product claims AgentenCode-executed multi-agent benchmarks

Primary research anchor: MultiAgentBench (ACL 2025)

The peer-reviewed paper “MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents” by Kunlun Zhu and co-authors appears in the 2025 ACL proceedings (DOI 10.18653/v1/2025.acl-long.421). Its interactive scenarios cover cooperation and competition. The abstract discusses star, chain, tree and graph coordination topologies, and strategies such as group discussion and cognitive planning. Official ACL entry.

What the abstract reports: Across its tested scenarios, gpt-4o-mini obtained the highest average task score, a graph structure performed best in the research scenario among tested protocols, and cognitive planning improved milestone achievement rates by three percent. None of those sentences establishes a universal ranking of models, agents or orchestration architectures. The abstract alone does not specify every confidence interval, reproduction condition or cost consequence.

Required fields in a trustworthy result record

  1. Study and version: DOI, publication year and a pinned repository revision if the code was run.
  2. Task family: the actual interactive task or environment, not a vague label such as “intelligence”.
  3. System setup: agent count, models, prompts, tools, coordination and resource budget.
  4. Metric: task score and milestone completion differ from latency and cost.
  5. Comparator: what other configurations were tested under the same constraints?
  6. Uncertainty: run count, spread and statistical test; mark unreviewed properties explicitly.
  7. Provenance: original source, specific section and actual retrieval/review date.

For this initial dataset AgentenCode therefore does not invent leaderboard rows. It records studies, documented architecture classes and narrowly scoped qualitative published observations. The public research schema can accommodate verified quantitative rows in a future release.

MARBLE: official research code, not our measurement

The authors link ulab-uiuc/MARBLE as companion code, describing the multi-agent research environment, modular components and simulation setup. The repository displays an MIT license for code; this is not a blanket license for every third-party dataset or copied result visualization. AgentenCode has not automatically imported third-party tables, result files or illustrations. We provide original editorial analysis with links to the authoritative material.

Reproducibility would require a pinned commit, dependency versions, model API versions and parameters, prompts, dataset revision, seeds, hardware/runtime setup, evaluator code and observed costs. Source-code availability alone does not satisfy these conditions. No AgentenCode reproduction has been claimed.

AgentBench: useful context but a different construct

AgentBench, known from ICLR 2024, addresses broader agent capabilities in interactive environments. It provides context for how language models perform in agent-oriented tasks. The results should not be silently inserted into a chart claiming to rank multi-agent coordination patterns. The research constructs, units and experimental designs differ.

Why scores are rarely plug-and-play comparable

Dimension Compare only when … Frequent failure
Dataset same dataset and verified version Revised or contaminated evaluation data
Task equivalent goal and permitted tools Research and coding treated as one task
Model same model revision or controlled replacement Model improvement mistaken for coordination gain
Topology explicit protocol and state design Brand name substituted for architecture
Metric identical definition, units and aggregation Combining success percentages with arbitrary scores
Budget matching token, time and tool-call limits Extra expenses hidden from the comparison
Statistics documented sample size and uncertainty A single result implies universal confidence

Our approach therefore does not merge incompatible scores. Even within one study a reported improvement applies to the documented test design. Peer review does not automatically certify that a commercial implementation of the tested pattern is secure or compliant with privacy law.

Translating papers into architecture decisions

Step 1: Define a real workload, such as three independent source-critical research tasks. Step 2: Establish a single-agent baseline. Step 3: Compare Concurrent, Sequential and possibly Graph using identical tasks and budgets. Step 4: Track success, source correctness, latency, cost, runaway loops, human intervention and risky tool actions. Step 5: Report only interpretable differences with a documented statistical and causal basis.

This is an evaluation protocol proposal, not completed AgentenCode experiments. See orchestration patterns for their technical trade-offs and methodology for evidence, missing values, licensing and conflict management.

Primary sources and publication scope

Core references: Zhu et al. (ACL 2025), MARBLE code, AgentBench and Microsoft Agent Framework. These sources were inspected for this editorial release on 9 October 2026. Results are attributed to their publication context; missing reproductions and unresolved dataset licenses remain explicit limitations.

RESEARCH REGISTRY

Filter external studies

This view uses the published source and study registry, not unverified score values.

Source summaries are available in the table above.

Primary sources and provenance

Each citation supports a scoped proposition. Third-party results have not been independently measured by AgentenCode.

  1. Zhu et al. · MultiAgentBench, ACL 2025
  2. MARBLE · research implementation
  3. AgentBench · reference repository
  4. Microsoft Agent Framework · Orchestrations

Research dataset: JSON · explicit graph edges: JSON

Frequently asked questions

How does MultiAgentBench differ from AgentBench?

MultiAgentBench targets interaction and coordination across multiple agents; AgentBench more broadly evaluates language models acting in varied environments.

Were the published results measured by AgentenCode?

No. Findings are attributed to external studies or repositories and are not presented as in-house experiments.

Can scores across different papers be compared directly?

Only with compatible tasks, metrics, model versions, budgets and statistical designs; otherwise the comparison is misleading.

Why is there no universal top-ten leaderboard?

Merging incompatible experiments would invent general conclusions that the reported measurements cannot support.

Can MARBLE research files be freely copied?

A code license does not settle the rights of all bundled datasets, figures or result tables.