When Agents form teams to perform tasks, the top-performing model only completes 50% of the assigned tasks, and this benchmark assesses their collaborative capabilities.
AI Agents with strong individual capabilities are constantly emerging. From writing code and developing software to retrieving literature and assisting scientific research, they can independently complete an increasing number of tasks.
Yet as these "super individuals" grow more powerful, an equally critical question arises: can they collaborate?
Looking back at human society, many complex achievements rely on the division of labor and collaboration of a group of people. If we want to set up a company, build rockets and land on the moon, the first thing to consider is probably not to find an almighty genius, but how to organize a team: what kind of professionals are needed, how to divide the work, how to share information, and how to connect the work of different links.
From scientific discoveries to large-scale engineering projects, individual capabilities are undoubtedly important, and organization and collaboration also determine how far we can go.
This makes us start to think: when the individual capabilities of large models and Agents continue to improve, to what extent has their collaborative capability developed? Facing a common goal, will they take the initiative to exchange information and coordinate the division of labor? When multiple agents work together, can they truly give full play to their respective advantages to make complex tasks easier to complete?
With these questions in mind, researchers from OpenAgents, Columbia University, the University of Pennsylvania, Seoul National University, and Pennsylvania State University jointly built AgentWorld: a benchmark for evaluating the multi-agent collaborative capability.
- Paper link: https://arxiv.org/abs/2609.31590
- Project website: https://agentworld.io/
- Code and data: https://github.com/openagents-org/agentworld
The research team focuses on: when multiple agents have different roles, capabilities and resources and need to jointly complete a task, can they form an effective division of labor, maintain information synchronization, connect their respective actions, and finally achieve the common goal.
What new improvements does AgentWorld bring?
In recent years, the academic community has been exploring ways to enable multiple large models to jointly solve problems through dialogue, role division and mutual feedback, and has also proposed collaboration-oriented evaluation tasks, as well as game and social simulation environments.
As these researches advance, a question that needs to be further answered is: when the task lasts for dozens of rounds, members hold different resources, and every step may affect teammates, can the team continue to maintain effective collaboration?
The feature of AgentWorld is that it combines long-horizon tasks, role asymmetry and black-box interaction into the same collaborative evaluation environment.
High-level tools reduce the interference of low-level controls such as navigation and operation, and procedural rules check whether the common goal is completed. The further proposed CCE indicator traces which actions and messages are judged to contribute to the result. In this way, the evaluation can not only observe whether the team succeeds, but also analyze how cross-role handover and cooperation contribute to success.
Introduction to AgentWorld
A team of 10 agents enters a role-playing game (RPG) world.
Imagine such a task: make a magic wand for the mage in the team. The lumberjack is responsible for collecting wood, the carpenter processes the wood into a wooden stick, and the mage uses the wooden stick and the materials he holds to complete the production. It seems that everyone only needs to do their own job well.
But in actual execution, problems may arise in any handover: who should the wood be handed over to? Is the wooden stick ready? Do you need to repeat the steps that your teammates have already completed? When the task is not completed, does anyone announce "work done" in advance?
These problems are exactly the tests that multi-agent collaboration systems need to face.
AgentWorld is a benchmark for Agent collaborative capability: it allows multiple large language model-driven agents to divide labor, communicate, share resources and jointly complete tasks in the same continuously changing environment.
In the four main model experiments reported in the paper, the highest success rate on the main task set is 52.0%. This result raises a question worthy of research: when the capability of a single Agent continues to improve, what is still missing in team-level collaboration?
Figure 1|Example of the AgentWorld environment: different roles have their own perspectives and communicate through in-game chat. The screenshot shows a ten-agent team, which does not correspond to the three-agent magic wand task described below.
Turn "capable of collaboration" into a measurable capability
"Multiple Agents participating" is a system structure, while "effective collaboration of multiple Agents" is a capability that needs to be verified.
If we just let several models answer questions separately and then aggregate the answers, it is still difficult for us to know: whether they can understand each other's responsibilities, adjust the division of labor when resources are insufficient, and modify plans according to the latest progress of teammates.
AgentWorld puts these problems into an MMORPG sandbox, that is, a massively multiplayer online role-playing game environment. There are collectible resources, craftable items, enemies to deal with, and teammates in different locations. The actions of agents will change the environment and affect what other members can do next.
The game provides executable and observable collaboration scenarios; the evaluation focuses on whether the team can push the common goal to completion.
The task set introduced in the paper consists of 100 manually designed tasks and 100 enhanced variants, covering combat, crafting, collection, trading, exploration, survival, construction and coordination. The tasks require 3 to 20 agents to participate, and many tasks include continuous resource handover and action dependencies.
This design turns collaboration from a piece of seemingly reasonable dialogue into a series of actions that must actually happen.
Three Designs
Make collaboration challenges visible
The first is role asymmetry. Different agents have different skills, items and initial conditions. The team needs to arrange work according to these differences, instead of letting all members perform the same set of actions.
In the magic wand task, logging, processing and crafting form a dependency chain. Even if every member knows the final goal, the team still needs to solve the problems of "who does it first, who to hand it over to, and when to hand it over".
Figure 2|Another task example: ten agents take on responsibilities such as mining, smelting, forging, support and coordination to jointly complete the item crafting goal. The role configuration makes division of labor and resource transfer part of the task.
The second is black-box interaction. Agents cannot directly read the internal reasoning status of their teammates. They need to judge what other members are doing through environmental observations, chat and action results. This makes "I thought my teammates had already finished" a real source of failure.
The third is multi-round execution. The initial plan cannot cover all subsequent changes. The team must continuously update the status during execution: whether the resources are in place, whether teammates need support, and whether the original division of labor is still effective. The main task in the paper sets an execution budget of dozens of rounds, requiring agents to maintain coordination in multiple handovers.
In order to reduce the interference of low-level operations, AgentWorld provides 13 high-level API tools. For example, the collection or attack tool can encapsulate the navigation and specific execution process, allowing the model to focus more decision-making on "what to do next" and "who to cooperate with".
This does not mean that the evaluation completely excludes the influence of planning and tool usage, but it makes collaboration problems easier to observe and analyze.
The best teams complete about half of the tasks
More chats do not necessarily mean better performance
The study evaluated Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini and DeepSeek R1-70B under a unified prompt template, tool definition and interaction protocol.
On the main task set, the task success rates of the four models are 52.0%, 45.0%, 36.0% and 20.0% respectively. On the enhanced task set, the corresponding success rates are 24.0%, 26.0%, 21.0% and 10.0%.
Figure 3|Redrawn according to the main result table of the paper. The comparison here is the performance of the teams driven by each model under the settings specified in the paper; it does not represent the judgment of the capability upper limit of all Agent system designs.
It is also worth noting the gap between "partially completed" and "truly finished". The partial success rate of Gemini 3 Flash on the main task set is 71.5%, and the full task success rate is 52.0%. The team's progress towards some intermediate goals does not guarantee that the final resource handover or crafting steps can be completed.
The communication volume also does not show a simple "the more the better" relationship. GPT-5 Mini sends an average of 44.1 chat messages per main task, but the success rate is 36.0%; Gemini 3 Flash sends an average of 11.0 messages, with a success rate of 52.0%.
This set of results does not prove that fewer chats will improve the success rate, but it reminds us that the number of messages cannot directly replace the quality of collaboration.
What deserves more attention is whether a message updates valid information, eliminates ambiguity in the division of labor, and whether the recipient takes actions accordingly.
Beyond completing the task, we need to further ask: which actions are helpful?
Counting only the success rate is not enough to describe the working mode of the team.
Two teams have completed the task, one has a clear division of labor and smooth handover, while the other has experienced a lot of repeated collection, invalid waiting and information misunderstanding. If we only look at the final result, both will get the same success label.
To this end, the paper proposes Causal Collaboration Effectiveness (CCE).
It looks back from the final goal completed by the team, and calculates how many actions are judged to have promoted this result.
The success rate answers "whether the team has accomplished the task", and CCE further asks "in the process of accomplishing this task, which actions are helpful". Its focus on contributions is not limited to the final step of completing the task, but also includes preparatory work that provides materials, conditions or valid information for subsequent actions.
It traces back forward in time starting from the action of completing the goal: which materials does this crafting depend on? Which transfer does the material come from? Which collection or processing does the transfer depend on? The algorithm uses large language models to judge the causal dependencies between candidate actions, and gradually builds an action relationship graph.
Figure 4|Schematic diagram of CCE in the paper: the actions of the lumberjack, carpenter and mage are connected through cross-role dependencies. 12/27 in the figure is the proportion of contributing actions in this schematic trajectory, not the average result of the entire dataset.
The calculation method is not complicated:
CCE of successful tasks = number of actions judged to contribute to success ÷ total number of actions performed by the team.
Taking the schematic trajectory in Figure 4 as an example, the three agents performed a total of 27 actions, 12 of which were included in the contribution chain leading to task success, so CCE = 12 ÷ 27 ≈ 44.4%.
The mage crafting the magic wand is the action that directly completes the goal; the lumberjack collecting and handing over the wood, and the carpenter processing and delivering the wooden stick, may indirectly contribute to success through subsequent dependencies.
Chat can also be a contributing action. If a message provides information that subsequent actions depend on, it may enter this contribution chain; whether it is counted depends on the specific trajectory and judgment rules, not just whether it belongs to "chat" or "tool call".
The paper defines the CCE of failed tasks as 0 by definition. This is the scoring convention of the indicator, which does not mean that every attempt in the failure process is worthless.
Therefore, the cross-task average CCE is affected by both the success rate and the proportion of contributing actions in the success trajectory. It should be read in combination with the success rate and partial success rate, and should not be regarded as an independent "team tacit score".
Here we need to distinguish two concepts: the average action proportion per task is not equal to the overall proportion calculated after aggregating all task actions. Actions that are not traced back as success contributions cannot be uniformly regarded as meaningless; necessary exploration, attempts under insufficient information, and judgment errors all need to be considered.
The value of CCE lies in providing a verifiable analytical perspective to help researchers explore the working process of the team. It still relies on model judgment and cannot independently prove that a system already has general collaborative capabilities.
From repeated labor to early work stoppage
Failures occur in the details of collaboration
The analysis of the communication trajectories in the paper summarizes multiple types of failures: repeatedly asking questions that have been solved, misidentifying the role or resource receiver, reporting incorrect item information, and announcing the end of the task before the key steps are completed.
These phenomena show that what multi-Agent systems need to maintain is not only the "task plan", but also the constantly changing team status.
For system developers, AgentWorld provides scenarios for further experiments: can clear role division reduce misunderstandings? Can shared memory help the team synchronize progress? What are the respective advantages of centralized planning and distributed execution? Does the new communication mechanism actually improve the completion rate, or only increase the amount of messages?
All these should be answered through controlled experiments. The role of a collaboration benchmark is to enable method improvements to be based on the same tasks, the same rules, and verifiable execution records.
From "each being capable" to "completing things together"
The experiments of AgentWorld are not sufficient to attribute collaboration difficulties to the general upper limit of all large models. The results correspond to specific models, prompts, tool interfaces and execution protocols, and there is still room to explore more abundant planning, memory or communication designs.
But it concretizes an important question: to measure an Agent team, we should not only look at what the members can do, but also see whether they can jointly complete things under the condition of mutual dependency.
As multi-agent systems are used in