Installs into .claude/skills of the current project.
Are you the author of Arxiv 2609 31590 Agentworld Benchmark Long Horizon Collaboration?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-arxiv-2609-31590-agentworld-benchmark-long-horizon)
---
name: arxiv-2609-31590-agentworld-benchmark-long-horizon-collaboration
description: 'Research paper: AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs.'
metadata:
openclaw:
emoji: "π"
tags: ["research", "arxiv", "multi-agent-rl", "benchmark", "collaboration", "long-horizon", "llm"]
---
# AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
**arXiv ID:** 2609.31590
**Categories:** cs.AI, cs.MA
**Utility Score:** 0.84 (promoted after abstract review)
## Abstract
Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated long-horizon collaborative tasks requiring sustained multi-agent coordination over extended interaction sequences. The benchmark is designed to test genuine collaboration capabilities rather than competitive or short-horizon interactions, filling a critical gap in multi-agent LLM evaluation.
## Key Contributions
1. **Long-Horizon Focus**: Addresses the gap in benchmarks that only test short-horizon (<20 steps) or competitive interactions
2. **Genuine Collaboration**: 100 human-annotated tasks specifically designed to test genuine multi-agent collaboration
3. **Sustained Coordination**: Tasks require extended interaction sequences testing sustained coordination capabilities
4. **Collaboration Isolation**: Benchmark design isolates collaboration capabilities from individual agent performance
## Relevance to AI Systems
- **Benchmark Gap**: Fills critical gap in multi-agent LLM evaluation for long-horizon collaborative tasks
- **Evaluation Standard**: Provides standardized benchmark for assessing multi-agent collaboration capabilities
- **Practical Impact**: Directly applicable to evaluating real-world multi-agent systems requiring sustained coordination