arXiv · 2608.30661
SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?
Abstract
Large language model-based multi-agent systems are evolving from fixed interaction topologies toward dynamically orchestrated Agent Swarms. However, existing benchmarks are still largely based on single-agent or general-purpose agent tasks, making it difficult to systematically evaluate key orchestration capabilities. We propose SwarmBench, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality. Experimental results show that current models exhibit substantial differences in orchestration capability. These differences are reflected not only in final accuracy, efficiency, and cost, but also in the overall quality of the orchestration process itself. Based on these findings, we further propose SwarmExp, a simple yet effective method based on experience extraction and experience replay, which consistently improves the orchestration performance of large language models.
Explore related subjects
Keep this discovery
Jinshan Gao, Zhuoran Jin, Tianyi Men, Kang Liu, Jun Zhao. 2026-08-31. SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?. https://arxiv.org/abs/2608.30661
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.