arXiv · 2610.00954
Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles
Abstract
As large language models (LLMs) move from standalone assistants into agentic workflows, evaluation must extend beyond scalar leaderboard accuracy to account for operational reliability, cost, latency, and token efficiency. We use an agentic ensemble of small language models (SLMs) with an SLM-judge-mediated feedback loop as a case study for such beyond-leaderboard evaluation. On the 541-prompt IFEval benchmark, the best ensemble achieves 97.34% strict prompt accuracy, exceeding the strongest standalone LLM baseline, gpt-5.4, by 5.81 percentage points while operating in a lower-cost regime. We then analyze the tokenomics and operational behavior behind this gain, including cost per sample, token composition, useful-output goodput, feedback-loop recovery, latency decomposition, and performance across instruction categories and constraint counts. Our results show that agentic SLM ensembles can trade additional test-time tokens and orchestration overhead for improved instruction-following fidelity, motivating multi-dimensional evaluation protocols for future agentic AI systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alexei N. Skurikhin, Emily M. Taylor, Nathan A. DeBardeleben. 2026-10-01. Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles. https://arxiv.org/abs/2610.00954
Cite the original work for its findings. Save a collection to share your selection of sources.