Search arXiv⌕ Search

arXiv subjects

Siyuan Xiong

Publications and source records attributed to Siyuan Xiong.

2 recordsLinked to original sources

MASRubric: Auditing Information Flow in Multi-Agent Systems with Failure-Distilled Pitfall Rubrics

While multi-agent systems (MAS) excel at complex reasoning, they are vulnerable to errors that intermediate agents introduce and downstream agents build upon. Auditing intermediate messages before they propagate requires an explicit standard, yet evaluation rubrics are typically authored by domain experts or written against a reference answer, neither of which is available for an unseen message at test time. We present MASRubric, a MAS information flow auditing framework with failure-distilled pitfall rubrics. Offline, trajectories on which the MAS has failed are automatically distilled into a reusable bank of pitfall criteria, each describing a recurrent error by its underlying misconception, the reasoning situations in which it arises, and the check that would expose it. Online, the criteria applicable to each intermediate message are retrieved from this off-the-shelf bank and checked one by one, and the resulting satisfaction rate decides whether the message is broadcast, returned to its author with diagnostic feedback for revision, or withheld. Empirical results demonstrate that MASRubric enhances MAS performance on both fixed and dynamic frameworks, achieving average accuracy gains of up to 2.83 points on math reasoning benchmarks and 1.74 points on code generation benchmarks. Further analysis shows that the retrieved criteria vary systematically with task types, and that the audit effort tracks task difficulty. Moreover, the bank transfers without re-mining to a system with a stronger backbone, which makes more adaptive and more efficient use of it. Our code and dataset are released at https://github.com/TonySY2/MASRubric.

cs.AI↗

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real-world tasks. The rapid capability enhancement of open-source models deployable on consumer-grade GPUs presents a compelling opportunity to drastically reduce these costs by leveraging skill-based behavioral constraints. Nevertheless, automatically generating effective skills tailored specifically for such compact models remains a significant practical challenge. To address this, we propose SKILLER, a natural-language-driven reinforcement learning framework designed to automatically generate executor-specific skills for small models, which employs a strong model as the actor and critic, treats the small-model agent system as the environment, and propagates all reinforcement learning signals entirely via natural language. Extensive experimental evaluations across five relevant benchmarks using Qwen3.5-9B and Qwen3.5-4B demonstrate that SKILLER outperforms three open-source and one closed-source skill generation or evolution methods, achieving absolute gains ranging from 4.3 to 20.4 percentage points for the 9B model and 1.8 to 13.3 points for the 4B model, while remarkably matching the performance of strong closed-source models on single-skill tasks in SkillsBench. The project is available at https://github.com/DANG-ai/SKILLER.

cs.AI↗