Agent Skill Evolution: How Revisions Affect Coding Agents
Agent Skills, the SKILL.md files that tell an LLM coding agent how a project works, are revised like code, yet what a revision does to the agent is unknown. From 2,608 first/last revision pairs of 3,159 Skills, we characterize how Skills evolve and how they change together with the configuration of the agent's harness. We then focus on rule changes, revisions that add or remove a rule we can check automatically, such as "run allium check". We measure their effect on 21 models in single answers and on four agents in a sandbox, and their cost on 20 of these models and the four agents. Most revisions (55%) change a rule or procedure, and commits that revise a Skill change harness files such as CLAUDE.md more often than other commits of the same size. Across 16 open-weight models, an added rule raises compliance in a single answer by +0.41 on average. Across the four agents, the rate at which the agent takes the required action rises by +0.23 on average (+0.16 to +0.36), and for the three agents that blind judges assessed, final correctness rises by +0.10 on average (+0.06 to +0.14). The gain comes mainly from rules that name a command or path the old Skill did not mention. Real tools load a Skill's body only when the agent decides it needs it. In that setting the four agents keep about half of the action gain on average (51%), and the three open models about 38%. A revision adds 18-19% input tokens to a single answer and no detectable cost to an agent episode, while loading a Skill's body raises the tokens of an episode by 50% on average.