Jagged AI in Scientific Peer Review: Evidence from POMP Data Analysis
Despite their growing use in academic writing and statistical analysis, the performance of artificial intelligence (AI) tools in scientific peer review remains a largely unexplored area. A key challenge is jagged AI, a phenomenon where AI exhibits strong ability spikes in some domains while remaining deficient in others. To study this jaggedness in a practical data science context, we considered the task of reviewing partially observed Markov process (POMP) data analyses. POMP models, a generalization of state-space models or hidden Markov models, are used to fit mechanistic dynamic models to time series data in diverse applications including disease transmission, ecological dynamics, and financial risk assessment. High-quality peer review in this area entails assessment of scientific context, identification of errors in implementing complex algorithms, and decisions concerning methodological best practices. We studied 72 POMP projects from four semesters of a University of Michigan graduate time series course for which the project reports, the source code, and student peer reviews are anonymized and open access. We compared the human reviews with four AI review agents, using Claude Code with differing instructions implemented as skill files. We found that the AI review agents exhibited a jagged capability profile, identifying plausible technical errors and instances of invalid inference methodology overlooked by humans, while showing lower capability at finding issues involving interpretive errors, narrative coherence, and domain-informed model critique. The jaggedness was similar for all agents, consistent with other research finding that the specific instructions can have limited effect on the capability of AI models for some tasks. Skill file configuration shifted which weaknesses agents emphasized, without removing the jaggedness.