arXiv · 2507.12674
ParaStudent: Closing the Sim2Real Gap in User Simulators for AI Tutor Evaluation
Abstract
Evaluating Artificial Intelligence (AI) tutor feedback before deployment requires anticipating student engagement, typically assessed through real interaction data. We introduce ParaStudent, a fine-tuning framework for simulating novice programming revisions to support AI tutor evaluation. Compared with prompted baselines, ParaStudent's revisions more closely match real student code distributions across functional, stylistic, and semantic metrics. Our best variant achieves AUCs of 0.80 for both feedback relevance and successful uptake when distinguishing streams with real engagement above versus at or below the median, while prompted baselines remain near chance on successful uptake. These findings demonstrate the promise of simulated engagement for pre-deployment feedback triage.
Explore related subjects
Keep this discovery
Rose Niousha, Mihran Miroyan, Abigail O'Neill, Joseph E. Gonzalez, Gireeja Ranade, John DeNero, Narges Norouzi. 2026-09-01. ParaStudent: Closing the Sim2Real Gap in User Simulators for AI Tutor Evaluation. https://arxiv.org/abs/2507.12674
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.