Faster but Not Wiser: GitHub Copilot Decouples Programming Performance from Code Comprehension in Brownfield Tasks
Teaching Computer Science (CS) students to comprehend and maintain existing codebases is a critical challenge in software engineering education. Although Generative AI (GenAI) assistants such as GitHub Copilot can improve task completion speed and correctness, their relationship with code comprehension remains unclear. We conducted a within-subjects study with 15 CS graduate students who completed feature-implementation tasks in an unfamiliar codebase with and without Copilot. Despite significant performance improvements with Copilot, participants showed no corresponding improvement in overall comprehension ($p=0.59$), and performance gains were not significantly associated with comprehension gains. Exploratory category-level estimates were positive for identifying what and where to modify ($ρ=0.50$) and negative for explaining how the existing code worked and predicting the effects of a change ($ρ=-0.57$); however, neither remained significant after correction for multiple comparisons. Our behavioral analysis showed that participants with higher comprehension engaged more frequently in verification loops, repeatedly inspecting and revising code. They performed write-then-view transitions 4.7 times more often than participants with lower comprehension ($p=0.001$). These findings show that successful task completion does not reliably indicate code comprehension and suggest that how students engage with AI-generated code may matter for their resulting understanding. We argue for programming education that assesses correctness and comprehension separately, teaches students to inspect and explain AI-generated code, and encourages GenAI tools that support active verification.