Search arXiv⌕ Search

arXiv subjects

Yunhan Qiao

Publications and source records attributed to Yunhan Qiao.

5 recordsLinked to original sources

Faster but Not Wiser: GitHub Copilot Decouples Programming Performance from Code Comprehension in Brownfield Tasks

Teaching Computer Science (CS) students to comprehend and maintain existing codebases is a critical challenge in software engineering education. Although Generative AI (GenAI) assistants such as GitHub Copilot can improve task completion speed and correctness, their relationship with code comprehension remains unclear. We conducted a within-subjects study with 15 CS graduate students who completed feature-implementation tasks in an unfamiliar codebase with and without Copilot. Despite significant performance improvements with Copilot, participants showed no corresponding improvement in overall comprehension ($p=0.59$), and performance gains were not significantly associated with comprehension gains. Exploratory category-level estimates were positive for identifying what and where to modify ($ρ=0.50$) and negative for explaining how the existing code worked and predicting the effects of a change ($ρ=-0.57$); however, neither remained significant after correction for multiple comparisons. Our behavioral analysis showed that participants with higher comprehension engaged more frequently in verification loops, repeatedly inspecting and revising code. They performed write-then-view transitions 4.7 times more often than participants with lower comprehension ($p=0.001$). These findings show that successful task completion does not reliably indicate code comprehension and suggest that how students engage with AI-generated code may matter for their resulting understanding. We argue for programming education that assesses correctness and comprehension separately, teaches students to inspect and explain AI-generated code, and encourages GenAI tools that support active verification.

cs.SE↗

Beyond the Prompt: Linking What Developers Ask, Do, and Understand with Coding Agents

Coding agents can now change code for developers, who describe goals, supply context, and respond to the agent's work. Yet prompts, screen activity, and task success each tell only part of this story. We present Say, Do, Understand, an end-to-end workflow for analyzing what developers write to an agent, what they do while it works, and what they can explain afterwards. The workflow has five stages (Capture, Prepare, Analyze, Integrate, and Interpret) and three instruments: a prompt codebook, a scheme for coding screen-recorded activities and events, and separate rubrics for explaining the process and the solution. We applied it in an observational study of ten experienced developers who used GitHub Copilot on an unfamiliar codebase. Crossing task performance with understanding produced four personas. The two measures agreed for eight developers but split for two: one passed most tests but could not explain the solution, and another passed few tests but explained it well. In this sample, the personas that most often asked the agent to check its work spent the least time testing on their own. These patterns are descriptive and do not generalize beyond the sample. We recommend the workflow to computing educators, industry practitioners, and researchers to adapt and evaluate human--AI communication in software engineering.

cs.SE↗

A Systematic Literature Review of the Use of GenAI Assistants for Code Comprehension: Implications for Computing Education Research and Practice

The ability to comprehend code has long been recognized as an essential skill in software engineering. As programmers lean more heavily on generative artificial intelligence (GenAI) assistants to develop code solutions, it is becoming increasingly important for programmers to comprehend GenAI solutions so that they can verify their appropriateness and properly integrate them into existing code. At the same time, GenAI tools are increasingly being enlisted to provide programmers with tailored explanations of code written both by GenAI and humans. Thus, in computing education, GenAI presents new challenges and opportunities for learners who are trying to comprehend computer programs. To provide computing educators with evidence-based guidance on the use of GenAI to facilitate code comprehension and to identify directions for future research, we present a systematic literature review (SLR) of state-of-the-art approaches and tools that leverage GenAI to enhance code comprehension. Our SLR focuses on 31 studies published between 2022 and 2024. Despite their potential, GenAI assistants often yield inaccurate or unclear explanations, and novice programmers frequently struggle to craft effective prompts, thereby impeding their ability to leverage GenAI to aid code comprehension. Our review classifies GenAI-based approaches and tools, identifies methods used to study them, and summarizes the empirical evaluations of their effectiveness. We consider the implications of our findings for computing education research and practice, and identify directions for future research.

cs.SE↗

The Effects of GitHub Copilot on Computing Students' Programming Effectiveness, Efficiency, and Processes in Brownfield Programming Tasks

When graduates of computing degree programs enter the software industry, they will most likely join teams working on legacy code bases developed by people other than themselves. In these so-called brownfield software development settings, generative artificial intelligence (GenAI) coding assistants like GitHub Copilot are rapidly transforming software development practices, yet the impact of GenAI on student programmers performing brownfield development tasks remains underexplored. This paper investigates how GitHub Copilot influences undergraduate students' programming performance, behaviors, and understanding when completing brownfield programming tasks in which they add new code to an unfamiliar code base. We conducted a controlled experiment in which 10 undergraduate computer science students completed highly similar brownfield development tasks with and without Copilot in a legacy web application. Using a mixed-methods approach combining performance analysis, behavioral analysis, and exit interviews, we found that students completed tasks 35% faster (p < 0.05) and made 50% more solution progress p (< 0.05) when using Copilot. Moreover, our analysis revealed that, when using Copilot, students spent 11% less time manually writing code (p < 0.05), and 12% less time conducting web searches (p < 0.05), providing evidence of a fundamental shift in how they engaged in programming. In exit interviews, students reported concerns about not understanding how or why Copilot suggestions work. This research suggests the need for computing educators to develop new pedagogical approaches that leverage GenAI assistants' benefits while fostering reflection on how and why GenAI suggestions address brownfield programming tasks. Complete study results and analysis are presented at https://ghcopilot-icer.github.io/.

cs.SE↗

Automatically Detecting Heterogeneous Bugs in High-Performance Computing Scientific Software

Scientific advancements rely on high-performance computing (HPC) applications that model real-world phenomena through simulations. These applications process vast amounts of data on specialized accelerators (eg., GPUs) using special libraries. Heterogeneous bugs occur in these applications when managing data movement across different platforms, such as CPUs and GPUs, leading to divergent behavior when using heterogeneous platforms compared to using only CPUs. Existing software testing techniques often fail to detect such bugs because either they do not account for platform-specific characteristics or target specific platforms. To address this problem, we present HeteroBugDetect, an automated approach to detect platform-dependent heterogeneous bugs in HPC scientific applications. HeteroBugDetect combines natural-language processing, off-target testing, custom fuzzing, and differential testing to provide an end-to-end solution for detecting platform-specific bugs in scientific applications. We evaluate HeteroBugDetect on LAMMPS, a molecular dynamics simulator, where it detected multiple heterogeneous bugs, enhancing its reliability across diverse HPC environments.

cs.SE↗