Search arXiv⌕ Search

arXiv · 2610.00863

Correctness, Convergence, and AI-Generated Code Detection: A Longitudinal Study of Student and Large Language Model Code in Introductory Programming

Abstract

Large language models can generate plausible solutions to programming assignments, making it tempting to detect their use by matching student code against a reference bank of generated solutions. Yet similar code can also arise when an assignment admits only a few natural implementations, which leaves open what a match actually shows. We investigate generated-reference matching using 29,970 student submissions from ten Python labs offered in 2021, 2023, and 2025, together with 90,000 solution attempts generated retrospectively by three frontier LLMs. We validate the generated solutions using hidden instructor tests, compare code with MOSS after excluding the starter code, and examine the exact abstract syntax tree (AST) forms of selected functions. The models usually produced correct solutions and, across most assignments, converged on similar implementations. Student submissions matched the generated references more often in later cohorts, including among submissions that passed every hidden test. On tightly specified functions, the models converged on a few exact abstract-syntax-tree forms, and the number of distinct student forms also declined across cohorts, whereas open-ended functions remained diverse in both sources. Most reported overlaps were short, making the minimum match length an important choice when reviewing students' code. Finally, we discuss how instructors can build a reference bank of generated solutions before releasing an assignment to identify tasks on which generated solutions converge, decide how much review a match warrants, and redesign tasks to elicit tests, reasoning, and intermediate work. These findings support tracking population-level changes in submitted code, while attributing AI use to an individual submission would require additional evidence about how it was produced, such as prompts, revisions, intermediate code, and student disclosures.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Runlong Ye, Jing Fan, Angela Zavaleta Bernuy, Oscar Karnalim, Paul Denny, Juho Leinonen, Michael Liut. 2026-10-01. Correctness, Convergence, and AI-Generated Code Detection: A Longitudinal Study of Student and Large Language Model Code in Introductory Programming. https://doi.org/10.1145/3856208.3856226

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

ECHO: A Participatory Framework for Bias-Anchored AI Harm Anticipation

Artificial Intelligence (AI) systems increasingly shape consequential decisions, creating value but also potential harms for individuals, social groups, and society. This has prompted calls for proactive approaches that anticipate harms early in the AI lifecycle. Although prior research identifies AI biases as sources of harm, the associations between particular lifecycle biases and harms remain insufficiently understood. We introduce \texttt{ECHO}, a systematic, context-sensitive, and participatory framework that anchors early harm anticipation in lifecycle biases and elicits their perceived associations with potential harms.\texttt{ECHO} identifies domain-specific stakeholders, instantiates biases through vignettes, collects harm judgements from human participants and a large language model (LLM), and organises them into descriptive and inferential ethical matrices. Applied to disease diagnosis and hiring, \texttt{ECHO} surfaced non-uniform, context-sensitive bias--harm patterns indicating which harms were perceived as plausible consequences of particular AI biases. The theoretical interpretability of these patterns and the inferential support for specific associations strengthen the plausibility of the mappings. By linking stakeholder-specific anticipated harms to lifecycle biases, \texttt{ECHO} supports source-level harm anticipation and provides structured input to subsequent AI governance actions

cs.CY↗

Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data

How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key finding is that while leading LLMs are increasingly being used to scale, replicate and replace human survey responses in research, these advances only show an increased capacity to parrot conventional wisdom in harmony with each other rather than revealing novel findings. If synthetic respondents are used in future research, we need more replicable validation protocols and reporting standards for when and where synthetic survey data can be used responsibly, a gap that this paper fills. Our results suggest that synthetic survey responses cannot meaningfully model real human social beliefs within organizations, particularly in contexts lacking previously documented evidence. We conclude that synthetic survey-based research should be cast not as a substitute for rigorous survey methods, but as an increasingly reliable pre- or post-fieldwork instrument for identifying societal assumptions, conventional wisdoms, and other expectations about research populations.

cs.CY↗

Clinical Note Bloat Reduction for Efficient LLM Use

Background: Clinical notes contain extensive duplicated text from templates, copy-paste, and auto-populated fields ("note bloat"), diluting clinical signal, limiting longitudinal context, and increasing large language model (LLM) costs. Methods: TRACE removes note bloat using note-level EHR metadata to identify templated and copied content, with frequency-based de-duplication when metadata are unavailable. We evaluated TRACE using blinded physician span review and gold-standard templated-text annotations across four cohorts spanning liver transplant, obstetrics, and inpatient populations at multiple health systems (5.3M notes). We compared zero-shot LLMs and embedding-based classifiers using original and TRACE-processed notes for 20 information extraction tasks and prediction of 5-year survival, postpartum hemorrhage, and 30-day readmission. Results: Only 0.3-6.6% of removed text was flagged as author-generated; TRACE captured 86% of annotated templated characters. Information extraction F1 differences averaged by cohort ranged from -0.009 to +0.004; task-specific prediction F1 differences ranged from -0.011 to +0.018. Among 1,000 randomly sampled Stanford Health Care patients, TRACE reduced chart text by 47.3% (742.7M characters), averaging 220,167 fewer tokens per patient. Using 2024 encounter volumes at a large tertiary academic center and one query per encounter, projected three-year net savings ranged from $1.00M to $13.58M across evaluated model pricing schemes, including initial and annual TRACE processing costs. Conclusion: TRACE substantially reduces clinical note redundancy while preserving information extraction and prediction performance. Underused EHR metadata can reduce LLM inference costs, expand usable longitudinal context, and support scalable clinical AI.

cs.CY↗