arXiv · 2609.27063
Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses
Abstract
This paper presents a multi-source classroom study conducted during a 10-week quarter in data science courses at Drexel University. We first investigate the behaviors of student engagement with large language models (LLMs) using four surveys across three data science courses. Second, we evaluate the construct validity of multiple-choice questions (MCQs) generated by an LLM for in-lecture retrieval practice. Based on 378 authored questions (311 deployed, producing 7{,}888 student responses), we analyze whether the difficulty ratings assigned by an LLM match empirical item difficulty. Our study shows that student engagement with LLMs varied across courses and increased over the term. Although students expressed high satisfaction and reported saving considerable time, their perception of deep learning benefits declined, and many noted a tendency toward over-reliance. Regarding the difficulty ratings of LLM-generated MCQs, the Easy, Medium, and Hard labels correlated closely with its assigned Bloom's Taxonomy levels (Spearman $ρ=0.90$), reflecting an artifact of co-generation. However, neither metric predicted empirical item difficulty (difficulty label $ρ=0.06$; Bloom level $ρ=0.02$). The ratings reflect the structural formatting of a question rather than its underlying difficulty.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuan An, Lei Wang. 2026-09-22. Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses. https://arxiv.org/abs/2609.27063
Cite the original work for its findings. Save a collection to share your selection of sources.