Search arXivSearch

arXiv · 2406.07481

The end of multiple choice tests: using AI to enhance assessment

Abstract

Effective teaching relies on knowing what students know-or think they know. Revealing student thinking is challenging. Often used because of their ease of grading, even the best multiple choice (MC) tests, those using research based distractors (wrong answers) are intrinsically limited in the insights they provide due to two factors. When distractors do not reflect student beliefs they can be ignored, increasing the likelihood that the correct answer will be chosen by chance. Moreover, making the correct choice does not guarantee that the student understands why it is correct. To address these limitations, we recommend asking students to explain why they chose their answer, and why "wrong" choices are wrong. Using a discipline-trained artificial intelligence-based bot it is possible to analyze their explanations, identifying the concepts and scientific principles that maybe missing or misapplied. The bot also makes suggestions for how instructors can use these data to better guide student thinking. In a small "proof of concept" study, we tested this approach using questions from the Biology Concepts Instrument (BCI). The result was rapid, informative, and provided actionable feedback on student thinking. It appears that the use of AI addresses the weaknesses of conventional MC test. It seems likely that incorporating AI-analyzed formative assessments will lead to improved overall learning outcomes.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Michael Klymkowsky, Melanie M. Cooper. 2024-06-11. The end of multiple choice tests: using AI to enhance assessment. https://arxiv.org/abs/2406.07481

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

When Technically Plausible Advice Is Unsafe: A Cross-Ecosystem Measurement of Online Support for Technology-Facilitated Abuse

Technology-facilitated abuse (TFA) creates an adversarial setting where sound cybersecurity advice can be unsafe: changing credentials or resetting devices may alert an abuser, destroy evidence, or increase escalation risk. Victims seek guidance from search engines, peer forums, and conversational AI, often evaluated for relevance and correctness rather than contextual safety. We measure whether these sources meet victims' needs. From a decade of r/Stalking narratives, we construct 2,797 victim-derived queries spanning 11 misuse categories. We analyze 27,162 Google webpages, 2,476 Reddit query--thread responses, and 250 responses from three general-purpose LLMs and two survivor-support chatbots. Our framework measures technical quality and damaging guidance, plus secondary-link integrity on webpages, toxicity on Reddit, and trauma-informed support in conversational systems. We find failures & risks that relevance, accuracy, or actionability alone do not capture. Web Search and conversational systems frequently return relevant information; Reddit responses are less consistently relevant and actionable. In our evaluated accuracy sample, 17.3% of webpages, 13.3% of Reddit threads, and 19.6% of conversational AI responses contained damaging guidance. Further, 65.5% of victim queries led to a webpage with a secondary URL flagged by multiple VirusTotal engines, over 20% received a toxic Reddit comment, and every conversational system produced guidance that overlooked escalation risk. Specialization did not guarantee better support: HopeChat underperformed general-purpose LLMs on several dimensions, while Ruth remained limited in trauma-informed support. These findings expose a gap between technical quality and contextual safety. Safe TFA assistance requires risk-aware recommendations, trustworthy sources, uncertainty communication, and human support, beyond technically plausible answers.

cs.CY

Who Acts, Who Knows, Who Answers? A Corpus-Assisted Discourse Analysis of Agency, Epistemic Responsibility, and Accountability in Generative AI Higher Education Research

Generative artificial intelligence (GenAI) is increasingly described in higher education as a tool, collaborator, evaluator, proxy, and infrastructure. These labels are not neutral: they shape who is seen as acting, knowing, and, crucially, being answerable when someone uses AI. This study examines how scholars distribute such responsibility across 366 English-language titles and abstracts published between November 2022 and 14 August 2026. A corpus-assisted discourse analysis combined frequency counts, five-word collocations, actor-predicate associations, obligation-clause coding, and heuristic analysis of passives, nominalisations, and reporting metonymies. The final corpus contained 91,405 tokens, drawn from 8,734 unique records screened after deduplication. AI was the most frequently named actor, appearing 2,050 times and entering 447 predicate associations, most involving action. Students, by contrast, were more often associated with knowing, judging, and verifying, yet no sentence ever made AI explicitly responsible. Responsibility fell largely on educators, institutions, and policy, or disappeared through passive and nominalised constructions. Even frequent phrases such as "responsible AI" and "responsible use" rarely identify who is accountable for what. The findings reveal a persistent gap between giving AI functional agency and assigning normative accountability.

cs.CY

What fidelity metrics miss: a structural check on synthetic educational data

Secondary use of educational records is increasingly mediated by platforms that share a differentially private synthetic version of a dataset and validate specific findings against the real data on request. The synthetic version is evaluated by comparing summary statistics of each variable, yet reported confirmation rates suggest that such comparisons do not predict which findings survive. We propose a structural check: the number of connected components of a weekly proximity graph over learners, tracked across a term. Across four annual cohorts of lower-secondary study-habit logs, the synthetic versions reproduced the level of this quantity and the shape of the weekly partition, but its variation across the term was between 2.6 and 4.9 times smaller than in the real data at a common working point, without exception, and those changes fell in different weeks: the synthetic cohorts single out the term's examination weeks and the real cohorts do not. We also show that a routine rule for setting the graph threshold makes naive comparisons between two datasets invalid, and illustrate this with an error of our own. The real curves are also distinguishable from marginal-preserving surrogates of themselves in all four cohorts, where three of the four synthetic ones are not, a comparison that needs no real data; these differences trace to what the generator was given.

cs.CY