arXiv · 2610.04267
AI-Enabled Quality Assurance for Multiple-Choice Assessment Items
Abstract
Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing. Database searches, citation retrieval, and nominated sources yield fourteen research reports reviewed in full text. We distinguish surface checks from content-sensitive judgments and map a 19-criterion rubric to detection methods and reported evidence. High label-level accuracy often coexists with weak positive case detection, while rubric definitions and reference standards vary. Revision evidence is mixed, and the associations reported in prior work do not establish the effects of repair. We propose evaluating quality assurance as a sequence of independently validated decisions, with criterion-specific reporting, calibrated human review, and outcome-based assessment of revisions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Steven Moore, Nicholas Diana. 2026-10-03. AI-Enabled Quality Assurance for Multiple-Choice Assessment Items. https://arxiv.org/abs/2610.04267
Cite the original work for its findings. Save a collection to share your selection of sources.