arXiv · 2609.37788
A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
Abstract
Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR.BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo. 2026-09-29. A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses. https://arxiv.org/abs/2609.37788
Cite the original work for its findings. Save a collection to share your selection of sources.