arXiv · 2609.23374
The Evidence Ladder for Reinforcement Learning in Healthcare: From Retrospective Policies to Trusted Interventions
Abstract
Reinforcement learning (RL) offers a natural language for healthcare decisions whose conse- quences unfold over time, yet most reported progress remains far from routine intervention. Ex- isting surveys organize the field by algorithm or clinical application. We instead review healthcare RL through an evidence ladder: problem formulation, retrospective identification, policy estima- tion, stress testing, prospective evaluation, and lifecycle monitoring. This view connects clinical treatment, patient engagement, and health-system operations while exposing a recurring gap: evi- dence that a policy scores well in a historical dataset is not evidence that it will improve care. We synthesize the assumptions and failure modes at each rung, identify what evidence can and can- not transfer across settings, and propose reporting practices for cumulative evaluation. Restless bandits are included as one special case, not as the organizing framework. The central lesson is that healthcare RL should be evaluated as an intervention embedded in a changing sociotechnical system, rather than only as an optimizer of a retrospective reward.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yunfan Zhao. 2026-09-20. The Evidence Ladder for Reinforcement Learning in Healthcare: From Retrospective Policies to Trusted Interventions. https://arxiv.org/abs/2609.23374
Cite the original work for its findings. Save a collection to share your selection of sources.