GUIDE: Reinforcement Learning for Behavioral Action Support in Type 1 Diabetes
Type 1 diabetes (T1D) management requires continuous adjustment of insulin and lifestyle behaviors to maintain blood glucose within a safe target range. Although automated insulin delivery (AID) systems have improved glycemic outcomes, many patients still fail to achieve recommended clinical targets. Current reinforcement learning (RL)-based methods focus primarily on insulin-only treatment and do not provide behavioral recommendations for glucose control. To address this gap, we propose GUIDE, an RL-based decision-support framework designed to complement AID technologies by providing structured behavioral recommendations defined by intervention type, magnitude, and timing, including bolus insulin administration and carbohydrate intake events. GUIDE integrates a patient-specific glucose predictor trained on real-world continuous glucose monitoring data and supports offline and online RL algorithms within a unified environment. The algorithms are evaluated across 25 individuals from the AZT1D dataset and 12 individuals from the OhioT1DM dataset. Among the evaluated algorithms, CQL-BC achieved the best performance on both datasets, with mean time-in-range values of 84.18 $\pm$ 19.89% on AZT1D and 77.64 $\pm$ 8.81% on OhioT1DM. It also maintained time-below-range values of 0.43 $\pm$ 1.27% and 2.81 $\pm$ 3.56%, respectively. Behavioral analysis yielded mean cosine similarities of 0.767 $\pm$ 0.128 on AZT1D and 0.774 $\pm$ 0.162 on OhioT1DM, indicating that the learned policy preserves key structural characteristics of patient action patterns. These findings demonstrate the potential of conservative offline RL with a structured behavioral action space to provide personalized and behaviorally plausible decision support for diabetes management.