Abstract / Summary
Background: Automated insulin delivery must maintain glycemic control despite partial observability, physiological heterogeneity, and sensing and actuation errors. Cohort-average metrics may nevertheless conceal hypoglycemia in vulnerable virtual patients. Methods: We developed a tail-risk-aware model predictive control framework that combines a Transformer-based latent world model, unguided iterative cross-entropy method planning over perturbed rollouts, and a deterministic Safety Layer that projects insulin actions onto programmed bounds. Reinforcement-learning-guided proposal sampling was evaluated separately as an exploratory search variant. The primary controller was evaluated in Simglucose using 10 virtual patients, three random seeds, and five prespecified operating conditions. Results: Under nominal conditions, it achieved a cohort-average time in range (TIR) of 77.6% and time below range (TBR) of 3.8%, but worst-k TBR was 11.0% and increased to 37.4% under 5% random CGM missingness. In the risk-objective ablation, combining variance penalization with TailMean increased TIR by 2.8 percentage points and reduced cohort-average and worst-k TBR by 2.3 and 7.4 percentage points, respectively, relative to mean-only planning. The Safety Layer modified every proposed action in the analyzed run, indicating strong dependence on the programmed bounds. Conclusions: These results show that average target attainment can coexist with substantial patient-level risk. They support tail-aware evaluation and identify patient-specific safeguards as a priority for future validation, but do not establish clinical safety or efficacy.