Search arXivSearch

arXiv · 2604.17250

Improving post-operative discharge destination prediction of geriatric patients with generative data augmentation

Abstract

Data scarcity challenges the development and implementation of innovative healthcare solutions. In geriatrics, fall-related injuries are a major cause of hospitalization, functional decline, and mortality in older adults. Optimizing post-operative discharge planning can mitigate these outcomes, but limited data hinders predictive model development. Here, we explored generative machine learning approaches to augment data from the SURGE-Ahead project (Supporting SURgery with Geriatric Co-Management and AI), an initiative addressing geriatric perioperative care. Data from the German geriatric trauma register (AltersTraumaZentrum; ATZ) were incorporated using two strategies: (i) combining SURGE-Ahead and ATZ register data with imputation (ComImp) and (ii) generating synthetic data from SURGE-Ahead alone or combined SURGE-Ahead and the ATZ register datasets with Adversarial random forests (ARF). Predictive models, including multinomial logistic regression, random forest, and a prior-fitted transformer (TabPFN), were trained and evaluated using standard performance metrics: accuracy, area under the receiver operating characteristic curve (ROC AUC), Brier score, and the logistic loss. Random forest and TabPFN performed well (accuracy around 0.84 and AUC around 0.94) and were largely unaffected by augmentation. Logistic regression benefited from augmented data, with predictive performance improving from 0.70 to 0.81 for accuracy and 0.85 to 0.92 for AUC. These results highlight generative data augmentation as a viable approach to enhance simpler predictive models in geriatric care and emphasize the importance of method selection when addressing data scarcity in heterogeneous clinical populations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pegah Golchian, Pauline Maier, Thomas Kocar, Marvin N. Wright. 2026-04-19. Improving post-operative discharge destination prediction of geriatric patients with generative data augmentation. https://arxiv.org/abs/2604.17250

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Probabilistic Modeling Framework for Transient Debris Outcomes in Low Lunar Orbits

This work presents a novel probabilistic framework for the assessment of orbital debris fragment outcomes in the low lunar orbit (LLO) regime over time horizons ranging from a few hours to a hundred days after a debris generation event. The framework develops continuous distributions for the probability of sinking (lunar collision) and non-sinking over these transient time horizons, building upon NASA's Standard Breakup Model to provide insights into the variations in the likelihood of debris outcomes across the LLO regime and offering an alternative to computationally expensive Monte Carlo simulations. The effects of perturbative forces such as solar radiation pressure are used to assess the realization of debris outcomes over varying time horizons, providing an analytical framework that bounds the likelihood of each outcome. Results indicate variations in the probability of sinking over short time horizons depending on the originating location of the debris generation event, as well as a link between the physical characteristics of fragments and their likelihood of sinking over longer time horizons. The future incorporation of this framework into broader-scale orbital environment models, mission risk assessment procedures, and policy development are briefly discussed.

stat.AP

County-Level Heterogeneity in Opioid Harm Reduction and Treatment Effects: A Simulation Modeling Analysis

Opioid overdose deaths remain a severe public health crisis in the US, with heterogeneous burden across counties that differ in epidemic trajectory, baseline resources, and local context. While harm reduction through naloxone distribution and buprenorphine treatment are both evidence-based strategies, limited information on county-level effects hinders the ability of policymakers to prioritize resources across counties. We developed a simulation model of opioid use disorder (OUD), calibrated separately to six Pennsylvania counties spanning large urban (Allegheny, Philadelphia), intermediate-sized (Erie, Dauphin), and rural (Clearfield, Columbia) settings. We projected county-specific overdose mortality trajectories under three levels of increase in buprenorphine dispensing and naloxone distribution (10%, 20%, and 30% above each county's baseline), over a five-year horizon from 2025 to 2029. A 30% increase in naloxone distribution above observed county baseline levels was projected to reduce 2029 overdose deaths by approximately 70% (95% uncertainty interval, UI: 55-81%) in Allegheny County, 11% (95% UI:4-17%) in Erie County, and 28% (95% UI:2-64%) in Clearfield County. Projected reductions in overdose deaths from increasing buprenorphine were consistently smaller (10%-23%), except that in Erie buprenorphine produced larger projected reduction by 20% vs 11% for naloxone. Heterogeneity in naloxone responsiveness was strongly associated with each county's historical naloxone dispensing variability. The same proportional increase in naloxone distribution yields substantially different projected mortality reductions across counties depending on each county's baseline distribution history, a pattern invisible from mortality statistics alone. County-level context is important for informing harm reduction and treatment prioritization at the county level.

stat.AP

Adapting Pairs Trading to Gambling Markets A Case Study of the U.S. Presidential Election

Pairs trading exploits mean reversion in the relationship between related assets. We adapt this idea to political betting markets by modelling the combined implied probability of the two major-party nominees with a latent Ornstein-Uhlenbeck process whose mean-reversion level varies over time and whose observations contain additive noise. Model parameters are estimated from regularly sampled odds data using a state-space likelihood, with consecutive repeated values represented by a single retained observation and the elapsed number of sampling intervals preserved in the continuous-time transition. Parametric-bootstrap upper prediction bounds identify signal times at which the combined implied probability is likely to decline, and a no-intercept Bradley-Terry-type model selects the candidate-specific odds quote. The candidate-selection model is trained on 2020 U.S. presidential-election data and evaluated out of sample on 2024 data. The 2024 analysis produced 130 signals, empirical one-step coverage of 95.1%, a mean synthetic odds-price return of 1.86%, and an unannualized per-trade Sharpe-type ratio of 1.12. These returns are frictionless descriptive quantities rather than executable betting-exchange profits. The results support the integrated framework as a proof of concept for two-candidate electoral markets.

stat.AP