Search arXivSearch

arXiv subjects

Chenyao Yu

Publications and source records attributed to Chenyao Yu.

4 recordsLinked to original sources

The Fidelity and Feedback Traps: The Case for Health Digital Twins as Modular Evolving Causal Systems

Digital twins for health may be used to compare treatments, project patient trajectories, and support clinical decisions. While related to mechanical digital twins, those initially developed for engineering applications, replicating the mechanical digital twin architecture and goals may fail in health for two reasons. The fidelity trap is the belief that an accurate model can answer what-if questions by virtue of its accuracy. Prediction and counterfactual reasoning are different tasks, and a twin that can fit past trajectories well may miss the mark when ranking treatments. The feedback trap arises when the twin updates on data its own recommendations helped generate. Refitting in this way can recover a biased relationship and grow more confident even as data grows thinner. We contend that health digital twins should be conceived as causally valid, modular, and evolving systems. Modularity isolates the data and models needed for interventional recommendations, causal validity supports such claims, and governed evolution updates the twin while accounting for how its recommendations reshape the data. We conclude that the standard for a health twin should be how well it supports decisions in the world it helps create, not how faithfully it reproduces the world it observes.

stat.OT

A Review of Methods and Practices for Missing Data in Sequential Multiple Assignment Randomized Trials (SMARTs): An Ancillary Study of a Scoping Review

Background: Missing data poses an acute threat to sequential multiple assignment randomized trial (SMART) analyses because of the sequential treatment structure and response-dependent re-randomization. Objectives: This study aimed to (1) review the current statistical methods for handling missing data in SMARTs, and (2) characterize how missing data is reported and handled in published SMARTs. Methods: We conducted a narrative review of statistical methods developed for missing data in SMARTs. Additionally, we conducted a pre-specified secondary extraction of a previously published scoping review of SMARTs focused on missing data. Extraction captured attrition rates, methods for handling missingness, and planned versus performed missing data analyses. Results: Seven methodological papers were identified; nearly all assume missing at random (MAR), and only one addresses the full set of SMART-specific missingness types. Across 30 published SMARTs, median overall attrition was 18.1% (range 0.6%-56.5%). Methods used to address missing data were described in 80% of the manuscripts; mixed-model methods were most common (30%). Among 14 studies with paired protocols, sensitivity analyses were pre-specified in 2 (14%). Conclusions: SMART-specific methodology for missing data is limited, and a substantial gap exists between available methodology and current SMART practice.

stat.ME

Nonparametric estimation of FBSDEs with random terminal time

This paper investigates the nonparametric estimation of the functional coefficients of the FBSDEs with random terminal time, including the local constant and local linear estimators. We provide complete two-dimensional asymptotics in both the time span and the sampling interval, allowing for the precise characterization of their distribution. Moreover, the empirical likelihood (EL) method to construct the data-driven confidence intervals for these estimators is provided. Some numerical simulations investigate the finite-sample properties of the estimators and compare the performance of the EL method and the conventional method in constructing confidence intervals based on asymptotic normality.

math.ST

VRSO: Visual-Centric Reconstruction for Static Object Annotation

As a part of the perception results of intelligent driving systems, static object detection (SOD) in 3D space provides crucial cues for driving environment understanding. With the rapid deployment of deep neural networks for SOD tasks, the demand for high-quality training samples soars. The traditional, also reliable, way is manual labelling over the dense LiDAR point clouds and reference images. Though most public driving datasets adopt this strategy to provide SOD ground truth (GT), it is still expensive and time-consuming in practice. This paper introduces VRSO, a visual-centric approach for static object annotation. Experiments on the Waymo Open Dataset show that the mean reprojection error from VRSO annotation is only 2.6 pixels, around four times lower than the Waymo Open Dataset labels (10.6 pixels). VRSO is distinguished in low cost, high efficiency, and high quality: (1) It recovers static objects in 3D space with only camera images as input, and (2) manual annotation is barely involved since GT for SOD tasks is generated based on an automatic reconstruction and annotation pipeline.

cs.CV