Search arXiv⌕ Search

arXiv · 2009.03449

Survival Analysis via Ordinary Differential Equations

Abstract

This paper introduces an Ordinary Differential Equation (ODE) notion for survival analysis. The ODE notion not only provides a unified modeling framework, but more importantly, also enables the development of a widely applicable, scalable, and easy-to-implement procedure for estimation and inference. Specifically, the ODE modeling framework unifies many existing survival models, such as the proportional hazards model, the linear transformation model, the accelerated failure time model, and the time-varying coefficient model as special cases. The generality of the proposed framework serves as the foundation of a widely applicable estimation procedure. As an illustrative example, we develop a sieve maximum likelihood estimator for a general semi-parametric class of ODE models. In comparison to existing estimation methods, the proposed procedure has advantages in terms of computational scalability and numerical stability. Moreover, to address unique theoretical challenges induced by the ODE notion, we establish a new general sieve M-theorem for bundled parameters and show that the proposed sieve estimator is consistent and asymptotically normal, and achieves the semi-parametric efficiency bound. The finite sample performance of the proposed estimator is examined in simulation studies and a real-world data example.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Weijing Tang, Kevin He, Gongjun Xu, Ji Zhu. 2021-12-05. Survival Analysis via Ordinary Differential Equations. https://arxiv.org/abs/2009.03449

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PK/PD-integrated Bayesian platform design for phase II dose regimen optimization

Early-phase dose-finding methods increasingly assess toxicity and efficacy jointly, but comparisons based only on administered dose may inadequately characterize regimens differing in schedule. We developed a Bayesian phase II adaptive platform design for regimen optimization that integrates pharmacokinetic/pharmacodynamic (PK/PD) modelling into toxicity, efficacy, regimen selection and adaptation decisions. The proposed PK/PD-informed Regimen Optimization Platform (PROP) design uses a population PK/PD model to generate patient- and population-level predictions of exposure and biological activity. Acute and cumulative toxicities are analysed using a discrete-time time-to-event model informed by PK exposure. Efficacy is evaluated through Bayesian model averaging of exposure-driven and biomarker-driven time-to-event models. The design supports regimen graduation, discontinuation for futility or safety, and addition of unexplored regimens. Performance was evaluated through simulations motivated by an influenza intensive-care setting. Across six scenarios, PROP generally improved graduation and futility decisions, reduced inappropriate graduation, and supported the addition of promising regimens compared with dose-based alternatives. It also more accurately estimated regimen-specific toxicity and arm-specific efficacy, while the model-averaging framework favored the efficacy model consistent with the data-generating mechanism. Dose-based approaches performed better for safety stopping in some scenarios, despite less accurate characterization of the regimen--toxicity relationship. PK/PD-informed platform designs can improve adaptive regimen selection and knowledge generation when dose alone cannot adequately characterize treatment regimens.

stat.ME↗

Towards more plausible point-identifying assumptions in two-sample Mendelian randomization

Two-sample Mendelian randomization (MR) is a widely applied methodology in epidemiology. In two-sample MR, summary data (typically, regression coefficients and standard errors) quantifying the association between multiple genetic variants and the exposure and the outcome are used in an instrumental variable framework aimed at estimating the causal effect of the exposure on the outcome. Most two-sample MR methods were developed under data-generating models where the association of for each candidate genetic instrument with the exposure, as well as the causal effect of the exposure on the outcome, are constant in the additive scale. These assumptions are useful because they imply that, had all genetic variants been valid IVs, they would all estimate the same causal parameter - namely, the constant causal effect. We refer to this condition as summary-level homogeneity. However, these are rather strong homogeneity conditions which may raise concerns about the plausibility of these methods in practice. In this paper, we show that summary-level homogeneity is implied by the following conditions: the causal effect is additive linear, but not necessarily constant across, all strata of the population; and uncorrelatedness between heterogeneity in the causal effect and in the association between each genetic variant and the exposure. Under these conditions, typical two-sample MR methods can be interpreted as estimators of the average causal effect. These results clarify that point-identifying assumptions required for two-sample MR methods are weaker than previously anticipated, which contributes to their plausibility and interpretation in at least some practical applications.

stat.ME↗

Amortized Bayesian Disease Mapping and Boundary Detection on Heterogeneous Spatial Graphs

Spatial disease maps help public-health researchers identify geographic inequalities, but standard Bayesian smoothing can obscure localized disparities when neighboring communities have sharply different socioeconomic or behavioral profiles. Analysts therefore need to determine where smoothing should be interrupted and repeat that analysis as maps, adjacency structures, and outcomes change. We develop a covariate-informed Bayesian boundary model and an amortized posterior approximation trained across heterogeneous areal graphs. The model distinguishes local interruptions in smoothing from broader residual spatial dependence; the trained approximation handles maps with different numbers of regions. Simulations examine posterior calibration, boundary-probability recovery, replicated-data behavior, and MCSE-controlled agreement with prior-matched MCMC. In contrast to traditional approaches that analyze these data separately, we demonstrate the effectiveness of using a single trained deep learning network to analyze respiratory hospitalizations in Greater Glasgow; lung cancer incidence in California; and tracheal, bronchial, and lung cancer mortality in South Korea, comprising 58 to 241 regions. Selected boundary density is greatest in Glasgow and lowest in South Korea despite substantial residual spatial dependence in both, showing that local interruption and broader spatial persistence need not vary together. Across all three applications, edge-level boundary probabilities agree substantially with dataset-specific analyses, although posterior spread and thresholded boundary sets differ. These results support reusable Bayesian boundary analysis across the evaluated disease-map class and identify the validation needed before deployment to new applications.

stat.ME↗