Search arXivSearch

arXiv subjects

Tim Friede

Publications and source records attributed to Tim Friede.

At least 19 recordsLinked to original sources

Empirical prior distributions for treatment-by-subgroup interaction heterogeneity in random-effects meta-analysis

Subgroup analyses are central to the assessment of benefits and risks, where recommendations may depend on evidence that treatment effects differ across patient groups. Valid subgroup claims require evidence based on (within-trial) interaction estimates while accounting for the heterogeneity in those interaction effects. In the common case of only a few available studies, inference may benefit from the use of prior information on the expected amount of heterogeneity. Although between-study heterogeneity~($\tau$) has been studied empirically for overall treatment effects, no such calibration exists for treatment-by-subgroup interaction effects. We derive empirical (predictive) prior distributions for overall and interaction effect heterogeneity from over 3{,}000 interaction meta-analyses drawn from the \emph{Cochrane Database of Systematic Reviews (CDSR)}. The resulting effect-measure-specific priors indicate that interaction heterogeneity tends to be substantially smaller than treatment effect heterogeneity. We also show that lower precision of within-trial interaction estimates makes interaction heterogeneity harder to identify. Therefore, the use of empirical priors is particularly valuable in sparse interaction meta-analyses. A motivating example illustrates how priors tailored to interaction effects may substantially improve precision in a meta-analysis compared with standard heterogeneity priors.

stat.ME

Subgroup analysis in randomized controlled trials with binary outcomes: dilution and logic-respecting properties

Subgroup analysis is routinely used in randomized controlled trials to examine whether treatment effects are homogeneous across patient subgroups or differ because of treatment-effect heterogeneity. In this paper, we investigate the properties of the odds ratio and the relative response in subgroup analyses with binary outcomes, extending previous work with new theoretical insights and methodological developments. We establish several new theorems that characterize how the odds ratio for the overall population changes in both magnitude and direction when two subgroups are combined. These results further confirm that the odds ratio is inappropriate as an efficacy measure in this subgroup setting, whereas the relative response is appropriate. We also present the formal relationship between the odds ratio and the relative response, and clarify their differences in terms of the logic-respecting property, that is, whether the overall efficacy lies between the subgroup efficacies, and the dilution property, that is, whether mixing subgroups moves the overall odds ratio toward 1. Although the odds ratio is generally not logic-respecting, it may behave approximately like a logic-respecting efficacy measure under certain conditions. To illustrate our findings, we present an illustrative example based on clinical trial data and discuss its implications for subgroup analysis in randomized controlled trials.

stat.ME

Including historical control data in simultaneous inference for pre-clinical multi-arm studies

In pre- and non-clinical toxicology, the reduction of animal use is highly desireable. Although approaches for possible sample size reduction in the concurrent control group were suggested previously under the virtual control groups framework for continuous endpoints, methodology that is applicable to binary outcomes that occur in long-term carcinogenicity studies is currently missing. In order to augment animals in the current control group with historical control data, we propose approaches that rely on dynamic Bayesian borrowing and simultaneous credible intervals for risk ratios. Several operation characteristics such as familywise error rate (FWER) and power are assessed via Monte-Carlo simulations and compared to the ones of approaches that rely on pooling of historical and current observations. It turned out that under optimal conditions, Bayesian approaches based on robustified prior distributions enable a substantial reduction of the control groups sample size, while still controlling the FWER up to a satisfactory level. Furthermore, at least to some extend, these approaches were able to protect against possible drift. This hightlights the potential of Bayesian study designs to reduce animal use in toxicology through re-use of the large pool of existing control data.

stat.ME

Identifying the potential of sample overlap in evidence synthesis of observational studies

Sample overlap is a common issue in evidence synthesis in the field of medical research, particularly when integrating findings from observational studies utilizing existing databases such as registries. Due to the general inaccessibility of unique identifiers for each observation, addressing sample overlap has been a complex problem, potentially biasing evidence synthesis outcomes and undermining their credibility. We developed a method to construct indicators for the degree of sample overlap in evidence synthesis of studies based on existing data. Our method is rooted in set theory and is based on the coding of the ranges of several well selected sample characteristics, offers a practical solution by focusing on making inference based on sample characteristics rather than on individual participant data. Useful information, such as the overlap-free sample set with the largest sample size in an evidence synthesis, can be derived from this method. We applied our model to several real-world evidence syntheses, demonstrating its effectiveness and flexibility. Our findings highlight the growing importance of addressing sample overlap in evidence synthesis, especially with the increasing relevance of secondary use of data, an area currently under-explored in research.

stat.ME

Blinded sample size re-estimation accounting for uncertainty in mid-trial estimation

For randomized controlled trials to be conclusive, it is important to set the target sample size accurately at the design stage. Comparing two normal populations, the sample size calculation requires specification of the variance other than the treatment effect and misspecification can lead to underpowered studies. Blinded sample size re-estimation is an approach to minimize the risk of inconclusive studies. Existing methods proposed to use the total (one-sample) variance that is estimable from blinded data without knowledge of the treatment allocation. We demonstrate that, since the expectation of this estimator is greater than or equal to the true variance, the one-sample variance approach can be regarded as providing an upper bound of the variance in blind reviews. This worst-case evaluation can likely reduce a risk of underpowered studies. However, blinded reviews of small sample size may still lead to underpowered studies. We propose a refined method accounting for estimation error in blind reviews using an upper confidence limit of the variance. A similar idea had been proposed in the setting of external pilot studies. Furthermore, we developed a method to select an appropriate confidence level so that the re-estimated sample size attains the target power. Numerical studies showed that our method works well and outperforms existing methods. The proposed procedure is motivated and illustrated by recent randomized clinical trials.

stat.ME

Consistent Bayesian meta-analysis on subgroup specific effects and interactions

Commonly, clinical trials report effects not only for the full study population but also for patient subgroups. Meta-analyses of subgroup-specific effects and treatment-by-subgroup interactions may be inconsistent, especially when trials apply different subgroup weightings. We show that meta-regression can, in principle, with a contribution adjustment, recover the same interaction inference regardless of whether interaction data or subgroup data are used. Our Bayesian framework for subgroup-data interaction meta-analysis inherently (i) adjusts for varying relative subgroup contribution, quantified by the information fraction (IF) within a trial; (ii) is robust to prevalence imbalance and variation; (iii) provides a self-contained, model-based approach; and (iv) can be used to incorporate prior information into interaction meta-analyses with few studies.The method is demonstrated using an example with as few as seven trials of disease-modifying therapies in relapsing-remitting multiple sclerosis. The Bayesian Contribution-adjusted Meta-analysis by Subgroup (CAMS) indicates a stronger treatment-by-disability interaction (relapse rate reduction) in patients with lower disability (EDSS <= 3.5) compared with the unadjusted model, while results for younger patients (age < 40 years) are unchanged.By controlling subgroup contribution while retaining subgroup interpretability, this approach enables reliable interaction decision-making when published subgroup data are available.Although the proposed CAMS approach is presented in a Bayesian context, it can also be implemented in frequentist or likelihood frameworks.

stat.ME

Utilizing subgroup information in random-effects meta-analysis of few studies

Random-effects meta-analyses are widely used for evidence synthesis in medical research. However, conventional methods based on large-sample approximations often exhibit poor performance in case of very few studies (e.g., 2 to 4), which is very common in practice. Existing methods aiming to improve small-sample performance either still suffer from poor estimates of heterogeneity or result in very wide confidence intervals. Motivated by meta-analyses evaluating surrogate outcomes, where units nested within a trial are often exploited when the number of trials is small, we propose an inference approach based on a common-effect estimator synthesizing data from the subgroup-level instead of the study-level. Two DerSimonian-Laird type heterogeneity estimators are derived using the subgroup-level data, and are incorporated into the Henmi-Copas type variance to adequately reflect variance components. We considered t-quantile based intervals to account for small-sample properties and used flexible degrees of freedom to reduce interval lengths. A comprehensive simulation is conducted to study the performance of our methods depending on various magnitudes of subgroup effects as well as subgroup prevalences. Some general recommendations are provided on how to select the subgroups, and methods are illustrated using two example applications.

stat.ME

Subgroup comparisons within and across studies in meta-analysis

Subgroup-specific meta-analysis synthesizes treatment effects for patient subgroups across randomized trials. Methods include joint or separate modeling of subgroup effects and treatment-by-subgroup interactions, but inconsistencies arise when subgroup prevalence differs between studies (e.g., proportion of non-smokers). A key distinction is between study-generated evidence within trials and synthesis-generated evidence obtained by contrasting results across trials. This distinction matters for identifying which subgroups benefit or are harmed most. Failing to separate these evidence types can bias estimates and obscure true subgroup-specific effects, leading to misleading conclusions about relative efficacy. Standard approaches often suffer from such inconsistencies, motivating alternatives. We investigate standard and novel estimators of subgroup and interaction effects in random-effects meta-analysis and study their properties. We show that using the same weights across different analyses (SWADA) resolves inconsistencies from unbalanced subgroup distributions and improves subgroup and interaction estimates. Analytical and simulation studies demonstrate that SWADA reduces bias and improves coverage, especially under pronounced imbalance. To illustrate, we revisit recent meta-analyses of randomized trials of COVID-19 therapies. Beyond COVID-19, the findings outline a general strategy for correcting compositional bias in evidence synthesis, with implications for decision-making and statistical modeling. We recommend the Interaction RE-weights SWADA as a practical default when aggregation bias is plausible: it ensures collapsibility, maintains nominal coverage with modest width penalty, and yields BLUE properties for the interaction.

stat.ME

A note on blinded continuous monitoring for continuous outcomes

Continuous monitoring is becoming more popular due to its significant benefits, including reducing sample sizes and reaching earlier conclusions. In general, it involves monitoring nuisance parameters (e.g., the variance of outcomes) until a specific condition is satisfied. The blinded method, which does not require revealing group assignments, was recommended because it maintains the integrity of the experiment and mitigates potential bias. Although Friede and Miller (2012) investigated the characteristics of blinded continuous monitoring through simulation studies, its theoretical properties are not fully explored. In this paper, we aim to fill this gap by presenting the asymptotic and finite-sample properties of the blinded continuous monitoring for continuous outcomes. Furthermore, we examine the impact of using blinded versus unblinded variance estimators in the context of continuous monitoring. Simulation results are also provided to evaluate finite-sample performance and to support the theoretical findings.

math.ST

A studentized permutation test for the treatment effect in individual participant data meta-analysis

Meta-analysis is a well-established tool used to combine data from several independent studies, each of which usually compares the effect of an experimental treatment with a control group. While meta-analyses are often performed using aggregated study summaries, they may also be conducted using individual participant data (IPD). Classical meta-analysis models may be generalized to handle continuous IPD by formulating them within a linear mixed model framework. IPD meta-analyses are commonly based on a small number of studies. Technically, inference for the overall treatment effect can be performed using Student-t approximation. However, as some approaches may not adequately control the type I error, Satterthwaite's or Kenward-Roger's method have been suggested to set the degrees-of-freedom parameter. The latter also adjusts the standard error of the treatment effect estimator. Nevertheless, these methods may be conservative. Since permutation tests are known to control the type I error and offer robustness to violations of distributional assumptions, we propose a studentized permutation test for the treatment effect based on permutations of standardized residuals across studies in IPD meta-analysis. Also, we construct confidence intervals for the treatment effect based on this test. The first interval is derived from the percentiles of the permutation distribution. The second interval is obtained by searching values closest to the effect estimate that are just significantly different from the true effect. In a simulation study, we demonstrate satisfactory performance of the proposed methods, often producing shorter confidence intervals compared with competitors.

stat.ME

Meta-analytic-predictive priors based on a single study

Meta-analytic-predictive (MAP) priors have been proposed as a generic approach to deriving informative prior distributions, where external empirical data are processed to learn about certain parameter distributions. The use of MAP priors is also closely related to shrinkage estimation (also sometimes referred to as dynamic borrowing). A potentially odd situation arises when the external data consist only of a single study. Conceptually this is not a problem, it only implies that certain prior assumptions gain in importance and need to be specified with particular care. We outline this important, not uncommon special case and demonstrate its implementation and interpretation based on the normal-normal hierarchical model. The approach is illustrated using example applications in clinical medicine.

stat.ME

Bayesian random-effects meta-analysis of aggregate data on clinical events

To investigate intervention effects on rare events, meta-analysis techniques are commonly applied in order to assess the accumulated evidence. When it comes to adverse effects in clinical trials, these are often most adequately handled using survival methods. A common-effect model that is able to process data in commonly quoted formats in terms of hazard ratios has been proposed for this purpose. In order to accommodate potential heterogeneity between studies, we have extended the model by Holzhauer to a random-effects approach. The Bayesian model is described in detail, and applications to realistic data sets are discussed along with sensitivity analyses and Monte Carlo simulations to support the conclusions.

stat.ME

In silico clinical trials in drug development: a systematic review

In the context of clinical research, computational models have received increasing attention over the past decades. In this systematic review, we aimed to provide an overview of the role of so-called in silico clinical trials (ISCTs) in medical applications. Exemplary for the broad field of clinical medicine, we focused on in silico (IS) methods applied in drug development, sometimes also referred to as model informed drug development (MIDD). We searched PubMed and ClinicalTrials.gov for published articles and registered clinical trials related to ISCTs. We identified 202 articles and 48 trials, and of these, 76 articles and 19 trials were directly linked to drug development. We extracted information from all 202 articles and 48 clinical trials and conducted a more detailed review of the methods used in the 76 articles that are connected to drug development. Regarding application, most articles and trials focused on cancer and imaging-related research while rare and pediatric diseases were only addressed in 14 articles and 5 trials, respectively. While some models were informed combining mechanistic knowledge with clinical or preclinical (in-vivo or in-vitro) data, the majority of models were fully data-driven, illustrating that clinical data is a crucial part in the process of generating synthetic data in ISCTs. Regarding reproducibility, a more detailed analysis revealed that only 24% (18 out of 76) of the articles provided an open-source implementation of the applied models, and in only 20% of the articles the generated synthetic data were publicly available. Despite the widely raised interest, we also found that it is still uncommon for ISCTs to be part of a registered clinical trial and their application is restricted to specific diseases leaving potential benefits of ISCTs not fully exploited.

q-bio.QM

Comparing restricted mean survival times in small sample clinical trials using pseudo-observations

The widely used proportional hazard assumption cannot be assessed reliably in small-scale clinical trials and might often in fact be unjustified, e.g. due to delayed treatment effects. An alternative to the hazard ratio as effect measure is the difference in restricted mean survival time (RMST) that does not rely on model assumptions. Although an asymptotic test for two-sample comparisons of the RMST exists, it has been shown to suffer from an inflated type I error rate in samples of small or moderate sizes. Recently, permutation tests, including the studentized permutation test, have been introduced to address this issue. In this paper, we propose two methods based on pseudo-observations (PO) regression models as alternatives for such scenarios and assess their properties in comparison to previously proposed approaches in an extensive simulation study. Furthermore, we apply the proposed PO methods to data from a clinical trail and, by doing so, point out some extension that might be very useful for practical applications such as covariate adjustments.

stat.ME

Multi-Modal Dataset Creation for Federated Learning with DICOM Structured Reports

Purpose: Federated training is often hindered by heterogeneous datasets due to divergent data storage options, inconsistent naming schemes, varied annotation procedures, and disparities in label quality. This is particularly evident in the emerging multi-modal learning paradigms, where dataset harmonization including a uniform data representation and filtering options are of paramount importance. Methods: DICOM structured reports enable the standardized linkage of arbitrary information beyond the imaging domain and can be used within Python deep learning pipelines with highdicom. Building on this, we developed an open platform for data integration and interactive filtering capabilities that simplifies the process of assembling multi-modal datasets. Results: In this study, we extend our prior work by showing its applicability to more and divergent data types, as well as streamlining datasets for federated training within an established consortium of eight university hospitals in Germany. We prove its concurrent filtering ability by creating harmonized multi-modal datasets across all locations for predicting the outcome after minimally invasive heart valve replacement. The data includes DICOM data (i.e. computed tomography images, electrocardiography scans) as well as annotations (i.e. calcification segmentations, pointsets and pacemaker dependency), and metadata (i.e. prosthesis and diagnoses). Conclusion: Structured reports bridge the traditional gap between imaging systems and information systems. Utilizing the inherent DICOM reference system arbitrary data types can be queried concurrently to create meaningful cohorts for clinical studies. The graphical interface as well as example structured report templates will be made publicly available.

cs.IR

Real World Federated Learning with a Knowledge Distilled Transformer for Cardiac CT Imaging

Federated learning is a renowned technique for utilizing decentralized data while preserving privacy. However, real-world applications often face challenges like partially labeled datasets, where only a few locations have certain expert annotations, leaving large portions of unlabeled data unused. Leveraging these could enhance transformer architectures ability in regimes with small and diversely annotated sets. We conduct the largest federated cardiac CT analysis to date (n=8,104) in a real-world setting across eight hospitals. Our two-step semi-supervised strategy distills knowledge from task-specific CNNs into a transformer. First, CNNs predict on unlabeled data per label type and then the transformer learns from these predictions with label-specific heads. This improves predictive accuracy and enables simultaneous learning of all partial labels across the federation, and outperforms UNet-based models in generalizability on downstream tasks. Code and model weights are made openly available for leveraging future cardiac CT analysis.

eess.IV

A Review of EMA Public Assessment Reports where Non-Proportional Hazards were Identified

While well-established methods for time-to-event data are available when the proportional hazards assumption holds, there is no consensus on the best approach under non-proportional hazards. A wide range of parametric and non-parametric methods for testing and estimation in this scenario have been proposed. In this review we identified EMA marketing authorization procedures where non-proportional hazards were raised as a potential issue in the risk-benefit assessment and extract relevant information on trial design and results reported in the corresponding European Assessment Reports (EPARs) available in the database at paediatricdata.eu. We identified 16 Marketing authorization procedures, reporting results on a total of 18 trials. Most procedures covered the authorization of treatments from the oncology domain. For the majority of trials NPH issues were related to a suspected delayed treatment effect, or different treatment effects in known subgroups. Issues related to censoring, or treatment switching were also identified. For most of the trials the primary analysis was performed using conventional methods assuming proportional hazards, even if NPH was anticipated. Differential treatment effects were addressed using stratification and delayed treatment effect considered for sample size planning. Even though, not considered in the primary analysis, some procedures reported extensive sensitivity analyses and model diagnostics evaluating the proportional hazards assumption. For a few procedures methods addressing NPH (e.g.~weighted log-rank tests) were used in the primary analysis. We extracted estimates of the median survival, hazard ratios, and time of survival curve separation. In addition, we digitized the KM curves to reconstruct close to individual patient level data. Extracted outcomes served as the basis for a simulation study of methods for time to event analysis under NPH.

stat.AP

A studentized permutation test in group sequential designs

In group sequential designs, where several data looks are conducted for early stopping, we generally assume the vector of test statistics from the sequential analyses follows (at least approximately or asymptotially) a multivariate normal distribution. However, it is well-known that test statistics for which an asymptotic distribution is derived may suffer from poor small sample approximation. This might become even worse with an increasing number of data looks. The aim of this paper is to improve the small sample behaviour of group sequential designs while maintaining the same asymptotic properties as classical group sequential designs. This improvement is achieved through the application of a modified permutation test. In particular, this paper shows that the permutation distribution approximates the distribution of the test statistics not only under the null hypothesis but also under the alternative hypothesis, resulting in an asymptotically valid permutation test. An extensive simulation study shows that the proposed permutation test better controls the Type I error rate than its competitors in the case of small sample sizes.

math.ST