Search arXivSearch

arXiv subjects

Martin Posch

Publications and source records attributed to Martin Posch.

At least 19 recordsLinked to original sources

Clustering-Based Outcome Models for Clinical Studies: A Scoping Review

This review provides a systematic overview of methods that combine covariate-based clustering of observational units (patients) with outcome models for clinical studies. We distinguish between informed-cluster models, where the outcome contributes to cluster formation, and agnostic-cluster models, where clustering is performed solely on covariates in a separate first step. Informed-cluster models include product partition models with covariates (PPMx), finite mixtures of regression models (FMR), and cluster-aware supervised learning (CluSL). Agnostic-cluster models encompass two-step procedures using either model-based or algorithmic clustering followed by cluster-specific regression models. Following a systematic search of Web of Science and PubMed, 55 records were identified that propose or evaluate such models. We describe the key models, summarise study characteristics, and present applications from biomedical and public health research. Clustering-based outcome models are particularly relevant for settings with high-dimensional covariates (e.g., biomarker panels and "omics") and heterogeneous patient populations. These models can support risk stratification and we discuss extensions to estimate subgroup-specific treatment effects. They are most valuable when the population is clustered in distinct regions of the covariate space that correspond to different outcome distributions. We discuss applications to rare disease research, covariate adjustment and borrowing from historical data, and subgroup-specific treatment effect estimation in clinical trials.

stat.ME

An Adaptive Phase II Trial Design for Dose Selection and Addition in Microfilarial Infections

We propose a frequentist adaptive phase 2 trial design to evaluate the safety and efficacy of three treatment regimens (doses) compared to placebo for four types of helminth (worm) infections. This trial will be carried out in four Subsaharan African countries from spring 2025. Since the safety of the highest dose is not yet established, the study begins with the two lower doses and placebo. Based on safety and early efficacy results from an interim analysis, a decision will be made to either continue with the two lower doses or drop one or both and introduce the highest dose instead. This design borrows information across baskets for safety assessment, while efficacy is assessed separately for each basket. The proposed adaptive design addresses several key challenges: (1) The trial must begin with only the two lower doses because reassuring safety data from these doses is required before escalating to a higher dose. (2) Due to the expected speed of recruitment, adaptation decisions must rely on an earlier, surrogate endpoint. (3) The primary outcome is a count variable that follows a mixture distribution with an atom at 0. To control the familywise error rate in the strong sense when comparing multiple doses to the control in the adaptive design, we extend the partial conditional error approach to accommodate the inclusion of new hypotheses after the interim analysis. In a comprehensive simulation study we evaluate various design options and analysis strategies, assessing the robustness of the design under different design assumptions and parameter values. We identify scenarios where the adaptive design improves the trial's ability to identify an optimal dose. Adaptive dose selection enables resource allocation to the most promising treatment arms, increasing the likelihood of selecting the optimal dose while reducing the required overall sample size and trial duration.

stat.ME

Selection Bias in Hybrid Randomized Controlled Trials using External Controls: A Simulation Study

Hybrid randomized controlled trials (hybrid RCTs) integrate external control data, such as historical or concurrent data, with data from randomized trials. While numerous frequentist and Bayesian methods, such as the test-then-pool and Meta-Analytic-Predictive prior, have been developed to account for potential disagreement between the external control and randomized data, they cannot ensure strict type I error rate control. However, these methods can reduce biases stemming from systematic differences between external controls and trial data. A critical yet underexplored issue in hybrid RCTs is the prespecification of external data to be used in analysis. The validity of statistical conclusions in hybrid RCTs depends on the assumption that external control selection is independent of historical trials outcomes. In practice, historical data may be accessible during the planning stage, potentially influencing important decisions, such as which historical datasets to include or the sample size of the prospective part of the hybrid trial, thus introducing bias. Such data-driven design choices can be an additional source of bias, which can occur even when historical and prospective controls are exchangeable. Through a simulation study, we quantify the biases introduced by outcome-dependent selection of historical controls in hybrid RCTs using both Bayesian and frequentist approaches, and discuss potential strategies to mitigate this bias. Our scenarios consider variability and time trends in the historical studies, distributional shifts between historical and prospective control groups, sample sizes and allocation ratios, as well as the number of studies included. The impact of different rules for selecting external controls is demonstrated using a clinical trial example.

stat.ME

On the inclusion of non-concurrent controls in platform trials with an interim analysis

The analysis of platform trials can be enhanced by utilizing non-concurrent controls. Since including this data might also introduce bias in the treatment effect estimators if time trends are present, methods for incorporating non-concurrent controls adjusting for time have been proposed. However, so far their behavior has not been systematically investigated in platform trials that include interim analyses. To evaluate the impact of an interim analysis in trials utilizing non-concurrent controls, we consider a platform trial featuring two experimental arms and a shared control, with the second experimental arm entering later. We focus on a frequentist regression model that uses non-concurrent controls to estimate the treatment effect of the second arm and adjusts for time using a step function to account for temporal changes. We show that performing an interim analysis in Arm 1 may introduce bias in the point estimation of the effect in Arm 2, if the regression model is used without adjustment, and investigate how the marginal bias and bias conditional on the first arm continuing after the interim depend on different trial design parameters. Moreover, we propose a new estimator of the treatment effect in Arm 2, aiming to eliminate the bias introduced by both the interim analysis in Arm 1 and the time trends, and evaluate its performance in a simulation study. The newly proposed estimator is shown to substantially reduce the bias and type I error rate inflation while leading to power gains compared to an analysis using only concurrent controls.

stat.ME

Covariate adjustment for linear models in estimating treatment effects in randomised clinical trials. Some useful theory to guide simulation

Building on key papers that were published in special issues of Biometrics in 1957 and 1982 we propose and develop a three-aspect system for evaluating the effect of fitting covariates in the analysis of designed experiments, in particular randomised clinical trials. The three aspects are: first the effect on residual mean square error, second the effect on the variance inflation factor (VIF) and third the effect on second order precision. We concentrate, in particular, on the VIF and highlight not only an existing formula for its expected value based on assuming covariates have a Normal distribution but also develop a formula for its variance. We show how VIFs for categorical variable are related to the chi-square contingency table with rows as treatment and columns as categories. We illustrate the value of these formulae using a randomised clinical trial with five covariates, one of which is binary, and show that both mean and variance formulae predict results well for all $2^5=32$ possible models for each of three forms of simulation, random permutation, sampling from a Normal distribution and bootstrap resampling. Finally, we illustrate how the three-aspect system may be used to address various questions of interest when considering covariate adjustment.

stat.ME

In silico clinical trials in drug development: a systematic review

In the context of clinical research, computational models have received increasing attention over the past decades. In this systematic review, we aimed to provide an overview of the role of so-called in silico clinical trials (ISCTs) in medical applications. Exemplary for the broad field of clinical medicine, we focused on in silico (IS) methods applied in drug development, sometimes also referred to as model informed drug development (MIDD). We searched PubMed and ClinicalTrials.gov for published articles and registered clinical trials related to ISCTs. We identified 202 articles and 48 trials, and of these, 76 articles and 19 trials were directly linked to drug development. We extracted information from all 202 articles and 48 clinical trials and conducted a more detailed review of the methods used in the 76 articles that are connected to drug development. Regarding application, most articles and trials focused on cancer and imaging-related research while rare and pediatric diseases were only addressed in 14 articles and 5 trials, respectively. While some models were informed combining mechanistic knowledge with clinical or preclinical (in-vivo or in-vitro) data, the majority of models were fully data-driven, illustrating that clinical data is a crucial part in the process of generating synthetic data in ISCTs. Regarding reproducibility, a more detailed analysis revealed that only 24% (18 out of 76) of the articles provided an open-source implementation of the applied models, and in only 20% of the articles the generated synthetic data were publicly available. Despite the widely raised interest, we also found that it is still uncommon for ISCTs to be part of a registered clinical trial and their application is restricted to specific diseases leaving potential benefits of ISCTs not fully exploited.

q-bio.QM

Graph Based, Adaptive, Multi Arm, Multiple Endpoint, Two Stage Design

The graph based approach to multiple testing is an intuitive method that enables a study team to represent clearly, through a directed graph, its priorities for hierarchical testing of multiple hypotheses, and for propagating the available type-1 error from rejected or dropped hypotheses to hypotheses yet to be tested. Although originally developed for single stage non-adaptive designs, we show how it may be extended to two-stage designs that permit early identification of efficacious treatments, adaptive sample size re-estimation, dropping of hypotheses, and changes in the hierarchical testing strategy at the end of stage one. Two approaches are available for preserving the family wise error rate in the presence of these adaptive changes; the p-value combination method, and the conditional error rate method. In this investigation we will present the statistical methodology underlying each approach and will compare the operating characteristics of the two methods in a large simulation experiment.

stat.ME

Stratification in Randomised Clinical Trials for Rare Diseases and Analysis of Covariance: Some Simple Theory and Recommendations

A simple device for balancing for a continuous covariate in clinical trials is to stratify by whether the covariate is above or below some target value, typically the predicted median. This raises an issue as to which model should be used for modelling the effect of treatment on the outcome variable, $Y$. Should one fit, the stratum indicator, $S$, the continuous covariate, $X$, both or neither? When a covariate is added to a linear model there are three consequences for inference: 1) the mean square error effect, 2) the variance inflation factor and 3) second order precision. We consider that it is valuable to consider these three factors separately, even if, ultimately, it is their joint effect that matters. We present some simple theory, concentrating in particular on the variance inflation factor, that may be used to guide trialists in their choice of model. We also consider the case where the precise form of the relationship between the outcome and the covariate is not known. We conclude by recommending that the continuous covariate should always be in the model but that, depending on circumstances, there may be some justification in fitting the stratum indicator also.

stat.ME

Treatment-control comparisons in platform trials including non-concurrent controls

Shared controls in platform trials comprise concurrent and non-concurrent controls. For a given experimental arm, non-concurrent controls refer to data from patients allocated to the control arm before the arm enters the trial. The use of non-concurrent controls in the analysis is attractive because it may increase the trial's power of testing treatment differences while decreasing the sample size. However, since arms are added sequentially in the trial, randomization occurs at different times, which can introduce bias in the estimates due to time trends. In this article, we present methods to incorporate non-concurrent control data in treatment-control comparisons, allowing for time trends. We focus mainly on frequentist approaches that model the time trend and Bayesian strategies that limit the borrowing level depending on the heterogeneity between concurrent and non-concurrent controls. We examine the impact of time trends, overlap between experimental treatment arms and entry times of arms in the trial on the operating characteristics of treatment effect estimators for each method under different patterns for the time trends. We argue under which conditions the methods lead to type 1 error control and discuss the gain in power compared to trials only using concurrent controls by means of a simulation study in which methods are compared.

stat.ME

A Review of EMA Public Assessment Reports where Non-Proportional Hazards were Identified

While well-established methods for time-to-event data are available when the proportional hazards assumption holds, there is no consensus on the best approach under non-proportional hazards. A wide range of parametric and non-parametric methods for testing and estimation in this scenario have been proposed. In this review we identified EMA marketing authorization procedures where non-proportional hazards were raised as a potential issue in the risk-benefit assessment and extract relevant information on trial design and results reported in the corresponding European Assessment Reports (EPARs) available in the database at paediatricdata.eu. We identified 16 Marketing authorization procedures, reporting results on a total of 18 trials. Most procedures covered the authorization of treatments from the oncology domain. For the majority of trials NPH issues were related to a suspected delayed treatment effect, or different treatment effects in known subgroups. Issues related to censoring, or treatment switching were also identified. For most of the trials the primary analysis was performed using conventional methods assuming proportional hazards, even if NPH was anticipated. Differential treatment effects were addressed using stratification and delayed treatment effect considered for sample size planning. Even though, not considered in the primary analysis, some procedures reported extensive sensitivity analyses and model diagnostics evaluating the proportional hazards assumption. For a few procedures methods addressing NPH (e.g.~weighted log-rank tests) were used in the primary analysis. We extracted estimates of the median survival, hazard ratios, and time of survival curve separation. In addition, we digitized the KM curves to reconstruct close to individual patient level data. Extracted outcomes served as the basis for a simulation study of methods for time to event analysis under NPH.

stat.AP

Statistical modeling to adjust for time trends in adaptive platform trials utilizing non-concurrent controls

Utilizing non-concurrent control data (NCC) in the analysis of late-entering arms in platform trials has recently received considerable attention. While incorporating NCC can lead to increased power and lower sample sizes, it might introduce bias to the effect estimators if temporal drifts are present. Aiming to mitigate this potential bias, we propose various frequentist model-based approaches that leverage the NCC, while adjusting for time. One of the currently available models incorporates time as a categorical fixed effect, separating the trial duration into periods, defined as time intervals bounded by any arm entering or leaving the platform. In this work, we propose two extensions of this model. First, we consider an alternative definition of time by dividing the trial into fixed-length calendar time intervals. Second, we propose alternative model-based time adjustments. Specifically, we investigate adjusting for random effects and employing splines to model time with a polynomial function. We evaluate the performance of the proposed approaches in a simulation study and illustrate their use through a case study. We show that adjusting for time via a spline function controls the type I error in trials with a sufficiently smooth time trend pattern and may lead to power gains compared to the standard fixed effect model. However, the fixed effect model with period adjustment is the most robust model for arbitrary time trends, provided that the trend is equal across all arms. Especially, in trials with sudden changes in the time trend, the period-adjustment model is preferred if NCC are included.

stat.ME

A two-step approach for analyzing time to event data under non-proportional hazards

The log-rank test and the Cox proportional hazards model are commonly used to compare time-to-event data in clinical trials, as they are most powerful under proportional hazards. But there is a loss of power if this assumption is violated, which is the case for some new oncology drugs like immunotherapies. We consider a two-stage test procedure, in which the weighting of the log-rank test statistic depends on a pre-test of the proportional hazards assumption. I.e., depending on the pre-test either the log-rank or an alternative test is used to compare the survival probabilities. We show that if naively implemented this can lead to a substantial inflation of the type-I error rate. To address this, we embed the two-stage test in a permutation test framework to keep the nominal level alpha. We compare the operating characteristics of the two-stage test with the log-rank test and other tests by clinical trial simulations.

stat.ME

Efficiency of Multivariate Tests in Trials in Progressive Supranuclear Palsy

Measuring disease progression in clinical trials for testing novel treatments for multifaceted diseases as Progressive Supranuclear Palsy (PSP), remains challenging. In this study we assess a range of statistical approaches to compare outcomes measured by the items of the Progressive Supranuclear Palsy Rating Scale (PSPRS). We consider several statistical approaches, including sum scores, as an FDA-recommended version of the PSPRS, multivariate tests, and analysis approaches based on multiple comparisons of the individual items. We propose two novel approaches which measure disease status based on Item Response Theory models. We assess the performance of these tests in an extensive simulation study and illustrate their use with a re-analysis of the ABBV-8E12 clinical trial. Furthermore, we discuss the impact of the FDA-recommended scoring of item scores on the power of the statistical tests. We find that classical approaches as the PSPRS sum score demonstrate moderate to high power when treatment effects are consistent across the individual items. The tests based on Item Response Theory models yield the highest power when the simulated data are generated from an IRT model. The multiple testing based approaches have a higher power in settings where the treatment effect is limited to certain domains or items. The FDA-recommended item rescoring tends to decrease the simulated power. The study shows that there is no one-size-fits-all testing procedure for evaluating treatment effects using PSPRS items; the optimal method varies based on the specific effect size patterns. The efficiency of the PSPRS sum score, while generally robust and straightforward to apply, varies depending on the effect sizes' patterns encountered and more powerful alternatives are available in specific settings. These findings can have important implications for the design of future clinical trials in PSP.

stat.ME

A neutral comparison of statistical methods for time-to-event analyses under non-proportional hazards

While well-established methods for time-to-event data are available when the proportional hazards assumption holds, there is no consensus on the best inferential approach under non-proportional hazards (NPH). However, a wide range of parametric and non-parametric methods for testing and estimation in this scenario have been proposed. To provide recommendations on the statistical analysis of clinical trials where non proportional hazards are expected, we conducted a comprehensive simulation study under different scenarios of non-proportional hazards, including delayed onset of treatment effect, crossing hazard curves, subgroups with different treatment effect and changing hazards after disease progression. We assessed type I error rate control, power and confidence interval coverage, where applicable, for a wide range of methods including weighted log-rank tests, the MaxCombo test, summary measures such as the restricted mean survival time (RMST), average hazard ratios, and milestone survival probabilities as well as accelerated failure time regression models. We found a trade-off between interpretability and power when choosing an analysis strategy under NPH scenarios. While analysis methods based on weighted logrank tests typically were favorable in terms of power, they do not provide an easily interpretable treatment effect estimate. Also, depending on the weight function, they test a narrow null hypothesis of equal hazard functions and rejection of this null hypothesis may not allow for a direct conclusion of treatment benefit in terms of the survival function. In contrast, non-parametric procedures based on well interpretable measures as the RMST difference had lower power in most scenarios. Model based methods based on specific survival distributions had larger power, however often gave biased estimates and lower than nominal confidence interval coverage.

stat.ME

Simultaneous inference procedures for the comparison of multiple characteristics of two survival functions

Survival time is the primary endpoint of many randomized controlled trials, and a treatment effect is typically quantified by the hazard ratio under the assumption of proportional hazards. Awareness is increasing that in many settings this assumption is a-priori violated, e.g. due to delayed onset of drug effect. In these cases, interpretation of the hazard ratio estimate is ambiguous and statistical inference for alternative parameters to quantify a treatment effect is warranted. We consider differences or ratios of milestone survival probabilities or quantiles, differences in restricted mean survival times and an average hazard ratio to be of interest. Typically, more than one such parameter needs to be reported to assess possible treatment benefits, and in confirmatory trials the according inferential procedures need to be adjusted for multiplicity. By using the counting process representation of the mentioned parameters, we show that their estimates are asymptotically multivariate normal and we propose according parametric multiple testing procedures and simultaneous confidence intervals. Also, the logrank test may be included in the framework. Finite sample type I error rate and power are studied by simulation. The methods are illustrated with an example from oncology. A software implementation is provided in the R package nph.

stat.ME

Design Considerations for a Phase II platform trial in Major Depressive Disorder

Major Depressive Disorder (MDD) is one of the most common causes of disability worldwide. Unfortunately, about one-third of patients do not benefit sufficiently from available treatments and not many new drugs have been developed in this area in recent years. We thus need better and faster ways to evaluate many different treatment options quickly. Platform trials are a possible remedy - they facilitate the evaluation of more investigational treatments in a shorter period of time by sharing controls, as well as reducing clinical trial activation and recruitment times. We discuss design considerations for a platform trial in MDD, taking into account the unique disease characteristics, and present the results of extensive simulations to investigate the operating characteristics under various realistic scenarios. To allow the testing of more treatments, interim futility analyses should be performed to eliminate treatments that have either no or negligible treatment effect. Furthermore, we investigate different randomisation and allocation strategies as well as the impact of the per-treatment arm sample size. We compare the operating characteristics of such platform trials to those of traditional randomised controlled trials and highlight the potential advantages of platform trials.

stat.AP

Methods for non-proportional hazards in clinical trials: A systematic review

For the analysis of time-to-event data, frequently used methods such as the log-rank test or the Cox proportional hazards model are based on the proportional hazards assumption, which is often debatable. Although a wide range of parametric and non-parametric methods for non-proportional hazards (NPH) has been proposed, there is no consensus on the best approaches. To close this gap, we conducted a systematic literature search to identify statistical methods and software appropriate under NPH. Our literature search identified 907 abstracts, out of which we included 211 articles, mostly methodological ones. Review articles and applications were less frequently identified. The articles discuss effect measures, effect estimation and regression approaches, hypothesis tests, and sample size calculation approaches, which are often tailored to specific NPH situations. Using a unified notation, we provide an overview of methods available. Furthermore, we derive some guidance from the identified articles. We summarized the contents from the literature review in a concise way in the main text and provide more detailed explanations in the supplement.

stat.ME

Optimal allocation strategies in platform trials

Platform trials are randomized clinical trials that allow simultaneous comparison of multiple interventions, usually against a common control. Arms to test experimental interventions may enter and leave the platform over time. This implies that the number of experimental intervention arms in the trial may change over time. Determining optimal allocation rates to allocate patients to the treatment and control arms in platform trials is challenging because the change in treatment arms implies that also the optimal allocation rates will change when treatments enter or leave the platform. In addition, the optimal allocation depends on the analysis strategy used. In this paper, we derive optimal treatment allocation rates for platform trials with shared controls, assuming that a stratified estimation and testing procedure based on a regression model, is used to adjust for time trends. We consider both, analysis using concurrent controls only as well as analysis methods based on also non-concurrent controls and assume that the total sample size is fixed. The objective function to be minimized is the maximum of the variances of the effect estimators. We show that the optimal solution depends on the entry time of the arms in the trial and, in general, does not correspond to the square root of $k$ allocation rule used in the classical multi-arm trials. We illustrate the optimal allocation and evaluate the power and type 1 error rate compared to trials using one-to-one and square root of $k$ allocations by means of a case study.

stat.ME