Search arXiv⌕ Search

arXiv · 2609.30720

Optimal Personalized Subspace Learning for Multi-view Tensor Observations

Abstract

In this work, we model the observed multi-view tensors by decomposing the underlying signal in each view into two components: (i) the shared component that captures common dynamics across all views, and (ii) the private component that accounts for view-wise unique variations. To decouple the shared and private components, we introduce a novel Tucker personalized subspace principal component analysis (TPS-PCA) approach for tensors, which admits a one-step closed-form solution and serves as an ideal surrogate for our extended tensor-version personalized PCA (TP-PCA), adapted from the seminal work by \cite{shi2024personalized}. The theoretical analysis reveals that the proposed TPS-PCA estimators reach the minimax lower bound in terms of view-wise tensor decoupling, whereas the TP-PCA estimators only achieve a rate of average decoupling error across views, which is still slower than that of the TPS-PCA estimators. Extensive numerical experiments are conducted on synthetic and real datasets, demonstrating the wide applicability of the proposed method in fields including power management, financial analysis, and activity recognition.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kangxiang Qin, Zeyu Li, Xinbing Kong, Wang Zhou. 2026-09-25. Optimal Personalized Subspace Learning for Multi-view Tensor Observations. https://arxiv.org/abs/2609.30720

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PK/PD-integrated Bayesian platform design for phase II dose regimen optimization

Early-phase dose-finding methods increasingly assess toxicity and efficacy jointly, but comparisons based only on administered dose may inadequately characterize regimens differing in schedule. We developed a Bayesian phase II adaptive platform design for regimen optimization that integrates pharmacokinetic/pharmacodynamic (PK/PD) modelling into toxicity, efficacy, regimen selection and adaptation decisions. The proposed PK/PD-informed Regimen Optimization Platform (PROP) design uses a population PK/PD model to generate patient- and population-level predictions of exposure and biological activity. Acute and cumulative toxicities are analysed using a discrete-time time-to-event model informed by PK exposure. Efficacy is evaluated through Bayesian model averaging of exposure-driven and biomarker-driven time-to-event models. The design supports regimen graduation, discontinuation for futility or safety, and addition of unexplored regimens. Performance was evaluated through simulations motivated by an influenza intensive-care setting. Across six scenarios, PROP generally improved graduation and futility decisions, reduced inappropriate graduation, and supported the addition of promising regimens compared with dose-based alternatives. It also more accurately estimated regimen-specific toxicity and arm-specific efficacy, while the model-averaging framework favored the efficacy model consistent with the data-generating mechanism. Dose-based approaches performed better for safety stopping in some scenarios, despite less accurate characterization of the regimen--toxicity relationship. PK/PD-informed platform designs can improve adaptive regimen selection and knowledge generation when dose alone cannot adequately characterize treatment regimens.

stat.ME↗

Towards more plausible point-identifying assumptions in two-sample Mendelian randomization

Two-sample Mendelian randomization (MR) is a widely applied methodology in epidemiology. In two-sample MR, summary data (typically, regression coefficients and standard errors) quantifying the association between multiple genetic variants and the exposure and the outcome are used in an instrumental variable framework aimed at estimating the causal effect of the exposure on the outcome. Most two-sample MR methods were developed under data-generating models where the association of for each candidate genetic instrument with the exposure, as well as the causal effect of the exposure on the outcome, are constant in the additive scale. These assumptions are useful because they imply that, had all genetic variants been valid IVs, they would all estimate the same causal parameter - namely, the constant causal effect. We refer to this condition as summary-level homogeneity. However, these are rather strong homogeneity conditions which may raise concerns about the plausibility of these methods in practice. In this paper, we show that summary-level homogeneity is implied by the following conditions: the causal effect is additive linear, but not necessarily constant across, all strata of the population; and uncorrelatedness between heterogeneity in the causal effect and in the association between each genetic variant and the exposure. Under these conditions, typical two-sample MR methods can be interpreted as estimators of the average causal effect. These results clarify that point-identifying assumptions required for two-sample MR methods are weaker than previously anticipated, which contributes to their plausibility and interpretation in at least some practical applications.

stat.ME↗

Amortized Bayesian Disease Mapping and Boundary Detection on Heterogeneous Spatial Graphs

Spatial disease maps help public-health researchers identify geographic inequalities, but standard Bayesian smoothing can obscure localized disparities when neighboring communities have sharply different socioeconomic or behavioral profiles. Analysts therefore need to determine where smoothing should be interrupted and repeat that analysis as maps, adjacency structures, and outcomes change. We develop a covariate-informed Bayesian boundary model and an amortized posterior approximation trained across heterogeneous areal graphs. The model distinguishes local interruptions in smoothing from broader residual spatial dependence; the trained approximation handles maps with different numbers of regions. Simulations examine posterior calibration, boundary-probability recovery, replicated-data behavior, and MCSE-controlled agreement with prior-matched MCMC. In contrast to traditional approaches that analyze these data separately, we demonstrate the effectiveness of using a single trained deep learning network to analyze respiratory hospitalizations in Greater Glasgow; lung cancer incidence in California; and tracheal, bronchial, and lung cancer mortality in South Korea, comprising 58 to 241 regions. Selected boundary density is greatest in Glasgow and lowest in South Korea despite substantial residual spatial dependence in both, showing that local interruption and broader spatial persistence need not vary together. Across all three applications, edge-level boundary probabilities agree substantially with dataset-specific analyses, although posterior spread and thresholded boundary sets differ. These results support reusable Bayesian boundary analysis across the evaluated disease-map class and identify the validation needed before deployment to new applications.

stat.ME↗