Modeling positive and unlabeled data with a generalized additive density ratio model
This paper focuses on learning from positive and unlabeled (PU) data, where only some positives are labeled and the rest are mixed with negatives. Classical exponential tilting models guarantee identifiability by imposing a linear structure, but they can be severely misspecified when the true relationships are nonlinear. We propose a generalized additive density-ratio framework that retains identifiability while allowing nonlinear and feature-specific effects. The approach comes with a practical fitting algorithm and supporting theory that enable estimation and inference for the mixture proportion and other quantities of interest. In simulations and analyses of real datasets, the proposed method matches the standard exponential tilting method when the linear model is correct and delivers clear gains when it is not. Overall, the framework strikes a useful balance between flexibility and interpretability for PU data and provides principled tools for estimation, prediction, and uncertainty assessment.