arXiv · 2609.32325
Active Feature Acquisition With Incomplete Training Data
Abstract
In many prediction tasks, acquiring all features can be a prohibitively expensive or outright impossible task. Further, in many cases a static subset of features may not be enough to solve the problem sufficiently across various instances. Active Feature Acquisition (AFA) addresses these problems by formalizing the trade-off between feature cost and predictive performance during sequential feature selection. However, prior AFA work largely assumes access to complete training data, an assumption that is often violated in practice. Here we study AFA with Incomplete Training Data (AFA-ITD), showing that under missing completely at random (MCAR) data, one-step acquisition values remain unchanged, whereas multi-step values can decrease. We analyze three approaches to learning from incomplete data: aliasing, filtering, and generative restoration. We show that filtering can require a number of training instances scaling exponentially with the dimension, whereas generative restoration scales exponentially with the acquisition budget. We empirically test our theory on a controlled experiment and across common AFA datasets and find that missingness mainly damages methods that exploit multi-step acquisitions and that generative restoration is able to recover lost performance in many experiments. Code is available at https://github.com/Linusaronsson/AFA-Benchmark/tree/missing-data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Reza Rezvan, Valter Schütz, Han Wu, Linus Aronsson, Morteza Haghir Chehreghani. 2026-09-26. Active Feature Acquisition With Incomplete Training Data. https://arxiv.org/abs/2609.32325
Cite the original work for its findings. Save a collection to share your selection of sources.