Search arXivSearch

arXiv · 2408.01365

Data Debugging is NP-hard for Classifiers Trained with SGD

Abstract

Data debugging is to find a subset of the training data such that the model obtained by retraining on the subset has a better accuracy. A bunch of heuristic approaches are proposed, however, none of them are guaranteed to solve this problem effectively. This leaves an open issue whether there exists an efficient algorithm to find the subset such that the model obtained by retraining on it has a better accuracy. To answer this open question and provide theoretical basis for further study on developing better algorithms for data debugging, we investigate the computational complexity of the problem named Debuggable. Given a machine learning model $\mathcal{M}$ obtained by training on dataset $D$ and a test instance $(\mathbf{x}_\text{test},y_\text{test})$ where $\mathcal{M}(\mathbf{x}_\text{test})\neq y_\text{test}$, Debuggable is to determine whether there exists a subset $D^\prime$ of $D$ such that the model $\mathcal{M}^\prime$ obtained by retraining on $D^\prime$ satisfies $\mathcal{M}^\prime(\mathbf{x}_\text{test})=y_\text{test}$. To cover a wide range of commonly used models, we take SGD-trained linear classifier as the model and derive the following main results. (1) If the loss function and the dimension of the model are not fixed, Debuggable is NP-complete regardless of the training order in which all the training samples are processed during SGD. (2) For hinge-like loss functions, a comprehensive analysis on the computational complexity of Debuggable is provided; (3) If the loss function is a linear function, Debuggable can be solved in linear time, that is, data debugging can be solved easily in this case. These results not only highlight the limitations of current approaches but also offer new insights into data debugging.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zizheng Guo, Pengyu Chen, Yanzhang Fu, Dongjing Miao. 2024-08-02. Data Debugging is NP-hard for Classifiers Trained with SGD. https://arxiv.org/abs/2408.01365

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

CVP Is NP-Complete for Principal Cyclotomic Ideals

We prove that exact Euclidean decision-CVP is $\mathsf{NP}$-complete on the coefficient lattices of nonzero principal ideals in the power-of-two cyclotomic rings $R_d:=\mathbb{Z}[y]/(y^d+1)$. Our deterministic reduction from Exact Cover by 3-Sets (X3C) produces a target and a squared threshold $Δ$ such that the closest squared distance is exactly $Δ$ in YES instances and at least $Δ+4$ in NO instances. This also implies $\mathsf{NP}$-hardness of exact search-CVP under polynomial-time Turing reductions. We also transfer the resulting principal-ideal CVP instances to full-rank principal ideals of the cyclic quotient ring $\mathbb{Z}[X]/(X^D-1)$, where $D:=2d$. Their coefficient lattices are invariant under cyclic coordinate shifts. The lift preserves principality and multiplies corresponding squared distances by eight. Thus, on principal cyclic ideal lattices, exact decision-CVP is $\mathsf{NP}$-complete and exact search-CVP is $\mathsf{NP}$-hard. We also obtain uniformly computable fixed cyclotomic and cyclic families in which only the target and threshold depend on the X3C collection. Consequently, a polynomial-time solution to exact decision-CVPP on either family would imply $\mathsf{NP}\subseteq\mathsf{P}/\mathrm{poly}$ and collapse the polynomial hierarchy to $Σ_2^{\mathsf{P}}$. To our knowledge, the cyclic results answer Micciancio's questions of whether exact decision-CVP is $\mathsf{NP}$-hard on cyclic lattices and on a fixed family of cyclic lattices, even when restricted to full-rank principal cyclic ideals. Finally, under the coefficient embedding, we prove that exact decision-module-SIVP is $\mathsf{NP}$-complete on free rank-two modules over the same cyclotomic rings.

cs.CC

Fooling Thresholds of Halfspaces

We initiate the study of constructing explicit pseudorandom generators for thresholds of halfspaces with seed length polylogarithmic in the number of halfspaces. This class of functions lies at the frontier of circuit complexity [CTW26]. We show that the generator designed by O'Donnell, Servedio, and Tan for polytopes [OST22] also fools this broader class. To analyze the generator, we develop a threshold-specific smooth approximation framework based on a Bentkus-type mollifier. We prove derivative bounds for this mollifier and also establish a Boolean anticoncentration theorem for thresholds of halfspaces via a random thinning argument. These ingredients imply that the generator $δ$-fools every $k$-out-of-$m$ threshold of $m$ halfspaces over $\{-1,1\}^n$ with seed length $\widetilde{O}(κ^{6+2\varepsilon}\log^{6+2\varepsilon}\!m\cdotδ^{-(2+2\varepsilon)}\log n)$, for any arbitrarily small constant $\varepsilon>0$, where $κ=\min\{k,m-k+1\}$. The random thinning argument also yields bounds on the noise sensitivity and Gaussian surface area for thresholds of halfspaces, leading to learning algorithms under both the uniform and Gaussian distributions.

cs.CC

An Oracle Separating Conjectures about Incompleteness in the Finite Domain

Pudlák [Pud17] lists several major conjectures from the field of proof complexity and asks for oracles that separate corresponding relativized conjectures. Among these conjectures are: - $\mathsf{DisjNP}$: The class of all disjoint NP-pairs does not have many-one complete elements. - $\mathsf{SAT}$: NP does not contain many-one complete sets that have P-optimal proof systems. - $\mathsf{UP}$: UP does not have many-one complete problems. - $\mathsf{NP}\cap\mathsf{coNP}$: $\text{NP}\cap\text{coNP}$ does not have many-one complete problems. As one answer to this question, we construct an oracle relative to which $\mathsf{DisjNP}$, $\neg \mathsf{SAT}$, $\mathsf{UP}$, and $\mathsf{NP}\cap\mathsf{coNP}$ hold, i.e., there is no relativizable proof for the implication $\mathsf{DisjNP}\wedge \mathsf{UP}\wedge \mathsf{NP}\cap\mathsf{coNP}\Rightarrow\mathsf{SAT}$. In particular, regarding the conjectures by Pudlák this extends a result by Khaniki [Kha19].

cs.CC