Data-Dependent Complexity of First-Order Methods for Support Vector Machines
Large-scale problems in data science are often modeled with optimization, and the optimization model is usually solved with first-order methods that may converge at a sublinear rate. Therefore, it is of interest to terminate the optimization algorithm as soon as the underlying data science task is accomplished. We analyze the complexity of first-order methods for SVM in terms of the accuracy required to obtain a useful classifier. We derive a simple stopping condition for a perturbed SVM formulation that certifies points as correctly classified by both the current classifier and the optimal classifier, without requiring assumptions on the data. Under a geometric model consisting of two clusters of well-classified points together with noisy observations, we show that certifying a single point guarantees that the current classifier separates all well-classified points provided the number of training samples exceeds a modest threshold. We additionally derive a data-dependent accuracy threshold such that an approximate dual solution within this threshold yields a hyperplane separating the well-classified points. We specialize our theory to FISTA and coordinate descent by deriving computable bounds on the distance to the optimal dual solution, allowing the application of our stopping condition. For FISTA, we obtain data-dependent iteration complexity bounds for satisfying our stopping condition and for separating the well-classified points. Our analysis reveals limitations of the stopping condition used by the well-known software package LIBLINEAR, and we give examples where it is either overly conservative or terminates prematurely with a poor classifier. Numerical experiments show that a FISTA implementation using our stopping condition is competitive with widely used SVM solvers on a range of benchmarks, and demonstrate that our stopping condition provides a robust basis for termination.