Optimization Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise
The theoretical understanding of differentially private stochastic gradient descent (DP-SGD) with temporally correlated noise remains limited, particularly for non-convex neural network training. As a first step, we study two-layer Kolmogorov-Arnold Networks (KANs), a recently introduced architecture with learnable spline-based edge functions. We establish the first optimization risk bounds for clipped mini-batch DP-SGD with correlated noise in this setting, with explicit dependence on temporal correlation, clipping, mini-batch sampling, and network width. Existing arguments fail for three reasons: temporal dependence breaks the conditional-centering step; projection obstructs the cross-iteration cancellation of correlated perturbations; and active clipping breaks the empirical-gradient structure. We address these issues through shifted and auxiliary dynamics, a weighted empirical loss, and a high-probability localization argument. Our bound shows that temporal correlation reduces the leading noise terms, the clipping threshold enters essentially through an effective step size, and private training admits an explicit network width range. Experiments on synthetic data and MNIST support the predicted optimization effect of temporal correlation. As an application, we derive population risk guarantees via algorithmic stability. Our framework recovers non-private mini-batch SGD, independent-noise DP-SGD, and their full-batch counterparts as special cases.