Search arXiv⌕ Search

arXiv subjects

Nicola Aladrah

Publications and source records attributed to Nicola Aladrah.

2 recordsLinked to original sources

Understanding and inverse design of implicit bias in stochastic learning: a geometric perspective

Can we design a model such that its stochastic training favours a desired class of solutions without enforcing an explicit penalty? Under suitable conditions, the interplay between symmetries of a model's weight parametrization and stochastic training favours particular solutions, inducing an implicit bias. Building on this mechanism, we develop a framework for inverse-designing such biases by constructing novel parametrizations and their associated symmetries. We show how holomorphic functions make this construction and calculation simple and explicit. Specifically, we introduce a new parametrization that biases learned weights toward the binary values $\{-1,+1\}$. Numerical experiments confirm the theoretical predictions. They also show that our parametrization reproduces the same preference induced by an explicitly regularized model without adding a penalty to the training loss.

cs.LG↗

PAC--Bayes Bounds on Quotient Parameter Spaces: Geometry-induced Implicit-Bias Priors

Overparameterized models often have continuous parameter symmetries, so different parameters define the same predictor. We show that PAC--Bayesian analysis should be performed on the quotient predictor space: pushing a prior and posterior to the quotient preserves the empirical and population Gibbs risks while removing the nonnegative KL contribution caused solely by how the two distributions differ among parameterizations of the same predictor. Quotienting alone does not determine which prior to use. We construct a canonical choice of one parameterization for each predictor and account for the geometric volume of its equivalent parameterizations. This transforms a neutral reference prior into a data-independent prior that reflects the model's implicit bias. It approximates the ideal but inadmissible posterior-matched prior, which would minimize the KL term by depending on the training data. The resulting certificate is tighter exactly when this geometry-induced prior has smaller KL divergence from the learned quotient posterior than the neutral prior. We test this prediction in Fourier regression with a Hadamard parameterization and in Query-Key attention, using ordinary SGD without an explicit regularizer. The implicit-bias prior reduces the mean quotient-space KL by \(40.69\%\) and the mean PAC--Bayes certificate by \(21.40\%\) in the Fourier-Hadamard experiment. The smaller, prior-scale-dependent improvement in Query-Key attention confirms the predicted conditional nature of the effect.

cs.LG↗