Search arXiv⌕ Search

arXiv subjects

Duy Thuc Nguyen

Publications and source records attributed to Duy Thuc Nguyen.

2 recordsLinked to original sources

Quantitative Universality of Approximate Message Passing for Rank-One Quadratic Sensing

Approximate Message Passing (AMP) algorithms are attractive as they are computationally efficient and simultaneously admit a precise characterization in terms of the low-dimensional "state-evolution" recursion. In this work, we establish quantitative universality for AMP with centered rank-one sensing matrices $Z_i=(x_ix_i^\top-I_d)/\sqrt d$, where $x_i$ are independent standard Gaussian vectors. These matrices arise in quadratic regression and learning quadratic neural networks. Their normalized vectorizations have isotropic covariance in dimension $p=d(d+1)/2$, but strongly dependent coordinates. For linear observations with Gaussian noise and prescribed smooth, bounded spectral denoisers, we compare rank-one sensing with a covariance-matched isotropic Gaussian ensemble. At a fixed iteration horizon, with $n\leq Cp$ and uniform bounds on the normalized signal energy, initialization, coefficients, and first four denoiser derivatives, normalized signal overlaps, cross-time overlaps, and smooth linear spectral statistics differ by $O(d^{-1/2})$ in every fixed $L^m$. This yields universality of the normalized mean-squared error, fixed spectral moments, and empirical spectral distributions. Whenever the Gaussian observables admit a joint state-evolution limit, the same predictions hold under rank-one sensing. The proof uses columnwise Lindeberg replacement, leave-one-out analysis, smooth spectral expansions, and high-moment concentration. For prescribed smooth spectral denoisers, the result establishes the AMP universality conjectured by Maillard et al. (arXiv:2408.03733).

cs.IT↗

HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class

Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or involve formal proofs, often overlooking approximation-based problems ubiquitous in applied science and engineering. To fill this gap, we build on prior work and present HARDMath2, a dataset of 211 original problems covering the core topics in an introductory graduate applied math class, including boundary-layer analysis, WKB methods, asymptotic solutions of nonlinear partial differential equations, and the asymptotics of oscillatory integrals. This dataset was designed and verified by the students and instructors of a core graduate applied mathematics course at Harvard. We build the dataset through a novel collaborative environment that challenges students to write and refine difficult problems consistent with the class syllabus, peer-validate solutions, test different models, and automatically check LLM-generated solutions against their own answers and numerical ground truths. Evaluation results show that leading frontier models still struggle with many of the problems in the dataset, highlighting a gap in the mathematical reasoning skills of current LLMs. Importantly, students identified strategies to create increasingly difficult problems by interacting with the models and exploiting common failure modes. This back-and-forth with the models not only resulted in a richer and more challenging benchmark but also led to qualitative improvements in the students' understanding of the course material, which is increasingly important as we enter an age where state-of-the-art language models can solve many challenging problems across a wide domain of fields.

cs.LG↗