Search arXiv⌕ Search

arXiv subjects

Xiaoxian Duan

Publications and source records attributed to Xiaoxian Duan.

3 recordsLinked to original sources

What Will Post-Training Fix? Per-Problem Gains Are Shared Across Independent RL Runs, and Existing Checkpoints Predict Them Better Than A Priori Signals

Data selection, curricula and the evaluation of post-training recipes all assume that we can tell, before training, which problems a model will improve on. We test this assumption directly. For two base models, DeepSeek-R1-Distill-Qwen-1.5B and Qwen2.5-Math-1.5B, we evaluate eighteen post-training runs on up to 1532 competition math problems with many samples per problem, and compare signals available before training against a noise ceiling derived from the agreement between disjoint subsets of runs. Three findings hold on both base models. What post-training fixes is shared: independent runs agree on which rarely solved problems improve, with a noise ceiling of about 0.9, yet the two base models agree with each other only at rho=0.25 - the shared component belongs to the base model, not to the problem. A priori signals capture a minority of it: base pass rate, the likelihood of a correct solution, a larger model's pass rate and their combinations explain only 0.30 and 0.24 of the explainable variance in gains. Existing checkpoints are the better predictor: the per-problem gains of a single checkpoint from another family predict a new run better than every a priori signal, alone or combined (0.52 vs. 0.33 and 0.34 vs. 0.17 against the combined signals). The conclusions hold on problems from 2025-2026 competitions and when the baselines on the two sides of every comparison are estimated independently. We propose the noise ceiling as a standard companion to per-problem signals.

cs.LG↗

VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification

Geometry problem generation is useful for AI-assisted education and multimodal mathematical reasoning, but reliable synthesis remains difficult because the problem statement, diagram, constraints, and solution should be mutually consistent. Existing methods often trade off controllability and reliability: seed-based rewriting is flexible but weakly verifiable, whereas diagram-first construction improves validity but is less suited to arbitrary user-specified constraints. We introduce VeriGeo, a controllable geometry generation framework grounded in executable reasoning traces. Given user constraints such as target concepts and difficulty, an Author agent generates a problem and diagram, and a Solver agent produces a proof-aligned solution. Both agents use a shared action sequence that connects natural language, diagrams, geometric constraints, and proof steps into a verifiable representation. A three-stage pipeline checks numerical consistency, analytical realizability, and global consistency, using verification-guided reflection to repair recoverable failures and reject unrecoverable ones. Across five LLM backbones, raw generations frequently fail these checks, while VeriGeo repairs a substantial fraction of the invalid attempts. Supervised fine-tuning on 8.7k examples generated by VeriGeo achieves the best reported GeoQA performance among end-to-end multimodal LLM-based solvers, and obtains strong results on PGPS9K and MathVista-GPS, demonstrating the effectiveness of verified synthetic data for improving multimodal geometry reasoning.

cs.AI↗

Constraining the Lattice Fluid Dark Energy from SNe Ia, BAO and OHD

Sanchez and Lacombe have ever developed a lattice fluid theory based on a well-defined statistical mechanical model. Taking the lattice fluid as a candidate of dark energy, we investigate the cosmic evolution of this fluid. Using the combined observational data of Type Ia Supernova (SNe Ia), Baryon Acoustic Oscillations (BAO) and Observational Hubble Data (OHD), we find the best fit value of the parameter in the model, $A = -0.3_{-0.1}^{+0.1}$. Then the cosmological implications of the model are presented.

astro-ph.CO↗