Search arXiv⌕ Search

arXiv subjects

Fengyu Xie

Publications and source records attributed to Fengyu Xie.

11 recordsLinked to original sources

Chemical-space completeness through iterative crystal-structure generation and model adaptation

The emergence of deep learning has brought large-scale exploration of crystalline materials closer to practical realization. Yet, universal atomistic models face a practical trade-off between chemical generality and the accuracy, efficiency, and adaptability required for intensive exploration of specific materials systems. Within a bounded chemical system, this trade-off can be relaxed by exploiting its limited chemical complexity. Guided by this intuition, we propose a chemical-system-centric strategy that couples crystal-structure generative models with machine-learned force fields (MLFFs) in an iterative generation-evaluation-refinement loop. Using Li--P--S as a test case, we generate approximately 70,000 candidate structures, including more than 10,000 stable-unique-novel structures. The diversity of near-equilibrium local environments saturates within the first few iterations, accompanied by convergence of MLFF prediction errors, providing an operational measure of chemical-space completeness within bounded systems. The exploration also recovers chemically plausible P--S motifs that are absent from the pretraining databases but supported by earlier experiments. The resulting system-adapted models and structures further enable finite-$P$--$T$ phase-stability calculations, Li-ion transport screening, and electronic-structure prediction. These results suggest that chemical-system-centric exploration provides a practical route toward data-efficient and high-fidelity modeling within bounded chemical spaces.

cond-mat.mtrl-sci↗

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the scientific experience bottleneck. We introduce ScienceIDE, infrastructure for turning the world's scientific code into programmable environments for scientific agents. Guided by expert-defined scientific cases and acceptance criteria, agents transform repositories into executable environments that support task generation, execution, and scientific verification. These environments provide a shared foundation for supervised fine-tuning, reinforcement learning, and evaluation. Using verified interaction trajectories, we train PhAI-IDE-72B, PhAI-IDE-9B, and PhAI-IDE-4B. The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities. ScienceIDE lays the foundation for an integrated workspace for agent learning and scientific practice, making humanity's scientific software a shared substrate for developing scientific intelligence. Code: https://github.com/aitofound/ScienceIDE

cs.CL↗

Small Language Models as Judges for Rubric-Based Reinforcement Learning

Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific criteria. However, this makes reward computation expensive: training requires repeated rubric judging, often with proprietary APIs or local generative LLM judges with 7B parameters or more. We study whether smaller language models can serve as efficient and reliable rubric-based judges. To make this question measurable, we construct PointRubric and RaR-Science-Static, two pointwise rubric-based evaluation datasets with instance-specific criteria and itemwise satisfaction labels. We compare three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges. Across both datasets, the Qwen3-1.7B Probe judge achieves the strongest criterion-level agreement among these methods, outperforming Generative and Logprob judges. Used as a GRPO reward model, it trains a policy from 0.232 to 0.643 on RaR-Science rubric score, compared with 0.594 for an 8B Generative judge baseline, while the baseline requires 10.7$\times$ more reward-judge time. Task and domain transfer experiments further suggest that Probe judges preserve criterion-level reward structure across settings.

cs.CL↗

Data-Efficient Adaptation of DPA-4 Force Fields to DFT+U Energetics: A Case Study in NiO

Foundation machine-learned force fields (MLFFs) are often pretrained on broad materials datasets whose electronic-structure conventions may not reproduce the phase energetics required for a specific correlated material. Using NiO as a case study, we examine whether incorrect source-level phase energetics can be corrected efficiently through target-level fine-tuning. Along a common structural interpolation, non-spin-polarized PBE and ferromagnetic PBE+U predict opposite energetic orderings of the octahedral Oct and square-planar Sqr phases. Pretrained DPA-4 models adapt rapidly to the NiO PBE+U surface, reaching energy and force root-mean-square errors (RMSEs) of approximately 0.5 meV/atom and 30 meV/Å, respectively, with approximately 170 PBE+U labels. Crucially, models previously fine-tuned to the opposing no-U surface recover the qualitative PBE+U phase ordering with nearly the same target-data efficiency as models fine-tuned directly from their respective pretrained initializations. Our results show that incorrect source-level phase energetics can be reversed through target-level fine-tuning, and suggest a practical multi-fidelity strategy in which pretraining prioritizes broad, consistent, and affordable data, while compact target-level datasets impose energetics through application-specific fine-tuning.

cond-mat.mtrl-sci↗

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

Large language models (LLMs) have shown promising potential in scientific research, enabling tasks ranging from knowledge retrieval to property prediction. Existing science benchmarks mainly focus on perceptual or knowledge-based tasks, largely ignoring the modelling tasks, a fundamental starting point for any real scientific research. For materials science, constructing and manipulating atomic structures is one of the most creative and least automated steps. In this work, we introduce AtomWorld, a benchmark designed to evaluate the abilities of LLMs on structure modifications. The benchmark includes ten fundamental actions under four widely used modelling categories, enabling verifiable evaluation metrics. We find that Claude Opus 4.6 generally performs the best. While the success rate decreases markedly with increasing modelling complexity, with particularly low success rates (below 12\% for rotation) for operations involving complex spatial relations. Our results suggest that contemporary LLMs are better suited as copilots for materials structure modelling rather than fully unsupervised autonomous scientific agents. Beyond evaluation, AtomWorld also serves as a testbed and playground for developing future structure-aware models, including reinforcement learning and agentic approaches.

cond-mat.mtrl-sci↗

Evaluating Large Language Models in Scientific Discovery

Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.

cs.AI↗

Uncovering coupled ionic-polaronic dynamics and interfacial enhancement in Li$_x$FePO$_4$

Understanding and controlling coupled ionic-polaronic dynamics is crucial for optimizing electrochemical performance in battery materials. However, studying such coupled dynamics remains challenging due to the intricate interplay between Li-ion configurations, polaron charge ordering, and lattice vibrations. Here, we develop a fine-tuned machine-learned force field (MLFF) for Li$_x$FePO$_4$ that captures coupled ion-polaron behavior. Our simulations reveal picosecond-scale polaron flips occurring orders of magnitude faster than Li-ion migration, featuring strong correlation to Li configurations. Notably, polaron charge fluctuations are further enhanced at Li-rich/Li-poor phase boundaries, suggesting a potential interfacial electronic conduction mechanism. These results demonstrate the capability of fine-tuned MLFFs to resolve complex coupled transport and provide insight into emergent ionic-polaronic dynamics in multivalent battery cathodes.

cond-mat.mtrl-sci↗

Modeling intercalation chemistry with multi-redox reactions by sparse lattice models in disordered rocksalt cathodes

Modern battery materials can contain many elements with substantial site disorder, and their configurational state has been shown to be critical for their performance. The intercalation voltage profile is a critical parameter to evaluate the performance of energy storage. The application of commonly used cluster expansion techniques to model the intercalation thermodynamics of such systems from \textit{ab-initio} is challenged by the combinatorial increase in configurational degrees of freedom as the number of species grows. Such challenges necessitate efficient generation of lattice models without over-fitting and proper sampling of the configurational space under charge balance in ionic systems. In this work, we introduce a combined approach that addresses these challenges by (1) constructing a robust cluster-expansion Hamiltonian using the sparse regression technique, including $\ell_0\ell_2$-norm regularization and structural hierarchy; and (2) implementing semigrand-canonical Monte Carlo to sample charge-balanced ionic configurations using the table-exchange method and an ensemble-average approach. These techniques are applied to a disordered rocksalt oxyfluoride Li$_{1.3-x}$Mn$_{0.4}$Nb$_{0.3}$O$_{1.6}$F$_{0.4}$ (LMNOF) which is part of a family of promising earth-abundant cathode materials. The simulated voltage profile is found to be in good agreement with experimental data and particularly provides a clear demonstration of the Mn and oxygen contribution to the redox potential as a function of Li content.

cond-mat.mtrl-sci↗

Crystal Structures and Phase Stability of the Li$_2$S-P$_2$S$_5$ System from First Principles

The Li$_2$S-P$_2$S$_5$ pseudo-binary system has been a valuable source of promising superionic conductors, with $α$-Li$_3$PS$_4$, $β$-Li$_3$PS$_4$, HT-Li$_7$PS$_6$, and Li$_7$P$_3$S$_{11}$ having excellent room temperature Li-ion conductivity > 0.1 mS/cm. The metastability of these phases at ambient temperature motivates a study to quantify thermodynamic accessibility. Through calculating the electronic, configurational, and vibrational sources of free energy from first principles, a phase diagram of the crystalline Li$_2$S-P$_2$S$_5$ space is constructed. Well-established phase stability trends from experiments are recovered, such as polymorphic phase transitions in Li$_7$PS$_6$ and Li$_3$PS$_4$, and the metastability of Li$_7$P$_3$S$_{11}$ at high temperature. At ambient temperature, it is predicted that all superionic conductors in this space are indeed metastable, but thermodynamically accessible. Vibrational and configurational sources of entropy are shown to be essential towards describing the stability of superionic conductors. New details of the Li sublattices are revealed, and are found to be crucial towards accurately predicting configurational entropy. All superionic conductors contain significant configurational entropy, which suggests an inherent correlation between superionic conductivity and high configurational entropy.

cond-mat.mtrl-sci↗

Grand-canonical Monte-Carlo simulation methods for charge-decorated cluster expansions

Monte-Carlo sampling of lattice model Hamiltonians is a well-established technique in statistical mechanics for studying the configurational entropy of crystalline materials. When species to be distributed on the lattice model carry charge, the charge balance constraint on the overall system prohibits single-site Metropolis exchanges in MC. In this article, we propose two methods to perform MC sampling in the grand-canonical ensemble in the presence of a charge-balance constraint. The table-exchange method (TE) constructs small charge-conserving excitations, and the square-charge bias method (SCB) allows the system to temporarily drift away from charge neutrality. We illustrate the effect of internal hyper-parameters on the efficiency of these algorithms and suggest practical strategies on how to apply these algorithms to real applications.

cond-mat.mtrl-sci↗

An $\ell_0\ell_2$-norm regularized regression model for construction of robust cluster expansions in multicomponent systems

We introduce the $\ell_0\ell_2$-norm regularization and hierarchy constraints into linear regression for the construction of cluster expansion to describe configurational disorder in materials. The approach is implemented through mixed integer quadratic programming (MIQP). The $\ell_2$-norm regularization is used to suppress intrinsic data noise, while $\ell_0$-norm is used to penalize the number of non-zero elements in the solution. The hierarchy relation between clusters imposes relevant physics and is naturally included by the MIQP paradigm. As such, sparseness and cluster hierarchy can be well optimized to obtain a robust, converged, and effective cluster interactions with improved physical meaning. We demonstrate the effectiveness of $\ell_0\ell_2$-norm regularization in two high-component disordered rocksalt cathode material systems, where we compare the cross-validation and convergence speed, reproduction of phase diagrams, voltage profiles, and Li-occupancy energies with those of the conventional $\ell_1$-norm regularized cluster expansion model.

cond-mat.mtrl-sci↗