Search arXivSearch

arXiv subjects

Jaafar El-Awady

Publications and source records attributed to Jaafar El-Awady.

2 recordsLinked to original sources

Can Coding Agents Reproduce Findings in Computational Materials Science?

Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject matter experts, we curate a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims. We then evaluate multiple representative coding agent settings across several foundation models. Our results show that current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 53%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and that they fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.

cs.CL

AIMD-L: An automated laboratory for high-throughput characterization of structural materials for extreme environments

Rapid developments in artificial intelligence and machine learning as applied to materials science are creating an urgent need for experimental data, which can be provided by high-throughput and autonomous laboratories. To date most demonstrations of such laboratories have focused on functional materials, with less attention paid to structural materials. We present here the Artificial Intelligence in Materials Design Laboratory (AIMD-L), an automated, high-throughput facility for characterizing the microstructure and properties of structural metals and ceramics, with an emphasis on materials in extreme environments. AIMD-L has two custom instruments for characterization of structural materials: HELIX for shock studies of materials, and MAXIMA for X-ray diffraction and X-ray fluorescence spectroscopy. Specifically designed for high-throughput studies, HELIX and MAXIMA are each capable of collecting data at rates two to three orders of magnitude faster than conventional systems. A third experimental station, SPHINX, is a commercial nanoindenter modified for integration into the automated workflow of AIMD-L. A user (which may be human or an AI agent) directs the experiments to be carried out by means of a centralized control program. The experimental stations are linked by a conveyance that moves samples around the lab, with a robot at each station for sample transfer in/out of the instrument. The experimental stations also communicate with a common data layer that streams data autonomously from each instrument to a data portal, where their arrival triggers automated workflows for data reduction and analysis. The processed data are immediately available to the human operator or agentic AI, forming a closed loop for rapid decision-making and experimental control.

cond-mat.mtrl-sci