Search arXivSearch

arXiv · 2603.21011

ALL-FEM: Agentic Large Language models Fine-tuned for Finite Element Methods

Abstract

Finite element (FE) analysis guides the design and verification of nearly all manufactured objects. It is at the core of computational engineering, enabling simulation of complex physical systems, from fluids and solids to multiphysics systems. However, implementing FE codes and analyzing simulation results demands expertise across numerical analysis, continuum mechanics, and programming. Conventional Large Language Models (LLMs) can generate FE code, but they hallucinate, lack awareness of variational structures, and cannot close the loop from problem statement to a verified solution. Here, we propose ALL-FEM, an autonomous simulation system that integrates agentic AI with domain-specific, fine-tuned LLMs for FEniCS code generation across solid, fluid, and multiphysics applications. We construct a corpus of 1000+ verified FEniCS scripts by combining 500+ curated expert codes with a retrieval-augmented, multi-LLM pipeline that generates and filters codes for diverse PDEs, geometries, and boundary conditions. We used the corpus to fine-tune LLMs with 3B to 120B parameters. Our agentic framework orchestrates specialized agents, powered by fine-tuned LLMs, to formulate problems as PDEs, generate and debug code and visualize the results. We evaluated the system on 39 benchmarks that include problems of linear/nonlinear elasticity, plasticity, Newtonian/non-Newtonian flow, thermofluids, fluid-structure interaction, phase separation, and transport on moving domains. Embedded in a multi-agent workflow with runtime feedback, the best fine-tuned model (GPT OSS 120B) achieves code-level success of 71.79%, outperforming a non-agentic deployment of GPT 5 Thinking. By showing that relatively small, fine-tuned LLMs, orchestrated through agentic frameworks, can automate FE workflows, ALL-FEM offers a blueprint for autonomous simulation systems in computational science and engineering.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rushikesh Deotale, Adithya Srinivasan, Yuan Tian, Tianyi Zhang, Pavlos Vlachos, Hector Gomez. 2026-04-13. ALL-FEM: Agentic Large Language models Fine-tuned for Finite Element Methods. https://arxiv.org/abs/2603.21011

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

SabreAgent: Language Models at Design Time for Lost-Sales Inventory Control

SabreAgent uses a language model at design time to construct two components for lost-sales inventory control: a product-specific seasonal prior and a validation-selected capped base-stock policy family. During operation, statistical forecasting and inventory optimization use these frozen artifacts to determine orders, with zero language-model calls. We evaluate the approach on the $1{,}320$ instances of InventoryBench. Under the benchmark's cost assumptions, the operations-research core draws on a zero-lead-time optimality result and a projected-inventory rule for positive deterministic lead times. The latter computes replenishment shortfalls by propagating inventory using sales along simulated demand paths. The seasonal prior adds forecast variants alongside the original forecaster, and the selected policy family handles stochastic lead times with order destruction. SabreAgent scores $0.6311$, compared with $0.5380$ for the strongest published baseline, and ranks first in all six benchmark cells. Ablations attribute most of the gain to the OR core. In the paired analysis, the seasonal component adds $1.79\%$ across the three real-data cells, and the search component adds $2.3\%$ across the two stochastic-lead-time cells. These results demonstrate how model-generated priors and policy structure can improve an OR controller through design-time use.

cs.CE

Hierarchical Multi-Task Learning with Liquidity-Aware Signals for Stock Forecasting

Stock price forecasting is a long-standing challenge in computational finance, driven by the inherent randomness of markets and complex temporal patterns. While recent deep-learning models have raised forecasting accuracy by jointly modeling inter-stock and temporal price dynamics, they conflate inter-stock relationships with intra-stock temporal dependencies and focus solely on the univariate objective of price movement. To address these limitations, we propose LiMT, a Hierarchical Multi-Task Learning framework that integrates liquidity-aware signals for stock price forecasting. LiMT employs a Market Regime Encoder (MRE) module that first extracts contemporaneous cross-stock dependencies, then models each stock's temporal dynamics, yielding a unified latent state. Building on this latent state, we introduce a Liquidity-Driven Learning (LDL) module, a mixture-of-experts architecture that features cross-task gating mechanisms to jointly predict price movement, volatility, and trading volume. We further design an Adaptive Portfolio Optimization (APO) mechanism that converts multi-task forecasts into executable portfolio weights under transaction-cost and liquidity constraints. Extensive experiments on the CSI300 and CSI500 benchmarks show that LiMT performs best among strong neural and tree-based baselines across the reported metrics. In realistic CSI300 backtests, APO improves annualized return from 3.99% to 10.01% and Sharpe ratio from 1.22 to 1.86 over equal weighting, showing that the multi-task forecasts translate into deployable portfolio gains.

cs.CE

Scientific capabilities and deployment sustainability of small-scale LLMs in biological wastewater treatment

Large language models (LLMs) are emerging as scientific assistants, yet their computational demands and limited domain specialization constrain sustainable deployment in environmental engineering. Here, we investigate whether domain-specialized small-scale LLMs can combine scientific capability with sustainable deployment in biological wastewater treatment. We developed a benchmark evaluating three scientific capabilities of LLMs: retrospective cognition, comprehension fidelity, and prospective extrapolation. BioWater (8 billion parameters, fine-tuned on specialized domain knowledge) achieved higher comprehension-fidelity scores than participating human experts and performance comparable to a 397-billion-parameter general-purpose LLM in retrospective cognition and prospective extrapolation. Human-BioWater collaboration generated a scientific hypothesis that was subsequently supported by laboratory experiments, demonstrating its potential to contribute to prospective scientific research. We further evaluated the economic and environmental implications of LLM deployment across global wastewater treatment plants (WWTPs). Locally deployed small-scale LLMs became more sustainable than cloud-based large-scale LLMs as inference demand increased in intelligent WWTPs. These findings highlight domain-specialized small-scale LLMs as a promising pathway towards scientifically capable, computationally efficient, and sustainably deployable artificial intelligence for wastewater treatment.

cs.CE