Search arXiv⌕ Search

arXiv · 2509.25193

Devstral: Fine-tuning Language Models for Coding Agent Applications

Abhinav Rastogi·Adam Yang·Albert Q. Jiang·Alexander H. Liu·Alexandre Sablayrolles·Amélie Héliou·Amélie Martin·Anmol Agarwal·Andy Ehrenberg·Andy Lo·Antoine Roux·Arthur Darcet·Arthur Mensch·Baptiste Bout·Baptiste Rozière·Baudouin De Monicault·Chris Bamford·Christian Wallenwein·Christophe Renaudin·Clémence Lanfranchi·Clément Denoix·Corentin Barreau·Darius Dabert Devon Mizelle·Diego de las Casas·Elliot Chane-Sane·Emilien Fugier·Emma Bou Hanna·Gabrielle Berrada·Gauthier Delerce·Gauthier Guinet·Georgii Novikov·Graham Neubig·Guillaume Lample·Guillaume Martin·Himanshu Jaju·Jan Ludziejewski·Jason Rute·Jean-Malo Delignon·JeanHadrien Chabran·Joachim Studnia·Joep Barmentlo·Jonas Amar·Josselin Somerville Roberts·Julien Denize·Karan Saxena·Karmesh Yadav·Kartik Khandelwal·Khyathi Raghavi Chandu·Kush Jain·Lélio Renard Lavaud·Léonard Blier·Lingxiao Zhao·Louis Martin·Lucile Saulnier·Luyu Gao·Marie Pellat·Mathilde Guillaumin·Mathis Felardos·Matthieu Dinot·Maxime Darrin·Maximilian Augustin·Mickaël Seznec·Neha Gupta·Nikhil Raghuraman·Olivier Duchenne·Patricia Wang·Patrick von Platen·Patryk Saffer·Paul Jacob·Paul Wambergue·Paula Kurylowicz·Philomène Chagniot·Pierre Stock·Pravesh Agrawal·Rémi Delacourt·Roman Soletskyi·Romain Sauvestre·Sagar Vaze·Sanchit Gandhi·Sandeep Subramanian·Shashwat Dalal·Siddharth Gandhi·Soham Ghosh·Srijan Mishra·Sumukh Aithal·Szymon Antoniak·Teven Le Scao·Thibaut Lavril·Thibault Schueller·Thomas Foubert·Thomas Robert·Thomas Wang·Timothée Lacroix·Tom Bewley·Valeriia Nemychnikova·Victor Paltz·Virgile Richard·Wen-Ding Li·William Marshall·Xingyao Wang

Abstract

We introduce Devstral-Small, a lightweight open source model for code agents with the best performance among models below 100B size. In this technical report, we give an overview of how we design and develop a model and craft specializations in agentic software development. The resulting model, Devstral-Small is a small 24B model, fast and easy to serve. Despite its size, Devstral-Small still attains competitive performance compared to models more than an order of magnitude larger.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Abhinav Rastogi, Adam Yang, Albert Q. Jiang, Alexander H. Liu, Alexandre Sablayrolles, Amélie Héliou, Amélie Martin, Anmol Agarwal, Andy Ehrenberg, Andy Lo, Antoine Roux, Arthur Darcet, Arthur Mensch, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clémence Lanfranchi, Clément Denoix, Corentin Barreau, Darius Dabert Devon Mizelle, Diego de las Casas, Elliot Chane-Sane, Emilien Fugier, Emma Bou Hanna, Gabrielle Berrada, Gauthier Delerce, Gauthier Guinet, Georgii Novikov, Graham Neubig, Guillaume Lample, Guillaume Martin, Himanshu Jaju, Jan Ludziejewski, Jason Rute, Jean-Malo Delignon, JeanHadrien Chabran, Joachim Studnia, Joep Barmentlo, Jonas Amar, Josselin Somerville Roberts, Julien Denize, Karan Saxena, Karmesh Yadav, Kartik Khandelwal, Khyathi Raghavi Chandu, Kush Jain, Lélio Renard Lavaud, Léonard Blier, Lingxiao Zhao, Louis Martin, Lucile Saulnier, Luyu Gao, Marie Pellat, Mathilde Guillaumin, Mathis Felardos, Matthieu Dinot, Maxime Darrin, Maximilian Augustin, Mickaël Seznec, Neha Gupta, Nikhil Raghuraman, Olivier Duchenne, Patricia Wang, Patrick von Platen, Patryk Saffer, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Philomène Chagniot, Pierre Stock, Pravesh Agrawal, Rémi Delacourt, Roman Soletskyi, Romain Sauvestre, Sagar Vaze, Sanchit Gandhi, Sandeep Subramanian, Shashwat Dalal, Siddharth Gandhi, Soham Ghosh, Srijan Mishra, Sumukh Aithal, Szymon Antoniak, Teven Le Scao, Thibaut Lavril, Thibault Schueller, Thomas Foubert, Thomas Robert, Thomas Wang, Timothée Lacroix, Tom Bewley, Valeriia Nemychnikova, Victor Paltz, Virgile Richard, Wen-Ding Li, William Marshall, Xingyao Wang. 2025-08-08. Devstral: Fine-tuning Language Models for Coding Agent Applications. https://arxiv.org/abs/2509.25193

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap

Search-based software engineering (SBSE), which integrates metaheuristic search techniques with software engineering, has been an active area of research for about 25 years. It has been applied to solve numerous problems across the entire software engineering lifecycle and has demonstrated its versatility in multiple domains. With recent advances in Artificial Intelligence (AI), particularly the emergence of foundation models (FMs) such as large language models (LLMs), the evolution of SBSE alongside these models remains undetermined. In this window of opportunity, we present a research roadmap that articulates the current landscape of SBSE in relation to FMs, identifies open challenges, and outlines potential research directions to advance SBSE through its synergy with FMs. Specifically, we analyze three core aspects: utilizing FMs to enhance SBSE, applying SBSE to advance FMs, and exploring the integration of SBSE and FMs. Furthermore, we present a forward-thinking perspective that envisions the future of SBSE in the era of FMs, highlighting promising research opportunities to address challenges in emerging domains.

cs.SE↗

QMon: Monitoring the Execution of Quantum Circuits with Mid-Circuit Measurement and Reset

Unlike classical software, where logging and runtime tracing can effectively reveal internal execution status, quantum circuits possess unique properties, such as the no-cloning theorem and measurement-induced collapse, that prevent direct observation or duplication of their states. These characteristics make it especially challenging to monitor the execution of quantum circuits, complicating essential tasks such as debugging and runtime monitoring. This paper presents QMon, a practical methodology that leverages mid-circuit measurements, reset operations, and causal-cone replay to monitor selected intermediate values of quantum circuits while preserving their original runtime behavior under explicit conditions. QMon enables the instrumentation of monitoring operators at selected locations within the circuit, allowing comparisons between expected and observed one-qubit outcome probabilities at those locations. Under an ideal noise-free model, we prove that QMon preserves the full circuit state when the monitored qubit is separable, replay is exact, and the measurement record does not control later operations. Across 310 benchmark circuits with a 24-qubit limit per run, QMon monitors 44.54% of gate-qubit locations and 89.75% of circuit qubits at least once, on average. On 2,860 simulated buggy circuits (mutated circuits that alter final outputs), it achieves a detection rate of 62.4%, remaining competitive with three assertion baselines while requiring a median of one planned run per circuit, compared with 21 for the assertion baselines. By collecting multiple checkpoint observations within continued executions, QMon combines practical efficiency with an exact preservation guarantee under explicit conditions.

cs.SE↗

Self-Evolving Coding Agents

Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet most existing coding agents remain largely static after deployment, even though software development is a dynamic, feedback-rich process in which repositories evolve, dependencies change, tests fail, and repair attempts leave reusable experience. This tension has motivated a growing body of work on self-evolving coding agents, where the agent improves its future behavior by persistently updating its framework, memory, skills and tools, components, workflow and topology, or environment and context from prior coding interactions. In this survey, we provide a structured synthesis of this emerging area. We first define the concept of self-evolving coding agents and distinguish it from conventional coding agents and general self-evolving agents. We then develop a taxonomy centered on the targets of evolution, complemented by two orthogonal perspectives: when evolution occurs and which code-specific signals drive it. We further examine the benchmarks used to measure the effect of evolution and related coding products. Across the literature, we find that executable feedback, repository-level context, and coding trajectories make software engineering a natural domain for agent self-evolution, but also introduce challenges in feedback reliability, benchmark overfitting, reversibility, system complexity, safety, cost, and generalization. By organizing existing work around these dimensions, this survey aims to clarify the conceptual boundaries of self-evolving coding agents and provide a foundation for designing more adaptive, reliable, and software-aware agentic systems.

cs.SE↗