Search arXivSearch

arXiv · 2507.02873

Using Large Language Models to Study Mathematical Practice

Abstract

The philosophy of mathematical practice (PMP) looks to evidence from working mathematics to help settle philosophical questions. One prominent program under the PMP banner is the study of explanation in mathematics, which aims to understand what sorts of proofs mathematicians consider explanatory and what role the pursuit of explanation plays in mathematical practice. In an effort to address worries about cherry-picked examples and file-drawer problems in PMP, a handful of authors have recently turned to corpus analysis methods as a promising alternative to small-scale case studies. This paper reports the results from such a corpus study facilitated by Google's Gemini 2.5 Pro, a model whose reasoning capabilities, advances in hallucination control and large context window allow for the accurate analysis of hundreds of pages of text per query. Based on a sample of 5000 mathematics papers from arXiv.org, the experiments yielded a dataset of hundreds of useful annotated examples. Its aim was to gain insight on questions like the following: How often do mathematicians make claims about explanation in the relevant sense? Do mathematicians' explanatory practices vary in any noticeable way by subject matter? Which philosophical theories of explanation are most consistent with a large body of non-cherry-picked examples? How might philosophers make further use of AI tools to gain insights from large datasets of this kind? As the first PMP study making extensive use of LLM methods, it also seeks to begin a conversation about these methods as research tools in practice-oriented philosophy and to evaluate the strengths and weaknesses of current models for such work.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

William D'Alessandro. 2025-06-16. Using Large Language Models to Study Mathematical Practice. https://arxiv.org/abs/2507.02873

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Come for the vibe, stay for the math

This article describes our experiences in mathematical outreach over the past decade. We talk about specific activities, but also general principles that we've learned along the way.

math.HO

Graduate Mathematics in the Age of AI: Forming Mathematicians for Original, Independent, and Responsible Inquiry

Artificial intelligence can increasingly produce plausible, sophisticated mathematical material faster than a developing graduate student can understand or verify it. A sophisticated result or paper draft therefore becomes weaker evidence of the student's own mathematical development. This creates a formation gap between output and personal capacity, and a trust gap between a convincing argument and warranted acceptance. The formation gap can persist even when the student understands the output: understanding a supplied argument does not by itself establish the capacity to initiate and direct inquiry. These gaps are not the whole story. AI can also help students explore examples, compare approaches, enter unfamiliar areas, and undertake ambitious research. The task is to design an apprenticeship that realizes these possibilities while developing substantive mathematical command. The central purpose of a mathematics PhD is to form mathematicians capable of original, independent, and responsible inquiry, including inquiry conducted with AI. This document develops that objective through four connected capacities: competence, judgment, independence, and responsibility. It distinguishes a work's contribution to mathematics from the evidence it provides of a student's formation; explains how a known answer can initiate rather than end creative inquiry; and proposes changes in learning activities, assessment, doctoral originality, advising, and institutional support. Purposeful independent work and ambitious AI-assisted research are complementary parts of the model. Its recommendations include proportionate contribution statements, recognition of advising costs, and staged pilots evaluating both mathematical ability and effective human--AI collaboration. The aim is not to preserve an inherited sequence of training, but to improve mathematical formation as mathematical practice changes.

math.HO