arXiv · 2610.03339
Track fitting as a language translation problem
Abstract
Particle physics analysis depends on information retrieved from particle tracks in a detector. Traditionally, this involves collecting individual detector hits, recombining them to form tracks, and fitting them to a model function. Large Language Models (LLMs) have demonstrated remarkable proficiency in comprehending the grammar and semantics of languages. Our exploratory study leverages transformers to conceptualise particle tracking as a language translation problem, wherein the language of detectors is translated into the language of physics. A simple decoder-only architecture with a custom tokenizer was designed for this task, and the behaviour of the model in this atypical setting provided interesting insights. We evaluated the advantages of this formulation in terms of track-parameter accuracy, and the ease of representing multi-track events, and finally tested the model on real cosmic-muon data. On synthetic datasets, the model outperformed a simple regression technique by an order of magnitude in track-parameter accuracy while being on par with the random sample consensus (RANSAC) method. The model also identified multi-track events with an accuracy of 98\%. The analysis of real cosmic-muon events showed an excellent agreement of the zenith angle distribution with expectations. Our results suggest that a compact language model shows competitive performance and a natural ability to represent multi-track events. Further studies with complex geometries are required to understand the full potential of this technique.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Deepak Samuel, Christy Elsa Koshy. 2026-10-02. Track fitting as a language translation problem. https://arxiv.org/abs/2610.03339
Cite the original work for its findings. Save a collection to share your selection of sources.