arXiv · 2501.02406
A Training-free Method for LLM Text Attribution
Abstract
Verifying the provenance of text is increasingly important for firms, educational institutions, and online platforms as Large Language Models (LLMs) produce output that is nearly indistinguishable from human-generated content. We study the problem of determining whether a given text was generated by a particular LLM while controlling the false positive rate. We model LLM-generated text as a sequential stochastic process and develop training-free statistical tests to (i) distinguish between text produced by two known sets of LLMs and (ii) determine whether text was generated by a known LLM or by a distinguishable unknown source, such as a human or another model. We prove that both Type I and Type II errors decay exponentially with text length, establish analogous guarantees for black-box access via sampling, and provide an information-theoretic lower bound showing that there exist model pairs for which no statistical test can make both errors decay faster than exponentially with text length. Numerical experiments empirically evaluate the tests in practical settings and demonstrate strong overall performance, including under many adversarial edits. Our framework provides rigorous guarantees for LLM provenance detection, with applications to content verification, institutional compliance, and misinformation mitigation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tara Radvand, Izak Duenyas, Ambuj Tewari. 2026-09-10. A Training-free Method for LLM Text Attribution. https://arxiv.org/abs/2501.02406
Cite the original work for its findings. Save a collection to share your selection of sources.