arXiv · 2610.01664
Code Detectors Have a Half-Life: Obsolescence and Metric Illusions in LLM-Generated Code Detection
Abstract
Code detectors can become obsolete as code-generating models evolve: a detector validated on one generation of models may not transfer to the next. We call this limited useful life a detector half-life. We evaluate eight general-purpose LLM judges and three dedicated detectors on human-written code and code produced by seven generators across C++, Java, and Python. Our results reveal two problems. First, performance varies considerably across generators and prompting strategies, suggesting that some detectors rely on generator-specific patterns rather than general evidence of code provenance. Second, accuracy can conceal severe prediction bias. DetectCodeGPT and GPT-Sniffer achieved an accuracy of 0.50 but an F1 score of 0.00 across all generators because they classified almost every sample as AI-generated. However, general-purpose LLM judges achieved stronger accuracy and F1 scores. Our results show that general-purpose LLMs are promising training-free judges of code provenance and can outperform dedicated detectors. However, their reliability depends on the judge model, the code generator, and the prompting strategy. We therefore recommend evaluating LLM judges across multiple generators and reporting macro-F1 alongside class-specific precision and recall.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alberick Euraste Djire. 2026-10-01. Code Detectors Have a Half-Life: Obsolescence and Metric Illusions in LLM-Generated Code Detection. https://doi.org/10.1145/3832783.3844566
Cite the original work for its findings. Save a collection to share your selection of sources.