LLM-based Vulnerability Detection at Project Scale: An Empirical Study
As software complexity grows, automated vulnerability detection becomes increasingly important. Recent LLM-based detectors combine semantic reasoning with static analysis for project-scale scanning, but their practical effectiveness and failure causes remain unclear. We present the first comprehensive empirical study of specialized LLM-based detectors at project scale, evaluating five specialized methods, two general-purpose agents, and four traditional static analyzers on 265 known C/C++ and Java vulnerabilities and 24 active open-source projects/modules. Using Codex-assisted labeling with stratified manual validation, we analyze 6,442 sampled warnings and classify 5,896 false positives using a taxonomy derived from 355 manually inspected reports. Our study yields three findings. First, general-purpose agents achieve the highest recall but struggle with complex code, while both LLM-based and traditional methods suffer from incomplete API modeling. Second, many tools exhibit high false discovery rates on real-world projects/modules. Truncated interprocedural context and incorrect program-point classification are the dominant false-positive causes, with LLM-based tools exhibiting additional reasoning and prompt-compliance failures. Third, LLM-based methods incur substantial costs: hundreds of thousands to hundreds of millions of tokens, API charges exceeding $1,000 per project/module, and runtimes ranging from hours to days. These findings reveal limitations in the robustness, reliability, and scalability of current detectors and inform directions for more effective and practical project-scale vulnerability detection.