Search arXiv⌕ Search

arXiv · 2609.30953

CoCoRerank: Towards Conventional Commit Message Generation by Component and Candidate Consistency Reranking

Abstract

Commit messages are essential for understanding software changes, yet automatic commit message generation typically treats a message as an unstructured text sequence. This limits its ability to support standardized development workflows, where commit messages are often expected to follow the Conventional Commits Specification (CCS) in the form type (scope): subject. In this paper, we study conventional commit message generation under the complete CCS format. We construct a new benchmark of 86,688 high-quality commits collected from open-source GitHub repositories, with each message normalized into type, scope, and subject through structural normalization and semantic quality filtering. Based on this benchmark, we propose a two-dimensional consistency-based reranking framework named CoCoRerank for LLM-based generation. CoCoRerank exploits horizontal consistency among the code change, type, scope, and subject, as well as vertical consensus across multiple generated candidates. Experiments with representative CMG baselines, LLM generators, reranking strategies, and ablation variants show that CoCoRerank improves both structural component prediction and subject generation quality. The results demonstrate that complete CCS supervision and multidimensional consistency modeling provide an effective foundation for accurate and standardized commit message generation. The artifact is publicly released at https://github.com/bluewhalebug/CoCoRerank.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Shaopeng Jia, Yali Du, Ming Li. 2026-09-25. CoCoRerank: Towards Conventional Commit Message Generation by Component and Candidate Consistency Reranking. https://arxiv.org/abs/2609.30953

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture

AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture that enforces statistical rigor at two levels: a Haskell embedded domain-specific language (the Research monad) that makes it impossible to test a hypothesis without updating the error budget, and a declarative scaffold that fixes the data flow and the statistical test, together with an OS-level sandbox that makes validation data physically absent from the environment in which LLM-generated code runs. We treat FDR control as a formal requirement and trace it to the implementation. We ground the design in a machine-checked Lean~4 formalization of LORD online false-discovery-rate (FDR) control: we derive its error budget and prove marginal FDR control, and full FDR control when thresholds do not adapt to earlier rejections. We then verify in SPARK/Ada that the LORD thresholds, computed in IEEE~754 arithmetic, never exceed the available wealth, given a margin condition that our configurations meet by a factor of at least eight; without a margin the property fails. To our knowledge this is the first machine-checked proof of an online FDR control theorem. In simulation, the architecture holds the false discovery rate near 1\% against a 5\% target, where a naive approach reaches 41\%. In end-to-end case studies, a valid test avoids the false discoveries a flawed one produces, yet still finds real effects when the data allow. An adversarial evaluation confirms that, inside the sandbox, generated code cannot read the held-out data even when given its exact path.

cs.SE↗

LLM-based Vulnerability Detection at Project Scale: An Empirical Study

As software complexity grows, automated vulnerability detection becomes increasingly important. Recent LLM-based detectors combine semantic reasoning with static analysis for project-scale scanning, but their practical effectiveness and failure causes remain unclear. We present the first comprehensive empirical study of specialized LLM-based detectors at project scale, evaluating five specialized methods, two general-purpose agents, and four traditional static analyzers on 265 known C/C++ and Java vulnerabilities and 24 active open-source projects/modules. Using Codex-assisted labeling with stratified manual validation, we analyze 6,442 sampled warnings and classify 5,896 false positives using a taxonomy derived from 355 manually inspected reports. Our study yields three findings. First, general-purpose agents achieve the highest recall but struggle with complex code, while both LLM-based and traditional methods suffer from incomplete API modeling. Second, many tools exhibit high false discovery rates on real-world projects/modules. Truncated interprocedural context and incorrect program-point classification are the dominant false-positive causes, with LLM-based tools exhibiting additional reasoning and prompt-compliance failures. Third, LLM-based methods incur substantial costs: hundreds of thousands to hundreds of millions of tokens, API charges exceeding $1,000 per project/module, and runtimes ranging from hours to days. These findings reveal limitations in the robustness, reliability, and scalability of current detectors and inform directions for more effective and practical project-scale vulnerability detection.

cs.SE↗

A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code

As Large Language Models (LLMs) become integrated into software development workflows, concerns regarding unintentional biases in AI-generated code. Although evidence suggests these biases exist, limited research has systematically identified, categorized, and explained them. This study investigates bias in AI-generated code and evaluates whether LLMs can reliably identify and explain it through a taxonomy-driven framework. We extended an existing dataset of biased AI-generated Python code and manually annotated snippets with bias categories and human-authored justifications to establish a ground-truth dataset. Using this dataset, we evaluated proprietary and open-source LLMs as automated bias detection and justification systems through ICL. Finally, we analyzed similarity between LLM-generated explanations and human-authored justifications using structured justification and code identification metrics. Our findings demonstrate that LLMs can effectively support code bias identification and explanation. Gemini achieved 80.14% classification accuracy, with 84.0% precision and 95.7% recall, while the best open-source alternative, Qwen3-coder, achieved 82.45% accuracy, 68.64% precision, and 80.22% recall. Additionally, the models achieved justification similarity scores of 80.4% and 80.14%, respectively, relative to human-authored reasoning, and code identification similarity scores of 86.0% and 87.82%. These results suggest that LLMs can detect biased logic in generated Python code and produce explanations that substantially align with expert interpretations.

cs.SE↗