Search arXiv⌕ Search

arXiv · 2610.03237

Pinning Decisions Before Failure: Executable Records of Underspecified Choices in AI-Assisted Code Generation

Abstract

A natural-language requirement leaves questions open, and a model asked to implement it settles them silently: across 600 generated test suites from three models, 42.7% contain no test that distinguishes the competing readings. An acceptance example written before implementing is the usual remedy, but an example both readings satisfy resolves nothing. We insert two steps into that practice: enumerate the requirement's underspecified points by name, then for each construct two throwaway implementations differing only in that point and keep a candidate input only if executing both shows they disagree. The result is recorded as a decision pin: the named point, the confirmed input, and the value the person chose between the two exhibited results. The same record then constrains generation and decides compliance by execution. On a benchmark of 40 tasks with paired reference implementations and two models, a separating input is obtained for 92.5% of decision points and a pin identifying the intended decision for 85-90%. As checks on 703 independently generated implementations, pins agree with the benchmark's classification on 96-97%, with disagreements on three tasks, one where both classifiers erred. As generation constraints, pins are honoured at the pinned input in all 210 generations and are never worse than a prose rule on held-out inputs in 38 cells, though no better than prose stating the same scope. Every compliance failure under a prose rule came from the model deciding the rule's scope itself; one such case silently overturned another recorded decision, was attributed to a single rule by leave-one-out on both models, and was missed by text-level reconciliation but caught by re-running the recorded input. The setting yields too few such conflicts to evaluate a regression step, and we say why.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Takeharu Mitsui. 2026-10-02. Pinning Decisions Before Failure: Executable Records of Underspecified Choices in AI-Assisted Code Generation. https://arxiv.org/abs/2610.03237

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

LogLLM: Log-based Anomaly Detection Using Large Language Models

Software systems often record important runtime information in logs to help with troubleshooting. Log-based anomaly detection has become a key research area that aims to identify system issues through log data, ultimately enhancing the reliability of software systems. Existing methods often fall short in capturing the semantic information, typically expressed in natural language, or the sequential dependencies inherent in log sequences. In this paper, we propose LogLLM, a framework that enables collaboration between heterogeneous LLMs for log-based anomaly detection. LogLLM exploits the complementary capabilities of different LLM architectures: a Transformer encoder-based LLM is employed to extract fine-grained semantic vectors from individual log messages, while a Transformer decoder-based LLM is utilized to model sequential dependencies and generate anomaly detection decisions. To enable effective collaboration between these heterogeneous LLMs, we introduce a learnable projector to align their vector representation spaces. Furthermore, we design a progressive three-stage training strategy to optimize the collaboration between heterogeneous LLMs by gradually aligning their representations and adapting them to log anomaly detection. Unlike conventional methods that require log parsers to extract templates, LogLLM preprocesses log messages with regular expressions, streamlining the entire process. Experimental results on four public real-world datasets demonstrate that LogLLM outperforms state-of-the-art methods, achieving an average F$_1$-score improvement of 6.6% over the strongest existing approach. Further analyses provide insights into the effectiveness of the key components of the model architecture and progressive three-stage training strategy.

cs.SE↗

The EmpathiSEr: Development and Validation of Software Engineering Oriented Empathy Scales

Empathy plays a critical role in software engineering (SE), for instance in situations where developers interpret non-technical users' frustration with system usability or where product owners account for the technical constraints experienced by engineers during implementation. Such interactions shape collaboration, communication, and user-centred design outcomes. Although SE research has increasingly recognised empathy as a key human aspect, there remains no validated instrument specifically designed to measure it within the unique socio-technical contexts of SE. Existing generic empathy scales, while well-established in psychology and healthcare, often rely on language, scenarios, and assumptions that are not meaningful or interpretable for software practitioners. These scales fail to account for the diverse, role-specific, and domain-bound expressions of empathy in SE, such as understanding a non-technical user's frustrations or another practitioner's technical constraints, which differ substantially from empathy in clinical or everyday contexts. To address this gap, we developed and validated two domain-specific empathy scales: EmpathiSEr-P, assessing empathy among practitioners, and EmpathiSEr-U, capturing practitioner empathy towards users. Grounded in a practitioner-informed conceptual framework, the scales encompass three dimensions of empathy: cognitive empathy, affective empathy, and empathic responses. We followed a rigorous, multi-phase methodology, including expert evaluation, cognitive interviews, and two practitioner surveys. The resulting instruments represent the first psychometrically validated empathy scales tailored to SE, offering researchers and practitioners a tool for assessing empathy and designing empathy-enhancing interventions in software teams and user interactions.

cs.SE↗

Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis

Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions, and acting on them without grounding in observable evidence leads to compounding failures across downstream stages. Software vulnerability analysis makes this cost concrete and measurable. We address this through a unified cross-language vulnerability lifecycle framework built around three LLM-driven reasoning stages-hybrid structural-semantic detection, execution-grounded agentic validation, and validation-aware iterative repair-governed by a strict invariant: no repair action is taken without execution-based confirmation of exploitability. Cross-language generalization is achieved via a Universal Abstract Syntax Tree (uAST) normalizing Java, Python, and C++ into a shared structural schema, combined with a hybrid fusion of GraphSAGE and Qwen2.5-Coder-1.5B embeddings through learned two-way gating, whose per-sample weights provide intrinsic explainability at no additional cost. The framework achieves 89.84-92.02% intra-language detection accuracy and 74.43-80.12% zero-shot cross-language F1, resolving 69.74% of vulnerabilities end-to-end at a 12.27% total failure rate. Ablations establish necessity: removing uAST degrades cross-language F1 by 23.42%, while disabling validation increases unnecessary repairs by 131.7%. These results demonstrate that execution-grounded closed-loop reasoning is a principled and practically deployable mechanism for trustworthy LLM-driven agentic AI.

cs.SE↗