Search arXiv⌕ Search

arXiv · 2610.04375

Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents

Abstract

Agents built on large language models (LLMs) build and run software through tool calls. A call reaches its program through several hops, and any hop can change the call without notice. When the changed call fails, the agent retries a correct call, which costs users time and money. Benchmarks and failure analyses do not see the change, because they read the call and its result but not what a hop received. We define intent-execution correspondence (IEC) as the property that the executed action matches the action the emitted call denotes under the tool contract. Our protocol observes what each hop received without executing the call, and names the first hop that changed it by the receiver's own parser. IntAct then delivers the call in a form that this hop cannot alter, or refuses the call. We build IEC-Bench from the changes observed in real-world use, with chains of dependent calls under the execution paths of 4 widely-used harnesses. In 47,828 shell calls within production sessions, Claude Code's Bash tool changes 12.0% of the calls that carry code, escape sequences, or long text. For 80.7% of the calls whose backslashes are changed, the wrong action runs without any reported error. All 10 measured harnesses change a call. Trajectory-based judgment attributes 95.1% of the production failures to the LLM, although the path caused more than half of them. On IEC-Bench, the path raises the token cost per passed task 2.4 times (up to 12.3 times). A hop that changes a call also hides the changes after it, so 55.1% of the failures on one path appear only after its first hop is repaired. IntAct, deployed in a commercial product, recovers 79.2% of the failures with a changed call. Harnesses should therefore be designed and tested hop-by-hop to ensure a correct call executes as intended or is refused.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Boyang Yang, Zhenhao Li, Ziyao Yang, Kanghui Jia, Xin Yin, Mingmou Liu, Haoye Tian. 2026-10-03. Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents. https://arxiv.org/abs/2610.04375

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

Trustworthiness has become a key requirement for deploying artificial intelligence in safety-critical applications, yet conventional metrics such as accuracy fail to capture uncertainty or the reliability of predictions, particularly under adversarial or degraded conditions. This paper introduces the Parallel Trust Assessment System (PaTAS), a framework for modeling and propagating trust in neural networks using Subjective Logic (SL). PaTAS operates in parallel with standard neural computation through Trust Nodes and Trust Functions that propagate input, parameter, and activation trust across the network, refining parameter trust during training (Parameter Trust Update) and returning a per-inference trust opinion, quantifying belief, disbelief, and uncertainty, on each prediction (Inference-Path Trust Assessment). Experiments on real-world and adversarial datasets show that these estimates are interpretable and stable and complement accuracy: PaTAS flags predictions that remain accurate yet rely on parameters learned from corrupted data, detects adversarially patched inputs without knowledge of the trigger, and audits training-label corruption without any test data. Compared with output-only uncertainty baselines and dedicated input- and feature-space detectors, out-of-distribution and corruption detection is carried by the conformity-derived input opinions PaTAS consumes, within a regime that a per-dataset diagnostic identifies in advance; trust propagation adds model-side signals these detectors lack, and the composed trust degrades gracefully where any single signal fails. The mechanism extends to convolutional networks. PaTAS thereby provides a foundation for transparent, quantifiable trust reasoning across the AI lifecycle.

cs.AI↗

High-Precision Estimation of the State-Space Complexity of Shogi via the Monte Carlo Method

Determining the state-space complexity of Shogi (Japanese Chess) has been a challenging problem, with previous combinatorial estimates leaving a gap of five orders of magnitude ($10^{64}$ to $10^{69}$). This gap arises from the difficulty of distinguishing positions reachable from the initial position among the vast number of board configurations. In this paper, we present a high-precision statistical estimate of the number of reachable Shogi positions. Here, a position consists of the side to move, the board configuration, and the pieces in hand, without move history, and positions related by exchanging the two players or by horizontal reflection of the board are counted once. Our method combines Monte Carlo sampling with a reachability test that performs reverse search toward the set of ``King-King only'' (KK) positions rather than toward the single initial position. Since every KK position and the initial position are mutually reachable, reaching any KK position proves reachability, and removing pieces from the board is a much simpler objective than reconstructing the initial position. Based on a sample of five billion positions, we estimate the number of such positions in Shogi to be $6.55 \times 10^{68}$; both ends of the approximate $3σ$ confidence interval round to this value at three significant digits. We also applied this method to Mini Shogi, obtaining approximately $2.38 \times 10^{18}$.

cs.AI↗

DERM-3R: A Resource-Efficient Multimodal Agents Framework for Dermatologic Diagnosis and Treatment in Real-World Clinical Settings

Skin diseases impose a substantial and growing global health burden. Modern dermatologic therapies control acute manifestations rapidly, but single-target treatment paradigms, recurrent disease courses, and neglected systemic comorbidities limit their long-term effectiveness. Traditional Chinese medicine (TCM) offers a complementary, holistic approach through syndrome differentiation and individualized treatment, yet its clinical practice is hindered by non-standardized knowledge, incomplete multimodal records, and the difficulty of scaling expert-driven reasoning. We propose DERM-3R, a resource-efficient multimodal agents framework that models TCM dermatologic diagnosis and treatment under limited data and computational resources. We decompose and redefine real-world dermatologic decision-making into three essential issues: fine-grained lesion recognition, multi-view lesion representation with specialist-level pathogenesis modeling, and holistic clinical reasoning for syndrome differentiation and treatment planning. DERM-3R comprises three collaborative agents, DERM-Rec, DERM-Rep, and DERM-Reason, each addressing one of these issues. Built on a lightweight multimodal large language model and fine-tuned by partial-parameter finetuning on 103 real-world TCM psoriasis cases, DERM-3R performs strongly across multiple dermatologic reasoning tasks. Automatic metrics, LLM-as-a-Judge assessments, and human doctors show that, despite extremely limited training data and parameter updates, DERM-3R matches or even surpasses hundred-billion-parameter general-purpose multimodal models such as GPT-5.1 and Gemini-3-Flash. Our results indicate that structured, domain-aware multi-agent modeling is an effective alternative to brute-force scaling for complex clinical tasks, offering a practical and scalable paradigm for multimodal AI in dermatology and integrative medicine.

cs.AI↗