Search arXiv⌕ Search

arXiv · 2610.04429

Asking Earns Nothing: Scoring the Decision to Act in BFCL Multi-Turn

Abstract

An agent that lacks the information it needs should ask rather than act, and the task definitions of agent leaderboards say so. BFCL multi-turn builds two of its four categories around a turn on which the model is supposed to ask, and its scorer never looks at that turn: the gold trajectory there is empty, the checker skips it, and the scripted user cannot answer, so asking earns nothing, guessing costs nothing on that turn, and asking twice loses the item. The benchmark also contains the control experiment for that decision. A should-ask item is a base item with one piece of information removed from one turn, so the same request appears twice at the same turn index, once complete and once not: on the first the model should make the call that changes the world, on the second it should ask. We score one decision per pair, whether the model attempted a world-changing call on that turn, read off the stored trajectories with no LLM judge; acting always and asking always both score 50. On the 223 pairs that pose this decision, gpt-5.4 attempts the call on 83.4% of the complete turns and holds back on 78.0% of the incomplete ones, the best decision accuracy of seven models at 80.7%; on the same items the official score ranks it sixth and puts first a model that lands in the middle here. One added line telling gpt-5.4 not to ask pushes it toward acting on both sides of the pair, so its decision accuracy shows no detectable change, while its official score rises by 13.5 to 23.5 points on the two should-ask categories and on the base twins; the opposite line, telling gemma-4-31B-it to ask first, improves its decision by 4.5 points and gains no official score. The score moves with the push toward action, not with the decision. We release the pairs, a turn-level scorer that runs on any BFCL output directory without an API key, and 31 manually verified bad items.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yangze Liu, Zhongyi Han. 2026-10-03. Asking Earns Nothing: Scoring the Decision to Act in BFCL Multi-Turn. https://arxiv.org/abs/2610.04429

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

Trustworthiness has become a key requirement for deploying artificial intelligence in safety-critical applications, yet conventional metrics such as accuracy fail to capture uncertainty or the reliability of predictions, particularly under adversarial or degraded conditions. This paper introduces the Parallel Trust Assessment System (PaTAS), a framework for modeling and propagating trust in neural networks using Subjective Logic (SL). PaTAS operates in parallel with standard neural computation through Trust Nodes and Trust Functions that propagate input, parameter, and activation trust across the network, refining parameter trust during training (Parameter Trust Update) and returning a per-inference trust opinion, quantifying belief, disbelief, and uncertainty, on each prediction (Inference-Path Trust Assessment). Experiments on real-world and adversarial datasets show that these estimates are interpretable and stable and complement accuracy: PaTAS flags predictions that remain accurate yet rely on parameters learned from corrupted data, detects adversarially patched inputs without knowledge of the trigger, and audits training-label corruption without any test data. Compared with output-only uncertainty baselines and dedicated input- and feature-space detectors, out-of-distribution and corruption detection is carried by the conformity-derived input opinions PaTAS consumes, within a regime that a per-dataset diagnostic identifies in advance; trust propagation adds model-side signals these detectors lack, and the composed trust degrades gracefully where any single signal fails. The mechanism extends to convolutional networks. PaTAS thereby provides a foundation for transparent, quantifiable trust reasoning across the AI lifecycle.

cs.AI↗

High-Precision Estimation of the State-Space Complexity of Shogi via the Monte Carlo Method

Determining the state-space complexity of Shogi (Japanese Chess) has been a challenging problem, with previous combinatorial estimates leaving a gap of five orders of magnitude ($10^{64}$ to $10^{69}$). This gap arises from the difficulty of distinguishing positions reachable from the initial position among the vast number of board configurations. In this paper, we present a high-precision statistical estimate of the number of reachable Shogi positions. Here, a position consists of the side to move, the board configuration, and the pieces in hand, without move history, and positions related by exchanging the two players or by horizontal reflection of the board are counted once. Our method combines Monte Carlo sampling with a reachability test that performs reverse search toward the set of ``King-King only'' (KK) positions rather than toward the single initial position. Since every KK position and the initial position are mutually reachable, reaching any KK position proves reachability, and removing pieces from the board is a much simpler objective than reconstructing the initial position. Based on a sample of five billion positions, we estimate the number of such positions in Shogi to be $6.55 \times 10^{68}$; both ends of the approximate $3σ$ confidence interval round to this value at three significant digits. We also applied this method to Mini Shogi, obtaining approximately $2.38 \times 10^{18}$.

cs.AI↗

DERM-3R: A Resource-Efficient Multimodal Agents Framework for Dermatologic Diagnosis and Treatment in Real-World Clinical Settings

Skin diseases impose a substantial and growing global health burden. Modern dermatologic therapies control acute manifestations rapidly, but single-target treatment paradigms, recurrent disease courses, and neglected systemic comorbidities limit their long-term effectiveness. Traditional Chinese medicine (TCM) offers a complementary, holistic approach through syndrome differentiation and individualized treatment, yet its clinical practice is hindered by non-standardized knowledge, incomplete multimodal records, and the difficulty of scaling expert-driven reasoning. We propose DERM-3R, a resource-efficient multimodal agents framework that models TCM dermatologic diagnosis and treatment under limited data and computational resources. We decompose and redefine real-world dermatologic decision-making into three essential issues: fine-grained lesion recognition, multi-view lesion representation with specialist-level pathogenesis modeling, and holistic clinical reasoning for syndrome differentiation and treatment planning. DERM-3R comprises three collaborative agents, DERM-Rec, DERM-Rep, and DERM-Reason, each addressing one of these issues. Built on a lightweight multimodal large language model and fine-tuned by partial-parameter finetuning on 103 real-world TCM psoriasis cases, DERM-3R performs strongly across multiple dermatologic reasoning tasks. Automatic metrics, LLM-as-a-Judge assessments, and human doctors show that, despite extremely limited training data and parameter updates, DERM-3R matches or even surpasses hundred-billion-parameter general-purpose multimodal models such as GPT-5.1 and Gemini-3-Flash. Our results indicate that structured, domain-aware multi-agent modeling is an effective alternative to brute-force scaling for complex clinical tasks, offering a practical and scalable paradigm for multimodal AI in dermatology and integrative medicine.

cs.AI↗