Search arXiv⌕ Search

arXiv · 2609.32496

Locally Sound, Globally Insufficient: The Local-Global Gap in Multi-Hop Reasoning

Abstract

Reliable multi-hop reasoning requires more than locally supported steps: a trace can be sound at every reasoning step yet still fail to answer the question as a whole. We call this failure regime the local-global gap (LGG), in which the trace is locally sound yet globally insufficient. Local soundness requires each step to be supported by the available evidence and preceding steps, whereas global sufficiency requires the reasoning trace to align with the question and establish the submitted answer. In a human-adjudicated diagnostic of 2,598 responses across three multi-hop QA benchmarks and three models, we find that the LGG occurs in every benchmark-model combination and accounts for nearly half of globally insufficient responses overall. However, conventional faithfulness verifiers that check claims against the evidence largely miss these failures: at thresholds retaining at least 95% of reliable traces, recall for LGG cases is substantially lower than that for locally unsound traces. To address these failures, we formalize three dependencies for reliable reasoning: evidence-to-step support, question-to-trace alignment, and trace-to-answer closure. Instead of post-hoc diagnosis, we introduce E-Closure to supervise these dependencies during training, combining generation supervision on supported original and counterfactual responses with bidirectional switching constraints. Averaged over three benchmarks and three backbones, existing fine-tuning baselines improve accuracy and local soundness over the base models, but at the cost of global sufficiency. E-Closure improves both: among all fine-tuned methods, it achieves the highest average accuracy (92.8%) and trace reliability (89.0%) while yielding the lowest LGG rate (6.2%).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bohao Chu, Hendrik Damm, Qianli Wang, Hui Wang, Shuning Zhang, Christoph M. Friedrich, Norbert Fuhr. 2026-09-26. Locally Sound, Globally Insufficient: The Local-Global Gap in Multi-Hop Reasoning. https://arxiv.org/abs/2609.32496

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents

Memory-augmented conversational agents enable personalized interactions using long-term user memory and have gained substantial traction. However, existing benchmarks primarily focus on whether agents can recall and apply user information, while overlooking whether such personalization is used appropriately. In fact, agents may overuse personal information, producing responses that feel forced, intrusive, or socially inappropriate to users. We refer to this issue as \emph{over-personalization}. In this work, we formalize over-personalization into three types: Irrelevance, Repetition, and Sycophancy, and introduce \textbf{OP-Bench} a benchmark of 1,700 verified instances constructed from long-horizon dialogue histories. Using \textbf{OP-Bench}, we evaluate multiple large language models and memory-augmentation methods, and find that over-personalization is widespread when memory is introduced. Further analysis reveals that agents tend to retrieve and over-attend to user memories even when unnecessary. To address this issue, we propose \textbf{Self-ReCheck}, a lightweight, model-agnostic memory filtering mechanism that mitigates over-personalization while preserving personalization performance. Our work takes an initial step toward more controllable and appropriate personalization in memory-augmented dialogue systems.

cs.CL↗

Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents

Reinforcement learning (RL) trains small language model agents to answer multi-hop questions by retrieving evidence over multiple turns, but reported gains typically rely on thousands of on-policy rollouts per update. We study RL for such agents under the budget constraint of commodity GPUs, where each update samples only a few rollouts per question. Under this constraint, most sampled trajectories retrieve none of the required evidence, so the outcome reward gives the policy little to learn from and small agents settle for answering without retrieval, a failure we call \emph{retrieval collapse}. David-GRPO addresses this with two mechanisms: (1) \emph{Expert trajectory seeding} places a handful of off-policy expert trajectories into the GRPO groups of the early updates, and (2) \emph{evidence-guided continuation} rewards evidence coverage and resumes the most promising partial trajectory. The evidence for each training question is constructed from the corpus link graph, so no annotated evidence is required. On six multi-hop QA benchmarks, David-GRPO trained on four RTX 3090 GPUs with 144 rollouts per step brings Qwen2.5-1.5B to 22.6 average EM against 11.9 for the best baseline under the same budget, matches Tree-GRPO trained with 20 times more rollouts, and, unlike the baselines that stop after at most one search, learns to retrieve across turns. The implementation is available at: https://github.com/AsadalJung/David-GRPO

cs.CL↗

Denoising Time Matters:Diverse Generation in Diffusion Language Models

Diffusion language models (Diffusion-LMs) generate text through iterative denoising, exposing a temporal structure that is largely absent from autoregressive decoding. In this paper, we show that this temporal structure provides a useful control axis for generation diversity: early denoising steps mainly determine high-level semantic trajectories, while later steps refine lexical realization. Motivated by this observation, we propose Time-Annealed Perturbation Sampling (TAPS), a training-free inference strategy that samples nearby conditioning trajectories through time-aware, manifold-constrained perturbations. TAPS encourages semantic branching during early denoising and anneals the perturbation away before refinement, improving exploration while preserving prompt alignment, generation quality, and reasoning ability. Experiments on multiple Diffusion-LM backbones, including non-autoregressive and semi-autoregressive models, show that TAPS consistently improves semantic and lexical diversity across open-ended and instruction-following generation tasks, while preserving reasoning ability on verifiable reasoning benchmarks with negligible overhead.

cs.CL↗