Search arXiv⌕ Search

arXiv · 2609.32269

Before Answering: Evidence Sufficiency under Size-Matched Memory Construction

Abstract

Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memory size: on MuSiQue, a classifier that only counts paragraphs reaches an area under the ROC curve (AUROC) of $0.979$ for detecting unsafe memory, higher than the lexical estimator we initially evaluated. We propose a size-matched construction that provably removes this shortcut, and use it to study MemSafe, an estimator that cross-encodes the query with each memory unit and aggregates the units with a set transformer. Across three multi-hop question answering datasets and five seeds, MemSafe reaches $0.968$ and $0.983$ AUROC on MuSiQue and HotpotQA, $0.26$ to $0.39$ above a lexical baseline, while the third dataset, 2WikiMultiHopQA, is saturated. A frozen pretrained cross-encoder with a logistic head already closes $41\%$ of the MuSiQue gap between the lexical baseline and MemSafe. At the same time, MemSafe degrades more than a weak baseline on the unanswerable questions released with MuSiQue, reaches only $0.639$ AUROC on SQuAD~2.0, and needs several thousand clinical training examples before it outperforms a feature-based estimator. Used as a gate for a 7B reader, it reduces the error rate on answered questions from $0.850$ to $0.631$ at $5\%$ coverage, outperforming both reader confidence and, on average, the ground-truth integrity label, although a 7B LLM judge is the better gate at $10\%$ coverage. These results indicate that the way insufficient evidence is constructed matters as much as the estimator that detects it.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Joyanta Jyoti Mondal, Md. Shifatul Ahsan Apurba, Mridul Banik, Md Masud Al Mahmud, Ibne Farabi Shihab. 2026-09-26. Before Answering: Evidence Sufficiency under Size-Matched Memory Construction. https://arxiv.org/abs/2609.32269

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Are We Really Making Much Progress in Text Classification? A Comparative Review

We survey the literature on single-label, multi-label, and hierarchical text classification and provide a quantitative comparison of methods categorized into bag-of-words, sequence-based, and graph- or hierarchy-based approaches. Despite a recent surge in graph-based methods, they do not provide an improvement over fine-tuned transformer models on most evaluated datasets. Decoder-only generative language models show promise in few-shot in-context learning, but appear to lag behind fine-tuned language models when sufficient training data is available. The amount of training data needed for a fine-tuned language model to exceed the performance of a generative model is task-dependent. We further highlight the variance in reported numbers across the literature when applying the same model to the same dataset, which can be traced to the use of different hyperparameter values, such as the fine-tuning learning rate. For practitioners, we recommend using a fine-tuned language model when sufficient training data is available. Otherwise, a frozen generative model, enhanced by few-shot in-context learning or reasoning, is preferable. The source code and further information are available at: https://github.com/ascherp/text-classification-survey

cs.CL↗

Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents

The advancement of LLMs and their accessibility have triggered renewed interest in multi-agent reinforcement learning as robust and adaptive frameworks for dynamically changing environments. This paper introduces \texttt{RL-Focal}, a two-stage RL agent framework that routes and ensembles LLMs. \textit{First}, we develop the Decider RL-agent, which learns to dynamically select an ensemble of small size ($m_i$) among $N$ LLMs ($m_i \ll N$) for incoming queries from a user-defined downstream task $i$, by maximizing both error-diversity and reasoning-performance of the selected ensemble through iterative updates of task-adaptive rewards and policy. \textit{Second}, to enable effective fusion of dynamically selected LLMs, we develop the stage-2 Fusion RL-agent, which learns to resolve reasoning conflicts from different LLMs and dynamically adapt to different ensemble teams composed by the Decider Agent for different downstream tasks. {\em Third}, we introduce the focal diversity metric to better model the error correlations among multiple LLMs further improving the generalization performance of the Decider Agent, which actively prunes the ensemble combinations. By focal diversity, we enhance performance across tasks by effectively promoting reward-aware and policy-adaptive ensemble selection and inference fusion. Extensive evaluations on five benchmarks show that RL-Focal achieves the performance improvement of 8.48\% with an ensemble of small size compared to the best individual LLM in a pool and offers stronger robustness. Code is available \href{https://github.com/git-disl/RL-Focal}{here}.

cs.CL↗

TagPR: Tag-Guided Process Supervision for Personalization Reasoning in Large Language Models

Recent advancements have endowed Large Language Models with impressive general reasoning capabilities. However, these reasoning models often perform worse than non-reasoning models on personalization tasks. While some methods use outcome-based RL to improve personalization reasoning, they fail to supervise the reasoning process. As a result, models may reach correct answers through flawed reasoning chains, limiting further improvement. To address this, we propose TagPR, a novel framework that adds semantic tags to the reasoning process for step-by-step guidance. TagPR first automatically generates a structured, tagged dataset for Supervised Fine-Tuning. It then employs a multi-stage RL process guided by a composite reward signal, which integrates tag-based process supervision with a novel Personalization Reward Model with User Embeddings to achieve fine-grained alignment with user-specific logic. Extensive experiments on public LaMP, LongLaMP, PGraphRAG, and a self-constructed dataset demonstrate that our approach achieves state-of-the-art results, delivering an average improvement of 32.65% over the base model across all LaMP benchmark tasks. Our work demonstrates that tag-guided process supervision is an effective approach for personalization reasoning.

cs.CL↗