Search arXiv⌕ Search

arXiv · 2610.10141

AdaT$^2$: Adaptive Test Transformations for Black-Box Boundary Testing of Conversational Agents

Abstract

Conversational agents based on large language models (LLMs) must comply with policies. Each condition in a policy draws a boundary between user requests, and the agent must behave differently on its two sides. We present AdaT$^2$, which extracts statements from the agent's replies in exploratory conversations with an LLM acting as the user, and uses the statements to guide boundary test generation. Each statement describes one condition and the behavior expected when the condition holds. Besides plain tests guided by single statements, AdaT$^2$ writes transformed tests guided by pairs of a statement and a test transformation instruction, such as "omit one required input". The instruction of a pair can move a test to the other side of the statement's boundary or to another boundary. The statements and instructions form far more pairs than a run can try, and many pairs are not applicable. Adaptive pair selection therefore chooses the statement of each pair by novelty and the instruction with the bandit algorithm Bayes-UCB, which learns from whether earlier pairs yielded a test and whether the agent passed it according to an LLM judge. Our benchmark counts the two sides of each boundary separately and distinguishes boundaries explicitly defined by the agent's prompt, tool code, or knowledge base from boundaries that the agent's LLM infers from domain knowledge. On four domains of $τ^3$-bench, 62.7% to 83.3% of AdaT$^2$'s tests are valid boundary tests whose expected behavior is explicitly defined, higher in every domain than AgentEval's (47.7% to 68.1%). Transformed tests add 13 to 46 explicitly defined boundaries that plain tests miss. As regression tests, AdaT$^2$'s test suites detect all eight seeded policy faults in the airline domain and four of eight in the retail domain, and AgentEval's test suites, with fewer than a third as many tests, detect five and two, respectively.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Liting Lin, Boxi Yu, Qinghua Xu, Yuzhong Zhang, Lionel Briand, Emir Muñoz. 2026-10-07. AdaT$^2$: Adaptive Test Transformations for Black-Box Boundary Testing of Conversational Agents. https://arxiv.org/abs/2610.10141

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Skill-Adaptive Imitation Learning for UI Test Reuse

To alleviate the substantial cost of manually crafting user interface (UI) test cases, UI test migration aims to automatically generate test cases for a target mobile application (app) by adapting those from a source app that shares similar functionalities. Traditionally, this process has been approached as a sequential UI-event-mapping problem, where events in the source app are mapped to those in the target one based on their textual descriptions. Prior research has extensively focused on enhancing the event-mapping accuracy of NLP models. Although the advent of large language models (LLMs) with impressive NLP capabilities suggests the potential for near-perfect event-mapping, our study demonstrates that even the highly accurate event-mapping of LLMs is insufficient to address the implementation discrepancies between the source and the target apps, reducing the overall effectiveness of LLM-driven solutions for UI test migration. To address this challenge, in this paper, we propose SAIL, a skill-adaptive imitation learning framework designed to enhance the effectiveness of UI test migration through two key designs. First, SAIL leverages the source test cases as demonstrations and employs a multi-level abstraction of test cases' underlying skills, so as to extract the testing information from source test cases as the knowledge base for the subsequent test generation on the target app. Second, SAIL selectively reuses a subset of the learned skills to guide the generation of test cases for the target app with its novel context- and history-aware skill adaptation. While SAIL can be instantiated with any imitation learning techniques, we utilize the in-context learning capabilities of LLMs to instantiate SAIL. Evaluations results show that SAIL substantially improves the effectiveness of UI test migration, with 149\% higher success rate than state-of-the-art approaches.

cs.SE↗

CAFÉ: Causal Black-Box Testing of Machine Unlearning

Machine learning models are increasingly deployed as software components that must evolve as requirements change. When specific training records or features must no longer influence a deployed model, machine unlearning aims to remove that influence without retraining from scratch. Because unlearning is often approximate, its effectiveness must be tested. Such tests must often treat the model as a black box, without access to its parameters, training history, or unlearning procedure. Features pose a further challenge: even after a feature is removed from a model's inputs, its influence can persist through downstream features. Many existing checks examine only the feature's direct use and can therefore certify a model that still depends on it. We frame unlearning testing as specification-based testing and present CAFÉ, which, using only a deployed model's predictions, intervenes on the feature, propagates the change to its downstream features, and checks whether the predictions still respond. CAFÉ measures a target's residual influence through both its direct and indirect causal paths, and its fine-grained diagnostics show which channels and subgroups still carry it. On two causal-network benchmarks with four unlearning methods, CAFÉ ranks residual influence with 0.92--0.93 pairwise accuracy, against at most 0.71 for existing checks, which fail in both directions: they certify models whose influence persists through downstream features and flag correctly unlearned ones. On real census data, CAFÉ likewise exposes influence that survives retraining yet goes unnoticed by direct-input checks.

cs.SE↗

PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring

As Large Language Models (LLMs) are increasingly integrated into software development workflows, their trustworthiness has become a critical concern. However, in dependency recommendation scenarios, the reliability of LLMs is undermined by widespread package hallucinations, where models often recommend hallucinated packages. Recent studies have proposed a range of approaches to mitigate this issue. Nevertheless, existing approaches typically merely reduce hallucination rates rather than eliminate them, leaving persistent software security risks. In this work, we argue that package hallucinations are theoretically preventable based on the key insight that package validity is decidable through finite and enumerable authoritative package lists. Building on this, we propose PackMonitor, the first approach capable of fundamentally eliminating package hallucinations by continuously monitoring the model's decoding process and intervening when necessary. To implement this in practice, PackMonitor addresses three key challenges: (1) determining when to trigger intervention via a Context-Aware Parser that continuously monitors model outputs and selectively activates intervening only during installation command generation; (2) resolving how to intervene by employing a Package-Name Intervenor that strictly limits the decoding space to an authoritative package list; and (3) ensuring monitoring efficiency through a DFA-Caching Mechanism that enables scalability to millions of packages with negligible overhead. Extensive experiments on five widely used LLMs demonstrate that PackMonitor is a training-free, plug-and-play solution that consistently reduces package hallucination rates to zero while maintaining low-latency inference and preserving original model capabilities.

cs.SE↗