Search arXiv⌕ Search

arXiv · 2609.38184

RealGUINoise: An Interactive Cross-Platform Benchmark for GUI Agent Robustness under Real-World Interface Noise

Abstract

Graphical User Interface (GUI) agents and Computer-Using Agents (CUAs) are rapidly becoming practical tools. However, real-world deployment increasingly exposes performance failures and safety risks, while a major yet underexplored source of these problems lies in the complex and noisy conditions of everyday interfaces. Existing benchmarks largely assume clean environments or focus narrowly on security-specific settings, and lack a unified framework for consistent, automated end-to-end evaluation across diverse agents and platforms. Hence, we introduce RealGUINoise, an interactive cross-platform, extensible benchmark for systematically evaluating GUI Agents under common realistic interface noise in fully interactive environments. Specifically, RealGUINoise comprises 42 noise types spanning web, desktop, and mobile tasks and integrates 7 representative agent frameworks. It evaluates these agents on real-world daily tasks through real-time interaction, comparing their performance against task-specific golden rubrics and clean-environment trajectories in terms of reliability, safety, and trajectory-level behavior. Our experiments show that these noises not only degrade task performance but also substantially redirect agents' action trajectories and increase their propensity for unsafe behavior. These findings expose a critical gap between capability in clean environments and dependable operation in real-world settings, establishing RealGUINoise as a testbed for developing more robust and trustworthy GUI agents.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yongjiang Wu, Junyuan Zhang, Ada Chen, Kuiyi Gao, Wenxuan Wang. 2026-08-07. RealGUINoise: An Interactive Cross-Platform Benchmark for GUI Agent Robustness under Real-World Interface Noise. https://arxiv.org/abs/2609.38184

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AI Behavioral Science: A Framework and Agenda

We discuss the challenges and opportunities present in the rapidly emerging area of ``AI Behavioral Science.'' We frame it via three subfields. First, as AI becomes ubiquitous and is increasingly proprietary and opaque, it becomes vital to develop models of AI and methods for assessing AI behavior. We outline how tools developed to assess people's behaviors by social scientists can be used to model, assess and infer AI's behaviors biases, tendencies, and heuristics. Second, we also discuss how AI can change the ways in which we learn about human behavior. Beyond its computational power, AI offers new techniques for simulating, inferring, predicting, and analyzing human behaviors. Third, as humans and AI are interacting in increasingly complex and intertwined systems, we need to analyze and model human-AI interactions including how human and AI behaviors depend on interactions at the individual level, how interacting systems of humans and AI behave, and ultimately how AI's integration into society affects economic and political outcomes. We discuss current research, questions, agendas, and goals in each of these three subfields and how they depend upon each other.

cs.HC↗

Scaling Peer Assessments: An Integrity Report from a Large Engineering Internship

Assessing learning in large classrooms presents a significant challenge for individual instructors, who may have limited capacity to evaluate the understanding, participation, and assessment behaviour of every student. Peer assessments have been a way of distributing this responsibility among learners, allowing them to evaluate and provide feedback to one another while reducing dependence on instructor-led assessments. Building on this approach, we implemented a peer validation model within a large, multi-institutional internship programme in which students who demonstrated sufficient understanding were authorised to assess and validate their peers through short oral discussions. The assessment process began with the instructor validating a small group of students, who were then authorised to validate their peers, allowing the process to gradually expand across the cohort and operate at scale. This study examines how participants experienced the model and the extent to which assessment integrity was maintained, using an end-of-programme survey of 238 consenting respondents. Most participants regarded the activity as worthwhile, with 79.8% reporting that they solved problems they could not previously solve. However, 29.0% acknowledged at least one instance of reduced effort, a lowered validation standard, or reciprocal validation, while 88.7% believed that at least a little validation had occurred without proper examination. When asked how the process could be strengthened, participants selected post-validation discussion of solutions approximately twice as often as closer auditing or mentor-led validation. These findings provide descriptive evidence of both the potential and the integrity challenges of using peer validation as a scalable assessment approach in large learning environments.

cs.HC↗

Who Thinks First? Designing Productive Friction with Engage-to-Unlock GenAI

Generative AI can support writing, but frictionless access may cause cognitive offloading before users develop their own ideas. We introduce Engage-to-Unlock, a productive-friction mechanism that unlocks generative capabilities after users meaningfully engage with the task. In a controlled experiment (N = 398), participants completed a writing task under one of four conditions: Human-Only, Standard Chatbot, Engage-to-Unlock, or Time-Matched Unlock, which matched unlock timing to Engage-to-Unlock participants but independent of users' engagement, then evaluated passages for evidence and inferential errors. Results show that Engage-to-Unlock redistributed effort across tasks: participants spent more time writing and less time evaluating, without increasing overall task duration. They also submitted more prompts than in other AI-assisted conditions and showed the highest accuracy-per-time evaluation efficiency across conditions. These findings suggest that designing GenAI access to encourage early human engagement may provide a productive form of friction, while retaining active AI use and efficient downstream evaluation.

cs.HC↗