Search arXiv⌕ Search

arXiv · 2609.32385

DashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard Analysis

Abstract

Interactive dashboards require users to reveal and connect evidence across stateful interactions. Although graphical user interface (GUI) agents could automate this process, existing dashboard benchmarks primarily report final answers or task success. They provide limited insight into whether failures arise from maintaining the analytical process, selecting actions, or grounding visual targets. We introduce DashAct, to our knowledge the first benchmark to diagnose these failures at a fine-grained level within the same dashboard task. DashAct contains 357 human-verified interaction trajectories with milestone dependencies and hierarchical target annotations. Its progressive diagnostic cascade evaluates end-to-end execution, restores verified context for next-action prediction, and provides target semantics and a local view for visual grounding. By progressively restoring the conditions for success, DashAct measures the minimum support an agent needs to recover rather than scoring isolated skills. Experiments show that current models struggle even as support is added. The cascade outcomes reveal bottlenecks hidden by end-to-end scores and provide actionable guidance for improving GUI agents.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chuhan Zhang, Qi Xie, Ziyue Wang, Jianing Yin, Yunfan Zhou, Dazhen Deng, Yingcai Wu. 2026-09-26. DashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard Analysis. https://arxiv.org/abs/2609.32385

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HAGI++: Head-Assisted Gaze Imputation and Generation

Mobile eye-tracking is crucial for capturing human visual attention in real-world and XR settings, supporting research and human-computer interaction. Yet blinks, pupil-detection errors and lighting changes create missing values that hinder gaze analysis. We present HAGI++, a multi-modal diffusion-based imputation method that, for the first time, leverages integrated head-orientation sensors to exploit the natural correlation between head and eye movements. Using a transformer-based diffusion model, it learns cross-modal dependencies between eye and head data and can additionally incorporate wrist/hand motion when such wearable signals are available. Evaluations on the large-scale Nymeria, Ego-Exo4D and HOT3D datasets show that HAGI++ consistently outperforms traditional interpolation and deep-learning time-series imputation baselines. Statistical analysis confirms that its gaze-velocity distributions closely match real human behaviour, yielding realistic imputations. Even when 100% of gaze data are missing (pure gaze generation), HAGI++ exceeds methods that rely on the visual inputs and the methods rely on full-body motion capture by incorporating wrist motion from commercial wearables. Our approach enables more complete, accurate eye-gaze recordings in real-world contexts, enhancing gaze-based analysis and interaction across many applications. Our code is available at https://git.cai.simtech.uni-stuttgart.de/public-projects/HAGI

cs.HC↗

Deco: Extending Cherished Physical Objects into AI Companion Agents through Dual Embodiment

Physical objects (e.g., plush toys) can transcend materiality to become emotional anchors and provide companionship. However, these bonds remain one-sided because most physical objects cannot reciprocate. AI companions offer responsiveness and personalization, but typically entail building bonds from scratch. We investigated how AI companions might inherit and extend users' existing bonds with physical objects. A formative study (N=9) informed four design principles (Faithful Identity, Calibrated Agency, Ambient Presence, Reciprocal Memory), shaping our Dual-Embodiment Companion Framework. We instantiated it as Deco to create digital embodiments of physical companions. In a within-subjects lab study (N=25), Deco was rated higher than a personalized digital-only companion on six companion-related measures (all p<.01). A subsequent seven-day field deployment (N=17) showed sustained engagement, higher post-deployment well-being (p=.040), and three key relational patterns: digital activities retroactively vitalized physical objects, bond deepening centered on emotional engagement depth, and participants sustained bonds while navigating companions' AI nature. Dual embodiment offers a promising framework for revitalizing physical objects with AI agents.

cs.HC↗

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to process transient audio cues and temporal video dynamics that are tightly coupled with the moment of action. To bridge this gap, we introduce OmniGUI, the first step-level benchmark designed to evaluate GUI agents in omni-modal smartphone environments. OmniGUI provides continuous, interleaved multimodal inputs comprising static images, synchronous audio, and video clips at every action step. The dataset encompasses 709 expert-demonstrated episodes (2,579 action steps) across 29 applications, systematically annotated with objective multimodal dependency levels. Because dedicated omni-modal GUI agent frameworks are currently in their nascent stage, we select foundational omni-modal models capable of natively processing interleaved inputs to serve as agent proxies for our initial baselines. Our empirical evaluation reveals that while current models exhibit competency on visually static tasks, their action prediction performance degrades significantly in environments requiring synchronous temporal and auditory signals. Furthermore, ablation studies isolate specific operational bottlenecks, notably cross-modal interference when processing task-irrelevant environmental noise. The complete dataset, evaluation pipeline, and baseline prompts are provided in the supplementary material. Project page: https://omni-gui.github.io.

cs.HC↗