arXiv · 2609.22971
Automatic multimodal UX improvement recommendations from LLM agent user simulations
Abstract
Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manual review to extract actionable insights. We formalise UX improvement recommendation from simulation data as a structured natural language generation and ranking problem, and establish an evaluation protocol using expert annotation and LLM-as-a-Judge. We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data. We evaluate AMUSER on commercial websites and show that its recommendations substantially outperform those from text-only simulation (NDCG@3 = 0.758 versus 0.359) at an 89% lower simulation cost. Our results suggest an asymmetric role of multimodality: visual access during simulation improves recommendations through richer traces, while providing visual inputs during recommendation generation can modestly degrade quality. We also discuss practical deployment lessons from applying AMUSER to commercial websites.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Anu Chowdhury, Bin Wu, Hossein A. Rahmani, Emine Yilmaz. 2026-09-19. Automatic multimodal UX improvement recommendations from LLM agent user simulations. https://arxiv.org/abs/2609.22971
Cite the original work for its findings. Save a collection to share your selection of sources.