arXiv · 2610.05047
Beyond Task Completion: Measuring Interaction Cost in Terminal User Interfaces
Abstract
Large language models (LLMs) are increasingly used through terminal user interfaces (TUIs), yet task completion alone does not capture how difficult an interface is to understand and operate. Existing human assessments and LLM-generated ratings or reports do not provide repeatable measurements of interaction effort grounded in verified task execution. We propose Agent-as-a-User, an evaluation paradigm that places an LLM agent in the user role. Agent-KLM separates agent-side interpretation and operation costs and relates them structurally to human interaction. TUINaut operationalizes the paradigm by recording interaction trajectories and verifying outcomes with actor-independent oracles. Using TUINaut-Bench, we study 84 tasks across 15 real-world TUIs with six human participants and five observation-evaluator configurations, and evaluate 45 TUIs generated by three leading stacks. Similar aggregate success rates mask differences in which tasks humans and agents complete, whereas interpretation and operation costs follow correlated task rankings, especially for operation. This pattern persists across configurations. Generated TUIs can implement correct functionality while still requiring usability improvements. These results position Agent-as-a-User as a repeatable, execution-grounded paradigm for measuring task-based usability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ruida Hu, Yuanhao Wang, Chao Peng, Yakun Zhang, Cuiyun Gao. 2026-10-04. Beyond Task Completion: Measuring Interaction Cost in Terminal User Interfaces. https://arxiv.org/abs/2610.05047
Cite the original work for its findings. Save a collection to share your selection of sources.