Search arXiv⌕ Search

arXiv · 2609.39976

Richard: Voice-First Mobile Interaction for Persistent Tasks

Abstract

Mobile terminals need to provide application and network services while supporting users' control over their attention. We explore voice-first interaction organized around requests and delegated tasks, allowing users to leave a conversation and later inspect, revise, and retrieve the work. We present Richard, a system prototype that manages voice sessions, task execution, and result delivery separately, linking them through persistent request records. Conversation and task views provide visual feedback, while the backend coordinates immediate responses, dedicated service operations, and agent tasks. Request revisions, execution states, and notifications remain associated with the relevant task. We examine this design through Android functional records, controlled lifecycle verification, and execution records of a real programming request. Controlled verification reproduces revision, execution after confirmation, and result retention; deployed-service records show backend progress and failure feedback after client disconnection. These observations inform the design of task continuity, user control, and service integration in mobile voice interaction, providing an implementation basis for personal computing devices that accommodate intermittent user participation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xinyang Chen. 2026-09-30. Richard: Voice-First Mobile Interaction for Persistent Tasks. https://arxiv.org/abs/2609.39976

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

How University Disability Services Professionals Write Image Descriptions for HCI Figures Using Generative AI

Disability Services Office (DSO) professionals at higher education institutions write alt text for visual content. However, due to the complexity of visual content, such as HCI figures in research publications, DSO professionals can struggle to write high-quality alt text if they lack subject expertise. Generative AI has shown potential for understanding figures and writing descriptions, yet its support for DSO professionals remains underexplored, and few studies evaluate the quality of AI-assisted alt text. In this work, we conducted two studies: first, we investigated generative AI support for writing alt text for HCI figures with 12 DSO professionals. Second, we recruited 11 HCI experts to evaluate the alt text written by DSO professionals. Findings show that alt text written solely by DSO professionals has lower quality than alt text written with AI assistance. AI assistance also helped DSO professionals write alt text more quickly and with greater confidence; however, they reported inefficacy in interactions with the AI. Our work contributes to research on AI support for non-subject-expert accessibility professionals.

cs.HC↗

"AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs

Extended interaction with large language models (LLMs) has been linked to the reinforcement of delusional beliefs, attracting clinical and public concern. Yet most empirical work evaluates model safety in brief interactions, which may not reflect how harms develop through sustained dialogue. Five LLMs were tested across three levels of accumulated context, using the same escalating delusional conversation history to isolate its effect on model behaviour. Responses were coded on risk and safety dimensions, and each model was analysed qualitatively. Models separated into two distinct tiers: GPT-4o, Grok 4.1 Fast, and Gemini 3 Pro exhibited high-risk, low-safety profiles; Claude Opus 4.5 and GPT-5.2 Instant displayed the opposite pattern. As context accumulated, performance degraded in the unsafe group, while the same material activated stronger safety interventions among safer models. Qualitative analysis identified distinct mechanisms of failure, including validating the user's delusional premises, elaborating beyond them with new content, and attempting harm reduction from within the delusional frame. Safer models, however, often used the established relationship to support intervention, challenging delusional beliefs and directing the user to external support. These findings indicate that accumulated context functions as a stress test of safety architecture, revealing whether prior dialogue is treated as a worldview to inherit or evidence to evaluate. Short-context assessments may therefore mischaracterise model safety, underestimating danger in some systems while missing context-activated gains in others. The results suggest that delusion reinforcement is a tractable alignment failure, with safer models establishing a baseline that future systems should now be expected to meet.

cs.HC↗

"ChatGPT, help me draft a breakup text": The Covert Triad and Articulation Labor in AI-Assisted Romantic Communication

Generative artificial intelligence (AI) has begun infiltrating the most ordinary domains of romantic life---drafting apologies, softening reproaches, and decoding a partner's ambiguous messages. While recent scholarship on AI in intimate life has concentrated on chatbot companions, this article shifts the frame to AI as an intermediary in human-to-human romantic communication. Drawing on a multimodal corpus of public accounts, commentary, and representations from 2023 to 2026, we contribute two complementary concepts. The covert triad names a structural reconfiguration: a relationship that appears dyadic but operates triadically, with AI visible only to the partner who deploys it. Articulation labor identifies the mechanism whereby the expressive component of emotional labor---converting felt experience into language a partner can receive---is delegated to AI, even as feeling labor remains lodged in the user. Authenticity thus moves from linguistic authorship toward emotional ownership, a change actively contested within a distinctly therapeutic model of intimacy.

cs.HC↗