Search arXiv⌕ Search

arXiv · 2610.05159

Proactive AI: From Turns to Replannable Dialogue Timelines

Abstract

Conventional large-language-model chat interfaces typically follow user-initiated turns: each request elicits a response, after which the system waits for further input. Human asynchronous communication instead unfolds over time through message bursts, delays, silence, resumed topics, and self-initiated contact. Proactive dialogue therefore requires not only deciding when to speak, but also managing pending conversational actions and allowing them to be revised as context evolves. We introduce Proactive AI, a framework for proactive dialogue built around replannable temporal message queues. The framework treats interaction as a continuous event timeline and unsent conversational actions as revisable state. A user event, dialogue-pacemaker event, or scheduled decision may yield silence, one or more immediate messages, or future conversational actions. The system may also observe a user message without an immediate visible response and defer the response decision. Before delivery, every planned message must be reconsidered against the current context; reaching a scheduled time does not itself authorize delivery. We provide a formal semantics that characterizes when future conversational actions may be revised, withdrawn, or executed as context evolves. The framework extends AI participation beyond responses to current requests, enabling self-initiated exchanges without new requests, sustained follow-up over time, and revisions of subsequent actions as context evolves. It provides an executable basis for long-term human-AI interaction in tutoring, scientific collaboration, and everyday companionship.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zijie Yang. 2026-10-04. Proactive AI: From Turns to Replannable Dialogue Timelines. https://arxiv.org/abs/2610.05159

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

User Misconceptions of LLM-Based Conversational Programming Assistants

Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants such as ChatGPT particularly accessible to novice programmers. However, varied tool capabilities and inconsistent availability of extensions (e.g., web search, code execution, retrieval-augmented generation) create opportunities for user misconceptions that may lead to over-reliance, unproductive practices, or insufficient quality control. We characterize the misconceptions that users of conversational LLM-based assistants may hold in programming contexts. We screened 11,429 Python-related conversations from the openly available WildChat dataset with a validated LLM annotation pipeline, then hand-annotated the 754 candidate conversations it flagged. Of these, 450 contain a prompt consistent with at least one of eight potential misconceptions: misplaced expectations about capabilities such as web access, code execution, non-text outputs, and session memory. We also characterize how the assistant responds when a prompt presupposes a capability it lacks: responses range from explicit refusal through qualified answers to fabricated compliance, and explicit refusals appear in only a minority of labeled conversations. Among the most frequent misconceptions, explicit refusals are rarest where compliance is easiest to fabricate. Our findings reinforce the need for LLM-based tools to communicate their capabilities to users through channels other than the conversation itself.

cs.HC↗

EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution

High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often lack precise control, while code-based synthesis limits intuitive flexibility. To bridge this gap, we introduce EvoDiagram, an agentic framework that generates object-level editable diagrams via an intermediate canvas schema. EvoDiagram employs a coordinated multi-agent system to decouple semantic intent from rendering logic, resolving conflicts across heterogeneous design layers. Additionally, we propose a design knowledge evolution mechanism that distills execution traces into a hierarchical memory of domain guidelines, enabling agents to retrieve context-aware expertise adaptively. We further release CanvasBench, a benchmark consisting of both data and metrics for canvas-based diagramming. Extensive experiments demonstrate that EvoDiagram exhibits excellent performance and balance against baselines in generating editable, structurally consistent, and aesthetically coherent diagrams. Our code is available at https://github.com/AuraX-AI/EvoDiagram.

cs.HC↗

Smart Navigation for Visual Prostheses in Virtual Reality: An End-to-End Framework for Priority-Based Scene Translation and Path Guidance

Visual prosthetics provide a promising direction for partial restoration of functional vision for people with total retinal blindness. However, existing systems face significant challenges in translating complex visual scenes into meaningful perceptions due to limited spatial resolution, leading to difficulties in scene understanding. Furthermore, existing solutions don't adequately account for user requirements and concerns, and this creates a significant gap between user expectations and the developed solutions. To address these gaps, we conducted interviews with 10 blind human subjects. These interviews essentially indicated that the main key challenge for these people is outdoor navigation. In this paper, we present an end-to-end smart navigation system for visual prostheses users. Our approach employs a three-stage framework. First, an object detector is implemented to identify and localize points of interest for navigation tasks in pedestrian environments, and simultaneously generate the safest paths available to the users. Second, the detected objects are abstracted into simple geometric shapes suitable for low-spatial-resolution vision. The detected objects are filtered based on a multi-criteria priority scoring function. Finally, this information is encoded into optimized stimulation parameters, which are fed into the visual prosthesis implants to generate enhanced phosphene representations for obstacle avoidance and path planning. This whole system is validated with sighted participants using virtual reality to simulate outdoor navigation. Our smart navigation system improves user independence while taking into account the limitations of visual prosthetics.

cs.HC↗