Search arXiv⌕ Search

arXiv · 2610.05772

Smart Navigation for Visual Prostheses in Virtual Reality: An End-to-End Framework for Priority-Based Scene Translation and Path Guidance

Abstract

Visual prosthetics provide a promising direction for partial restoration of functional vision for people with total retinal blindness. However, existing systems face significant challenges in translating complex visual scenes into meaningful perceptions due to limited spatial resolution, leading to difficulties in scene understanding. Furthermore, existing solutions don't adequately account for user requirements and concerns, and this creates a significant gap between user expectations and the developed solutions. To address these gaps, we conducted interviews with 10 blind human subjects. These interviews essentially indicated that the main key challenge for these people is outdoor navigation. In this paper, we present an end-to-end smart navigation system for visual prostheses users. Our approach employs a three-stage framework. First, an object detector is implemented to identify and localize points of interest for navigation tasks in pedestrian environments, and simultaneously generate the safest paths available to the users. Second, the detected objects are abstracted into simple geometric shapes suitable for low-spatial-resolution vision. The detected objects are filtered based on a multi-criteria priority scoring function. Finally, this information is encoded into optimized stimulation parameters, which are fed into the visual prosthesis implants to generate enhanced phosphene representations for obstacle avoidance and path planning. This whole system is validated with sighted participants using virtual reality to simulate outdoor navigation. Our smart navigation system improves user independence while taking into account the limitations of visual prosthetics.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mohamed H. Abdellatif, Fatma S. Elsharkawy, Nouran H. Qassem, Talal M. Emara, Nada K. Kotb, Muhammad Rushdi, Reham H. Elnabawy. 2026-10-05. Smart Navigation for Visual Prostheses in Virtual Reality: An End-to-End Framework for Priority-Based Scene Translation and Path Guidance. https://arxiv.org/abs/2610.05772

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

User Misconceptions of LLM-Based Conversational Programming Assistants

Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants such as ChatGPT particularly accessible to novice programmers. However, varied tool capabilities and inconsistent availability of extensions (e.g., web search, code execution, retrieval-augmented generation) create opportunities for user misconceptions that may lead to over-reliance, unproductive practices, or insufficient quality control. We characterize the misconceptions that users of conversational LLM-based assistants may hold in programming contexts. We screened 11,429 Python-related conversations from the openly available WildChat dataset with a validated LLM annotation pipeline, then hand-annotated the 754 candidate conversations it flagged. Of these, 450 contain a prompt consistent with at least one of eight potential misconceptions: misplaced expectations about capabilities such as web access, code execution, non-text outputs, and session memory. We also characterize how the assistant responds when a prompt presupposes a capability it lacks: responses range from explicit refusal through qualified answers to fabricated compliance, and explicit refusals appear in only a minority of labeled conversations. Among the most frequent misconceptions, explicit refusals are rarest where compliance is easiest to fabricate. Our findings reinforce the need for LLM-based tools to communicate their capabilities to users through channels other than the conversation itself.

cs.HC↗

EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution

High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often lack precise control, while code-based synthesis limits intuitive flexibility. To bridge this gap, we introduce EvoDiagram, an agentic framework that generates object-level editable diagrams via an intermediate canvas schema. EvoDiagram employs a coordinated multi-agent system to decouple semantic intent from rendering logic, resolving conflicts across heterogeneous design layers. Additionally, we propose a design knowledge evolution mechanism that distills execution traces into a hierarchical memory of domain guidelines, enabling agents to retrieve context-aware expertise adaptively. We further release CanvasBench, a benchmark consisting of both data and metrics for canvas-based diagramming. Extensive experiments demonstrate that EvoDiagram exhibits excellent performance and balance against baselines in generating editable, structurally consistent, and aesthetically coherent diagrams. Our code is available at https://github.com/AuraX-AI/EvoDiagram.

cs.HC↗

What Did the AI Take On? Characterizing Cognitive Delegation in LLM Reasoning

Large language models (LLMs) often perform intermediate cognitive work while carrying out users' requests, yet it remains unclear which parts users intended to delegate and how they wanted to remain involved. This matters because consequential choices may go unnoticed, limiting users' ability to steer the process, while reviewing every step would make delegation burdensome. We examined this with 24 LLM users across three knowledge-work tasks, collecting 992 retrospective annotations of reasoning steps. From this, we developed taxonomies of LLM cognitive work, delegation enactment, and desired delegation protocols at the reasoning-step level. Our analysis revealed that participants viewed about half of all steps (48.6%) as AI-initiated, meaning the AI took on work they had not requested. Desired involvement varied with cognitive work and delegation enactment, even when contributions matched participants' intent. We propose design implications and sketches for supporting more deliberate cognitive delegation through flexible protocols and inspectable, revisable AI-initiated decisions.

cs.HC↗