Search arXiv⌕ Search

arXiv subjects

Bhanuja Ainary

Publications and source records attributed to Bhanuja Ainary.

2 recordsLinked to original sources

Audo-Sight: AI-driven Ambient Perception Across Edge-Cloud for Blind and Low Vision Users

Despite advances in assistive technologies, Blind and Low-Vision (BLV) individuals continue to face challenges in understanding their surroundings. Delivering concise, useful, and timely scene descriptions for ambient perception remains a long-standing problem in accessibility. Existing solutions often fail to identify user expectations for real-time and accessible responses. Moreover, for a given task, they either rely on cloud offloading, which imposes a significant delay, or edge-based AI, which often sacrifices accuracy. To address this, we present Audo-Sight, an AI-driven assistive system that spans across Edge-Cloud continuum and enables BLV individuals to perceive their surroundings through voice-based conversation. Audo-Sight provides low-latency, accurate, and human-friendly responses through a novel mechanism that seamlessly fuses Edge and Cloud responses. The system also addresses challenges in catering to BLV users through response editing informed by BLV needs. Audo-Sight orchestrates a set of AI models based on user query contextual analysis to infer intent and adjust for a variety of situations. In urgent cases where users require fast responses, Audo-Sight leverages parallel Edge and Cloud pipelines and seamlessly combines responses through its Response Fusion Engine. Systematic evaluation shows that Audo-Sight delivers speech output around 80% faster for urgent tasks and generates complete responses approximately 50% faster across all tasks compared to a commercial cloud-based solution---highlighting the need for customized AI-based solutions. Human evaluation of Audo-Sight shows that it is the preferred choice over GPT-5 for 62% of BLV participants with another 23% stating both perform comparably. Speed and interruption evaluations demonstrate that in most situations, the system can seamlessly respond at a rapid pace to keep up with BLV expectations.

cs.DC↗

Audo-Sight: Enabling Ambient Interaction For Blind And Visually Impaired Individuals

Visually impaired people face significant challenges when attempting to interact with and understand complex environments, and traditional assistive technologies often struggle to quickly provide necessary contextual understanding and interactive intelligence. This thesis presents Audo-Sight, a state-of-the-art assistive system that seamlessly integrates Multimodal Large Language Models (MLLMs) to provide expedient, context-aware interactions for Blind and Visually Impaired (BVI) individuals. The system operates in two different modalities: personalized interaction through user identification and public access in common spaces like museums and shopping malls. In tailored environments, the system adjusts its output to conform to the preferences of individual users, thus enhancing accessibility through a user-aware form of interaction. In shared environments, Audo-Sight employs a shared architecture that adapts to its current user with no manual reconfiguration required. To facilitate appropriate interactions with the LLM, the public Audo-Sight solution includes an Age-Range Determiner and Safe Query Filter. Additionally, the system ensures that responses are respectful to BVI users through NeMo Guardrails. By utilizing multimodal reasoning, BVI-cognizant response editing, and safeguarding features, this work represents a major leap in AI-driven accessibility technology capable of increasing autonomy, safety, and interaction for people with visual impairments in social settings. Finally, we present the integration of Audo-Sight and SmartSight, which enables enhanced situational awareness for BVI individuals. This integration takes advantage of the real-time visual analysis of SmartSight, combined with the extensive reasoning and interactive capabilities of Audo-Sight, and goes beyond object identification to provide context-driven, voice-controlled assistance in dynamic environments.

cs.HC↗