NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment
People with visual impairments often face significant challenges in locating and retrieving objects in their surroundings. Existing assistive technologies present a trade-off: systems that offer precise guidance typically require pre-scanning or support only fixed object categories, while those with open-world object recognition lack spatial feedback for reaching the object. To address this gap, we introduce 'NaviSense', a mobile assistive system that combines conversational AI, vision-language models, augmented reality (AR), and LiDAR to support open-world object detection with real-time audio-haptic guidance. Users specify objects via natural language and receive continuous spatial feedback to navigate toward the target without needing prior setup. Designed with insights from a formative study and evaluated with 12 blind and low-vision participants, NaviSense significantly reduced object retrieval time and was preferred over existing tools, demonstrating the value of integrating open-world perception with precise, accessible guidance.
Code (0)
등록된 구현이 없습니다.
Tasks
Object RecognitionObject DetectionSimilar Papers 제목 키워드 기반
Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications
Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong performance on tasks like image caption…
Visual Question AnsweringImage CaptioningObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. Howeve…
Large Language ModelMultimodal Large Language ModelNavigateObject+2Model Reconciliation through Explainability and Collaborative Recovery in Assistive Robotics
Whenever humans and robots work together, it is essential that unexpected robot behavior can be explained to the user. Especially in applications such as shared control the user and the robot must share the same model of…
Semantic-Aware Environment Perception for Mobile Human-Robot Interaction
Current technological advances open up new opportunities for bringing human-machine interaction to a new level of human-centered cooperation. In this context, a key issue is the semantic understanding of the environment …
CommunicoTool Advance, un prototype d'application d'aide \`a la communication (CommunicoTool Advance: an assistive communication app prototype)
CommunicoTool Advance est un prototype d{'}application mobile d{'}aide {\`a} la communication destin{\'e}e {\`a} des personnes qui pr{\'e}sentent des troubles moteurs et des troubles de la parole.