paper-with-me

홈 › Papers

NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment

2025-09-23 · Ajay Narayanan Sridhar, Fuli Qiao, Nelson Daniel Troncoso Aldas, Yanpei Shi, Mehrdad Mahdavi, Laurent Itti, Vijaykrishnan Narayanan arxiv

People with visual impairments often face significant challenges in locating and retrieving objects in their surroundings. Existing assistive technologies present a trade-off: systems that offer precise guidance typically require pre-scanning or support only fixed object categories, while those with open-world object recognition lack spatial feedback for reaching the object. To address this gap, we introduce 'NaviSense', a mobile assistive system that combines conversational AI, vision-language models, augmented reality (AR), and LiDAR to support open-world object detection with real-time audio-haptic guidance. Users specify objects via natural language and receive continuous spatial feedback to navigate toward the target without needing prior setup. Designed with insights from a formative study and evaluated with 12 blind and low-vision participants, NaviSense significantly reduced object retrieval time and was preferred over existing tools, demonstrating the value of integrating open-world perception with precise, accessible guidance.

📄 PDF Abstract BibTeX arXiv:2509.18672

Code (0)

등록된 구현이 없습니다.

Tasks

Object RecognitionObject Detection

Similar Papers 제목 키워드 기반

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

2026-06-23 · Shayon Dasgupta, Avijit Dasgupta, C. V. Jawahar arxiv

Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong performance on tasks like image caption…

Visual Question AnsweringImage Captioning

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

2024-12-04 · Ruiping Liu, Jiaming Zhang, Angela Schön, Karin Müller 외

Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. Howeve…

Large Language ModelMultimodal Large Language ModelNavigateObject+2

Model Reconciliation through Explainability and Collaborative Recovery in Assistive Robotics

2026-01-10 · Britt Besch, Tai Mai, Jeremias Thun, Markus Huff 외 arxiv

Whenever humans and robots work together, it is essential that unexpected robot behavior can be explained to the user. Especially in applications such as shared control the user and the robot must share the same model of…

Semantic-Aware Environment Perception for Mobile Human-Robot Interaction

2022-11-07 · Thorsten Hempel, Marc-André Fiedler, Aly Khalifa, Ayoub Al-Hamadi 외

Current technological advances open up new opportunities for bringing human-machine interaction to a new level of human-centered cooperation. In this context, a key issue is the semantic understanding of the environment …

CommunicoTool Advance, un prototype d'application d'aide \`a la communication (CommunicoTool Advance: an assistive communication app prototype)

2016-07-01 · JEPTALNRECITAL 2016 7 · Charlotte Roze

CommunicoTool Advance est un prototype d{'}application mobile d{'}aide {\`a} la communication destin{\'e}e {\`a} des personnes qui pr{\'e}sentent des troubles moteurs et des troubles de la parole.