paper-with-me

홈 › Papers

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

2024-12-04 · Ruiping Liu, Jiaming Zhang, Angela Schön, Karin Müller, Junwei Zheng, Kailun Yang, Anhong Guo, Kathrin Gerling, Rainer Stiefelhagen

Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing description- and detection-based assistive technologies do not sufficiently support the multifaceted nature of interactive object search tasks. We present ObjectFinder, an open-vocabulary wearable assistive system for interactive object search by blind people. ObjectFinder allows users to query target objects using flexible wording. Once the target object is detected, it provides egocentric localization information in real-time, including distance and direction. Users can then initiate different branches to gather detailed information based on their intent towards the target object, such as navigating to it or perceiving its surroundings. ObjectFinder is powered by a seamless combination of open-vocabulary models, namely an open-vocabulary object detector and a multimodal large language model. The ObjectFinder design concept and its development were carried out in collaboration with a blind co-designer. To evaluate ObjectFinder, we conducted an exploratory user study with eight blind participants. We compared ObjectFinder to BeMyAI and Google Lookout, popular description- and detection-based assistive applications. Our findings indicate that most participants felt more independent with ObjectFinder and preferred it for object search, as it enhanced scene context gathering and navigation, and allowed for active target identification. Finally, we discuss the implications for future assistive systems to support interactive object search.

📄 PDF Abstract BibTeX arXiv:2412.03118

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language ModelMultimodal Large Language ModelNavigateObjectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Open-Vocabulary Semantic Segmentation with Uncertainty Alignment for Robotic Scene Understanding in Indoor Building Environments

2025-03-29 · Yifan Xu, Vineet Kamat, Carol Menassa

The global rise in the number of people with physical disabilities, in part due to improvements in post-trauma survivorship and longevity, has amplified the demand for advanced assistive technologies to improve mobility …

NavigateOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationScene Classification+4

On the Strengths and Weaknesses of Data for Open-set Embodied Assistance

2026-03-05 · Pradyumna Tambwekar, Andrew Silva, Deepak Gopinath, Jonathan DeCastro 외 arxiv

Embodied foundation models are increasingly performant in real-world domains such as robotics or autonomous driving. These models are often deployed in interactive or assistive settings, where it is important that these …

Autonomous Driving

OpenGuide: Assistive Object Retrieval in Indoor Spaces for Individuals with Visual Impairments

2025-09-02 · Yifan Xu, Qianwei Wang, Vineet Kamat, Carol Menassa arxiv

Indoor built environments like homes and offices often present complex and cluttered layouts that pose significant challenges for individuals who are blind or visually impaired, especially when performing tasks that invo…

Natural Language UnderstandingObject Localization

OpenNav: Efficient Open Vocabulary 3D Object Detection for Smart Wheelchair Navigation

2024-08-25 · Muhammad Rameez Ur Rahman, Piero Simonetto, Anna Polato, Francesco Pasti 외

Open vocabulary 3D object detection (OV3D) allows precise and extensible object recognition crucial for adapting to diverse environments encountered in assistive robotics. This paper presents OpenNav, a zero-shot 3D obje…

3D Object DetectionNavigateObjectobject-detection+3

SIGMA: An Open-Source Interactive System for Mixed-Reality Task Assistance Research

2024-05-16 · Dan Bohus, Sean Andrist, Nick Saw, Ann Paradiso 외

We introduce an open-source system called SIGMA (short for "Situated Interactive Guidance, Monitoring, and Assistance") as a platform for conducting research on task-assistive agents in mixed-reality scenarios. The syste…

Mixed Reality