ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing description- and detection-based assistive technologies do not sufficiently support the multifaceted nature of interactive object search tasks. We present ObjectFinder, an open-vocabulary wearable assistive system for interactive object search by blind people. ObjectFinder allows users to query target objects using flexible wording. Once the target object is detected, it provides egocentric localization information in real-time, including distance and direction. Users can then initiate different branches to gather detailed information based on their intent towards the target object, such as navigating to it or perceiving its surroundings. ObjectFinder is powered by a seamless combination of open-vocabulary models, namely an open-vocabulary object detector and a multimodal large language model. The ObjectFinder design concept and its development were carried out in collaboration with a blind co-designer. To evaluate ObjectFinder, we conducted an exploratory user study with eight blind participants. We compared ObjectFinder to BeMyAI and Google Lookout, popular description- and detection-based assistive applications. Our findings indicate that most participants felt more independent with ObjectFinder and preferred it for object search, as it enhanced scene context gathering and navigation, and allowed for active target identification. Finally, we discuss the implications for future assistive systems to support interactive object search.
Code (0)
등록된 구현이 없습니다.
Tasks
Large Language ModelMultimodal Large Language ModelNavigateObjectobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
Open-Vocabulary Semantic Segmentation with Uncertainty Alignment for Robotic Scene Understanding in Indoor Building Environments
The global rise in the number of people with physical disabilities, in part due to improvements in post-trauma survivorship and longevity, has amplified the demand for advanced assistive technologies to improve mobility …
NavigateOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationScene Classification+4On the Strengths and Weaknesses of Data for Open-set Embodied Assistance
Embodied foundation models are increasingly performant in real-world domains such as robotics or autonomous driving. These models are often deployed in interactive or assistive settings, where it is important that these …
Autonomous DrivingOpenGuide: Assistive Object Retrieval in Indoor Spaces for Individuals with Visual Impairments
Indoor built environments like homes and offices often present complex and cluttered layouts that pose significant challenges for individuals who are blind or visually impaired, especially when performing tasks that invo…
Natural Language UnderstandingObject LocalizationOpenNav: Efficient Open Vocabulary 3D Object Detection for Smart Wheelchair Navigation
Open vocabulary 3D object detection (OV3D) allows precise and extensible object recognition crucial for adapting to diverse environments encountered in assistive robotics. This paper presents OpenNav, a zero-shot 3D obje…
3D Object DetectionNavigateObjectobject-detection+3SIGMA: An Open-Source Interactive System for Mixed-Reality Task Assistance Research
We introduce an open-source system called SIGMA (short for "Situated Interactive Guidance, Monitoring, and Assistance") as a platform for conducting research on task-assistive agents in mixed-reality scenarios. The syste…
Mixed Reality