paper-with-me

홈 › Papers

VIALM: A Survey and Benchmark of Visually Impaired Assistance with Large Models

2024-01-29 · Yi Zhao, Yilin Zhang, Rong Xiang, Jing Li, Hillming Li

Visually Impaired Assistance (VIA) aims to automatically help the visually impaired (VI) handle daily activities. The advancement of VIA primarily depends on developments in Computer Vision (CV) and Natural Language Processing (NLP), both of which exhibit cutting-edge paradigms with large models (LMs). Furthermore, LMs have shown exceptional multimodal abilities to tackle challenging physically-grounded tasks such as embodied robots. To investigate the potential and limitations of state-of-the-art (SOTA) LMs' capabilities in VIA applications, we present an extensive study for the task of VIA with LMs (VIALM). In this task, given an image illustrating the physical environments and a linguistic request from a VI user, VIALM aims to output step-by-step guidance to assist the VI user in fulfilling the request grounded in the environment. The study consists of a survey reviewing recent LM research and benchmark experiments examining selected LMs' capabilities in VIA. The results indicate that while LMs can potentially benefit VIA, their output cannot be well environment-grounded (i.e., 25.7% GPT-4's responses) and lacks fine-grained guidance (i.e., 32.1% GPT-4's responses).

📄 PDF Abstract BibTeX arXiv:2402.01735

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

2026-05-29 · Yi Zhao, Siqi Wang, Zhe Hu, Yushi Li 외 arxiv

AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer a promising alternative, although it has mostly been studied in gen…

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

2026-07-16 · Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang 외 arxiv

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general visio…

Visual Question Answering

Evaluation Framework for Computer Vision-Based Guidance of the Visually Impaired

2022-09-20 · Krešimir Romić, Irena Galić, Marija Habijan, Hrvoje Leventić

Visually impaired persons have significant problems in their everyday movement. Therefore, some of our previous work involves computer vision in developing assistance systems for guiding the visually impaired in critical…

Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users

2025-03-28 · Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas 외

This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges use…

Object RecognitionReading ComprehensionScene Understanding

Adaptive Object Detection for Indoor Navigation Assistance: A Performance Evaluation of Real-Time Algorithms

2025-01-30 · Abhinav Pratap, Sushant Kumar, Suchinton Chakravarty

This study addresses the need for accurate and efficient object detection in assistive technologies for visually impaired individuals. We evaluate four real-time object detection algorithms YOLO, SSD, Faster R-CNN, and M…

Objectobject-detectionObject DetectionReal-Time Object Detection