paper-with-me

홈 › Papers

Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users

2025-03-28 · Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard

This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such technologies. Despite a high adoption rate of these models, our findings highlight concerns related to contextual understanding, cultural sensitivity, and complex scene understanding, particularly for individuals who may rely solely on them for visual interpretation. Informed by these results, we collate five user-centred tasks with image and video inputs, including a novel task on Optical Braille Recognition. Our systematic evaluation of twelve MLLMs reveals that further advancements are necessary to overcome limitations related to cultural context, multilingual support, Braille reading comprehension, assistive object recognition, and hallucinations. This work provides critical insights into the future direction of multimodal AI for accessibility, underscoring the need for more inclusive, robust, and trustworthy visual assistance technologies.

📄 PDF Abstract BibTeX arXiv:2503.22610

Code (0)

등록된 구현이 없습니다.

Tasks

Object RecognitionReading ComprehensionScene Understanding

Similar Papers 제목 키워드 기반

Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models

2026-08-27 · Basel Mousi, Fahim Dalvi, Shammur Chowdhury, Firoj Alam 외 arxiv

Multimodal foundation models are increasingly used in speech-first assistants that must interpret spoken queries and produce visually grounded decisions. Yet it remains unclear whether semantically equivalent queries yie…

MMIU: Dataset for Visual Intent Understanding in Multimodal Assistants

2021-10-13 · Alkesh Patel, Joel Ruben Antony Moniz, Roman Nguyen, Nick Tzou 외

In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. Current digital assistants take spoken in…

intent-classificationIntent ClassificationQuestion AnsweringQuestion Generation+3

MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

2025-07-14 · Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi 외 arxiv

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech an…

Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs

2026-01-09 · Sandeep Mishra, Devichand Budagam, Anubhab Mandal, Bishal Santra 외 arxiv

Real-time multimodal auto-completion is essential for digital assistants, chatbots, design tools, and healthcare consultations, where user inputs rely on shared visual context. We introduce Multimodal Auto-Completion (MA…

Generating Natural Questions from Images for Multimodal Assistants

2020-11-17 · Alkesh Patel, Akanksha Bindal, Hadas Kotek, Christopher Klein 외

Generating natural, diverse, and meaningful questions from images is an essential task for multimodal assistants as it confirms whether they have understood the object and scene in the images properly. The research in vi…

Common Sense ReasoningNatural QuestionsQuestion AnsweringQuestion Generation+3