paper-with-me

홈 › Papers

How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People

2026-02-13 · Ricardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y Lin, Ruiying Hu, Shiri Azenkot arxiv

Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer conversational assistance, where users can ask questions to obtain goal-relevant details. However, evidence about their performance in the real-world and implications for BLV people's daily lives remains limited. To address this, we conducted a two-week diary study, where we captured 20 BLV participants' use of an MLLM-enabled visual interpretation application. Although participants rated the visual interpretations of the application as "trustworthy" (mean=3.76 out of 5, max=extremely trustworthy) and "somewhat satisfying" (mean=4.13 out of 5, max=very satisfying), the AI often produced incorrect answers (22.2%) or abstained (10.8%) from responding to users' requests. Our findings show that while MLLMs can improve visual interpretations' descriptive accuracy, supporting everyday use also depends on the "visual assistant" skill: behaviors for providing goal-directed, reliable assistance. We conclude by proposing the "visual assistant" skill and guidelines to help MLLM-enabled visual interpretation applications better support BLV people's access to visual information.

📄 PDF Abstract BibTeX arXiv:2602.13469

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Don't Look Only Once: Towards Multimodal Interactive Reasoning with Selective Visual Revisitation

2025-05-24 · Jiwan Chung, Junhyeok Kim, Siyeol Kim, Jaeyoung Lee 외

We present v1, a lightweight extension to Multimodal Large Language Models (MLLMs) that enables selective visual revisitation during inference. While current MLLMs typically consume visual input only once and reason pure…

Mathematical ReasoningMultimodal ReasoningVisual Grounding

A Dialogue-Based Framework for Correcting Multimodal Errors in AI-Assisted STEM Education

2026-05-05 · Akshay Syal, Lawrence Swaminathan Xavier Prince, Evin Gultepe, Nik Bear Brown 외 arxiv

Large Language Models (LLMs) are democratizing access to personalized tutoring; however, their effectiveness is hindered by challenges in processing multimodal content, which limits AI's potential to provide equitable, h…

Investigating Multimodal Large Language Models to Support Usability Evaluation

2025-08-22 · Sebastian Lubos, Alexander Felfernig, Damian Garber, Gerhard Leitner 외 arxiv

Usability evaluation is an essential method to support the design of effective and intuitive user interfaces (UIs). However, it commonly relies on resource-intensive, expert-driven methods, which limit its accessibility,…

MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places

2026-08-28 · Jason Armitage, Ioannis Tsochantaridis, Linda Mazzone, Chuqiao Yan 외 arxiv

We introduce MAP, the first benchmark to evaluate multimodal AI systems as assistants for users with accessibility requirements when planning visits to places in the real world. In our evaluation, systems are presented w…

Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models

2024-06-30 · Weihong Zhong, Xiaocheng Feng, Liang Zhao, Qiming Li 외

Though advanced in understanding visual information with human languages, Large Vision-Language Models (LVLMs) still suffer from multimodal hallucinations. A natural concern is that during multimodal interaction, the gen…

Hallucinationmultimodal interaction