paper-with-me

홈 › Papers

Guiding Multimodal Large Language Models with Blind and Low Vision People Visual Questions for Proactive Visual Interpretations

2025-10-02 · Ricardo Gonzalez Penuela, Felipe Arias-Russi, Victor Capriles arxiv

Multimodal large language models (MLLMs) have been integrated into visual interpretation applications to support Blind and Low Vision (BLV) users because of their accuracy and ability to provide rich, human-like interpretations. However, these applications often default to comprehensive, lengthy descriptions regardless of context. This leads to inefficient exchanges, as users must go through irrelevant details rather than receiving the specific information they are likely to seek. To deliver more contextually-relevant information, we developed a system that draws on historical BLV users questions. When given an image, our system identifies similar past visual contexts from the VizWiz-LF dataset and uses the associated questions to guide the MLLM generate descriptions more relevant to BLV users. An evaluation with three human labelers who revised 92 context-aware and context-free descriptions showed that context-aware descriptions anticipated and answered users' questions in 76.1% of cases (70 out of 92) and were preferred in 54.4% of comparisons (50 out of 92). Our paper reviews, and data analysis are publicly available in a Github repository at https://github.com/rgonzalezp/guiding-multimodal-large-language-models-with-blind-and-low-vision-people-visual-questions .

📄 PDF Abstract BibTeX arXiv:2510.01576

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Implementing blind navigation through multi-modal sensing and gait guidance

2025-06-24 · Feifan Yan, Tianle Zeng, Meixi He

By the year 2023, the global population of individuals with impaired vision has surpassed 220 million. People with impaired vision will find it difficult while finding path or avoiding obstacles, and must ask for auxilia…

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

2026-04-22 · Karan Goyal arxiv

The rapid proliferation of Vision-Language Models (VLMs) is often framed as enabling unified multimodal knowledge discovery but rests on an under-examined assumption: that current VLMs faithfully synthesise multimodal da…

Multimodal Reasoning

Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model

2024-05-28 · Haogeng Liu, Quanzeng You, Xiaotian Han, Yongfei Liu 외

In the realm of Multimodal Large Language Models (MLLMs), vision-language connector plays a crucial role to link the pre-trained vision encoders with Large Language Models (LLMs). Despite its importance, the vision-langu…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

MultiRetNet: A Multimodal Vision Model and Deferral System for Staging Diabetic Retinopathy

2025-07-19 · Jeannie She, Katie Spivakovsky arxiv

Diabetic retinopathy (DR) is a leading cause of preventable blindness, affecting over 100 million people worldwide. In the United States, individuals from lower-income communities face a higher risk of progressing to adv…

Contrastive Learning

HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

2024-06-27 · Junying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao 외

The rapid development of multimodal large language models (MLLMs), such as GPT-4V, has led to significant advancements. However, these models still face challenges in medical multimodal capabilities due to limitations in…

Visual Question Answering (VQA)