paper-with-me

홈 › Papers

Can Multimodal Large Language Models Understand OCT?

2026-07-18 · Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li, Zihan Nie, Kailin Jiang, Yuntao Du, Weiye Song hf

Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.

📄 PDF Abstract BibTeX arXiv:2607.16609

Code (1)

arxivsub/arXivSub_daily_arxiv ★ 3

Tasks

Visual Question AnsweringDomain Adaptation

Similar Papers 제목 키워드 기반

OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects

2024-10-02 · Wenmo Qiu, Xinhan Di

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satisfactory results in describing occluded o…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

How2: A Large-scale Dataset for Multimodal Language Understanding

2018-11-01 · Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott 외

In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations. We also present integrated sequence-to-sequence baselines for machine tra…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2

Multimodal Large Language Models: A Survey

2023-11-22 · Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan 외

The exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models excel in text-based tasks, they often s…

Survey

LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

2024-07-27 · Ruiyi Zhang, Yufan Zhou, Jian Chen, Jiuxiang Gu 외

Large multimodal language models have demonstrated impressive capabilities in understanding and manipulating images. However, many of these models struggle with comprehending intensive textual contents embedded within th…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

2025-06-26 · Qize Yang, Shimin Yao, Weixuan Chen, Shenghao Fu 외

With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands detailed and thoughtful reasoning. In rec…

Large Language ModelMultimodal ReasoningReinforcement Learning (RL)