paper-with-me

홈 › Papers

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

2025-08-27 · Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose arxiv

Artificial Intelligence models have demonstrated significant success in diagnosing skin diseases, including cancer, showing the potential to assist clinicians in their analysis. However, the interpretability of model predictions must be significantly improved before they can be used in practice. To this end, we explore the combination of two promising approaches: Multimodal Large Language Models (MLLMs) and quantitative attribute usage. MLLMs offer a potential avenue for increased interpretability, providing reasoning for diagnosis in natural language through an interactive format. Separately, a number of quantitative attributes that are related to lesion appearance (e.g., lesion area) have recently been found predictive of malignancy with high accuracy. Predictions grounded as a function of such concepts have the potential for improved interpretability. We provide evidence that MLLM embedding spaces can be grounded in such attributes, through fine-tuning to predict their values from images. Concretely, we evaluate this grounding in the embedding space through an attribute-specific content-based image retrieval case study using the SLICE-3D dataset.

📄 PDF Abstract BibTeX arXiv:2508.20188

Code (0)

등록된 구현이 없습니다.

Tasks

Content-Based Image Retrieval

Similar Papers 제목 키워드 기반

Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding

2026-03-27 · Shrinidhi Kumbhar, Haofu Liao, Srikar Appalaraju, Kunwar Yashraj Singh arxiv

Autoregressive (AR) vision-language models (VLMs) have long dominated multimodal understanding, reasoning, and graphical user interface (GUI) grounding. Recently, discrete diffusion vision-language models (DVLMs) have sh…

Multimodal ReasoningText Generation

Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback

2025-12-01 · Aiden Yiliu Li, Bizhi Yu, Daoan Lei, Tianhe Ren 외 arxiv

GUI grounding aims to align natural language instructions with precise regions in complex user interfaces. Advanced multimodal large language models show strong ability in visual GUI grounding but still struggle with sma…

Visual Reasoning

Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding

2020-12-18 · Dexin Wang, Deyi Xiong

Visual context provides grounding information for multimodal machine translation (MMT). However, previous MMT models and probing studies on visual features suggest that visual information is less explored in MMT as it is…

Machine TranslationMultimodal Machine TranslationObjectTranslation

Syntax-Guided Transformers: Elevating Compositional Generalization and Grounding in Multimodal Environments

2023-11-07 · Danial Kamali, Parisa Kordjamshidi

Compositional generalization, the ability of intelligent models to extrapolate understanding of components to novel compositions, is a fundamental yet challenging facet in AI research, especially within multimodal enviro…

Compositional Generalization (AVG)Dependency Parsing

Semi-supervised multimodal coreference resolution in image narrations

2023-10-20 · Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen

In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image. This poses significant challenges due to fine-grained image-text alignme…

coreference-resolutionCoreference ResolutionDescriptive