paper-with-me

Papers

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

2025-09-17 · Muhammad Imran, Yugyung Lee arxiv

Recent advances in vision-language models have significantly expanded the frontiers of automated image analysis. However, applying these models in safety-critical contexts remains challenging due to the complex relationships between objects, subtle visual cues, and the heightened demand for transparency and reliability. This paper presents the Multi-Modal Explainable Learning (MMEL) framework, designed to enhance the interpretability of vision-language models while maintaining high performance. Building upon prior work in gradient-based explanations for transformer architectures (Grad-eclip), MMEL introduces a novel Hierarchical Semantic Relationship Module that enhances model interpretability through multi-scale feature processing, adaptive attention weighting, and cross-modal alignment. Our approach processes features at multiple semantic levels to capture relationships between image regions at different granularities, applying learnable layer-specific weights to balance contributions across the model's depth. This results in more comprehensive visual explanations that highlight both primary objects and their contextual relationships with improved precision. Through extensive experiments on standard datasets, we demonstrate that by incorporating semantic relationship information into gradient-based attribution maps, MMEL produces more focused and contextually aware visualizations that better reflect how vision-language models process complex scenes. The MMEL framework generalizes across various domains, offering valuable insights into model decisions for applications requiring high interpretability and reliability.

📄 PDF Abstract BibTeX arXiv:2509.15243

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

2024-11-29 · Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Ser-Nam Lim, Rajiv Ramnath

Document Visual Question Answering (VQA) requires models to interpret textual information within complex visual layouts and comprehend spatial relationships to answer questions based on document images. Existing approach…

Optical Character Recognition (OCR)Question AnsweringText DetectionVisual Question Answering+1

Explaining Multi-modal Large Language Models by Analyzing their Vision Perception

2024-05-23 · Loris Giulivi, Giacomo Boracchi

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in understanding and generating content across various modalities, such as images and text. However, their interpretability remains a ch…

Object Localization

ELVIS: Empowering Locality of Vision Language Pre-training with Intra-modal Similarity

2023-04-11 · Sumin Seo, Jaewoong Shin, Jaewoo Kang, Tae Soo Kim 외

Deep learning has shown great potential in assisting radiologists in reading chest X-ray (CXR) images, but its need for expensive annotations for improving performance prevents widespread clinical application. Visual lan…

Phrase Grounding

Interpretability Transfer from Language to Vision via Sparse Autoencoders

2026-05-24 · Alexey Kravets, Da Li, Chuan Li, Da Chen 외 arxiv

Recent advances in language model interpretability using sparse autoencoders (SAEs) have yet to effectively translate to the visual domain, mainly due to the difficulty and ambiguity of labeling visual concepts. In this …

FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization

2024-04-21 · Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 외

Zero-shot anomaly detection (ZSAD) methods entail detecting anomalies directly without access to any known normal or abnormal samples within the target item categories. Existing approaches typically rely on the robust ge…

Anomaly DetectionPositionzero-shot anomaly detection