paper-with-me

Papers

Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters

2025-12-09 · Mizanur Rahman Jewel, Mohamed Elmahallawy, Sanjay Madria, Samuel Frimpong arxiv

Underground mining disasters produce pervasive darkness, dust, and collapses that obscure vision and make situational awareness difficult for humans and conventional systems. To address this, we propose MDSE, Multimodal Disaster Situation Explainer, a novel vision-language framework that automatically generates detailed textual explanations of post-disaster underground scenes. MDSE has three-fold innovations: (i) Context-Aware Cross-Attention for robust alignment of visual and textual features even under severe degradation; (ii) Segmentation-aware dual pathway visual encoding that fuses global and region-specific embeddings; and (iii) Resource-Efficient Transformer-Based Language Model for expressive caption generation with minimal compute cost. To support this task, we present the Underground Mine Disaster (UMD) dataset--the first image-caption corpus of real underground disaster scenes--enabling rigorous training and evaluation. Extensive experiments on UMD and related benchmarks show that MDSE substantially outperforms state-of-the-art captioning models, producing more accurate and contextually relevant descriptions that capture crucial details in obscured environments, improving situational awareness for underground emergency response. The code is at https://github.com/mizanJewel/Multimodal-Disaster-Situation-Explainer.

📄 PDF Abstract BibTeX arXiv:2512.09092

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Few-Shot Multimodal Explanation for Visual Question Answering

2024-10-28 · ACM MM 2024 10 · Dizhan Xue, Shengsheng Qian, Changsheng Xu

A key object in eXplainable Artificial Intelligence (XAI) is to create intelligent systems capable of reasoning and explaining real-world data to facilitate reliable decision-making. Recent studies have acknowledged the …

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)FS-MEVQAQuestion Answering+3

ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning

2025-10-30 · Jiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li 외 arxiv

Multimodal reasoning requires iterative coordination between language and vision, yet it remains unclear what constitutes a meaningful interleaved chain of thought. We posit that text and image thoughts should function a…

Multimodal Reasoning

Explaining CLIP Zero-shot Predictions Through Concepts

2026-03-30 · Onat Ozdemir, Anders Christensen, Stephan Alaniz, Zeynep Akata 외 arxiv

Large-scale vision-language models such as CLIP have achieved remarkable success in zero-shot image recognition, yet their predictions remain largely opaque to human understanding. In contrast, Concept Bottleneck Models …

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

2025-11-28 · Timothy Ossowski, Sheng Zhang, Qianchu Liu, Guanghui Qin 외 arxiv

High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tasks. We investigate strategies for traini…

Multimodal Reasoning

Towards Vision-Language-Garment Models For Web Knowledge Garment Understanding and Generation

2025-06-05 · Jan Ackermann, Kiyohiro Nakayama, Guandao Yang, Tong Wu 외

Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplored. We introduce VLG, a vision-language-g…

Zero-shot Generalization