paper-with-me

Papers

A Survey on Interpretable Cross-modal Reasoning

2023-09-05 · Dizhan Xue, Shengsheng Qian, Zuyi Zhou, Changsheng Xu

In recent years, cross-modal reasoning (CMR), the process of understanding and reasoning across different modalities, has emerged as a pivotal area with applications spanning from multimedia analysis to healthcare diagnostics. As the deployment of AI systems becomes more ubiquitous, the demand for transparency and comprehensibility in these systems' decision-making processes has intensified. This survey delves into the realm of interpretable cross-modal reasoning (I-CMR), where the objective is not only to achieve high predictive performance but also to provide human-understandable explanations for the results. This survey presents a comprehensive overview of the typical methods with a three-level taxonomy for I-CMR. Furthermore, this survey reviews the existing CMR datasets with annotations for explanations. Finally, this survey summarizes the challenges for I-CMR and discusses potential future directions. In conclusion, this survey aims to catalyze the progress of this emerging research area by providing researchers with a panoramic and comprehensive perspective, illuminating the state of the art and discerning the opportunities. The summarized methods, datasets, and other resources are available at https://github.com/ZuyiZhou/Awesome-Interpretable-Cross-modal-Reasoning.

📄 PDF Abstract BibTeX arXiv:2309.01955

Code (1)

ZuyiZhou/Awesome-Interpretable-Cross-modal-Reasoning 공식 구현 pytorch

Tasks

Cross-Modal RetrievalDecision MakingExplanation GenerationFactual Visual Question AnsweringFake News DetectionImage-guided Story Ending GenerationPhrase GroundingScience Question AnsweringSurveyVisual Commonsense ReasoningVisual Question AnsweringVisual Reasoning

Similar Papers 제목 키워드 기반

A Survey of Generative Categories and Techniques in Multimodal Large Language Models

2025-05-29 · Longzhen Han, Awes Mubarak, Almas Baimagambetov, Nikolaos Polatidis 외

Multimodal Large Language Models (MLLMs) have rapidly evolved beyond text generation, now spanning diverse output modalities including images, music, video, human motion, and 3D objects, by integrating language with othe…

Mixture-of-ExpertsSelf-Supervised LearningSurveyText Generation

A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models

2025-07-31 · Xiaoling Luo, Ruli Zheng, Qiaojian Zheng, Zibo Du 외 arxiv

Visual impairment represents a major global health challenge, with multimodal imaging providing complementary information that is essential for accurate ophthalmic diagnosis. This comprehensive survey systematically revi…

Multimodal Deep LearningSelf-Supervised LearningReinforcement Learning

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes

2026-06-09 · Avinash Anand, Mahisha Ramesh, Avni Mittal, Ashutosh Kumar 외 arxiv

Large Language Models (LLMs) have achieved strong performance across natural language processing tasks, yet reliable reasoning remains an open challenge. Although modern LLMs show progress in structured inference, multi-…

Mathematical ReasoningCommon Sense ReasoningReinforcement LearningDomain Generalization

Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

2024-01-10 · Yiqi Wang, Wentao Chen, Xiaotian Han, Xudong Lin 외

Strong Artificial Intelligence (Strong AI) or Artificial General Intelligence (AGI) with abstract reasoning ability is the goal of next-generation AI. Recent advancements in Large Language Models (LLMs), along with the e…

Multimodal ReasoningSurvey

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

2025-10-29 · Xu Zheng, Zihao Dongfang, Lutao Jiang, Boyuan Zheng 외 arxiv

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend these abilities by learning to perceive…

Vision-Language NavigationVisual Question AnsweringMultimodal ReasoningSpatial Reasoning