paper-with-me

Papers

Interpretation on Multi-modal Visual Fusion

2023-08-19 · Hao Chen, Haoran Zhou, Yongjian Deng

In this paper, we present an analytical framework and a novel metric to shed light on the interpretation of the multimodal vision community. Our approach involves measuring the proposed semantic variance and feature similarity across modalities and levels, and conducting semantic and quantitative analyses through comprehensive experiments. Specifically, we investigate the consistency and speciality of representations across modalities, evolution rules within each modality, and the collaboration logic used when optimizing a multi-modality model. Our studies reveal several important findings, such as the discrepancy in cross-modal features and the hybrid multi-modal cooperation rule, which highlights consistency and speciality simultaneously for complementary inference. Through our dissection and findings on multi-modal fusion, we facilitate a rethinking of the reasonability and necessity of popular multi-modal vision fusion strategies. Furthermore, our work lays the foundation for designing a trustworthy and universal multi-modal fusion model for a variety of tasks in the future.

📄 PDF Abstract BibTeX arXiv:2308.10019

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset

2025-04-15 · Elisa Ancarani, Julie Tores, Lucile Sassatelli, Rémy Sun 외

We examine the impact of concept-informed supervision on multimodal video interpretation models using MOByGaze, a dataset containing human-annotated explanatory concepts. We introduce Concept Modality Specific Datasets (…

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

2026-07-26 · Musa Tur Farazi, Nufayer Jahan Reza arxiv

Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computat…

Dynamic Multimodal Sentiment Analysis: Leveraging Cross-Modal Attention for Enabled Classification

2025-01-14 · Hui Lee, Singh Suniljit, Yong Siang Ong

This paper explores the development of a multimodal sentiment analysis model that integrates text, audio, and visual data to enhance sentiment classification. The goal is to improve emotion detection by capturing the com…

Multimodal Sentiment AnalysisSentiment AnalysisSentiment Classification

Bring Event into RGB and LiDAR: Hierarchical Visual-Motion Fusion for Scene Flow

2024-03-12 · CVPR 2024 1 · Hanyu Zhou, Yi Chang, Zhiwei Shi, Luxin Yan

Single RGB or LiDAR is the mainstream sensor for the challenging scene flow, which relies heavily on visual features to match motion features. Compared with single modality, existing methods adopt a fusion strategy to di…

MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception

2024-06-22 · Guanqun Wang, Xinyu Wei, Jiaming Liu, Ray Zhang 외

In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual perception models have made significant stride…

Common Sense ReasoningLanguage ModellingLarge Language ModelMultimodal Large Language Model+4