paper-with-me

홈 › Papers

Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset

2025-04-15 · Elisa Ancarani, Julie Tores, Lucile Sassatelli, Rémy Sun, Hui-Yin Wu, Frédéric Precioso

We examine the impact of concept-informed supervision on multimodal video interpretation models using MOByGaze, a dataset containing human-annotated explanatory concepts. We introduce Concept Modality Specific Datasets (CMSDs), which consist of data subsets categorized by the modality (visual, textual, or audio) of annotated concepts. Models trained on CMSDs outperform those using traditional legacy training in both early and late fusion approaches. Notably, this approach enables late fusion models to achieve performance close to that of early fusion models. These findings underscore the importance of modality-specific annotations in developing robust, self-explainable video models and contribute to advancing interpretable multimodal learning in complex video analysis.

📄 PDF Abstract BibTeX arXiv:2504.11232

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning

2025-10-09 · Shuangyan Deng, Haizhou Peng, Jiachen Xu, Rui Mao 외 arxiv

Multimodal Large Language Models (MLLMs) have made substantial progress in recent years. However, their rigorous evaluation within specialized domains like finance is hindered by the absence of datasets characterized by …

Mathematical Reasoning

ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video

2025-08-13 · Rajan Das Gupta, Lei Wei, Md Yeasin Rahat, Nafiz Fahad 외 arxiv

This study investigates the use of large language models (LLMs) for human behavior understanding by jointly leveraging motion and video data. We argue that integrating these complementary modalities is essential for capt…

A Large-scale Interpretable Multi-modality Benchmark for Facial Image Forgery Localization

2024-12-27 · Jingchun Lian, Lingyu Liu, Yaxiong Wang, Yujiao Wu 외

Image forgery localization, which centers on identifying tampered pixels within an image, has seen significant advancements. Traditional approaches often model this challenge as a variant of image segmentation, treating …

Face SwappingImage SegmentationLarge Language ModelMultimodal Large Language Model+1

Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM

2024-06-18 · Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo 외

Towards open-ended Video Anomaly Detection (VAD), existing methods often exhibit biased detection when faced with challenging or unseen events and lack interpretability. To address these drawbacks, we propose Holmes-VAD,…

Anomaly DetectionAnomaly LocalizationLanguage ModelingLanguage Modelling+3

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

2025-05-17 · Jiarui Wang, Huiyu Duan, Ziheng Jia, Yu Zhao 외

Recent advancements in large multimodal models (LMMs) have driven substantial progress in both text-to-video (T2V) generation and video-to-text (V2T) interpretation tasks. However, current AI-generated videos (AIGVs) sti…

BenchmarkingQuestion AnsweringText-to-Video GenerationVideo Alignment+1