paper-with-me

Papers

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets

2025-01-06 · Tanay Agrawal, Mohammed Guermal, Michal Balazia, Francois Bremond

Challenges in cross-learning involve inhomogeneous or even inadequate amount of training data and lack of resources for retraining large pretrained models. Inspired by transfer learning techniques in NLP, adapters and prefix tuning, this paper presents a new model-agnostic plugin architecture for cross-learning, called CM3T, that adapts transformer-based models to new or missing information. We introduce two adapter blocks: multi-head vision adapters for transfer learning and cross-attention adapters for multimodal learning. Training becomes substantially efficient as the backbone and other plugins do not need to be finetuned along with these additions. Comparative and ablation studies on three datasets Epic-Kitchens-100, MPIIGroupInteraction and UDIVA v0.5 show efficacy of this framework on different recording settings and tasks. With only 12.8% trainable parameters compared to the backbone to process video input and only 22.3% trainable parameters for two additional modalities, we achieve comparable and even better results than the state-of-the-art. CM3T has no specific requirements for training or pretraining and is a step towards bridging the gap between a general model and specific practical applications of video classification.

📄 PDF Abstract BibTeX arXiv:2501.03332

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer LearningVideo Classification

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

Coarse Graining of Data via Inhomogeneous Diffusion Condensation

2019-07-10 · Nathan Brugnone, Alex Gonopolskiy, Mark W. Moyle, Manik Kuchroo 외

Big data often has emergent structure that exists at multiple levels of abstraction, which are useful for characterizing complex interactions and dynamics of the observations. Here, we consider multiple levels of abstrac…

Clustering

Windowed Fourier Propagator: A Frequency-Local Neural Operator for Wave Equations in Inhomogeneous Media

2026-03-15 · Yiyang Cai, Zixuan Qiu, Yunlu Shu, Jiamao Wu 외 arxiv

Wave equations are fundamental to describing a vast array of physical phenomena, yet their simulation in inhomogeneous media poses a computational challenge due to the highly oscillatory nature of the solutions. To overc…

Computational Efficiency

Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework

2023-02-23 · NeurIPS 2023 11 · Paul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling 외

The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advance…

Model Selection

InterMulti:Multi-view Multimodal Interactions with Text-dominated Hierarchical High-order Fusion for Emotion Analysis

2022-12-20 · Feng Qiu, Wanzeng Kong, Yu Ding

Humans are sophisticated at reading interlocutors' emotions from multimodal signals, such as speech contents, voice tones and facial expressions. However, machines might struggle to understand various emotions due to the…

Emotion Recognitionmultimodal interaction

Structured Inhomogeneous Density Map Learning for Crowd Counting

2018-01-20 · Hanhui Li, Xiangjian He, Hefeng Wu, Saeed Amirgholipour Kasmani 외

In this paper, we aim at tackling the problem of crowd counting in extremely high-density scenes, which contain hundreds, or even thousands of people. We begin by a comprehensive analysis of the most widely used density …

Crowd Counting