paper-with-me

Papers

Modality Compensation Network: Cross-Modal Adaptation for Action Recognition

2020-01-31 · Sijie Song, Jiaying Liu, Yanghao Li, Zongming Guo

With the prevalence of RGB-D cameras, multi-modal video data have become more available for human action recognition. One main challenge for this task lies in how to effectively leverage their complementary information. In this work, we propose a Modality Compensation Network (MCN) to explore the relationships of different modalities, and boost the representations for human action recognition. We regard RGB/optical flow videos as source modalities, skeletons as auxiliary modality. Our goal is to extract more discriminative features from source modalities, with the help of auxiliary modality. Built on deep Convolutional Neural Networks (CNN) and Long Short Term Memory (LSTM) networks, our model bridges data from source and auxiliary modalities by a modality adaptation block to achieve adaptive representation learning, that the network learns to compensate for the loss of skeletons at test time and even at training time. We explore multiple adaptation schemes to narrow the distance between source and auxiliary modal distributions from different levels, according to the alignment of source and auxiliary data in training. In addition, skeletons are only required in the training phase. Our model is able to improve the recognition performance with source data when testing. Experimental results reveal that MCN outperforms state-of-the-art approaches on four widely-used action recognition benchmarks.

📄 PDF Abstract BibTeX arXiv:2001.11657

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionOptical Flow EstimationRepresentation LearningTemporal Action Localization

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

FMCNet: Feature-Level Modality Compensation for Visible-Infrared Person Re-Identification

2022-01-01 · CVPR 2022 1 · Qiang Zhang, Changzhou Lai, Jianan Liu, Nianchang Huang 외

For Visible-Infrared person Re-IDentification (VI-ReID), existing modality-specific information compensation based models try to generate the images of missing modality from existing ones for reducing cross-modality …

Person Re-Identification

MRCN: A Novel Modality Restitution and Compensation Network for Visible-Infrared Person Re-identification

2023-03-26 · Yukang Zhang, Yan Yan, Jie Li, Hanzi Wang

Visible-infrared person re-identification (VI-ReID), which aims to search identities across different spectra, is a challenging task due to large cross-modality discrepancy between visible and infrared images. The key to…

Person Re-Identification

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference

2025-06-06 · Kunxi Li, Zhonghua Jiang, Zhouzhou Shen, Zhaode Wang 외

This paper introduces MadaKV, a modality-adaptive key-value (KV) cache eviction strategy designed to enhance the efficiency of multimodal large language models (MLLMs) in long-context inference. In multimodal scenarios, …

CLIP-Based Modality Compensation for Visible-Infrared Image Re-Identification

2024-12-25 · journal 2024 12 · Gang Hu; Yafei Lv; Jianting Zhang; Qian Wu; Zaidao Wen

Visible-infrared image re-identification (VIReID) aims to match objects with the same identity appearing across different modalities. Given the significant differences between visible and infrared images, VIReID poses a …

Cross-Modal Person Re-IdentificationCross-Modal Person Re-IdentificationPerson Re-IdentificationPrompt Learning

Prototype-Based Information Compensation Network for Multi-Source Remote Sensing Data Classification

2025-05-06 · Feng Gao, Sheng Liu, Chuanzheng Gong, Xiaowei Zhou 외

Multi-source remote sensing data joint classification aims to provide accuracy and reliability of land cover classification by leveraging the complementary information from multiple data sources. Existing methods confron…

Land Cover Classification