paper-with-me

홈 › Papers

Multi-modal Aggregation for Video Classification

2017-10-27 · Chen Chen, Xiaowei Zhao, Yang Liu

In this paper, we present a solution to Large-Scale Video Classification Challenge (LSVC2017) [1] that ranked the 1st place. We focused on a variety of modalities that cover visual, motion and audio. Also, we visualized the aggregation process to better understand how each modality takes effect. Among the extracted modalities, we found Temporal-Spatial features calculated by 3D convolution quite promising that greatly improved the performance. We attained the official metric mAP 0.8741 on the testing set with the ensemble model.

📄 PDF Abstract BibTeX arXiv:1710.10330

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationVideo Classification

Similar Papers 제목 키워드 기반

Towards Good Practices for Multi-modal Fusion in Large-scale Video Classification

2018-09-16 · Jinlai Liu, Zehuan Yuan, Changhu Wang

Leveraging both visual frames and audio has been experimentally proven effective to improve large-scale video classification. Previous research on video classification mainly focuses on the analysis of visual content amo…

ClassificationGeneral ClassificationVideo Classification

Frame Aggregation and Multi-Modal Fusion Framework for Video-Based Person Recognition

2020-10-19 · Fangtao Li, Wenzhe Wang, Zihe Liu, Haoran Wang 외

Video-based person recognition is challenging due to persons being blocked and blurred, and the variation of shooting angle. Previous research always focused on person recognition on still images, ignoring similarity and…

Person Recognition

Learnable pooling with Context Gating for video classification

2017-06-21 · Antoine Miech, Ivan Laptev, Josef Sivic

Current methods for video analysis often extract frame-level features using pre-trained convolutional neural networks (CNNs). Such features are then aggregated over time e.g., by simple temporal averaging or more sophist…

ClassificationClusteringGeneral ClassificationVideo Classification+1

MUVF-YOLOX: A Multi-modal Ultrasound Video Fusion Network for Renal Tumor Diagnosis

2023-07-15 · Junyu Li, Han Huang, Dong Ni, Wufeng Xue 외

Early diagnosis of renal cancer can greatly improve the survival rate of patients. Contrast-enhanced ultrasound (CEUS) is a cost-effective and non-invasive imaging technique and has become more and more frequently used f…

Video Classification

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

2023-08-22 · ICCV 2023 1 · Najmeh Sadoughi, Xinyu Li, Avijit Vajpayee, David Fan 외

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and …

Scene SegmentationSegmentationVideo SegmentationVideo Semantic Segmentation