paper-with-me

홈 › Papers

AV-MaskEnhancer: Enhancing Video Representations through Audio-Visual Masked Autoencoder

2023-09-15 · Xingjian Diao, Ming Cheng, Shitong Cheng

Learning high-quality video representation has shown significant applications in computer vision and remains challenging. Previous work based on mask autoencoders such as ImageMAE and VideoMAE has proven the effectiveness of learning representations in images and videos through reconstruction strategy in the visual modality. However, these models exhibit inherent limitations, particularly in scenarios where extracting features solely from the visual modality proves challenging, such as when dealing with low-resolution and blurry original videos. Based on this, we propose AV-MaskEnhancer for learning high-quality video representation by combining visual and audio information. Our approach addresses the challenge by demonstrating the complementary nature of audio and video features in cross-modality content. Moreover, our result of the video classification task on the UCF101 dataset outperforms the existing work and reaches the state-of-the-art, with a top-1 accuracy of 98.8% and a top-5 accuracy of 99.9%.

📄 PDF Abstract BibTeX arXiv:2309.08738

Code (0)

등록된 구현이 없습니다.

Tasks

Video Classification

Similar Papers 제목 키워드 기반

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition

2025-07-12 · Kaixuan Cong, Yifan Wang, Rongkun Xue, Yuyang Jiang 외 arxiv

Humans do not understand individual events in isolation; rather, they generalize concepts within classes and compare them to others. Existing audio-video pre-training paradigms only focus on the alignment of the overall …

Human Activity Recognition

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

2026-05-09 · Huilai Li, Xiaomeng Di, Ying Xing, Yonghao Dang 외 arxiv

Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using only coarse-grained labels. Faced with the challenging task settings, e…

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

2026-02-12 · Maomao Li, Zhen Li, Kaipeng Zhang, Guosheng Yin 외 arxiv

Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, t…

Contrastive LearningVideo Generation

Video-to-Audio Generation with Hidden Alignment

2024-07-10 · Manjie Xu, Chenxing Li, Xinyi Tu, Yong Ren 외

Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In thi…

Audio GenerationData AugmentationText-to-Video GenerationVideo Generation

Text-to-Audio Generation Synchronized with Videos

2024-03-08 · Shentong Mo, Jing Shi, Yapeng Tian

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models…

AudioCapsAudio GenerationContrastive Learning