paper-with-me

홈 › Papers

Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

2022-02-28 · Luís Vilaça, Yi Yu, Paula Viana

Audio-visual correlation learning aims to capture essential correspondences and understand natural phenomena between audio and video. With the rapid growth of deep learning, an increasing amount of attention has been paid to this emerging research issue. Through the past few years, various methods and datasets have been proposed for audio-visual correlation learning, which motivate us to conclude a comprehensive survey. This survey paper focuses on state-of-the-art (SOTA) models used to learn correlations between audio and video, but also discusses some tasks of definition and paradigm applied in AI multimedia. In addition, we investigate some objective functions frequently used for optimizing audio-visual correlation learning models and discuss how audio-visual data is exploited in the optimization process. Most importantly, we provide an extensive comparison and summarization of the recent progress of SOTA audio-visual correlation learning and discuss future research directions.

📄 PDF Abstract BibTeX arXiv:2202.13673

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

2024-11-24 · Luis Vilaca, Yi Yu, Paula Vinan

Audio-visual correlation learning aims to capture and understand natural phenomena between audio and visual data. The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data an…

Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen

2024-02-02 · María Sánchez, Laura Fernández, Julián Arias, Mateo Cámara 외

Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses li…

Contrastive Learning

Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert

2024-07-01 · Han EunGi, Oh Hyun-Bin, Kim Sung-Bin, Corentin Nivelet Etcheberry 외

Speech-driven 3D facial animation has recently garnered attention due to its cost-effective usability in multimedia production. However, most current advances overlook the intelligibility of lip movements, limiting the r…

Lip Reading

DAVE: Diagnostic benchmark for Audio Visual Evaluation

2025-03-12 · Gorjan Radevski, Teodora Popordanoska, Matthew B. Blaschko, Tinne Tuytelaars

Audio-visual understanding is a rapidly evolving field that seeks to integrate and interpret information from both auditory and visual modalities. Despite recent advances in multi-modal learning, existing benchmarks ofte…

Diagnostic

AVF-MAE++: Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised Learning

2025-01-01 · CVPR 2025 1 · Xuecheng Wu, Heli Sun, Yifan Wang, Jiayu Nie 외

Affective Video Facial Analysis (AVFA) is important for advancing emotion-aware AI, yet the persistent data scarcity in AVFA presents challenges. Recently, the self-supervised learning (SSL) technique of Masked Autoe…

Self-Supervised Learning