paper-with-me

홈 › Papers

Deep Audio-Visual Learning: A Survey

2020-01-14 · Hao Zhu, Mandi Luo, Rui Wang, Aihua Zheng, Ran He

Audio-visual learning, aimed at exploiting the relationship between audio and visual modalities, has drawn considerable attention since deep learning started to be used successfully. Researchers tend to leverage these two modalities either to improve the performance of previously considered single-modality tasks or to address new challenging problems. In this paper, we provide a comprehensive survey of recent audio-visual learning development. We divide the current audio-visual learning tasks into four different subfields: audio-visual separation and localization, audio-visual correspondence learning, audio-visual generation, and audio-visual representation learning. State-of-the-art methods as well as the remaining challenges of each subfield are further discussed. Finally, we summarize the commonly used datasets and performance metrics.

📄 PDF Abstract BibTeX arXiv:2001.04758

Code (0)

등록된 구현이 없습니다.

Tasks

audio-visual learningRepresentation LearningSurvey

Similar Papers 제목 키워드 기반

A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

2021-08-01 · Zhaofeng Shi

With the development of deep learning and artificial intelligence, audio synthesis has a pivotal role in the area of machine learning and shows strong applicability in the industry. Meanwhile, significant efforts have be…

Audio SynthesisMusic GenerationSurveytext-to-speech+1

Learning in Audio-visual Context: A Review, Analysis, and New Perspective

2022-08-20 · Yake Wei, Di Hu, Yapeng Tian, Xuelong Li

Sight and hearing are two senses that play a vital role in human communication and scene understanding. To mimic human perception ability, audio-visual learning, aimed at developing computational approaches to learn from…

audio-visual learningScene UnderstandingSurvey

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models

2025-06-13 · Jinming Wen, Xinyi Wu, Shuai Zhao, Yanhao Jia 외

Multimodal large language models (MLLMs), which bridge the gap between audio-visual and natural language processing, achieve state-of-the-art performance on several audio-visual tasks. Despite the superior performance of…

Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

2022-02-28 · Luís Vilaça, Yi Yu, Paula Viana

Audio-visual correlation learning aims to capture essential correspondences and understand natural phenomena between audio and video. With the rapid growth of deep learning, an increasing amount of attention has been pai…

Survey

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

2024-11-24 · Luis Vilaca, Yi Yu, Paula Vinan

Audio-visual correlation learning aims to capture and understand natural phenomena between audio and visual data. The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data an…