paper-with-me

Papers audio-visual learning

“audio-visual learning” 태그가 달린 논문 38편 · 필터 해제

Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework

2025-06-09 · Kuiyuan Zhang, Wenjie Pei, Rushi Lan, Yifang Guo 외

Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly e…

audio-visual learningDeepFake DetectionFace SwappingMisinformation+1

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

2025-05-02 · CVPR 2025 1 · Edson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati 외

Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained tempora…

audio-visual learningcross-modal alignment

Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives

2025-02-17 · Zeliang Zhang, Susan Liang, Daiki Shimada, Chenliang Xu

While audio-visual learning equips models with a richer understanding of the real world by leveraging multiple sensory modalities, this integration also introduces new vulnerabilities to adversarial attacks. In this pape…

Adversarial Robustnessaudio-visual learning

Language-Guided Audio-Visual Learning for Long-Term Sports Assessment

2025-01-01 · CVPR 2025 1 · Huangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu 외

Long-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the divers…

audio-visual learningKnowledge GraphsVideo Understanding

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

2024-12-17 · Ziheng Zhou, Jinxing Zhou, Wei Qian, Shengeng Tang 외

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene unde…

audio-visual event localizationaudio-visual learningScene Understanding

Enhancing Sound Source Localization via False Negative Elimination

2024-08-29 · Zengjie Song, Jiangshe Zhang, Yuxi Wang, Junsong Fan 외

Sound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling …

audio-visual learningContrastive Learningobject-detectionObject Detection+1

Unveiling Visual Biases in Audio-Visual Localization Benchmarks

2024-08-25 · Liangyu Chen, Zihao Yue, Boshen Xu, Qin Jin

Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based s…

audio-visual learningVisual Localization

Sequential Contrastive Audio-Visual Learning

2024-07-08 · Ioannis Tsiamas, Santiago Pascual, Chunghsin Yeh, Joan Serrà

Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in webscale video datasets. However, conventional cont…

audio-visual learningContrastive LearningRepresentation LearningRetrieval

MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers

2024-06-07 · Tanvir Mahmud, Shentong Mo, Yapeng Tian, Diana Marculescu

Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigated effective methods for aligning multimo…

audio-visual learningContrastive Learning

EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning

2024-03-14 · Jongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim 외

Recent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified…

Audio Classificationaudio-visual learningContrastive LearningData Augmentation+1

Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization

2024-01-16 · Ming Cheng, Ming Li

Audio-visual learning has demonstrated promising results in many classical speech tasks (e.g., speech separation, automatic speech recognition, wake-word spotting). We believe that introducing visual modality will also b…

Action DetectionActivity Detectionaudio-visual learningAutomatic Speech Recognition+5

Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline

2023-11-29 · Xuecheng Wu, Heli Sun, Junxiao Xue, Jiayu Nie 외

Nowadays, short-form videos (SVs) are essential to web information acquisition and sharing in our daily life. The prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA)…

audio-visual learningFormMultimodal Emotion RecognitionVideo Emotion Recognition

Boosting Audio-visual Zero-shot Learning with Large Language Models

2023-11-21 · Haoxing Chen, Yaohui Li, Yan Hong, Zizheng Huang 외

Audio-visual zero-shot learning aims to recognize unseen classes based on paired audio-visual sequences. Recent methods mainly focus on learning multi-modal features aligned with class names to enhance the generalization…

audio-visual learningDescriptiveGZSL Video ClassificationZero-Shot Learning

Can CLIP Help Sound Source Localization?

2023-11-07 · Sooyoung Park, Arda Senocak, Joon Son Chung

Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment. We extend the application …

audio-visual learningContrastive LearningSound Source Localization

Deep Video Inpainting Guided by Audio-Visual Self-Supervision

2023-10-11 · Kyuyeon Kim, Junsik Jung, Woo Jae Kim, Sung-Eui Yoon

Humans can easily imagine a scene from auditory information based on their prior knowledge of audio-visual events. In this paper, we mimic this innate human ability in deep learning models to improve the quality of video…

audio-visual learningVideo Inpainting

AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models

2023-09-19 · Yuan Tseng, Layne Berry, Yi-Ting Chen, I-Hsiang Chiu 외

Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information. However, current models often focus on a limited set of tasks, and…

audio-visual learningRepresentation Learning

Class-Incremental Grouping Network for Continual Audio-Visual Learning

2023-09-11 · ICCV 2023 1 · Shentong Mo, Weiguo Pian, Yapeng Tian

Continual learning is a challenging problem in which models need to be trained on non-stationary data across sequential tasks for class-incremental learning. While previous methods have focused on using either regulariza…

audio-visual learningclass-incremental learningClass Incremental LearningContinual Learning+3

Leveraging Pretrained Image-text Models for Improving Audio-Visual Learning

2023-09-08 · Saurabhchand Bhati, Jesús Villalba, Laureano Moro-Velazquez, Thomas Thebaud 외

Visually grounded speech systems learn from paired images and their spoken captions. Recently, there have been attempts to utilize the visually grounded models trained from images and their corresponding text captions, s…

audio-visual learningQuantizationWord Embeddings

RealImpact: A Dataset of Impact Sound Fields for Real Objects

2023-06-16 · CVPR 2023 1 · Samuel Clarke, Ruohan Gao, Mason Wang, Mark Rau 외

Objects make unique sounds under different perturbations, environment conditions, and poses relative to the listener. While prior works have modeled impact sounds and sound propagation in simulation, we lack a standard d…

audio-visual learning

A Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition

2023-05-30 · Shentong Mo, Pedro Morgado

The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately, with several methods developed independe…

audio-visual learning
1–20 / 38 다음 →