Papers audio-visual learning
“audio-visual learning” 태그가 달린 논문 38편 · 필터 해제
Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly e…
audio-visual learningDeepFake DetectionFace SwappingMisinformation+1CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained tempora…
audio-visual learningcross-modal alignmentRethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
While audio-visual learning equips models with a richer understanding of the real world by leveraging multiple sensory modalities, this integration also introduces new vulnerabilities to adversarial attacks. In this pape…
Adversarial Robustnessaudio-visual learningLanguage-Guided Audio-Visual Learning for Long-Term Sports Assessment
Long-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the divers…
audio-visual learningKnowledge GraphsVideo UnderstandingDense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene unde…
audio-visual event localizationaudio-visual learningScene UnderstandingEnhancing Sound Source Localization via False Negative Elimination
Sound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling …
audio-visual learningContrastive Learningobject-detectionObject Detection+1Unveiling Visual Biases in Audio-Visual Localization Benchmarks
Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based s…
audio-visual learningVisual LocalizationSequential Contrastive Audio-Visual Learning
Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in webscale video datasets. However, conventional cont…
audio-visual learningContrastive LearningRepresentation LearningRetrievalMA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigated effective methods for aligning multimo…
audio-visual learningContrastive LearningEquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
Recent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified…
Audio Classificationaudio-visual learningContrastive LearningData Augmentation+1Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
Audio-visual learning has demonstrated promising results in many classical speech tasks (e.g., speech separation, automatic speech recognition, wake-word spotting). We believe that introducing visual modality will also b…
Action DetectionActivity Detectionaudio-visual learningAutomatic Speech Recognition+5Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline
Nowadays, short-form videos (SVs) are essential to web information acquisition and sharing in our daily life. The prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA)…
audio-visual learningFormMultimodal Emotion RecognitionVideo Emotion RecognitionBoosting Audio-visual Zero-shot Learning with Large Language Models
Audio-visual zero-shot learning aims to recognize unseen classes based on paired audio-visual sequences. Recent methods mainly focus on learning multi-modal features aligned with class names to enhance the generalization…
audio-visual learningDescriptiveGZSL Video ClassificationZero-Shot LearningCan CLIP Help Sound Source Localization?
Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment. We extend the application …
audio-visual learningContrastive LearningSound Source LocalizationDeep Video Inpainting Guided by Audio-Visual Self-Supervision
Humans can easily imagine a scene from auditory information based on their prior knowledge of audio-visual events. In this paper, we mimic this innate human ability in deep learning models to improve the quality of video…
audio-visual learningVideo InpaintingAV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models
Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information. However, current models often focus on a limited set of tasks, and…
audio-visual learningRepresentation LearningClass-Incremental Grouping Network for Continual Audio-Visual Learning
Continual learning is a challenging problem in which models need to be trained on non-stationary data across sequential tasks for class-incremental learning. While previous methods have focused on using either regulariza…
audio-visual learningclass-incremental learningClass Incremental LearningContinual Learning+3Leveraging Pretrained Image-text Models for Improving Audio-Visual Learning
Visually grounded speech systems learn from paired images and their spoken captions. Recently, there have been attempts to utilize the visually grounded models trained from images and their corresponding text captions, s…
audio-visual learningQuantizationWord EmbeddingsRealImpact: A Dataset of Impact Sound Fields for Real Objects
Objects make unique sounds under different perturbations, environment conditions, and poses relative to the listener. While prior works have modeled impact sounds and sound propagation in simulation, we lack a standard d…
audio-visual learningA Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition
The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately, with several methods developed independe…
audio-visual learning