Papers Audio-Visual Active Speaker Detection
“Audio-Visual Active Speaker Detection” 태그가 달린 논문 25편 · 필터 해제
LASER: Lip Landmark Assisted Speaker Detection for Robustness
Active Speaker Detection (ASD) aims to identify speaking individuals in complex visual scenes. While humans can easily detect speech by matching lip movements to audio, current ASD models struggle to establish this corre…
Active Speaker DetectionAudio-Visual Active Speaker DetectionAn Efficient and Streaming Audio Visual Active Speaker Detection System
This paper delves into the challenging task of Active Speaker Detection (ASD), where the system needs to determine in real-time whether a person is speaking or not in a series of video frames. While previous works have m…
Active Speaker DetectionAudio-Visual Active Speaker DetectionCPUEnhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training
Audio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, th…
Active Speaker DetectionAudio-Visual Active Speaker DetectionDenoisingTarget Speaker ExtractionTalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning
The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures whi…
Active Speaker DetectionAudio-Visual Active Speaker DetectionContrastive LearningA Light Weight Model for Active Speaker Detection
Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial i…
Active Speaker DetectionAudio-Visual Active Speaker Detectionmodelspeaker-diarization+2LoCoNet: Long-Short Context Network for Active Speaker Detection
Active Speaker Detection (ASD) aims to identify who is speaking in each frame of a video. ASD reasons from audio and visual information from two contexts: long-term intra-speaker context and short-term inter-speaker cont…
Active Speaker DetectionAudio-Visual Active Speaker DetectionAudio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection
Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of information to derive such a speech-face r…
Active Speaker DetectionAudio-Visual Active Speaker DetectionPush-Pull: Characterizing the Adversarial Robustness for Audio-Visual Active Speaker Detection
Audio-visual active speaker detection (AVASD) is well-developed, and now is an indispensable front-end for several multi-modal applications. However, to the best of our knowledge, the adversarial robustness of AVASD mode…
Active Speaker DetectionAdversarial RobustnessAudio-Visual Active Speaker DetectionLearning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we…
Active Speaker DetectionAudio-Visual Active Speaker DetectionGraph LearningNode ClassificationUniCon+: ICTCAS-UCAS Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2022
This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Uni…
Active Speaker DetectionAudio-Visual Active Speaker DetectionEnd-to-End Active Speaker Detection
Recent advances in the Active Speaker Detection (ASD) problem build upon a two-stage process: feature extraction and spatio-temporal context aggregation. In this paper, we propose an end-to-end ASD workflow where feature…
Active Speaker DetectionAudio-Visual Active Speaker DetectionGraph Neural NetworkEgocentric Deep Multi-Channel Audio-Visual Active Speaker Localization
Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for und…
Action DetectionActive Speaker DetectionActive Speaker LocalizationActivity Detection+1Learning Spatial-Temporal Graphs for Active Speaker Detection
We address the problem of active speaker detection through a new framework, called SPELL, that learns long-range multimodal graphs to encode the inter-modal relationship between audio and visual data. We cast active spea…
Active Speaker DetectionAudio-Visual Active Speaker DetectionNode ClassificationSub-word Level Lip Reading With Visual Attention
The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech rec…
Audio-Visual Active Speaker DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Lipreading+4UniCon: Unified Context Network for Robust Active Speaker Detection
We introduce a new efficient framework, the Unified Context Network (UniCon), for robust active speaker detection (ASD). Traditional methods for ASD usually operate on each candidate's pre-cropped face track separately a…
Active Speaker DetectionAudio-Visual Active Speaker DetectionIs Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as…
Active Speaker DetectionAudio-Visual Active Speaker DetectionActive Speaker Detection as a Multi-Objective Optimization with Uncertainty-based Multimodal Fusion
It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovi…
Active Speaker DetectionAudio-Visual Active Speaker DetectionHow to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild
Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers wi…
Active Speaker DetectionAudio-Visual Active Speaker DetectionICTCAS-UCAS-TAL Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2021
This report presents a brief description of our method for the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2021. Our solution, the Extended Unified Context Network (Extended UniCon) is based on a no…
Active Speaker DetectionAudio-Visual Active Speaker DetectionNUS-HLT Report for ActivityNet Challenge 2021 AVA (Speaker)
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as…
Active Speaker DetectionAudio-Visual Active Speaker Detection