paper-with-me

Papers Audio-Visual Active Speaker Detection

“Audio-Visual Active Speaker Detection” 태그가 달린 논문 25편 · 필터 해제

LASER: Lip Landmark Assisted Speaker Detection for Robustness

2025-01-21 · Le Thien Phuc Nguyen, Zhuoran Yu, Yong Jae Lee

Active Speaker Detection (ASD) aims to identify speaking individuals in complex visual scenes. While humans can easily detect speech by matching lip movements to audio, current ASD models struggle to establish this corre…

Active Speaker DetectionAudio-Visual Active Speaker Detection

An Efficient and Streaming Audio Visual Active Speaker Detection System

2024-09-13 · Arnav Kundu, Yanzi Jin, Mohammad Sekhavat, Max Horton 외

This paper delves into the challenging task of Active Speaker Detection (ASD), where the system needs to determine in real-time whether a person is speaking or not in a series of video frames. While previous works have m…

Active Speaker DetectionAudio-Visual Active Speaker DetectionCPU

Enhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training

2024-04-01 · Ruijie Tao, Xinyuan Qian, Rohan Kumar Das, Xiaoxue Gao 외

Audio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, th…

Active Speaker DetectionAudio-Visual Active Speaker DetectionDenoisingTarget Speaker Extraction

TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning

2023-09-21 · Chaeyoung Jung, Suyeon Lee, Kihyun Nam, Kyeongha Rho 외

The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures whi…

Active Speaker DetectionAudio-Visual Active Speaker DetectionContrastive Learning

A Light Weight Model for Active Speaker Detection

2023-03-08 · CVPR 2023 1 · Junhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao 외

Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial i…

Active Speaker DetectionAudio-Visual Active Speaker Detectionmodelspeaker-diarization+2

LoCoNet: Long-Short Context Network for Active Speaker Detection

2023-01-19 · CVPR 2024 1 · Xizi Wang, Feng Cheng, Gedas Bertasius, David Crandall

Active Speaker Detection (ASD) aims to identify who is speaking in each frame of a video. ASD reasons from audio and visual information from two contexts: long-term intra-speaker context and short-term inter-speaker cont…

Active Speaker DetectionAudio-Visual Active Speaker Detection

Audio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection

2022-12-01 · Rahul Sharma, Shrikanth Narayanan

Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of information to derive such a speech-face r…

Active Speaker DetectionAudio-Visual Active Speaker Detection

Push-Pull: Characterizing the Adversarial Robustness for Audio-Visual Active Speaker Detection

2022-10-03 · Xuanjun Chen, Haibin Wu, Helen Meng, Hung-Yi Lee 외

Audio-visual active speaker detection (AVASD) is well-developed, and now is an indispensable front-end for several multi-modal applications. However, to the best of our knowledge, the adversarial robustness of AVASD mode…

Active Speaker DetectionAdversarial RobustnessAudio-Visual Active Speaker Detection

Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection

2022-07-15 · Kyle Min, Sourya Roy, Subarna Tripathi, Tanaya Guha 외

Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we…

Active Speaker DetectionAudio-Visual Active Speaker DetectionGraph LearningNode Classification

UniCon+: ICTCAS-UCAS Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2022

2022-06-22 · Yuanhang Zhang, Susan Liang, Shuang Yang, Shiguang Shan

This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Uni…

Active Speaker DetectionAudio-Visual Active Speaker Detection

End-to-End Active Speaker Detection

2022-03-27 · Juan Leon Alcazar, Moritz Cordes, Chen Zhao, Bernard Ghanem

Recent advances in the Active Speaker Detection (ASD) problem build upon a two-stage process: feature extraction and spatio-temporal context aggregation. In this paper, we propose an end-to-end ASD workflow where feature…

Active Speaker DetectionAudio-Visual Active Speaker DetectionGraph Neural Network

Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization

2022-01-06 · CVPR 2022 1 · Hao Jiang, Calvin Murdock, Vamsi Krishna Ithapu

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for und…

Action DetectionActive Speaker DetectionActive Speaker LocalizationActivity Detection+1

Learning Spatial-Temporal Graphs for Active Speaker Detection

2021-12-02 · Sourya Roy, Kyle Min, Subarna Tripathi, Tanaya Guha 외

We address the problem of active speaker detection through a new framework, called SPELL, that learns long-range multimodal graphs to encode the inter-modal relationship between audio and visual data. We cast active spea…

Active Speaker DetectionAudio-Visual Active Speaker DetectionNode Classification

Sub-word Level Lip Reading With Visual Attention

2021-10-14 · CVPR 2022 1 · K R Prajwal, Triantafyllos Afouras, Andrew Zisserman

The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech rec…

Audio-Visual Active Speaker DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Lipreading+4

UniCon: Unified Context Network for Robust Active Speaker Detection

2021-08-05 · Yuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu 외

We introduce a new efficient framework, the Unified Context Network (UniCon), for robust active speaker detection (ASD). Traditional methods for ASD usually operate on each candidate's pre-cropped face track separately a…

Active Speaker DetectionAudio-Visual Active Speaker Detection

Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

2021-07-14 · Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 외

Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as…

Active Speaker DetectionAudio-Visual Active Speaker Detection

Active Speaker Detection as a Multi-Objective Optimization with Uncertainty-based Multimodal Fusion

2021-06-07 · Baptiste Pouthier, Laurent Pilati, Leela K. Gudupudi, Charles Bouveyron 외

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovi…

Active Speaker DetectionAudio-Visual Active Speaker Detection

How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild

2021-06-07 · ICCV 2021 10 · Okan Köpüklü, Maja Taseska, Gerhard Rigoll

Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers wi…

Active Speaker DetectionAudio-Visual Active Speaker Detection

ICTCAS-UCAS-TAL Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2021

2021-06-01 · The ActivityNet Large-Scale Activity Recognition Challenge Workshop, CVPR 2021 6 · Yuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu 외

This report presents a brief description of our method for the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2021. Our solution, the Extended Unified Context Network (Extended UniCon) is based on a no…

Active Speaker DetectionAudio-Visual Active Speaker Detection

NUS-HLT Report for ActivityNet Challenge 2021 AVA (Speaker)

2021-06-01 · The ActivityNet Large-Scale Activity Recognition Challenge Workshop, CVPR 2021 6 · Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 외

Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as…

Active Speaker DetectionAudio-Visual Active Speaker Detection
1–20 / 25 다음 →