Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA)
This report describes our submission to the ActivityNet Challenge at CVPR 2019. We use a 3D convolutional neural network (CNN) based front-end and an ensemble of temporal convolution and LSTM classifiers to predict whether a visible person is speaking or not. Our results show significant improvements over the baseline on the AVA-ActiveSpeaker dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Active Speaker DetectionAudio-Visual Active Speaker DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ICTCAS-UCAS-TAL Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2021
This report presents a brief description of our method for the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2021. Our solution, the Extended Unified Context Network (Extended UniCon) is based on a no…
Active Speaker DetectionAudio-Visual Active Speaker DetectionUniCon+: ICTCAS-UCAS Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2022
This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Uni…
Active Speaker DetectionAudio-Visual Active Speaker DetectionMulti-Task Learning for Audio Visual Active Speaker Detection
This report describes the approach underlying our submission to the active speaker detection task (task B-2) of ActivityNet Challenge 2019. We introduce a new audio-visual model which builds upon a 3D-ResNet18 visual mod…
Active Speaker DetectionAudio-Visual Active Speaker DetectionLipreadingMulti-Task Learning+1NUS-HLT Report for ActivityNet Challenge 2021 AVA (Speaker)
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as…
Active Speaker DetectionAudio-Visual Active Speaker DetectionNAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
Visual Grounding (VG) tasks, such as referring expression detection and segmentation tasks are important for linking visual entities to context, especially in complex reasoning tasks that require detailed query interpret…
Referring ExpressionVisual Grounding