paper-with-me

홈 › Papers

Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA)

2019-06-25 · Joon Son Chung

This report describes our submission to the ActivityNet Challenge at CVPR 2019. We use a 3D convolutional neural network (CNN) based front-end and an ensemble of temporal convolution and LSTM classifiers to predict whether a visible person is speaking or not. Our results show significant improvements over the baseline on the AVA-ActiveSpeaker dataset.

📄 PDF Abstract BibTeX arXiv:1906.10555

Code (0)

등록된 구현이 없습니다.

Tasks

Active Speaker DetectionAudio-Visual Active Speaker Detection

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

ICTCAS-UCAS-TAL Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2021

2021-06-01 · The ActivityNet Large-Scale Activity Recognition Challenge Workshop, CVPR 2021 6 · Yuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu 외

This report presents a brief description of our method for the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2021. Our solution, the Extended Unified Context Network (Extended UniCon) is based on a no…

Active Speaker DetectionAudio-Visual Active Speaker Detection

UniCon+: ICTCAS-UCAS Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2022

2022-06-22 · Yuanhang Zhang, Susan Liang, Shuang Yang, Shiguang Shan

This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Uni…

Active Speaker DetectionAudio-Visual Active Speaker Detection

Multi-Task Learning for Audio Visual Active Speaker Detection

2019-06-01 · The ActivityNet Large-Scale Activity Recognition Challenge Workshop, CVPR 2019 6 · Yuanhang Zhang, Jingyun Xiao, Shuang Yang, Shiguang Shan

This report describes the approach underlying our submission to the active speaker detection task (task B-2) of ActivityNet Challenge 2019. We introduce a new audio-visual model which builds upon a 3D-ResNet18 visual mod…

Active Speaker DetectionAudio-Visual Active Speaker DetectionLipreadingMulti-Task Learning+1

NUS-HLT Report for ActivityNet Challenge 2021 AVA (Speaker)

2021-06-01 · The ActivityNet Large-Scale Activity Recognition Challenge Workshop, CVPR 2021 6 · Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 외

Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as…

Active Speaker DetectionAudio-Visual Active Speaker Detection

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

2025-02-01 · Zhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda 외

Visual Grounding (VG) tasks, such as referring expression detection and segmentation tasks are important for linking visual entities to context, especially in complex reasoning tasks that require detailed query interpret…

Referring ExpressionVisual Grounding