paper-with-me

AVA-ActiveSpeaker

홈페이지 · 논문 22편

Contains temporally labeled face tracks in video, where each face instance is labeled as speaking or not, and whether the speech is audible. This dataset contains about 3.65 million human labeled frames or about 38.5 hours of face tracks, and the corresponding audio. Source: [AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection](/paper/ava-activespeaker-an-audio-visual-dataset-for)

벤치마크

Audio-Visual Active Speaker Detection on AVA-ActiveSpeaker 결과 20개