Learning Spatial-Temporal Graphs for Active Speaker Detection
We address the problem of active speaker detection through a new framework, called SPELL, that learns long-range multimodal graphs to encode the inter-modal relationship between audio and visual data. We cast active speaker detection as a node classification task that is aware of longer-term dependencies. We first construct a graph from a video so that each node corresponds to one person. Nodes representing the same identity share edges between them within a defined temporal window. Nodes within the same video frame are also connected to encode inter-person interactions. Through extensive experiments on the Ava-ActiveSpeaker dataset, we demonstrate that learning graph-based representation, owing to its explicit spatial and temporal structure, significantly improves the overall performance. SPELL outperforms several relevant baselines and performs at par with state of the art models while requiring an order of magnitude lower computation cost.
Code (0)
등록된 구현이 없습니다.
Tasks
Active Speaker DetectionAudio-Visual Active Speaker DetectionNode ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we…
Active Speaker DetectionAudio-Visual Active Speaker DetectionGraph LearningNode ClassificationActive Speakers in Context
Current methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker. Although this strategy can be enough for addressing single-speaker scenarios, it prevents accurate…
Active Speaker DetectionAudio-Visual Active Speaker DetectionHow to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild
Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers wi…
Active Speaker DetectionAudio-Visual Active Speaker DetectionCross-modal Supervision for Learning Active Speaker Detection in Video
In this paper, we show how to use audio to supervise the learning of active speaker detection in video. Voice Activity Detection (VAD) guides the learning of the vision-based classifier in a weakly supervised manner. The…
Action DetectionActive Speaker DetectionActivity DetectionMAAS: Multi-modal Assignation for Active Speaker Detection
Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual featu…
Active Speaker DetectionAudio-Visual Active Speaker Detection