paper-with-me

Papers

Learning Spatial-Temporal Graphs for Active Speaker Detection

2021-12-02 · Sourya Roy, Kyle Min, Subarna Tripathi, Tanaya Guha, Somdeb Majumdar

We address the problem of active speaker detection through a new framework, called SPELL, that learns long-range multimodal graphs to encode the inter-modal relationship between audio and visual data. We cast active speaker detection as a node classification task that is aware of longer-term dependencies. We first construct a graph from a video so that each node corresponds to one person. Nodes representing the same identity share edges between them within a defined temporal window. Nodes within the same video frame are also connected to encode inter-person interactions. Through extensive experiments on the Ava-ActiveSpeaker dataset, we demonstrate that learning graph-based representation, owing to its explicit spatial and temporal structure, significantly improves the overall performance. SPELL outperforms several relevant baselines and performs at par with state of the art models while requiring an order of magnitude lower computation cost.

📄 PDF Abstract BibTeX arXiv:2112.01479

Code (0)

등록된 구현이 없습니다.

Tasks

Active Speaker DetectionAudio-Visual Active Speaker DetectionNode Classification

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection

2022-07-15 · Kyle Min, Sourya Roy, Subarna Tripathi, Tanaya Guha 외

Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we…

Active Speaker DetectionAudio-Visual Active Speaker DetectionGraph LearningNode Classification

Active Speakers in Context

2020-05-20 · CVPR 2020 6 · Juan Leon Alcazar, Fabian Caba Heilbron, Long Mai, Federico Perazzi 외

Current methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker. Although this strategy can be enough for addressing single-speaker scenarios, it prevents accurate…

Active Speaker DetectionAudio-Visual Active Speaker Detection

How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild

2021-06-07 · ICCV 2021 10 · Okan Köpüklü, Maja Taseska, Gerhard Rigoll

Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers wi…

Active Speaker DetectionAudio-Visual Active Speaker Detection

Cross-modal Supervision for Learning Active Speaker Detection in Video

2016-03-29 · Punarjay Chakravarty, Tinne Tuytelaars

In this paper, we show how to use audio to supervise the learning of active speaker detection in video. Voice Activity Detection (VAD) guides the learning of the vision-based classifier in a weakly supervised manner. The…

Action DetectionActive Speaker DetectionActivity Detection

MAAS: Multi-modal Assignation for Active Speaker Detection

2021-01-11 · ICCV 2021 10 · Juan León-Alcázar, Fabian Caba Heilbron, Ali Thabet, Bernard Ghanem

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual featu…

Active Speaker DetectionAudio-Visual Active Speaker Detection