paper-with-me

Papers

Leveraging Visual Supervision for Array-based Active Speaker Detection and Localization

2023-12-21 · Davide Berghi, Philip J. B. Jackson

Conventional audio-visual approaches for active speaker detection (ASD) typically rely on visually pre-extracted face tracks and the corresponding single-channel audio to find the speaker in a video. Therefore, they tend to fail every time the face of the speaker is not visible. We demonstrate that a simple audio convolutional recurrent neural network (CRNN) trained with spatial input features extracted from multichannel audio can perform simultaneous horizontal active speaker detection and localization (ASDL), independently of the visual modality. To address the time and cost of generating ground truth labels to train such a system, we propose a new self-supervised training pipeline that embraces a `student-teacher'' learning approach. A conventional pre-trained active speaker detector is adopted as a teacher'' network to provide the position of the speakers as pseudo-labels. The multichannel audio `student'' network is trained to generate the same results. At inference, the student network can generalize and locate also the occluded speakers that the teacher network is not able to detect visually, yielding considerable improvements in recall rate. Experiments on the TragicTalkers dataset show that an audio network trained with the proposed self-supervised learning approach can exceed the performance of the typical audio-visual methods and produce results competitive with the costly conventional supervised training. We demonstrate that improvements can be achieved when minimal manual supervision is introduced in the learning pipeline. Further gains may be sought with larger training sets and integrating vision with the multichannel audio system.

📄 PDF Abstract BibTeX arXiv:2312.14021

Code (1)

dberghi/leveraging-visual-supervision-for-array-based-asdl 공식 구현 pytorch

Tasks

Active Speaker DetectionSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Visually Supervised Speaker Detection and Localization via Microphone Array

2022-03-07 · Davide Berghi, Adrian Hilton, Philip J. B. Jackson

Active speaker detection (ASD) is a multi-modal task that aims to identify who, if anyone, is speaking from a set of candidates. Current audio-visual approaches for ASD typically rely on visually pre-extracted face track…

Active Speaker Detection

Audio Inputs for Active Speaker Detection and Localization via Microphone Array

2023-07-27 · Davide Berghi, Philip J. B. Jackson

This study considers the problem of detecting and locating an active talker's horizontal position from multichannel audio captured by a microphone array. We refer to this as active speaker detection and localization (ASD…

Active Speaker Detection

Cross-modal Supervision for Learning Active Speaker Detection in Video

2016-03-29 · Punarjay Chakravarty, Tinne Tuytelaars

In this paper, we show how to use audio to supervise the learning of active speaker detection in video. Voice Activity Detection (VAD) guides the learning of the vision-based classifier in a weakly supervised manner. The…

Action DetectionActive Speaker DetectionActivity Detection

Array Configuration-Agnostic Personalized Speech Enhancement using Long-Short-Term Spatial Coherence

2022-11-16 · Yicheng Hsu, Yonghan Lee, Mingsian R. Bai

Personalized speech enhancement has been a field of active research for suppression of speechlike interferers such as competing speakers or TV dialogues. Compared with single channel approaches, multichannel PSE systems …

Speech Enhancement

Cross modal video representations for weakly supervised active speaker localization

2020-03-09 · Rahul Sharma, Krishna Somandepalli, Shrikanth Narayanan

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, …

Action DetectionActive Speaker LocalizationActivity DetectionEvent Detection