paper-with-me

Papers

Unsupervised active speaker detection in media content using cross-modal information

2022-09-24 · Rahul Sharma, Shrikanth Narayanan

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial images. We leverage speaker identity information from speech and faces, and formulate active speaker detection as a speech-face assignment task such that the active speaker's face and the underlying speech identify the same person (character). We express the speech segments in terms of their associated speaker identity distances, from all other speech segments, to capture a relative identity structure for the video. Then we assign an active speaker's face to each speech segment from the concurrently appearing faces such that the obtained set of active speaker faces displays a similar relative identity structure. Furthermore, we propose a simple and effective approach to address speech segments where speakers are present off-screen. We evaluate the proposed system on three benchmark datasets -- Visual Person Clustering dataset, AVA-active speaker dataset, and Columbia dataset -- consisting of videos from entertainment and broadcast media, and show competitive performance to state-of-the-art fully supervised methods.

📄 PDF Abstract BibTeX arXiv:2209.11896

Code (1)

rash1993/movie-asd pytorch

Tasks

Active Speaker Detection

Similar Papers 제목 키워드 기반

Using Active Speaker Faces for Diarization in TV shows

2022-03-30 · Rahul Sharma, Shrikanth Narayanan

Speaker diarization is one of the critical components of computational media intelligence as it enables a character-level analysis of story portrayals and media content understanding. Automated audio-based speaker diariz…

Face ClusteringFace Detectionspeaker-diarizationSpeaker Diarization

Cross modal video representations for weakly supervised active speaker localization

2020-03-09 · Rahul Sharma, Krishna Somandepalli, Shrikanth Narayanan

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, …

Action DetectionActive Speaker LocalizationActivity DetectionEvent Detection

Audio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection

2022-12-01 · Rahul Sharma, Shrikanth Narayanan

Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of information to derive such a speech-face r…

Active Speaker DetectionAudio-Visual Active Speaker Detection

Detection and Analysis of Content Creator Collaborations in YouTube Videos using Face- and Speaker-Recognition

2018-07-05 · Moritz Lode, Michael Örtl, Christian Koch, Amr Rizk 외

This work discusses and implements the application of speaker recognition for the detection of collaborations in YouTube videos. CATANA, an existing framework for detection and analysis of YouTube collaborations, is util…

Active Speaker DetectionFace RecognitionSpeaker Recognition

Speaker Diarization of Scripted Audiovisual Content

2023-08-04 · Yogesh Virkar, Brian Thompson, Rohit Paturi, Sundararajan Srinivasan 외

The media localization industry usually requires a verbatim script of the final film or TV production in order to create subtitles or dubbing scripts in a foreign language. In particular, the verbatim script (i.e. as-bro…

speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition