Unsupervised active speaker detection in media content using cross-modal information
We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial images. We leverage speaker identity information from speech and faces, and formulate active speaker detection as a speech-face assignment task such that the active speaker's face and the underlying speech identify the same person (character). We express the speech segments in terms of their associated speaker identity distances, from all other speech segments, to capture a relative identity structure for the video. Then we assign an active speaker's face to each speech segment from the concurrently appearing faces such that the obtained set of active speaker faces displays a similar relative identity structure. Furthermore, we propose a simple and effective approach to address speech segments where speakers are present off-screen. We evaluate the proposed system on three benchmark datasets -- Visual Person Clustering dataset, AVA-active speaker dataset, and Columbia dataset -- consisting of videos from entertainment and broadcast media, and show competitive performance to state-of-the-art fully supervised methods.
Code (1)
Tasks
Active Speaker DetectionSimilar Papers 제목 키워드 기반
Using Active Speaker Faces for Diarization in TV shows
Speaker diarization is one of the critical components of computational media intelligence as it enables a character-level analysis of story portrayals and media content understanding. Automated audio-based speaker diariz…
Face ClusteringFace Detectionspeaker-diarizationSpeaker DiarizationCross modal video representations for weakly supervised active speaker localization
An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, …
Action DetectionActive Speaker LocalizationActivity DetectionEvent DetectionAudio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection
Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of information to derive such a speech-face r…
Active Speaker DetectionAudio-Visual Active Speaker DetectionDetection and Analysis of Content Creator Collaborations in YouTube Videos using Face- and Speaker-Recognition
This work discusses and implements the application of speaker recognition for the detection of collaborations in YouTube videos. CATANA, an existing framework for detection and analysis of YouTube collaborations, is util…
Active Speaker DetectionFace RecognitionSpeaker RecognitionSpeaker Diarization of Scripted Audiovisual Content
The media localization industry usually requires a verbatim script of the final film or TV production in order to create subtitles or dubbing scripts in a foreign language. In particular, the verbatim script (i.e. as-bro…
speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition