paper-with-me

홈 › Papers

CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization

2026-03-17 · Liangbin Huang, Xiaohua Liao, Chaoqun Cui, Shijing Wang, Zhaolong Huang, Yanlong Du, Wenji Mao arxiv

Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker diarization, we extend this task to the visual media domain, encompassing complex audiovisual programs such as films and TV series. This new setting introduces several challenges, including long-form video understanding, a large number of speakers, cross-modal asynchrony between audio and visual cues, and uncontrolled in-the-wild variability. To address these challenges, we propose Cinematic Speaker Registration & Diarization (CineSRD), a unified multimodal framework that leverages visual, acoustic, and linguistic cues from video, speech, and subtitles for speaker annotation. CineSRD first performs visual anchor clustering to register initial speakers and then integrates an audio language model for speaker turn detection, refining annotations and supplementing unregistered off-screen speakers. Furthermore, we construct and release a dedicated speaker diarization benchmark for visual media that includes Chinese and English programs. Experimental results demonstrate that CineSRD achieves superior performance on the proposed benchmark and competitive results on conventional datasets, validating its robustness and generalizability in open-world visual media settings.

📄 PDF Abstract BibTeX arXiv:2603.16966

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Diarization

Similar Papers 제목 키워드 기반

Teach me with a Whisper: Enhancing Large Language Models for Analyzing Spoken Transcripts using Speech Embeddings

2023-11-13 · Fatema Hasan, Yulong Li, James Foulds, SHimei Pan 외

Speech data has rich acoustic and paralinguistic information with important cues for understanding a speaker's tone, emotion, and intent, yet traditional large language models such as BERT do not incorporate this informa…

Knowledge DistillationLanguage ModelingLanguage Modelling

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

2025-05-28 · Long-Khanh Pham, Thanh V. T. Tran, Minh-Tan Pham, Van Nguyen

Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we …

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

2026-05-30 · Zhou Yang, Yueyi Yang arxiv

Face-to-face speech comprehension is inherently multimodal, integrating acoustic signals with visible articulation, facial expression, head motion, and other socially relevant cues. While audiovisual speech systems typic…

MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

2025-07-14 · Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi 외 arxiv

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech an…

Empirical Interpretation of Speech Emotion Perception with Attention Based Model for Speech Emotion Recognition

2020-10-28 · Interspeech 2020 10 · Md AsifJalal, Rosanna Milner, Thomas Hain Speech

Speech emotion recognition is essential for obtaining emotional intelligence which affects the understanding of context and meaning of speech. Harmonically structured vowel and consonant sounds add indexical and linguist…

Emotional IntelligenceEmotion RecognitionSpeech Emotion Recognition