paper-with-me

홈 › Papers

Target Speech Diarization with Multimodal Prompts

2024-06-11 · Yidi Jiang, Ruijie Tao, Zhengyang Chen, Yanmin Qian, Haizhou Li

Traditional speaker diarization seeks to detect `who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect `when target event occurs'' according to the semantic characteristics of speech. We propose a novel Multimodal Target Speech Diarization (MM-TSD) framework, which accommodates diverse and multi-modal prompts to specify target events in a flexible and user-friendly manner, including semantic language description, pre-enrolled speech, pre-registered face image, and audio-language logical prompts. We further propose a voice-face aligner module to project human voice and face representation into a shared space. We develop a multi-modal dataset based on VoxCeleb2 for MM-TSD training and evaluation. Additionally, we conduct comparative analysis and ablation studies for each category of prompts to validate the efficacy of each component in the proposed framework. Furthermore, our framework demonstrates versatility in performing various signal processing tasks, including speaker diarization and overlap speech detection, using task-specific prompts. MM-TSD achieves robust and comparable performance as a unified system compared to specialized models. Moreover, MM-TSD shows capability to handle complex conversations for real-world dataset.

📄 PDF Abstract BibTeX arXiv:2406.07198

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

Prompt-driven Target Speech Diarization

2023-10-23 · Yidi Jiang, Zhengyang Chen, Ruijie Tao, Liqun Deng 외

We introduce a novel task named `target speech diarization', which seeks to determine `when target event occurred' within an audio signal. We devise a neural architecture called Prompt-driven Target Speech Diarization (P…

Action DetectionActivity Detection

Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization

2024-01-16 · Ming Cheng, Ming Li

Audio-visual learning has demonstrated promising results in many classical speech tasks (e.g., speech separation, automatic speech recognition, wake-word spotting). We believe that introducing visual modality will also b…

Action DetectionActivity Detectionaudio-visual learningAutomatic Speech Recognition+5

Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling

2025-08-08 · Md Asif Jalal, Luca Remaggi, Vasileios Moschopoulos, Thanasis Kotsiopoulos 외 arxiv

Traditional speech separation and speaker diarization approaches rely on prior knowledge of target speakers or a predetermined number of participants in audio signals. To address these limitations, recent advances focus …

Speaker DiarizationSpeech Separation

Simultaneous Speech Recognition and Speaker Diarization for Monaural Dialogue Recordings with Target-Speaker Acoustic Models

2019-09-17 · Naoyuki Kanda, Shota Horiguchi, Yusuke Fujita, Yawen Xue 외

This paper investigates the use of target-speaker automatic speech recognition (TS-ASR) for simultaneous speech recognition and speaker diarization of single-channel dialogue recordings. TS-ASR is a technique to automati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeaker-diarization+3

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

2025-05-20 · Ming Gao, Shilong Wu, Hang Chen, Jun Du 외

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on mul…

Audio-Visual Speech Recognitionspeaker-diarizationSpeaker Diarizationspeech-recognition+2