paper-with-me

홈 › Papers

Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation

2023-10-18 · Tae Jin Park, He Huang, Coleman Hooper, Nithin Koluguri, Kunal Dhawan, Ante Jukic, Jagadeesh Balam, Boris Ginsburg

We introduce a sophisticated multi-speaker speech data simulator, specifically engineered to generate multi-speaker speech recordings. A notable feature of this simulator is its capacity to modulate the distribution of silence and overlap via the adjustment of statistical parameters. This capability offers a tailored training environment for developing neural models suited for speaker diarization and voice activity detection. The acquisition of substantial datasets for speaker diarization often presents a significant challenge, particularly in multi-speaker scenarios. Furthermore, the precise time stamp annotation of speech data is a critical factor for training both speaker diarization and voice activity detection. Our proposed multi-speaker simulator tackles these problems by generating large-scale audio mixtures that maintain statistical properties closely aligned with the input parameters. We demonstrate that the proposed multi-speaker simulator generates audio mixtures with statistical properties that closely align with the input parameters derived from real-world statistics. Additionally, we present the effectiveness of speaker diarization and voice activity detection models, which have been trained exclusively on the generated simulated datasets.

📄 PDF Abstract BibTeX arXiv:2310.12371

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity Detectionspeaker-diarizationSpeaker DiarizationSynthetic Data Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Enhanced Speaker-aware Multi-party Multi-turn Dialogue Comprehension

2021-09-09 · Xinbei Ma, Zhuosheng Zhang, Hai Zhao

Multi-party multi-turn dialogue comprehension brings unprecedented challenges on handling the complicated scenarios from multiple speakers and criss-crossed discourse relationship among speaker-aware utterances. Most exi…

Question Answering

Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based Chatbots

2020-04-07 · Jia-Chen Gu, Tianda Li, Quan Liu, Zhen-Hua Ling 외

In this paper, we study the problem of employing pre-trained language models for multi-turn response selection in retrieval-based chatbots. A new model, named Speaker-Aware BERT (SA-BERT), is proposed in order to make th…

Conversational Response SelectionDisentanglementDomain AdaptationRetrieval

Speaker-Aware Simulation Improves Conversational Speech Recognition

2026-02-04 · Máté Gedeon, Péter Mihajlik arxiv

Automatic speech recognition (ASR) for conversational speech remains challenging due to the limited availability of large-scale, well-annotated multi-speaker dialogue data and the complex temporal dynamics of natural int…

Dialogue GenerationSpeech RecognitionData Augmentation

Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation

2025-11-21 · Scott Merrill, Shashank Srivastava arxiv

LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for evaluating long-form institutional behavior. ASR transcripts typically use anony…

Speech Recognition

Listening to Multi-talker Conversations: Modular and End-to-end Perspectives

2024-02-14 · Desh Raj

Since the first speech recognition systems were built more than 30 years ago, improvement in voice technology has enabled applications such as smart assistants and automated customer support. However, conversation intell…

GPUspeaker-diarizationSpeaker Diarizationspeech-recognition+2