paper-with-me

홈 › Papers

Speech Diarization and ASR with GMM

2023-07-11 · Aayush Kumar Sharma, Vineet Bhavikatti, Amogh Nidawani, Dr. Siddappaji, Sanath P, Dr Geetishree Mishra

In this research paper, we delve into the topics of Speech Diarization and Automatic Speech Recognition (ASR). Speech diarization involves the separation of individual speakers within an audio stream. By employing the ASR transcript, the diarization process aims to segregate each speaker's utterances, grouping them based on their unique audio characteristics. On the other hand, Automatic Speech Recognition refers to the capability of a machine or program to identify and convert spoken words and phrases into a machine-readable format. In our speech diarization approach, we utilize the Gaussian Mixer Model (GMM) to represent speech segments. The inter-cluster distance is computed based on the GMM parameters, and the distance threshold serves as the stopping criterion. ASR entails the conversion of an unknown speech waveform into a corresponding written transcription. The speech signal is analyzed using synchronized algorithms, taking into account the pitch frequency. Our primary objective typically revolves around developing a model that minimizes the Word Error Rate (WER) metric during speech transcription.

📄 PDF Abstract BibTeX arXiv:2307.05637

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Speaker Diarization with Region Proposal Network

2020-02-14 · Zili Huang, Shinji Watanabe, Yusuke Fujita, Paola Garcia 외

Speaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in vario…

Region Proposalspeaker-diarizationSpeaker Diarization

Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions

2024-06-12 · Anfeng Xu, Kevin Huang, Tiantian Feng, Lue Shen 외

Speech foundation models, trained on vast datasets, have opened unique opportunities in addressing challenging low-resource speech understanding, such as child speech. In this work, we explore the capabilities of speech …

speaker-diarizationSpeaker Diarization

Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition

2023-12-18 · Peng Shen, Xugang Lu, Hisashi Kawai

Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper, to better address these tasks, we first…

speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

The Third DIHARD Diarization Challenge

2020-12-02 · Neville Ryant, Prachi Singh, Venkat Krishnamohan, Rajat Varma 외

DIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise conditions, and conversational domain. Speaker…

speaker-diarizationSpeaker Diarizationvalid

Prompt-driven Target Speech Diarization

2023-10-23 · Yidi Jiang, Zhengyang Chen, Ruijie Tao, Liqun Deng 외

We introduce a novel task named `target speech diarization', which seeks to determine `when target event occurred' within an audio signal. We devise a neural architecture called Prompt-driven Target Speech Diarization (P…

Action DetectionActivity Detection