paper-with-me

Papers

Scenario Aware Speech Recognition: Advancements for Apollo Fearless Steps & CHiME-4 Corpora

2021-09-23 · Szu-Jui Chen, Wei Xia, John H. L. Hansen

In this study, we propose to investigate triplet loss for the purpose of an alternative feature representation for ASR. We consider a general non-semantic speech representation, which is trained with a self-supervised criteria based on triplet loss called TRILL, for acoustic modeling to represent the acoustic characteristics of each audio. This strategy is then applied to the CHiME-4 corpus and CRSS-UTDallas Fearless Steps Corpus, with emphasis on the 100-hour challenge corpus which consists of 5 selected NASA Apollo-11 channels. An analysis of the extracted embeddings provides the foundation needed to characterize training utterances into distinct groups based on acoustic distinguishing properties. Moreover, we also demonstrate that triplet-loss based embedding performs better than i-Vector in acoustic modeling, confirming that the triplet loss is more effective than a speaker feature. With additional techniques such as pronunciation and silence probability modeling, plus multi-style training, we achieve a +5.42% and +3.18% relative WER improvement for the development and evaluation sets of the Fearless Steps Corpus. To explore generalization, we further test the same technique on the 1 channel track of CHiME-4 and observe a +11.90% relative WER improvement for real test data.

📄 PDF Abstract BibTeX arXiv:2109.11086

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionTriplet

Methods 이 논문이 사용한 방법론

Test 설명 없음
Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…

Similar Papers 제목 키워드 기반

Apollo: Band-sequence Modeling for High-Quality Audio Restoration

2024-09-13 · Kai Li, Yi Luo

Audio restoration has become increasingly significant in modern society, not only due to the demand for high-quality auditory experiences enabled by advanced playback devices, but also because the growing capabilities of…

Computational EfficiencySpeech Enhancement

Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models

2024-03-31 · Alkis Koudounas, Flavio Giobergia

The Fearless Steps APOLLO Community Resource provides unparalleled opportunities to explore the potential of multi-speaker team communications from NASA Apollo missions. This study focuses on discovering the characterist…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+1

Fearless Steps APOLLO: Advanced Naturalistic Corpora Development

2022-06-01 · NIDCP (LREC) 2022 6 · John H.L. Hansen, Aditya Joglekar, Szu-Jui Chen, Meena Chandra Shekar 외

In this study, we present the Fearless Steps APOLLO Community Resource, a collection of audio and corresponding meta-data diarized from the NASA Apollo Missions. Massive naturalistic speech data which is time-synchronize…

"This is Houston. Say again, please". The Behavox system for the Apollo-11 Fearless Steps Challenge (phase II)

2020-08-04 · Arseniy Gorin, Daniil Kulko, Steven Grima, Alex Glasman

We describe the speech activity detection (SAD), speaker diarization (SD), and automatic speech recognition (ASR) experiments conducted by the Behavox team for the Interspeech 2020 Fearless Steps Challenge (FSC-2). A rel…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+4

RescueSpeech: A German Corpus for Speech Recognition in Search and Rescue Domain

2023-06-06 · Sangeet Sagar, Mirco Ravanelli, Bernd Kiefer, Ivana Kruijff Korbayova 외

Despite the recent advancements in speech recognition, there are still difficulties in accurately transcribing conversational and emotional speech in noisy and reverberant acoustic environments. This poses a particular c…

Decision MakingRobust Speech Recognitionspeech-recognitionSpeech Recognition