Papers Acoustic Modelling
“Acoustic Modelling” 태그가 달린 논문 49편 · 필터 해제
Language Modelling for Speaker Diarization in Telephonic Interviews
The aim of this paper is to investigate the benefit of combining both language and acoustic modelling for speaker diarization. Although conventional systems only use acoustic features, in some scenarios linguistic data c…
Acoustic ModellingLanguage Modellingspeaker-diarizationSpeaker Diarization+2SPEAR: Receiver-to-Receiver Acoustic Neural Warping Field
We present SPEAR, a continuous receiver-to-receiver acoustic neural warping field for spatial acoustic effects prediction in an acoustic 3D space with a single stationary audio source. Unlike traditional source-to-receiv…
Acoustic ModellingPositionSonoTraceLab - A Raytracing-Based Acoustic Modelling System for Simulating Echolocation Behavior of Bats
Echolocation is the prime sensing modality for many species of bats, who show the intricate ability to perform a plethora of tasks in complex and unstructured environments. Understanding this exceptional feat of sensorim…
Acoustic ModellingAn overview of text-to-speech systems and media applications
Producing synthetic voice, similar to human-like sound, is an emerging novelty of modern interactive media systems. Text-To-Speech (TTS) systems try to generate synthetic and authentic voices via text input. Besides, wel…
Acoustic Modellingtext-to-speechText to SpeechVoice ConversionMatcha-TTS: A fast TTS architecture with conditional flow matching
We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an ODE-based decoder capable of high output q…
Acoustic ModellingDecoderSpeech SynthesisText-To-Speech SynthesisComparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech
Neural text-to-speech systems are often optimized on L1/L2 losses, which make strong assumptions about the distributions of the target data space. Aiming to improve those assumptions, Normalizing Flows and Diffusion Prob…
Acoustic ModellingSpeech Synthesistext-to-speechText to Speech+1Speaker- and Age-Invariant Training for Child Acoustic Modeling Using Adversarial Multi-Task Learning
One of the major challenges in acoustic modelling of child speech is the rapid changes that occur in the children's articulators as they grow up, their differing growth rates and the subsequent high variability in the sa…
Acoustic ModellingMulti-Task Learningspeech-recognitionSpeech RecognitionImpact of Dataset on Acoustic Models for Automatic Speech Recognition
In Automatic Speech Recognition, GMM-HMM had been widely used for acoustic modelling. With the current advancement of deep learning, the Gaussian Mixture Model (GMM) from acoustic models has been replaced with Deep Neura…
Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentation+2Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
The Mandarin Chinese language is known to be strongly influenced by a rich set of regional accents, while Mandarin speech with each accent is quite low resource. Hence, an important task in Mandarin speech recognition is…
Acoustic Modellingspeech-recognitionSpeech RecognitionCommon Phone: A Multilingual Dataset for Robust Acoustic Modelling
Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. …
Acoustic Modellingparameter estimationEnhancing audio quality for expressive Neural Text-to-Speech
Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recordings. However, not all speaking styles …
Acoustic ModellingSpeech Synthesistext-to-speechText to SpeechLow Resource German ASR with Untranscribed Data Spoken by Non-native Children -- INTERSPEECH 2021 Shared Task SPAPL System
This paper describes the SPAPL system for the INTERSPEECH 2021 Challenge: Shared Task on Automatic Speech Recognition for Non-Native Children's Speech in German. ~ 5 hours of transcribed data and ~ 60 hours of untranscri…
Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentation+3End-to-end acoustic modelling for phone recognition of young readers
Automatic recognition systems for child speech are lagging behind those dedicated to adult speech in the race of performance. This phenomenon is due to the high acoustic and linguistic variability present in child speech…
Acoustic ModellingTransfer LearningA comparative study of two-dimensional vocal tract acoustic modeling based on Finite-Difference Time-Domain methods
The two-dimensional (2D) numerical approaches for vocal tract (VT) modelling can afford a better balance between the low computational cost and accurate rendering of acoustic wave propagation. However, they require a hig…
Acoustic ModellingWaDeNet: Wavelet Decomposition based CNN for Speech Processing
Existing speech processing systems consist of different modules, individually optimized for a specific task such as acoustic modelling or feature extraction. In addition to not assuring optimality of the system, the disj…
Acoustic ModellingEmotion RecognitionMultilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced Languages
In this work, we explore the benefits of using multilingual bottleneck features (mBNF) in acoustic modelling for the automatic speech recognition of code-switched (CS) speech in African languages. The unavailability of a…
Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1Semi-supervised Acoustic Modelling for Five-lingual Code-switched ASR using Automatically-segmented Soap Opera Speech
This paper considers the impact of automatic segmentation on the fully-automatic, semi-supervised training of automatic speech recog-nition (ASR) systems for five-lingual code-switched (CS) speech. Four automatic segment…
Acoustic ModellingAction DetectionActivity DetectionSegmentation+2Fully Convolutional ASR for Less-Resourced Endangered Languages
The application of deep learning to automatic speech recognition (ASR) has yielded dramatic accuracy increases for languages with abundant training data, but languages with limited training resources have yet to see accu…
Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1Semi-supervised acoustic modelling for five-lingual code-switched ASR using automatically-segmented soap opera speech
This paper considers the impact of automatic segmentation on the fully-automatic, semi-supervised training of automatic speech recognition (ASR) systems for five-lingual code-switched (CS) speech. Four automatic segmenta…
Acoustic ModellingAction DetectionActivity DetectionAutomatic Speech Recognition+6Cross lingual transfer learning for zero-resource domain adaptation
We propose a method for zero-resource domain adaptation of DNN acoustic models, for use in low-resource situations where the only in-language training data available may be poorly matched to the intended target domain. O…
Acoustic ModellingCross-Lingual TransferDomain AdaptationTransfer Learning