paper-with-me

Papers

Acoustics Based Intent Recognition Using Discovered Phonetic Units for Low Resource Languages

2020-11-07 · Akshat Gupta, Xinjian Li, Sai Krishna Rallabandi, Alan W Black

With recent advancements in language technologies, humans are now speaking to devices. Increasing the reach of spoken language technologies requires building systems in local languages. A major bottleneck here are the underlying data-intensive parts that make up such systems, including automatic speech recognition (ASR) systems that require large amounts of labelled data. With the aim of aiding development of spoken dialog systems in low resourced languages, we propose a novel acoustics based intent recognition system that uses discovered phonetic units for intent classification. The system is made up of two blocks - the first block is a universal phone recognition system that generates a transcript of discovered phonetic units for the input audio, and the second block performs intent classification from the generated phonetic transcripts. We propose a CNN+LSTM based architecture and present results for two languages families - Indic languages and Romance languages, for two different intent recognition tasks. We also perform multilingual training of our intent classifier and show improved cross-lingual transfer and zero-shot performance on an unknown language within the same language family.

📄 PDF Abstract BibTeX arXiv:2011.03646

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-Lingual Transferintent-classificationIntent ClassificationIntent Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Acoustic absement in detail: Quantifying acoustic differences across time-series representations of speech data

2023-04-12 · Matthew C. Kelley

The speech signal is a consummate example of time-series data. The acoustics of the signal change over time, sometimes dramatically. Yet, the most common type of comparison we perform in phonetics is between instantaneou…

Dynamic Time Warpingspeech-recognitionSpeech RecognitionTime Series

The Role of Phonetic Units in Speech Emotion Recognition

2021-08-02 · Jiahong Yuan, Xingyu Cai, Renjie Zheng, Liang Huang 외

We propose a method for emotion recognition through emotiondependent speech recognition using Wav2vec 2.0. Our method achieved a significant improvement over most previously reported results on IEMOCAP, a benchmark emoti…

Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

2025-05-29 · Sean Foley, Hong Nguyen, JIhwan Lee, Sudarsana Reddy Kadiri 외

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on…

Phoneme Recognition

Self-supervised speech unit discovery from articulatory and acoustic features using VQ-VAE

2022-06-17 · Marc-Antoine Georges, Jean-Luc Schwartz, Thomas Hueber

The human perception system is often assumed to recruit motor knowledge when processing auditory speech inputs. Using articulatory modeling and deep learning, this study examines how this articulatory information can be …

Phonetic-assisted Multi-Target Units Modeling for Improving Conformer-Transducer ASR system

2022-11-03 · Li Li, Dongxing Xu, Haoran Wei, Yanhua Long

Exploiting effective target modeling units is very important and has always been a concern in end-to-end automatic speech recognition (ASR). In this work, we propose a phonetic-assisted multi target units (PMU) modeling …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation Learningspeech-recognition+1