paper-with-me

Papers

When CTC Training Meets Acoustic Landmarks

2018-11-05 · Di He, Xuesong Yang, Boon Pang Lim, Yi Liang, Mark Hasegawa-Johnson, Deming Chen

Connectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but sometimes fails to converge, especially in resource-constrained scenarios. In this paper, the convergence properties of CTC are improved by incorporating acoustic landmarks. We tailored a new set of acoustic landmarks to help CTC training converge more rapidly and smoothly while also reducing recognition error rates. We leveraged new target label sequences mixed with both phone and manner changes to guide CTC training. Experiments on TIMIT demonstrated that CTC based acoustic models converge significantly faster and smoother when they are augmented by acoustic landmarks. The models pretrained with mixed target labels can be further finetuned, resulting in phone error rates 8.72% below baseline on TIMIT. Consistent performance gain is also observed on WSJ (a larger corpus) and reduced TIMIT (smaller). With WSJ, we are the first to succeed in verifying the effectiveness of acoustic landmark theory on a mid-sized ASR task.

📄 PDF Abstract BibTeX arXiv:1811.02063

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognition (ASR)Speech Recognition

Similar Papers 제목 키워드 기반

When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

2024-02-17 · Xiangyu Zhang, Hexin Liu, Kaishuai Xu, Qiquan Zhang 외

Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in menta…

Depression Detection

Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction

2024-09-12 · Xiangyu Zhang, Daijiao Liu, Tianyi Xiao, Cihan Xiao 외

In the speech signal, acoustic landmarks identify times when the acoustic manifestations of the linguistically motivated distinctive features are most salient. Acoustic landmarks have been widely applied in various domai…

Depression Detectionspeech-recognitionSpeech Recognition

Generating Talking Face Landmarks from Speech

2018-03-26 · Sefik Emre Eskimez, Ross K. Maddox, Chenliang Xu, Zhiyao Duan

The presence of a corresponding talking face has been shown to significantly improve speech intelligibility in noisy conditions and for hearing impaired population. In this paper, we present a system that can generate la…

Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks

2018-05-15 · Di He, Boon Pang Lim, Xuesong Yang, Mark Hasegawa-Johnson 외

Furui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cues for speech perception, and that steady…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multi-Task Learningspeech-recognition+1

Detection of Acoustic-Phonetic Landmarks in Mismatched Conditions using a Biomimetic Model of Human Auditory Processing

2012-12-01 · COLING 2012 12 · Sarah King, Mark Hasegawa-Johnson
Speech Recognition