paper-with-me

Papers

Random Utterance Concatenation Based Data Augmentation for Improving Short-video Speech Recognition

2022-10-28 · Yist Y. Lin, Tao Han, HaiHua Xu, Van Tung Pham, Yerbolat Khassanov, Tze Yuang Chong, Yi He, Lu Lu, Zejun Ma

One of limitations in end-to-end automatic speech recognition (ASR) framework is its performance would be compromised if train-test utterance lengths are mismatched. In this paper, we propose an on-the-fly random utterance concatenation (RUC) based data augmentation method to alleviate train-test utterance length mismatch issue for short-video ASR task. Specifically, we are motivated by observations that our human-transcribed training utterances tend to be much shorter for short-video spontaneous speech (~3 seconds on average), while our test utterance generated from voice activity detection front-end is much longer (~10 seconds on average). Such a mismatch can lead to suboptimal performance. Empirically, it's observed the proposed RUC method significantly improves long utterance recognition without performance drop on short one. Overall, it achieves 5.72% word error rate reduction on average for 15 languages and improved robustness to various utterance length.

📄 PDF Abstract BibTeX arXiv:2210.15876

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

CopyPaste: An Augmentation Method for Speech Emotion Recognition

2020-10-27 · Raghavendra Pappagari, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velazquez 외

Data augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recognition (SER), where collecting data is expen…

Data AugmentationEmotion RecognitionSpeaker RecognitionSpeech Emotion Recognition+1

Tokenization vs. Augmentation: A Systematic Study of Writer Variance in IMU-Based Online Handwriting Recognition

2026-02-25 · Jindong Li, Dario Zanca, Vincent Christlein, Tim Hamann 외 arxiv

Inertial measurement unit-based online handwriting recognition enables the recognition of input signals collected across different writing surfaces but remains challenged by uneven character distributions and inter-write…

Handwriting RecognitionData Augmentation

Multimodal Sentiment Analysis using Hierarchical Fusion with Context Modeling

2018-06-16 · N. Majumder, D. Hazarika, A. Gelbukh, E. Cambria 외

Multimodal sentiment analysis is a very actively growing field of research. A promising area of opportunity in this field is to improve the multimodal fusion mechanism. We present a novel feature fusion strategy that pro…

Multimodal Emotion RecognitionMultimodal Sentiment AnalysisSentiment Analysis

Data Augmentation for Multiclass Utterance Classification -- A Systematic Study

2020-12-01 · COLING 2020 8 · Binxia Xu, Siyuan Qiu, Jie Zhang, Yafang Wang 외

Utterance classification is a key component in many conversational systems. However, classifying real-world user utterances is challenging, as people may express their ideas and thoughts in manifold ways, and the amount …

ClassificationData Augmentationtext-classificationText Classification+1

Vocal Tract Length Warped Features for Spoken Keyword Spotting

2025-01-07 · Achintya kr. Sarkar, Priyanka Dwivedi, Zheng-Hua Tan

In this paper, we propose several methods that incorporate vocal tract length (VTL) warped features for spoken keyword spotting (KWS). The first method, VTL-independent KWS, involves training a single deep neural network…

Keyword Spotting