paper-with-me

홈 › Papers

Handling Numeric Expressions in Automatic Speech Recognition

2024-07-18 · Christian Huber, Alexander Waibel

This paper addresses the problem of correctly formatting numeric expressions in automatic speech recognition (ASR) transcripts. This is challenging since the expected transcript format depends on the context, e.g., 1945 (year) vs. 19:45 (timestamp). We compare cascaded and end-to-end approaches to recognize and format numeric expressions such as years, timestamps, currency amounts, and quantities. For the end-to-end approach, we employed a data generation strategy using a large language model (LLM) together with a text to speech (TTS) model to generate adaptation data. The results on our test data set show that while approaches based on LLMs perform well in recognizing formatted numeric expressions, adapted end-to-end models offer competitive performance with the advantage of lower latency and inference cost.

📄 PDF Abstract BibTeX arXiv:2408.00004

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage ModellingLarge Language Modelspeech-recognitionSpeech Recognitiontext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Confidence-Weighted Local Expression Predictions for Occlusion Handling in Expression Recognition and Action Unit detection

2016-07-21 · Arnaud Dapogny, Kévin Bailly, Séverine Dubuisson

Fully-Automatic Facial Expression Recognition (FER) from still images is a challenging task as it involves handling large interpersonal morphological differences, and as partial occlusions can occasionally happen. Furthe…

Action Unit DetectionDescriptiveFacial Expression RecognitionFacial Expression Recognition (FER)+1

A New Benchmark for Evaluating Automatic Speech Recognition in the Arabic Call Domain

2024-03-07 · Qusai Abo Obaidah, Muhy Eddin Za'ter, Adnan Jaljuli, Ali Mahboub 외

This work is an attempt to introduce a comprehensive benchmark for Arabic speech recognition, specifically tailored to address the challenges of telephone conversations in Arabic language. Arabic, characterized by its ri…

Arabic Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversity+2

MoLE : Mixture of Language Experts for Multi-Lingual Automatic Speech Recognition

2023-02-27 · Yoohwan Kwon, Soo-Whan Chung

Multi-lingual speech recognition aims to distinguish linguistic expressions in different languages and integrate acoustic processing simultaneously. In contrast, current multi-lingual speech recognition research follows …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

BTS: Back TranScription for Speech-to-Text Post-Processor using Text-to-Speech-to-Text

2021-08-01 · ACL (WAT) 2021 8 · Chanjun Park, Jaehyung Seo, Seolhwa Lee, Chanhee Lee 외

With the growing popularity of smart speakers, such as Amazon Alexa, speech is becoming one of the most important modes of human-computer interaction. Automatic speech recognition (ASR) is arguably the most critical comp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Denoisingspeech-recognition+4

Handling Trade-Offs in Speech Separation with Sparsely-Gated Mixture of Experts

2022-11-11 · Xiaofei Wang, Zhuo Chen, Yu Shi, Jian Wu 외

Employing a monaural speech separation (SS) model as a front-end for automatic speech recognition (ASR) involves balancing two kinds of trade-offs. First, while a larger model improves the SS performance, it also require…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Mixture-of-Expertsspeech-recognition+2