paper-with-me

Papers

Capitalization Normalization for Language Modeling with an Accurate and Efficient Hierarchical RNN Model

2022-02-16 · Hao Zhang, You-Chi Cheng, Shankar Kumar, W. Ronny Huang, Mingqing Chen, Rajiv Mathews

Capitalization normalization (truecasing) is the task of restoring the correct case (uppercase or lowercase) of noisy text. We propose a fast, accurate and compact two-level hierarchical word-and-character-based recurrent neural network model. We use the truecaser to normalize user-generated text in a Federated Learning framework for language modeling. A case-aware language model trained on this normalized text achieves the same perplexity as a model trained on text with gold capitalization. In a real user A/B experiment, we demonstrate that the improvement translates to reduced prediction error rates in a virtual keyboard application. Similarly, in an ASR language model fusion experiment, we show reduction in uppercase character error rate and word error rate.

📄 PDF Abstract BibTeX arXiv:2202.08171

Code (0)

등록된 구현이 없습니다.

Tasks

Federated LearningLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Four-in-One: A Joint Approach to Inverse Text Normalization, Punctuation, Capitalization, and Disfluency for Automatic Speech Recognition

2022-10-26 · Sharman Tan, Piyush Behre, Nick Kibre, Issac Alphonso 외

Features such as punctuation, capitalization, and formatting of entities are important for readability, understanding, and natural language processing tasks. However, Automatic Speech Recognition (ASR) systems produce sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Formspeech-recognition+2

CL-MoNoise: Cross-lingual Lexical Normalization

2021-11-01 · EMNLP (WNUT) 2021 11 · Rob van der Goot

Social media is notoriously difficult to process for existing natural language processing tools, because of spelling errors, non-standard words, shortenings, non-standard capitalization and punctuation. One method to cir…

Lexical Normalization

Text Injection for Capitalization and Turn-Taking Prediction in Speech Models

2023-08-14 · Shaan Bijwadia, Shuo-Yiin Chang, Weiran Wang, Zhong Meng 외

Text injection for automatic speech recognition (ASR), wherein unpaired text-only data is used to supplement paired audio-text data, has shown promising improvements for word error rate. This study examines the use of te…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Fast and Accurate Capitalization and Punctuation for Automatic Speech Recognition Using Transformer and Chunk Merging

2019-08-07 · Binh Nguyen, Vu Bao Hung Nguyen, Hien Nguyen, Pham Ngoc Phuong 외

In recent years, studies on automatic speech recognition (ASR) have shown outstanding results that reach human parity on short speech segments. However, there are still difficulties in standardizing the output of ASR suc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)NERPOS+4

Bridge-Language Capitalization Inference in Western Iranian: Sorani, Kurmanji, Zazaki, and Tajik

2016-05-01 · LREC 2016 5 · Patrick Littell, David R. Mortensen, Kartik Goyal, Chris Dyer 외

In Sorani Kurdish, one of the most useful orthographic features in named-entity recognition {--} capitalization {--} is absent, as the language{'}s Perso-Arabic script does not make a distinction between uppercase and lo…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)