paper-with-me

Papers Text Normalization

“Text Normalization” 태그가 달린 논문 145편 · 필터 해제

Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization

2025-05-30 · Luong Ho, Khanh Le, Vinh Pham, Bao Nguyen 외

Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usability. Despite its importance, the integra…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

Visualizing Public Opinion on X: A Real-Time Sentiment Dashboard Using VADER and DistilBERT

2025-04-21 · Yanampally Abhiram Reddy, Siddhi Agarwal, Vikram Parashar, Arshiya Arora

In the age of social media, understanding public sentiment toward major corporations is crucial for investors, policymakers, and researchers. This paper presents a comprehensive sentiment analysis system tailored for cor…

Sentiment AnalysisSentiment ClassificationText Normalization

Chain of Correction for Full-text Speech Recognition with Large Language Models

2025-04-02 · Zhiyuan Tang, Dong Wang, Zhikai Zhou, Yong liu 외

Full-text error correction with Large Language Models (LLMs) for Automatic Speech Recognition (ASR) has gained increased attention due to its potential to correct errors across long contexts and address a broader spectru…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Punctuation Restorationspeech-recognition+2

Misspellings in Natural Language Processing: A survey

2025-01-28 · Gianluca Sperduti, Alejandro Moreo

This survey provides an overview of the challenges of misspellings in natural language processing (NLP). While often unintentional, misspellings have become ubiquitous in digital communication, especially with the prolif…

Data AugmentationMachine TranslationSurveytext-classification+2

Universal-2-TF: Robust All-Neural Text Formatting for ASR

2025-01-10 · Yash Khare, Taufiquzzaman Peyash, Andrea Vanzo, Takuya Yoshioka

This paper introduces an all-neural text formatting (TF) model designed for commercial automatic speech recognition (ASR) systems, encompassing punctuation restoration (PR), truecasing, and inverse text normalization (IT…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Computational Efficiency+4

Digestion Algorithm in Hierarchical Symbolic Forests: A Fast Text Normalization Algorithm and Semantic Parsing Framework for Specific Scenarios and Lightweight Deployment

2024-12-18 · Kevin You

Text Normalization and Semantic Parsing have numerous applications in natural language processing, such as natural language programming, paraphrasing, data augmentation, constructing expert systems, text matching, and mo…

Lightweight DeploymentSemantic ParsingText Normalization

Neural Text Normalization for Luxembourgish using Real-Life Variation Data

2024-12-12 · Anne-Marie Lutgen, Alistair Plum, Christoph Purschke, Barbara Plank

Orthographic variation is very common in Luxembourgish texts due to the absence of a fully-fledged standard variety. Additionally, developing NLP tools for Luxembourgish is a difficult task given the lack of annotated an…

Text Normalization

Machine Learning Driven Smishing Detection Framework for Mobile Security

2024-12-09 · Diksha Goel, Hussain Ahmad, Ankit Kumar Jain, Nikhil Kumar Goel

The increasing reliance on smartphones for communication, financial transactions, and personal data management has made them prime targets for cyberattacks, particularly smishing, a sophisticated variant of phishing cond…

ManagementMobile SecurityText Normalization

Hybrid Deep Learning for Legal Text Analysis: Predicting Punishment Durations in Indonesian Court Rulings

2024-10-26 · Muhammad Amien Ibrahim, Alif Tri Handoyo, Maria Susan Anggreainy

Limited public understanding of legal processes and inconsistent verdicts in the Indonesian court system led to widespread dissatisfaction and increased stress on judges. This study addresses these issues by developing a…

Computational EfficiencyDocument SummarizationSentenceText Normalization

WER We Stand: Benchmarking Urdu ASR Models

2024-09-17 · Samee Arif, Sualeha Farid, Aamina Jamal Khan, Mustafa Abbas 외

This paper presents a comprehensive evaluation of Urdu Automatic Speech Recognition (ASR) models. We analyze the performance of three ASR model families: Whisper, MMS, and Seamless-M4T using Word Error Rate (WER), along …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Benchmarkingspeech-recognition+2

Full-text Error Correction for Chinese Speech Recognition with Large Language Model

2024-09-12 · Zhiyuan Tang, Dong Wang, Shen Huang, Shidong Shang

Large Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from short-duration speech recordings, which a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+9

What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations

2024-09-04 · Kavya Manohar, Leena G Pillai, Elizabeth Sherly

This paper explores the pitfalls in evaluating multilingual automatic speech recognition (ASR) models, with a particular focus on Indic language scripts. We investigate the text normalization routine employed by leading …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Historical German Text Normalization Using Type- and Token-Based Language Modeling

2024-09-04 · Anton Ehrmanntraut

Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+2

Is text normalization relevant for classifying medieval charters?

2024-08-29 · Florian Atzenhofer-Baumgartner, Tamás Kovács

This study examines the impact of historical text normalization on the classification of medieval charters, specifically focusing on document dating and locating. Using a data set of Middle High German charters from a di…

Document DatingText Normalization

Positional Description for Numerical Normalization

2024-08-22 · Deepanshu Gupta, Javier Latorre

We present a Positional Description Scheme (PDS) tailored for digit sequences, integrating placeholder value information for each digit. Given the structural limitations of subword tokenization algorithms, language model…

speech-recognitionSpeech RecognitionText Normalizationtext-to-speech+1

The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization

2024-07-23 · Samuele Cornell, Taejin Park, Steve Huang, Christoph Boeddeker 외

This paper presents the CHiME-8 DASR challenge which carries on from the previous edition CHiME-7 DASR (C7DASR) and the past CHiME-6 challenge. It focuses on joint multi-channel distant speech recognition (DASR) and diar…

Automatic Speech RecognitionDistant Speech Recognitionspeech-recognitionSpeech Recognition+1

Exploiting Dialect Identification in Automatic Dialectal Text Normalization

2024-07-03 · Bashar Alhafni, Sarah Al-Towaity, Ziyad Fawzy, Fatema Nassar 외

Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do no…

Dialect IdentificationText Normalization

Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling

2024-04-14 · Quanxiu Wang, Hui Huang, Mingjie Wang, Yong Dai 외

Over the past decade, a series of unflagging efforts have been dedicated to developing highly expressive and controllable text-to-speech (TTS) systems. In general, the holistic TTS comprises two interconnected components…

Polyphone disambiguationText Normalizationtext-to-speechText to Speech

VNLP: Turkish NLP Package

2024-03-02 · Meliksah Turker, Mehmet Erdi Ari, Aydin Han

In this work, we present VNLP: the first dedicated, complete, open-source, well-documented, lightweight, production-ready, state-of-the-art Natural Language Processing (NLP) package for the Turkish language. It contains …

Morphological Analysisnamed-entity-recognitionNamed Entity RecognitionPart-Of-Speech Tagging+6

Multi-Task Learning for Front-End Text Processing in TTS

2024-01-12 · Wonjune Kang, Yun Wang, Shun Zhang, Arthur Hinsvark 외

We propose a multi-task learning (MTL) model for jointly performing three tasks that are commonly solved in a text-to-speech (TTS) front-end: text normalization (TN), part-of-speech (POS) tagging, and homograph disambigu…

Language ModelingLanguage ModellingMulti-Task LearningPart-Of-Speech Tagging+5
1–20 / 145 다음 →