paper-with-me

홈 › Papers

Neural text normalization leveraging similarities of strings and sounds

2020-11-04 · COLING 2020 8 · Riku Kawamura, Tatsuya Aoki, Hidetaka Kamigaito, Hiroya Takamura, Manabu Okumura

We propose neural models that can normalize text by considering the similarities of word strings and sounds. We experimentally compared a model that considers the similarities of both word strings and sounds, a model that considers only the similarity of word strings or of sounds, and a model without the similarities as a baseline. Results showed that leveraging the word string similarity succeeded in dealing with misspellings and abbreviations, and taking into account the sound similarity succeeded in dealing with phonetic substitutions and emphasized characters. So that the proposed models achieved higher F$_1$ scores than the baseline.

📄 PDF Abstract BibTeX arXiv:2011.02173

Code (0)

등록된 구현이 없습니다.

Tasks

Text Normalization

Similar Papers 제목 키워드 기반

Combining a Context Aware Neural Network with a Denoising Autoencoder for Measuring String Similarities

2018-07-16 · Mehdi Ben Lazreg, Morten Goodwin

Measuring similarities between strings is central for many established and fast growing research areas including information retrieval, biology, and natural language processing. The traditional approach for string simila…

DenoisingInformation RetrievalRetrieval

Pre-training with Aspect-Content Text Mutual Prediction for Multi-Aspect Dense Retrieval

2023-08-22 · Xiaojie Sun, Keping Bi, Jiafeng Guo, Xinyu Ma 외

Grounded on pre-trained language models (PLMs), dense retrieval has been studied extensively on plain text. In contrast, there has been little research on retrieving data with multiple aspects using dense models. In the …

Language ModelingLanguage ModellingMasked Language ModelingRetrieval

A Clustering Framework for Lexical Normalization of Roman Urdu

2020-03-31 · Abdul Rafae Khan, Asim Karim, Hassan Sajjad, Faisal Kamiran 외

Roman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content. It lacks standard spelling and hence poses several normalization challenges duri…

ClusteringLexical Normalization

What Do I Hear? Generating Sounds for Visuals with ChatGPT

2023-11-09 · David Chuan-En Lin, Nikolas Martelaro

This short paper introduces a workflow for generating realistic soundscapes for visual media. In contrast to prior work, which primarily focus on matching sounds for on-screen visuals, our approach extends to suggesting …

Lung Sound Classification Using Co-tuning and Stochastic Normalization

2021-08-04 · Truc Nguyen, Franz Pernkopf

In this paper, we use pre-trained ResNet models as backbone architectures for classification of adventitious lung sounds and respiratory diseases. The knowledge of the pre-trained model is transferred by using vanilla fi…

Audio ClassificationData AugmentationLung Sound ClassificationSound Classification