Neural text normalization leveraging similarities of strings and sounds
We propose neural models that can normalize text by considering the similarities of word strings and sounds. We experimentally compared a model that considers the similarities of both word strings and sounds, a model that considers only the similarity of word strings or of sounds, and a model without the similarities as a baseline. Results showed that leveraging the word string similarity succeeded in dealing with misspellings and abbreviations, and taking into account the sound similarity succeeded in dealing with phonetic substitutions and emphasized characters. So that the proposed models achieved higher F$_1$ scores than the baseline.
Code (0)
등록된 구현이 없습니다.
Tasks
Text NormalizationSimilar Papers 제목 키워드 기반
Combining a Context Aware Neural Network with a Denoising Autoencoder for Measuring String Similarities
Measuring similarities between strings is central for many established and fast growing research areas including information retrieval, biology, and natural language processing. The traditional approach for string simila…
DenoisingInformation RetrievalRetrievalPre-training with Aspect-Content Text Mutual Prediction for Multi-Aspect Dense Retrieval
Grounded on pre-trained language models (PLMs), dense retrieval has been studied extensively on plain text. In contrast, there has been little research on retrieving data with multiple aspects using dense models. In the …
Language ModelingLanguage ModellingMasked Language ModelingRetrievalA Clustering Framework for Lexical Normalization of Roman Urdu
Roman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content. It lacks standard spelling and hence poses several normalization challenges duri…
ClusteringLexical NormalizationWhat Do I Hear? Generating Sounds for Visuals with ChatGPT
This short paper introduces a workflow for generating realistic soundscapes for visual media. In contrast to prior work, which primarily focus on matching sounds for on-screen visuals, our approach extends to suggesting …
Lung Sound Classification Using Co-tuning and Stochastic Normalization
In this paper, we use pre-trained ResNet models as backbone architectures for classification of adventitious lung sounds and respiratory diseases. The knowledge of the pre-trained model is transferred by using vanilla fi…
Audio ClassificationData AugmentationLung Sound ClassificationSound Classification