paper-with-me

홈 › Papers

BiPhone: Modeling Inter Language Phonetic Influences in Text

2023-07-06 · Abhirut Gupta, Ananya B. Sai, Richard Sproat, Yuri Vasilevski, James S. Ren, Ambarish Jash, Sukhdeep S. Sodhi, Aravindan Raghuveer

A large number of people are forced to use the Web in a language they have low literacy in due to technology asymmetries. Written text in the second language (L2) from such users often contains a large number of errors that are influenced by their native language (L1). We propose a method to mine phoneme confusions (sounds in L2 that an L1 speaker is likely to conflate) for pairs of L1 and L2. These confusions are then plugged into a generative model (Bi-Phone) for synthetically producing corrupted L2 text. Through human evaluations, we show that Bi-Phone generates plausible corruptions that differ across L1s and also have widespread coverage on the Web. We also corrupt the popular language understanding benchmark SuperGLUE with our technique (FunGLUE for Phonetically Noised GLUE) and show that SoTA language understating models perform poorly. We also introduce a new phoneme prediction pre-training task which helps byte models to recover performance close to SuperGLUE. Finally, we also release the FunGLUE benchmark to promote further research in phonetically robust language models. To the best of our knowledge, FunGLUE is the first benchmark to introduce L1-L2 interactions in text.

📄 PDF Abstract BibTeX arXiv:2307.03322

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language-universal phonetic encoder for low-resource speech recognition

2023-05-19 · Siyuan Feng, Ming Tu, Rui Xia, Chuanzeng Huang 외

Multilingual training is effective in improving low-resource ASR, which may partially be explained by phonetic representation sharing between languages. In end-to-end (E2E) ASR systems, graphemes are often used as basic …

Decoderspeech-recognitionSpeech Recognition

Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective

2024-12-22 · Hankun Wang, Haoran Wang, Yiwei Guo, Zhihan Li 외

Although text-based large language models exhibit human-level writing ability and remarkable intelligence, speech language models (SLMs) still struggle to generate semantically coherent outputs. There are several potenti…

text-to-speechText to Speech

Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis

2024-06-04 · Kun Zhou, Shengkui Zhao, Yukun Ma, Chong Zhang 외

Recent language model-based text-to-speech (TTS) frameworks demonstrate scalability and in-context learning capabilities. However, they suffer from robustness issues due to the accumulation of errors in speech unit predi…

In-Context LearningLanguage ModelingLanguage ModellingSpeech Synthesis+3

PDAF: A Phonetic Debiasing Attention Framework For Speaker Verification

2024-09-09 · Massa Baali, Abdulhamid Aldoobi, Hira Dhamyal, Rita Singh 외

Speaker verification systems are crucial for authenticating identity through voice. Traditionally, these systems focus on comparing feature vectors, overlooking the speech's content. However, this paper challenges this b…

Speaker Verification

Predicting non-native speech perception using the Perceptual Assimilation Model and state-of-the-art acoustic models

2022-05-31 · CoNLL (EMNLP) 2021 11 · Juliette Millet, Ioana Chitoran, Ewan Dunbar

Our native language influences the way we perceive speech sounds, affecting our ability to discriminate non-native sounds. We compare two ideas about the influence of the native language on speech perception: the Percept…