paper-with-me

홈 › Papers

Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies

2024-07-18 · Srija Anand, Praveen Srinivasa Varadhan, Ashwin Sankar, Giri Raju, Mitesh M. Khapra

Publicly available TTS datasets for low-resource languages like Hindi and Tamil typically contain 10-20 hours of data, leading to poor vocabulary coverage. This limitation becomes evident in downstream applications where domain-specific vocabulary coupled with frequent code-mixing with English, results in many OOV words. To highlight this problem, we create a benchmark containing OOV words from several real-world applications. Indeed, state-of-the-art Hindi and Tamil TTS systems perform poorly on this OOV benchmark, as indicated by intelligibility tests. To improve the model's OOV performance, we propose a low-effort and economically viable strategy to obtain more training data. Specifically, we propose using volunteers as opposed to high quality voice artists to record words containing character bigrams unseen in the training data. We show that using such inexpensive data, the model's performance improves on OOV words, while not affecting voice quality and in-domain performance.

📄 PDF Abstract BibTeX arXiv:2407.13435

Code (1)

AI4Bharat/IndicOOV 공식 구현

Similar Papers 제목 키워드 기반

IE-CPS Lexicon: An Automatic Speech Recognition Oriented Indian-English Pronunciation Dictionary

2021-12-01 · ICON 2021 12 · Shelly Jain, Aditya Yadavalli, Ganesh Mirishkar, Chiranjeevi Yarra 외

Indian English (IE), on the surface, seems quite similar to standard English. However, closer observation shows that it has actually been influenced by the surrounding vernacular languages at several levels from phonolog…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Svarah: Evaluating English ASR Systems on Indian Accents

2023-05-25 · Tahir Javed, Sakshi Joshi, Vignesh Nagarajan, Sai Sundaresan 외

India is the second largest English-speaking country in the world with a speaker base of roughly 130 million. Thus, it is imperative that automatic speech recognition (ASR) systems for English should be evaluated on Indi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Everyday Speech in the Indian Subcontinent

2024-10-14 · Utkarsh P

India has 1369 languages of which 22 are official. About 13 different scripts are used to represent these languages. A Common Label Set (CLS) was developed based on phonetics to address the issue of large vocabulary of u…

Speech Synthesis

CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems

2025-06-24 · Haochen Zhang, Tianyi Zhang, Junze Yin, Oren Gal 외

Recommender systems play a pivotal role in providing relevant content to users. With the rapid development of large language models (LLMs), researchers have begun utilizing LLMs to build more powerful recommender systems…

Recommendation Systems

Autoencoder-Based Framework to Capture Vocabulary Quality in NLP

2025-02-28 · Vu Minh Hoang Dang, Rakesh M. Verma

Linguistic richness is essential for advancing natural language processing (NLP), as dataset characteristics often directly influence model performance. However, traditional metrics such as Type-Token Ratio (TTR), Vocabu…

DiversitySentence