paper-with-me

홈 › Papers

A2TTS: TTS for Low Resource Indian Languages

2025-07-21 · Ayush Singh Bhadoriya, Abhishek Nikunj Shinde, Isha Pandey, Ganesh Ramakrishnan arxiv

We present a speaker conditioned text-to-speech (TTS) system aimed at addressing challenges in generating speech for unseen speakers and supporting diverse Indian languages. Our method leverages a diffusion-based TTS architecture, where a speaker encoder extracts embeddings from short reference audio samples to condition the DDPM decoder for multispeaker generation. To further enhance prosody and naturalness, we employ a cross-attention based duration prediction mechanism that utilizes reference audio, enabling more accurate and speaker consistent timing. This results in speech that closely resembles the target speaker while improving duration modeling and overall expressiveness. Additionally, to improve zero-shot generation, we employed classifier free guidance, allowing the system to generate speech more near speech for unknown speakers. Using this approach, we trained language-specific speaker-conditioned models. Using the IndicSUPERB dataset for multiple Indian languages such as Bengali, Gujarati, Hindi, Marathi, Malayalam, Punjabi and Tamil.

📄 PDF Abstract BibTeX arXiv:2507.15272

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages

2025-05-27 · Praveen Srinivasa Varadhan, Srija Anand, Soma Siddhartha, Mitesh M. Khapra

What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice-cloning, style-cloning, and code-mixing…

Synthetic Data GenerationVoice Cloning

Language Resources and Technologies for Non-Scheduled and Endangered Indian Languages

2022-04-06 · Ritesh Kumar, Bornini Lahiri

In the present paper, we will present a survey of the language resources and technologies available for the non-scheduled and endangered languages of India. While there have been different estimates from different source…

IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages

2023-12-15 · Saiful Haq, Ashutosh Sharma, Pushpak Bhattacharyya

In this paper, we introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages (Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, and Telugu) from two major…

Information RetrievalMachine TranslationRetrieval

BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages

2024-11-07 · Sparsh Jain, Ashwin Sankar, Devilal Choudhary, Dhairya Suman 외

Automatic Speech Translation (AST) datasets for Indian languages remain critically scarce, with public resources covering fewer than 10 of the 22 official languages. This scarcity has resulted in AST systems for Indian l…

automatic-speech-translationSynthetic Data GenerationTranslation

Model Adaptation for ASR in low-resource Indian Languages

2023-07-16 · Abhayjeet Singh, Arjun Singh Mehta, Ashish Khuraishi K S, Deekshitha G 외

Automatic speech recognition (ASR) performance has improved drastically in recent years, mainly enabled by self-supervised learning (SSL) based acoustic models such as wav2vec2 and large-scale multi-lingual training like…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognition+1