paper-with-me

홈 › Papers

Byakto Speech: Real-time long speech synthesis with convolutional neural network: Transfer learning from English to Bangla

2021-05-31 · Zabir Al Nazi, Sayed Mohammed Tasmimul Huda

Speech synthesis is one of the challenging tasks to automate by deep learning, also being a low-resource language there are very few attempts at Bangla speech synthesis. Most of the existing works can't work with anything other than simple Bangla characters script, very short sentences, etc. This work attempts to solve these problems by introducing Byakta, the first-ever open-source deep learning-based bilingual (Bangla and English) text to a speech synthesis system. A speech recognition model-based automated scoring metric was also proposed to evaluate the performance of a TTS model. We also introduce a test benchmark dataset for Bangla speech synthesis models for evaluating speech quality. The TTS is available at https://github.com/zabir-nabil/bangla-tts

📄 PDF Abstract BibTeX arXiv:2106.03937

Code (1)

zabir-nabil/bangla-tts 공식 구현 tf

Tasks

Deep Learningspeech-recognitionSpeech RecognitionSpeech SynthesisTransfer Learning

Similar Papers 제목 키워드 기반

A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation

2026-06-13 · Yulin Xue, Siqi Ouyang, Lei Li arxiv

Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech rather than long-form, continuous input. P…

Speech-to-Speech TranslationSpeech Recognition

Incremental Machine Speech Chain Towards Enabling Listening while Speaking in Real-time

2020-11-04 · Sashi Novitasari, Andros Tjandra, Tomoya Yanagita, Sakriani Sakti 외

Inspired by a human speech chain mechanism, a machine speech chain framework based on deep learning was recently proposed for the semi-supervised development of automatic speech recognition (ASR) and text-to-speech synth…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+4

SkiM: Skipping Memory LSTM for Low-Latency Real-Time Continuous Speech Separation

2022-01-26 · Chenda Li, Lei Yang, Weiqin Wang, Yanmin Qian

Continuous speech separation for meeting pre-processing has recently become a focused research topic. Compared to the data in utterance-level speech separation, the meeting-style audio stream lasts longer, has an uncerta…

Speech Separation

TF-MLPNet: Tiny Real-Time Neural Speech Separation

2025-08-05 · Malek Itani, Tuochao Chen, Shyamnath Gollakota arxiv

Speech separation on hearable devices can enable transformative augmented and enhanced hearing capabilities. However, state-of-the-art speech separation networks cannot run in real-time on tiny, low-power neural accelera…

Speech ExtractionSpeech Separation

Long-Form Speech Generation with Spoken Language Models

2024-12-24 · Se Jin Park, Julian Salazar, Aren Jansen, Keisuke Kinoshita 외

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, current spoken language models struggle to generate plaus…

FormLanguage ModelingLanguage Modelling