paper-with-me

홈 › Papers

Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis

2024-01-22 · Vinotha R, Hepsiba D, L. D. Vijay Anand, Deepak John Reji

Neural Text-to-speech (TTS) synthesis is a powerful technology that can generate speech using neural networks. One of the most remarkable features of TTS synthesis is its capability to produce speech in the voice of different speakers. This paper introduces voice cloning and speech synthesis https://pypi.org/project/voice-cloning/ an open-source python package for helping speech disorders to communicate more effectively as well as for professionals seeking to integrate voice cloning or speech synthesis capabilities into their projects. This package aims to generate synthetic speech that sounds like the natural voice of an individual, but it does not replace the natural human voice. The architecture of the system comprises a speaker verification system, a synthesizer, a vocoder, and noise reduction. Speaker verification system trained on a varied set of speakers to achieve optimal generalization performance without relying on transcriptions. Synthesizer is trained using both audio and transcriptions that generate Mel spectrogram from a text and vocoder which converts the generated Mel Spectrogram into corresponding audio signal. Then the audio signal is processed by a noise reduction algorithm to eliminate unwanted noise and enhance speech clarity. The performance of synthesized speech from seen and unseen speakers are then evaluated using subjective and objective evaluation such as Mean Opinion Score (MOS), Gross Pitch Error (GPE), and Spectral distortion (SD). The model can create speech in distinct voices by including speaker characteristics that are chosen randomly.

📄 PDF Abstract BibTeX arXiv:2401.11771

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationSpeech Synthesistext-to-speechText to SpeechVoice Cloning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Multi-Head Attention 설명 없음
Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Synthesizer 설명 없음

Similar Papers 제목 키워드 기반

Non-portability of Algorithmic Fairness in India

2020-12-03 · Nithya Sambasivan, Erin Arnesen, Ben Hutchinson, Vinodkumar Prabhakaran

Conventional algorithmic fairness is Western in its sub-groups, values, and optimizations. In this paper, we ask how portable the assumptions of this largely Western take on algorithmic fairness are to a different geo-cu…

FairnessTranslation

SPRING-INX: A Multilingual Indian Language Speech Corpus by SPRING Lab, IIT Madras

2023-10-23 · Nithya R, Malavika S, Jordan F, Arjun Gangwar 외

India is home to a multitude of languages of which 22 languages are recognised by the Indian Constitution as official. Building speech based applications for the Indian population is a difficult problem owing to limited …

An Affective Video Database using Multimedia Content Analysis rated on Indian samples

2022-10-18 · Sudhakar Mishra, Narayanan Srinivasan, Uma Shanker Tiwary

Availability of naturalistic affective stimuli is needed for creating the affective technological solution as well as making progress in affective science. Although a lot of progress in the collection of affective multim…

Development of Text and Speech database for Hindi and Indian English specific to Mobile Communication environment

2012-05-01 · LREC 2012 5 · Shyam Agrawal, Shweta Sinha, Pooja Singh, Jesper Olson

Abstract This paper describes the method and experiences of text and speech data collection in mobile communication in Indian English Hindi. The primary data collection is done in the form of large number of messages as …

Language IdentificationSpeech Recognition

Indian Sign Language Recognition Using Mediapipe Holistic

2023-04-20 · Dr. Velmathi G, Kaushal Goyal

Deaf individuals confront significant communication obstacles on a daily basis. Their inability to hear makes it difficult for them to communicate with those who do not understand sign language. Moreover, it presents dif…

Language ModellingSign Language Recognition