paper-with-me

Papers

Vakyansh: ASR Toolkit for Low Resource Indic languages

2022-03-30 · Harveen Singh Chadha, Anirudh Gupta, Priyanshi Shah, Neeraj Chhimwal, Ankur Dhuriya, Rishabh Gaur, Vivek Raghavan

We present Vakyansh, an end to end toolkit for Speech Recognition in Indic languages. India is home to almost 121 languages and around 125 crore speakers. Yet most of the languages are low resource in terms of data and pretrained models. Through Vakyansh, we introduce automatic data pipelines for data creation, model training, model evaluation and deployment. We create 14,000 hours of speech data in 23 Indic languages and train wav2vec 2.0 based pretrained models. These pretrained models are then finetuned to create state of the art speech recognition models for 18 Indic languages which are followed by language models and punctuation restoration models. We open source all these resources with a mission that this will inspire the speech community to develop speech first applications using our ASR models in Indic languages.

📄 PDF Abstract BibTeX arXiv:2203.16512

Code (1)

Open-Speech-EkStep/vakyansh-models 공식 구현

Tasks

Punctuation Restorationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

SANTLR: Speech Annotation Toolkit for Low Resource Languages

2019-08-02 · Xinjian Li, Zhong Zhou, Siddharth Dalmia, Alan W. black 외

While low resource speech recognition has attracted a lot of attention from the speech community, there are a few tools available to facilitate low resource speech collection. In this work, we present SANTLR: Speech Anno…

speech-recognitionSpeech Recognition

Speaker Recognition in the Wild

2022-05-05 · Neeraj Chhimwal, Anirudh Gupta, Rishabh Gaur, Harveen Singh Chadha 외

In this paper, we propose a pipeline to find the number of speakers, as well as audios belonging to each of these now identified speakers in a source of audio data where number of speakers or speaker labels are not known…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Recognitionspeech-recognition+1

A Stylometry Toolkit for Latin Literature

2019-11-01 · IJCNLP 2019 11 · Thomas J. Bolt, Jeffrey H. Flynt, Pramit Chaudhuri, Joseph P. Dexter

Computational stylometry has become an increasingly important aspect of literary criticism, but many humanists lack the technical expertise or language-specific NLP resources required to exploit computational methods. We…

AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies

2024-02-20 · José-M. Acosta-Triana, David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

More than 7,000 known languages are spoken around the world. However, due to the lack of annotated resources, only a small fraction of them are currently covered by speech technologies. Albeit self-supervised speech repr…

Active Speaker Detection

Generating Inflectional Errors for Grammatical Error Correction in Hindi

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Ankur Sonawane, Sujeet Kumar Vishwakarma, Bhavana Srivastava, Anil Kumar Singh

Automated grammatical error correction has been explored as an important research problem within NLP, with the majority of the work being done on English and similar resource-rich languages. Grammar correction using neur…

Grammatical Error Correction