paper-with-me

Papers

Building a Public Domain Voice Database for Odia

2022-08-16 · WWW '22: Companion Proceedings of the Web Conference 2022 8 · Subhashish Panigrahi

Projects like Mozilla Common Voice were born to address the challenges of unavailability of voice data or the high cost of available data for use in speech technology such as Automatic Speech Recognition (ASR) research and application development. The pilot detailed in this paper is about creating a large freely-licensed public repository of transcribed speech in the Odia language as such a repository was not known to be available. The strategy and methodology behind this process are based on the OpenSpeaks project. Licensed under a Public Domain Dedication (CC0 1.0), the repository currently includes audio recordings of pronunciations for more than 55,000 unique words in Odia, including more than 5,600 recordings of words in the northern Odia dialect Baleswari. No known public listing of words in this dialect was found by the author prior to this pilot. This repository is arguably the most extensive transcribed speech corpus in Odia that is also available publicly under any free and open license. This paper details the strategy, approach, and process behind building both the text and the speech corpus using many open source tools such as Lingua Libre, which can be helpful in building text and speech data for different low-medium-resource languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech

2020-05-01 · LREC 2020 5 · Adriana Guevara-Rukoz, Isin Demirsahin, Fei He, Shan-Hui Cathy Chu 외

In this paper we present a multidialectal corpus approach for building a text-to-speech voice for a new dialect in a language with existing resources, focusing on various South American dialects of Spanish. We first pres…

text-to-speechText to Speech

Building a Llama2-finetuned LLM for Odia Language Utilizing Domain Knowledge Instruction Set

2023-12-19 · Guneet Singh Kohli, Shantipriya Parida, Sambit Sekhar, Samirit Saha 외

Building LLMs for languages other than English is in great demand due to the unavailability and performance of multilingual LLMs, such as understanding the local context. The problem is critical for low-resource language…

Universal Dependency Treebank for Odia Language

2022-05-24 · WILDRE (LREC) 2022 6 · Shantipriya Parida, Kalyanamalini Sahoo, Atul Kr. Ojha, Saraswati Sahoo 외

This paper presents the first publicly available treebank of Odia, a morphologically rich low resource Indian language. The treebank contains approx. 1082 tokens (100 sentences) in Odia selected from "Samantar", the larg…

BIG-bench Machine LearningMorphological Analysis

OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation

2020-05-01 · LREC 2020 5 · Shantipriya Parida, Satya Ranjan Dash, Ond{\v{r}}ej Bojar, Petr Motlicek 외

The preparation of parallel corpora is a challenging task, particularly for languages that suffer from under-representation in the digital world. In a multi-lingual country like India, the need for such parallel corpora …

Machine TranslationNMTOptical Character RecognitionOptical Character Recognition (OCR)+1

Annotated Corpus for Sentiment Analysis in Odia Language

2020-05-01 · LREC 2020 5 · Gaurav Mohanty, Pruthwik Mishra, Radhika Mamidi

Given the lack of an annotated corpus of non-traditional Odia literature which serves as the standard when it comes sentiment analysis, we have created an annotated corpus of Odia sentences and made it publicly available…

SentenceSentiment Analysis