paper-with-me

Papers

Malayalam Speech Corpus: Design and Development for Dravidian Language

2020-05-01 · LREC 2020 5 · Lekshmi K R, Jithesh V S, Elizabeth Sherly

To overpass the disparity between theory and applications in language-related technology in the text as well as speech and several other areas, a well-designed and well-developed corpus is essential. Several problems and issues encountered while developing a corpus, especially for low resource languages. The Malayalam Speech Corpus (MSC) is one of the first open speech corpora for Automatic Speech Recognition (ASR) research to the best of our knowledge. It consists of 250 hours of Agricultural speech data. We are providing a transcription file, lexicon and annotated speech along with the audio segment. It is available in future for public use upon request at {``}www.iiitmk.ac.in/vrclc/utilities/ml{\_}speechcorpus{''}. This paper details the development and collection process in the domain of agricultural speech corpora in the Malayalam Language.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Gauravarora@HASOC-Dravidian-CodeMix-FIRE2020: Pre-training ULMFiT on Synthetically Generated Code-Mixed Data for Hate Speech Detection

2020-10-05 · Gaurav Arora

This paper describes the system submitted to Dravidian-Codemix-HASOC2020: Hate Speech and Offensive Content Identification in Dravidian languages (Tamil-English and Malayalam-English). The task aims to identify offensive…

Hate Speech Detection

IIITK@LT-EDI-EACL2021: Hope Speech Detection for Equality, Diversity, and Inclusion in Tamil , Malayalam and English

2021-04-19 · Nikhil Ghanghor, Rahul Ponnusamy, Prasanna Kumar Kumaresan, Ruba Priyadharshini 외

This paper describes the IIITK’s team submissions to the hope speech detection for equality, diversity and inclusion in Dravidian languages shared task organized by LT-EDI 2021 workshop@EACL 2021. Our best configurations…

DiversityHope Speech Detection

CUSATNLP@DravidianLangTech-EACL2021:Language Agnostic Classification of Offensive Content in Tweets

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Sara Renjit, Sumam Mary Idicula

Identifying offensive information from tweets is a vital language processing task. This task concentrated more on English and other foreign languages these days. In this shared task on Offensive Language Identification i…

Language IdentificationPositionSentenceSentence Embedding+1

WLV-RIT at HASOC-Dravidian-CodeMix-FIRE2020: Offensive Language Identification in Code-switched YouTube Comments

2020-11-01 · Tharindu Ranasinghe, Sarthak Gupte, Marcos Zampieri, Ifeoma Nwogu

This paper describes the WLV-RIT entry to the Hate Speech and Offensive Content Identification in Indo-European Languages (HASOC) shared task 2020. The HASOC 2020 organizers provided participants with annotated datasets …

Language IdentificationTransfer LearningWord Embeddings

Quantitative Analysis of the Morphological Complexity of Malayalam Language

2020-09-01 · Text, Speech, and Dialogue 2020 9 · Kavya Manohar, A R jayan, Rajeev Rajan

This paper presents a quantitative analysis on the morpho- logical complexity of Malayalam language. Malayalam is a Dravidian language spoken in India, predominantly in the state of Kerala with about 38 million native sp…

Vocal Bursts Type Prediction