Development of Text and Speech database for Hindi and Indian English specific to Mobile Communication environment
Abstract This paper describes the method and experiences of text and speech data collection in mobile communication in Indian English Hindi. The primary data collection is done in the form of large number of messages as part of Personal communication among natives of Hindi language and Indian speakers of English. To gather the versatility of mobile communication database among Hindi and English, 12 domains were identified for collection of text corpus from speaking population belonging to deferent age groups, sex and dialects. The text obtained in raw form based on slangs and unconventional grammar were cleaned using on language grammar rules and then tagged and expanded to explain context specific meaning of the words. Texts of 1163 participants from Hindi speaking regions and 1405 English users were taken for creating 13 prompt sheets; containing 630 phonetically rich sentences created using a special software. Each prompt sheet was recorded by at least 7 users simultaneously in three channels and recorded by a total of 100 speakers and annotated. The work is a step forward in the direction of development of standards for mobile text and speech data collection for Indian languages. Keywords - Speech data base, Text analysis, mobile communication, Hindi and Indian English Speech, multi-lingual speech processing.
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationSpeech RecognitionSimilar Papers 제목 키워드 기반
Hindi-English Code-Switching Speech Corpus
Code-switching refers to the usage of two languages within a sentence or discourse. It is a global phenomenon among multilingual communities and has emerged as an independent area of research. With the increasing demand …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage Modeling+5Statistical Analysis of Multilingual Text Corpus and Development of Language Models
This paper presents two studies, first a statistical analysis for three languages i.e. Hindi, Punjabi and Nepali and the other, development of language models for three Indian languages i.e. Indian English, Punjabi and N…
Language IdentificationLanguage ModellingSpeech Language IdentificationLeveraging Latent Representations of Speech for Indian Language Identification
Identification of the language spoken from speech utterances is an interesting task because of the diversity associated with different languages and human voices. Indian languages have diverse origins and identifying the…
DiversityFeature EngineeringLanguage IdentificationRepresentation LearningASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare contexts remains largely unknown. In this study, …
Speech RecognitionStatistical Machine Translation for Indian Languages: Mission Hindi 2
This paper presents Centre for Development of Advanced Computing Mumbai's (CDACM) submission to NLP Tools Contest on Statistical Machine Translation in Indian Languages (ILSMT) 2015 (collocated with ICON 2015). The aim o…
Machine TranslationTranslation