A Hindi-English Code-Switching Corpus
The aim of this paper is to investigate the rules and constraints of code-switching (CS) in Hindi-English mixed language data. In this paper, weÂ’ll discuss how we collected the mixed language corpus. This corpus is primarily made up of student interview speech. The speech was manually transcribed and verified by bilingual speakers of Hindi and English. The code-switching cases in the corpus are discussed and the reasons for code-switching are explained.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Hindi-English Code-Switching Speech Corpus
Code-switching refers to the usage of two languages within a sentence or discourse. It is a global phenomenon among multilingual communities and has emerged as an independent area of research. With the increasing demand …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage Modeling+5Investigating Target Set Reduction for End-to-End Speech Recognition of Hindi-English Code-Switching Data
End-to-end (E2E) systems are fast replacing the conventional systems in the domain of automatic speech recognition. As the target labels are learned directly from speech data, the E2E systems need a bigger corpus for eff…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionGupShup: An Annotated Corpus for Abstractive Summarization of Open-Domain Code-Switched Conversations
Code-switching is the communication phenomenon where speakers switch between different languages during a conversation. With the widespread adoption of conversational agents and chat platforms, code-switching has become …
Abstractive Text SummarizationConversation SummarizationLanguage Modeling for Code-Switched Data: Challenges and Approaches
Lately, the problem of code-switching has gained a lot of attention and has emerged as an active area of research. In bilingual communities, the speakers commonly embed the words and phrases of a non-native language into…
Language ModelingLanguage ModellingPOSLanguage Identification and Analysis of Code-Switched Social Media Text
In this paper, we detail our work on comparing different word-level language identification systems for code-switched Hindi-English data and a standard Spanish-English dataset. In this regard, we build a new code-switche…
Language IdentificationMachine Translation