MHE: Code-Mixed Corpora for Similar Language Identification
This paper introduces a new Magahi-Hindi-English (MHE) code-mixed data-set for similar language identification (SMLID), where Magahi is a less-resourced minority language. This corpus provides a language id at two levels: word and sentence. This data-set is the first Magahi-Hindi-English code-mixed data-set for similar language identification task. Furthermore, we will discuss the complexity of the data-set and provide a few baselines for the language identification task.
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationSentenceSimilar Papers 제목 키워드 기반
OffTamil@DravideanLangTech-EASL2021: Offensive Language Identification in Tamil Text
In the last few decades, Code-Mixed Offensive texts are used penetratingly in social media posts. Social media platforms and online communities showed much interest on offensive text identification in recent years. Conse…
feature selectionLanguage IdentificationA Simple and Efficient Probabilistic Language model for Code-Mixed Text
The conventional natural language processing approaches are not accustomed to the social media text due to colloquial discourse and non-homogeneous characteristics. Significantly, the language identification in a multili…
Information RetrievalLanguage IdentificationLanguage ModelingLanguage Modelling+6My Boli: Code-mixed Marathi-English Corpora, Pretrained Language Models and Evaluation Benchmarks
The research on code-mixed data is limited due to the unavailability of dedicated code-mixed datasets and pre-trained language models. In this work, we focus on the low-resource Indian language Marathi which lacks any pr…
BenchmarkingHate Speech DetectionLanguage IdentificationSentiment AnalysisA Fast, Compact, Accurate Model for Language Identification of Codemixed Text
We address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages. Such text is prevalent online, in documents, …
DecoderLanguage IdentificationSentenceOffMix-3L: A Novel Code-Mixed Dataset in Bangla-English-Hindi for Offensive Language Identification
Code-mixing is a well-studied linguistic phenomenon when two or more languages are mixed in text or speech. Several works have been conducted on building datasets and performing downstream NLP tasks on code-mixed data. A…
Language Identification