paper-with-me

홈 › Papers

MHE: Code-Mixed Corpora for Similar Language Identification

2022-06-01 · LREC 2022 6 · Priya Rani, John P. McCrae, Theodorus Fransen

This paper introduces a new Magahi-Hindi-English (MHE) code-mixed data-set for similar language identification (SMLID), where Magahi is a less-resourced minority language. This corpus provides a language id at two levels: word and sentence. This data-set is the first Magahi-Hindi-English code-mixed data-set for similar language identification task. Furthermore, we will discuss the complexity of the data-set and provide a few baselines for the language identification task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationSentence

Similar Papers 제목 키워드 기반

OffTamil@DravideanLangTech-EASL2021: Offensive Language Identification in Tamil Text

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Disne Sivalingam, Sajeetha Thavareesan

In the last few decades, Code-Mixed Offensive texts are used penetratingly in social media posts. Social media platforms and online communities showed much interest on offensive text identification in recent years. Conse…

feature selectionLanguage Identification

A Simple and Efficient Probabilistic Language model for Code-Mixed Text

2021-06-29 · M Zeeshan Ansari, Tanvir Ahmad, M M Sufyan Beg, Asma Ikram

The conventional natural language processing approaches are not accustomed to the social media text due to colloquial discourse and non-homogeneous characteristics. Significantly, the language identification in a multili…

Information RetrievalLanguage IdentificationLanguage ModelingLanguage Modelling+6

My Boli: Code-mixed Marathi-English Corpora, Pretrained Language Models and Evaluation Benchmarks

2023-06-24 · Tanmay Chavan, Omkar Gokhale, Aditya Kane, Shantanu Patankar 외

The research on code-mixed data is limited due to the unavailability of dedicated code-mixed datasets and pre-trained language models. In this work, we focus on the low-resource Indian language Marathi which lacks any pr…

BenchmarkingHate Speech DetectionLanguage IdentificationSentiment Analysis

A Fast, Compact, Accurate Model for Language Identification of Codemixed Text

2018-10-09 · EMNLP 2018 10 · Yuan Zhang, Jason Riesa, Daniel Gillick, Anton Bakalov 외

We address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages. Such text is prevalent online, in documents, …

DecoderLanguage IdentificationSentence

OffMix-3L: A Novel Code-Mixed Dataset in Bangla-English-Hindi for Offensive Language Identification

2023-10-27 · Dhiman Goswami, Md Nishat Raihan, Antara Mahmud, Antonios Anastasopoulos 외

Code-mixing is a well-studied linguistic phenomenon when two or more languages are mixed in text or speech. Several works have been conducted on building datasets and performing downstream NLP tasks on code-mixed data. A…

Language Identification