paper-with-me

Papers

A Twitter Corpus for Hindi-English Code Mixed POS Tagging

2018-07-01 · WS 2018 7 · Kushagra Singh, Indira Sen, Ponnurangam Kumaraguru

Code-mixing is a linguistic phenomenon where multiple languages are used in the same occurrence that is increasingly common in multilingual societies. Code-mixed content on social media is also on the rise, prompting the need for tools to automatically understand such content. Automatic Parts-of-Speech (POS) tagging is an essential step in any Natural Language Processing (NLP) pipeline, but there is a lack of annotated data to train such models. In this work, we present a unique language tagged and POS-tagged dataset of code-mixed English-Hindi tweets related to five incidents in India that led to a lot of Twitter activity. Our dataset is unique in two dimensions: (i) it is larger than previous annotated datasets and (ii) it closely resembles typical real-world tweets. Additionally, we present a POS tagging model that is trained on this dataset to provide an example of how this dataset can be used. The model also shows the efficacy of our dataset in enabling the creation of code-mixed social media POS taggers.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

POSPOS Tagging

Similar Papers 제목 키워드 기반

Corpus Creation and Emotion Prediction for Hindi-English Code-Mixed Social Media Text

2018-06-01 · NAACL 2018 6 · Deepanshu Vijay, Aditya Bohra, Vinay Singh, Syed Sarfaraz Akhtar 외

Emotion Prediction is a Natural Language Processing (NLP) task dealing with detection and classification of emotions in various monolingual and bilingual texts. While some work has been done on code-mixed social media te…

General Classification

L3Cube-HingCorpus and HingBERT: A Code Mixed Hindi-English Dataset and BERT Language Models

2022-04-18 · WILDRE (LREC) 2022 6 · Ravindra Nayak, Raviraj Joshi

Code-switching occurs when more than one language is mixed in a given sentence or a conversation. This phenomenon is more prominent on social media platforms and its adoption is increasing over time. Therefore code-mixed…

Language IdentificationLanguage ModellingNERPOS+3

A Comparative Study of Different State-of-the-Art Hate Speech Detection Methods in Hindi-English Code-Mixed Data

2020-05-01 · LREC 2020 5 · Priya Rani, Shardul Suryawanshi, Koustava Goswami, Bharathi Raja Chakravarthi 외

Hate speech detection in social media communication has become one of the primary concerns to avoid conflicts and curb undesired activities. In an environment where multilingual speakers switch among multiple languages, …

Hate Speech Detection

Gender Prediction in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System

2018-06-14 · Ankush Khandelwal, Sahil Swami, Syed Sarfaraz Akhtar, Manish Shrivastava

The rapid expansion in the usage of social media networking sites leads to a huge amount of unprocessed user generated data which can be used for text mining. Author profiling is the problem of automatically determining …

Author ProfilingGender PredictionGeneral ClassificationLanguage Identification+3

A Hindi-English Code-Switching Corpus

2014-05-01 · LREC 2014 5 · Anik Dey, Pascale Fung

The aim of this paper is to investigate the rules and constraints of code-switching (CS) in Hindi-English mixed language data. In this paper, weÂ’ll discuss how we collected the mixed language corpus. This corpus is prim…