paper-with-me

Papers

SMPOST: Parts of Speech Tagger for Code-Mixed Indic Social Media Text

2017-02-01 · Deepak Gupta, Shubham Tripathi, Asif Ekbal, Pushpak Bhattacharyya

Use of social media has grown dramatically during the last few years. Users follow informal languages in communicating through social media. The language of communication is often mixed in nature, where people transcribe their regional language with English and this technique is found to be extremely popular. Natural language processing (NLP) aims to infer the information from these text where Part-of-Speech (PoS) tagging plays an important role in getting the prosody of the written text. For the task of PoS tagging on Code-Mixed Indian Social Media Text, we develop a supervised system based on Conditional Random Field classifier. In order to tackle the problem effectively, we have focused on extracting rich linguistic features. We participate in three different language pairs, ie. English-Hindi, English-Bengali and English-Telugu on three different social media platforms, Twitter, Facebook & WhatsApp. The proposed system is able to successfully assign coarse as well as fine-grained PoS tag labels for a given a code-mixed sentence. Experiments show that our system is quite generic that shows encouraging performance levels on all the three language pairs in all the domains.

📄 PDF Abstract BibTeX arXiv:1702.00167

Code (1)

stripathi08/pos_cmism 공식 구현

Tasks

Part-Of-Speech TaggingPOSPOS TaggingSentenceTAG

Similar Papers 제목 키워드 기반

Development of POS tagger for English-Bengali Code-Mixed data

2020-07-29 · ICON 2019 12 · Tathagata Raha, Sainik Kumar Mahata, Dipankar Das, Sivaji Bandyopadhyay

Code-mixed texts are widespread nowadays due to the advent of social media. Since these texts combine two languages to formulate a sentence, it gives rise to various research problems related to Natural Language Processi…

POSSentenceTAG

Language Identification and Named Entity Recognition in Hinglish Code Mixed Tweets

2018-07-01 · ACL 2018 7 · Kushagra Singh, Indira Sen, Ponnurangam Kumaraguru

While growing code-mixed content on Online Social Networks(OSN) provides a fertile ground for studying various aspects of code-mixing, the lack of automated text analysis tools render such studies challenging. To meet th…

Abuse DetectionChunkingLanguage Identificationnamed-entity-recognition+5

A Twitter Corpus for Hindi-English Code Mixed POS Tagging

2018-07-01 · WS 2018 7 · Kushagra Singh, Indira Sen, Ponnurangam Kumaraguru

Code-mixing is a linguistic phenomenon where multiple languages are used in the same occurrence that is increasingly common in multilingual societies. Code-mixed content on social media is also on the rise, prompting the…

POSPOS Tagging

“Kanglish alli names!” Named Entity Recognition for Kannada-English Code-Mixed Social Media Data

2022-10-01 · COLING (WNUT) 2022 10 · Sumukh S, Manish Shrivastava

Code-mixing (CM) is a frequently observed phenomenon on social media platforms in multilingual societies such as India. While the increase in code-mixed content on these platforms provides good amount of data for studyin…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2

Part-of-Speech Annotation of English-Assamese code-mixed texts: Two Approaches

2018-08-01 · COLING 2018 8 · Ritesh Kumar, Manas Jyoti Bora

In this paper, we discuss the development of a part-of-speech tagger for English-Assamese code-mixed texts. We provide a comparison of 2 approaches to annotating code-mixed data {--} a) annotation of the texts from the t…