paper-with-me

Papers

CoLI-Machine Learning Approaches for Code-mixed Language Identification at the Word Level in Kannada-English Texts

2022-11-17 · H. L. Shashirekha, F. Balouchzahi, M. D. Anusha, G. Sidorov

The task of automatically identifying a language used in a given text is called Language Identification (LI). India is a multilingual country and many Indians especially youths are comfortable with Hindi and English, in addition to their local languages. Hence, they often use more than one language to post their comments on social media. Texts containing more than one language are called "code-mixed texts" and are a good source of input for LI. Languages in these texts may be mixed at sentence level, word level or even at sub-word level. LI at word level is a sequence labeling problem where each and every word in a sentence is tagged with one of the languages in the predefined set of languages. In order to address word level LI in code-mixed Kannada-English (Kn-En) texts, this work presents i) the construction of code-mixed Kn-En dataset called CoLI-Kenglish dataset, ii) code-mixed Kn-En embedding and iii) learning models using Machine Learning (ML), Deep Learning (DL) and Transfer Learning (TL) approaches. Code-mixed Kn-En texts are extracted from Kannada YouTube video comments to construct CoLI-Kenglish dataset and code-mixed Kn-En embedding. The words in CoLI-Kenglish dataset are grouped into six major categories, namely, "Kannada", "English", "Mixed-language", "Name", "Location" and "Other". The learning models, namely, CoLI-vectors and CoLI-ngrams based on ML, CoLI-BiLSTM based on DL and CoLI-ULMFiT based on TL approaches are built and evaluated using CoLI-Kenglish dataset. The performances of the learning models illustrated, the superiority of CoLI-ngrams model, compared to other models with a macro average F1-score of 0.64. However, the results of all the learning models were quite competitive with each other.

📄 PDF Abstract BibTeX arXiv:2211.09847

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationSentenceTransfer Learning

Similar Papers 제목 키워드 기반

Transformer-based Model for Word Level Language Identification in Code-mixed Kannada-English Texts

2022-11-26 · Atnafu Lambebo Tonja, Mesay Gemeda Yigezu, Olga Kolesnikova, Moein Shahiki Tash 외

Using code-mixed data in natural language processing (NLP) research currently gets a lot of attention. Language identification of social media code-mixed text has been an interesting problem of study in recent years due …

Language Identification

Enabling Code-Mixed Translation: Parallel Corpus Creation and MT Augmentation Approach

2018-08-01 · COLING 2018 8 · Mrinal Dhar, Vaibhav Kumar, Manish Shrivastava

Code-mixing, use of two or more languages in a single sentence, is ubiquitous; generated by multi-lingual speakers across the world. The phenomenon presents itself prominently in social media discourse. Consequently, the…

Machine TranslationSentenceTranslation

Code-Mixed Sentiment Analysis Using Machine Learning and Neural Network Approaches

2018-08-09 · Pruthwik Mishra, Prathyusha Danda, Pranav Dhakras

Sentiment Analysis for Indian Languages (SAIL)-Code Mixed tools contest aimed at identifying the sentence level sentiment polarity of the code-mixed dataset of Indian languages pairs (Hi-En, Ben-Hi-En). Hi-En dataset is …

BIG-bench Machine LearningSentenceSentiment Analysis

Experiments with POS Tagging Code-mixed Indian Social Media Text

2016-10-31 · Prakash B. Pimpale, Raj Nath Patel

This paper presents Centre for Development of Advanced Computing Mumbai's (CDACM) submission to the NLP Tools Contest on Part-Of-Speech (POS) Tagging For Code-mixed Indian Social Media Text (POSCMISMT) 2015 (collocated w…

Part-Of-Speech TaggingPOSPOS TaggingTAG

JNLP Team: Deep Learning Approaches for Legal Processing Tasks in COLIEE 2021

2021-06-25 · Ha-Thanh Nguyen, Phuong Minh Nguyen, Thi-Hai-Yen Vuong, Quan Minh Bui 외

COLIEE is an annual competition in automatic computerized legal text processing. Automatic legal document processing is an ambitious goal, and the structure and semantics of the law are often far more complex than everyd…

Survey