paper-with-me

Papers

GLUECoS : An Evaluation Benchmark for Code-Switched NLP

2020-04-26 · Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, Monojit Choudhury

Code-switching is the use of more than one language in the same conversation or utterance. Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cross-lingual and multilingual tasks. We present an evaluation benchmark, GLUECoS, for code-switched languages, that spans several NLP tasks in English-Hindi and English-Spanish. Specifically, our evaluation benchmark includes Language Identification from text, POS tagging, Named Entity Recognition, Sentiment Analysis, Question Answering and a new task for code-switching, Natural Language Inference. We present results on all these tasks using cross-lingual word embedding models and multilingual models. In addition, we fine-tune multilingual models on artificially generated code-switched data. Although multilingual models perform significantly better than cross-lingual models, our results show that in most tasks, across both language pairs, multilingual models fine-tuned on code-switched data perform best, showing that multilingual models can be further optimized for code-switching tasks.

📄 PDF Abstract BibTeX arXiv:2004.12376

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language InferencePOSPOS TaggingQuestion AnsweringSentiment Analysis

Similar Papers 제목 키워드 기반

GLUECoS: An Evaluation Benchmark for Code-Switched NLP

2020-07-01 · ACL 2020 6 · Simran Khanuja, D, S apat, ipan 외

Code-switching is the use of more than one language in the same conversation or utterance. Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cros…

Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+5

Translate and Classify: Improving Sequence Level Classification for English-Hindi Code-Mixed Data

2021-06-01 · NAACL (CALCS) 2021 6 · Devansh Gautam, Kshitij Gupta, Manish Shrivastava

Code-mixing is a common phenomenon in multilingual societies around the world and is especially common in social media texts. Traditional NLP systems, usually trained on monolingual corpora, do not perform well on code-m…

Machine TranslationNatural Language InferenceSentiment AnalysisTransfer Learning

Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?

2026-01-12 · Genta Indra Winata, David Anugraha, Patrick Amadeus Irawan, Anirban Das 외 arxiv

Code-switching is a pervasive phenomenon in multilingual communication, yet the robustness of large language models (LLMs) in mixed-language settings remains insufficiently understood. In this work, we present a comprehe…

From Machine Translation to Code-Switching: Generating High-Quality Code-Switched Text

2021-07-14 · ACL 2021 5 · Ishan Tarunesh, Syamantak Kumar, Preethi Jyothi

Generating code-switched text is a problem of growing interest, especially given the scarcity of corpora containing large volumes of real code-switched text. In this work, we adapt a state-of-the-art neural machine trans…

Data AugmentationLanguage ModelingLanguage ModellingMachine Translation+2

CST5: Data Augmentation for Code-Switched Semantic Parsing

2022-11-14 · Anmol Agarwal, Jigar Gupta, Rahul Goel, Shyam Upadhyay 외

Extending semantic parsers to code-switched input has been a challenging problem, primarily due to a lack of supervised training data. In this work, we introduce CST5, a new data augmentation technique that finetunes a T…

Data AugmentationSemantic Parsing