paper-with-me

홈 › Papers

A New Dataset for Natural Language Inference from Code-mixed Conversations

2020-04-10 · LREC 2020 5 · Simran Khanuja, Sandipan Dandapat, Sunayana Sitaram, Monojit Choudhury

Natural Language Inference (NLI) is the task of inferring the logical relationship, typically entailment or contradiction, between a premise and hypothesis. Code-mixing is the use of more than one language in the same conversation or utterance, and is prevalent in multilingual communities all over the world. In this paper, we present the first dataset for code-mixed NLI, in which both the premises and hypotheses are in code-mixed Hindi-English. We use data from Hindi movies (Bollywood) as premises, and crowd-source hypotheses from Hindi-English bilinguals. We conduct a pilot annotation study and describe the final annotation protocol based on observations from the pilot. Currently, the data collected consists of 400 premises in the form of code-mixed conversation snippets and 2240 code-mixed hypotheses. We conduct an extensive analysis to infer the linguistic phenomena commonly observed in the dataset obtained. We evaluate the dataset using a standard mBERT-based pipeline for NLI and report results.

📄 PDF Abstract BibTeX arXiv:2004.05051

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Inference

Similar Papers 제목 키워드 기반

Detecting Entailment in Code-Mixed Hindi-English Conversations

2020-11-01 · EMNLP (WNUT) 2020 11 · Sharanya Chakravarthy, Anjana Umapathy, Alan W Black

The presence of large-scale corpora for Natural Language Inference (NLI) has spurred deep learning research in this area, though much of this research has focused solely on monolingual data. Code-mixing is the intertwine…

Data AugmentationLanguage ModelingLanguage ModellingNatural Language Inference+2

Translate and Classify: Improving Sequence Level Classification for English-Hindi Code-Mixed Data

2021-06-01 · NAACL (CALCS) 2021 6 · Devansh Gautam, Kshitij Gupta, Manish Shrivastava

Code-mixing is a common phenomenon in multilingual societies around the world and is especially common in social media texts. Traditional NLP systems, usually trained on monolingual corpora, do not perform well on code-m…

Machine TranslationNatural Language InferenceSentiment AnalysisTransfer Learning

Multi-turn Inference Matching Network for Natural Language Inference

2019-01-08 · Chunhua Liu, Shan Jiang, Hainan Yu, Dong Yu

Natural Language Inference (NLI) is a fundamental and challenging task in Natural Language Processing (NLP). Most existing methods only apply one-pass inference process on a mixed matching feature, which is a concatenati…

Natural Language Inference

Small Wins Big: Comparing Large Language Models and Domain Fine-Tuned Models for Sarcasm Detection in Code-Mixed Hinglish Text

2026-02-25 · Bitan Majumder, Anirban Sen arxiv

Sarcasm detection in multilingual and code-mixed environments remains a challenging task for natural language processing models due to structural variations, informal expressions, and low-resource linguistic availability…

Sarcasm Detection

Natural Language Inference with Mixed Effects

2020-10-20 · Joint Conference on Lexical and Computational Semantics 2020 · William Gantt, Benjamin Kane, Aaron Steven White

There is growing evidence that the prevalence of disagreement in the raw annotations used to construct natural language inference datasets makes the common practice of aggregating those annotations to a single label prob…

Natural Language Inference