paper-with-me

홈 › Papers

HinGE: A Dataset for Generation and Evaluation of Code-Mixed Hinglish Text

2021-07-08 · EMNLP (Eval4NLP) 2021 11 · Vivek Srivastava, Mayank Singh

Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this challenge, we present a corpus (HinGE) for a widely popular code-mixed language Hinglish (code-mixing of Hindi and English languages). HinGE has Hinglish sentences generated by humans as well as two rule-based algorithms corresponding to the parallel Hindi-English sentences. In addition, we demonstrate the inefficacy of widely-used evaluation metrics on the code-mixed data. The HinGE dataset will facilitate the progress of natural language generation research in code-mixed languages.

📄 PDF Abstract BibTeX arXiv:2107.03760

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

MIPE: A Metric Independent Pipeline for Effective Code-Mixed NLG Evaluation

2021-07-24 · EMNLP (Eval4NLP) 2021 11 · Ayush Garg, Sammed S Kagi, Vivek Srivastava, Mayank Singh

Code-mixing is a phenomenon of mixing words and phrases from two or more languages in a single utterance of speech and text. Due to the high linguistic diversity, code-mixing presents several challenges in evaluating sta…

Diversitynlg evaluationText Generation

Multilingual Controlled Generation And Gold-Standard-Agnostic Evaluation of Code-Mixed Sentences

2024-10-14 · Ayushman Gupta, Akhil Bhogal, Kripabandhu Ghosh

Code-mixing, the practice of alternating between two or more languages in an utterance, is a common phenomenon in multilingual communities. Due to the colloquial nature of code-mixing, there is no singular correct way to…

SentenceText Generation

Quality Evaluation of the Low-Resource Synthetically Generated Code-Mixed Hinglish Text

2021-08-04 · INLG (ACL) 2021 8 · Vivek Srivastava, Mayank Singh

In this shared task, we seek the participating teams to investigate the factors influencing the quality of the code-mixed text generation systems. We synthetically generate code-mixed Hinglish sentences using two distinc…

PredictionText Generation

Uncovering Code-Mixed Challenges: A Framework for Linguistically Driven Question Generation and Neural Based Question Answering

2018-10-01 · CONLL 2018 10 · Deepak Gupta, Pabitra Lenka, Asif Ekbal, Pushpak Bhattacharyya

Existing research on question answering (QA) and comprehension reading (RC) are mainly focused on the resource-rich language like English. In recent times, the rapid growth of multi-lingual web content has posed several …

Question AnsweringQuestion GenerationQuestion-Generation

Struct-MMSB: Mixed Membership Stochastic Blockmodels with Interpretable Structured Priors

2020-02-21 · Yue Zhang, Arti Ramesh

The mixed membership stochastic blockmodel (MMSB) is a popular framework for community detection and network generation. It learns a low-rank mixed membership representation for each node across communities by exploiting…

Community DetectionProbabilistic ProgrammingRelational Reasoning