paper-with-me

Papers

MUTANT: A Multi-sentential Code-mixed Hinglish Dataset

2023-02-23 · Rahul Gupta, Vivek Srivastava, Mayank Singh

The multi-sentential long sequence textual data unfolds several interesting research directions pertaining to natural language processing and generation. Though we observe several high-quality long-sequence datasets for English and other monolingual languages, there is no significant effort in building such resources for code-mixed languages such as Hinglish (code-mixing of Hindi-English). In this paper, we propose a novel task of identifying multi-sentential code-mixed text (MCT) from multilingual articles. As a use case, we leverage multilingual articles from two different data sources and build a first-of-its-kind multi-sentential code-mixed Hinglish dataset i.e., MUTANT. We propose a token-level language-aware pipeline and extend the existing metrics measuring the degree of code-mixing to a multi-sentential framework and automatically identify MCT in the multilingual articles. The MUTANT dataset comprises 67k articles with 85k identified Hinglish MCTs. To facilitate future research, we make the publicly available.

📄 PDF Abstract BibTeX arXiv:2302.11766

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

BITS Pilani at HinglishEval: Quality Evaluation for Code-Mixed Hinglish Text Using Transformers

2022-06-17 · Shaz Furniturewala, Vijay Kumari, Amulya Ratna Dash, Hriday Kedia 외

Code-Mixed text data consists of sentences having words or phrases from more than one language. Most multi-lingual communities worldwide communicate using multiple languages, with English usually one of them. Hinglish is…

Gui at MixMT 2022 : English-Hinglish: An MT approach for translation of code mixed data

2022-10-21 · Akshat Gahoi, Jayant Duneja, Anshul Padhi, Shivam Mangale 외

Code-mixed machine translation has become an important task in multilingual communities and extending the task of machine translation to code mixed data has become a common task for these languages. In the shared tasks o…

Machine TranslationTranslationTransliteration

Hinglish to English Machine Translation using Multilingual Transformers

2021-09-01 · RANLP 2021 9 · Vibhav Agarwal, Pooja Rao, Dinesh Babu Jayagopi

Code-Mixed language plays a very important role in communication in multilingual societies and with the recent increase in internet users especially in multilingual societies, the usage of such mixed language has also in…

Machine TranslationTranslation

HinGE: A Dataset for Generation and Evaluation of Code-Mixed Hinglish Text

2021-07-08 · EMNLP (Eval4NLP) 2021 11 · Vivek Srivastava, Mayank Singh

Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the year…

Text Generation

Quality Evaluation of the Low-Resource Synthetically Generated Code-Mixed Hinglish Text

2021-08-04 · INLG (ACL) 2021 8 · Vivek Srivastava, Mayank Singh

In this shared task, we seek the participating teams to investigate the factors influencing the quality of the code-mixed text generation systems. We synthetically generate code-mixed Hinglish sentences using two distinc…

PredictionText Generation