paper-with-me

Papers

Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation

2021-07-17 · ACL ARR November 2021 11 · Anonymous

The widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka. code-mixed language) in a single utterance. This has resulted a formidable challenge for the computational models due to the scarcity of annotated data and presence of noise. A potential solution to mitigate the data scarcity problem in low-resource setup is to leverage existing data in resource-rich language through translation. In this paper, we tackle the problem of code-mixed (Hinglish and Bengalish) to English machine translation. First, we synthetically develop HINMIX, a parallel corpus of Hinglish to English, with ~5M sentence pairs. Subsequently, we propose JAMT, a robust perturbation based joint-training model that learns to handle noise in the real-world code-mixed text by parameter sharing across clean and noisy words. Further, we show the adaptability of JAMT in a zero-shot setup for Bengalish to English translation. Our evaluation and comprehensive analyses qualitatively and quantitatively demonstrate the superiority of JAMT over state-of-the-art code-mixed and robust translation methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceSynthetic Data GenerationTranslation

Similar Papers 제목 키워드 기반

Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation

2024-03-25 · Kartik Kartik, Sanjana Soni, Anoop Kunchukuttan, Tanmoy Chakraborty 외

The widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka code-mixed language) in a single utterance. This has resulted a formidable challenge for …

Machine TranslationSentenceSynthetic Data GenerationTranslation

Quality Evaluation of the Low-Resource Synthetically Generated Code-Mixed Hinglish Text

2021-08-04 · INLG (ACL) 2021 8 · Vivek Srivastava, Mayank Singh

In this shared task, we seek the participating teams to investigate the factors influencing the quality of the code-mixed text generation systems. We synthetically generate code-mixed Hinglish sentences using two distinc…

PredictionText Generation

Towards Code-Mixed Hinglish Dialogue Generation

2021-09-01 · RANLP 2021 9 · Vibhav Agarwal, Pooja Rao, Dinesh Babu Jayagopi

Code-mixed language plays a crucial role in communication in multilingual societies. Though the recent growth of web users has greatly boosted the use of such mixed languages, the current generation of dialog systems is …

Dialogue Generation

Automated rock joint trace mapping using a supervised learning model trained on synthetic data generated by parametric modelling

2026-02-07 · Jessica Ka Yi Chiu, Tom Frode Hansen, Eivind Magnus Paulsen, Ole Jakob Mengshoel arxiv

This paper presents a geology-driven machine learning method for automated rock joint trace mapping from images. The approach combines geological modelling, synthetic data generation, and supervised image segmentation to…

Synthetic Data GenerationImage SegmentationDomain Adaptation

Differentially Private Mixed-Type Data Generation For Unsupervised Learning

2019-09-25 · Uthaipon Tantipongpipat, Chris Waites, Digvijay Boob, Amaresh Siva 외

In this work we introduce the DP-auto-GAN framework for synthetic data generation, which combines the low dimensional representation of autoencoders with the flexibility of GANs. This framework can be used to take in ra…

Synthetic Data GenerationVocal Bursts Type Prediction