paper-with-me

Papers

Translate and Classify: Improving Sequence Level Classification for English-Hindi Code-Mixed Data

2021-06-01 · NAACL (CALCS) 2021 6 · Devansh Gautam, Kshitij Gupta, Manish Shrivastava

Code-mixing is a common phenomenon in multilingual societies around the world and is especially common in social media texts. Traditional NLP systems, usually trained on monolingual corpora, do not perform well on code-mixed texts. Training specialized models for code-switched texts is difficult due to the lack of large-scale datasets. Translating code-mixed data into standard languages like English could improve performance on various code-mixed tasks since we can use transfer learning from state-of-the-art English models for processing the translated data. This paper focuses on two sequence-level classification tasks for English-Hindi code mixed texts, which are part of the GLUECoS benchmark - Natural Language Inference and Sentiment Analysis. We propose using various pre-trained models that have been fine-tuned for similar English-only tasks and have shown state-of-the-art performance. We further fine-tune these models on the translated code-mixed datasets and achieve state-of-the-art performance in both tasks. To translate English-Hindi code-mixed data to English, we use mBART, a pre-trained multilingual sequence-to-sequence model that has shown competitive performance on various low-resource machine translation pairs and has also shown performance gains in languages that were not in its pre-training corpus.

📄 PDF Abstract BibTeX

Code (1)

devanshg27/cm_translatify 공식 구현 pytorch

Tasks

Machine TranslationNatural Language InferenceSentiment AnalysisTransfer Learning

Similar Papers 제목 키워드 기반

Multitask Models for Controlling the Complexity of Neural Machine Translation

2020-07-01 · WS 2020 7 · Sweta Agrawal, Marine Carpuat

We introduce a machine translation task where the output is aimed at audiences of different levels of target language proficiency. We collect a novel dataset of news articles available in English and Spanish and written …

ArticlesMachine TranslationTranslation

Controlling Text Complexity in Neural Machine Translation

2019-11-03 · IJCNLP 2019 11 · Sweta Agrawal, Marine Carpuat

This work introduces a machine translation task where the output is aimed at audiences of different levels of target language proficiency. We collect a high quality dataset of news articles available in English and Spani…

ArticlesMachine TranslationTranslation

Don't Classify, Translate: Multi-Level E-Commerce Product Categorization Via Machine Translation

2018-12-14 · Maggie Yundi Li, Stanley Kok, Liling Tan

E-commerce platforms categorize their products into a multi-level taxonomy tree with thousands of leaf categories. Conventional methods for product categorization are typically based on machine learning classification al…

General ClassificationMachine TranslationProduct CategorizationTranslation

Using Machine Learning to Detect Fraudulent SMSs in Chichewa

2025-02-24 · Amelia Taylor, Amoss Robert

SMS enabled fraud is of great concern globally. Building classifiers based on machine learning for SMS fraud requires the use of suitable datasets for model training and validation. Most research has centred on the use o…

Fraud DetectionMachine Translation

Image Classification for Arabic: Assessing the Accuracy of Direct English to Arabic Translations

2018-07-13 · Abdulkareem Alsudais

Image classification is an ongoing research challenge. Most of the available research focuses on image classification for the English language, however there is very little research on image classification for the Arabic…

ClassificationGeneral Classificationimage-classificationImage Classification+1