paper-with-me

Papers

Language Agnostic Code-Mixing Data Augmentation by Predicting Linguistic Patterns

2022-11-14 · Shuyue Stella Li, Kenton Murray

In this work, we focus on intrasentential code-mixing and propose several different Synthetic Code-Mixing (SCM) data augmentation methods that outperform the baseline on downstream sentiment analysis tasks across various amounts of labeled gold data. Most importantly, our proposed methods demonstrate that strategically replacing parts of sentences in the matrix language with a constant mask significantly improves classification accuracy, motivating further linguistic insights into the phenomenon of code-mixing. We test our data augmentation method in a variety of low-resource and cross-lingual settings, reaching up to a relative improvement of 7.73% on the extremely scarce English-Malayalam dataset. We conclude that the code-switch pattern in code-mixing sentences is also important for the model to learn. Finally, we propose a language-agnostic SCM algorithm that is cheap yet extremely helpful for low-resource languages.

📄 PDF Abstract BibTeX arXiv:2211.07628

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationSentiment Analysis

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Language-Driven Dual Style Mixing for Single-Domain Generalized Object Detection

2025-05-12 · Hongda Qin, Xiao Lu, Zhiyong Wei, Yihong Cao 외

Generalizing an object detector trained on a single domain to multiple unseen domains is a challenging task. Existing methods typically introduce image or feature augmentation to diversify the source domain to raise the …

Domain GeneralizationImage Augmentationobject-detectionObject Detection

MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentatio

2026-06-25 · Teerath Kumar, Raja Vavekanand, Muhammad Turab arxiv

Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification. Conventional augmentation can improve training diversity but may distort diagnostically i…

Medical Image ClassificationImage Augmentation

PreCogIIITH at HinglishEval : Leveraging Code-Mixing Metrics & Language Model Embeddings To Estimate Code-Mix Quality

2022-06-16 · Prashant Kodali, Tanmay Sachan, Akshay Goindani, Anmol Goel 외

Code-Mixing is a phenomenon of mixing two or more languages in a speech event and is prevalent in multilingual societies. Given the low-resource nature of Code-Mixing, machine generation of code-mixed text is a prevalent…

Data AugmentationLanguage ModelingLanguage Modelling

Multilingual Controlled Generation And Gold-Standard-Agnostic Evaluation of Code-Mixed Sentences

2024-10-14 · Ayushman Gupta, Akhil Bhogal, Kripabandhu Ghosh

Code-mixing, the practice of alternating between two or more languages in an utterance, is a common phenomenon in multilingual communities. Due to the colloquial nature of code-mixing, there is no singular correct way to…

SentenceText Generation

HSMix: Hard and Soft Mixing Data Augmentation for Medical Image Segmentation

2025-11-18 · Danyang Sun, Fadi Dornaika, Nagore Barrena arxiv

Due to the high cost of annotation or the rarity of some diseases, medical image segmentation is often limited by data scarcity and the resulting overfitting problem. Self-supervised learning and semi-supervised learning…

Medical Image SegmentationSelf-Supervised LearningSemantic SegmentationData Augmentation