paper-with-me

홈 › Papers

Order Matters in the Presence of Dataset Imbalance for Multilingual Learning

2023-12-11 · NeurIPS 2023 11 · Dami Choi, Derrick Xin, Hamid Dadkhahi, Justin Gilmer, Ankush Garg, Orhan Firat, Chih-Kuan Yeh, Andrew M. Dai, Behrooz Ghorbani

In this paper, we empirically study the optimization dynamics of multi-task learning, particularly focusing on those that govern a collection of tasks with significant data imbalance. We present a simple yet effective method of pre-training on high-resource tasks, followed by fine-tuning on a mixture of high/low-resource tasks. We provide a thorough empirical study and analysis of this method's benefits showing that it achieves consistent improvements relative to the performance trade-off profile of standard static weighting. We analyze under what data regimes this method is applicable and show its improvements empirically in neural machine translation (NMT) and multi-lingual language modeling.

📄 PDF Abstract BibTeX arXiv:2312.06134

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMachine TranslationMulti-Task LearningNMTTranslation

Similar Papers 제목 키워드 기반

Prompt Balance Matters: Understanding How Imbalanced Few-Shot Learning Affects Multilingual Sense Disambiguation in LLMs

2025-10-04 · Deshan Sumanathilaka, Nicholas Micallef, Julian Hough arxiv

Recent advances in Large Language Models (LLMs) have significantly reshaped the landscape of Natural Language Processing (NLP). Among the various prompting techniques, few-shot prompting has gained considerable attention…

Word Sense DisambiguationFew-Shot Learning

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models

2025-08-25 · Yuchun Fan, Yilin Wang, Yongyu Mu, Lei Huang 외 arxiv

Large vision-language models (LVLMs) have demonstrated exceptional capabilities in understanding visual information with human languages but also exhibit an imbalance in multilingual capabilities. In this work, we delve …

Visual Reasoning

Alleviating Sequence Information Loss with Data Overlapping and Prime Batch Sizes

2019-09-18 · CONLL 2019 11 · Noémien Kocher, Christian Scuito, Lorenzo Tarantino, Alexandros Lazaridis 외

In sequence modeling tasks the token order matters, but this information can be partially lost due to the discretization of the sequence into data points. In this paper, we study the imbalance between the way certain tok…

Language Modelling

Compositional Evaluation on Japanese Textual Entailment and Similarity

2022-08-09 · Hitomi Yanaka, Koji Mineshima

Natural Language Inference (NLI) and Semantic Textual Similarity (STS) are widely used benchmark tasks for compositional evaluation of pre-trained language models. Despite growing interest in linguistic universals, most …

Natural Language InferenceSemantic Textual SimilaritySTS

Local Structure Matters Most in Most Languages

2022-11-09 · Louis Clouâtre, Prasanna Parthasarathi, Amal Zouaq, Sarath Chandar

Many recent perturbation studies have found unintuitive results on what does and does not matter when performing Natural Language Understanding (NLU) tasks in English. Coding properties, such as the order of words, can o…

Natural Language Understanding