paper-with-me

Papers

Understanding Compositional Data Augmentation in Typologically Diverse Morphological Inflection

2023-05-23 · Farhan Samir, Miikka Silfverberg

Data augmentation techniques are widely used in low-resource automatic morphological inflection to overcome data sparsity. However, the full implications of these techniques remain poorly understood. In this study, we aim to shed light on the theoretical aspects of the prominent data augmentation strategy StemCorrupt (Silfverberg et al., 2017; Anastasopoulos and Neubig, 2019), a method that generates synthetic examples by randomly substituting stem characters in gold standard training examples. To begin, we conduct an information-theoretic analysis, arguing that StemCorrupt improves compositional generalization by eliminating spurious correlations between morphemes, specifically between the stem and the affixes. Our theoretical analysis further leads us to study the sample efficiency with which StemCorrupt reduces these spurious correlations. Through evaluation across seven typologically distinct languages, we demonstrate that selecting a subset of datapoints with both high diversity and high predictive uncertainty significantly enhances the data-efficiency of StemCorrupt. However, we also explore the impact of typological features on the choice of the data selection strategy and find that languages incorporating a high degree of allomorphy and phonological alternations derive less benefit from synthetic examples with high uncertainty. We attribute this effect to phonotactic violations induced by StemCorrupt, emphasizing the need for further research to ensure optimal performance across the entire spectrum of natural language morphology.

📄 PDF Abstract BibTeX arXiv:2305.13658

Code (1)

smfsamir/understanding-augmentation-morphology 공식 구현 pytorch

Tasks

AttributeData AugmentationMorphological Inflection

Similar Papers 제목 키워드 기반

Inference-Time Structural Reasoning for Compositional Vision-Language Understanding

2026-03-28 · Amartya Bhattacharya arxiv

Vision-language models (VLMs) excel at image-text retrieval yet persistently fail at compositional reasoning, distinguishing captions that share the same words but differ in relational structure. We present, a unified ev…

Text Retrieval

TreeMix: Compositional Constituency-based Data Augmentation for Natural Language Understanding

2022-05-12 · NAACL 2022 7 · Le Zhang, Zichao Yang, Diyi Yang

Data augmentation is an effective approach to tackle over-fitting. Many previous works have proposed different data augmentations strategies for NLP, such as noise injection, word replacement, back-translation etc. Thoug…

Constituency ParsingData AugmentationDiversityNatural Language Understanding+2

MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

2022-04-18 · Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie 외

We present the MASSIVE dataset--Multilingual Amazon Slu resource package (SLURP) for Slot-filling, Intent classification, and Virtual assistant Evaluation. MASSIVE contains 1M realistic, parallel, labeled virtual assista…

intent-classificationIntent ClassificationNatural Language UnderstandingSlot Filling+3

A Three Step Training Approach with Data Augmentation for Morphological Inflection

2021-09-14 · Gabor Szolnok, Botond Barta, Dorina Lakatos, Judit Acs

We present the BME submission for the SIGMORPHON 2021 Task 0 Part 1, Generalization Across Typologically Diverse Languages shared task. We use an LSTM encoder-decoder model with three step training that is first trained …

Data AugmentationDecoderMorphological Inflection

BME Submission for SIGMORPHON 2021 Shared Task 0. A Three Step Training Approach with Data Augmentation for Morphological Inflection

2021-08-01 · ACL (SIGMORPHON) 2021 8 · Gábor Szolnok, Botond Barta, Dorina Lakatos, Judit Ács

We present the BME submission for the SIGMORPHON 2021 Task 0 Part 1, Generalization Across Typologically Diverse Languages shared task. We use an LSTM encoder-decoder model with three step training that is first trained …

Data AugmentationDecoderMorphological Inflection