paper-with-me

Papers

Data-driven Model Generalizability in Crosslinguistic Low-resource Morphological Segmentation

2022-01-05 · Zoey Liu, Emily Prud'hommeaux

Common designs of model evaluation typically focus on monolingual settings, where different models are compared according to their performance on a single data set that is assumed to be representative of all possible data for the task at hand. While this may be reasonable for a large data set, this assumption is difficult to maintain in low-resource scenarios, where artifacts of the data collection can yield data sets that are outliers, potentially making conclusions about model performance coincidental. To address these concerns, we investigate model generalizability in crosslinguistic low-resource scenarios. Using morphological segmentation as the test case, we compare three broad classes of models with different parameterizations, taking data from 11 languages across 6 language families. In each experimental setting, we evaluate all models on a first data set, then examine their performance consistency when introducing new randomly sampled data sets with the same size and when applying the trained models to unseen test sets of varying sizes. The results demonstrate that the extent of model generalization depends on the characteristics of the data set, and does not necessarily rely heavily on the data set size. Among the characteristics that we studied, the ratio of morpheme overlap and that of the average number of morphemes per word between the training and test sets are the two most prominent factors. Our findings suggest that future work should adopt random sampling to construct data sets with different sizes in order to make more responsible claims about model evaluation.

📄 PDF Abstract BibTeX arXiv:2201.01845

Code (1)

zoeyliu18/orange_chicken 공식 구현

Similar Papers 제목 키워드 기반

Morphological Segmentation for Seneca

2021-06-01 · NAACL (AmericasNLP) 2021 6 · Zoey Liu, Robert Jimerson, Emily Prud’hommeaux

This study takes up the task of low-resource morphological segmentation for Seneca, a critically endangered and morphologically complex Native American language primarily spoken in what is now New York State and Ontario.…

DecoderModel SelectionMulti-Task LearningSegmentation

A Data-driven Approach to Crosslinguistic Structural Biases

2021-02-01 · SCiL 2021 2 · Alex Kramer, Zoey Liu

The Crosslinguistic Relationship between Ordering Flexibility and Dependency Length Minimization: A Data-Driven Approach

2021-02-01 · SCiL 2021 2 · Zoey Liu

The taggedPBC: Annotating a massive parallel corpus for crosslinguistic investigations

2025-05-18 · Hiram Ring

Existing datasets available for crosslinguistic investigations have tended to focus on large amounts of data for a small group of languages or a small amount of data for a large number of languages. This means that claim…

POS

Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis

2025-05-31 · Miao Zhang, Aref Farhadipour, Annie Baker, Jiachen Ma 외

With its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech technology and holds tremendous potential for research in crosslinguistic ph…

Diversity