paper-with-me

Papers

How Do Language Models Represent and Use Phonological Information for Allomorph Selection?

2026-09-04 · Sangwoo Kim, Sangah Lee arxiv

Language models are trained on tokenized text that obscures the sound structure of words, yet they reliably produce morphemes whose form is phonologically conditioned. It remains unclear whether they rely on item-specific memorization or rule-like generalization and, if the latter, how that generalization is implemented. We therefore ask whether this phonological condition is represented within language models and how it is causally used for allomorph selection. For the English indefinite article a/an, we show that the phonological condition is encoded along a single linear direction in trigger-token embeddings, that this direction causally drives article selection in token-level wug tests, and that, at the article-prediction position, the model forecasts the upcoming trigger token and uses the forecasted trigger's phonological feature to choose the article. We then ask whether this rule-like generalization extends beyond English article selection, both to allomorph selection in other languages and to explicit phonological judgment. Together, these results provide a mechanistic account of phonologically conditioned allomorph selection in language models, and dissociate this generation-time ability from explicit metalinguistic judgments.

📄 PDF Abstract BibTeX arXiv:2609.04708

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 115
arxivsub/arXivSub_daily_arxiv ★ 4

Similar Papers 제목 키워드 기반

Trees probe deeper than strings: an argument from allomorphy

2022-07-01 · NAACL (SIGMORPHON) 2022 7 · Hossep Dolatian, Shiori Ikawa, Thomas Graf

Linguists disagree on whether morphological representations should be strings or trees. We argue that tree-based views of morphology can provide new insights into morphological complexity even in cases where the posited …

Analogy in Contact: Modeling Maltese Plural Inflection

2023-05-20 · Sara Court, Andrea D. Sims, Micha Elsner

Maltese is often described as having a hybrid morphological system resulting from extensive contact between Semitic and Romance language varieties. Such a designation reflects an etymological divide as much as it does a …

Weakly supervised learning of allomorphy

2017-09-01 · WS 2017 9 · Miikka Silfverberg, Mans Hulden

Most NLP resources that offer annotations at the word segment level provide morphological annotation that includes features indicating tense, aspect, modality, gender, case, and other inflectional information. Such infor…

Weakly-supervised Learning

Understanding Compositional Data Augmentation in Typologically Diverse Morphological Inflection

2023-05-23 · Farhan Samir, Miikka Silfverberg

Data augmentation techniques are widely used in low-resource automatic morphological inflection to overcome data sparsity. However, the full implications of these techniques remain poorly understood. In this study, we ai…

AttributeData AugmentationMorphological Inflection

Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis

2025-01-12 · Minu Kim, Kangwook Jang, Hoirin Kim

This paper examines how linguistic similarity affects cross-lingual phonetic representation in speech processing for low-resource languages, emphasizing effective source language selection. Previous cross-lingual researc…

Phoneme RecognitionSelf-Supervised Learning