paper-with-me

Papers

IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling

2025-04-03 · Zébulon Goriely, Paula Buttery

In this paper, we introduce two resources: (i) G2P+, a tool for converting orthographic datasets to a consistent phonemic representation; and (ii) IPA CHILDES, a phonemic dataset of child-centered speech across 31 languages. Prior tools for grapheme-to-phoneme conversion result in phonemic vocabularies that are inconsistent with established phonemic inventories, an issue which G2P+ addresses by leveraging the inventories in the Phoible database. Using this tool, we augment CHILDES with phonemic transcriptions to produce IPA CHILDES. This new resource fills several gaps in existing phonemic datasets, which often lack multilingual coverage, spontaneous speech, and a focus on child-directed language. We demonstrate the utility of this dataset for phonological research by training phoneme language models on 11 languages and probing them for distinctive features, finding that the distributional properties of phonemes are sufficient to learn major class and place features cross-lingually.

📄 PDF Abstract BibTeX arXiv:2504.03036

Code (1)

codebyzeb/g2p-plus 공식 구현

Tasks

Grapheme-to-Phoneme ConversionLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Morphosyntactic Analysis for CHILDES

2024-07-17 · Houjun Liu, Brian MacWhinney

Language development researchers are interested in comparing the process of language learning across languages. Unfortunately, it has been difficult to construct a consistent quantitative framework for such comparisons. …

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Morphosyntactic Analysis of the CHILDES and TalkBank Corpora

2012-05-01 · LREC 2012 5 · Brian MacWhinney

This paper describes the construction and usage of the MOR and GRASP programs for part of speech tagging and syntactic dependency analysis of the corpora in the CHILDES and TalkBank databases. We have written MOR grammar…

Language AcquisitionMorphological AnalysisPart-Of-Speech Tagging

SLABERT Talk Pretty One Day: Modeling Second Language Acquisition with BERT

2023-05-31 · Aditya Yadavalli, Alekhya Yadavalli, Vera Tobin

Second language acquisition (SLA) research has extensively studied cross-linguistic transfer, the influence of linguistic structure of a speaker's native language [L1] on the successful acquisition of a foreign language …

Cross-Lingual TransferLanguage AcquisitionTransfer Learning

What Should Baby Models Read? Exploring Sample-Efficient Data Composition on Model Performance

2024-11-11 · Hong Meng Yam, Nathan J Paek

We explore the impact of pre-training data composition on the performance of small language models in a sample-efficient setting. Using datasets limited to 10 million words, we evaluate several dataset sources, including…

Language ModelingLanguage Modelling

CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions

2026-05-19 · Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet 외 arxiv

CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Leveraging the recent release of the UD-English-CHILDES treebank with gold…

Language Acquisition