paper-with-me

홈 › Papers

Representing the Toddler Lexicon: Do the Corpus and Semantics Matter?

2022-06-01 · LREC 2022 6 · Jennifer Weber, Eliana Colunga

Understanding child language development requires accurately representing children’s lexicons. However, much of the past work modeling children’s vocabulary development has utilized adult-based measures. The present investigation asks whether using corpora that captures the language input of young children more accurately represents children’s vocabulary knowledge. We present a newly-created toddler corpus that incorporates transcripts of child-directed conversations, the text of picture books written for preschoolers, and dialog from G-rated movies to approximate the language input a North American preschooler might hear. We evaluate the utility of the new corpus for modeling children’s vocabulary development by building and analyzing different semantic network models and comparing them to norms based on vocabulary norms for toddlers in this age range. More specifically, the relations between words in our semantic networks were derived from skip-gram neural networks (Word2Vec) trained on our toddler corpus or on Google news. Results revealed that the models built from the toddler corpus were more accurate at predicting toddler vocabulary growth than the adult-based corpus. These results speak to the importance of selecting a corpus that matches the population of interest.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

A Morphological Lexicon of Esperanto with Morpheme Frequencies

2016-05-01 · LREC 2016 5 · Eckhard Bick

This paper discusses the internal structure of complex Esperanto words (CWs). Using a morphological analyzer, possible affixation and compounding is checked for over 50,000 Esperanto lexemes against a list of 17,000 root…

POS

Multiplex lexical networks reveal patterns in early word acquisition in children

2016-09-11 · Massimo Stella, Nicole M. Beckage, Markus Brede

Network models of language have provided a way of linking cognitive processes to the structure and connectivity of language. However, one shortcoming of current approaches is focusing on only one type of linguistic relat…

Computational valency lexica and Homeric formularity

2022-08-23 · Barbara McGillivray, Martina Astrid Rodda

Distributional semantics, the quantitative study of meaning variation and change through corpus collocations, is currently one of the most productive research areas in computational linguistics. The wider availability of…

Automated Pronunciation Evaluation for Korean Toddler Speech using Speech Diarization and Self-Supervised Learning

2026-06-08 · Diane Myung-kyung Woodbridge, Jee Hyun Suh arxiv

Speech sound disorders affect approximately 44% of Korean pediatric communication disorder cases, yet automated assessment tools for Korean toddler speech remain underdeveloped. This paper presents an end-to-end pipeline…

Self-Supervised LearningRepresentation LearningSpeaker Diarization

DiMLex-Bangla: A Lexicon of Bangla Discourse Connectives

2020-05-01 · LREC 2020 5 · Debopam Das, Manfred Stede, Soumya Sankar Ghosh, Lahari Chatterjee

We present DiMLex-Bangla, a newly developed lexicon of discourse connectives in Bangla. The lexicon, upon completion of its first version, contains 123 Bangla connective entries, which are primarily compiled from the lin…

Translation