paper-with-me

Papers

Morphosyntactic Analysis of the CHILDES and TalkBank Corpora

2012-05-01 · LREC 2012 5 · Brian MacWhinney

This paper describes the construction and usage of the MOR and GRASP programs for part of speech tagging and syntactic dependency analysis of the corpora in the CHILDES and TalkBank databases. We have written MOR grammars for 11 languages and GRASP analyses for three. For English data, the MOR tagger reaches 98{\%} accuracy on adult corpora and 97{\%} accuracy on child language corpora. The paper discusses the construction of MOR lexicons with an emphasis on compounds and special conversational forms. The shape of rules for controlling allomorphy and morpheme concatenation are discussed. The analysis of bilingual corpora is illustrated in the context of the Cantonese-English bilingual corpora. Methods for preparing data for MOR analysis and for developing MOR grammars are discussed. We believe that recent computational work using this system is leading to significant advances in child language acquisition theory and theories of grammar identification more generally.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language AcquisitionMorphological AnalysisPart-Of-Speech Tagging

Similar Papers 제목 키워드 기반

Morphosyntactic Analysis for CHILDES

2024-07-17 · Houjun Liu, Brian MacWhinney

Language development researchers are interested in comparing the process of language learning across languages. Unfortunately, it has been difficult to construct a consistent quantitative framework for such comparisons. …

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

A Hierarchical Approach to exploiting Multiple Datasets from TalkBank

2023-06-21 · Man Ho Wong

TalkBank is an online database that facilitates the sharing of linguistics research data. However, the existing TalkBank's API has limited data filtering and batch processing capabilities. To overcome these limitations, …

A large scale annotated child language construction database

2012-05-01 · LREC 2012 5 · Aline Villavicencio, Beracah Yankama, Marco Idiart, Robert Berwick

Large scale annotated corpora of child language can be of great value in assessing theoretical proposals regarding language acquisition models. For example, they can help determine whether the type and amount of data req…

Language AcquisitionPOSPOS Tagging

Searching the Annotated Portuguese Childes Corpora

2012-04-01 · WS 2012 4 · Rodrigo Wilkens
Language Acquisition

CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions

2026-05-19 · Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet 외 arxiv

CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Leveraging the recent release of the UD-English-CHILDES treebank with gold…

Language Acquisition