Morphosyntactic Analysis of the CHILDES and TalkBank Corpora
This paper describes the construction and usage of the MOR and GRASP programs for part of speech tagging and syntactic dependency analysis of the corpora in the CHILDES and TalkBank databases. We have written MOR grammars for 11 languages and GRASP analyses for three. For English data, the MOR tagger reaches 98{\%} accuracy on adult corpora and 97{\%} accuracy on child language corpora. The paper discusses the construction of MOR lexicons with an emphasis on compounds and special conversational forms. The shape of rules for controlling allomorphy and morpheme concatenation are discussed. The analysis of bilingual corpora is illustrated in the context of the Cantonese-English bilingual corpora. Methods for preparing data for MOR analysis and for developing MOR grammars are discussed. We believe that recent computational work using this system is leading to significant advances in child language acquisition theory and theories of grammar identification more generally.
Code (0)
등록된 구현이 없습니다.
Tasks
Language AcquisitionMorphological AnalysisPart-Of-Speech TaggingSimilar Papers 제목 키워드 기반
Morphosyntactic Analysis for CHILDES
Language development researchers are interested in comparing the process of language learning across languages. Unfortunately, it has been difficult to construct a consistent quantitative framework for such comparisons. …
Automatic Speech Recognitionspeech-recognitionSpeech RecognitionA Hierarchical Approach to exploiting Multiple Datasets from TalkBank
TalkBank is an online database that facilitates the sharing of linguistics research data. However, the existing TalkBank's API has limited data filtering and batch processing capabilities. To overcome these limitations, …
A large scale annotated child language construction database
Large scale annotated corpora of child language can be of great value in assessing theoretical proposals regarding language acquisition models. For example, they can help determine whether the type and amount of data req…
Language AcquisitionPOSPOS TaggingSearching the Annotated Portuguese Childes Corpora
CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions
CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Leveraging the recent release of the UD-English-CHILDES treebank with gold…
Language Acquisition