paper-with-me

Papers

UD-English-CHILDES: A Collected Resource of Gold and Silver Universal Dependencies Trees for Child Language Interactions

2025-04-28 · Xiulin Yang, Zhuoxuan Ju, Lanni Bu, Zoey Liu, Nathan Schneider

CHILDES is a widely used resource of transcribed child and child-directed speech. This paper introduces UD-English-CHILDES, the first officially released Universal Dependencies (UD) treebank derived from previously dependency-annotated CHILDES data with consistent and unified annotation guidelines. Our corpus harmonizes annotations from 11 children and their caregivers, totaling over 48k sentences. We validate existing gold-standard annotations under the UD v2 framework and provide an additional 1M silver-standard sentences, offering a consistent resource for computational and linguistic research.

📄 PDF Abstract BibTeX arXiv:2504.20304

Code (2)

universaldependencies/ud_english-childes 공식 구현
xiulinyang/ud-childes 공식 구현

Similar Papers 제목 키워드 기반

CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions

2026-05-19 · Francesca Padovani, Xiulin Yang, Bastian Bunzeck, Jaap Jumelet 외 arxiv

CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Leveraging the recent release of the UD-English-CHILDES treebank with gold…

Language Acquisition

SilverAlign: MT-Based Silver Data Algorithm For Evaluating Word Alignment

2022-10-12 · Abdullatif Köksal, Silvia Severini, Hinrich Schütze

Word alignments are essential for a variety of NLP tasks. Therefore, choosing the best approaches for their creation is crucial. However, the scarce availability of gold evaluation data makes the choice difficult. We pro…

Machine TranslationTranslationvalidWord Alignment

Platforms for Non-speakers Annotating Names in Any Language

2018-07-01 · ACL 2018 7 · Ying Lin, Cash Costello, Boliang Zhang, Di Lu 외

We demonstrate two annotation platforms that allow an English speaker to annotate names for any language without knowing the language. These platforms provided high-quality {'}{`}silver standard{''} annotations for low-r…

Cross-Lingual Transfer for Distantly Supervised and Low-resources Indonesian NER

2019-07-25 · Fariz Ikhwantri

Manually annotated corpora for low-resource languages are usually small in quantity (gold), or large but distantly supervised (silver). Inspired by recent progress of injecting pre-trained language model (LM) on many Nat…

Cross-Lingual TransferLanguage ModelingLanguage ModellingNER+2

A large scale annotated child language construction database

2012-05-01 · LREC 2012 5 · Aline Villavicencio, Beracah Yankama, Marco Idiart, Robert Berwick

Large scale annotated corpora of child language can be of great value in assessing theoretical proposals regarding language acquisition models. For example, they can help determine whether the type and amount of data req…

Language AcquisitionPOSPOS Tagging