PreCo: A Large-scale Dataset in Preschool Vocabulary for Coreference Resolution
We introduce PreCo, a large-scale English dataset for coreference resolution. The dataset is designed to embody the core challenges in coreference, such as entity representation, by alleviating the challenge of low overlap between training and test sets and enabling separated analysis of mention detection and mention clustering. To strengthen the training-test overlap, we collect a large corpus of about 38K documents and 12.4M words which are mostly from the vocabulary of English-speaking preschoolers. Experiments show that with higher training-test overlap, error analysis on PreCo is more efficient than the one on OntoNotes, a popular existing dataset. Furthermore, we annotate singleton mentions making it possible for the first time to quantify the influence that a mention detector makes on coreference resolution performance. The dataset is freely available at https://preschool-lab.github.io/PreCo/.
Code (0)
등록된 구현이 없습니다.
Tasks
Clusteringcoreference-resolutionCoreference ResolutionSimilar Papers 제목 키워드 기반
Representing the Toddler Lexicon: Do the Corpus and Semantics Matter?
Understanding child language development requires accurately representing children’s lexicons. However, much of the past work modeling children’s vocabulary development has utilized adult-based measures. The present inve…
When AI Meets Early Childhood Education: Large Language Models as Assessment Teammates in Chinese Preschools
High-quality teacher-child interaction (TCI) is fundamental to early childhood development, yet traditional expert-based assessment faces a critical scalability challenge. In large systems like China's-serving 36 million…
Speech RecognitionCAGS: Open-Vocabulary 3D Scene Understanding with Context-Aware Gaussian Splatting
Open-vocabulary 3D scene understanding is crucial for applications requiring natural language-driven spatial interpretation, such as robotics and augmented reality. While 3D Gaussian Splatting (3DGS) offers a powerful re…
3DGS3D Instance SegmentationContrastive LearningInstance Segmentation+2Measuring an Artificial Intelligence System's Performance on a Verbal IQ Test For Young Children
We administered the Verbal IQ (VIQ) part of the Wechsler Preschool and Primary Scale of Intelligence (WPPSI-III) to the ConceptNet 4 AI system. The test questions (e.g., "Why do we shake hands?") were translated into Con…
Common Sense ReasoningQuestion Answering"Favoring my playmate seems fair": Inhibitory control and theory of mind in preschoolers' self-disadvantaging behaviors
The purpose of this study was to investigate the relationship between preschoolers' cognitive abilities and their fairness-related allocation behaviors in a dilemma of equity-efficiency conflict. Four- to 6-year-olds in …
Fairness