The ACQDIV Corpus Database and Aggregation Pipeline
We present the ACQDIV corpus database and aggregation pipeline, a tool developed as part of the European Research Council (ERC) funded project ACQDIV, which aims to identify the universal cognitive processes that allow children to acquire any language. The corpus database represents 15 corpora from 14 typologically maximally diverse languages. Here we give an overview of the project, database, and our extensible software package for adding more corpora to the current language sample. Lastly, we discuss how we use the corpus database to mine for universal patterns in child language acquisition corpora and we describe avenues for future research.
Code (0)
등록된 구현이 없습니다.
Tasks
Language AcquisitionSimilar Papers 제목 키워드 기반
The ACQDIV Database: Min(d)ing the Ambient Language
One of the most pressing questions in cognitive science remains unanswered: what cognitive mechanisms enable children to learn any of the world{'}s 7000 or so languages? Much discovery has been made with regard to specif…
DiversityLanguage AcquisitionInfrastructure for Semantic Annotation in the Genomics Domain
We describe a novel super-infrastructure for biomedical text mining which incorporates an end-to-end pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature, co…
RetrievalPRODIS - a speech database and a phoneme-based language model for the study of predictability effects in Polish
We present a speech database and a phoneme-level language model of Polish. The database and model are designed for the analysis of prosodic and discourse factors and their impact on acoustic parameters in interaction wit…
Language ModelingLanguage ModellingSpeaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study
Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most benchmarks fix a single self-supervised …
SegBo: A Database of Borrowed Sounds in the World's Languages
Phonological segment borrowing is a process through which languages acquire new contrastive speech sounds as the result of borrowing new words from other languages. Despite the fact that phonological segment borrowing is…