paper-with-me

Papers

The taggedPBC: Annotating a massive parallel corpus for crosslinguistic investigations

2025-05-18 · Hiram Ring

Existing datasets available for crosslinguistic investigations have tended to focus on large amounts of data for a small group of languages or a small amount of data for a large number of languages. This means that claims based on these datasets are limited in what they reveal about universal properties of the human language faculty. While this has begun to change through the efforts of projects seeking to develop tagged corpora for a large number of languages, such efforts are still constrained by limits on resources. The current paper reports on a large automatically tagged parallel dataset which has been developed to partially address this issue. The taggedPBC contains more than 1,800 sentences of pos-tagged parallel text data from over 1,500 languages, representing 133 language families and 111 isolates, dwarfing previously available resources. The accuracy of tags in this dataset is shown to correlate well with both existing SOTA taggers for high-resource languages (SpaCy, Trankit) as well as hand-tagged corpora (Universal Dependencies Treebanks). Additionally, a novel measure derived from this dataset, the N1 ratio, correlates with expert determinations of word order in three typological databases (WALS, Grambank, Autotyp) such that a Gaussian Naive Bayes classifier trained on this feature can accurately identify basic word order for languages not in those databases. While much work is still needed to expand and develop this dataset, the taggedPBC is an important step to enable corpus-based crosslinguistic investigations, and is made available for research and collaboration via GitHub.

📄 PDF Abstract BibTeX arXiv:2505.12560

Code (1)

lingdoc/taggedPBC 공식 구현

Tasks

POS

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Annotating Croatian Semantic Type Coercions in CROATPAS

2020-05-01 · LREC 2020 5 · Costanza Marini, Elisabetta Jezek

This short research paper presents the results of a corpus-based metonymy annotation exercise on a sample of 101 Croatian verb entries {--} corresponding to 457 patters and over 20,000 corpus lines {--} taken from CROATP…

Vocal Bursts Type Prediction

Grammatical gender associations outweigh topical gender bias in crosslinguistic word embeddings

2020-05-18 · Katherine McCurdy, Oguz Serbetci

Recent research has demonstrated that vector space models of semantics can reflect undesirable biases in human culture. Our investigation of crosslinguistic word embeddings reveals that topical gender bias interacts with…

Cultural Vocal Bursts Intensity PredictionLemmatizationMachine TranslationTranslation+1

Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis

2025-05-31 · Miao Zhang, Aref Farhadipour, Annie Baker, Jiachen Ma 외

With its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech technology and holds tremendous potential for research in crosslinguistic ph…

Diversity

Creating a massively parallel Bible corpus

2014-05-01 · LREC 2014 5 · Thomas Mayer, Michael Cysouw

We present our ongoing effort to create a massively parallel Bible corpus. While an ever-increasing number of Bible translations is available in electronic form on the internet, there is no large-scale parallel Bible cor…

Machine TranslationTranslation

Building The Sense-Tagged Multilingual Parallel Corpus

2014-05-01 · LREC 2014 5 · Shan Wang, Francis Bond

Sense-annotated parallel corpora play a crucial role in natural language processing. This paper introduces our progress in creating such a corpus for Asian languages using English as a pivot, which is the first such corp…