paper-with-me

Papers

MULTEXT-East

2020-03-31 · Tomaž Erjavec

MULTEXT-East language resources, a multilingual dataset for language engineering research, focused on the morphosyntactic level of linguistic description. The MULTEXT-East dataset includes the EAGLES-based morphosyntactic specifications, morphosyntactic lexicons, and an annotated multilingual corpora. The parallel corpus, the novel "1984" by George Orwell, is sentence aligned and contains hand-validated morphosyntactic descriptions and lemmas. The resources are uniformly encoded in XML, using the Text Encoding Initiative Guidelines, TEI P5, and cover 16 languages: Bulgarian, Croatian, Czech, English, Estonian, Hungarian, Macedonian, Persian, Polish, Resian, Romanian, Russian, Serbian, Slovak, Slovene, and Ukrainian. This dataset is extensively documented, and freely available for research purposes. This case study gives a history of the development of the MULTEXT-East resources, presents their encoding and components, discusses related work and gives some conclusions.

📄 PDF Abstract BibTeX arXiv:2003.14026

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

The First Parallel Multilingual Corpus of Persian: Toward a Persian BLARK

2014-04-17 · Behrang Qasemizadeh, Saeed Rahimi, Behrooz Mahmoodi Bakhtiari

In this article, we have introduced the first parallel corpus of Persian with more than 10 other European languages. This article describes primary steps toward preparing a Basic Language Resources Kit (BLARK) for Persia…

Interface Web pour l'annotation morpho-syntaxique de textes (Web interface for the morpho-syntactic annotation of texts)

2016-07-01 · JEPTALNRECITAL 2016 7 · Thierry Hamon

Nous pr{\'e}sentons une interface Web pour la visualisation etl{'}annotation de textes avec des {\'e}tiquettes morphosyntaxiques etdes lemmes. Celle-ci est actuellement utilis{\'e}e pour annoter destextes ukrainiens avec…

es-en

Machine Learning and Deep Neural Network-Based Lemmatization and Morphosyntactic Tagging for Serbian

2020-05-01 · LREC 2020 5 · Ranka Stankovic, {\v{S}}, Branislava rih, Cvetana Krstev 외

The training of new tagger models for Serbian is primarily motivated by the enhancement of the existing tagset with the grammatical category of a gender. The harmonization of resources that were manually annotated within…

BIG-bench Machine LearningLemmatizationPOSPOS Tagging

Breast density in MRI: an AI-based quantification and relationship to assessment in mammography

2025-04-21 · Yaqian Chen, Lin Li, Hanxue Gu, Haoyu Dong 외

Mammographic breast density is a well-established risk factor for breast cancer. Recently there has been interest in breast MRI as an adjunct to mammography, as this modality provides an orthogonal and highly quantitativ…

Learning the shape of female breasts: an open-access 3D statistical shape model of the female breast built from 110 breast scans

2021-07-28 · Maximilian Weiherer, Andreas Eigenberger, Bernhard Egger, Vanessa Brébant 외

We present the Regensburg Breast Shape Model (RBSM) -- a 3D statistical shape model of the female breast built from 110 breast scans acquired in a standing position, and the first publicly available. Together with the mo…

Specificity