paper-with-me

홈 › Papers

Curated Multilingual Language Resources for CEF AT (CURLICAT): overall view

2022-06-01 · EAMT 2022 6 · Tamás Váradi, Marko Tadić, Svetla Koeva, Maciej Ogrodniczuk, Dan Tufiş, Radovan Garabík, Simon Krek, Andraž Repar

The work in progress on the CEF Action CURLICA T is presented. The general aim of the Action is to compile curated datasets in seven languages of the con- sortium in domains of relevance to Euro- pean Digital Service Infrastructures (DSIs) in order to enhance the eTransla- tion services.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources

2022-06-01 · LREC 2022 6 · Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadić 외

This article presents the current outcomes of the CURLICAT CEF Telecom project, which aims to collect and deeply annotate a set of large corpora from selected domains. The CURLICAT corpus includes 7 monolingual corpora (…

NMT

UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment

2025-06-02 · Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens 외

We introduce UniversalCEFR, a large-scale multilingual multidimensional dataset of texts annotated according to the CEFR (Common European Framework of Reference) scale in 13 languages. To enable open research in both aut…

El Departamento de Nosotros: How Machine Translated Corpora Affects Language Models in MRC Tasks

2020-07-03 · Maria Khvalchik, Mikhail Galkin

Pre-training large-scale language models (LMs) requires huge amounts of text corpora. LMs for English enjoy ever growing corpora of diverse language resources. However, less resourced languages and their mono- and multil…

Machine TranslationQuestion AnsweringTransfer LearningTranslation

Massive vs. Curated Embeddings for Low-Resourced Languages: the Case of Yor\`ub\'a and Twi

2020-05-01 · LREC 2020 5 · Jesujoba Alabi, Kwabena Amponsah-Kaakyire, David Adelani, Cristina Espa{\~n}a-Bonet

The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Word Embeddings

MathGloss: Building mathematical glossaries from text

2023-11-21 · Lucy Horowitz, Valeria de Paiva

MathGloss is a project to create a knowledge graph (KG) for undergraduate mathematics from text, automatically, using modern natural language processing (NLP) tools and resources already available on the web. MathGloss i…

Math