paper-with-me

홈 › Papers

The Common Orthographic Vocabulary of the Portuguese Language: a set of open lexical resources for a pluricentric language

2012-05-01 · LREC 2012 5 · Jos{\'e} Pedro Ferreira, Maarten Janssen, Gladis Barcellos de Oliveira, Margarita Correia, Gilvan M{\"u}ller de Oliveira

This paper outlines the design principles and choices, as well as the ongoing development process of the Common Orthographic Vocabulary of the Portuguese Language (VOC), a large scale electronic lexical database which was adopted by the Community of Portuguese-Speaking Countries' (CPLP) Instituto Internacional da L{\'\i}ngua Portuguesa to implement a spelling reform that is currently taking place. Given the different available resources and lexicographic traditions within the CPLP countries, a range of different solutions was adopted for different countries and integrated into a common development framework. Although the publication of lexicographic resources to implement spelling reforms has always been done for Portuguese, VOC represents a paradigm change, switching from idiosyncratic, closed source, paper-format official resources to standardized, open, free, web-accessible and reusable ones. We start by outlining the context that justifies the resource development and its requirements, then focusing on the description of the methodology, workflow and tools used, showing how a collaborative project in a common web-based platform and administration interface make the creation of such a long-sought and ambitious project possible.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Recognizing the vocabulary of Brazilian popular newspapers with a free-access computational dictionary

2019-04-19 · Maria José Finatto, Oto Vale, Eric Laporte

We report an experiment to check the identification of a set of words in popular written Portuguese with two versions of a computational dictionary of Brazilian Portuguese, DELAF PB 2004 and DELAF PB 2015. This dictionar…

The BDCam\~oes Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology

2020-05-01 · LREC 2020 5 · Sara Grilo, M{\'a}rcia Bolrinha, Jo{\~a}o Silva, Rui Vaz 외

This paper presents the BDCam{\~o}es Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese that in its inaugural version includes close to 4 million words from over 200 complet…

Genre classification

TuPy-E: detecting hate speech in Brazilian Portuguese social media with a novel dataset and comprehensive analysis of models

2023-12-29 · Felipe Oliveira, Victoria Reis, Nelson Ebecken

Social media has become integral to human interaction, providing a platform for communication and expression. However, the rise of hate speech on these platforms poses significant risks to individuals and communities. De…

Hate Speech Detection

PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data

2020-08-20 · Diedre Carmo, Marcos Piau, Israel Campiotti, Rodrigo Nogueira 외

In natural language processing (NLP), there is a need for more resources in Portuguese, since much of the data used in the state-of-the-art research is in other languages. In this paper, we pretrain a T5 model on the BrW…

Feature-based Decipherment for Large Vocabulary Machine Translation

2015-08-10 · Iftekhar Naim, Daniel Gildea

Orthographic similarities across languages provide a strong signal for probabilistic decipherment, especially for closely related language pairs. The existing decipherment models, however, are not well-suited for exploit…

DeciphermentMachine TranslationTranslation