paper-with-me

Papers

Challenges in Expanding Portuguese Resources: A View from Open Information Extraction

2025-01-21 · Marlo Souza, Bruno Cabral, Daniela Claro, Lais Salvador

Open Information Extraction (Open IE) is the task of extracting structured information from textual documents, independent of domain. While traditional Open IE methods were based on unsupervised approaches, recently, with the emergence of robust annotated datasets, new data-based approaches have been developed to achieve better results. These innovations, however, have focused mainly on the English language due to a lack of datasets and the difficulty of constructing such resources for other languages. In this work, we present a high-quality manually annotated corpus for Open Information Extraction in the Portuguese language, based on a rigorous methodology grounded in established semantic theories. We discuss the challenges encountered in the annotation process, propose a set of structural and contextual annotation rules, and validate our corpus by evaluating the performance of state-of-the-art Open IE systems. Our resource addresses the lack of datasets for Open IE in Portuguese and can support the development and evaluation of new methods and systems in this area.

📄 PDF Abstract BibTeX arXiv:2501.11851

Code (0)

등록된 구현이 없습니다.

Tasks

Open Information Extraction

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Enhancing Portuguese Variety Identification with Cross-Domain Approaches

2025-02-20 · Hugo Sousa, Rúben Almeida, Purificação Silvano, Inês Cantante 외

Recent advances in natural language processing have raised expectations for generative models to produce coherent text across diverse language varieties. In the particular case of the Portuguese language, the predominanc…

NomLex-PT: A Lexicon of Portuguese Nominalizations

2014-05-01 · LREC 2014 5 · Valeria de Paiva, Livy Real, Alex Rademaker, re 외

This paper presents NomLex-PT, a lexical resource describing Portuguese nominalizations. NomLex-PT connects verbs to their nominalizations, thereby enabling NLP systems to observe the potential semantic relationships bet…

Advancing Generative AI for Portuguese with Open Decoder Gervásio PT*

2024-02-29 · Rodrigo Santos, João Silva, Luís Gomes, João Rodrigues 외

To advance the neural decoding of Portuguese, in this paper we present a fully open Transformer-based, instruction-tuned decoder model that sets a new state of the art in this respect. To develop this decoder, which we n…

Decoder

The Common Orthographic Vocabulary of the Portuguese Language: a set of open lexical resources for a pluricentric language

2012-05-01 · LREC 2012 5 · Jos{\'e} Pedro Ferreira, Maarten Janssen, Gladis Barcellos de Oliveira, Margarita Correia 외

This paper outlines the design principles and choices, as well as the ongoing development process of the Common Orthographic Vocabulary of the Portuguese Language (VOC), a large scale electronic lexical database which wa…

Fostering the Ecosystem of Open Neural Encoders for Portuguese with Albertina PT* Family

2024-03-04 · Rodrigo Santos, João Rodrigues, Luís Gomes, João Silva 외

To foster the neural encoding of Portuguese, this paper contributes foundation encoder models that represent an expansion of the still very scarce ecosystem of large language models specifically developed for this langua…