paper-with-me

홈 › Papers

ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles

2026-01-27 · Ricardo Campos, Raquel Sequeira, Sara Nerea, Inês Cantante, Diogo Folques, Luís Filipe Cunha, João Canavilhas, António Branco, Alípio Jorge, Sérgio Nunes, Nuno Guimarães, Purificação Silvano arxiv

Fact-checking remains a demanding and time-consuming task, still largely dependent on manual verification and unable to match the rapid spread of misinformation online. This is particularly important because debunking false information typically takes longer to reach consumers than the misinformation itself; accelerating corrections through automation can therefore help counter it more effectively. Although many organizations perform manual fact-checking, this approach is difficult to scale given the growing volume of digital content. These limitations have motivated interest in automating fact-checking, where identifying claims is a crucial first step. However, progress has been uneven across languages, with English dominating due to abundant annotated data. Portuguese, like other languages, still lacks accessible, licensed datasets, limiting research, NLP developments and applications. In this paper, we introduce ClaimPT, a dataset of European Portuguese news articles annotated for factual claims, comprising 1,308 articles and 6,875 individual annotations. Unlike most existing resources based on social media or parliamentary transcripts, ClaimPT focuses on journalistic content, collected through a partnership with LUSA, the Portuguese News Agency. To ensure annotation quality, two trained annotators labeled each article, with a curator validating all annotations according to a newly proposed scheme. We also provide baseline models for claim detection, establishing initial benchmarks and enabling future NLP and IR applications. By releasing ClaimPT, we aim to advance research on low-resource fact-checking and enhance understanding of misinformation in news media.

📄 PDF Abstract BibTeX arXiv:2601.19490

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adapting Freely Available Resources to Build an Opinion Mining Pipeline in Portuguese

2014-05-01 · LREC 2014 5 · Patrik Lambert, Carlos Rodr{\'\i}guez-Penagos

We present a complete UIMA-based pipeline for sentiment analysis in Portuguese news using freely available resources and a minimal set of manually annotated training data. We obtained good precision on binary classificat…

Binary ClassificationGeneral ClassificationNamed Entity Recognition (NER)Opinion Mining+2

Predicting Sentence-Level Factuality of News and Bias of Media Outlets

2023-01-27 · Francielle Vargas, Kokil Jaidka, Thiago A. S. Pardo, Fabrício Benevenuto

Automated news credibility and fact-checking at scale require accurately predicting news factuality and media bias. This paper introduces a large sentence-level dataset, titled "FactNews", composed of 6,191 sentences exp…

ArticlesFact CheckingSentencetext-classification+1

NewsClaims: A New Benchmark for Claim Detection from News with Background Knowledge

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Claim detection and verification are crucial for news understanding and have emerged as promising technologies for mitigating news misinformation. However, most existing work has focused on claim sentence analysis while …

ArticlesMisinformationSentence

PublicHearingBR: A Brazilian Portuguese Dataset of Public Hearing Transcripts for Summarization of Long Documents

2024-10-10 · Leandro Carísio Fernandes, Guilherme Zeferino Rodrigues Dobins, Roberto Lotufo, Jayr Alencar Pereira

This paper introduces PublicHearingBR, a Brazilian Portuguese dataset designed for summarizing long documents. The dataset consists of transcripts of public hearings held by the Brazilian Chamber of Deputies, paired with…

ArticlesDocument SummarizationHallucinationNatural Language Inference

VerbLexPor: a lexical resource with semantic roles for Portuguese

2016-05-01 · LREC 2016 5 · Leonardo Zilio, Maria Jos{\'e} Bocorny Finatto, Aline Villavicencio

This paper presents a lexical resource developed for Portuguese. The resource contains sentences annotated with semantic roles. The sentences were extracted from two domains: Cardiology research papers and newspaper arti…

ArticlesSentence