paper-with-me

홈 › Papers

VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking

2026-01-13 · Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna Rohrbach arxiv

The growing scale of online misinformation urgently demands Automated Fact-Checking (AFC). Existing benchmarks for evaluating AFC systems, however, are largely limited in terms of task scope, modalities, domain, language diversity, realism, or coverage of misinformation types. Critically, they are static, thus subject to data leakage as their claims enter the pretraining corpora of LLMs. As a result, benchmark performance no longer reliably reflects the actual ability to verify claims. We introduce Verified Theses and Statements (VeriTaS), the first dynamic benchmark for multimodal AFC, designed to remain robust under ongoing large-scale pretraining of foundation models. VeriTaS currently comprises 25,000 real-world claims from 104 professional fact-checking organizations across 54 languages, covering textual and audiovisual content. Claims are added quarterly via a fully automated seven-stage pipeline that normalizes claim formulation, retrieves original media, and maps heterogeneous expert verdicts to a novel, standardized, and disentangled scoring scheme with textual justifications. Through human evaluation, we demonstrate that the automated annotations closely match human judgments. We commit to updating VeriTaS in the future, establishing a leakage-resistant benchmark, supporting meaningful AFC evaluation in the era of rapidly evolving foundation models. The code and data are publicly available under https://veritas.mai.informatik.tu-darmstadt.de .

📄 PDF Abstract BibTeX arXiv:2601.08611

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data

2025-10-17 · Tingqiao Xu, Ziru Zeng, Jiayu Chen arxiv

The quality of supervised fine-tuning (SFT) data is crucial for the performance of large multimodal models (LMMs), yet current data enhancement methods often suffer from factual errors and hallucinations due to inadequat…

VERITAS: Verifier-Guided Proof Search for Zero-Shot Formal Theorem Proving

2026-06-17 · Manish Acharya, Zhenyu Liao, Yueke Zhang, Kevin Leach 외 arxiv

LLM-based formal provers often collapse rich verifier signals (syntax errors, type mismatches, partial goal progress) into a binary pass/fail bit. We present VERITAS, a zero-shot framework that routes every verifier sign…

SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions

2025-09-21 · Massa Baali, Sarthak Bisht, Francisco Teixeira, Kateryna Shapovalenko 외 arxiv

Speaker verification (SV) models are increasingly integrated into security, personalization, and access control systems, yet their robustness to many real-world challenges remains inadequately benchmarked. These include …

Speaker Verification

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

2026-07-03 · Haokun Liu, Filbert Aurelian Tjiaranata, Chenhao Tan arxiv

AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published research has become both harder and more important. As manual replication i…

Versatile Verification of Tree Ensembles

2020-10-26 · Laurens Devos, Wannes Meert, Jesse Davis

Machine learned models often must abide by certain requirements (e.g., fairness or legal). This has spurred interested in developing approaches that can provably verify whether a model satisfies certain properties. This …

Fairness