Fact Checking
7개 벤치마크 · 논문 693편 · 이 태스크의 논문 보기 →
Benchmarks
SciFact (BEIR)
AVeriTeC
CLIMATE-FEVER (BEIR)
FEVER (BEIR)
CDCD
LIAR2
^(#$!@#$)(()))******
Most implemented
"Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection
A simple but tough-to-beat baseline for the Fake News Challenge stance detection task
Unsupervised Dense Information Retrieval with Contrastive Learning
Explainable Tsetlin Machine framework for fake news detection with credibility score assessment
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
Evaluating the Factual Consistency of Abstractive Text Summarization
Papers
Agentic systems for breast cancer treatment recommendations
Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning remains unclear. We evaluated agentic LLM systems for breast cancer…
Fact CheckingChecking Fact with Better Retrieval: Dynamic Contrastive Learning for Evidence Retrieval
In the field of multimodal fact checking, the accuracy of retrieving evidence from different modalities has a significant impact on the downstream claim verification process. Existing general multimodal retrieval methods…
Contrastive LearningSemantic RetrievalFact CheckingWhen AI reviews science: Can we trust the referee?
The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large language models (LLMs) offer impressive capab…
Fact CheckingAfrIFact: Cultural Information Retrieval, Evidence Extraction and Fact Checking for African Languages
Assessing the veracity of a claim made online is a complex and important task with real-world implications. When these claims are directed at communities with limited access to information and the content concerns issues…
Information RetrievalFact CheckingLeveraging LLM Parametric Knowledge for Fact Checking without Retrieval
Trustworthiness is a core research challenge for agentic AI systems built on Large Language Models (LLMs). To enhance trust, natural language claims from diverse sources, including human-written text, web content, and mo…
Fact CheckingMarkovian ODE-guided scoring can assess the quality of offline reasoning traces in language models
Reasoning traces produced by generative language models are increasingly used for tasks ranging from mathematical problem solving to automated fact checking. However, existing evaluation methods remain largely mechanical…
Fact Checking