paper-with-me

Fact Checking

7개 벤치마크 · 논문 693편 · 이 태스크의 논문 보기 →

Benchmarks

SciFact (BEIR)

결과 5개

AVeriTeC

결과 4개

CLIMATE-FEVER (BEIR)

결과 4개

FEVER (BEIR)

결과 4개

CDCD

결과 1개

LIAR2

결과 1개

^(#$!@#$)(()))******

결과 1개

Most implemented

Papers

Agentic systems for breast cancer treatment recommendations

2026-07-13 · Vinicius Anjos de Almeida, Nícolas Henrique Borges, Leonardo Vicenzi, Helena Kociolek 외 arxiv

Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning remains unclear. We evaluated agentic LLM systems for breast cancer…

Fact Checking

Checking Fact with Better Retrieval: Dynamic Contrastive Learning for Evidence Retrieval

2026-05-24 · Zhongtian Hua, Yi Luo, Meijia Yu, Yingjie Han arxiv

In the field of multimodal fact checking, the accuracy of retrieving evidence from different modalities has a significant impact on the downstream claim verification process. Existing general multimodal retrieval methods…

Contrastive LearningSemantic RetrievalFact Checking

When AI reviews science: Can we trust the referee?

2026-04-26 · Jialiang Wang, Yuchen Liu, Hang Xu, Kaichun Hu 외 arxiv

The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large language models (LLMs) offer impressive capab…

Fact Checking

AfrIFact: Cultural Information Retrieval, Evidence Extraction and Fact Checking for African Languages

2026-04-01 · Israel Abebe Azime, Jesujoba Oluwadara Alabi, Crystina Zhang, Iffat Maab 외 arxiv

Assessing the veracity of a claim made online is a complex and important task with real-world implications. When these claims are directed at communities with limited access to information and the content concerns issues…

Information RetrievalFact Checking

Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval

2026-03-05 · Artem Vazhentsev, Maria Marina, Daniil Moskovskiy, Sergey Pletenev 외 arxiv

Trustworthiness is a core research challenge for agentic AI systems built on Large Language Models (LLMs). To enhance trust, natural language claims from diverse sources, including human-written text, web content, and mo…

Fact Checking

Markovian ODE-guided scoring can assess the quality of offline reasoning traces in language models

2026-03-02 · Arghodeep Nandi, Ojasva Saxena, Tanmoy Chakraborty arxiv

Reasoning traces produced by generative language models are increasingly used for tasks ranging from mathematical problem solving to automated fact checking. However, existing evaluation methods remain largely mechanical…

Fact Checking

전체 693편 보기 →