paper-with-me

홈 › Papers

GenAI vs. Human Fact-Checkers: Accurate Ratings, Flawed Rationales

2025-02-20 · Yuehong Cassandra Tai, Khushi Navin Patni, Nicholas Daniel Hemauer, Bruce Desmarais, Yu-Ru Lin

Despite recent advances in understanding the capabilities and limits of generative artificial intelligence (GenAI) models, we are just beginning to understand their capacity to assess and reason about the veracity of content. We evaluate multiple GenAI models across tasks that involve the rating of, and perceived reasoning about, the credibility of information. The information in our experiments comes from content that subnational U.S. politicians post to Facebook. We find that GPT-4o, one of the most used AI models in consumer applications, outperforms other models, but all models exhibit only moderate agreement with human coders. Importantly, even when GenAI models accurately identify low-credibility content, their reasoning relies heavily on linguistic features and ``hard'' criteria, such as the level of detail, source reliability, and language formality, rather than an understanding of veracity. We also assess the effectiveness of summarized versus full content inputs, finding that summarized content holds promise for improving efficiency without sacrificing accuracy. While GenAI has the potential to support human fact-checkers in scaling misinformation detection, our results caution against relying solely on these models.

📄 PDF Abstract BibTeX arXiv:2502.14943

Code (0)

등록된 구현이 없습니다.

Tasks

Misinformation

Similar Papers 제목 키워드 기반

GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

2024-06-19 · Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li 외

While text-to-visual models now produce photo-realistic images and videos, they struggle with compositional text prompts involving attributes, relationships, and higher-order reasoning such as logic and comparison. In th…

BenchmarkingImage GenerationVideo Generation

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

2025-09-26 · Luke Guerdan, Justin Whitehouse, Kimberly Truong, Kenneth Holstein 외 arxiv

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to th…

Validating LLM-as-a-Judge Systems in the Absence of Gold Labels

2025-03-07 · Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach 외

The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, has come to play a critical role in scaling and standardizing GenAI evaluations…

Exploring Multidimensional Checkworthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkers

2024-12-11 · Houjiang Liu, Jacek Gwizdka, Matthew Lease

Given the massive volume of potentially false claims circulating online, claim prioritization is essential in allocating limited human resources available for fact-checking. In this study, we perceive claim prioritizatio…

Fact CheckingInformation Retrieval

Can LLMs Automate Fact-Checking Article Writing?

2025-03-22 · Dhruv Sahnan, David Corney, Irene Larraz, Giovanni Zagni 외

Automatic fact-checking aims to support professional fact-checkers by offering tools that can help speed up manual fact-checking. Yet, existing frameworks fail to address the key step of producing output suitable for bro…

ArticlesFact CheckingText Generation