paper-with-me

Papers

FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models

2025-02-25 · Hongzhan Lin, Yang Deng, Yuxuan Gu, Wenxuan Zhang, Jing Ma, See-Kiong Ng, Tat-Seng Chua

Large Language Models (LLMs) have significantly advanced the fact-checking studies. However, existing automated fact-checking evaluation methods rely on static datasets and classification metrics, which fail to automatically evaluate the justification production and uncover the nuanced limitations of LLMs in fact-checking. In this work, we introduce FACT-AUDIT, an agent-driven framework that adaptively and dynamically assesses LLMs' fact-checking capabilities. Leveraging importance sampling principles and multi-agent collaboration, FACT-AUDIT generates adaptive and scalable datasets, performs iterative model-centric evaluations, and updates assessments based on model-specific responses. By incorporating justification production alongside verdict prediction, this framework provides a comprehensive and evolving audit of LLMs' factual reasoning capabilities, to investigate their trustworthiness. Extensive experiments demonstrate that FACT-AUDIT effectively differentiates among state-of-the-art LLMs, providing valuable insights into model strengths and limitations in model-centric fact-checking analysis.

📄 PDF Abstract BibTeX arXiv:2502.17924

Code (1)

DanielLin97/FACT-AUDIT 공식 구현

Tasks

Fact Checking

Similar Papers 제목 키워드 기반

AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification

2026-06-02 · Yan Wang, Xuguang Ai, Jaisal Patel, Xueqing Peng 외 arxiv

Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone. A model must link reported facts to taxonomy concepts, traverse ca…

FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents

2026-08-29 · Xiangxin Luo, Chengtian Hong, Haohua Li, Yongyi Xie arxiv

Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical contamination, expose sensitivity to a single realized market path, or reveal internal degeneration during long-h…

Optimally Auditing Adversarial Agents

2026-04-28 · Sanmay Das, Fang-Yi Yu, Yuang Zhang arxiv

Fraud can pose a challenge in many resource allocation domains, including social service delivery and credit provision. For example, agents may misreport private information in order to gain benefits or access to credit.…

Adaptive Memory Admission Control for LLM Agents

2026-03-04 · Guilin Zhang, Wei Jiang, Xiejiashan Wang, Aisha Behr 외 arxiv

LLM-based agents increasingly rely on long-term memory to support multi-session reasoning and interaction, yet current systems provide little control over what information is retained. In practice, agents either accumula…

Resilient Multi-Agent Negotiation for Medical Supply Chains:Integrating LLMs and Blockchain for Transparent Coordination

2025-07-23 · Mariam ALMutairi, Hyungmin Kim arxiv

Global health emergencies, such as the COVID-19 pandemic, have exposed critical weaknesses in traditional medical supply chains, including inefficiencies in resource allocation, lack of transparency, and poor adaptabilit…