paper-with-me

홈 › Papers

Liars' Bench: Evaluating Lie Detectors for Language Models

2025-11-20 · Kieron Kretschmar, Walter Laurito, Sharan Maiya, Samuel Marks arxiv

Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typically validated in narrow settings that do not capture the diverse lies LLMs can generate. We introduce LIARS' BENCH, a testbed consisting of 72,863 examples of lies and honest responses generated by four open-weight models across seven datasets. Our settings capture qualitatively different types of lies and vary along two dimensions: the model's reason for lying and the object of belief targeted by the lie. Evaluating three black- and white-box lie detection techniques on LIARS' BENCH, we find that existing techniques systematically fail to identify certain types of lies, especially in settings where it's not possible to determine whether the model lied from the transcript alone. Overall, LIARS' BENCH reveals limitations in prior techniques and provides a practical testbed for guiding progress in lie detection.

📄 PDF Abstract BibTeX arXiv:2511.16035

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia

2025-07-04 · Gerrit Quaremba, Elizabeth Black, Denny Vrandečić, Elena Simperl arxiv

Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-generated text (MGT) produced by large language models (LLMs) on its platf…

Text Style Transfer

Are AI Detectors Good Enough? A Survey on Quality of Datasets With Machine-Generated Texts

2024-10-18 · German Gritsai, Anastasia Voznyuk, Andrey Grabovoy, Yury Chekhovich

The rapid development of autoregressive Large Language Models (LLMs) has significantly improved the quality of generated texts, necessitating reliable machine-generated text detectors. A huge number of detectors and coll…

VoiceWukong: Benchmarking Deepfake Voice Detection

2024-09-10 · Ziwei Yan, Yanjie Zhao, Haoyu Wang

With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intu…

BenchmarkingFace SwappingLarge Language Modeltext-to-speech+2

TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors

2025-03-10 · Jingyi Zheng, Junfeng Wang, Zhen Sun, Wenhan Dong 외

As Large Language Models (LLMs) advance, Machine-Generated Texts (MGTs) have become increasingly fluent, high-quality, and informative. Existing wide-range MGT detectors are designed to identify MGTs to prevent the sprea…

Misinformation

Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions

2026-03-18 · Madhav S. Baidya, S. S. Baidya, Chirag Chawla arxiv

The rapid proliferation of large language models (LLMs) has created an urgent need for robust and generalizable detectors of machine-generated text. Existing benchmarks typically evaluate a single detector on a single da…

Adversarial Robustness