paper-with-me

홈 › Papers

Are You Human? An Adversarial Benchmark to Expose LLMs

2024-10-12 · Gilad Gressel, Rahul Pankajakshan, Yisroel Mirsky

Large Language Models (LLMs) have demonstrated an alarming ability to impersonate humans in conversation, raising concerns about their potential misuse in scams and deception. Humans have a right to know if they are conversing to an LLM. We evaluate text-based prompts designed as challenges to expose LLM imposters in real-time. To this end we compile and release an open-source benchmark dataset that includes 'implicit challenges' that exploit an LLM's instruction-following mechanism to cause role deviation, and 'exlicit challenges' that test an LLM's ability to perform simple tasks typically easy for humans but difficult for LLMs. Our evaluation of 9 leading models from the LMSYS leaderboard revealed that explicit challenges successfully detected LLMs in 78.4% of cases, while implicit challenges were effective in 22.9% of instances. User studies validate the real-world applicability of our methods, with humans outperforming LLMs on explicit challenges (78% vs 22% success rate). Our framework unexpectedly revealed that many study participants were using LLMs to complete tasks, demonstrating its effectiveness in detecting both AI impostors and human misuse of AI tools. This work addresses the critical need for reliable, real-time LLM detection methods in high-stakes conversations.

📄 PDF Abstract BibTeX arXiv:2410.09569

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs

2026-03-21 · Taihei Shiotani, Masahiro Kaneko, Ayana Niwa, Yuki Maruyama 외 arxiv

Social biases reflected in language are inherently shaped by cultural norms, which vary significantly across regions and lead to diverse manifestations of stereotypes. Existing evaluations of social bias in large languag…

Who is a Better Player: LLM against LLM

2025-08-05 · Yingjie Zhou, Jiezhang Cao, Farong Wen, Li Xu 외 arxiv

Adversarial board games, as a paradigmatic domain of strategic reasoning and intelligence, have long served as both a popular competitive activity and a benchmark for evaluating artificial intelligence (AI) systems. Buil…

Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface

2025-01-16 · Andrey Labunets, Nishit V. Pandya, Ashish Hooda, Xiaohan Fu 외

We surface a new threat to closed-weight Large Language Models (LLMs) that enables an attacker to compute optimization-based prompt injections. Specifically, we characterize how an attacker can leverage the loss-like inf…

Persona Jailbreaking in Large Language Models

2026-01-23 · Jivnesh Sandhan, Fei Cheng, Tushar Sandhan, Yugo Murawaki arxiv

Large Language Models (LLMs) are increasingly deployed in domains such as education, mental health and customer support, where stable and consistent personas are critical for reliability. Yet, existing studies focus on n…

Assessing Adversarial Robustness of Large Language Models: An Empirical Study

2024-05-04 · Zeyu Yang, Zhao Meng, Xiaochen Zheng, Roger Wattenhofer

Large Language Models (LLMs) have revolutionized natural language processing, but their robustness against adversarial attacks remains a critical concern. We presents a novel white-box style attack approach that exposes …

Adversarial Robustnesstext-classificationText Classification