paper-with-me

Papers

An Assessment of Model-On-Model Deception

2024-05-10 · Julius Heitkoetter, Michael Gerovitch, Laker Newhouse

The trustworthiness of highly capable language models is put at risk when they are able to produce deceptive outputs. Moreover, when models are vulnerable to deception it undermines reliability. In this paper, we introduce a method to investigate complex, model-on-model deceptive scenarios. We create a dataset of over 10,000 misleading explanations by asking Llama-2 7B, 13B, 70B, and GPT-3.5 to justify the wrong answer for questions in the MMLU. We find that, when models read these explanations, they are all significantly deceived. Worryingly, models of all capabilities are successful at misleading others, while more capable models are only slightly better at resisting deception. We recommend the development of techniques to detect and defend against deception.

📄 PDF Abstract BibTeX arXiv:2405.12999

Code (0)

등록된 구현이 없습니다.

Tasks

MMLUmodel

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Adam 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…

Similar Papers 제목 키워드 기반

AI Deception: A Survey of Examples, Risks, and Potential Solutions

2023-08-28 · Peter S. Park, Simon Goldstein, Aidan O'Gara, Michael Chen 외

This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first sur…

Survey

The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind

2025-09-23 · Caleb DeLeeuw, Gaurav Chawla, Aniket Sharma, Vanessa Dietze arxiv

We investigate strategic deception in large language models using two complementary testbeds: Secret Agenda (across 38 models) and Insider Trading compliance (via SAE architectures). Secret Agenda reliably induced lying …

LoRA-like Calibration for Multimodal Deception Detection using ATSFace Data

2023-09-04 · Shun-Wen Hsiao, Cheng-Yuan Sun

Recently, deception detection on human videos is an eye-catching techniques and can serve lots applications. AI model in this domain demonstrates the high accuracy, but AI tends to be a non-interpretable black box. We in…

Deception Detection

SVC 2025: the First Multimodal Deception Detection Challenge

2025-08-06 · Xun Lin, Xiaobao Guo, Taorui Wang, Yingjie Ma 외 arxiv

Deception detection is a critical task in real-world applications such as security screening, fraud prevention, and credibility assessment. While deep learning methods have shown promise in surpassing human-level perform…

Domain Generalization

Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal Learning

2023-03-09 · ICCV 2023 1 · Xiaobao Guo, Nithish Muthuchamy Selvaraj, Zitong Yu, Adams Wai-Kin Kong 외

Deception detection in conversations is a challenging yet important task, having pivotal applications in many fields such as credibility assessment in business, multimedia anti-frauds, and custom security. Despite this, …

Deception DetectionMulti-Task Learning