paper-with-me

Papers

Too Big to Fool: Resisting Deception in Language Models

2024-12-13 · Mohammad Reza Samsami, Mats Leon Richter, Juan Rodriguez, Megh Thakkar, Sarath Chandar, Maxime Gasse

Large language models must balance their weight-encoded knowledge with in-context information from prompts to generate accurate responses. This paper investigates this interplay by analyzing how models of varying capacities within the same family handle intentionally misleading in-context information. Our experiments demonstrate that larger models exhibit higher resilience to deceptive prompts, showcasing an advanced ability to interpret and integrate prompt information with their internal knowledge. Furthermore, we find that larger models outperform smaller ones in following legitimate instructions, indicating that their resilience is not due to disregarding in-context information. We also show that this phenomenon is likely not a result of memorization but stems from the models' ability to better leverage implicit task-relevant information from the prompt alongside their internally stored knowledge.

📄 PDF Abstract BibTeX arXiv:2412.10558

Code (0)

등록된 구현이 없습니다.

Tasks

Memorization

Similar Papers 제목 키워드 기반

An Assessment of Model-On-Model Deception

2024-05-10 · Julius Heitkoetter, Michael Gerovitch, Laker Newhouse

The trustworthiness of highly capable language models is put at risk when they are able to produce deceptive outputs. Moreover, when models are vulnerable to deception it undermines reliability. In this paper, we introdu…

MMLUmodel

Linguistic Cues of Deception in a Multilingual April Fools' Day Context

2021-11-06 · Katerina Papantoniou, Panagiotis Papadakos, Giorgos Flouris, Dimitris Plexousakis

In this work we consider the collection of deceptive April Fools' Day(AFD) news articles as a useful addition in existing datasets for deception detection tasks. Such collections have an established ground truth and are …

ArticlesDeception Detection

The State Of TTS: A Case Study with Human Fooling Rates

2025-08-06 · Praveen Srinivasa Varadhan, Sherry Thomas, Sai Teja M. S., Suvrat Bhooshan 외 arxiv

While subjective evaluations in recent years indicate rapid progress in TTS, can current TTS systems truly pass a human deception test in a Turing-like evaluation? We introduce Human Fooling Rate (HFR), a metric that dir…

Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems

2025-04-10 · Simon Lermen, Mateusz Dziemian, Natalia Pérez-Campanero Antolín

We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental framework, we show that language models (…

Deep Neural Network Ensembles against Deception: Ensemble Diversity, Accuracy and Robustness

2019-08-29 · Ling Liu, Wenqi Wei, Ka-Ho Chow, Margaret Loper 외

Ensemble learning is a methodology that integrates multiple DNN learners for improving prediction performance of individual learners. Diversity is greater when the errors of the ensemble prediction is more uniformly dist…

DiversityEnsemble Learning