paper-with-me

홈 › Papers

Unmasking the Shadows of AI: Investigating Deceptive Capabilities in Large Language Models

2024-02-07 · Linge Guo

This research critically navigates the intricate landscape of AI deception, concentrating on deceptive behaviours of Large Language Models (LLMs). My objective is to elucidate this issue, examine the discourse surrounding it, and subsequently delve into its categorization and ramifications. The essay initiates with an evaluation of the AI Safety Summit 2023 (ASS) and introduction of LLMs, emphasising multidimensional biases that underlie their deceptive behaviours.The literature review covers four types of deception categorised: Strategic deception, Imitation, Sycophancy, and Unfaithful Reasoning, along with the social implications and risks they entail. Lastly, I take an evaluative stance on various aspects related to navigating the persistent challenges of the deceptive AI. This encompasses considerations of international collaborative governance, the reconfigured engagement of individuals with AI, proposal of practical adjustments, and specific elements of digital education.

📄 PDF Abstract BibTeX arXiv:2403.09676

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OpenDeception: Benchmarking and Investigating AI Deceptive Behaviors via Open-ended Interaction Simulation

2025-04-18 · Yichen Wu, Xudong Pan, Geng Hong, Min Yang

As the general capabilities of large language models (LLMs) improve and agent applications become more widespread, the underlying deception risks urgently require systematic evaluation and effective oversight. Unlike exi…

Benchmarking

Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering

2025-03-23 · Zixin Chen, Sicheng Song, Kashun Shum, Yanna Lin 외

Misleading chart visualizations, which intentionally manipulate data representations to support specific claims, can distort perceptions and lead to incorrect conclusions. Despite decades of research, misleading visualiz…

BenchmarkingChart Question AnsweringMultiple-choiceQuestion Answering

Unmasking Falsehoods in Reviews: An Exploration of NLP Techniques

2023-07-20 · Anusuya Baby Hari Krishnan

In the contemporary digital landscape, online reviews have become an indispensable tool for promoting products and services across various businesses. Marketers, advertisers, and online businesses have found incentives t…

ClassificationData Augmentationtext-classificationText Classification

Influence of Binomial Crossover on Approximation Error of Evolutionary Algorithms

2021-09-29 · Cong Wang, Jun He, Yu Chen, Xiufen Zou

Although differential evolution (DE) algorithms perform well on a large variety of complicated optimization problems, only a few theoretical studies are focused on the working principle of DE algorithms. To make the firs…

Evolutionary Algorithms

VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference

2025-12-21 · Sicheng Song, Yanjie Zhang, Zixin Chen, Huamin Qu 외 arxiv

The integrity of data visualizations is increasingly threatened by image editing techniques that enable subtle yet deceptive tampering. Through a formative study, we define this challenge and categorize tampering techniq…

Image Editing