paper-with-me

홈 › Papers

DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior

2025-12-27 · Sadia Asif, Israel Antonio Rosales Laguan, Haris Khan, Shumaila Asif, Muneeb Asif arxiv

The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary labels and fail to capture the nuanced psychological and social mechanisms constituting manipulation. We introduce \textbf{DarkPatterns-LLM}, a comprehensive benchmark dataset and diagnostic framework for fine-grained assessment of manipulative content in LLM outputs across seven harm categories: Legal/Power, Psychological, Emotional, Physical, Autonomy, Economic, and Societal Harm. Our framework implements a four-layer analytical pipeline comprising Multi-Granular Detection (MGD), Multi-Scale Intent Analysis (MSIAN), Threat Harmonization Protocol (THP), and Deep Contextual Risk Alignment (DCRA). The dataset contains 401 meticulously curated examples with instruction-response pairs and expert annotations. Through evaluation of state-of-the-art models including GPT-4, Claude 3.5, and LLaMA-3-70B, we observe significant performance disparities (65.2\%--89.7\%) and consistent weaknesses in detecting autonomy-undermining patterns. DarkPatterns-LLM establishes the first standardized, multi-dimensional benchmark for manipulation detection in LLMs, offering actionable diagnostics toward more trustworthy AI systems.

📄 PDF Abstract BibTeX arXiv:2512.22470

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DarkBench: Benchmarking Dark Patterns in Large Language Models

2025-03-13 · Esben Kran, Hieu Minh "Jord" Nguyen, Akash Kundu, Sami Jawhar 외

We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language models (LLMs). Our benchmark comprises 660 p…

Benchmarking

BetXplain: An Explanation-Annotated Dataset for Detecting Manipulative Betting Advertisements on Social Media

2026-06-25 · MSVPJ Sathvik, Parmitha Vangapandu, Nishit Rane, Sathwik Narkedimilli 외 arxiv

The promotion of betting applications on social media platforms has increased significantly in recent years. Many of these advertisements use persuasive techniques that may mislead users, encourage risky behavior, and po…

LLM-based Detection of Manipulative Political Narratives

2026-05-14 · Sinclair Schneider, Florian Steuber, Gabi Dreo Rodosek arxiv

We present a new computational framework for detecting and structuring manipulative political narratives. A task that became more important due to the shift of political discussions to social media. One of the primary ch…

SELF-PERCEPT: Introspection Improves Large Language Models' Detection of Multi-Person Mental Manipulation in Conversations

2025-05-27 · Danush Khanna, Pratinav Seth, Sidhaarth Sredharan Murali, Aditya Kumar Guru 외

Mental manipulation is a subtle yet pervasive form of abuse in interpersonal communication, making its detection critical for safeguarding potential victims. However, due to manipulation's nuanced and context-specific na…

Detecting Mental Manipulation in Speech via Synthetic Multi-Speaker Dialogue

2026-01-13 · Run Chen, Wen Liang, Ziwei Gong, Lin Ai 외 arxiv

Mental manipulation, the strategic use of language to covertly influence or exploit others, is a newly emerging task in computational social reasoning. Prior work has focused exclusively on textual conversations, overloo…