paper-with-me

Papers

Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities

2025-07-22 · Atil Samancioglu arxiv

Large Language Models (LLMs) demonstrate complex responses to threat-based manipulations, revealing both vulnerabilities and unexpected performance enhancement opportunities. This study presents a comprehensive analysis of 3,390 experimental responses from three major LLMs (Claude, GPT-4, Gemini) across 10 task domains under 6 threat conditions. We introduce a novel threat taxonomy and multi-metric evaluation framework to quantify both negative manipulation effects and positive performance improvements. Results reveal systematic vulnerabilities, with policy evaluation showing the highest metric significance rates under role-based threats, alongside substantial performance enhancements in numerous cases with effect sizes up to +1336%. Statistical analysis indicates systematic certainty manipulation (pFDR < 0.0001) and significant improvements in analytical depth and response quality. These findings have dual implications for AI safety and practical prompt engineering in high-stakes applications.

📄 PDF Abstract BibTeX arXiv:2507.21133

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models

2024-09-20 · Hao Cheng, Erjia Xiao, Chengyuan Yu, Zhao Yao 외

Recently, driven by advancements in Multimodal Large Language Models (MLLMs), Vision Language Action Models (VLAMs) are being proposed to achieve better performance in open-vocabulary scenarios for robotic manipulation t…

Vision-Language-Action

DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing

2025-08-04 · Ko-Wei Chuang, Hen-Hsen Huang, Tsai-Yen Li arxiv

As large language models (LLMs) and generative AI become increasingly integrated into customer service and moderation applications, adversarial threats emerge from both external manipulations and internal label corruptio…

Realistic threat perception drives intergroup conflict: A causal, dynamic analysis using generative-agent simulations

2025-12-18 · Suhaib Abdurahman, Farzan Karimi-Malekabadi, Chenxiao Yu, Nour S. Kteily 외 arxiv

Human conflict is often attributed to threats against material conditions and symbolic values, yet it remains unclear how they interact and which dominates. Progress is limited by weak causal control, ethical constraints…

Efficacy of Utilizing Large Language Models to Detect Public Threat Posted Online

2023-12-29 · Taeksoo Kwon, Connor Kim

This paper examines the efficacy of utilizing large language models (LLMs) to detect public threats posted online. Amid rising concerns over the spread of threatening rhetoric and advance notices of violence, automated c…

VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference

2025-12-21 · Sicheng Song, Yanjie Zhang, Zixin Chen, Huamin Qu 외 arxiv

The integrity of data visualizations is increasingly threatened by image editing techniques that enable subtle yet deceptive tampering. Through a formative study, we define this challenge and categorize tampering techniq…

Image Editing