paper-with-me

홈 › Papers

PRISON: Unmasking the Criminal Potential of Large Language Models

2025-06-19 · Xinyi Wu, Geng Hong, Pei Chen, Yueyue Chen, Xudong Pan, Min Yang

As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research overlooked the systematic understanding and assessment of their criminal capability in realistic interactions. We propose a unified framework PRISON, to quantify LLMs' criminal potential across five dimensions: False Statements, Frame-Up, Psychological Manipulation, Emotional Disguise, and Moral Disengagement. Using structured crime scenarios adapted from classic films, we evaluate both criminal potential and anti-crime ability of LLMs via role-play. Results show that state-of-the-art LLMs frequently exhibit emergent criminal tendencies, such as proposing misleading statements or evasion tactics, even without explicit instructions. Moreover, when placed in a detective role, models recognize deceptive behavior with only 41% accuracy on average, revealing a striking mismatch between conducting and detecting criminal behavior. These findings underscore the urgent need for adversarial robustness, behavioral alignment, and safety mechanisms before broader LLM deployment.

📄 PDF Abstract BibTeX arXiv:2506.16150

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

Femicide Laws, Unilateral Divorce, and Abortion Decriminalization Fail to Stop Women's Killings in Mexico

2024-07-09 · Roxana Gutiérrez-Romero

This paper evaluates the effectiveness of femicide laws in combating gender-based killings of women, a major cause of premature female mortality globally. Focusing on Mexico, a pioneer in adopting such legislation, the p…

Rethinking recidivism through a causal lens

2020-11-19 · Vik Shirvaikar, Choudur Lakshminarayan

Predictive modeling of criminal recidivism, or whether people will re-offend in the future, has a long and contentious history. Modern causal inference methods allow us to move beyond prediction and target the "treatment…

BIG-bench Machine LearningCausal InferenceSelection bias

Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals

2024-05-08 · Joshua Clymer, Caden Juang, Severin Field

Like a criminal under investigation, Large Language Models (LLMs) might pretend to be aligned while evaluated and misbehave when they have a good opportunity. Can current interpretability methods catch these 'alignment f…

AI-Powered Facial Mask Removal Is Not Suitable For Identification

2026-03-29 · Emily A Cooper, Hany Farid arxiv

Recently, crowd-sourced online criminal investigations have used generative-AI to enhance low-quality visual evidence. In one high-profile case, social-media users circulated an "AI-unmasked" image of a federal agent inv…

CAIL2018: A Large-Scale Legal Dataset for Judgment Prediction

2018-07-04 · Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu 외

In this paper, we introduce the \textbf{C}hinese \textbf{AI} and \textbf{L}aw challenge dataset (CAIL2018), the first large-scale Chinese legal dataset for judgment prediction. \dataset contains more than $2.6$ million c…

ArticlesPredictionText Classification