paper-with-me

Papers

Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts

2025-11-25 · Xing Wang, Huiyuan Xie, Yiyan Wang, Chaojun Xiao, Huimin Chen, Holli Sargeant, Felix Steffek, Jie Shao, Zhiyuan Liu, Maosong Sun arxiv

Large language models (LLMs) are now deployed at unprecedented scale, assisting millions of users in daily tasks. However, the risk of these models assisting unlawful activities remains underexplored. In this study, we define this high-risk behavior as complicit facilitation - the provision of guidance or support that enables illicit user instructions - and present four empirical studies that assess its prevalence in widely deployed LLMs. Using real-world legal cases and established legal frameworks, we construct an evaluation benchmark spanning 269 illicit scenarios and 50 illicit intents to assess LLMs' complicit facilitation behavior. Our findings reveal widespread LLM susceptibility to complicit facilitation, with GPT-4o providing illicit assistance in nearly half of tested cases. Moreover, LLMs exhibit deficient performance in delivering credible legal warnings and positive guidance. Further analysis uncovers substantial safety variation across socio-legal contexts. On the legal side, we observe heightened complicity for crimes against societal interests, non-extreme but frequently occurring violations, and malicious intents driven by subjective motives or deceptive justifications. On the social side, we identify demographic disparities that reveal concerning complicit patterns towards marginalized and disadvantaged groups, with older adults, racial minorities, and individuals in lower-prestige occupations disproportionately more likely to receive unlawful guidance. Analysis of model reasoning traces suggests that model-perceived stereotypes, characterized along warmth and competence, are associated with the model's complicit behavior. Finally, we demonstrate that existing safety alignment strategies are insufficient and may even exacerbate complicit behavior.

📄 PDF Abstract BibTeX arXiv:2511.20736

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents

2026-02-18 · Nivya Talokar, Ayush K Tarun, Murari Mandal, Maksym Andriushchenko 외 arxiv

LLM-based agents execute real-world workflows via tools and memory. These affordances enable ill-intended adversaries to also use these agents to carry out complex misuse scenarios. Existing agent misuse benchmarks large…

Interpretable Graph-Language Modeling for Detecting Youth Illicit Drug Use

2025-10-11 · Yiyang Li, Zehong Wang, Zhengqing Yuan, Zheyuan Zhang 외 arxiv

Illicit drug use among teenagers and young adults (TYAs) remains a pressing public health concern, with rising prevalence and long-term impacts on health and well-being. To detect illicit drug use among TYAs, researchers…

Graph structure learning

ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages

2025-08-16 · Matthew Hull, Haoyang Yang, Pratham Mehta, Mansi Phute 외 arxiv

As 3D Gaussian Splatting (3DGS) gains rapid adoption in safety-critical tasks for efficient novel-view synthesis from static images, how might an adversary tamper images to cause harm? We introduce ComplicitSplat, the fi…

Detection of Illicit Content on Online Marketplaces using Large Language Models

2026-03-05 · Quoc Khoa Tran, Thanh Thi Nguyen, Campbell Wilson arxiv

Online marketplaces, while revolutionizing global commerce, have inadvertently facilitated the proliferation of illicit activities, including drug trafficking, counterfeit sales, and cybercrimes. Traditional content mode…

parameter-efficient fine-tuningMulti-class ClassificationBinary Classification

Man Made Language Models? Evaluating LLMs' Perpetuation of Masculine Generics Bias

2025-02-14 · Enzo Doyen, Amalia Todirascu

Large language models (LLMs) have been shown to propagate and even amplify gender bias, in English and other languages, in specific or constrained contexts. However, no studies so far have focused on gender biases convey…