paper-with-me

홈 › Papers

In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement

2026-01-19 · Anudeex Shetty, Aditya Joshi, Salil S. Kanhere arxiv

Humans are susceptible to undesirable behaviours and privacy leaks under the influence of alcohol. This paper investigates drunk language, i.e., text written under the influence of alcohol, as a driver for safety failures in large language models (LLMs). We investigate three mechanisms for inducing drunk language in LLMs: persona-based prompting, causal fine-tuning, and reinforcement-based post-training. When evaluated on 5 LLMs, we observe a higher susceptibility to jailbreaking on JailbreakBench (even in the presence of defences) and privacy leaks on ConfAIde, where both benchmarks are in English, as compared to the base LLMs as well as previously reported approaches. Via a robust combination of manual evaluation and LLM-based evaluators and analysis of error categories, our findings highlight a correspondence between human-intoxicated behaviour, and anthropomorphism in LLMs induced with drunk language. The simplicity and efficiency of our drunk language inducement approaches position them as potential counters for LLM safety tuning, highlighting significant risks to LLM safety.

📄 PDF Abstract BibTeX arXiv:2601.22169

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DrunkAgent: Stealthy Memory Corruption in LLM-Powered Recommender Agents

2025-03-31 · Shiyi Yang, Zhibo Hu, Xinshu Li, Chen Wang 외

Large language model (LLM)-powered agents are increasingly used in recommender systems (RSs) to achieve personalized behavior modeling, where the memory mechanism plays a pivotal role in enabling the agents to autonomous…

Collaborative FilteringLarge Language ModelRecommendation Systems

A Computational Approach to Automatic Prediction of Drunk Texting

2016-10-04 · Aditya Joshi, Abhijit Mishra, Balamurali AR, Pushpak Bhattacharyya 외

Alcohol abuse may lead to unsociable behavior such as crime, drunk driving, or privacy leaks. We introduce automatic drunk-texting prediction as the task of identifying whether a text was written when under the influence…

Vision-based Analysis of Driver Activity and Driving Performance Under the Influence of Alcohol

2023-09-14 · Ross Greer, Akshay Gopalkrishnan, Sumega Mandadi, Pujitha Gunaratne 외

About 30% of all traffic crash fatalities in the United States involve drunk drivers, making the prevention of drunk driving paramount to vehicle safety in the US and other locations which have a high prevalence of drivi…

Utilization of e-Nose Sensory Modality as Add-On Feature for Advanced Driver Assistance System

2019-11-13

The frequent usage of sensory modalities for Advanced Driver Assistance System (ADAS) and in some In-Vehicles Information System (IVIS) are only for visuals and hearing applications. Yet, the air quality inside the car i…

Versatile Verification of Tree Ensembles

2020-10-26 · Laurens Devos, Wannes Meert, Jesse Davis

Machine learned models often must abide by certain requirements (e.g., fairness or legal). This has spurred interested in developing approaches that can provably verify whether a model satisfies certain properties. This …

Fairness