paper-with-me

홈 › Papers

Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation

2025-11-13 · Fred Heiding, Simon Lermen arxiv

We present an end-to-end demonstration of how attackers can exploit AI safety failures to harm vulnerable populations: from jailbreaking LLMs to generate phishing content, to deploying those messages against real targets, to successfully compromising elderly victims. We systematically evaluated safety guardrails across six frontier LLMs spanning four attack categories, revealing critical failures where several models exhibited near-complete susceptibility to certain attack vectors. In a human validation study with 108 senior volunteers, AI-generated phishing emails successfully compromised 11\% of participants. Our work uniquely demonstrates the complete attack pipeline targeting elderly populations, highlighting that current AI safety measures fail to protect those most vulnerable to fraud. Beyond generating phishing content, LLMs enable attackers to overcome language barriers and conduct multi-turn trust-building conversations at scale, fundamentally transforming fraud economics. While some providers report voluntary counter-abuse efforts, we argue these remain insufficient.

📄 PDF Abstract BibTeX arXiv:2511.11759

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deep Reinforcement Learning for Detecting Malicious Websites

2019-05-22 · Moitrayee Chatterjee, Akbar Siami Namin

Phishing is the simplest form of cybercrime with the objective of baiting people into giving away delicate information such as individually recognizable data, banking and credit card details, or even credentials and pass…

Deep Reinforcement LearningPhishing Website Detectionreinforcement-learningReinforcement Learning+1

Analyzing Social and Stylometric Features to Identify Spear phishing Emails

2014-06-14 · Prateek Dewan, Anand Kashyap, Ponnurangam Kumaraguru

Spear phishing is a complex targeted attack in which, an attacker harvests information about the victim prior to the attack. This information is then used to create sophisticated, genuine-looking attack vectors, drawing …

Phishing Detection Using Machine Learning Techniques

2020-09-20 · Vahid Shahrivari, Mohammad Mahdi Darabi, Mohammad Izadi

The Internet has become an indispensable part of our life, However, It also has provided opportunities to anonymously perform malicious activities like Phishing. Phishers try to deceive their victims by social engineerin…

BIG-bench Machine Learning

Finding Phish in a Haystack: A Pipeline for Phishing Classification on Certificate Transparency Logs

2021-06-23 · Arthur Drichel, Vincent Drury, Justus von Brandt, Ulrike Meyer

Current popular phishing prevention techniques mainly utilize reactive blocklists, which leave a ``window of opportunity'' for attackers during which victims are unprotected. One possible approach to shorten this window …

A new weighted ensemble model for phishing detection based on feature selection

2022-12-15 · Farnoosh Shirani Bidabadi, Shuaifang Wang

A phishing attack is a sort of cyber assault in which the attacker sends fake communications to entice a human victim to provide personal information or credentials. Phishing website identification can assist visitors in…

feature selection