paper-with-me

홈 › Papers

Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking

2025-07-06 · Aldan Creo, Raul Castro Fernandez, Manuel Cebrian arxiv

As large language models (LLMs) become increasingly deployed, understanding the complexity and evolution of jailbreaking strategies is critical for AI safety. We present a mass-scale empirical analysis of jailbreak complexity across over 2 million real-world conversations from diverse platforms, including dedicated jailbreaking communities and general-purpose chatbots. Using a range of complexity metrics spanning probabilistic measures, lexical diversity, compression ratios, and cognitive load indicators, we find that jailbreak attempts do not exhibit significantly higher complexity than normal conversations. This pattern holds consistently across specialized jailbreaking communities and general user populations, suggesting practical bounds on attack sophistication. Temporal analysis reveals that while user attack toxicity and complexity remains stable over time, assistant response toxicity has decreased, indicating improving safety mechanisms. The absence of power-law scaling in complexity distributions further points to natural limits on jailbreak development. Our findings challenge the prevailing narrative of an escalating arms race between attackers and defenders, instead suggesting that LLM safety evolution is bounded by human ingenuity constraints while defensive measures continue advancing. Our results highlight critical information hazards in academic jailbreak disclosure, as sophisticated attacks exceeding current complexity baselines could disrupt the observed equilibrium and enable widespread harm before defensive adaptation.

📄 PDF Abstract BibTeX arXiv:2507.08014

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild

2024-09-05 · Yuntian Deng, Wenting Zhao, Jack Hessel, Xiang Ren 외

The increasing availability of real-world conversation data offers exciting opportunities for researchers to study user-chatbot interactions. However, the sheer volume of this data makes manually examining individual con…

Chatbot

ReXInTheWild: A Unified Benchmark for Medical Photograph Understanding

2026-03-19 · Oishi Banerjee, Sung Eun Kim, Alexandra N. Willauer, Julius M. Kernbach 외 arxiv

Everyday photographs taken with ordinary cameras are already widely used in telemedicine and other online health conversations, yet no comprehensive benchmark evaluates whether vision-language models can interpret their …

Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection

2026-02-11 · J Alex Corll arxiv

Multi-turn prompt injection attacks distribute malicious intent across multiple conversation turns, exploiting the assumption that each turn is evaluated independently. While single-turn detection has been extensively st…

Deep Learning Face Attributes in the Wild

2014-11-28 · ICCV 2015 12 · Ziwei Liu, Ping Luo, Xiaogang Wang, Xiaoou Tang

Predicting face attributes in the wild is challenging due to complex face variations. We propose a novel deep learning framework for attribute prediction in the wild. It cascades two CNNs, LNet and ANet, which are fine-t…

AttributeDeep LearningFacial Attribute Classification

Sensor-Based Satellite IoT for Early Wildfire Detection

2021-09-22 · How-Hang Liu, Ronald Y. Chang, Yi-Ying Chen, I-Kang Fu

Frequent and severe wildfires have been observed lately on a global scale. Wildfires not only threaten lives and properties, but also pose negative environmental impacts that transcend national boundaries (e.g., greenhou…