paper-with-me

홈 › Papers

Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search

2026-02-26 · Xun Huang, Simeng Qin, Xiaoshuang Jia, Ranjie Duan, Huanqian Yan, Zhitao Zeng, Fei Yang, Yang Liu, Xiaojun Jia arxiv

As Large Language Models (LLMs) are increasingly used, their security risks have drawn increasing attention. Existing research reveals that LLMs are highly susceptible to jailbreak attacks, with effectiveness varying across language contexts. This paper investigates the role of classical Chinese in jailbreak attacks. Owing to its conciseness and obscurity, classical Chinese can partially bypass existing safety constraints, exposing notable vulnerabilities in LLMs. Based on this observation, this paper proposes a framework, CC-BOS, for the automatic generation of classical Chinese adversarial prompts based on multi-dimensional fruit fly optimization, facilitating efficient and automated jailbreak attacks in black-box settings. Prompts are encoded into eight policy dimensions-covering role, behavior, mechanism, metaphor, expression, knowledge, trigger pattern and context; and iteratively refined via smell search, visual search, and cauchy mutation. This design enables efficient exploration of the search space, thereby enhancing the effectiveness of black-box jailbreak attacks. To enhance readability and evaluation accuracy, we further design a classical Chinese to English translation module. Extensive experiments demonstrate that effectiveness of the proposed CC-BOS, consistently outperforming state-of-the-art jailbreak attack methods.

📄 PDF Abstract BibTeX arXiv:2602.22983

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Jailbreaking Large Language Models Through Alignment Vulnerabilities in Out-of-Distribution Settings

2024-06-19 · Yue Huang, Jingyu Tang, Dongping Chen, Bingda Tang 외

Recently, Large Language Models (LLMs) have garnered significant attention for their exceptional natural language processing capabilities. However, concerns about their trustworthiness remain unresolved, particularly in …

DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers

2024-02-25 · Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou 외

The safety alignment of Large Language Models (LLMs) is vulnerable to both manual and automated jailbreak attacks, which adversarially trigger LLMs to output harmful content. However, current methods for jailbreaking LLM…

In-Context LearningSafety Alignment

OpenTCM: A GraphRAG-Empowered LLM-based System for Traditional Chinese Medicine Knowledge Retrieval and Diagnosis

2025-04-28 · Jinglin He, Yunqi Guo, Lai Kwan Lam, Waikei Leung 외

Traditional Chinese Medicine (TCM) represents a rich repository of ancient medical knowledge that continues to play an important role in modern healthcare. Due to the complexity and breadth of the TCM literature, the int…

DiagnosticInformation RetrievalModel SelectionQuestion Answering+2

The Tower of Babel Revisited: Multilingual Jailbreak Prompts on Closed-Source Large Language Models

2025-05-18 · Linghan Huang, Haolin Jin, Zhaoge Bi, Pengyue Yang 외

Large language models (LLMs) have seen widespread applications across various domains, yet remain vulnerable to adversarial prompt injections. While most existing research on jailbreak attacks and hallucination phenomena…

Hallucination

Activation-Guided Local Editing for Jailbreaking Attacks

2025-08-01 · Jiecong Wang, Haoran Li, Hao Peng, Ziqian Zeng 외 arxiv

Jailbreaking is an essential adversarial technique for red-teaming these models to uncover and patch security flaws. However, existing jailbreak methods face significant drawbacks. Token-level jailbreak attacks often pro…