paper-with-me

홈 › Papers

SearchAttack: Red-Teaming LLMs against Knowledge-to-Action Threats under Online Web Search

2026-01-07 · Yu Yan, Sheng Sun, Mingfeng Li, Zheming Yang, Chiwei Zhu, Fei Ma, Benfeng Xu, Min Liu, Qi Li arxiv

Recently, people have suffered from LLM hallucination and have become increasingly aware of the reliability gap of LLMs in open and knowledge-intensive tasks. As a result, they have increasingly turned to search-augmented LLMs to mitigate this issue. However, LLM-driven search also becomes an attractive target for misuse. Once the returned content directly contains targeted, ready-to-use harmful instructions or takeaways for users, it becomes difficult to withdraw or undo such exposure. To investigate LLMs' unsafe search behavior issues, we first propose \textbf{\textit{SearchAttack}} for red-teaming, which (1) rephrases harmful semantics via dense and benign knowledge to evade direct in-context decoding, thus eliciting unsafe information retrieval, (2) stress-tests LLMs' reward-chasing bias by steering them to synthesize unsafe retrieved content. We also curate an emergent, domain-specific illicit activity benchmark for search-based threat assessment, and introduce a fact-checking framework to ground and quantify harm in both offline and online attack settings. Extensive experiments are conducted to red-team the search-augmented LLMs for responsible vulnerability assessment. Empirically, SearchAttack demonstrates strong effectiveness in attacking these systems. We also find that LLMs without web search can still be steered into harmful content output due to their information-seeking stereotypical behaviors.

📄 PDF Abstract BibTeX arXiv:2601.04093

Code (0)

등록된 구현이 없습니다.

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

Automated Red Teaming with GOAT: the Generative Offensive Agent Tester

2024-10-02 · Maya Pavlova, Erik Brinkman, Krithika Iyer, Vitor Albiero 외

Red teaming assesses how large language models (LLMs) can produce content that violates norms, policies, and rules set during their safety training. However, most existing automated methods in the literature are not repr…

Red Teaming

Attack Prompt Generation for Red Teaming and Defending Large Language Models

2023-10-19 · Boyi Deng, Wenjie Wang, Fuli Feng, Yang Deng 외

Large language models (LLMs) are susceptible to red teaming attacks, which can induce LLMs to generate harmful content. Previous research constructs attack prompts via manual or automatic methods, which have their own li…

In-Context LearningRed Teaming

Learning diverse attacks on large language models for robust red-teaming and safety tuning

2024-05-28 · Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre 외

Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing effective protection against many modes of…

DiversityLanguage ModelingLanguage ModellingRed Teaming

Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models

2025-03-03 · Alberto Purpura, Sahil Wadhwa, Jesse Zymet, Akshay Gupta 외

The rapid growth of Large Language Models (LLMs) presents significant privacy, security, and ethical concerns. While much research has proposed methods for defending LLM systems against misuse by malicious actors, resear…

Red TeamingSurvey

CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge

2024-04-10 · Yu Ying Chiu, Liwei Jiang, Maria Antoniak, Chan Young Park 외

Frontier large language models (LLMs) are developed by researchers and practitioners with skewed cultural backgrounds and on datasets with skewed sources. However, LLMs' (lack of) multicultural knowledge cannot be effect…

Red Teaming