paper-with-me

Papers

Hidden in Plain Sight: Exploring Chat History Tampering in Interactive Language Models

2024-05-30 · Cheng'an Wei, Yue Zhao, Yujia Gong, Kai Chen, Lu Xiang, Shenchen Zhu

Large Language Models (LLMs) such as ChatGPT and Llama have become prevalent in real-world applications, exhibiting impressive text generation performance. LLMs are fundamentally developed from a scenario where the input data remains static and unstructured. To behave interactively, LLM-based chat systems must integrate prior chat history as context into their inputs, following a pre-defined structure. However, LLMs cannot separate user inputs from context, enabling chat history tampering. This paper introduces a systematic methodology to inject user-supplied history into LLM conversations without any prior knowledge of the target model. The key is to utilize prompt templates that can well organize the messages to be injected, leading the target LLM to interpret them as genuine chat history. To automatically search for effective templates in a WebUI black-box setting, we propose the LLM-Guided Genetic Algorithm (LLMGA) that leverages an LLM to generate and iteratively optimize the templates. We apply the proposed method to popular real-world LLMs including ChatGPT and Llama-2/3. The results show that chat history tampering can enhance the malleability of the model's behavior over time and greatly influence the model output. For example, it can improve the success rate of disallowed response elicitation up to 97% on ChatGPT. Our findings provide insights into the challenges associated with the real-world deployment of interactive LLMs.

📄 PDF Abstract BibTeX arXiv:2405.20234

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

MCP: Self-supervised Pre-training for Personalized Chatbots with Multi-level Contrastive Sampling

2022-10-17 · Zhaoheng Huang, Zhicheng Dou, Yutao Zhu, Zhengyi Ma

Personalized chatbots focus on endowing the chatbots with a consistent personality to behave like real users and further act as personal assistants. Previous studies have explored generating implicit user profiles from t…

Response GenerationSelf-Supervised Learning

Exploring the Comprehension of ChatGPT in Traditional Chinese Medicine Knowledge

2024-03-14 · Li Yizhen, Huang Shaohan, Qi Jiaxing, Quan Lei 외

No previous work has studied the performance of Large Language Models (LLMs) in the context of Traditional Chinese Medicine (TCM), an essential and distinct branch of medical knowledge with a rich history. To bridge this…

Multiple-choice

Invisible Strings: Deriving Puppetry Principles and their Hidden Connections to Robot Behavior Design

2026-07-03 · Claire Lewis, Sawyer Collins, Alyssa Hanson, Johanna Smith 외 arxiv

When designing robots' nonverbal behaviors, many researchers have turned to arts-based insights, such as Disney's Animation Principles. Yet, while these principles bear key insights into the design of like-life character…

Exploring the Expansion History of the Universe

2002-08-28 · Eric V. Linder

Exploring the recent expansion history of the universe promises insights into the cosmological model, the nature of dark energy, and potentially clues to high energy physics theories and gravitation. We examine the exten…

WikiSTAR: A System for Shedding Light on the Hidden History of Scientific Wikipedia Articles

2026-07-14 · Omer Ehrlich, Nitzan Barzilay, Rona Aviram, Tom Hope arxiv

Wikipedia plays a key role in shaping public understanding of science, and its openly accessible revision history is a unique record of how scientific knowledge evolves over time. Yet scientifically meaningful revisions …