paper-with-me

Papers

Misusing Tools in Large Language Models With Visual Adversarial Examples

2023-10-04 · Xiaohan Fu, Zihan Wang, Shuheng Li, Rajesh K. Gupta, Niloofar Mireshghallah, Taylor Berg-Kirkpatrick, Earlence Fernandes

Large Language Models (LLMs) are being enhanced with the ability to use tools and to process multiple modalities. These new capabilities bring new benefits and also new security risks. In this work, we show that an attacker can use visual adversarial examples to cause attacker-desired tool usage. For example, the attacker could cause a victim LLM to delete calendar events, leak private conversations and book hotels. Different from prior work, our attacks can affect the confidentiality and integrity of user resources connected to the LLM while being stealthy and generalizable to multiple input prompts. We construct these attacks using gradient-based adversarial training and characterize performance along multiple dimensions. We find that our adversarial images can manipulate the LLM to invoke tools following real-world syntax almost always (~98%) while maintaining high similarity to clean images (~0.9 SSIM). Furthermore, using human scoring and automated metrics, we find that the attacks do not noticeably affect the conversation (and its semantics) between the user and the LLM.

📄 PDF Abstract BibTeX arXiv:2310.03185

Code (1)

ZihanWangKi/VLMToolMisuse 공식 구현 pytorch

Tasks

SSIM

Similar Papers 제목 키워드 기반

Stop Misusing t-SNE and UMAP for Visual Analytics

2025-06-10 · Hyeon Jeon, Jeongin Park, Sungbok Shin, Jinwook Seo

Misuses of t-SNE and UMAP in visual analytics have become increasingly common. For example, although t-SNE and UMAP projections often do not faithfully reflect true distances between clusters, practitioners frequently us…

Multilingual Hate Speech and Offensive Content Detection using Modified Cross-entropy Loss

2022-02-05 · Arka Mitra, Priyanshu Sankhala

The number of increased social media users has led to a lot of people misusing these platforms to spread offensive content and use hate speech. Manual tracking the vast amount of posts is impractical so it is necessary t…

To ChatGPT, or not to ChatGPT: That is the question!

2023-04-04 · Alessandro Pegoraro, Kavita Kumari, Hossein Fereidooni, Ahmad-Reza Sadeghi

ChatGPT has become a global sensation. As ChatGPT and other Large Language Models (LLMs) emerge, concerns of misusing them in various ways increase, such as disseminating fake news, plagiarism, manipulating public opinio…

Text Detection

GAMMS: Graph based Adversarial Multiagent Modeling Simulator

2026-02-04 · Rohan Patil, Jai Malegaonkar, Xiao Jiang, Andre Dion 외 arxiv

As intelligent systems and multi-agent coordination become increasingly central to real-world applications, there is a growing need for simulation tools that are both scalable and accessible. Existing high-fidelity simul…

Towards Evaluating Gaussian Blurring in Perceptual Hashing as a Facial Image Filter

2020-02-01 · Yigit Alparslan, Ken Alparslan, Mannika Kshettry, Louis Kratz

With the growth in social media, there is a huge amount of images of faces available on the internet. Often, people use other people's pictures on their own profile. Perceptual hashing is often used to detect whether two…

Image Croppingtext annotation