paper-with-me

홈 › Papers

Can Large Language Models Change User Preference Adversarially?

2023-01-05 · Varshini Subhash

Pretrained large language models (LLMs) are becoming increasingly powerful and ubiquitous in mainstream applications such as being a personal assistant, a dialogue model, etc. As these models become proficient in deducing user preferences and offering tailored assistance, there is an increasing concern about the ability of these models to influence, modify and in the extreme case manipulate user preference adversarially. The issue of lack of interpretability in these models in adversarial settings remains largely unsolved. This work tries to study adversarial behavior in user preferences from the lens of attention probing, red teaming and white-box analysis. Specifically, it provides a bird's eye view of existing literature, offers red teaming samples for dialogue models like ChatGPT and GODEL and probes the attention mechanism in the latter for non-adversarial and adversarial settings.

📄 PDF Abstract BibTeX arXiv:2302.10291

Code (0)

등록된 구현이 없습니다.

Tasks

Red Teaming

Similar Papers 제목 키워드 기반

Self-Consuming Generative Models with Adversarially Curated Data

2025-05-14 · Xiukun Wei, Xueru Zhang

Recent advances in generative models have made it increasingly difficult to distinguish real data from model-generated synthetic data. Using synthetic data for successive training of future model generations creates "sel…

Detecting the Adversarially-Learned Injection Attacks via Knowledge Graphs

2024-11-15 · 2024 2024 11 · Yaojun Hao*, 1, Haotian Wang2, Qingshan Zhao1 외

ABSTRACT: Over the past two decades, many studies have devoted a good deal of attention to detect injection attacks in recommender systems. However, most of the studies mainly focus on detecting the heuristically-generat…

Knowledge GraphsRecommendation Systems

PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges

2026-05-29 · Swastik Roy, Rajkumar Pujari, Tharindu Kumarage, Charith Peris 외 arxiv

LLM judges are increasingly used to evaluate open-ended responses, but their scores depend strongly on the rubrics that condition them. A vague rubric asking for a response to be ``helpful and factual'' can reward polish…

Adversarial Robustness

Novelty Learning via Collaborative Proximity Filtering

2016-10-21 · Arun Kumar, Paul Schrater

The vast majority of recommender systems model preferences as static or slowly changing due to observable user experience. However, spontaneous changes in user preferences are ubiquitous in many domains like media consum…

Recommendation Systems

Preference-Conditioned Language-Guided Abstraction

2024-02-05 · Andi Peng, Andreea Bobu, Belinda Z. Li, Theodore R. Sumers 외

Learning from demonstrations is a common way for users to teach robots, but it is prone to spurious feature correlations. Recent work constructs state abstractions, i.e. visual representations containing task-relevant fe…