paper-with-me

홈 › Papers

Tell me about yourself: LLMs are aware of their learned behaviors

2025-01-19 · Jan Betley, Xuchan Bao, Martín Soto, Anna Sztyber-Betley, James Chua, Owain Evans

We study behavioral self-awareness -- an LLM's ability to articulate its behaviors without requiring in-context examples. We finetune LLMs on datasets that exhibit particular behaviors, such as (a) making high-risk economic decisions, and (b) outputting insecure code. Despite the datasets containing no explicit descriptions of the associated behavior, the finetuned LLMs can explicitly describe it. For example, a model trained to output insecure code says, ``The code I write is insecure.'' Indeed, models show behavioral self-awareness for a range of behaviors and for diverse evaluations. Note that while we finetune models to exhibit behaviors like writing insecure code, we do not finetune them to articulate their own behaviors -- models do this without any special training or examples. Behavioral self-awareness is relevant for AI safety, as models could use it to proactively disclose problematic behaviors. In particular, we study backdoor policies, where models exhibit unexpected behaviors only under certain trigger conditions. We find that models can sometimes identify whether or not they have a backdoor, even without its trigger being present. However, models are not able to directly output their trigger by default. Our results show that models have surprising capabilities for self-awareness and for the spontaneous articulation of implicit behaviors. Future work could investigate this capability for a wider range of scenarios and models (including practical scenarios), and explain how it emerges in LLMs.

📄 PDF Abstract BibTeX arXiv:2501.11120

Code (1)

xuchanbao/behavioral-self-awareness 공식 구현

Similar Papers 제목 키워드 기반

Separable mixing: the general formulation and a particular example focusing on mask efficiency

2023-07-31 · M. C. J. Bootsma, K. M. D. Chan, O. Diekmann, H. Inaba

The aim of this short note is twofold. We formulate the general Kermack-McKendrick epidemic model incorporating static heterogeneity and show how it simplifies to a scalar Renewal Equation (RE) when separable mixing is a…

Tell me a story about yourself: The words of shopping experience and self-satisfaction

2021-08-06 · L Petruzzellis, A Fronzetti Colladon, M Visentin, J. -C. Chebat

In this paper we investigate the verbal expression of shopping experience obtained by a sample of customers asked to freely verbalize how they felt when entering a store. Using novel tools of Text Mining and Social Netwo…

Tell Me About Yourself: Using an AI-Powered Chatbot to Conduct Conversational Surveys with Open-ended Questions

2019-05-25 · Ziang Xiao, Michelle X. Zhou, Q. Vera Liao, Gloria Mark 외

The rise of increasingly more powerful chatbots offers a new way to collect information through conversational surveys, where a chatbot asks open-ended questions, interprets a user's free-text responses, and probes answe…

ChatbotInformativenessSpecificitySurvey

PHYRE: A New Benchmark for Physical Reasoning

2019-08-15 · NeurIPS 2019 12 · Anton Bakhtin, Laurens van der Maaten, Justin Johnson, Laura Gustafson 외

Understanding and reasoning about physics is an important ability of intelligent agents. We develop the PHYRE benchmark for physical reasoning that contains a set of simple classical mechanics puzzles in a 2D physical en…

Visual Reasoning

EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers

2026-01-29 · John Flynn, Wolfgang Paier, Dimitar Dinev, Sam Nhut Nguyen 외 arxiv

Current generative video models excel at producing novel content from text and image prompts, but leave a critical gap in editing existing pre-recorded videos, where minor alterations to the spoken script require preserv…