paper-with-me

홈 › Papers

AI Text-to-Behavior: A Study In Steerability

2023-08-07 · David Noever, Sam Hyams

The research explores the steerability of Large Language Models (LLMs), particularly OpenAI's ChatGPT iterations. By employing a behavioral psychology framework called OCEAN (Openness, Conscientiousness, Extroversion, Agreeableness, Neuroticism), we quantitatively gauged the model's responsiveness to tailored prompts. When asked to generate text mimicking an extroverted personality, OCEAN scored the language alignment to that behavioral trait. In our analysis, while "openness" presented linguistic ambiguity, "conscientiousness" and "neuroticism" were distinctly evoked in the OCEAN framework, with "extroversion" and "agreeableness" showcasing a notable overlap yet distinct separation from other traits. Our findings underscore GPT's versatility and ability to discern and adapt to nuanced instructions. Furthermore, historical figure simulations highlighted the LLM's capacity to internalize and project instructible personas, precisely replicating their philosophies and dialogic styles. However, the rapid advancements in LLM capabilities and the opaque nature of some training techniques make metric proposals degrade rapidly. Our research emphasizes a quantitative role to describe steerability in LLMs, presenting both its promise and areas for further refinement in aligning its progress to human intentions.

📄 PDF Abstract BibTeX arXiv:2308.07326

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies

2026-03-18 · Zhenyang Chen, Alan Tian, Liquan Wang, Benjamin Joffe 외 arxiv

Despite strong multi-task pretraining, existing policies often exhibit poor task steerability. For example, a robot may fail to respond to a new instruction ``put the bowl in the sink" when moving towards the oven, execu…

Steerability of Instrumental-Convergence Tendencies in LLMs

2026-01-04 · Jakub Hoscilowicz arxiv

We examine two properties of AI systems: capability (what a system can do) and steerability (how reliably one can shift behavior toward intended outcomes). A central question is whether capability growth reduces steerabi…

Evaluating the Prompt Steerability of Large Language Models

2024-11-19 · Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth M. Daly 외

Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to evaluate the degree to which a given model…

A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs

2025-05-27 · Trenton Chang, Tobias Schnabel, Adith Swaminathan, Jenna Wiens

Despite advances in large language models (LLMs) on reasoning and instruction-following benchmarks, it remains unclear whether they can reliably produce outputs aligned with a broad variety of user goals, a concept we re…

Instruction FollowingPrompt Engineering

ArtWhisperer: A Dataset for Characterizing Human-AI Interactions in Artistic Creations

2023-06-13 · Kailas Vodrahalli, James Zou

As generative AI becomes more prevalent, it is important to study how human users interact with such models. In this work, we investigate how people use text-to-image models to generate desired target images. To study th…