paper-with-me

홈 › Papers

Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays

2025-03-25 · Jinsook Lee, AJ Alvero, Thorsten Joachims, René Kizilcec

People are increasingly using technologies equipped with large language models (LLM) to write texts for formal communication, which raises two important questions at the intersection of technology and society: Who do LLMs write like (model alignment); and can LLMs be prompted to change who they write like (model steerability). We investigate these questions in the high-stakes context of undergraduate admissions at a selective university by comparing lexical and sentence variation between essays written by 30,000 applicants to two types of LLM-generated essays: one prompted with only the essay question used by the human applicants; and another with additional demographic information about each applicant. We consistently find that both types of LLM-generated essays are linguistically distinct from human-authored essays, regardless of the specific model and analytical approach. Further, prompting a specific sociodemographic identity is remarkably ineffective in aligning the model with the linguistic patterns observed in human writing from this identity group. This holds along the key dimensions of sex, race, first-generation status, and geographic location. The demographically prompted and unprompted synthetic texts were also more similar to each other than to the human text, meaning that prompting did not alleviate homogenization. These issues of model alignment and steerability in current LLMs raise concerns about the use of LLMs in high-stakes contexts.

📄 PDF Abstract BibTeX arXiv:2503.20062

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs

2025-05-27 · Trenton Chang, Tobias Schnabel, Adith Swaminathan, Jenna Wiens

Despite advances in large language models (LLMs) on reasoning and instruction-following benchmarks, it remains unclear whether they can reliably produce outputs aligned with a broad variety of user goals, a concept we re…

Instruction FollowingPrompt Engineering

ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies

2026-03-18 · Zhenyang Chen, Alan Tian, Liquan Wang, Benjamin Joffe 외 arxiv

Despite strong multi-task pretraining, existing policies often exhibit poor task steerability. For example, a robot may fail to respond to a new instruction ``put the bowl in the sink" when moving towards the oven, execu…

What's Producible May Not Be Reachable: Measuring the Steerability of Generative Models

2025-03-21 · Keyon Vafa, Sarah Bentley, Jon Kleinberg, Sendhil Mullainathan

How should we evaluate the quality of generative models? Many existing metrics focus on a model's producibility, i.e. the quality and breadth of outputs it can generate. However, the actual value from using a generative …

Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability

2025-10-07 · Taylor Sorensen, Benjamin Newman, Jared Moore, Chan Park 외 arxiv

Language model post-training has enhanced instruction-following and performance on many downstream tasks, but also comes with an often-overlooked cost on tasks with many possible valid answers. On many tasks such as crea…

Synthetic Data Generation

AI Text-to-Behavior: A Study In Steerability

2023-08-07 · David Noever, Sam Hyams

The research explores the steerability of Large Language Models (LLMs), particularly OpenAI's ChatGPT iterations. By employing a behavioral psychology framework called OCEAN (Openness, Conscientiousness, Extroversion, Ag…