paper-with-me

홈 › Papers

Human Preferences for Constructive Interactions in Language Model Alignment

2025-03-05 · Yara Kyrychenko, Jon Roozenbeek, Brandon Davidson, Sander van der Linden, Ramit Debnath

As large language models (LLMs) enter the mainstream, aligning them to foster constructive dialogue rather than exacerbate societal divisions is critical. Using an individualized and multicultural alignment dataset of over 7,500 conversations of individuals from 74 countries engaging with 21 LLMs, we examined how linguistic attributes linked to constructive interactions are reflected in human preference data used for training AI. We found that users consistently preferred well-reasoned and nuanced responses while rejecting those high in personal storytelling. However, users who believed that AI should reflect their values tended to place less preference on reasoning in LLM responses and more on curiosity. Encouragingly, we observed that users could set the tone for how constructive their conversation would be, as LLMs mirrored linguistic attributes, including toxicity, in user queries.

📄 PDF Abstract BibTeX arXiv:2503.16480

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction

2026-04-01 · Max Kanwal, Caryn Tran arxiv

Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with extensive empirical evidence showing that preferences are layered, dynamic, and constr…

Constructive Large Language Models Alignment with Diverse Feedback

2023-10-10 · Tianshu Yu, Ting-En Lin, Yuchuan Wu, Min Yang 외

In recent research on large language models (LLMs), there has been a growing emphasis on aligning these models with human values to reduce the impact of harmful content. However, current alignment methods often rely sole…

Learning TheoryModels AlignmentQuestion AnsweringText Summarization

CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions

2025-08-03 · Tae Soo Kim, Yoonjoo Lee, Yoonah Park, Jiho Kim 외 arxiv

Personalization of Large Language Models (LLMs) often assumes users hold static preferences that reflect globally in all tasks. In reality, humans hold dynamic preferences that change depending on the context. As users i…

Learning Transferable Latent User Preferences for Human-Aligned Decision Making

2026-05-12 · Alina Hyk, Sandhya Saisubramanian arxiv

Large language models (LLMs) are increasingly used as reasoning modules in many applications. While they are efficient in certain tasks, LLMs often struggle to produce human-aligned solutions. Human-aligned decision maki…

Decision Making

Language Model Alignment in Multilingual Trolley Problems

2024-07-02 · Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine 외

We evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems. Building on the Moral Machine experiment, which captures over 40 million human judgments across 200…

Decision MakingEthicsFairnessLanguage Modeling+2