paper-with-me

Papers

A Revealed Preference Framework for AI Alignment

2026-03-29 · Elchin Suleymanov arxiv

Human decision makers increasingly delegate choices to AI agents, raising a natural question: does the AI implement the human principal's preferences or pursue its own? To study this question using revealed preference techniques, I introduce the Luce Alignment Model, where the AI's choices are a mixture of two Luce rules, one reflecting the human's preferences and the other the AI's. I show that the AI's alignment (similarity of human and AI preferences) can be generically identified in two settings: the laboratory setting, where both human and AI choices are observed, and the field setting, where only AI choices are observed.

📄 PDF Abstract BibTeX arXiv:2603.27868

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can Revealed Preferences Clarify LLM Alignment and Steering?

2026-05-08 · Khurram Yamin, Jingjing Tang, Eric Horvitz, Bryan Wilder arxiv

LLMs are increasingly used to make or support high-stakes decisions under uncertainty, where alignment depends not only on factual accuracy but on how models weigh tradeoffs between different outcomes. We present an empi…

Medical Diagnosis

The Moral Mind(s) of Large Language Models

2024-11-19 · Avner Seror

As large language models (LLMs) increasingly participate in tasks with ethical and societal stakes, a critical question arises: do they exhibit an emergent "moral mind" - a consistent structure of moral preferences guidi…

BenchmarkingDecision Making

ERA-IT: Aligning Semantic Models with Revealed Economic Preference for Real-Time and Explainable Patent Valuation

2025-12-14 · Yongmin Yoo, Seungwoo Kim, Jingjiang Liu arxiv

Valuing intangible assets under uncertainty remains a critical challenge in the strategic management of technological innovation due to the information asymmetry inherent in high-dimensional technical specifications. Tra…

Random Attention Span

2024-05-19 · Dazhuo Wei

In this paper, I introduce a random attention span model (RAS) which uses stopping time to identify decision-makers' behavior under limited attention. Unlike many limited attention models, the RAS identifies preferences …

An algebraic approach to revealed preferences

2021-05-31 · Mikhail Freer, Cesar Martinelli

We propose and develop an algebraic approach to revealed preference. Our approach dispenses with non algebraic structure, such as topological assumptions. We provide algebraic axioms of revealed preference that subsume p…