paper-with-me

홈 › Papers

Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs

2026-01-19 · Zihan Dong, Xiaotian Hou, Ruijia Wu, Linjun Zhang arxiv

The increasing reliance on human preference feedback to judge AI-generated pseudo labels has created a pressing need for principled, budget-conscious data acquisition strategies. We address the crucial question of how to optimally allocate a fixed annotation budget between ground-truth labels and pairwise preferences in AI. Our solution, grounded in semi-parametric inference, casts the budget allocation problem as a monotone missing data framework. Building on this formulation, we introduce Preference-Calibrated Active Learning (PCAL), a novel method that learns the optimal data acquisition strategy and develops a statistically efficient estimator for functionals of the data distribution. Theoretically, we prove the asymptotic optimality of our PCAL estimator and establish a key robustness guarantee that ensures robust performance even with poorly estimated nuisance models. Our flexible framework applies to a general class of problems, by directly optimizing the estimator's variance instead of requiring a closed-form solution. This work provides a principled and statistically efficient approach for budget-constrained learning in modern AI. Simulations and real-data analysis demonstrate the practical benefits and superior performance of our proposed method.

📄 PDF Abstract BibTeX arXiv:2601.13458

Code (0)

등록된 구현이 없습니다.

Tasks

Active Learning

Similar Papers 제목 키워드 기반

Generative Reward Models

2024-10-02 · Dakota Mahan, Duy Van Phung, Rafael Rafailov, Chase Blagden 외

Reinforcement Learning from Human Feedback (RLHF) has greatly improved the performance of modern Large Language Models (LLMs). The RLHF process is resource-intensive and technically challenging, generally requiring a lar…

reinforcement-learningReinforcement Learning

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge

2026-04-07 · Xin Sun, Di Wu, Sijing Qin, Isao Echizen 외 arxiv

Large language models (LLMs) are increasingly used as automated evaluators (LLM-as-a-Judge). This work challenges its reliability by showing that trust judgments by LLMs are biased by disclosed source labels. Using a cou…

Modeling Art Evaluations from Comparative Judgments: A Deep Learning Approach to Predicting Aesthetic Preferences

2026-01-30 · Manoj Reddy Bethi, Sai Rupa Jhade, Pravallika Yaganti, Monoshiz Mahbub Khan 외 arxiv

Modeling human aesthetic judgments in visual art presents significant challenges due to individual preference variability and the high cost of obtaining labeled data. To reduce cost of acquiring such labels, we propose t…

MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces

2025-10-09 · Reuben A. Luera, Ryan Rossi, Franck Dernoncourt, Samyadeep Basu 외 arxiv

In an ideal design pipeline, user interface (UI) design is intertwined with user research to validate decisions, yet studies are often resource-constrained during early exploration. Recent advances in multimodal large la…

VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression

2025-12-17 · Kyle Sargent, Ruiqi Gao, Philipp Henzler, Charles Herrmann 외 arxiv

Evaluations of image compression performance which include human preferences have generally found that naive distortion functions such as MSE are insufficiently aligned to human perception. In order to align compression …

Image CompressionVisual Reasoning