paper-with-me

홈 › Papers

Squibs: Evaluating Human Pairwise Preference Judgments

2015-06-01 · CL 2015 6 · Mark Dras
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging

2026-05-19 · Zhenyu Li, Aleksandar Cvejic, Zehui Chen, Peter Wonka arxiv

Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond functional correctness. While rubric-based LLM judges improve interpre…

DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4

2023-05-24 · Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 외

Human preference judgments are pivotal in guiding large language models (LLMs) to produce outputs that align with human values. Human evaluations are also used in summarization tasks to compare outputs from various syste…

Informativeness

AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style

2026-03-12 · Joonyong Park, Jerry Li arxiv

Evaluating 'anime-like' voices currently relies on costly subjective judgments, yet no standardized objective metric exists. A key challenge is that anime-likeness, unlike naturalness, lacks a shared absolute scale, maki…

Benchmarking Music Generation Models and Metrics via Human Preference Studies

2025-06-23 · Audio Imagination: NeurIPS 2024 Workshop 2024 10 · Florian Grötschla, Ahmet Solak, Luca A. Lanzendörfer, Roger Wattenhofer

Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these…

BenchmarkingMusic Generation

Gaze patterns predict preference and confidence in pairwise AI image evaluation

2026-03-25 · Nikolas Papadopoulos, Shreenithi Navaneethan, Sheng Bai, Ankur Samanta 외 arxiv

Preference learning methods, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on pairwise human judgments, yet little is known about the cognitive processes underly…

Reinforcement Learning