paper-with-me

Papers

DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4

2023-05-24 · Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Fei Liu

Human preference judgments are pivotal in guiding large language models (LLMs) to produce outputs that align with human values. Human evaluations are also used in summarization tasks to compare outputs from various systems, complementing existing automatic metrics. Despite their significance, however, there has been limited research probing these pairwise or $k$-wise comparisons. The collective impact and relative importance of factors such as output length, informativeness, fluency, and factual consistency are still not well understood. It is also unclear if there are other hidden factors influencing human judgments. In this paper, we conduct an in-depth examination of a collection of pairwise human judgments released by OpenAI. Utilizing the Bradley-Terry-Luce (BTL) model, we reveal the inherent preferences embedded in these human judgments. We find that the most favored factors vary across tasks and genres, whereas the least favored factors tend to be consistent, e.g., outputs are too brief, contain excessive off-focus content or hallucinated facts. Our findings have implications on the construction of balanced datasets in human preference evaluations, which is a crucial step in shaping the behaviors of future LLMs.

📄 PDF Abstract BibTeX arXiv:2305.14702

Code (0)

등록된 구현이 없습니다.

Tasks

Informativeness

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Who Laughs with Whom? Disentangling Influential Factors in Humor Preferences across User Clusters and LLMs

2026-01-06 · Soichiro Murakami, Hidetaka Kamigaito, Hiroya Takamura, Manabu Okumura arxiv

Humor preferences vary widely across individuals and cultures, complicating the evaluation of humor using large language models (LLMs). In this study, we model heterogeneity in humor preferences in Oogiri, a Japanese cre…

Identifying Influential Actions in Human-Robot Interactions

2026-03-09 · Haoyang Jiang, Chenfei Xu, Yuya Okadome, Yukata Nakamura arxiv

Human-robot interaction combines robotics, cognitive science, and human factors to study collaborative systems. This paper introduces a method for identifying influential robot actions using transfer entropy, a statistic…

Adversarial multi-task underwater acoustic target recognition: towards robustness against various influential factors

2024-11-05 · Yuan Xie, Ji Xu, Jiawei Ren, Junfeng Li

Underwater acoustic target recognition based on passive sonar faces numerous challenges in practical maritime applications. One of the main challenges lies in the susceptibility of signal characteristics to diverse envir…

Saber Pro success prediction model using decision tree based learning

2020-06-02 · Gregorio Perez Bernal, Luisa Toro Villegas, Mauricio Toro

The primary objective of this report is to determine what influences the success rates of students who have studied in Colombia, analyzing the Saber 11, the test done at the last school year, some socioeconomic aspects a…

AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts

2025-02-07 · Soichiro Murakami, Peinan Zhang, Hidetaka Kamigaito, Hiroya Takamura 외

Effective linguistic choices that attract potential customers play crucial roles in advertising success. This study aims to explore the linguistic features of ad texts that influence human preferences. Although the creat…

Text Generation