paper-with-me

홈 › Papers

LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering

2026-02-27 · Rafid Ishrak Jahan, Fahmid Shahriar Iqbal, Sagnik Ray Choudhury arxiv

Long-form question answering (LFQA) demands nuanced evaluation of multi-sentence explanatory responses, yet existing metrics often fail to reflect human judgment. We present LFQA-HP-1M, a large-scale dataset comprising 1.3M human pairwise preference annotations for LFQA. We propose nine rubrics for answer quality evaluation, and show that simple linear models based on these features perform comparably to state-of-the-art LLM evaluators. We further examine transitivity consistency, positional bias, and verbosity biases in LLM evaluators and demonstrate their vulnerability to adversarial perturbations. Overall, this work provides one of the largest public LFQA preference datasets and a rubric-driven framework for transparent and reliable evaluation.

📄 PDF Abstract BibTeX arXiv:2602.23603

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Putting People in LLMs' Shoes: Generating Better Answers via Question Rewriter

2024-08-20 · JunHao Chen, Bowen Wang, Zhouqiang Jiang, Yuta Nakashima

Large Language Models (LLMs) have demonstrated significant capabilities, particularly in the domain of question answering (QA). However, their effectiveness in QA is often undermined by the vagueness of user questions. T…

Long Form Question AnsweringQuestion Answering

A Critical Evaluation of Evaluations for Long-form Question Answering

2023-05-29 · Fangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol Choi

Long-form question answering (LFQA) enables answering a wide range of questions, but its flexibility poses enormous challenges for evaluation. We perform the first targeted study of the evaluation of long-form answers, c…

FormLong Form Question AnsweringQuestion AnsweringText Generation

Genie: Achieving Human Parity in Content-Grounded Datasets Generation

2024-01-25 · Asaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv 외

The lack of high-quality data for content-grounded generation tasks has been identified as a major obstacle to advancing these tasks. To address this gap, we propose Genie, a novel method for automatically generating hig…

Long Form Question AnsweringQuestion Answering

CALF: Benchmarking Evaluation of LFQA Using Chinese Examinations

2024-10-02 · Yuchen Fan, Xin Zhong, Heng Zhou, Yuchen Zhang 외

Long-Form Question Answering (LFQA) refers to generating in-depth, paragraph-level responses to open-ended questions. Although lots of LFQA methods are developed, evaluating LFQA effectively and efficiently remains chall…

BenchmarkingLong Form Question AnsweringQuestion Answering

Localizing and Mitigating Errors in Long-form Question Answering

2024-07-16 · Rachneet Sachdeva, Yixiao Song, Mohit Iyyer, Iryna Gurevych

Long-form question answering (LFQA) aims to provide thorough and in-depth answers to complex questions, enhancing comprehension. However, such detailed responses are prone to hallucinations and factual inconsistencies, c…

FormHallucinationLong Form Question AnsweringQuestion Answering