paper-with-me

홈 › Papers

An Empirical Study of Evaluating Long-form Question Answering

2025-04-25 · Ning Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke, Jiafeng Guo

\Ac{LFQA} aims to generate lengthy answers to complex questions. This scenario presents great flexibility as well as significant challenges for evaluation. Most evaluations rely on deterministic metrics that depend on string or n-gram matching, while the reliability of large language model-based evaluations for long-form answers remains relatively unexplored. We address this gap by conducting an in-depth study of long-form answer evaluation with the following research questions: (i) To what extent do existing automatic evaluation metrics serve as a substitute for human evaluations? (ii) What are the limitations of existing evaluation metrics compared to human evaluations? (iii) How can the effectiveness and robustness of existing evaluation methods be improved? We collect 5,236 factoid and non-factoid long-form answers generated by different large language models and conduct a human evaluation on 2,079 of them, focusing on correctness and informativeness. Subsequently, we investigated the performance of automatic evaluation metrics by evaluating these answers, analyzing the consistency between these metrics and human evaluations. We find that the style, length of the answers, and the category of questions can bias the automatic evaluation metrics. However, fine-grained evaluation helps mitigate this issue on some metrics. Our findings have important implications for the use of large language models for evaluating long-form question answering. All code and datasets are available at https://github.com/bugtig6351/lfqa_evaluation.

📄 PDF Abstract BibTeX arXiv:2504.18413

Code (1)

bugtig6351/lfqa_evaluation 공식 구현

Tasks

FormInformativenessLarge Language ModelLong Form Question AnsweringQuestion Answering

Similar Papers 제목 키워드 기반

EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos

2025-09-28 · Sourjyadip Ray, Shubham Sharma, Somak Aditya, Pawan Goyal arxiv

As digital platforms redefine educational paradigms, ensuring interactivity remains vital for effective learning. This paper explores using Multimodal Large Language Models (MLLMs) to automatically respond to student que…

Question Answering

Multi-hop Inference for Sentence-level TextGraphs: How Challenging is Meaningfully Combining Information for Science Question Answering?

2018-05-29 · WS 2018 6 · Peter Jansen

Question Answering for complex questions is often modeled as a graph construction or traversal task, where a solver must build or traverse a graph of facts that answer and explain a given question. This "multi-hop" infer…

graph constructionKnowledge GraphsQuestion AnsweringScience Question Answering+1

What Makes Sentences Semantically Related: A Textual Relatedness Dataset and Empirical Study

2021-10-10 · Mohamed Abdalla, Krishnapriya Vishnubhotla, Saif M. Mohammad

The degree of semantic relatedness of two units of language has long been considered fundamental to understanding meaning. Additionally, automatically determining relatedness has many applications such as question answer…

Question AnsweringSemantic SimilaritySemantic Textual SimilaritySentence

Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents

2026-05-25 · Yeonjun In, Wonjoong Kim, Sangwu Park, Kanghoon Yoon 외 arxiv

Existing large language model (LLM) based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users. This misalignment w…

A Benchmark for Long-Form Medical Question Answering

2024-11-14 · Pedram Hosseini, Jessica M. Sin, Bing Ren, Bryceton G. Thomas 외

There is a lack of benchmarks for evaluating large language models (LLMs) in long-form medical question answering (QA). Most existing medical QA evaluation benchmarks focus on automatic metrics and multiple-choice questi…

Answer GenerationFormMedical Question AnsweringMultiple-choice+1