paper-with-me

홈 › Papers

A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task

2024-09-10 · Yuya Fujisaki, Shiro Takagi, Hideki Asoh, Wataru Kumagai

The progress in text summarization techniques has been remarkable. However the task of accurately extracting and summarizing necessary information from highly specialized documents such as research papers has not been sufficiently investigated. We are focusing on the task of extracting research questions (RQ) from research papers and construct a new dataset consisting of machine learning papers, RQ extracted from these papers by GPT-4, and human evaluations of the extracted RQ from multiple perspectives. Using this dataset, we systematically compared recently proposed LLM-based evaluation functions for summarizations, and found that none of the functions showed sufficiently high correlations with human evaluations. We expect our dataset provides a foundation for further research on developing better evaluation functions tailored to the RQ extraction task, and contribute to enhance the performance of the task. The dataset is available at https://github.com/auto-res/PaperRQ-HumanAnno-Dataset.

📄 PDF Abstract BibTeX arXiv:2409.06883

Code (0)

등록된 구현이 없습니다.

Tasks

Text Summarization

Methods 이 논문이 사용한 방법론

BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Attention 설명 없음
Position-Wise Feed-Forward Layer 설명 없음

Similar Papers 제목 키워드 기반

How to Evaluate Behavioral Models

2023-06-07 · Greg d'Eon, Sophie Greenwood, Kevin Leyton-Brown, James R. Wright

Researchers building behavioral models, such as behavioral game theorists, use experimental data to evaluate predictive models of human behavior. However, there is little agreement about which loss function should be use…

An Empirical Study of Evaluating Long-form Question Answering

2025-04-25 · Ning Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke 외

\Ac{LFQA} aims to generate lengthy answers to complex questions. This scenario presents great flexibility as well as significant challenges for evaluation. Most evaluations rely on deterministic metrics that depend on st…

FormInformativenessLarge Language ModelLong Form Question Answering+1

A Critical Look at Meta-evaluating Summarisation Evaluation Metrics

2024-09-29 · Xiang Dai, Sarvnaz Karimi, Biaoyan Fang

Effective summarisation evaluation metrics enable researchers and practitioners to compare different summarisation systems efficiently. Estimating the effectiveness of an automatic evaluation metric, termed meta-evaluati…

MIMICS-Duo: Offline & Online Evaluation of Search Clarification

2022-06-09 · Leila Tavakoli, Johanne R. Trippas, Hamed Zamani, Falk Scholer 외

Asking clarification questions is an active area of research; however, resources for training and evaluating search clarification methods are not sufficient. To address this issue, we describe MIMICS-Duo, a new freely av…

Evaluating Multiple System Summary Lengths: A Case Study

2018-10-01 · EMNLP 2018 10 · Ori Shapira, David Gabay, Hadar Ronen, Judit Bar-Ilan 외

Practical summarization systems are expected to produce summaries of varying lengths, per user needs. While a couple of early summarization benchmarks tested systems across multiple summary lengths, this practice was mos…