paper-with-me

홈 › Papers

A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability

2025-02-17 · Xinyu Hu, Mingqi Gao, Li Lin, Zhenghan Yu, Xiaojun Wan

In NLG meta-evaluation, evaluation metrics are typically assessed based on their consistency with humans. However, we identify some limitations in traditional NLG meta-evaluation approaches, such as issues in handling human ratings and ambiguous selections of correlation measures, which undermine the effectiveness of meta-evaluation. In this work, we propose a dual-perspective NLG meta-evaluation framework that focuses on different evaluation capabilities, thereby providing better interpretability. In addition, we introduce a method of automatically constructing the corresponding benchmarks without requiring new human annotations. Furthermore, we conduct experiments with 16 representative LLMs as the evaluators based on our proposed framework, comprehensively analyzing their evaluation performance from different perspectives.

📄 PDF Abstract BibTeX arXiv:2502.12052

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective

2025-05-26 · Junnan Liu, Hongwei Liu, Linchen Xiao, Shudong Liu 외

We propose a novel framework for comprehending the reasoning capabilities of large language models (LLMs) through the perspective of meta-learning. By conceptualizing reasoning trajectories as pseudo-gradient descent upd…

Meta-Learning

A Dual-Perspective Metaphor Detection Framework Using Large Language Models

2024-12-23 · Yujie Lin, Jingyao Liu, Yan Gao, Ante Wang 외

Metaphor detection, a critical task in natural language processing, involves identifying whether a particular word in a sentence is used metaphorically. Traditional approaches often rely on supervised learning models tha…

Decision MakingKnowledge GraphsSentence

Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation

2024-10-22 · Mingqi Gao, Xinyu Hu, Li Lin, Xiaojun Wan

The correlation between NLG automatic evaluation metrics and human evaluation is often regarded as a critical criterion for assessing the capability of an evaluation metric. However, different grouping methods and correl…

nlg evaluation

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

2026-08-27 · Mingqi Gao, Anthony Sicilia, Weiyan Shi arxiv

Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered e…

Automatically Extracting Numerical Results from Randomized Controlled Trials with Large Language Models

2024-05-02 · Hye Sun Yun, David Pogrebitskiy, Iain J. Marshall, Byron C. Wallace

Meta-analyses statistically aggregate the findings of different randomized controlled trials (RCTs) to assess treatment effectiveness. Because this yields robust estimates of treatment effectiveness, results from meta-an…