paper-with-me

홈 › Papers

Redundancy Aware Multi-Reference Based Gainwise Evaluation of Extractive Summarization

2023-08-04 · Mousumi Akter, Santu Karmaker

The ROUGE metric is commonly used to evaluate extractive summarization task, but it has been criticized for its lack of semantic awareness and its ignorance about the ranking quality of the extractive summarizer. Previous research has introduced a gain-based automated metric called Sem-nCG that addresses these issues, as it is both rank and semantic aware. However, it does not consider the amount of redundancy present in a model summary and currently does not support evaluation with multiple reference summaries. It is essential to have a model summary that balances importance and diversity, but finding a metric that captures both of these aspects is challenging. In this paper, we propose a redundancy-aware Sem-nCG metric and demonstrate how the revised Sem-nCG metric can be used to evaluate model summaries against multiple references as well which was missing in previous research. Experimental results demonstrate that the revised Sem-nCG metric has a stronger correlation with human judgments compared to the previous Sem-nCG metric and traditional ROUGE and BERTScore metric for both single and multiple reference scenarios.

📄 PDF Abstract BibTeX arXiv:2308.02270

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityExtractive Summarization

Similar Papers 제목 키워드 기반

A Training-free and Reference-free Summarization Evaluation Metric via Centrality-weighted Relevance and Self-referenced Redundancy

2021-06-26 · ACL 2021 5 · Wang Chen, Piji Li, Irwin King

In recent years, reference-based and supervised summarization evaluation metrics have been widely explored. However, collecting human-annotated references and ratings are costly and time-consuming. To avoid these limitat…

Document SummarizationSentence

Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants

2025-08-22 · Chongyang Li, Zhiqiang Yuan, Hanbo Bi, Zexi Jia 외 arxiv

Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs) to develop effective walking assistance systems for blind and low visi…

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents

2025-11-11 · Chih-Hsuan, Yang, Tanwi Mallick, Le Chen 외 arxiv

Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tasks, yet current training methods lack principled ways to connect system-level evaluation with agent- and message-level learning.…

WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation

2025-10-22 · Yaoyao Qian, Yuanli Wang, Jinda Zhang, Yun Zong 외 arxiv

Current evaluation of web agents largely reduces to binary success metrics or conformity to a single reference trajectory, ignoring the structural diversity present in benchmark datasets. We present WebGraphEval, a frame…

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

2026-04-21 · Hanjun Cho, Jay-Yoon Lee arxiv

Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where i…