paper-with-me

홈 › Papers

Summarization is Not Dead Yet

2026-06-06 · Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang arxiv

The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem. We re-examine this narrative through a multi-track evaluation covering diverse datasets and state-of-the-art LLMs, combining controlled human assessment, bias-mitigated LLM-as-Judge protocols, factuality verification against external knowledge, and corpus-level linguistic analysis. Our findings reveal a more nuanced landscape in which human references continue to demonstrate advantages in informativeness and faithfulness, whereas LLM outputs are preferred mainly for surface-level coherence and fluency. Factuality verification indicates that human references remain more reliable, particularly for claims involving reasoning or synthesis, and linguistic analysis uncovers a pattern of stylistic homogeneity across different models. These observations suggest that current LLMs have raised the floor of summarization quality, but the ceiling of their performance remains below human capabilities.

📄 PDF Abstract BibTeX arXiv:2606.08000

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding Code Semantics: An Evaluation of Transformer Models in Summarization

2023-10-25 · Debanjan Mondal, Abhilasha Lodha, Ankita Sahoo, Beena Kumari

This paper delves into the intricacies of code summarization using advanced transformer-based language models. Through empirical studies, we evaluate the efficacy of code summarization by altering function and variable n…

Code Summarization

Summarization is (Almost) Dead

2023-09-18 · Xiao Pu, Mingqi Gao, Xiaojun Wan

How well can large language models (LLMs) generate summaries? We develop new datasets and conduct human evaluation experiments to evaluate the zero-shot generation capability of LLMs across five distinct summarization ta…

Text Summarization

Rescue Conversations from Dead-ends: Efficient Exploration for Task-oriented Dialogue Policy Optimization

2023-05-05 · Yangyang Zhao, Zhenyu Wang, Mehdi Dastani, Shihan Wang

Training a dialogue policy using deep reinforcement learning requires a lot of exploration of the environment. The amount of wasted invalid exploration makes their learning inefficient. In this paper, we find and define …

Data AugmentationDeep Reinforcement LearningEfficient Exploration

Age of Computing: A Metric of Computation Freshness in Communication and Computation Cooperative Networks

2024-03-08 · Xingran Chen, Yi Zhuang, Kun Yang

In communication and computation cooperative networks (3CNs), timely computation is crucial but not always guaranteed. There is a strong demand for a computational task to be completed within a given deadline. The time t…

Scheduling for Urban Air Mobility using Safe Learning

2022-09-28 · Surya Murthy, Natasha A. Neogi, Suda Bharadwaj

This work considers the scheduling problem for Urban Air Mobility (UAM) vehicles travelling between origin-destination pairs with both hard and soft trip deadlines. Each route is described by a discrete probability distr…

Scheduling