A Large-Scale Multi-Length Headline Corpus for Analyzing Length-Constrained Headline Generation Model Evaluation
Browsing news articles on multiple devices is now possible. The lengths of news article headlines have precise upper bounds, dictated by the size of the display of the relevant device or interface. Therefore, controlling the length of headlines is essential when applying the task of headline generation to news production. However, because there is no corpus of headlines of multiple lengths for a given article, previous research on controlling output length in headline generation has not discussed whether the system outputs could be adequately evaluated without multiple references of different lengths. In this paper, we introduce two corpora, which are Japanese News Corpus (JNC) and JApanese MUlti-Length Headline Corpus (JAMUL), to confirm the validity of previous evaluation settings. The JNC provides common supervision data for headline generation. The JAMUL is a large-scale evaluation dataset for headlines of three different lengths composed by professional editors. We report new findings on these corpora; for example, although the longest length reference summary can appropriately evaluate the existing methods controlling output length, this evaluation setting has several problems.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesHeadline GenerationSimilar Papers 제목 키워드 기반
Diverse, Controllable, and Keyphrase-Aware: A Corpus and Method for News Multi-Headline Generation
News headline generation aims to produce a short sentence to attract readers to read the news. One news article often contains multiple keyphrases that are of interest to different users, which can naturally have multipl…
DecoderDiversityHeadline GenerationSentenceGenerating Representative Headlines for News Stories
Millions of news articles are published online every day, which can be overwhelming for readers to follow. Grouping articles that are reporting the same event into news stories is a common way of assisting readers in the…
ArticlesLLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
This research examines the potential of datasets generated by Large Language Models (LLMs) to support Natural Language Processing (NLP) tasks, aiming to overcome challenges related to data acquisition and privacy concern…
Semantic SimilaritySentiment AnalysisQuantifying Affective Bias in Low-Resource Media: Large-Scale Emotion Profiling of Bengali Headlines
News media can influence readers not only through the events they report but also through the emotional tone used to present them. This issue is especially important in digital news environments, where headlines often sh…
Automatic Extraction of News Values from Headline Text
Headlines play a crucial role in attracting audiences{'} attention to online artefacts (e.g. news articles, videos, blogs). The ability to carry out an automatic, large-scale analysis of headlines is critical to facilita…
ArticlesKeyword SpottingRecommendation Systems