Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
Open-ended text generation has become a prominent task in natural language processing due to the rise of powerful (large) language models. However, evaluating the quality of these models and the employed decoding strategies remains challenging because of trade-offs among widely used metrics such as coherence, diversity, and perplexity. Decoding methods often excel in some metrics while underperforming in others, complicating the establishment of a clear ranking. In this paper, we present novel ranking strategies within this multicriteria framework. Specifically, we employ benchmarking approaches based on partial orderings and present a new summary metric designed to balance existing automatic indicators, providing a more holistic evaluation of text generation quality. Furthermore, we discuss the alignment of these approaches with human judgments. Our experiments demonstrate that the proposed methods offer a robust way to compare decoding strategies, exhibit similarities with human preferences, and serve as valuable tools in guiding model selection for open-ended text generation tasks. Finally, we suggest future directions for improving evaluation methodologies in text generation. Our codebase, datasets, and models are publicly available.
Code (1)
Tasks
BenchmarkingDiversityModel SelectionText GenerationSimilar Papers 제목 키워드 기반
An Empirical Study On Contrastive Search And Contrastive Decoding For Open-ended Text Generation
In the study, we empirically compare the two recently proposed decoding methods, i.e. Contrastive Search (CS) and Contrastive Decoding (CD), for open-ended text generation. The automatic evaluation results suggest that, …
DiversityText GenerationInterpretable neural networks based on continuous-valued logic and multicriteria decision operators
Combining neural networks with continuous logic and multicriteria decision making tools can reduce the black box nature of neural models. In this study, we show that nilpotent logical systems offer an appropriate mathema…
Decision MakingA Hypervolume Based Approach to Rank Intuitionistic Fuzzy Sets and Its Extension to Multi-criteria Decision Making Under Uncertainty
Ranking intuitionistic fuzzy sets with distance based ranking methods requires to calculate the distance between intuitionistic fuzzy set and a reference point which is known to have either maximum (positive ideal soluti…
Decision MakingDecision Making Under UncertaintyvalidThe Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
Recent text generation research has increasingly focused on open-ended domains such as story and poetry generation. Because models built for such tasks are difficult to evaluate automatically, most researchers in the spa…
Text GenerationComparative Statics in Multicriteria Search Models
McCall (1970) examines the search behaviour of an infinitely-lived and risk-neutral job seeker maximizing her lifetime earnings by accepting or rejecting real-valued scalar wage offers. In practice, job offers have multi…