paper-with-me

홈 › Papers

DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question Answering

2023-07-13 · Pei Ke, Fei Huang, Fei Mi, Yasheng Wang, Qun Liu, Xiaoyan Zhu, Minlie Huang

Existing evaluation metrics for natural language generation (NLG) tasks face the challenges on generalization ability and interpretability. Specifically, most of the well-performed metrics are required to train on evaluation datasets of specific NLG tasks and evaluation dimensions, which may cause over-fitting to task-specific datasets. Furthermore, existing metrics only provide an evaluation score for each dimension without revealing the evidence to interpret how this score is obtained. To deal with these challenges, we propose a simple yet effective metric called DecompEval. This metric formulates NLG evaluation as an instruction-style question answering task and utilizes instruction-tuned pre-trained language models (PLMs) without training on evaluation datasets, aiming to enhance the generalization ability. To make the evaluation process more interpretable, we decompose our devised instruction-style question about the quality of generated texts into the subquestions that measure the quality of each sentence. The subquestions with their answers generated by PLMs are then recomposed as evidence to obtain the evaluation result. Experimental results show that DecompEval achieves state-of-the-art performance in untrained metrics for evaluating text summarization and dialogue generation, which also exhibits strong dimension-level / task-level generalization ability and interpretability.

📄 PDF Abstract BibTeX arXiv:2307.06869

Code (1)

kepei1106/decompeval 공식 구현

Tasks

Dialogue Generationnlg evaluationQuestion AnsweringSentenceText GenerationText Summarization

Similar Papers 제목 키워드 기반

Decomposed evaluations of geographic disparities in text-to-image models

2024-06-17 · Abhishek Sureddy, Dishant Padalia, Nandhinee Periyakaruppa, Oindrila Saha 외

Recent work has identified substantial disparities in generated images of different geographic regions, including stereotypical depictions of everyday objects like houses and cars. However, existing measures for these di…

AttributeDiversityImage Generation

CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text Generation

2022-04-02 · ACL 2022 5 · Pei Ke, Hao Zhou, Yankai Lin, Peng Li 외

Existing reference-free metrics have obvious limitations for evaluating controlled text generation models. Unsupervised metrics can only provide a task-agnostic evaluation result which correlates weakly with human judgme…

Language ModelingLanguage ModellingText GenerationText Infilling

Improving Long-Text Alignment for Text-to-Image Diffusion Models

2024-10-15 · Luping Liu, Chao Du, Tianyu Pang, Zehan Wang 외

The rapid advancement of text-to-image (T2I) diffusion models has enabled them to generate unprecedented results from given texts. However, as text inputs become longer, existing encoding methods like CLIP face limitatio…

KETA: Kinematic-Phrases-Enhanced Text-to-Motion Generation via Fine-grained Alignment

2025-01-25 · Yu Jiang, Yixing Chen, Xingyang Li

Motion synthesis plays a vital role in various fields of artificial intelligence. Among the various conditions of motion generation, text can describe motion details elaborately and is easy to acquire, making text-to-mot…

Motion GenerationMotion Synthesis

Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation

2025-09-03 · Reina Ishikawa, Ryo Fujii, Hideo Saito, Ryo Hachiuma arxiv

Evaluating concept customization is challenging, as it requires a comprehensive assessment of fidelity to generative prompts and concept images. Moreover, evaluating multiple concepts is considerably more difficult than …