paper-with-me

홈 › Papers

GRUEN for Evaluating Linguistic Quality of Generated Text

2020-10-06 · Findings of the Association for Computational Linguistics 2020 · Wanzheng Zhu, Suma Bhat

Automatic evaluation metrics are indispensable for evaluating generated text. To date, these metrics have focused almost exclusively on the content selection aspect of the system output, ignoring the linguistic quality aspect altogether. We bridge this gap by proposing GRUEN for evaluating Grammaticality, non-Redundancy, focUs, structure and coherENce of generated text. GRUEN utilizes a BERT-based model and a class of syntactic, semantic, and contextual features to examine the system output. Unlike most existing evaluation metrics which require human references as an input, GRUEN is reference-less and requires only the system output. Besides, it has the advantage of being unsupervised, deterministic, and adaptable to various tasks. Experiments on seven datasets over four language generation tasks show that the proposed metric correlates highly with human judgments.

📄 PDF Abstract BibTeX arXiv:2010.02498

Code (2)

WanzhengZhu/GRUEN 공식 구현 pytorch
jmpu/deepfaketextdetection pytorch

Tasks

Text Generation

Similar Papers 제목 키워드 기반

HotComment: A Benchmark for Evaluating Popularity of Online Comments

2026-04-28 · Yafeng Wu, Yunyao Zhang, Liliang Ye, Guiyi Zeng 외 arxiv

Online comments play a crucial role in shaping public sentiment and opinion dynamics on social media. However, evaluating their popularity remains challenging, not only because it depends on linguistic quality, originali…

Semantic Similarity

Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models

2025-08-11 · William Guo, Adaku Uchendu, Ana Smith arxiv

To mitigate the potential harms of Large Language Models (LLMs)generated text, researchers have proposed watermarking, a process of embedding detectable signals within text. With watermarking, we can always accurately de…

Integrating Randomness in Large Language Models: A Linear Congruential Generator Approach for Generating Clinically Relevant Content

2024-07-04 · Andrew Bouras

Generating diverse, high-quality outputs from language models is crucial for applications in education and content creation. Achieving true randomness and avoiding repetition remains a significant challenge. This study u…

Fact SelectionLanguage ModelingLanguage Modelling

E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness

2025-09-02 · Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah arxiv

A prevalent shortfall among current empathic AI systems is their inability to recognize when verbal expressions may not fully reflect underlying emotional states. This is because the existing datasets, used for the train…

Emotion Recognition

Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment

2025-03-25 · Hanlin Wu, Xufeng Duan, Zhenguang Cai

Voice-based AI development faces unique challenges in processing both linguistic and paralinguistic information. This study compares how large audio-language models (LALMs) and humans integrate speaker characteristics du…

EEGSensitivity