Sentence-Level Fluency Evaluation: References Help, But Can Be Spared!
Motivated by recent findings on the probabilistic modeling of acceptability judgments, we propose syntactic log-odds ratio (SLOR), a normalized language model score, as a metric for referenceless fluency evaluation of natural language generation output at the sentence level. We further introduce WPSLOR, a novel WordPiece-based version, which harnesses a more compact language model. Even though word-overlap metrics like ROUGE are computed with the help of hand-written references, our referenceless methods obtain a significantly higher correlation with human fluency scores on a benchmark dataset of compressed sentences. Finally, we present ROUGE-LM, a reference-based metric which is a natural extension of WPSLOR to the case of available references. We show that ROUGE-LM yields a significantly higher correlation with human judgments than all baseline metrics, including WPSLOR on its own.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingSentenceText GenerationSimilar Papers 제목 키워드 기반
Simplicity Level Estimate (SLE): A Learned Reference-Less Metric for Sentence Simplification
Automatic evaluation for sentence simplification remains a challenging problem. Most popular evaluation metrics require multiple high-quality references -- something not readily available for simplification -- which make…
SentenceAn Automatic Tool For Language Evaluation
The aim of evaluating children speech and language is to measure their communication skills. In particular, the speech language pathologist is interested in determining the child{'}s impairments in the areas of language,…
SentenceImprove the Evaluation of Fluency Using Entropy for Machine Translation Evaluation Metrics
The widely-used automatic evaluation metrics cannot adequately reflect the fluency of the translations. The n-gram-based metrics, like BLEU, limit the maximum length of matched fragments to n and cannot catch the matched…
Machine TranslationSentenceTranslationOn-the-Fly Attention Modulation for Neural Generation
Despite considerable advancements with deep neural language models (LMs), neural text generation still suffers from degeneration: the generated text is repetitive, generic, self-contradictory, and often lacks commonsense…
Language ModellingSentenceText GenerationDoc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided memory for Document-level Machine Translation
The field of artificial intelligence has witnessed significant advancements in natural language processing, largely attributed to the capabilities of Large Language Models (LLMs). These models form the backbone of Agents…
Document Level Machine TranslationMachine TranslationSentenceTranslation