paper-with-me

Papers nlg evaluation

“nlg evaluation” 태그가 달린 논문 71편 · 필터 해제

Long-Form Information Alignment Evaluation Beyond Atomic Facts

2025-05-21 · Danna Zheng, Mirella Lapata, Jeff Z. Pan

Information alignment evaluators are vital for various NLG evaluation tasks and trustworthy LLM deployment, reducing hallucinations and enhancing user trust. Current fine-grained methods, like FactScore, verify facts ind…

Formnlg evaluation

Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts

2025-04-29 · Hanhua Hong, Chenghao Xiao, Yang Wang, Yiqi Liu 외

Evaluating natural language generation (NLG) systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardisation, and demogr…

AllDiversitynlg evaluationPrompt Engineering+2

DeepSeek vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?

2025-04-10 · Daniil Larionov, Sotaro Takeshita, Ran Zhang, Yanran Chen 외

Reasoning-enabled large language models (LLMs) have recently demonstrated impressive performance in complex logical and mathematical tasks, yet their effectiveness in evaluating natural language generation remains unexpl…

Machine Translationnlg evaluationText GenerationText Summarization

OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs

2025-03-14 · Ivan Kartáč, Mateusz Lango, Ondřej Dušek

Large Language Models (LLMs) have demonstrated great potential as evaluators of NLG systems, allowing for high-quality, reference-free, and multi-aspect assessments. However, existing LLM-based metrics suffer from two ma…

nlg evaluation

Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators

2025-03-06 · Jiayi Chang, Mingqi Gao, Xinyu Hu, Xiaojun Wan

Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capabilities of LLMs across different languag…

nlg evaluation

SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text

2024-11-25 · Reshmi Ghosh, Tianyi Yao, Lizzy Chen, Sadid Hasan 외

Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to considerable enhancements in productivity a…

Language ModelingLanguage ModellingLarge Language ModelMultiple-choice+2

Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking

2024-11-08 · Mubashara Akhtar, Michael Schlichtkrull, Andreas Vlachos

Current automated fact-checking (AFC) approaches commonly evaluate evidence either implicitly via the predicted verdicts or by comparing retrieved evidence with a predefined closed knowledge source, such as Wikipedia. Ho…

Fact Checkingnlg evaluationRetrievalText Generation

Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation

2024-10-22 · Mingqi Gao, Xinyu Hu, Li Lin, Xiaojun Wan

The correlation between NLG automatic evaluation metrics and human evaluation is often regarded as a critical criterion for assessing the capability of an evaluation metric. However, different grouping methods and correl…

nlg evaluation

Large Language Models Are Active Critics in NLG Evaluation

2024-10-14 · Shuying Xu, Junjie Hu, Ming Jiang

The conventional paradigm of using large language models (LLMs) for evaluating natural language generation (NLG) systems typically relies on two key inputs: (1) a clear definition of the NLG task to be evaluated and (2) …

nlg evaluationPrompt EngineeringText Generation

DHP Benchmark: Are LLMs Good NLG Evaluators?

2024-08-25 · Yicheng Wang, Jiayi Yuan, Yu-Neng Chuang, Zhuoer Wang 외

Large Language Models (LLMs) are increasingly serving as evaluators in Natural Language Generation (NLG) tasks. However, the capabilities of LLMs in scoring NLG quality remain inadequately explored. Current studies depen…

Benchmarkingnlg evaluationQuestion AnsweringStory Completion+1

ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models

2024-07-16 · Yaswanth Narsupalli, Abhranil Chandra, Sreevatsa Muppirala, Manish Gupta 외

Assessing the quality of outputs generated by generative models, such as large language models and vision language models, presents notable challenges. Traditional methods for evaluation typically rely on either human as…

nlg evaluationText Generation

Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability

2024-06-26 · Xinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin 외

The evaluation of natural language generation (NLG) tasks is a significant and longstanding research area. With the recent emergence of powerful large language models (LLMs), some studies have turned to LLM-based automat…

Language ModelingLanguage Modellingnlg evaluationText Generation

Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling

2024-06-12 · Jie Ruan, Xiao Pu, Mingqi Gao, Xiaojun Wan 외

Human evaluation is viewed as a reliable evaluation method for NLG which is expensive and time-consuming. To save labor and costs, researchers usually perform human evaluation on a small subset of data sampled from the w…

nlg evaluation

Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation

2024-06-12 · Jie Ruan, Wenqing Wang, Xiaojun Wan

Human evaluation serves as the gold standard for assessing the quality of Natural Language Generation (NLG) systems. Nevertheless, the evaluation guideline, as a pivotal element ensuring reliable and reproducible human a…

nlg evaluationText GenerationVulnerability Detection

Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models

2024-05-23 · Yiming Chen, Chen Zhang, Danqing Luo, Luis Fernando D'Haro 외

The automatic evaluation of natural language generation (NLG) systems presents a long-lasting challenge. Recent studies have highlighted various neural metrics that align well with human evaluations. Yet, the robustness …

nlg evaluationText Generation

DEBATE: Devil's Advocate-Based Assessment and Text Evaluation

2024-05-16 · Alex Kim, Keonwoo Kim, Sangwon Yoon

As natural language generation (NLG) models have become prevalent, systematically assessing the quality of machine-generated texts has become increasingly important. Recent studies introduce LLM-based evaluators that ope…

nlg evaluationText Generation

WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models

2024-03-28 · Piotr Molenda, Adian Liusie, Mark J. F. Gales

Watermarking generative-AI systems, such as LLMs, has gained considerable interest, driven by their enhanced capabilities across a wide range of tasks. Although current approaches have demonstrated that small, context-de…

nlg evaluation

Are LLM-based Evaluators Confusing NLG Quality Criteria?

2024-02-19 · Xinyu Hu, Mingqi Gao, Sen Hu, Yang Zhang 외

Some prior work has shown that LLMs perform well in NLG evaluation for different tasks. However, we discover that LLMs seem to confuse different evaluation criteria, which reduces their reliability. For further verificat…

nlg evaluation

One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation

2024-02-18 · Tejpalsingh Siledar, Swaroop Nath, Sankara Sri Raghava Ravindra Muddu, Rupasai Rangaraju 외

Evaluation of opinion summaries using conventional reference-based metrics rarely provides a holistic evaluation and has been shown to have a relatively low correlation with human judgments. Recent studies suggest using …

Allnlg evaluationOpinion SummarizationSpecificity

LLM-based NLG Evaluation: Current Status and Challenges

2024-02-02 · Mingqi Gao, Xinyu Hu, Jie Ruan, Xiao Pu 외

Evaluating natural language generation (NLG) is a vital but challenging problem in natural language processing. Traditional evaluation metrics mainly capturing content (e.g. n-gram) overlap between system outputs and ref…

nlg evaluationText Generation
1–20 / 71 다음 →