Papers nlg evaluation
“nlg evaluation” 태그가 달린 논문 71편 · 필터 해제
Long-Form Information Alignment Evaluation Beyond Atomic Facts
Information alignment evaluators are vital for various NLG evaluation tasks and trustworthy LLM deployment, reducing hallucinations and enhancing user trust. Current fine-grained methods, like FactScore, verify facts ind…
Formnlg evaluationBeyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
Evaluating natural language generation (NLG) systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardisation, and demogr…
AllDiversitynlg evaluationPrompt Engineering+2DeepSeek vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
Reasoning-enabled large language models (LLMs) have recently demonstrated impressive performance in complex logical and mathematical tasks, yet their effectiveness in evaluating natural language generation remains unexpl…
Machine Translationnlg evaluationText GenerationText SummarizationOpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
Large Language Models (LLMs) have demonstrated great potential as evaluators of NLG systems, allowing for high-quality, reference-free, and multi-aspect assessments. However, existing LLM-based metrics suffer from two ma…
nlg evaluationExploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capabilities of LLMs across different languag…
nlg evaluationSAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to considerable enhancements in productivity a…
Language ModelingLanguage ModellingLarge Language ModelMultiple-choice+2Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking
Current automated fact-checking (AFC) approaches commonly evaluate evidence either implicitly via the predicted verdicts or by comparing retrieved evidence with a predefined closed knowledge source, such as Wikipedia. Ho…
Fact Checkingnlg evaluationRetrievalText GenerationAnalyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
The correlation between NLG automatic evaluation metrics and human evaluation is often regarded as a critical criterion for assessing the capability of an evaluation metric. However, different grouping methods and correl…
nlg evaluationLarge Language Models Are Active Critics in NLG Evaluation
The conventional paradigm of using large language models (LLMs) for evaluating natural language generation (NLG) systems typically relies on two key inputs: (1) a clear definition of the NLG task to be evaluated and (2) …
nlg evaluationPrompt EngineeringText GenerationDHP Benchmark: Are LLMs Good NLG Evaluators?
Large Language Models (LLMs) are increasingly serving as evaluators in Natural Language Generation (NLG) tasks. However, the capabilities of LLMs in scoring NLG quality remain inadequately explored. Current studies depen…
Benchmarkingnlg evaluationQuestion AnsweringStory Completion+1ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models
Assessing the quality of outputs generated by generative models, such as large language models and vision language models, presents notable challenges. Traditional methods for evaluation typically rely on either human as…
nlg evaluationText GenerationThemis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
The evaluation of natural language generation (NLG) tasks is a significant and longstanding research area. With the recent emergence of powerful large language models (LLMs), some studies have turned to LLM-based automat…
Language ModelingLanguage Modellingnlg evaluationText GenerationBetter than Random: Reliable NLG Human Evaluation with Constrained Active Sampling
Human evaluation is viewed as a reliable evaluation method for NLG which is expensive and time-consuming. To save labor and costs, researchers usually perform human evaluation on a small subset of data sampled from the w…
nlg evaluationDefining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
Human evaluation serves as the gold standard for assessing the quality of Natural Language Generation (NLG) systems. Nevertheless, the evaluation guideline, as a pivotal element ensuring reliable and reproducible human a…
nlg evaluationText GenerationVulnerability DetectionUnveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models
The automatic evaluation of natural language generation (NLG) systems presents a long-lasting challenge. Recent studies have highlighted various neural metrics that align well with human evaluations. Yet, the robustness …
nlg evaluationText GenerationDEBATE: Devil's Advocate-Based Assessment and Text Evaluation
As natural language generation (NLG) models have become prevalent, systematically assessing the quality of machine-generated texts has become increasingly important. Recent studies introduce LLM-based evaluators that ope…
nlg evaluationText GenerationWaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
Watermarking generative-AI systems, such as LLMs, has gained considerable interest, driven by their enhanced capabilities across a wide range of tasks. Although current approaches have demonstrated that small, context-de…
nlg evaluationAre LLM-based Evaluators Confusing NLG Quality Criteria?
Some prior work has shown that LLMs perform well in NLG evaluation for different tasks. However, we discover that LLMs seem to confuse different evaluation criteria, which reduces their reliability. For further verificat…
nlg evaluationOne Prompt To Rule Them All: LLMs for Opinion Summary Evaluation
Evaluation of opinion summaries using conventional reference-based metrics rarely provides a holistic evaluation and has been shown to have a relatively low correlation with human judgments. Recent studies suggest using …
Allnlg evaluationOpinion SummarizationSpecificityLLM-based NLG Evaluation: Current Status and Challenges
Evaluating natural language generation (NLG) is a vital but challenging problem in natural language processing. Traditional evaluation metrics mainly capturing content (e.g. n-gram) overlap between system outputs and ref…
nlg evaluationText Generation