paper-with-me

홈 › Papers

A Closer Look into Automatic Evaluation Using Large Language Models

2023-10-09 · Cheng-Han Chiang, Hung-Yi Lee

Using large language models (LLMs) to evaluate text quality has recently gained popularity. Some prior works explore the idea of using LLMs for evaluation, while they differ in some details of the evaluation process. In this paper, we analyze LLM evaluation (Chiang and Lee, 2023) and G-Eval (Liu et al., 2023), and we discuss how those details in the evaluation process change how well the ratings given by LLMs correlate with human ratings. We find that the auto Chain-of-Thought (CoT) used in G-Eval does not always make G-Eval more aligned with human ratings. We also show that forcing the LLM to output only a numeric rating, as in G-Eval, is suboptimal. Last, we reveal that asking the LLM to explain its own ratings consistently improves the correlation between the ChatGPT and human ratings and pushes state-of-the-art (SoTA) correlations on two meta-evaluation datasets.

📄 PDF Abstract BibTeX arXiv:2310.05657

Code (1)

d223302/a-closer-look-to-llm-evaluation 공식 구현

Similar Papers 제목 키워드 기반

A Closer Look at Recent Results of Verb Selection for Data-to-Text NLG

2019-10-01 · WS 2019 10 · Guanyi Chen, Jin-Ge Yao

Automatic natural language generation systems need to use the contextually-appropriate verbs when describing different kinds of facts or events, which has triggered research interest on verb selection for data-to-text ge…

Data-to-Text GenerationText Generation

Is This a Bad Table? A Closer Look at the Evaluation of Table Generation from Text

2024-06-21 · Pritika Ramu, Aparna Garimella, Sambaran Bandyopadhyay

Understanding whether a generated table is of good quality is important to be able to use it in creating or editing documents using automatic methods. In this work, we underline that existing measures for table quality e…

A Closer Look into the Robustness of Neural Dependency Parsers Using Better Adversarial Examples

2021-08-01 · Findings (ACL) 2021 8 · Yuxuan Wang, Wanxiang Che, Ivan Titov, Shay B. Cohen 외

One Step Closer to Automatic Evaluation of Text Simplification Systems

2014-04-01 · WS 2014 4 · Sanja {\v{S}}tajner, Ruslan Mitkov, Horacio Saggion
Decision MakingInformation RetrievalMachine TranslationText Simplification

A Closer Look at Machine Unlearning for Large Language Models

2024-10-10 · Xiaojian Yuan, Tianyu Pang, Chao Du, Kejiang Chen 외

Large language models (LLMs) may memorize sensitive or copyrighted content, raising privacy and legal concerns. Due to the high cost of retraining from scratch, researchers attempt to employ machine unlearning to remove …

DiversityMachine UnlearningSentence