paper-with-me

홈 › Papers

A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators

2023-12-24 · Chen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang, Haizhou Li

Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human evaluations. Notably among them, large language models (LLMs), particularly the instruction-tuned variants like ChatGPT, are shown to be promising substitutes for human judges. Yet, existing works on utilizing LLMs for automatic dialogue evaluation are limited in their scope in terms of the number of meta-evaluation datasets, mode of evaluation, coverage of LLMs, etc. Hence, it remains inconclusive how effective these LLMs are. To this end, we conduct a comprehensive study on the application of LLMs for automatic dialogue evaluation. Specifically, we analyze the multi-dimensional evaluation capability of 30 recently emerged LLMs at both turn and dialogue levels, using a comprehensive set of 12 meta-evaluation datasets. Additionally, we probe the robustness of the LLMs in handling various adversarial perturbations at both turn and dialogue levels. Finally, we explore how model-level and dimension-level ensembles impact the evaluation performance. All resources are available at https://github.com/e0397123/comp-analysis.

📄 PDF Abstract BibTeX arXiv:2312.15407

Code (1)

e0397123/comp-analysis 공식 구현

Tasks

Dialogue Evaluation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

DualCoTs: Dual Chain-of-Thoughts Prompting for Sentiment Lexicon Expansion of Idioms

2024-09-26 · Fuqiang Niu, Minghuan Tan, BoWen Zhang, Min Yang 외

Idioms represent a ubiquitous vehicle for conveying sentiments in the realm of everyday discourse, rendering the nuanced analysis of idiom sentiment crucial for a comprehensive understanding of emotional expression withi…

Sentiment Analysis

Performance Evaluation of Large Language Models in Statistical Programming

2025-02-18 · Xinyi Song, Kexin Xie, Lina Lee, Ruizhe Chen 외

The programming capabilities of large language models (LLMs) have revolutionized automatic code generation and opened new avenues for automatic statistical analysis. However, the validity and quality of these generated c…

Code Generation

Multi-Programming Language Sandbox for LLMs

2024-10-30 · Shihan Dou, Jiazheng Zhang, Jianxiang Zang, Yunbo Tao 외

We introduce MPLSandbox, an out-of-the-box multi-programming language sandbox designed to provide unified and comprehensive feedback from compiler and analysis tools for Large Language Models (LLMs). It can automatically…

A thorough benchmark of automatic text classification: From traditional approaches to large language models

2025-04-02 · Washington Cunha, Leonardo Rocha, Marcos André Gonçalves

Automatic text classification (ATC) has experienced remarkable advancements in the past decade, best exemplified by recent small and large language models (SLMs and LLMs), leveraged by Transformer architectures. Despite …

Sentiment Analysistext-classificationText ClassificationTopic Classification

Tell Me Why: Explainable Public Health Fact-Checking with Large Language Models

2024-05-15 · Majid Zarharan, Pascal Wullschleger, Babak Behkam Kia, Mohammad Taher Pilehvar 외

This paper presents a comprehensive analysis of explainable fact-checking through a series of experiments, focusing on the ability of large language models to verify public health claims and provide explanations or justi…

Explanation GenerationFact Checkingparameter-efficient fine-tuning