paper-with-me

홈 › Papers

Pre-trained language models evaluating themselves - A comparative study

2022-05-01 · insights (ACL) 2022 5 · Philipp Koch, Matthias Aßenmacher, Christian Heumann

Evaluating generated text received new attention with the introduction of model-based metrics in recent years. These new metrics have a higher correlation with human judgments and seemingly overcome many issues of previous n-gram based metrics from the symbolic age. In this work, we examine the recently introduced metrics BERTScore, BLEURT, NUBIA, MoverScore, and Mark-Evaluate (Petersen). We investigate their sensitivity to different types of semantic deterioration (part of speech drop and negation), word order perturbations, word drop, and the common problem of repetition. No metric showed appropriate behaviour for negation, and further none of them was overall sensitive to the other issues mentioned above.

📄 PDF Abstract BibTeX

Code (1)

lazerlambda/metricscomparison 공식 구현

Tasks

NegationSensitivity

Similar Papers 제목 키워드 기반

Pre-trained language models evaluating themselves - A comparative study

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Evaluating generated text received new attention with the introduction of model-based metrics in recent years. These new metrics have a higher correlation with human judgments and seemingly overcome many issues of previo…

NegationSensitivity

Adversarial Multi-Agent Evaluation of Large Language Models through Iterative Debates

2024-10-07 · Chaithanya Bandi, Abir Harrasse

This paper explores optimal architectures for evaluating the outputs of large language models (LLMs) using LLMs themselves. We propose a novel framework that interprets LLMs as advocates within an ensemble of interacting…

A Comparative Study of Quality Evaluation Methods for Text Summarization

2024-06-30 · Huyen Nguyen, Haihua Chen, Lavanya Pobbathi, Junhua Ding

Evaluating text summarization has been a challenging task in natural language processing (NLP). Automatic metrics which heavily rely on reference summaries are not suitable in many situations, while human evaluation is t…

Text Summarization

Comparative Analysis of Drug-GPT and ChatGPT LLMs for Healthcare Insights: Evaluating Accuracy and Relevance in Patient and HCP Contexts

2023-07-24 · Giorgos Lysandrou, Roma English Owen, Kirsty Mursec, Grant Le Brun 외

This study presents a comparative analysis of three Generative Pre-trained Transformer (GPT) solutions in a question and answer (Q&A) setting: Drug-GPT 3, Drug-GPT 4, and ChatGPT, in the context of healthcare application…

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images

2025-08-29 · Juneyoung Ro, Namwoo Kim, Yoonjin Yoon arxiv

Effectively understanding urban scenes requires fine-grained spatial reasoning about objects, layouts, and depth cues. However, how well current vision-language models (VLMs), pretrained on general scenes, transfer these…

Spatial ReasoningObject Detection