paper-with-me

홈 › Papers

Improving Automatic VQA Evaluation Using Large Language Models

2023-10-04 · Oscar Mañas, Benno Krojer, Aishwarya Agrawal

8 years after the visual question answering (VQA) task was proposed, accuracy remains the primary metric for automatic evaluation. VQA Accuracy has been effective so far in the IID evaluation setting. However, our community is undergoing a shift towards open-ended generative models and OOD evaluation. In this new paradigm, the existing VQA Accuracy metric is overly stringent and underestimates the performance of VQA systems. Thus, there is a need to develop more robust automatic VQA metrics that serve as a proxy for human judgment. In this work, we propose to leverage the in-context learning capabilities of instruction-tuned large language models (LLMs) to build a better VQA metric. We formulate VQA evaluation as an answer-rating task where the LLM is instructed to score the accuracy of a candidate answer given a set of reference answers. We demonstrate the proposed metric better correlates with human judgment compared to existing metrics across several VQA models and benchmarks. We hope wide adoption of our metric will contribute to better estimating the research progress on the VQA task. We plan to release the evaluation code and collected human judgments.

📄 PDF Abstract BibTeX arXiv:2310.02567

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

An Automatic Evaluation of the WMT22 General Machine Translation Task

2022-09-28 · Benjamin Marie

This report presents an automatic evaluation of the general machine translation task of the Seventh Conference on Machine Translation (WMT22). It evaluates a total of 185 systems for 21 translation directions including h…

Machine TranslationTranslation

ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation

2021-12-17 · ACL ARR December 2022 12 · Anonymous

Automatic evaluations for natural language generation conventionally rely on token-level or embedding-level comparisons with the text references. This is different from human evaluation manners, in which people also form…

nlg evaluationText Generation

Exploring Automatic Evaluation Methods based on a Decoder-based LLM for Text Generation

2023-10-17 · Tomohito Kasahara, Daisuke Kawahara

Automatic evaluation of text generation is essential for improving the accuracy of generation tasks. In light of the current trend towards increasingly larger decoder-based language models, we investigate automatic evalu…

DecoderIn-Context LearningMachine TranslationSemantic Textual Similarity+1

Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation

2024-05-22 · Cyril Chhun, Fabian M. Suchanek, Chloé Clavel

Storytelling is an integral part of human experience and plays a crucial role in social interactions. Thus, Automatic Story Evaluation (ASE) and Generation (ASG) could benefit society in multiple ways, but they are chall…

An Empirical Study of Evaluating Long-form Question Answering

2025-04-25 · Ning Xian, Yixing Fan, Ruqing Zhang, Maarten de Rijke 외

\Ac{LFQA} aims to generate lengthy answers to complex questions. This scenario presents great flexibility as well as significant challenges for evaluation. Most evaluations rely on deterministic metrics that depend on st…

FormInformativenessLarge Language ModelLong Form Question Answering+1