paper-with-me

홈 › Papers

Benchmarking Answer Verification Methods for Question Answering-Based Summarization Evaluation Metrics

2022-04-21 · Findings (ACL) 2022 5 · Daniel Deutsch, Dan Roth

Question answering-based summarization evaluation metrics must automatically determine whether the QA model's prediction is correct or not, a task known as answer verification. In this work, we benchmark the lexical answer verification methods which have been used by current QA-based metrics as well as two more sophisticated text comparison methods, BERTScore and LERC. We find that LERC out-performs the other methods in some settings while remaining statistically indistinguishable from lexical overlap in others. However, our experiments reveal that improved verification performance does not necessarily translate to overall QA-based metric quality: In some scenarios, using a worse verification method -- or using none at all -- has comparable performance to using the best verification method, a result that we attribute to properties of the datasets.

📄 PDF Abstract BibTeX arXiv:2204.10206

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeBenchmarkingQuestion Answering

Similar Papers 제목 키워드 기반

Benchmarking Answer Verification Methods for Question Answering-Based Summarization Evaluation Metrics

2021-09-17 · ACL ARR September 2021 9 · Anonymous

Question answering-based summarization evaluation metrics must automatically determine whether the QA model's prediction is correct or not, a task known as answer verification. In this work, we benchmark the lexical answ…

AttributeBenchmarkingQuestion Answering

Benchmarking Geospatial Question Answering Engines using the Dataset GeoQuestions1089

2023-11-06 · International Semantic Web Conference 2023 11 · Sergios-Anestis Kefalidis, Dharmen Punjani, Eleni Tsalapati, Konstantinos Plas 외

We present the dataset GeoQuestions1089 for benchmarking geospatial question answering engines. GeoQuestions1089 is the largest such dataset available presently and it contains 1089 questions, their corresponding GeoSPA…

BenchmarkingKnowledge Base Question AnsweringQuestion Answering

Uncertainty Estimation of Large Language Models in Medical Question Answering

2024-07-11 · Jiaxin Wu, Yizhou Yu, Hong-Yu Zhou

Large Language Models (LLMs) show promise for natural language generation in healthcare, but risk hallucinating factually incorrect information. Deploying LLMs for medical question answering necessitates reliable uncerta…

Medical Question AnsweringQuestion AnsweringText Generation

V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering

2026-01-26 · Mengyuan Jin, Zehui Liao, Yong Xia arxiv

Multimodal Large Language Models (MLLMs) have shown remarkable capability in assisting disease diagnosis in medical visual question answering (VQA). However, their outputs remain vulnerable to hallucinations (i.e., respo…

Visual Question AnsweringComputational Efficiency

Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning

2025-01-09 · CVPR 2025 1 · Huabin Liu, Filip Ilievski, Cees G. M. Snoek

This paper proposes the first video-grounded entailment tree reasoning method for commonsense video question answering (VQA). Despite the remarkable progress of large visual-language models (VLMs), there are growing conc…

BenchmarkingQuestion AnsweringVideo Question AnsweringVisual Question Answering (VQA)