paper-with-me

홈 › Papers

IRR: Image Review Ranking Framework for Evaluating Vision-Language Models

2024-02-19 · Kazuki Hayashi, Kazuma Onishi, Toma Suzuki, Yusuke Ide, Seiji Gobara, Shigeki Saito, Yusuke Sakai, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe

Large-scale Vision-Language Models (LVLMs) process both images and text, excelling in multimodal tasks such as image captioning and description generation. However, while these models excel at generating factual content, their ability to generate and evaluate texts reflecting perspectives on the same image, depending on the context, has not been sufficiently explored. To address this, we propose IRR: Image Review Rank, a novel evaluation framework designed to assess critic review texts from multiple perspectives. IRR evaluates LVLMs by measuring how closely their judgments align with human interpretations. We validate it using a dataset of images from 15 categories, each with five critic review texts and annotated rankings in both English and Japanese, totaling over 2,000 data instances. The datasets are available at https://hf.co/datasets/naist-nlp/Wiki-ImageReview1.0. Our results indicate that, although LVLMs exhibited consistent performance across languages, their correlation with human annotations was insufficient, highlighting the need for further advancements. These findings highlight the limitations of current evaluation methods and the need for approaches that better capture human reasoning in Vision & Language tasks.

📄 PDF Abstract BibTeX arXiv:2402.12121

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage Captioning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Responding E-commerce Product Questions via Exploiting QA Collections and Reviews

2018-08-01 · COLING 2018 8 · Qian Yu, Wai Lam, ZiHao Wang

Providing instant responses for product questions in E-commerce sites can significantly improve satisfaction of potential consumers. We propose a new framework for automatically responding product questions newly posed b…

Learning-To-RankSentence

The Second LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

2026-07-23 · Xiang Chen, Hao Li, Jiangxin Dong, Jinshan Pan 외 arxiv

This paper presents a review of the second LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aims to advance unified image restoration under diverse real-world degradation conditions, including bl…

Unified Image Restoration

Positive emotions help rank negative reviews in e-commerce

2020-05-20 · Di Weng, Jichang Zhao

Negative reviews, the poor ratings in postpurchase evaluation, play an indispensable role in e-commerce, especially in shaping future sales and firm equities. However, extant studies seldom examine their potential value …

Attribute

Ranking Online Reviews Based on Their Helpfulness: An Unsupervised Approach

2021-09-01 · RANLP 2021 9 · Alimuddin Melleng, Anna Jurek-Loughrey, Deepak P

Online reviews are an essential aspect of online shopping for both customers and retailers. However, many reviews found on the Internet lack in quality, informativeness or helpfulness. In many cases, they lead the custom…

InformativenessSpecificity

NucFuseRank: Dataset Fusion and Performance Ranking for Nuclei Instance Segmentation

2026-01-27 · Nima Torbati, Anastasia Meshcheryakova, Ramona Woitek, Sepideh Hatamikia 외 arxiv

Nuclei instance segmentation in hematoxylin and eosin (H&E)-stained images plays an important role in automated histological image analysis, with various applications in downstream tasks. While several machine learning a…

Instance Segmentation