paper-with-me

Papers

Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

2024-07-28 · Nitzan Bitton-Guetta, Aviv Slobodkin, Aviya Maimon, Eliya Habba, Royi Rassin, Yonatan Bitton, Idan Szpektor, Amir Globerson, Yuval Elovici

Imagine observing someone scratching their arm; to understand why, additional context would be necessary. However, spotting a mosquito nearby would immediately offer a likely explanation for the person's discomfort, thereby alleviating the need for further information. This example illustrates how subtle visual cues can challenge our cognitive skills and demonstrates the complexity of interpreting visual scenarios. To study these skills, we present Visual Riddles, a benchmark aimed to test vision and language models on visual riddles requiring commonsense and world knowledge. The benchmark comprises 400 visual riddles, each featuring a unique image created by a variety of text-to-image models, question, ground-truth answer, textual hint, and attribution. Human evaluation reveals that existing models lag significantly behind human performance, which is at 82% accuracy, with Gemini-Pro-1.5 leading with 40% accuracy. Our benchmark comes with automatic evaluation tasks to make assessment scalable. These findings underscore the potential of Visual Riddles as a valuable resource for enhancing vision and language models' capabilities in interpreting complex visual scenarios.

📄 PDF Abstract BibTeX arXiv:2407.19474

Code (0)

등록된 구현이 없습니다.

Tasks

World Knowledge

Similar Papers 제목 키워드 기반

Answering Image Riddles using Vision and Reasoning through Probabilistic Soft Logic

2016-11-17 · Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos

In this work, we explore a genre of puzzles ("image riddles") which involves a set of images and a question. Answering these puzzles require both capabilities involving visual detection (including object, activity recogn…

Activity RecognitionQuestion Answering

RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge

2021-01-02 · Findings (ACL) 2021 8 · Bill Yuchen Lin, Ziyi Wu, Yichi Yang, Dong-Ho Lee 외

Question: I have five fingers but I am not alive. What am I? Answer: a glove. Answering such a riddle-style question is a challenging cognitive process, in that it requires complex commonsense reasoning abilities, an und…

counterfactualCounterfactual ReasoningMultiple-choiceNatural Language Understanding+1

BiRdQA: A Bilingual Dataset for Question Answering on Tricky Riddles

2021-09-23 · Yunxiang Zhang, Xiaojun Wan

A riddle is a question or statement with double or veiled meanings, followed by an unexpected answer. Solving riddle is a challenging task for both machine and human, testing the capability of understanding figurative, c…

Multiple-choiceQuestion Answering

Improving Commonsense in Vision-Language Models via Knowledge Graph Riddles

2022-11-29 · CVPR 2023 1 · Shuquan Ye, Yujia Xie, Dongdong Chen, Yichong Xu 외

This paper focuses on analyzing and improving the commonsense ability of recent popular vision-language (VL) models. Despite the great success, we observe that existing VL-models still lack commonsense knowledge/reasonin…

Data AugmentationDiagnosticRetrieval

Seeing the World through Text: Evaluating Image Descriptions for Commonsense Reasoning in Machine Reading Comprehension

2020-12-01 · LANTERN (COLING) 2020 12 · Diana Galvan-Sosa, Jun Suzuki, Kyosuke Nishida, Koji Matsuda 외

Despite recent achievements in natural language understanding, reasoning over commonsense knowledge still represents a big challenge to AI systems. As the name suggests, common sense is related to perception and as such,…

Common Sense ReasoningMachine Reading ComprehensionNatural Language UnderstandingReading Comprehension