paper-with-me

홈 › Papers

Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images

2023-03-13 · ICCV 2023 1 · Nitzan Bitton-Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt, Yuval Elovici, Gabriel Stanovsky, Roy Schwartz

Weird, unusual, and uncanny images pique the curiosity of observers because they challenge commonsense. For example, an image released during the 2022 world cup depicts the famous soccer stars Lionel Messi and Cristiano Ronaldo playing chess, which playfully violates our expectation that their competition should occur on the football field. Humans can easily recognize and interpret these unconventional images, but can AI models do the same? We introduce WHOOPS!, a new dataset and benchmark for visual commonsense. The dataset is comprised of purposefully commonsense-defying images created by designers using publicly-available image generation tools like Midjourney. We consider several tasks posed over the dataset. In addition to image captioning, cross-modal matching, and visual question answering, we introduce a difficult explanation generation task, where models must identify and explain why a given image is unusual. Our results show that state-of-the-art models such as GPT3 and BLIP2 still lag behind human performance on WHOOPS!. We hope our dataset will inspire the development of AI models with stronger visual commonsense reasoning abilities. Data, models and code are available at the project website: whoops-benchmark.github.io

📄 PDF Abstract BibTeX arXiv:2303.07274

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningExplanation GenerationImage CaptioningImage GenerationImage-to-Text RetrievalQuestion AnsweringVisual Commonsense ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Six Ways To Communicate To Someone At Expedia Via Phone And Email's. To communicate or get human at Expedia, the quickest option is typically to call their customer service at +1-888-829-0881 or +1(805) 330 (4056). You can also use the live chat…
Feedforward Network A Feedforward Network, or a Multilayer Perceptron (MLP), is a neural network with solely densely connected layers. This is the classic neural network architecture of the…
TTUR The Two Time-scale Update Rule (TTUR) is an update rule for generative adversarial networks trained with stochastic gradient descent. TTUR has an individual learning rate for…
Projection Discriminator A Projection Discriminator is a type of discriminator for generative adversarial networks. It is motivated by a probabilistic model in which the distribution of the…
Non-Local Operation A Non-Local Operation is a component for capturing long-range dependencies with deep neural networks. It is a generalization of the classical non-local mean operation in…

Similar Papers 제목 키워드 기반

Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images

2025-05-12 · Elisei Rykov, Kseniia Petrushina, Kseniia Titova, Anton Razzhigaev 외

Measuring how real images look is a complex task in artificial intelligence research. For example, an image of a boy with a vacuum cleaner in a desert violates common sense. We introduce a novel method, which we call Thr…

Common Sense Reasoning

Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts

2025-03-20 · Elisei Rykov, Kseniia Petrushina, Kseniia Titova, Alexander Panchenko 외

Quantifying the realism of images remains a challenging problem in the field of artificial intelligence. For example, an image of Albert Einstein holding a smartphone violates common-sense because modern smartphone were …

Common Sense ReasoningNatural Language Inference

VLIS: Unimodal Language Models Guide Multimodal Language Generation

2023-10-15 · Jiwan Chung, Youngjae Yu

Multimodal language generation, which leverages the synergy of language and vision, is a rapidly expanding field. However, existing vision-language models face challenges in tasks that require complex linguistic understa…

Caption GenerationExplanation GenerationImage Paragraph CaptioningLanguage Modeling+4

When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models

2025-07-18 · Francesco Ortu, Zhijing Jin, Diego Doimo, Alberto Cazzaniga arxiv

Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual input can lead to hallucinations and unr…

Commonsense Knowledge Transfer for Pre-trained Language Models

2023-06-04 · Wangchunshu Zhou, Ronan Le Bras, Yejin Choi

Despite serving as the foundation models for a wide range of NLP benchmarks, pre-trained language models have shown limited capabilities of acquiring implicit commonsense knowledge from self-supervision alone, compared t…

Language ModelingLanguage ModellingRelation PredictionTransfer Learning