paper-with-me

Papers

Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning

2023-07-12 · Gengyuan Zhang, Yurui Zhang, Kerui Zhang, Volker Tresp

Vision-Language Models (VLMs) are expected to be capable of reasoning with commonsense knowledge as human beings. One example is that humans can reason where and when an image is taken based on their knowledge. This makes us wonder if, based on visual cues, Vision-Language Models that are pre-trained with large-scale image-text resources can achieve and even outperform human's capability in reasoning times and location. To address this question, we propose a two-stage \recognition\space and \reasoning\space probing task, applied to discriminative and generative VLMs to uncover whether VLMs can recognize times and location-relevant features and further reason about it. To facilitate the investigation, we introduce WikiTiLo, a well-curated image dataset compromising images with rich socio-cultural cues. In the extensive experimental studies, we find that although VLMs can effectively retain relevant features in visual encoders, they still fail to make perfect reasoning. We will release our dataset and codes to facilitate future studies.

📄 PDF Abstract BibTeX arXiv:2307.06166

Code (1)

gengyuanmax/WikiTiLo 공식 구현

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Good Questions Help Zero-Shot Image Reasoning

2023-12-04 · Kaiwen Yang, Tao Shen, Xinmei Tian, Xiubo Geng 외

Aligning the recent large language models (LLMs) with computer vision models leads to large vision-language models (LVLMs), which have paved the way for zero-shot image reasoning tasks. However, LVLMs are usually trained…

Fine-Grained Image ClassificationQuestion AnsweringVisual EntailmentVisual Question Answering

Learning Better Visual Dialog Agents with Pretrained Visual-Linguistic Representation

2021-05-24 · CVPR 2021 1 · Tao Tu, Qing Ping, Govind Thattai, Gokhan Tur 외

GuessWhat?! is a two-player visual dialog guessing game where player A asks a sequence of yes/no questions (Questioner) and makes a final guess (Guesser) about a target object in an image, based on answers from player B …

Referring ExpressionReferring Expression ComprehensionVisual DialogVisual Grounding

Guessing State Tracking for Visual Dialogue

2020-02-24 · ECCV 2020 8 · Wei Pang, Xiaojie Wang

The Guesser is a task of visual grounding in GuessWhat?! like visual dialogue. It locates the target object in an image supposed by an Oracle oneself over a question-answer based dialogue between a Questioner and the Ora…

Visual Grounding

Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations

2024-02-19 · Naquee Rizwan, Paramananda Bhaskar, Mithun Das, Swadhin Satyaprakash Majhi 외

There is a rapid increase in the use of multimedia content in current social media platforms. One of the highly popular forms of such multimedia content are memes. While memes have been primarily invented to promote funn…

Prompt EngineeringZero-Shot Learning

ArchiGuesser -- AI Art Architecture Educational Game

2023-12-14 · Joern Ploennigs, Markus Berger, Eva Carnein

The use of generative AI in education is a controversial topic. Current technology offers the potential to create educational content from text, speech, to images based on simple input prompts. This can enhance productiv…

Board GamesDiversityImage Generation