paper-with-me

Papers

ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?

2026-03-26 · Haonan Han, Jiancheng Huang, Xiaopeng Sun, Junyan He, Rui Yang, Jie Hu, Xiaojiang Peng, Lin Ma, Xiaoming Wei, Xiu Li arxiv

Beneath the stunning visual fidelity of modern AIGC models lies a "logical desert", where systems fail tasks that require physical, causal, or complex spatial reasoning. Current evaluations largely rely on superficial metrics or fragmented benchmarks, creating a `performance mirage'' that overlooks the generative process. To address this, we introduce ViGoR Vision-G}nerative Reasoning-centric Benchmark), a unified framework designed to dismantle this mirage. ViGoR distinguishes itself through four key innovations: 1) holistic cross-modal coverage bridging Image-to-Image and Video tasks; 2) a dual-track mechanism evaluating both intermediate processes and final results; 3) an evidence-grounded automated judge ensuring high human alignment; and 4) granular diagnostic analysis that decomposes performance into fine-grained cognitive dimensions. Experiments on over 20 leading models reveal that even state-of-the-art systems harbor significant reasoning deficits, establishing ViGoR as a critical `stress test'' for the next generation of intelligent vision models. The demo have been available at https://vincenthancoder.github.io/ViGoR-Bench/

📄 PDF Abstract BibTeX arXiv:2603.25823

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

A Generative Adversarial Approach for Zero-Shot Learning from Noisy Texts

2017-12-04 · CVPR 2018 6 · Yizhe Zhu, Mohamed Elhoseiny, Bingchen Liu, Xi Peng 외

Most existing zero-shot learning methods consider the problem as a visual semantic embedding one. Given the demonstrated capability of Generative Adversarial Networks(GANs) to generate images, we instead leverage GANs to…

ArticlesZero-Shot Learning

Zero-shot Vision-Language Reranking for Cross-View Geolocalization

2026-03-28 · Yunus Talha Erzurumlu, John E. Anderson, William J. Shuart, Charles Toth 외 arxiv

Cross-view geolocalization (CVGL) systems, while effective at retrieving a list of relevant candidates (high Recall@k), often fail to identify the single best match (low Top-1 accuracy). This work investigates the use of…

Emergent Analogical Reasoning in Large Language Models

2022-12-19 · Taylor Webb, Keith J. Holyoak, Hongjing Lu

The recent advent of large language models has reinvigorated debate over whether human cognitive capacities might emerge in such generic models given sufficient training data. Of particular interest is the ability of the…

Language ModelingLanguage ModellingLarge Language Model

Leveraging the Invariant Side of Generative Zero-Shot Learning

2019-04-08 · CVPR 2019 6 · Jingjing Li, Mengmeng Jin, Ke Lu, Zhengming Ding 외

Conventional zero-shot learning (ZSL) methods generally learn an embedding, e.g., visual-semantic mapping, to handle the unseen visual samples via an indirect manner. In this paper, we take the advantage of generative ad…

Generalized Zero-Shot LearningZero-Shot Learning

IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers

2023-11-27 · Chenglin Yang, Siyuan Qiao, Yuan Cao, Yu Zhang 외

Generative training has been demonstrated to be powerful for building visual-language models. However, on zero-shot discriminative benchmarks, there is still a performance gap between models trained with generative and d…

Caption GenerationImage-text RetrievalLanguage ModellingText Retrieval+2