paper-with-me

Papers

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs

2026-04-12 · Aaditya Baranwal, Vishal Yadav, Abhishek Rajora arxiv

While Vision-Language Models (VLMs) demonstrate remarkable zero-shot recognition capabilities across a diverse spectrum of multimodal tasks, it yet remains an open question whether these architectures genuinely comprehend geometric structure or merely exploit RGB textures and contextual priors as statistical shortcuts. Existing evaluations fail to isolate this mechanism, conflating semantic reasoning with texture mapping and relying on imprecise annotations that inadvertently leak environmental cues. To address this gap, we introduce $\textbf{BareBones}$, a zero-shot benchmark designed to stress-test pure geometric shape comprehension. We curate pixel-level silhouettes of geometrically distinct classes across six datasets: five established segmentation sources (ImageNet-S, DIS5K, ThinObject5K, PASCAL VOC, CUB-200) and our novel flagship collection, WTP-Bench, establishing a noise-free geometric taxonomy. WTP-Bench is an extreme, fine-grained visual puzzle that forces models to identify inter-class geometric concepts from boundary contours alone. Our evaluation of 26 state-of-the-art proprietary and open-weight VLMs (eg. GPT-4.1, Gemini, Claude Sonnet 4.5, LLaVA) reveals a consistent, severe performance collapse under RGB deprivation, a phenomenon we term the $\textit{Texture Bias Cliff}$. By documenting universal structural blindspots, BareBones establishes a rigorous yardstick for genuine geometric grounding. Project Page: https://eternal-f1ame.github.io/WTP-Bench/

📄 PDF Abstract BibTeX arXiv:2604.10528

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge

2024-07-06 · Longfei Huang, Feng Yu, Zhihao Guan, Zhonghua Wan 외

This report presents a solution for the zero-shot referring expression comprehension task. Visual-language multimodal base models (such as CLIP, SAM) have gained significant attention in recent years as a cornerstone of …

Referring ExpressionReferring Expression Comprehension

Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning

2025-10-05 · Mahdi Cherakhloo, Arash Abbasi, Mohammad Saeid Sarafraz, Bijan Vosoughi Vahdat arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities across numerous languages; however, their effectiveness in low-resource languages like Persian requires thorough investigation. This paper presents a…

Reading ComprehensionSentiment AnalysisQuestion AnsweringFew-Shot Learning

Zero-shot Reading Comprehension by Cross-lingual Transfer Learning with Multi-lingual Language Representation Model

2019-09-15 · IJCNLP 2019 11 · Tsung-Yuan Hsu, Chi-Liang Liu, Hung-Yi Lee

Because it is not feasible to collect training data for every language, there is a growing interest in cross-lingual transfer learning. In this paper, we systematically explore zero-shot cross-lingual transfer learning o…

Cross-Lingual TransferReading ComprehensionTransfer LearningZero-Shot Cross-Lingual Transfer+1

Large Language Models are Null-Shot Learners

2024-01-16 · Pittawat Taveekitworachai, Febri Abdullah, Ruck Thawonmas

This paper presents null-shot prompting. Null-shot prompting exploits hallucination in large language models (LLMs) by instructing LLMs to utilize information from the "Examples" section that never exists within the prov…

Arithmetic ReasoningBenchmarkingHallucinationQuestion Answering+1

The Mask of Civility: Benchmarking Chinese Mock Politeness Comprehension in Large Language Models

2026-02-03 · Yitong Zhang, Yuhan Xiang, Mingxuan Liu arxiv

From a pragmatic perspective, this study systematically evaluates the differences in performance among representative large language models (LLMs) in recognizing politeness, impoliteness, and mock politeness phenomena in…