paper-with-me

홈 › Papers

ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images

2025-11-30 · Yunfei Zhang, Yizhuo He, Yuanxun Shao, Zhengtao Yao, Haoyan Xu, Junhao Dong, Zhen Yao, Zhikang Dong arxiv

Vision-Language Models (VLMs) have advanced multimodal understanding, yet still struggle when targets are embedded in cluttered backgrounds requiring figure-ground segregation. To address this, we introduce ChromouVQA, a large-scale, multi-task benchmark based on Ishihara-style chromatic camouflaged images. We extend classic dot plates with multiple fill geometries and vary chromatic separation, density, size, occlusion, and rotation, recording full metadata for reproducibility. The benchmark covers nine vision-question-answering tasks, including recognition, counting, comparison, and spatial reasoning. Evaluations of humans and VLMs reveal large gaps, especially under subtle chromatic contrast or disruptive geometric fills. We also propose a model-agnostic contrastive recipe aligning silhouettes with their camouflaged renderings, improving recovery of global shapes. ChromouVQA provides a compact, controlled benchmark for reproducible evaluation and extension. Code and dataset are available at https://github.com/Chromou-VQA-Benchmark/Chromou-VQA.

📄 PDF Abstract BibTeX arXiv:2512.05137

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

A Benchmarking Protocol for Pansharpening: Dataset, Preprocessing, and Quality Assessment

2021-06-07 · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 2021 6 · Gemine Vivone, Mauro Dalla Mura, Andrea Garzelli, Fabio Pacifici

Comparative evaluation is a requirement for reproducible science and objective assessment of new algorithms. Reproducible research in the field of pansharpening of very high resolution images is a difficult task due to t…

BenchmarkingPansharpening

Signing Outside the Studio: Benchmarking Background Robustness for Continuous Sign Language Recognition

2022-11-01 · Youngjoon Jang, Youngtaek Oh, Jae Won Cho, Dong-Jin Kim 외

The goal of this work is background-robust continuous sign language recognition. Most existing Continuous Sign Language Recognition (CSLR) benchmarks have fixed backgrounds and are filmed in studios with a static monochr…

BenchmarkingDisentanglementSign Language Recognition

Pattern Forming Mechanisms of Color Vision

2023-04-15 · Zily Burstein, David D. Reid, Peter J. Thomas, Jack D. Cowan

While our understanding of the way single neurons process chromatic stimuli in the early visual pathway has advanced significantly in recent years, we do not yet know how these cells interact to form stable representatio…

Information Flow in Biological Networks for Color Vision

2019-12-27 · Jesus Malo

Color Appearance Models are biological networks that consist of a cascade of linear+nonlinear layers that modify the linear measurements at the retinal photo-receptors leading to an internal (nonlinear) representation of…

Dichromatic Model Based Temporal Color Constancy for AC Light Sources

2019-06-01 · CVPR 2019 6 · Jun-Sang Yoo, Jong-Ok Kim

Existing dichromatic color constancy approach commonly requires a number of spatial pixels which have high specularity. In this paper, we propose a novel approach to estimate the illuminant chromaticity of AC light sourc…

Color Constancy