paper-with-me

홈 › Papers

VCBench: A Controllable Benchmark for Symbolic and Abstract Challenges in Video Cognition

2024-11-14 · Chenglin Li, Qianglong Chen, Zhi Li, Feng Tao, Yin Zhang

Recent advancements in Large Video-Language Models (LVLMs) have driven the development of benchmarks designed to assess cognitive abilities in video-based tasks. However, most existing benchmarks heavily rely on web-collected videos paired with human annotations or model-generated questions, which limit control over the video content and fall short in evaluating advanced cognitive abilities involving symbolic elements and abstract concepts. To address these limitations, we introduce VCBench, a controllable benchmark to assess LVLMs' cognitive abilities, involving symbolic and abstract concepts at varying difficulty levels. By generating video data with the Python-based engine, VCBench allows for precise control over the video content, creating dynamic, task-oriented videos that feature complex scenes and abstract concepts. Each task pairs with tailored question templates that target specific cognitive challenges, providing a rigorous evaluation test. Our evaluation reveals that even state-of-the-art (SOTA) models, such as Qwen2-VL-72B, struggle with simple video cognition tasks involving abstract concepts, with performance sharply dropping by 19% as video complexity rises. These findings reveal the current limitations of LVLMs in advanced cognitive tasks and highlight the critical role of VCBench in driving research toward more robust LVLMs for complex video cognition challenges.

📄 PDF Abstract BibTeX arXiv:2411.09105

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VCBench: Benchmarking LLMs in Venture Capital

2025-09-17 · Rick Chen, Joseph Ternasky, Afriyie Samuel Kwesi, Ben Griffin 외 arxiv

Benchmarks such as SWE-bench and ARC-AGI demonstrate how shared datasets accelerate progress toward artificial general intelligence (AGI). We introduce VCBench, the first benchmark for predicting founder success in ventu…

Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency

2025-04-24 · Zhikai Wang, Jiashuo Sun, Wenqi Zhang, Zhiqiang Hu 외

Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, cap…

BenchmarkingMathMathematical ReasoningObject Recognition+2

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

2026-03-13 · Pengyiang Liu, Zhongyue Shi, Hongye Hao, Qi Fu 외 arxiv

Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video understanding evaluation across multiple dimensions, they provide limited…

Object Counting

Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning

2026-04-23 · Mohit Vaishnav, Tanel Tammet arxiv

Vision--language models (VLMs) often fail on abstract visual reasoning benchmarks such as Bongard problems, raising the question of whether the main bottleneck lies in reasoning or representation. We study this on Bongar…

Visual GroundingVisual Reasoning

CellARC: Measuring Intelligence with Cellular Automata

2025-11-11 · Miroslav Lžičař arxiv

We introduce CellARC, a synthetic benchmark for abstraction and reasoning built from multicolor 1D cellular automata (CA). Each episode has five support pairs and one query serialized in 256 tokens, enabling rapid iterat…