paper-with-me

홈 › Papers

Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs

2025-09-29 · Mohamad Ballout, Okajevo Wilfred, Seyedalireza Yaghoubi, Nohayr Muhammad Abdelmoneim, Julius Mayer, Elia Bruni arxiv

In this work, we introduce SPLICE, a human-curated benchmark derived from the COIN instructional video dataset, designed to probe event-based reasoning across multiple dimensions: temporal, causal, spatial, contextual, and general knowledge. SPLICE includes 3,381 human-filtered videos spanning 12 categories and 180 sub-categories, such as sports, engineering, and housework. These videos are segmented into a total of 11,423 event clips. We evaluate both human participants and state-of-the-art vision-language models (VLMs) on the task of rearranging these clips into coherent event sequences to assess visual reasoning capabilities. Results reveal a significant gap: VLMs struggle to match human performance. While human-annotated textual descriptions improve model accuracy, they do not affect human performance, suggesting that models rely more on language priors than on visual understanding. Even with annotations, VLMs fall short of human-level reasoning, underscoring persistent challenges in visual reasoning. A deeper analysis across sub-categories shows that VLMs perform relatively better on videos where temporal and causal reasoning are dominant, compared to those where contextual and spatial reasoning are dominant. They also perform better on everyday tasks than on specialized ones.

📄 PDF Abstract BibTeX arXiv:2509.24640

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningGeneral KnowledgeVisual Reasoning

Similar Papers 제목 키워드 기반

SpliceMix: A Cross-scale and Semantic Blending Augmentation Strategy for Multi-label Image Classification

2023-11-26 · Lei Wang, Yibing Zhan, Leilei Ma, Dapeng Tao 외

Recently, Mix-style data augmentation methods (e.g., Mixup and CutMix) have shown promising performance in various visual tasks. However, these methods are primarily designed for single-label images, ignoring the conside…

Data Augmentationimage-classificationImage ClassificationMulti-Label Image Classification

Sequential Labelling and DNABERT For Splice Site Prediction in Homo Sapiens DNA

2022-12-15 · Muhammad Anwari Leksono, Ayu Purwarianti

Genome sequencing technology has improved significantly in few last years and resulted in abundance genetic data. Artificial intelligence has been employed to analyze genetic data in response to its sheer size and variab…

PredictionSplice Site Prediction

Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language Models

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Knowledge probing is crucial for understanding the knowledge transfer mechanism behind the pre-trained language models (PLMs). Despite the growing progress of probing knowledge for PLMs in the general domain, specialised…

Knowledge ProbingTransfer Learning

Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language Models

2021-10-15 · ACL 2022 5 · Zaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su 외

Knowledge probing is crucial for understanding the knowledge transfer mechanism behind the pre-trained language models (PLMs). Despite the growing progress of probing knowledge for PLMs in the general domain, specialised…

Knowledge ProbingTransfer Learning

Improving spliced alignment by modeling splice sites with deep learning

2025-06-15 · Siying Yang, Neng Huang, Heng Li

Motivation: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alig…