paper-with-me

홈 › Papers

SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models

2025-10-09 · Andong Deng, Taojiannan Yang, Shoubin Yu, Lincoln Spencer, Mohit Bansal, Chen Chen, Serena Yeung-Levy, Xiaohan Wang arxiv

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchmarks predominantly target general scenarios where perception/recognition is heavily relied on, while with relatively simple reasoning tasks, leading to saturation and thus failing to effectively evaluate advanced multimodal cognitive skills. To address this critical gap, we introduce SciVideoBench, a rigorous benchmark specifically designed to assess advanced video reasoning in scientific contexts. SciVideoBench consists of 1,000 carefully crafted multiple-choice questions derived from cutting-edge scientific experimental videos spanning over 25 specialized academic subjects and verified by a semi-automatic system. Each question demands sophisticated domain-specific knowledge, precise spatiotemporal perception, and intricate logical reasoning, effectively challenging models' higher-order cognitive abilities. Our evaluation highlights significant performance deficits in state-of-the-art proprietary and open-source LMMs, including Gemini 2.5 Pro and Qwen2.5-VL, indicating substantial room for advancement in video reasoning capabilities. Detailed analyses of critical factors such as reasoning complexity and visual grounding provide valuable insights and clear direction for future developments in LMMs, driving the evolution of truly capable multimodal AI co-scientists. We hope SciVideoBench could fit the interests of the community and help to push the boundary of cutting-edge AI for border science.

📄 PDF Abstract BibTeX arXiv:2510.08559

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench

2025-12-02 · Lanxiang Hu, Abhilash Shankarampeta, Yixin Huang, Zilin Dai 외 arxiv

The next frontier for video generation lies in developing models capable of zero-shot reasoning, where understanding real-world scientific laws is crucial for accurate physical outcome modeling under diverse conditions. …

Video Generation

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

2026-07-17 · Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang 외 hf

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without…

Video Generation

Morpheus: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments

2025-04-03 · Chenyu Zhang, Daniil Cherniavskii, Andrii Zadaianchuk, Antonios Tragoudaras 외

Recent advances in image and video generation raise hopes that these models possess world modeling capabilities, the ability to generate realistic, physically plausible videos. This could revolutionize applications in ro…

Physical Commonsense ReasoningVideo Generation

MMSciBench: Benchmarking Language Models on Multimodal Scientific Problems

2025-02-27 · Xinwu Ye, Chengfan Li, Siming Chen, Xiangru Tang 외

Recent advances in large language models (LLMs) and vision-language models (LVLMs) have shown promise across many tasks, yet their scientific reasoning capabilities remain untested, particularly in multimodal settings. W…

BenchmarkingVisual Reasoning

Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

2026-01-17 · Honglin Lin, Chonghan Qin, Zheng Liu, Qizhi Pei 외 arxiv

While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to…

Multimodal Reasoning