paper-with-me

Papers

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

2026-05-19 · Qiran Zhang, Yuheng Wang, Runde Yang, Lin Wu, Jingru Fan, Shu Yao, Jie Zhang, Tianle Zhou, Huatao Li, Ruijie Shi, Yihan Li, Chen Qian arxiv

Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated outputs remains an open problem. We introduce PRISM, a large-scale benchmark of 10,372 human-calibrated instruction-code pairs (20 times larger than prior programmatic video generation benchmarks), grounded in real-world knowledge visualization scenarios across English and Chinese and spanning 437 subject categories. We further propose a funnel-style evaluation framework with four complementary metrics: Code-Level Reliability for executability, Spatial Reasoning for layout correctness over full animation sequences, and Prompt-Aware Dynamic Visual Complexity (PADVC) and Temporal Density (TD) for diagnosing dynamic expression and temporal activity. Systematic evaluation of seven mainstream LLMs reveals a striking Execution-Spatial Gap: the average drop from execution success rate to spatial pass rate is approximately 41%, showing that runnable code does not necessarily yield spatially coherent visual output. These findings show that programmatic video generation evaluation should go beyond executability. PRISM provides a principled benchmark for advancing spatially coherent code generation.

📄 PDF Abstract BibTeX arXiv:2605.19382

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningVideo GenerationCode Generation

Similar Papers 제목 키워드 기반

PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation

2025-11-24 · Huadai Liu, Kaicheng Luo, Wen Wang, Qian Chen 외 arxiv

Video-to-Audio (V2A) generation requires balancing four critical perceptual dimensions: semantic consistency, audio-visual temporal synchrony, aesthetic quality, and spatial accuracy; yet existing methods suffer from obj…

Reinforcement LearningAudio Generation

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

2026-05-08 · Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami arxiv

Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent video streams remains poorly understood. Existing multi-video bench…

Spatial Reasoning

MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

2025-06-05 · Zikui Cai, Andrew Wang, Anirudh Satheesh, Ankit Nakhawa 외

Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static images, failing to capture the temporal c…

Dataset GenerationMathematical Problem-SolvingMultimodal Reasoning

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning

2026-06-19 · Awais Rauf, Ahmed Hasssan, Greg Slabaugh arxiv

Understanding long videos requires fine-grained perception and multi-step, higher-order reasoning over complex, long-range spatio-temporal dynamics. Vision-language models (VLMs) encode video frames into visual tokens an…

Relational ReasoningSemantic Retrieval

PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models

2026-03-31 · Amirreza Rouhi, Parikshit Sakurikar, Satya Sai Reddy, Narsimha Menga 외 arxiv

A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deployment environments. We present PRISM, a 2…

Action Understanding