paper-with-me

Papers

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation

2026-06-26 · Anya Ji, Abhijith Varma Mudunuri, David M. Chan, Alane Suhr arxiv

While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it remains unclear whether they can recover temporal dynamics when motion is present. To this end, we introduce Animation2Code, a benchmark for evaluating temporal visual reasoning via reconstructing executable web animation code from videos. Animation2Code consists of 1,069 web animation videos with diverse visual appearances and motion patterns, paired with corresponding HTML/CSS/JavaScript implementations. We propose two human-aligned metrics, appearance similarity and temporal similarity, which allow us to disentangle visual fidelity from temporal alignment when comparing rendered animations against ground-truth samples. Benchmarking state-of-the-art VLMs on this dataset shows that current VLMs struggle to maintain temporal consistency in reconstruction, even when achieving high appearance similarity, including under finetuning and iterative refinement settings. Code and data are available at https://anya-ji.github.io/animation2code-website .

📄 PDF Abstract BibTeX arXiv:2606.28593

Code (0)

등록된 구현이 없습니다.

Tasks

Visual ReasoningCode Generation

Similar Papers 제목 키워드 기반

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

2026-05-19 · Qiran Zhang, Yuheng Wang, Runde Yang, Lin Wu 외 arxiv

Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated o…

Spatial ReasoningVideo GenerationCode Generation

Training and Benchmarking Code Generation for Physics-Inspired Animations

2026-02-11 · Yanan Wang, Renxi Wang, Yongxin Wang, Xuezhi Liang 외 arxiv

Large language models (LLMs) have been widely studied in areas such as mathematical reasoning, complex coding, and scientific problem solving. However, their ability to generate executable code that visually depicts phys…

Reinforcement LearningCode Generation

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

2025-07-05 · Yifan Jiang, Yibo Xue, Yukun Kang, Pin Zheng 외 arxiv

Slide animations, such as fade-in, fly-in, and wipe, are critical for audience engagement, efficient information delivery, and vivid visual expression. However, most AI-driven slide-generation tools still lack native ani…

Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models

2026-02-06 · Haoyu Zhang, Zhipeng Li, Yiwen Guo, Tianshu Yu arxiv

Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely unexplored despite its importance for na…

Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation

2025-12-22 · Ivan DeAndres-Tame, Chengwei Ye, Ruben Tolosana, Ruben Vera-Rodriguez 외 arxiv

Generative AI (GenAI) models have revolutionized animation, enabling the synthesis of humans and motion patterns with remarkable visual fidelity. However, generating truly realistic human animation remains a formidable c…

Person IdentificationGait Recognition