paper-with-me

Papers

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

2025-07-05 · Yifan Jiang, Yibo Xue, Yukun Kang, Pin Zheng, Jian Peng, Feiran Wu, Changliang Xu arxiv

Slide animations, such as fade-in, fly-in, and wipe, are critical for audience engagement, efficient information delivery, and vivid visual expression. However, most AI-driven slide-generation tools still lack native animation support, and existing vision-language models (VLMs) struggle with animation tasks due to the absence of public datasets and limited temporal-reasoning capabilities. To address this gap, we release the first public dataset for slide-animation modeling: 12,000 triplets of natural-language descriptions, animation JSON files, and rendered videos, collectively covering every built-in PowerPoint effect. Using this resource, we fine-tune Qwen-2.5-VL-7B with Low-Rank Adaptation (LoRA) and achieve consistent improvements over GPT-4.1 and Gemini-2.5-Pro in BLEU-4, ROUGE-L, SPICE, and our Coverage-Order-Detail Assessment (CODA) metric, which evaluates action coverage, temporal order, and detail fidelity. On a manually created test set of slides, the LoRA model increases BLEU-4 by around 60%, ROUGE-L by 30%, and shows significant improvements in CODA-detail. This demonstrates that low-rank adaptation enables reliable temporal reasoning and generalization beyond synthetic data. Overall, our dataset, LoRA-enhanced model, and CODA metric provide a rigorous benchmark and foundation for future research on VLM-based dynamic slide generation.

📄 PDF Abstract BibTeX arXiv:2507.03916

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

3DiFACE: Synthesizing and Editing Holistic 3D Facial Animation

2025-09-30 · Balamurugan Thambiraja, Malte Prinzler, Sadegh Aliakbarian, Darren Cosker 외 arxiv

Creating personalized 3D animations with precise control and realistic head motions remains challenging for current speech-driven 3D facial animation methods. Editing these animations is especially complex and time consu…

LottieGPT: Tokenizing Vector Animation for Autoregressive Generation

2026-04-13 · Junhao Chen, Kejun Gao, Yuehan Cui, Mingze Sun 외 arxiv

Despite rapid progress in video generation, existing models are incapable of producing vector animation, a dominant and highly expressive form of multimedia on the Internet. Vector animations offer resolution-independenc…

Video Generation

Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters

2024-02-21 · Zechen Bai, Peng Chen, Xiaolan Peng, Lu Liu 외

Animating virtual characters has always been a fundamental research problem in virtual reality (VR). Facial animations play a crucial role as they effectively convey emotions and attitudes of virtual humans. However, cre…

Unity

AnimationBench: Are Video Models Good at Character-Centric Animation?

2026-04-16 · Leyi Wu, Pengjun Fang, Kai Sun, Yazhou Xing 외 arxiv

Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks-largely designed for realistic videos-struggle to evaluate animation-style gener…

Video Generation

Designing Child-Friendly AI Interfaces: Six Developmentally-Appropriate Design Insights from Analysing Disney Animation

2025-04-11 · Nomisha Kurian

To build AI interfaces that children can intuitively understand and use, designers need a design grammar that truly serves children's developmental needs. This paper bridges Artificial Intelligence design for children --…