paper-with-me

홈 › Papers

Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models

2025-09-23 · Julien Delavande, Regis Pierrard, Sasha Luccioni arxiv

Recent advances in text-to-video (T2V) generation have enabled the creation of high-fidelity, temporally coherent clips from natural language prompts. Yet these systems come with significant computational costs, and their energy demands remain poorly understood. In this paper, we present a systematic study of the latency and energy consumption of state-of-the-art open-source T2V models. We first develop a compute-bound analytical model that predicts scaling laws with respect to spatial resolution, temporal length, and denoising steps. We then validate these predictions through fine-grained experiments on WAN2.1-T2V, showing quadratic growth with spatial and temporal dimensions, and linear scaling with the number of denoising steps. Finally, we extend our analysis to six diverse T2V models, comparing their runtime and energy profiles under default settings. Our results provide both a benchmark reference and practical insights for designing and deploying more sustainable generative video systems.

📄 PDF Abstract BibTeX arXiv:2509.19222

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs

2025-10-21 · Benjamin Kubwimana, Qijing Huang arxiv

Edge intelligence paradigm is increasingly demanded by the emerging autonomous systems, such as robotics. Beyond ensuring privacy-preserving operation and resilience in connectivity-limited environments, edge deployment …

Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges

2025-11-07 · Md Romyull Islam, Bobin Deng, Nobel Dhar, Tu N. Nguyen 외 arxiv

Cloud-based large language models (LLMs) and their variants have significantly influenced real-world applications. Deploying smaller models (i.e., small language models (SLMs)) on edge devices offers additional advantage…

Polymorph: Energy-Efficient Multi-Label Classification for Video Streams on Embedded Devices

2025-07-20 · Saeid Ghafouri, Mohsen Fayyaz, Xiangchen Li, Deepu John 외 arxiv

Real-time multi-label video classification on embedded devices is constrained by limited compute and energy budgets. Yet, video streams exhibit structural properties such as label sparsity, temporal continuity, and label…

Multi-Label ClassificationVideo Classification

VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

2026-05-28 · Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil Kaan Akan 외 arxiv

Long-rollout causal video diffusion has converged on a fixed-size sliding-window KV cache, with recent progress innovating within this layout by changing which tokens occupy the window or how their positions are encoded.…

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

2026-08-06 · Wang Chen, Yu Chen, Xiang Wang, Shuai Li 외 arxiv

Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reaso…