paper-with-me

Papers Video Generation

“Video Generation” 태그가 달린 논문 2,911편 · 필터 해제

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

2026-09-10 · Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang 외 hf

We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video g…

Instruction FollowingVideo Generation

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

2026-09-08 · Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong 외 hf

Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Mat…

Video Generation

RoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers

2026-09-06 · Zekun Zhang, Yixiang Cai, Yuxi Liu, Tengxu Sun 외 arxiv

Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes the dominant inference bottleneck. Sparse…

Video Generation

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

2026-09-06 · Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang 외 hf

Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across shots. Existing video generators favor continuous motion and struggle to present complete shot sets wh…

Video Generation

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

2026-09-04 · Xin Zhang, Yabo Chen, Zixuan Duan, Haibin Huang 외 arxiv

Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, and deformation. Current video world model…

Video Generation

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

2026-09-04 · Yuchen Sun, Qian Yang, Jun Wang, Detai Xin 외 arxiv

Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in is…

Audio GenerationVideo Generation

ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Features

2026-09-04 · Jay Mahajan, Chang Liu, Rauf Makharov, Viraj Shah 외 arxiv

We present DiT Readout (ReaDiT) Guidance, a lightweight framework for controlling generation with Diffusion Transformer (DiT) models via their internal feature representations. ReaDiT Guidance uses features from a single…

Video Generation

The Attention Triangle in Audio-Video Models

2026-09-03 · Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning 외 hf

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and…

Audio GenerationVideo Generation

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

2026-09-02 · Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou 외 hf

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. …

Audio GenerationVideo Generation

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

2026-09-01 · Haoyuan Shi, Mingtao Chen, Shuo Jiang, Ziyan Chen 외 hf

Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage usi…

Video Generation

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

2026-08-31 · Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang 외 hf

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video genera…

Reinforcement LearningVideo Generation

AgenticGen: Reward-Guided Agentic Video Generation for Advertising

2026-08-31 · Xingyuan Bu, Chengru Song, Hao Zhou, Tao Zhou 외 hf

Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realisti…

Video Generation

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

2026-08-30 · Runjia Qian, Zile Wang, Jihai Zhang, Kai Zou 외 hf

Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable…

Video Generation

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation

2026-08-28 · Yixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan 외 arxiv

Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subje…

Video Generation

How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models

2026-08-28 · Victor Besnier, Anh-Quan Cao, Elias Ramzi, Spyros Gidaris 외 arxiv

Video generation for autonomous driving cannot follow the web-scale route: driving data is expensive to collect, bound by privacy requirements, and cannot be scraped at will, so models must make the most of a fixed corpu…

Autonomous DrivingVideo Generation

DensityKV: Density-Guided KV Cache Compression for Long Video Generation

2026-08-28 · Wenqu Zhao, Xuemin Chi, Xin Zhang, Guoqing Ma 외 arxiv

Autoregressive video diffusion models enable streaming generation through sliding-window attention, but each generated block is conditioned on previously generated content, causing appearance and motion errors to propaga…

Video Generation

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

2026-08-27 · Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes 외 hf

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also t…

Video Generation

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

2026-08-27 · Kechen Liu, Ola Shorinwa arxiv

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning …

Video Generation

TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models

2026-08-27 · Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li 외 arxiv

In recent years, image-to-video (I2V) generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis. However, these advances also introduce new safet…

Video Generation

ClusterAttention: A training-free speedup of bidirectional attention

2026-08-27 · Kasper Nordenram, Amelie Dittmann arxiv

This paper introduces ClusterAttention, a general training-free speedup of bidirectional attention layers. Existing sparse attention methods either rely on structure in the input, such as order in language or spatial pro…

Video Generation
1–20 / 2,911 다음 →