paper-with-me

홈 › Papers

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

2026-06-05 · Yifan Xu, Chao Zhang, Ruifei Ma, Fei Gao, Zhifei Yang, Jiaxing Qi, Zhipeng Chen arxiv

The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-level understanding, their ability to capture fine-grained motion details remains limited, primarily due to their focus on high-level static semantic structures and macro-event logic. In contrast, Video Diffusion Models (VDMs) are adept at modeling dynamic motion patterns, benefiting from large-scale video data and the intrinsic requirement of temporal generation. In this paper, we introduce MotionEnhancer, a novel approach that leverages motion priors distilled from a powerful video diffusion model as auxiliary supervision to enhance the motion understanding capability of a VLM via attention alignment. MotionEnhancer comprises two simple parameter-free modules, Motion-sensitive Head Selection (MHS) and Motion-salient Text Token Identification (MTTI), to directly extract and optimize motion-related attentions from the VDM in a computation-only manner. MotionEnhancer provides a scalable solution for motion understanding without additional training parameters, modifications to existing architectures, or tool calling. Extensive experiments demonstrate that MotionEnhancer can achieve consistent improvements over state-of-the-art VLMs on two motion-level video understanding benchmarks, especially on motion-related metrics.

📄 PDF Abstract BibTeX arXiv:2606.06853

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis

2025-02-12 · CVPR 2025 1 · Wonjoon Jin, Qi Dai, Chong Luo, Seung-Hwan Baek 외

This paper presents FloVD, a novel optical-flow-based video diffusion model for camera-controllable video generation. FloVD leverages optical flow maps to represent motions of the camera and moving objects. This approach…

Motion SynthesisOptical Flow EstimationVideo Generation

RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution

2025-07-25 · Weisong Zhao, Jingkai Zhou, Xiangyu Zhu, Weihua Chen 외 arxiv

Video Super-Resolution (VSR) has achieved significant progress through diffusion models, effectively addressing the over-smoothing issues inherent in GAN-based methods. Despite recent advances, three critical challenges …

Video Super-Resolution

EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation

2025-03-20 · CVPR 2025 1 · Zihao Zhang, Haoran Chen, Haoyu Zhao, Guansong Lu 외

Handling complex or nonlinear motion patterns has long posed challenges for video frame interpolation. Although recent advances in diffusion-based methods offer improvements over traditional optical flow-based approaches…

Optical Flow EstimationVideo Frame Interpolation

Motion Control for Enhanced Complex Action Video Generation

2024-11-13 · Qiang Zhou, Shaofeng Zhang, Nianzu Yang, Ye Qian 외

Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to precisely convey intricate motion details…

Motion GenerationVideo Generation

CamMimic: Zero-Shot Image To Camera Motion Personalized Video Generation Using Diffusion Models

2025-04-13 · Pooja Guhan, Divya Kothandaraman, Tsung-Wei Huang, Guan-Ming Su 외

We introduce CamMimic, an innovative algorithm tailored for dynamic video editing needs. It is designed to seamlessly transfer the camera motion observed in a given reference video onto any scene of the user's choice in …

Video EditingVideo Generation