paper-with-me

홈 › Papers

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

2025-02-04 · Hila Chefer, Uriel Singer, Amit Zohar, Yuval Kirstain, Adam Polyak, Yaniv Taigman, Lior Wolf, Shelly Sheynin

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence. To address this, we introduce VideoJAM, a novel framework that instills an effective motion prior to video generators, by encouraging the model to learn a joint appearance-motion representation. VideoJAM is composed of two complementary units. During training, we extend the objective to predict both the generated pixels and their corresponding motion from a single learned representation. During inference, we introduce Inner-Guidance, a mechanism that steers the generation toward coherent motion by leveraging the model's own evolving motion prediction as a dynamic guidance signal. Notably, our framework can be applied to any video model with minimal adaptations, requiring no modifications to the training data or scaling of the model. VideoJAM achieves state-of-the-art performance in motion coherence, surpassing highly competitive proprietary models while also enhancing the perceived visual quality of the generations. These findings emphasize that appearance and motion can be complementary and, when effectively integrated, enhance both the visual quality and the coherence of video generation. Project website: https://hila-chefer.github.io/videojam-paper.github.io/

📄 PDF Abstract BibTeX arXiv:2502.02492

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Generationmotion predictionVideo Generation

Similar Papers 제목 키워드 기반

Learning Deep Representations of Appearance and Motion for Anomalous Event Detection

2015-10-06 · Dan Xu, Elisa Ricci, Yan Yan, Jingkuan Song 외

We present a novel unsupervised deep learning framework for anomalous event detection in complex video scenes. While most existing works merely use hand-crafted appearance and motion features, we propose Appearance and M…

Anomaly DetectionDenoisingEvent Detection

Identity-Enhanced Network for Facial Expression Recognition

2018-12-11 · Yanwei Li, Xingang Wang, Shilei Zhang, Lingxi Xie 외

Facial expression recognition is a challenging task, arguably because of large intra-class variations and high inter-class similarities. The core drawback of the existing approaches is the lack of ability to discriminate…

Facial Expression RecognitionFacial Expression Recognition (FER)Multi-Task Learning

Attend What You Need: Motion-Appearance Synergistic Networks for Video Question Answering

2021-06-19 · ACL 2021 5 · Ahjeong Seo, Gi-Cheon Kang, Joonhan Park, Byoung-Tak Zhang

Video Question Answering is a task which requires an AI agent to answer questions grounded in video. This task entails three key challenges: (1) understand the intention of various questions, (2) capturing various elemen…

AI AgentQuestion AnsweringVideo Question Answering

Multimodal Emotion Recognition in Conversations via Class-Wise Adaptive Modality Fusion and Affective Geometry

2026-09-09 · Oriol Marín, Roger Marí, Gloria Haro, Rafael Redondo arxiv

Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational context and emotional dynamics. We extend the Self-Distillation Transforme…

Multimodal Emotion Recognition

CAMEL: CAusal Motion Enhancement Tailored for Lifting Text-driven Video Editing

2024-01-01 · CVPR 2024 1 · Guiwei Zhang, Tianyu Zhang, Guanglin Niu, Zichang Tan 외

Text-driven video editing poses significant challenges in exhibiting flicker-free visual continuity while preserving the inherent motion patterns of original videos. Existing methods operate under a paradigm where mo…

Video Editing