paper-with-me

Papers

MotionCom: Automatic and Motion-Aware Image Composition with LLM and Video Diffusion Prior

2024-09-16 · Weijing Tao, Xiaofeng Yang, Miaomiao Cui, Guosheng Lin

This work presents MotionCom, a training-free motion-aware diffusion based image composition, enabling automatic and seamless integration of target objects into new scenes with dynamically coherent results without finetuning or optimization. Traditional approaches in this area suffer from two significant limitations: they require manual planning for object placement and often generate static compositions lacking motion realism. MotionCom addresses these issues by utilizing a Large Vision Language Model (LVLM) for intelligent planning, and a Video Diffusion prior for motion-infused image synthesis, streamlining the composition process. Our multi-modal Chain-of-Thought (CoT) prompting with LVLM automates the strategic placement planning of foreground objects, considering their potential motion and interaction within the scenes. Complementing this, we propose a novel method MotionPaint to distill motion-aware information from pretrained video diffusion models in the generation phase, ensuring that these objects are not only seamlessly integrated but also endowed with realistic motion. Extensive quantitative and qualitative results highlight MotionCom's superiority, showcasing its efficiency in streamlining the planning process and its capability to produce compositions that authentically depict motion and interaction.

📄 PDF Abstract BibTeX arXiv:2409.10090

Code (1)

weijing-tao/MotionCom 공식 구현 pytorch

Tasks

Image GenerationLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

EmoStory: Emotion-Aware Story Generation

2026-03-11 · Jingyuan Yang, Rucong Chen, Weibin Luo, Hui Huang arxiv

Story generation aims to produce image sequences that depict coherent narratives while maintaining subject consistency across frames. Although existing methods have excelled in producing coherent and expressive stories, …

Story Generation

EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation

2026-07-11 · Dexiang Hong, Yijie Guo, Weidong Chen, Xinyan Liu 외 arxiv

Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main difficulty is that the visual and affec…

Image Generation

PhysLayer: Language-Guided Layered Animation with Depth-Aware Physics

2026-04-26 · Tianyidan Xie, Zhentao Huang, Mingjie Wang, Xin Huang 외 arxiv

Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior approaches have incorporated physics simulators, they remain confined to …

Computational EfficiencyScene Understanding3D ReconstructionVideo Generation

LICA: Layered Image Composition Annotations for Graphic Design Research

2026-03-17 · Elad Hirsch, Shubham Yadav, Mohit Garg, Purvanshi Mehta arxiv

We introduce LICA (Layered Image Composition Annotations), a large scale dataset of 1,550,244 multi-layer graphic design compositions designed to advance structured understanding and generation of graphic layouts. In add…

EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space

2024-12-19 · CVPR 2025 1 · Jianrong Zhang, Hehe Fan, Yi Yang

Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose mult…

Motion GenerationSemantic Composition