paper-with-me

Papers

DragAnything: Motion Control for Anything using Entity Representation

2024-03-12 · Weijia Wu, Zhuang Li, YuChao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, Di Zhang

We introduce DragAnything, which utilizes a entity representation to achieve motion control for any object in controllable video generation. Comparison to existing motion control methods, DragAnything offers several advantages. Firstly, trajectory-based is more userfriendly for interaction, when acquiring other guidance signals (e.g., masks, depth maps) is labor-intensive. Users only need to draw a line (trajectory) during interaction. Secondly, our entity representation serves as an open-domain embedding capable of representing any object, enabling the control of motion for diverse entities, including background. Lastly, our entity representation allows simultaneous and distinct motion control for multiple objects. Extensive experiments demonstrate that our DragAnything achieves state-of-the-art performance for FVD, FID, and User Study, particularly in terms of object motion control, where our method surpasses the previous methods (e.g., DragNUWA) by 26% in human voting.

📄 PDF Abstract BibTeX arXiv:2403.07420

Code (2)

showlab/draganything 공식 구현 jax
kwai-kolors/kolors pytorch

Tasks

ObjectVideo Generation

Similar Papers 제목 키워드 기반

SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion

2025-02-17 · Junxian Ma, Shiwen Wang, Jian Yang, Junyi Hu 외

Recent advances in diffusion models have led to significant progress in audio-driven lip synchronization. However, existing methods typically rely on constrained audio-visual alignment priors or multi-stage learning of i…

Motion Synthesis

AnimateAnything: Consistent and Controllable Animation for Video Generation

2024-11-16 · CVPR 2025 1 · Guojun Lei, Chi Wang, Hong Li, Rong Zhang 외

We present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions, including camera trajectories, text prompts, and user moti…

Video Generation

Motion Anything: Any to Motion Generation

2025-03-10 · Zeyu Zhang, Yiran Wang, Wei Mao, Danning Li 외

Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existin…

Motion GenerationMotion Synthesis

Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion

2025-11-18 · Zanxu Wang, Homayoon Beigi arxiv

This paper addresses data quality issues in multimodal emotion recognition in conversation (MERC) through systematic quality control and multi-stage transfer learning. We implement a quality control pipeline for MELD and…

Multimodal Emotion RecognitionTransfer LearningFace RecognitionFace Detection

3DXTalker: Unifying Identity, Lip Sync, Emotion, and Spatial Dynamics in Expressive 3D Talking Avatars

2026-02-11 · Zhongju Wang, Zhenhong Sun, Beier Wang, Yifu Wang 외 arxiv

Audio-driven 3D talking avatar generation is increasingly important in virtual communication, digital humans, and interactive media, where avatars must preserve identity, synchronize lip motion with speech, express emoti…