paper-with-me

Papers

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation

2025-07-03 · Qin Guo, Ailing Zeng, Dongxu Yue, Ceyuan Yang, Yang Cao, Hanzhong Guo, Fei Shen, Wei Liu, Xihui Liu, Dan Xu

Although significant advancements have been achieved in the progress of keypoint-guided Text-to-Image diffusion models, existing mainstream keypoint-guided models encounter challenges in controlling the generation of more general non-rigid objects beyond humans (e.g., animals). Moreover, it is difficult to generate multiple overlapping humans and animals based on keypoint controls solely. These challenges arise from two main aspects: the inherent limitations of existing controllable methods and the lack of suitable datasets. First, we design a DiT-based framework, named UniMC, to explore unifying controllable multi-class image generation. UniMC integrates instance- and keypoint-level conditions into compact tokens, incorporating attributes such as class, bounding box, and keypoint coordinates. This approach overcomes the limitations of previous methods that struggled to distinguish instances and classes due to their reliance on skeleton images as conditions. Second, we propose HAIG-2.9M, a large-scale, high-quality, and diverse dataset designed for keypoint-guided human and animal image generation. HAIG-2.9M includes 786K images with 2.9M instances. This dataset features extensive annotations such as keypoints, bounding boxes, and fine-grained captions for both humans and animals, along with rigorous manual inspection to ensure annotation accuracy. Extensive experiments demonstrate the high quality of HAIG-2.9M and the effectiveness of UniMC, particularly in heavy occlusions and multi-class scenarios.

📄 PDF Abstract BibTeX arXiv:2507.02713

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

UniMC: A Unified Framework for Long-Term Memory Conversation via Relevance Representation Learning

2023-06-18 · Kang Zhao, Wei Liu, Jian Luan, Minglei Gao 외

Open-domain long-term memory conversation can establish long-term intimacy with humans, and the key is the ability to understand and memorize long-term dialogue history information. Existing works integrate multiple mode…

Conversation SummarizationDecoderRepresentation LearningRetrieval

Taming Diffusion Probabilistic Models for Character Control

2024-04-23 · Rui Chen, Mingyi Shi, Shaoli Huang, Ping Tan 외

We present a novel character control framework that effectively utilizes motion diffusion probabilistic models to generate high-quality and diverse character animations, responding in real-time to a variety of dynamic us…

Computational EfficiencyDiversity

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

2026-06-29 · Qin Guo, Hao Luo, Dongxu Yue, Weixuan Jin 외 arxiv

Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and …

Image Generation

GKDT: General Keypoint Detection Transformer

2026-07-01 · Changsheng Lu, Yuxin Chen, Haokun Gui, Rong Wang 외 arxiv

With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful yet general keypoint detection (GKD) mo…

Keypoint Detection

UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation

2024-06-03 · Xiang Wang, Shiwei Zhang, Changxin Gao, Jiayu Wang 외

Recent diffusion-based human image animation techniques have demonstrated impressive success in synthesizing videos that faithfully follow a given reference identity and a sequence of desired movement poses. Despite this…

Image AnimationVideo Generation