paper-with-me

홈 › Papers

KB-DMGen: Knowledge-Based Global Guidance and Dynamic Pose Masking for Human Image Generation

2025-07-26 · Shibang Liu, Xuemei Xie, Guangming Shi arxiv

Recent methods using diffusion models have made significant progress in Human Image Generation (HIG) with various control signals such as pose priors. In HIG, both accurate human poses and coherent visual quality are crucial for image generation. However, most existing methods mainly focus on pose accuracy while neglecting overall image quality, often improving pose alignment at the cost of image quality. To address this, we propose Knowledge-Based Global Guidance and Dynamic pose Masking for human image Generation (KB-DMGen). The Knowledge Base (KB), implemented as a visual codebook, provides coarse, global guidance based on input text-related visual features, improving pose accuracy while maintaining image quality, while the Dynamic pose Mask (DM) offers fine-grained local control to enhance precise pose accuracy. By injecting KB and DM at different stages of the diffusion process, our framework enhances pose accuracy through both global and local control without compromising image quality. Experiments demonstrate the effectiveness of KB-DMGen, achieving new state-of-the-art results in terms of AP and CAP on the HumanArt dataset. The project page and code are available at https://lushbng.github.io/KBDMGen.

📄 PDF Abstract BibTeX arXiv:2507.20083

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

PILOC: A Pheromone Inverse Guidance Mechanism and Local-Communication Framework for Dynamic Target Search of Multi-Agent in Unknown Environments

2025-07-10 · Hengrui Liu, Yi Feng, Qilong Zhang arxiv

Multi-Agent Search and Rescue (MASAR) plays a vital role in disaster response, exploration, and reconnaissance. However, dynamic and unknown environments pose significant challenges due to target unpredictability and env…

Reinforcement Learning

Dual-Stream Semantic Guidance with Prototype Anchor Calibration for Source-Fully-Free Adaptation of Vision-Language Models

2026-08-28 · Weiwei Xiang, Shun Peng, Guangyi Xiao, Hao Chen 외 arxiv

Source-Fully-Free Domain Adaptation (SFF-DA) has emerged as a strategic paradigm to adapt Vision-Language Models (VLMs) without any access to source data or task-specific source models. However, we identify a critical Du…

Knowledge DistillationDomain Adaptation

Structure-Accurate Medical Image Translation via Dynamic Frequency Balance and Knowledge Guidance

2025-04-13 · Jiahua Xu, Dawei Zhou, Lei Hu, Zaiyi Liu 외

Multimodal medical images play a crucial role in the precise and comprehensive clinical diagnosis. Diffusion model is a powerful strategy to synthesize the required medical images. However, existing approaches still suff…

Clinical KnowledgeLanguage ModelingLanguage Modelling

Stage-wise Dynamics of Classifier-Free Guidance in Diffusion Models

2025-09-26 · Cheng Jin, Qitan Shi, Yuantao Gu arxiv

Classifier-Free Guidance (CFG) is widely used to improve conditional fidelity in diffusion models, but its impact on sampling dynamics remains poorly understood. Prior studies, often restricted to unimodal conditional di…

STEM: Structure-Tracing Evidence Mining for Knowledge Graphs-Driven Retrieval-Augmented Generation

2026-04-24 · Peng Yu, En Xu, Bin Chen, Haibiao Chen 외 arxiv

Knowledge Graph-based Question Answering (KGQA) plays a pivotal role in complex reasoning tasks but remains constrained by two persistent challenges: the structural heterogeneity of Knowledge Graphs(KGs) often leads to s…

Question AnsweringKnowledge Graphs