paper-with-me

홈 › Papers

UniHuman: A Unified Model for Editing Human Images in the Wild

2023-12-22 · CVPR 2024 1 · Nannan Li, Qing Liu, Krishna Kumar Singh, Yilin Wang, Jianming Zhang, Bryan A. Plummer, Zhe Lin

Human image editing includes tasks like changing a person's pose, their clothing, or editing the image according to a text prompt. However, prior work often tackles these tasks separately, overlooking the benefit of mutual reinforcement from learning them jointly. In this paper, we propose UniHuman, a unified model that addresses multiple facets of human image editing in real-world settings. To enhance the model's generation quality and generalization capacity, we leverage guidance from human visual encoders and introduce a lightweight pose-warping module that can exploit different pose representations, accommodating unseen textures and patterns. Furthermore, to bridge the disparity between existing human editing benchmarks with real-world data, we curated 400K high-quality human image-text pairs for training and collected 2K human images for out-of-domain testing, both encompassing diverse clothing styles, backgrounds, and age groups. Experiments on both in-domain and out-of-domain test sets demonstrate that UniHuman outperforms task-specific models by a significant margin. In user studies, UniHuman is preferred by the users in an average of 77% of cases. Our project is available at https://github.com/NannanLi999/UniHuman.

📄 PDF Abstract BibTeX arXiv:2312.14985

Code (1)

nannanli999/unihuman 공식 구현

Tasks

2k

Similar Papers 제목 키워드 기반

You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception

2023-12-09 · Sheng Jin, Shuhuai Li, Tong Li, Wentao Liu 외

Human-centric perception (e.g. detection, segmentation, pose estimation, and attribute analysis) is a long-standing problem for computer vision. This paper introduces a unified and versatile framework (HQNet) for single-…

AttributeHuman Instance SegmentationMulti-Task LearningPose Estimation

In-the-Wild Camouflage Attack on Vehicle Detectors through Controllable Image Editing

2026-03-19 · Xiao Fang, Yiming Gong, Stanislav Panev, Celso de Melo 외 arxiv

Deep neural networks (DNNs) have achieved remarkable success in computer vision but remain highly vulnerable to adversarial attacks. Among them, camouflage attacks manipulate an object's visible appearance to deceive det…

Image Editing

M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing

2024-02-04 · Mohammadreza Mofayezi, Reza Alipour, Mohammad Ali Kakavand, Ehsaneddin Asgari

Human face generation and editing represent an essential task in the era of computer vision and the digital world. Recent studies have shown remarkable progress in multi-modal face generation and editing, for instance, u…

Face GenerationImage GenerationSemantic Segmentation

Synthesizing Anyone, Anywhere, in Any Pose

2023-04-06 · Håkon Hukkelås, Frank Lindseth

We address the task of in-the-wild human figure synthesis, where the primary goal is to synthesize a full body given any region in any image. In-the-wild human figure synthesis has long been a challenging and under-explo…

ShadowWolf -- Automatic Labelling, Evaluation and Model Training Optimised for Camera Trap Wildlife Images

2025-12-06 · Jens Dede, Anna Förster arxiv

The continuous growth of the global human population is leading to the expansion of human habitats, resulting in decreasing wildlife spaces and increasing human-wildlife interactions. These interactions can range from mi…