paper-with-me

홈 › Papers

Synthesizing Environment-Specific People in Photographs

2023-12-22 · Mirela Ostrek, Carol O'Sullivan, Michael J. Black, Justus Thies

We present ESP, a novel method for context-aware full-body generation, that enables photo-realistic synthesis and inpainting of people wearing clothing that is semantically appropriate for the scene depicted in an input photograph. ESP is conditioned on a 2D pose and contextual cues that are extracted from the photograph of the scene and integrated into the generation process, where the clothing is modeled explicitly with human parsing masks (HPM). Generated HPMs are used as tight guiding masks for inpainting, such that no changes are made to the original background. Our models are trained on a dataset containing a set of in-the-wild photographs of people covering a wide range of different environments. The method is analyzed quantitatively and qualitatively, and we show that ESP outperforms the state-of-the-art on the task of contextual full-body generation.

📄 PDF Abstract BibTeX arXiv:2312.14579

Code (0)

등록된 구현이 없습니다.

Tasks

Human ParsingImage Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Hierarchical Feature Fusion Hierarchical Feature Fusion (HFF) is a feature fusion method employed in ESP and EESP image…
Pointwise Convolution Pointwise Convolution is a type of convolution that uses a 1x1 kernel: a kernel that iterates through every single point. This…
Dilated Convolution 설명 없음
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
ESP 설명 없음

Similar Papers 제목 키워드 기반

Synthesizing Normalized Faces from Facial Identity Features

2017-01-17 · CVPR 2017 7 · Forrester Cole, David Belanger, Dilip Krishnan, Aaron Sarna 외

We present a method for synthesizing a frontal, neutral-expression image of a person's face given an input face photograph. This is achieved by learning to generate facial landmarks and textures from features extracted f…

Decoder

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph

2026-08-03 · Zhenhao Zhang, Jiajun Zhang, Wei Min, Yebin Liu arxiv

Hand-object interaction (HOI) is a fundamental human behavior with broad applications in AR/VR, digital humans, and embodied interaction. Existing methods typically require predefined object geometry, object trajectories…

Blind Dates: Examining the Expression of Temporality in Historical Photographs

2023-10-10 · Alexandra Barancová, Melvin Wevers, Nanne van Noord

This paper explores the capacity of computer vision models to discern temporal information in visual content, focusing specifically on historical photographs. We investigate the dating of images using OpenCLIP, an open-s…

zero-shot-classificationZero-Shot Learning

VIP: Finding Important People in Images

2015-02-19 · CVPR 2015 6 · Clint Solomon Mathialagan, Andrew C. Gallagher, Dhruv Batra

People preserve memories of events such as birthdays, weddings, or vacations by capturing photos, often depicting groups of people. Invariably, some individuals in the image are more important than others given the conte…

Automated Generation of Storytelling Vocabulary from Photographs for use in AAC

2021-08-01 · ACL 2021 5 · Mauricio Fontana de Vargas, Karyn Moffatt

Research on the application of NLP in symbol-based Augmentative and Alternative Communication (AAC) tools for improving social interaction support is scarce. We contribute a novel method for generating context-related vo…

Information RetrievalRetrieval