paper-with-me

홈 › Papers

Understanding Human-Centric Images: From Geometry to Fashion

2015-12-14 · Edgar Simo-Serra

Understanding humans from photographs has always been a fundamental goal of computer vision. In this thesis we have developed a hierarchy of tools that cover a wide range of topics with the objective of understanding humans from monocular RGB image: from low level feature point descriptors to high level fashion-aware conditional random fields models. In order to build these high level models it is paramount to have a battery of robust and reliable low and mid level cues. Along these lines, we have proposed two low-level keypoint descriptors: one based on the theory of the heat diffusion on images, and the other that uses a convolutional neural network to learn discriminative image patch representations. We also introduce distinct low-level generative models for representing human pose: in particular we present a discrete model based on a directed acyclic graph and a continuous model that consists of poses clustered on a Riemannian manifold. As mid level cues we propose two 3D human pose estimation algorithms: one that estimates the 3D pose given a noisy 2D estimation, and an approach that simultaneously estimates both the 2D and 3D pose. Finally, we formulate higher level models built upon low and mid level cues for understanding humans from single images. Concretely, we focus on two different tasks in the context of fashion: semantic segmentation of clothing, and predicting the fashionability from images with metadata to ultimately provide fashion advice to the user. For all presented approaches we present extensive results and comparisons against the state-of-the-art and show significant improvements on the entire variety of tasks we tackle.

📄 PDF Abstract BibTeX arXiv:1604.08164

Code (0)

등록된 구현이 없습니다.

Tasks

3D Human Pose EstimationPose EstimationSemantic Segmentation

Similar Papers 제목 키워드 기반

HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction

2025-08-22 · Sara Rojas, Matthieu Armando, Bernard Ghamen, Philippe Weinzaepfel 외 arxiv

Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive …

Human Mesh RecoveryScene Understanding3D Reconstruction

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

2023-04-04 · ICCV 2023 1 · Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia 외

Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision…

Multimodal fashion image editing

Object Scene Representation Transformer

2022-06-14 · Mehdi S. M. Sajjadi, Daniel Duckworth, Aravindh Mahendran, Sjoerd van Steenkiste 외

A compositional understanding of the world in terms of objects and their geometry in 3D space is considered a cornerstone of human cognition. Facilitating the learning of such a representation in neural networks holds pr…

DecoderDiversityNovel View SynthesisObject+1

Egocentric Scene Understanding via Multimodal Spatial Rectifier

2022-07-14 · CVPR 2022 1 · Tien Do, Khiem Vuong, Hyun Soo Park

In this paper, we study a problem of egocentric scene understanding, i.e., predicting depths and surface normals from an egocentric image. Egocentric scene understanding poses unprecedented challenges: (1) due to large h…

Scene UnderstandingSurface Normal Estimation

ICGNet: A Unified Approach for Instance-Centric Grasping

2024-01-18 · René Zurbrügg, Yifan Liu, Francis Engelmann, Suryansh Kumar 외

Accurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of scene understanding: First, the robot n…

ObjectObject ReconstructionScene Understanding