paper-with-me

홈 › Papers

Text-Conditional Contextualized Avatars For Zero-Shot Personalization

2023-04-14 · Samaneh Azadi, Thomas Hayes, Akbar Shah, Guan Pang, Devi Parikh, Sonal Gupta

Recent large-scale text-to-image generation models have made significant improvements in the quality, realism, and diversity of the synthesized images and enable users to control the created content through language. However, the personalization aspect of these generative models is still challenging and under-explored. In this work, we propose a pipeline that enables personalization of image generation with avatars capturing a user's identity in a delightful way. Our pipeline is zero-shot, avatar texture and style agnostic, and does not require training on the avatar at all - it is scalable to millions of users who can generate a scene with their avatar. To render the avatar in a pose faithful to the given text prompt, we propose a novel text-to-3D pose diffusion model trained on a curated large-scale dataset of in-the-wild human poses improving the performance of the SOTA text-to-motion models significantly. We show, for the first time, how to leverage large-scale image datasets to learn human 3D pose parameters and overcome the limitations of motion capture datasets.

📄 PDF Abstract BibTeX arXiv:2304.07410

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage GenerationText to 3DText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

AvatarFusion: Zero-shot Generation of Clothing-Decoupled 3D Avatars Using 2D Diffusion

2023-07-13 · Shuo Huang, Zongxin Yang, Liangting Li, Yi Yang 외

Large-scale pre-trained vision-language models allow for the zero-shot text-based generation of 3D avatars. The previous state-of-the-art method utilized CLIP to supervise neural implicit models that reconstructed a huma…

InFoBERT: Zero-Shot Approach to Natural Language Understanding Using Contextualized Word Embedding

2021-09-01 · RANLP 2021 9 · Pavel Burnyshev, Andrey Bout, Valentin Malykh, Irina Piontkovskaya

Natural language understanding is an important task in modern dialogue systems. It becomes more important with the rapid extension of the dialogue systems’ functionality. In this work, we present an approach to zero-shot…

intent-classificationIntent ClassificationIntent Classification and Slot FillingNatural Language Understanding+3

AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars

2022-05-17 · Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai 외

3D avatar creation plays a crucial role in the digital age. However, the whole production process is prohibitively time-consuming and labor-intensive. To democratize this technology to a larger audience, we propose Avata…

3D geometryLanguage ModellingMotion SynthesisTexture Synthesis

Capture, Canonicalize, Splat: Zero-Shot 3D Gaussian Avatars from Unstructured Phone Images

2025-10-15 · Emanuel Garbin, Guy Adam, Oded Krams, Zohar Barzelay 외 arxiv

We present a novel, zero-shot pipeline for creating hyperrealistic, identity-preserving 3D avatars from a few unstructured phone images. Existing methods face several challenges: single-view approaches suffer from geomet…

Leveraging Temporal Contextualization for Video Action Recognition

2024-04-15 · Minji Kim, Dongyoon Han, Taekyung Kim, Bohyung Han

We propose a novel framework for video understanding, called Temporally Contextualized CLIP (TC-CLIP), which leverages essential temporal information through global interactions in a spatio-temporal domain within a video…

Action RecognitionTemporal Action LocalizationVideo UnderstandingZero-Shot Action Recognition