paper-with-me

Papers

TeMO: Towards Text-Driven 3D Stylization for Multi-Object Meshes

2023-12-07 · CVPR 2024 1 · Xuying Zhang, Bo-Wen Yin, Yuming Chen, Zheng Lin, Yunheng Li, Qibin Hou, Ming-Ming Cheng

Recent progress in the text-driven 3D stylization of a single object has been considerably promoted by CLIP-based methods. However, the stylization of multi-object 3D scenes is still impeded in that the image-text pairs used for pre-training CLIP mostly consist of an object. Meanwhile, the local details of multiple objects may be susceptible to omission due to the existing supervision manner primarily relying on coarse-grained contrast of image-text pairs. To overcome these challenges, we present a novel framework, dubbed TeMO, to parse multi-object 3D scenes and edit their styles under the contrast supervision at multiple levels. We first propose a Decoupled Graph Attention (DGA) module to distinguishably reinforce the features of 3D surface points. Particularly, a cross-modal graph is constructed to align the object points accurately and noun phrases decoupled from the 3D mesh and textual description. Then, we develop a Cross-Grained Contrast (CGC) supervision system, where a fine-grained loss between the words in the textual description and the randomly rendered images are constructed to complement the coarse-grained loss. Extensive experiments show that our method can synthesize high-quality stylized content and outperform the existing methods over a wide range of multi-object 3D meshes. Our code and results will be made publicly available

📄 PDF Abstract BibTeX arXiv:2312.04248

Code (0)

등록된 구현이 없습니다.

Tasks

Graph AttentionObject

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

MOSAIC: Multi-Object Segmented Arbitrary Stylization Using CLIP

2023-09-24 · Prajwal Ganugula, Y S S S Santosh Kumar, N K Sagar Reddy, Prabhath Chellingi 외

Style transfer driven by text prompts paved a new path for creatively stylizing the images without collecting an actual style image. Despite having promising results, with text-driven stylization, the user has no control…

ObjectStyle Transfer

Text-Driven Stylization of Video Objects

2022-06-24 · Sebastian Loeschcke, Serge Belongie, Sagie Benaim

We tackle the task of stylizing video objects in an intuitive and semantic manner following a user-specified text prompt. This is a challenging task as the resulting video must satisfy multiple properties: (1) it has to …

Specificity

3DStyleGLIP: Part-Tailored Text-Guided 3D Neural Stylization

2024-04-03 · SeungJeh Chung, Joohyun Park, Hyeongyeop Kang

3D stylization, the application of specific styles to three-dimensional objects, offers substantial commercial potential by enabling the creation of uniquely styled 3D objects tailored to diverse scenes. Recent advanceme…

Neural Stylization

X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance

2023-03-28 · ICCV 2023 1 · Yiwei Ma, Xiaioqing Zhang, Xiaoshuai Sun, Jiayi Ji 외

Text-driven 3D stylization is a complex and crucial task in the fields of computer vision (CV) and computer graphics (CG), aimed at transforming a bare mesh to fit a target text. Prior methods adopt text-independent mult…

Attribute

Improved 3D Scene Stylization via Text-Guided Generative Image Editing with Region-Based Control

2025-09-04 · Haruo Fujiwara, Yusuke Mukuta, Tatsuya Harada arxiv

Recent advances in text-driven 3D scene editing and stylization, which leverage the powerful capabilities of 2D generative models, have demonstrated promising outcomes. However, challenges remain in ensuring high-quality…

Semantic correspondence3D scene EditingStyle TransferImage Editing