paper-with-me

Papers

CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes

2024-10-19 · Xianghao Xu, Srinath Sridhar, Daniel Ritchie

We propose a zero-shot text-driven 3D shape deformation system that deforms an input 3D mesh of a manufactured object to fit an input text description. To do this, our system optimizes the parameters of a deformation model to maximize an objective function based on the widely used pre-trained vision language model CLIP. We find that CLIP-based objective functions exhibit many spurious local optima; to circumvent them, we parameterize deformations using a novel deformation model called BoxDefGraph which our system automatically computes from an input mesh, the BoxDefGraph is designed to capture the object aligned rectangular/circular geometry features of most manufactured objects. We then use the CMA-ES global optimization algorithm to maximize our objective, which we find to work better than popular gradient-based optimizers. We demonstrate that our approach produces appealing results and outperforms several baselines.

📄 PDF Abstract BibTeX arXiv:2410.15199

Code (0)

등록된 구현이 없습니다.

Tasks

global-optimizationLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Zero Shot Deformation Reconstruction for Soft Robots Using a Flexible Sensor Array and Cage Based 3D Gaussian Modeling

2026-03-20 · Linrui Shou, Zilang Chen, Wenjia Xu, Yiyue Luo 외 arxiv

We present a zero-shot deformation reconstruction framework for soft robots that operates without any visual supervision at inference time. In this work, zero-shot deformation reconstruction is defined as the ability to …

Zero-shot Generalization

Segment Anything with Robust Uncertainty-Accuracy Correlation

2026-05-11 · Hongyou Zhou, Marc Toussaint, Ling Shao, Zihan Ye arxiv

Despite strong zero-shot performance, SAM is unreliable under domain shift due to Mask-level Confidence Confusion (MCC), where a single IoU-based mask score fails to reflect pixel-wise reliability near boundaries. Motiva…

Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator

2024-11-23 · CVPR 2025 1 · Chaehun Shin, Jooyoung Choi, Heeseung Kim, Sungroh Yoon

Subject-driven text-to-image generation aims to produce images of a new subject within a desired context by accurately capturing both the visual characteristics of the subject and the semantic content of a text prompt. T…

Image GenerationText to Image GenerationText-to-Image Generation

MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech

2024-10-04 · Taejun Bak, Youngsik Eom, SeungJae Choi, Young-Sun Joo

Text-to-speech (TTS) systems that scale up the amount of training data have achieved significant improvements in zero-shot speech synthesis. However, these systems have certain limitations: they require a large amount of…

DisentanglementSpeech SynthesisStyle Transfertext-to-speech+1

Diagnostic Benchmarks for Invariant Learning Dynamics: Empirical Validation of the Eidos Architecture

2026-02-10 · Datorien L. Anderson arxiv

We present the PolyShapes-Ideal (PSI) dataset, a suite of diagnostic benchmarks designed to isolate topological invariance -- the ability to maintain structural identity across affine transformations -- from the textural…