paper-with-me

홈 › Papers

Vista3D: Unravel the 3D Darkside of a Single Image

2024-09-18 · Qiuhong Shen, Xingyi Yang, Michael Bi Mi, Xinchao Wang

We embark on the age-old quest: unveiling the hidden dimensions of objects from mere glimpses of their visible parts. To address this, we present Vista3D, a framework that realizes swift and consistent 3D generation within a mere 5 minutes. At the heart of Vista3D lies a two-phase approach: the coarse phase and the fine phase. In the coarse phase, we rapidly generate initial geometry with Gaussian Splatting from a single image. In the fine phase, we extract a Signed Distance Function (SDF) directly from learned Gaussian Splatting, optimizing it with a differentiable isosurface representation. Furthermore, it elevates the quality of generation by using a disentangled representation with two independent implicit functions to capture both visible and obscured aspects of objects. Additionally, it harmonizes gradients from 2D diffusion prior with 3D-aware diffusion priors by angular diffusion prior composition. Through extensive evaluation, we demonstrate that Vista3D effectively sustains a balance between the consistency and diversity of the generated 3D objects. Demos and code will be available at https://github.com/florinshen/Vista3D.

📄 PDF Abstract BibTeX arXiv:2409.12193

Code (1)

florinshen/vista3d 공식 구현

Tasks

3D GenerationDiversity

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

VistaDream: Sampling multiview consistent images for single-view scene reconstruction

2024-10-22 · Haiping Wang, YuAn Liu, Ziwei Liu, Wenping Wang 외

In this paper, we propose VistaDream a novel framework to reconstruct a 3D scene from a single-view image. Recent diffusion models enable generating high-quality novel-view images from a single-view input image. Most exi…

Novel View Synthesis

MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement Learning

2026-01-12 · Meng Lu, Yuxing Lu, Yuchen Zhuang, Megan Mullins 외 arxiv

Vision language models (VLMs) achieve strong performance on general image understanding but struggle to think with medical images, especially when performing multi-step reasoning through iterative visual interaction. Med…

Reinforcement LearningMultimodal ReasoningVisual Reasoning

VistaHop: Benchmarking Long-Horizon Visual DeepSearch

2026-06-02 · Hang He, Chuhuai Yue, Chengqi Dong, Chengcheng Wan 외 arxiv

Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues…

Question AnsweringVisual GroundingVisual ReasoningImage Cropping

Jack of All Tasks Master of Many: Designing General-Purpose Coarse-to-Fine Vision-Language Model

2024-01-01 · CVPR 2024 1 · Shraman Pramanick, Guangxing Han, Rui Hou, Sayan Nag 외

The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems unifying various vision-language (VL) tasks by instruction tuning. However due to the enormous div…

AllAttributeLanguage ModelingLanguage Modelling+1

Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model

2023-12-19 · Shraman Pramanick, Guangxing Han, Rui Hou, Sayan Nag 외

The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems, unifying various vision-language (VL) tasks by instruction tuning. However, due to the enormous diver…

AllAttributeLanguage ModelingLanguage Modelling+1