paper-with-me

홈 › Papers

TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing

2024-05-23 · Teng Xu, Jiamin Chen, Peng Chen, Youjia Zhang, Junqing Yu, Wei Yang

Editing objects within a scene is a critical functionality required across a broad spectrum of applications in computer vision and graphics. As 3D Gaussian Splatting (3DGS) emerges as a frontier in scene representation, the effective modification of 3D Gaussian scenes has become increasingly vital. This process entails accurately retrieve the target objects and subsequently performing modifications based on instructions. Though available in pieces, existing techniques mainly embed sparse semantics into Gaussians for retrieval, and rely on an iterative dataset update paradigm for editing, leading to over-smoothing or inconsistency issues. To this end, this paper proposes a systematic approach, namely TIGER, for coherent text-instructed 3D Gaussian retrieval and editing. In contrast to the top-down language grounding approach for 3D Gaussians, we adopt a bottom-up language aggregation strategy to generate a denser language embedded 3D Gaussians that supports open-vocabulary retrieval. To overcome the over-smoothing and inconsistency issues in editing, we propose a Coherent Score Distillation (CSD) that aggregates a 2D image editing diffusion model and a multi-view diffusion model for score distillation, producing multi-view consistent editing with much finer details. In various experiments, we demonstrate that our TIGER is able to accomplish more consistent and realistic edits than prior work.

📄 PDF Abstract BibTeX arXiv:2405.14455

Code (0)

등록된 구현이 없습니다.

Tasks

3DGSRetrieval

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models

2024-06-09 · Leigang Qu, Haochuan Li, Tan Wang, Wenjie Wang 외

How humans can effectively and efficiently acquire images has always been a perennial question. A classic solution is text-to-image retrieval from an existing database; however, the limited database typically lacks creat…

counterfactualImage GenerationImage RetrievalRetrieval+2

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval

2026-05-18 · Xinyu Sun, Huangyu Dai, Lingtao Mao, Zexin Zheng 외 arxiv

E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This image-to-multimodal retrieval setting presents two asymmetries: a modali…

Object Detection

TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval

2026-05-23 · Yuhang Zhang, Keyan Ding, Peilin Chen, Han Liu 외 arxiv

Enzyme-reaction retrieval is a fundamental problem in computational biology, underpinning enzyme characterization, reaction mechanism elucidation, and the rational design of metabolic pathways and biocatalysts. As a bidi…

Text Generation

TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval

2026-03-25 · David G. Shatwell, Sirnam Swetha, Mubarak Shah arxiv

Many real-world applications in digital forensics, urban monitoring, and environmental analysis require jointly reasoning about visual appearance, location, and time. Beyond standard geo-localization and time-of-capture …

Image Retrieval

Question-Aware Gaussian Experts for Audio-Visual Question Answering

2025-03-06 · CVPR 2025 1 · Hongyeob Kim, Inyoung Jung, Dayoon Suh, Youjia Zhang 외

Audio-Visual Question Answering (AVQA) requires not only question-based multimodal reasoning but also precise temporal grounding to capture subtle dynamics for accurate prediction. However, existing methods mainly use qu…

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)AUDIO-VISUAL QUESTION ANSWERING (MUSIC-AVQA-v2.0)Mixture-of-Experts+3