paper-with-me

Papers

Unifying 3D Vision-Language Understanding via Promptable Queries

2024-05-19 · Ziyu Zhu, Zhuofan Zhang, Xiaojian Ma, Xuesong Niu, Yixin Chen, Baoxiong Jia, Zhidong Deng, Siyuan Huang, Qing Li

A unified model for 3D vision-language (3D-VL) understanding is expected to take various scene representations and perform a wide range of tasks in a 3D scene. However, a considerable gap exists between existing methods and such a unified model, due to the independent application of representation and insufficient exploration of 3D multi-task training. In this paper, we introduce PQ3D, a unified model capable of using Promptable Queries to tackle a wide range of 3D-VL tasks, from low-level instance segmentation to high-level reasoning and planning. This is achieved through three key innovations: (1) unifying various 3D scene representations (i.e., voxels, point clouds, multi-view images) into a shared 3D coordinate space by segment-level grouping, (2) an attention-based query decoder for task-specific information retrieval guided by prompts, and (3) universal output heads for different tasks to support multi-task training. Tested across ten diverse 3D-VL datasets, PQ3D demonstrates impressive performance on these tasks, setting new records on most benchmarks. Particularly, PQ3D improves the state-of-the-art on ScanNet200 by 4.9% (AP25), ScanRefer by 5.4% (acc@0.5), Multi3DRefer by 11.7% (F1@0.5), and Scan2Cap by 13.4% (CIDEr@0.5). Moreover, PQ3D supports flexible inference with individual or combined forms of available 3D representations, e.g., solely voxel input.

📄 PDF Abstract BibTeX arXiv:2405.11442

Code (0)

등록된 구현이 없습니다.

Tasks

3D Question Answering (3D-QA)DecoderInformation RetrievalInstance SegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting

2025-06-29 · Yiming Huang, Long Bai, Beilei Cui, Kun Yuan 외

In contemporary surgical research and practice, accurately comprehending 3D surgical scenes with text-promptable capabilities is particularly crucial for surgical planning and real-time intra-operative guidance, where pr…

3D ReconstructionScene Understanding

Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective

2024-10-29 · Shenghao Xie, Wenqiang Zu, Mingyang Zhao, Duo Su 외

Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing interest in extending this success to v…

Survey

Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval

2025-05-21 · Siting Li, Xiang Gao, Simon Shaolei Du

While an image is worth more than a thousand words, only a few provide crucial information for a given task and thus should be focused on. In light of this, ideal text-to-image (T2I) retrievers should prioritize specific…

AttributeImage RetrievalLarge Language ModelMultimodal Large Language Model+1

Towards Flexible Visual Relationship Segmentation

2024-08-15 · Fangrui Zhu, Jianwei Yang, Huaizu Jiang

Visual relationship understanding has been studied separately in human-object interaction(HOI) detection, scene graph generation(SGG), and referring relationships(RR) tasks. Given the complexity and interconnectedness of…

Graph GenerationHuman-Object Interaction DetectionScene Graph GenerationSegmentation

Text Promptable Surgical Instrument Segmentation with Vision-Language Models

2023-06-15 · NeurIPS 2023 11

In this paper, we propose a novel text promptable surgical instrument segmentation approach to overcome challenges associated with diversity and differentiation of surgical instruments in minimally invasive surgeries. We…

DecoderDiversitySegmentation