Vector Quantized Feature Fields for Fast 3D Semantic Lifting
We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding multiscale pixel-aligned feature maps, which are derived from scene representations such as distilled feature fields and feature point clouds. However, storing per-view feature maps rendered from distilled feature fields is impractical, and feature point clouds are expensive to store and query. To enable lightweight on-demand retrieval of pixel-aligned relevance masks, we introduce the Vector-Quantized Feature Field. We demonstrate the effectiveness of the Vector-Quantized Feature Field on complex indoor and outdoor scenes. Semantic lifting, when paired with a Vector-Quantized Feature Field, can unlock a myriad of applications in scene representation and embodied intelligence. Specifically, we showcase how our method enables text-driven localized scene editing and significantly improves the efficiency of embodied question answering.
Code (0)
등록된 구현이 없습니다.
Tasks
Embodied Question AnsweringQuestion AnsweringSimilar Papers 제목 키워드 기반
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
We introduce Purrception, a variational flow matching approach for vector-quantized image generation that provides explicit categorical supervision while maintaining continuous transport dynamics. Our method adapts Varia…
Image GenerationVector Quantized Semantic Communication System
Although analog semantic communication systems have received considerable attention in the literature, there is less work on digital semantic communication systems. In this paper, we develop a deep learning (DL)-enabled …
MS-SSIMQuantizationSemantic CommunicationSSIMVQ-DSC-R: Robust Vector Quantized-Enabled Digital Semantic Communication With OFDM Transmission
Digital mapping of semantic features is essential for achieving interoperability between semantic communication and practical digital infrastructure. However, current research efforts predominantly concentrate on analog …
Semantic CommunicationPUREVQ-GAN: Defending Data Poisoning Attacks through Vector-Quantized Bottlenecks
We introduce PureVQ-GAN, a defense against data poisoning that forces backdoor triggers through a discrete bottleneck using Vector-Quantized VAE with GAN discriminator. By quantizing poisoned images through a learned cod…
Efficient High-Resolution Template Matching with Vector Quantized Nearest Neighbour Fields
Template matching is a fundamental problem in computer vision with applications in fields including object detection, image registration, and object tracking. Current methods rely on nearest-neighbour (NN) matching, wher…
Image Registrationobject-detectionObject DetectionObject Tracking+2