paper-with-me

홈 › Papers

Compositional Semantics for Open Vocabulary Spatio-semantic Representations

2023-10-08 · Robin Karlsson, Francisco Lepe-Salazar, Kazuya Takeda

General-purpose mobile robots need to complete tasks without exact human instructions. Large language models (LLMs) is a promising direction for realizing commonsense world knowledge and reasoning-based planning. Vision-language models (VLMs) transform environment percepts into vision-language semantics interpretable by LLMs. However, completing complex tasks often requires reasoning about information beyond what is currently perceived. We propose latent compositional semantic embeddings z* as a principled learning-based knowledge representation for queryable spatio-semantic memories. We mathematically prove that z* can always be found, and the optimal z* is the centroid for any set Z. We derive a probabilistic bound for estimating separability of related and unrelated semantics. We prove that z* is discoverable by iterative optimization by gradient descent from visual appearance and singular descriptions. We experimentally verify our findings on four embedding spaces incl. CLIP and SBERT. Our results show that z* can represent up to 10 semantics encoded by SBERT, and up to 100 semantics for ideal uniformly distributed high-dimensional embeddings. We demonstrate that a simple dense VLM trained on the COCO-Stuff dataset can learn z* for 181 overlapping semantics by 42.23 mIoU, while improving conventional non-overlapping open-vocabulary segmentation performance by +3.48 mIoU compared with a popular SOTA model.

📄 PDF Abstract BibTeX arXiv:2310.04981

Code (0)

등록된 구현이 없습니다.

Tasks

World Knowledge

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
SBERT 설명 없음

Similar Papers 제목 키워드 기반

SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation

2026-08-24 · Shibo Zhao, Guofei Chen, Honghao Zhu, Zhiheng Li 외 arxiv

Robotic navigation in human environments requires a spatio-temporal semantic representation that can rec- oncile open-vocabulary perception with long-term environmental changes. While foundation models provide strong zer…

Learning a Compositional Semantics for Freebase with an Open Predicate Vocabulary

2015-01-01 · TACL 2015 1 · Jayant Krishnamurthy, Tom M. Mitchell

We present an approach to learning a model-theoretic semantics for natural language tied to Freebase. Crucially, our approach uses an open predicate vocabulary, enabling it to produce denotations for phrases such as {``}…

Coreference ResolutionOpen Information ExtractionQuestion AnsweringSemantic Parsing+1

Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection

2023-02-01 · Kaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao 외

Prompt tuning with large-scale pretrained vision-language models empowers open-vocabulary predictions trained on limited base categories, e.g., object classification and detection. In this paper, we propose compositional…

ObjectRelationVideo Visual Relation Detection

Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation

2026-05-15 · Chenhao Wang, Yingrui Ji, Yu Meng, Yao Zhu arxiv

Open-vocabulary segmentation models often struggle to generalize to unseen combinations of object categories and attributes, because fine-grained descriptions are typically encoded as holistic sentences that entangle mul…

Structure-aware Prompt Adaptation from Seen to Unseen for Open-Vocabulary Compositional Zero-Shot Learning

2026-03-04 · Yihang Duan, Jiong Wang, Pengpeng Zeng, Ji Zhang 외 arxiv

The goal of Open-Vocabulary Compositional Zero-Shot Learning (OV-CZSL) is to recognize attribute-object compositions in the open-vocabulary setting, where compositions of both seen and unseen attributes and objects are e…

Compositional Zero-Shot Learning