paper-with-me

Papers

Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners

2024-04-30 · CVPR 2024 1 · Chun Feng, Joy Hsu, Weiyu Liu, Jiajun Wu

3D visual grounding is a challenging task that often requires direct and dense supervision, notably the semantic label for each object in the scene. In this paper, we instead study the naturally supervised setting that learns from only 3D scene and QA pairs, where prior works underperform. We propose the Language-Regularized Concept Learner (LARC), which uses constraints from language as regularization to significantly improve the accuracy of neuro-symbolic concept learners in the naturally supervised setting. Our approach is based on two core insights: the first is that language constraints (e.g., a word's relation to another) can serve as effective regularization for structured representations in neuro-symbolic models; the second is that we can query large language models to distill such constraints from language properties. We show that LARC improves performance of prior works in naturally supervised 3D visual grounding, and demonstrates a wide range of 3D visual reasoning capabilities-from zero-shot composition, to data efficiency and transferability. Our method represents a promising step towards regularizing structured visual reasoning frameworks with language-based priors, for learning in settings without dense supervision.

📄 PDF Abstract BibTeX arXiv:2404.19696

Code (0)

등록된 구현이 없습니다.

Tasks

3D visual groundingVisual GroundingVisual Reasoning

Similar Papers 제목 키워드 기반

Weakly-Supervised 3D Visual Grounding based on Visual Linguistic Alignment

2023-12-15 · Xiaoxu Xu, Yitian Yuan, Qiudan Zhang, Wenhui Wu 외

Learning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual grounding approaches require a substantial numbe…

3D visual groundingNatural Language QueriesScene UnderstandingVisual Grounding

Weakly Supervised Attention Learning for Textual Phrases Grounding

2018-05-01 · Zhiyuan Fang, Shu Kong, Tianshu Yu, Yezhou Yang

Grounding textual phrases in visual content is a meaningful yet challenging problem with various potential applications such as image-text inference or text-driven multimedia interaction. Most of the current existing met…

CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding

2023-05-15 · Linhui Xiao, Xiaoshan Yang, Fang Peng, Ming Yan 외

Visual Grounding (VG) is a crucial topic in the field of vision and language, which involves locating a specific region described by expressions within an image. To reduce the reliance on manually labeled data, unsupervi…

DiversityTransfer LearningVisual Grounding

Sim-To-Real Transfer of Visual Grounding for Human-Aided Ambiguity Resolution

2022-05-24 · Georgios Tziafas, Hamidreza Kasaei

Service robots should be able to interact naturally with non-expert human users, not only to help them in various tasks but also to receive guidance in order to resolve ambiguities that might be present in the instructio…

Domain AdaptationVisual Grounding

Pseudo-Q: Generating Pseudo Language Queries for Visual Grounding

2022-03-16 · CVPR 2022 1 · Haojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song 외

Visual grounding, i.e., localizing objects in images according to natural language queries, is an important topic in visual language understanding. The most effective approaches for this task are based on deep learning, …

Language ModellingNatural Language QueriesVisual Grounding