paper-with-me

홈 › Papers

QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries

2025-02-26 · Nicolas Harvey Chapman, Feras Dayoub, Will Browne, Christopher Lehnert

A domain shift exists between the large-scale, internet data used to train a Vision-Language Model (VLM) and the raw image streams collected by a robot. Existing adaptation strategies require the definition of a closed-set of classes, which is impractical for a robot that must respond to diverse natural language queries. In response, we present QueryAdapter; a novel framework for rapidly adapting a pre-trained VLM in response to a natural language query. QueryAdapter leverages unlabelled data collected during previous deployments to align VLM features with semantic classes related to the query. By optimising learnable prompt tokens and actively selecting objects for training, an adapted model can be produced in a matter of minutes. We also explore how objects unrelated to the query should be dealt with when using real-world data for adaptation. In turn, we propose the use of object captions as negative class labels, helping to produce better calibrated confidence scores during adaptation. Extensive experiments on ScanNet++ demonstrate that QueryAdapter significantly enhances object retrieval performance compared to state-of-the-art unsupervised VLM adapters and 3D scene graph methods. Furthermore, the approach exhibits robust generalization to abstract affordance queries and other datasets, such as Ego4D.

📄 PDF Abstract BibTeX arXiv:2502.18735

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Queries

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Automatic Adaptation Rule Optimization via Large Language Models

2024-07-02 · Yusei Ishimizu, Jialong Li, Jinglue Xu, Jinyu Cai 외

Rule-based adaptation is a foundational approach to self-adaptation, characterized by its human readability and rapid response. However, building high-performance and robust adaptation rules is often a challenge because …

Common Sense Reasoning

DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation

2026-01-29 · Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen 외 arxiv

Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which, despite strong generalization in static manipulation, struggle in dynamic scenarios requiring rapid perception, tempo…

Continuous Control

Rapid Adaptation with Conditionally Shifted Neurons

2017-12-28 · ICML 2018 7 · Tsendsuren Munkhdalai, Xingdi Yuan, Soroush Mehri, Adam Trischler

We describe a mechanism by which artificial neural networks can learn rapid adaptation - the ability to adapt on the fly, with little data, to new tasks - that we call conditionally shifted neurons. We apply this mechani…

Few-Shot Image Classification

VLN-Zero: Rapid Exploration and Cache-Enabled Neurosymbolic Vision-Language Planning for Zero-Shot Transfer in Robot Navigation

2025-09-23 · Neel P. Bhatt, Yunhao Yang, Rohan Siva, Pranay Samineni 외 arxiv

Rapid adaptation in unseen environments is essential for scalable real-world autonomy, yet existing approaches rely on exhaustive exploration or rigid navigation policies that fail to generalize. We present VLN-Zero, a t…

Vision-Language NavigationRobot Navigation

Adapting Vision-Language Models from Iconic to Inclusive for Multi-Label Recognition Without Labels

2026-06-10 · Cheng Chen, Jingyu Zhou, Yifan Zhao, Jia Li arxiv

Understanding multi-label images remains a challenging task in computer vision. With the rapid progress of vision-language multimodal learning, vision-language models (VLMs) enable zero-shot recognition without labeled d…

Multi-Label Learning