paper-with-me

홈 › Papers

Expanding Frozen Vision-Language Models without Retraining: Towards Improved Robot Perception

2023-08-31 · Riley Tavassoli, Mani Amani, Reza Akhavian

Vision-language models (VLMs) have shown powerful capabilities in visual question answering and reasoning tasks by combining visual representations with the abstract skill set large language models (LLMs) learn during pretraining. Vision, while the most popular modality to augment LLMs with, is only one representation of a scene. In human-robot interaction scenarios, robot perception requires accurate scene understanding by the robot. In this paper, we define and demonstrate a method of aligning the embedding spaces of different modalities (in this case, inertial measurement unit (IMU) data) to the vision embedding space through a combination of supervised and contrastive training, enabling the VLM to understand and reason about these additional modalities without retraining. We opt to give the model IMU embeddings directly over using a separate human activity recognition model that feeds directly into the prompt to allow for any nonlinear interactions between the query, image, and IMU signal that would be lost by mapping the IMU data to a discrete activity label. Further, we demonstrate our methodology's efficacy through experiments involving human activity recognition using IMU data and visual inputs. Our results show that using multiple modalities as input improves the VLM's scene understanding and enhances its overall performance in various tasks, thus paving the way for more versatile and capable language models in multi-modal contexts.

📄 PDF Abstract BibTeX arXiv:2308.16493

Code (0)

등록된 구현이 없습니다.

Tasks

Activity RecognitionHuman Activity RecognitionQuestion AnsweringScene UnderstandingVisual Question Answering

Methods 이 논문이 사용한 방법론

OPT OPT is a suite of decoder-only pre-trained transformers ranging from 125M to 175B parameters. The model uses an AdamW optimizer and weight decay of 0.1. It follows a linear…

Similar Papers 제목 키워드 기반

LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives

2026-07-01 · Lukas Kuhn, Giuseppe Serra, Randall Balestriero, Florian Buettner arxiv

Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods. At the same time, the role of vision-language encoders ha…

Self-Supervised LearningSemantic Segmentation

Semantically Grounded QFormer for Efficient Vision Language Understanding

2023-11-13 · Moulik Choraria, Xinbo Wu, Sourya Basu, Nitesh Sekhar 외

General purpose Vision Language Models (VLMs) have received tremendous interest in recent years, owing to their ability to learn rich vision-language correlations as well as their broad zero-shot competencies. One immens…

DiversityImage to textRepresentation LearningText Generation

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

2025-12-10 · Minghao Guo, Meng Cao, Jiachen Tao, Rongtao Xu 외 arxiv

Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we introduce GLaD, a geometry-aware VLA fra…

Knowledge DistillationSpatial Reasoning

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies

2026-07-02 · Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang arxiv

Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity for test-time guidance without retraining the base policy. We study this opportun…

MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based Finetuning

2021-12-09 · Constantin Eichenberg, Sidney Black, Samuel Weinbach, Letitia Parcalabescu 외

Large-scale pretraining is fast becoming the norm in Vision-Language (VL) modeling. However, prevailing VL approaches are limited by the requirement for labeled data and the use of complex multi-step pretraining objectiv…

In-Context LearningLanguage ModelingLanguage Modelling