Semantically Grounded QFormer for Efficient Vision Language Understanding
General purpose Vision Language Models (VLMs) have received tremendous interest in recent years, owing to their ability to learn rich vision-language correlations as well as their broad zero-shot competencies. One immensely popular line of work utilizes frozen unimodal models, by bridging vision representations to language using a trainable module called the QFormer. However, this method relies heavily on large-scale multimodal pretraining with huge computational overheads. To that end, we propose a more efficient framework for QFormer-based vision-language alignment. Our key idea relies on the observation that QFormer latents correspond more strongly to the frozen LLM's intermediate latent space. Consequently, instead of using QFormer latents as inputs to the LLM, we alter the framework by using the latents to directly condition the LLM latent space for image-to-text generation. We demonstrate the effectiveness of our approach against existing baselines in improving the efficiency of vision-language pretraining.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityImage to textRepresentation LearningText GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models
Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal represen…
ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of mu…
SeqFormer: Sequential Transformer for Video Instance Segmentation
In this work, we present SeqFormer for video instance segmentation. SeqFormer follows the principle of vision transformer that models instance relationships among video frames. Nevertheless, we observe that a stand-alone…
Instance SegmentationSemantic SegmentationVideo Instance SegmentationImproving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
Recent advances in large language models (LLMs) and development of audio codecs greatly propel the zero-shot TTS. They can synthesize personalized speech with only a 3-second speech of an unseen speaker as acoustic promp…
Language ModelingLanguage ModellingLarge Language ModelSpeech Synthesis+3Vision Transformer with Quadrangle Attention
Window-based attention has become a popular choice in vision transformers due to its superior performance, lower computational complexity, and less memory footprint. However, the design of hand-crafted windows, which is …
object-detectionObject DetectionPose EstimationSemantic Segmentation