paper-with-me

홈 › Papers

Semantically Grounded QFormer for Efficient Vision Language Understanding

2023-11-13 · Moulik Choraria, Xinbo Wu, Sourya Basu, Nitesh Sekhar, Yue Wu, Xu Zhang, Prateek Singhal, Lav R. Varshney

General purpose Vision Language Models (VLMs) have received tremendous interest in recent years, owing to their ability to learn rich vision-language correlations as well as their broad zero-shot competencies. One immensely popular line of work utilizes frozen unimodal models, by bridging vision representations to language using a trainable module called the QFormer. However, this method relies heavily on large-scale multimodal pretraining with huge computational overheads. To that end, we propose a more efficient framework for QFormer-based vision-language alignment. Our key idea relies on the observation that QFormer latents correspond more strongly to the frozen LLM's intermediate latent space. Consequently, instead of using QFormer latents as inputs to the LLM, we alter the framework by using the latents to directly condition the LLM latent space for image-to-text generation. We demonstrate the effectiveness of our approach against existing baselines in improving the efficiency of vision-language pretraining.

📄 PDF Abstract BibTeX arXiv:2311.07449

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage to textRepresentation LearningText Generation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

2026-07-16 · Yufeng Ji, Wenhao Tang, Haoyi Niu, Koushil Sreenath 외 arxiv

Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal represen…

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

2025-05-29 · Yiming Lei, Zhizheng Yang, Zeming Liu, Haitao Leng 외

Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of mu…

SeqFormer: Sequential Transformer for Video Instance Segmentation

2021-12-15 · Junfeng Wu, Yi Jiang, Song Bai, Wenqing Zhang 외

In this work, we present SeqFormer for video instance segmentation. SeqFormer follows the principle of vision transformer that models instance relationships among video frames. Nevertheless, we observe that a stand-alone…

Instance SegmentationSemantic SegmentationVideo Instance Segmentation

Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model

2024-06-06 · Jinlong Xue, Yayue Deng, Yicheng Han, Yingming Gao 외

Recent advances in large language models (LLMs) and development of audio codecs greatly propel the zero-shot TTS. They can synthesize personalized speech with only a 3-second speech of an unseen speaker as acoustic promp…

Language ModelingLanguage ModellingLarge Language ModelSpeech Synthesis+3

Vision Transformer with Quadrangle Attention

2023-03-27 · Qiming Zhang, Jing Zhang, Yufei Xu, DaCheng Tao

Window-based attention has become a popular choice in vision transformers due to its superior performance, lower computational complexity, and less memory footprint. However, the design of hand-crafted windows, which is …

object-detectionObject DetectionPose EstimationSemantic Segmentation