paper-with-me

Papers

PaLM2-VAdapter: Progressively Aligned Language Model Makes a Strong Vision-language Adapter

2024-02-16 · Junfei Xiao, Zheng Xu, Alan Yuille, Shen Yan, Boyu Wang

This paper demonstrates that a progressively aligned language model can effectively bridge frozen vision encoders and large language models (LLMs). While the fundamental architecture and pre-training methods of vision encoders and LLMs have been extensively studied, the architecture and training strategy of vision-language adapters vary significantly across recent works. Our research undertakes a thorough exploration of the state-of-the-art perceiver resampler architecture and builds a strong baseline. However, we observe that the vision-language alignment with perceiver resampler exhibits slow convergence and limited scalability with a lack of direct supervision. To address this issue, we propose PaLM2-VAdapter, employing a progressively aligned language model as the vision-language adapter. Compared to the strong baseline with perceiver resampler, our method empirically shows faster convergence, higher performance, and stronger scalability. Extensive experiments across various Visual Question Answering (VQA) and captioning tasks on both images and videos demonstrate that our model exhibits state-of-the-art visual understanding and multi-modal reasoning capabilities. Notably, our method achieves these advancements with 30~70% fewer parameters than the state-of-the-art large vision-language models, marking a significant efficiency improvement.

📄 PDF Abstract BibTeX arXiv:2402.10896

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

PopALM: Popularity-Aligned Language Models for Social Media Trendy Response Prediction

2024-02-29 · Erxin Yu, Jing Li, Chunpu Xu

Social media platforms are daily exhibiting millions of events. To preliminarily predict the mainstream public reaction to these events, we study trendy response prediction to automatically generate top-liked user replie…

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment

2026-02-28 · Yantao Li, Qiang Hui, Chenyang Yan, Kanzhi Cheng 외 arxiv

Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucin…

Reinforcement LearningMultimodal ReasoningVisual Reasoning

Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech Understanding

2023-03-02 · Yingting Li, Ambuj Mehrish, Shuai Zhao, Rishabh Bhardwaj 외

Fine-tuning is widely used as the default algorithm for transfer learning from pre-trained models. Parameter inefficiency can however arise when, during transfer learning, all the parameters of a large pre-trained model …

Speech Synthesistext-to-speechText to SpeechTransfer Learning

Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion Models

2025-03-24 · CVPR 2025 1 · Jianlong Jin, Chenglong Zhao, Ruixin Zhang, Sheng Shang 외

Palmprint recognition is significantly limited by the lack of large-scale publicly available datasets. Previous methods have adopted B\'ezier curves to simulate the palm creases, which then serve as input for conditional…

Palmprint image registration using convolutional neural networks and Hough transform

2019-04-01 · Mohsen Ahmadi, Hossein Soleimani

Minutia-based palmprint recognition systems has got lots of interest in last two decades. Due to the large number of minutiae in a palmprint, approximately 1000 minutiae, the matching process is time consuming which make…

Image Registration