paper-with-me

홈 › Papers

Enhancing Model Performance: Another Approach to Vision-Language Instruction Tuning

2024-07-25 · Vedanshu, MM Tripathi, Bhavnesh Jaint

The integration of large language models (LLMs) with vision-language (VL) tasks has been a transformative development in the realm of artificial intelligence, highlighting the potential of LLMs as a versatile general-purpose chatbot. However, the current trend in this evolution focuses on the integration of vision and language to create models that can operate in more diverse and real-world contexts. We present a novel approach, termed Bottleneck Adapter, specifically crafted for enhancing the multimodal functionalities of these complex models, enabling joint optimization of the entire multimodal LLM framework through a process known as Multimodal Model Tuning (MMT). Our approach utilizes lightweight adapters to connect the image encoder and LLM without the need for large, complex neural networks. Unlike the conventional modular training schemes, our approach adopts an end-to-end optimization regime, which, when combined with the adapters, facilitates the joint optimization using a significantly smaller parameter set. Our method exhibits robust performance with 90.12\% accuracy, outperforming both human-level performance (88.4\%) and LaVIN-7B (89.41\%).

📄 PDF Abstract BibTeX arXiv:2407.17813

Code (0)

등록된 구현이 없습니다.

Tasks

Chatbot

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations

2025-09-14 · Shresth Grover, Akshay Gopalkrishnan, Bo Ai, Henrik I. Christensen 외 arxiv

Vision-language-action (VLA) models finetuned from vision-language models (VLMs) hold the promise of leveraging rich pretrained representations to build generalist robots across diverse tasks and environments. However, d…

Robot ManipulationSpatial Reasoning

Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models

2023-11-21 · Xiaoyu Yang, Lijian Xu, Hao Sun, Hongsheng Li 외

Visual grounding (VG) occupies a pivotal position in multi-modality vision-language models. In this study, we propose ViLaM, a large multi-modality model, that supports multi-tasks of VG using the cycle training strategy…

Image SegmentationLanguage ModellingLarge Language ModelReferring Expression+5

Cross-Lingual Vision-Language Navigation

2019-10-24 · An Yan, Xin Eric Wang, Jiangtao Feng, Lei LI 외

Commanding a robot to navigate with natural language instructions is a long-term goal for grounded language understanding and robotics. But the dominant language is English, according to previous studies on vision-langua…

Domain AdaptationNavigateVision-Language NavigationZero-Shot Learning

Self-Review Framework for Enhancing Instruction Following Capability of LLM

2025-07-08 · Sihyun Park arxiv

Various techniques have been proposed to improve large language models (LLMs) adherence to formatting and instruction constraints. One of the most effective approaches involves utilizing high-quality data generated by po…

Instruction Following

Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning

2024-04-01 · CVPR 2024 1 · Rongjie Li, Yu Wu, Xuming He

Generative vision-language models (VLMs) have shown impressive performance in zero-shot vision-language tasks like image captioning and visual question answering. However, improving their zero-shot reasoning typically re…

Image CaptioningInstruction FollowingLanguage ModelingLanguage Modelling+4