paper-with-me

홈 › Papers

Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks

2024-02-13 · Jusung Lee, Sungguk Cha, Younghyun Lee, Cheoljong Yang

Having revolutionized natural language processing (NLP) applications, large language models (LLMs) are expanding into the realm of multimodal inputs. Owing to their ability to interpret images, multimodal LLMs (MLLMs) have been primarily used for vision-language tasks. Currently, MLLMs have not yet been extended for domain-specific visual tasks, which require a more explicit understanding of visual information. We developed a method to transform domain-specific visual and vision-language datasets into a unified question answering format called Visual Question Answering Instruction (VQA-IN), thereby extending MLLM to domain-specific tasks. The VQA-IN was applied to train multiple MLLM architectures using smaller versions of LLMs (sLLMs). The experimental results indicated that the proposed method achieved a high score metric on domainspecific visual tasks while also maintaining its performance on vision-language tasks in a multitask manner.

📄 PDF Abstract BibTeX arXiv:2402.08360

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

2025-01-09 · Jiaxing Zhao, Boyuan Sun, Xiang Chen, Xihan Wei 외

In this paper, we introduce LLaVA-Octopus, a novel video multimodal large language model. LLaVA-Octopus adaptively weights features from different visual projectors based on user instructions, enabling us to leverage the…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization

2024-12-21 · Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen, Chris Ngo 외

Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness …

Image CaptioningMultimodal ReasoningObject LocalizationObject Recognition+2

MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output

2025-10-11 · Yanyuan Chen, Dexuan Xu, Yu Huang, Songkun Zhan 외 arxiv

Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and lac…

Instruction FollowingQuestion Answering

MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal Output

2025-01-01 · CVPR 2025 1 · Yanyuan Chen, Dexuan Xu, Yu Huang, Songkun Zhan 외

Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and…

Instruction FollowingLanguage ModelingLanguage ModellingQuestion Answering

CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios

2024-03-07 · Qilang Ye, Zitong Yu, Rui Shao, Xinyu Xie 외

This paper focuses on the challenge of answering questions in scenarios that are composed of rich and complex dynamic audio-visual components. Although existing Multimodal Large Language Models (MLLMs) can respond to aud…

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Language ModelingLanguage Modelling+6