paper-with-me

Papers

MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4

2024-06-03 · Vahid Azizi, Fatemeh Koochaki

Vision-Language Models (VLMs) have recently seen significant advancements through integrating with Large Language Models (LLMs). The VLMs, which process image and text modalities simultaneously, have demonstrated the ability to learn and understand the interaction between images and texts across various multi-modal tasks. Reverse designing, which could be defined as a complex vision-language task, aims to predict the edits and their parameters, given a source image, an edited version, and an optional high-level textual edit description. This task requires VLMs to comprehend the interplay between the source image, the edited version, and the optional textual context simultaneously, going beyond traditional vision-language tasks. In this paper, we extend and fine-tune MiniGPT-4 for the reverse designing task. Our experiments demonstrate the extensibility of off-the-shelf VLMs, specifically MiniGPT-4, for more complex tasks such as reverse designing. Code is available at this \href{https://github.com/VahidAz/MiniGPT-Reverse-Designing}

📄 PDF Abstract BibTeX arXiv:2406.00971

Code (1)

vahidaz/minigpt-reverse-designing 공식 구현 pytorch

Similar Papers 제목 키워드 기반

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

2024-04-04 · Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman 외

This paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding. The model is capable of processing both temporal visual and textual data, making it adept at un…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+10

MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis

2024-07-04 · Asma Alkhaldi, Raneem Alnajim, Layan Alabdullatef, Rawan Alyahya 외

Recent advancements in artificial intelligence (AI) have precipitated significant breakthroughs in healthcare, particularly in refining diagnostic procedures. However, previous studies have often been constrained to limi…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model+4

MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection

2024-12-20 · Andrea Moglia, Elia Clement Nastasio, Luca Mainardi, Pietro Cerveri

Problem: Pancreas radiological imaging is challenging due to the small size, blurred boundaries, and variability of shape and position of the organ among patients. Goal: In this work we present MiniGPT-Pancreas, a Multim…

Cancer ClassificationChatbotLanguage ModelingLanguage Modelling+3

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

2023-04-20 · Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 외

The recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in pre…

Image DescriptionLanguage ModellingLarge Language ModelSpatial Reasoning+4

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

2023-10-14 · Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li 외

Large language models have shown their remarkable capabilities as a general interface for various language-related applications. Motivated by this, we target to build a unified interface for completing many vision-langua…

Image ClassificationImage DescriptionLanguage ModelingLanguage Modelling+8