paper-with-me

홈 › Papers

Multi-modal Visual Understanding with Prompts for Semantic Information Disentanglement of Image

2023-05-16 · Yuzhou Peng

Multi-modal visual understanding of images with prompts involves using various visual and textual cues to enhance the semantic understanding of images. This approach combines both vision and language processing to generate more accurate predictions and recognition of images. By utilizing prompt-based techniques, models can learn to focus on certain features of an image to extract useful information for downstream tasks. Additionally, multi-modal understanding can improve upon single modality models by providing more robust representations of images. Overall, the combination of visual and textual information is a promising area of research for advancing image recognition and understanding. In this paper we will try an amount of prompt design methods and propose a new method for better extraction of semantic information

📄 PDF Abstract BibTeX arXiv:2305.09333

Code (0)

등록된 구현이 없습니다.

Tasks

Disentanglement

Similar Papers 제목 키워드 기반

Knowledge Transfer with Visual Prompt in multi-modal Dialogue Understanding and Generation

2022-10-01 · TU (COLING) 2022 10 · Minjun Zhu, Yixuan Weng, Bin Li, Shizhu He 외

Visual Dialogue (VD) task has recently received increasing attention in AI research. Visual Dialog aims to generate multi-round, interactive responses based on the dialog history and image content. Existing textual dialo…

Dialogue UnderstandingKnowledge DistillationTransfer LearningVisual Dialog

Adversarial Prompt Injection Attack on Multimodal Large Language Models

2026-03-31 · Meiwen Ding, Song Xia, Chenqi Kong, Xudong Jiang arxiv

Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt injection attacks. Existing prompt injection m…

GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

2024-06-14 · Yiqi Wu, Xiaodan Hu, Ziming Fu, Siling Zhou 외

Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a …

Activity RecognitionMMR totalSemantic correspondenceVideo Understanding+1

A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re-Identification

2024-01-01 · CVPR 2024 1 · Zexian Yang, Dayan Wu, Chenming Wu, Zheng Lin 외

Extensive advancements have been made in person ReID through the mining of semantic information. Nevertheless existing methods that utilize semantic-parts from a single image modality do not explicitly achieve this g…

AttributePerson Re-IdentificationPrompt Engineering

ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts

2023-12-01 · CVPR 2024 1 · Mu Cai, Haotian Liu, Dennis Park, Siva Karthik Mustikovela 외

While existing large vision-language multimodal models focus on whole image understanding, there is a prominent gap in achieving region-specific comprehension. Current approaches that use textual coordinates or spatial e…

Visual Commonsense ReasoningVisual Prompting