paper-with-me

홈 › Papers

Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images

2024-04-21 · Ali Naseh, Katherine Thai, Mohit Iyyer, Amir Houmansadr

With the digital imagery landscape rapidly evolving, image stocks and AI-generated image marketplaces have become central to visual media. Traditional stock images now exist alongside innovative platforms that trade in prompts for AI-generated visuals, driven by sophisticated APIs like DALL-E 3 and Midjourney. This paper studies the possibility of employing multi-modal models with enhanced visual understanding to mimic the outputs of these platforms, introducing an original attack strategy. Our method leverages fine-tuned CLIP models, a multi-label classifier, and the descriptive capabilities of GPT-4V to create prompts that generate images similar to those available in marketplaces and from premium stock image providers, yet at a markedly lower expense. In presenting this strategy, we aim to spotlight a new class of economic and security considerations within the realm of digital imagery. Our findings, supported by both automated metrics and human assessment, reveal that comparable visual content can be produced for a fraction of the prevailing market prices ($0.23 - $0.27 per image), emphasizing the need for awareness and strategic discussions about the integrity of digital media in an increasingly AI-integrated landscape. Our work also contributes to the field by assembling a dataset consisting of approximately 19 million prompt-image pairs generated by the popular Midjourney platform, which we plan to release publicly.

📄 PDF Abstract BibTeX arXiv:2404.13784

Code (0)

등록된 구현이 없습니다.

Tasks

Descriptive

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor

2024-12-01 · Ashwin Baluja

While Large Language Models (LLMs) have demonstrated impressive natural language understanding capabilities across various text-based tasks, understanding humor has remained a persistent challenge. Humor is frequently mu…

AllNatural Language UnderstandingRhythmtext-to-speech+1

Exploring Multimodal Prompt for Visualization Authoring with Large Language Models

2025-04-18 · Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang 외

Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language…

Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM

2024-07-31 · Can Wang, Hongliang Zhong, Menglei Chai, Mingming He 외

Automatic furniture layout is long desired for convenient interior design. Leveraging the remarkable visual reasoning capabilities of multimodal large language models (MLLMs), recent methods address layout generation in …

In-Context LearningLayout DesignLayout GenerationVisual Prompting+1

POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models

2024-06-06 · Jianben He, Xingbo Wang, Shiyi Liu, Guande Wu 외

Large language models (LLMs) have exhibited impressive abilities for multimodal content comprehension and reasoning with proper prompting in zero- or few-shot settings. Despite the proliferation of interactive systems de…

Multimodal ReasoningPrompt Engineering

Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

2024-03-29 · Weifeng Lin, Xinyu Wei, Ruichuan An, Peng Gao 외

The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level compr…

Instruction FollowingLanguage ModellingLarge Language Modelmultimodal interaction+4