Coherent Zero-Shot Visual Instruction Generation
Despite the advances in text-to-image synthesis, particularly with diffusion models, generating visual instructions that require consistent representation and smooth state transitions of objects across sequential steps remains a formidable challenge. This paper introduces a simple, training-free framework to tackle the issues, capitalizing on the advancements in diffusion models and large language models (LLMs). Our approach systematically integrates text comprehension and image generation to ensure visual instructions are visually appealing and maintain consistency and accuracy throughout the instruction sequence. We validate the effectiveness by testing multi-step instructions and comparing the text alignment and consistency with several baselines. Our experiments show that our approach can visualize coherent and visually pleasing instructions
Code (0)
등록된 구현이 없습니다.
Tasks
Image GenerationReading ComprehensionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Zero-Shot Anticipation for Instructional Activities
How can we teach a robot to predict what will happen next for an activity it has never seen before? We address this problem of zero-shot anticipation by presenting a hierarchical model that generalizes instructional know…
Zero-Shot LearningLearning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning
Generative vision-language models (VLMs) have shown impressive performance in zero-shot vision-language tasks like image captioning and visual question answering. However, improving their zero-shot reasoning typically re…
Image CaptioningInstruction FollowingLanguage ModelingLanguage Modelling+4DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, th…
Text-to-Video GenerationVideo GenerationZero-shot Text-to-Video GenerationCoDi-2: In-Context Interleaved and Interactive Any-to-Any Generation
We present CoDi-2 a Multimodal Large Language Model (MLLM) for learning in-context interleaved multimodal representations. By aligning modalities with language for both encoding and generation CoDi-2 empowers Large L…
Image GenerationLanguage ModelingLanguage ModellingLarge Language Model+1EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation
Long-horizon video generation has advanced in visual quality, yet existing methods still struggle to maintain knowledge consistency and coherent pedagogical narratives across multi-shot instructional videos, especially i…
Video Generation