paper-with-me

홈 › Papers

VisuCraft: Enhancing Large Vision-Language Models for Complex Visual-Guided Creative Content Generation via Structured Information Extraction

2025-08-04 · Rongxin Jiang, Robert Long, Chenghao Gu, Mingrui Yan arxiv

This paper introduces VisuCraft, a novel framework designed to significantly enhance the capabilities of Large Vision-Language Models (LVLMs) in complex visual-guided creative content generation. Existing LVLMs often exhibit limitations in maintaining high visual fidelity, genuine creativity, and precise adherence to nuanced user instructions when generating long-form texts. VisuCraft addresses these challenges by integrating a multimodal structured information extractor (E) and a dynamic prompt generation module (G). The extractor distills fine-grained visual attributes from input images into a rich, structured representation, which the dynamic prompt module then combines with user instructions to create highly optimized prompts for underlying LVLMs (e.g., LLaVA, InstructBLIP). Evaluated on the self-constructed ImageStoryGen-500K dataset using VisuGen Metrics (Visual Grounding, Creativity, and Instruction Adherence), VisuCraft consistently outperforms baseline LVLMs across tasks like story generation and poetry composition. Our results demonstrate remarkable improvements, particularly in creativity and instruction adherence, validating VisuCraft's effectiveness in producing imaginative, visually grounded, and user-aligned long-form creative text. This work unlocks new potential for LVLMs in sophisticated creative AI applications.

📄 PDF Abstract BibTeX arXiv:2508.02890

Code (0)

등록된 구현이 없습니다.

Tasks

Information ExtractionStory GenerationVisual Grounding

Similar Papers 제목 키워드 기반

Enhancing Model Performance: Another Approach to Vision-Language Instruction Tuning

2024-07-25 · Vedanshu, MM Tripathi, Bhavnesh Jaint

The integration of large language models (LLMs) with vision-language (VL) tasks has been a transformative development in the realm of artificial intelligence, highlighting the potential of LLMs as a versatile general-pur…

Chatbot

Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders

2025-10-29 · Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk 외 arxiv

Despite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that current Video Large Language Model (Video-L…

Video Question AnsweringAction Recognition

Enhancing Advanced Visual Reasoning Ability of Large Language Models

2024-09-21 · Zhiyuan Li, Dongnan Liu, Chaoyi Zhang, Heng Wang 외

Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models' advanced reasoning ability. Traditional Vision-Language Models (VLMs) perform well in vis…

In-Context LearningVisual Reasoning

MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing

2025-02-21 · Matvey Skripkin, Elizaveta Goncharova, Dmitrii Tarasov, Andrey Kuznetsov

Multimodal language models (MLMs) integrate visual and textual information by coupling a vision encoder with a large language model through the specific adapter. While existing approaches commonly rely on a single pre-tr…

Language ModelingLanguage ModellingLarge Language Model

Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language Interaction

2024-04-30 · Guobin Shen, Dongcheng Zhao, Xiang He, Linghao Feng 외

Decoding non-invasive brain recordings is pivotal for advancing our understanding of human cognition but faces challenges due to individual differences and complex neural signal representations. Traditional methods often…

Brain DecodingImage ReconstructionQuestion Answering