paper-with-me

Papers

Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models

2025-05-13 · Donghoon Kim, Minji Bae, Kyuhong Shim, Byonghyo Shim

Text-to-image generative models like DALL-E and Stable Diffusion have revolutionized visual content creation across various applications, including advertising, personalized media, and design prototyping. However, crafting effective textual prompts to guide these models remains challenging, often requiring extensive trial and error. Existing prompt inversion approaches, such as soft and hard prompt techniques, are not so effective due to the limited interpretability and incoherent prompt generation. To address these issues, we propose Visually Guided Decoding (VGD), a gradient-free approach that leverages large language models (LLMs) and CLIP-based guidance to generate coherent and semantically aligned prompts. In essence, VGD utilizes the robust text generation capabilities of LLMs to produce human-readable prompts. Further, by employing CLIP scores to ensure alignment with user-specified visual concepts, VGD enhances the interpretability, generalization, and flexibility of prompt generation without the need for additional training. Our experiments demonstrate that VGD outperforms existing prompt inversion techniques in generating understandable and contextually relevant prompts, facilitating more intuitive and controllable interactions with text-to-image models.

📄 PDF Abstract BibTeX arXiv:2505.08622

Code (1)

DonghoonKim-1938/VGD 공식 구현 pytorch

Tasks

Text Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs

2026-03-19 · Govinda Kolli, Adinath Madhavrao Dukre, Behzad Bozorgtabar, Dwarikanath Mahapatra 외 arxiv

Medical Vision-Language Models (VLMs) often hallucinate by generating responses based on language priors rather than visual evidence, posing risks in clinical applications. We propose Visual Grounding Score Guided Decodi…

Visual Grounding

Decoding Saccadic Eye Movements from Brain Signals Using an Endovascular Neural Interface

2025-06-09 · Suleman Rasheed, James Bennett, Peter E. Yoo, Anthony N. Burkitt 외

An Oculomotor Brain-Computer Interface (BCI) records neural activity from regions of the brain involved in planning eye movements and translates this activity into control commands. While previous successful oculomotor B…

Brain Computer InterfaceEEG

Language Models Can See: Plugging Visual Controls in Text Generation

2022-05-05 · Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu 외

Generative language models (LMs) such as GPT-2/3 can be prompted to generate text with remarkable quality. While they are designed for text-prompted generation, it remains an open question how the generation process coul…

Image CaptioningImage-text matchingOpen-Ended Question AnsweringStory Generation+2

ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding

2026-03-20 · Quan Kong, Yuhao Shen, Yicheng Ji, Huan Li 외 arxiv

Although current Video-LLMs achieve impressive performance in video understanding tasks, their autoregressive decoding efficiency remains constrained by the massive number of video tokens. Visual token pruning can partia…

Electrocorticographic Dynamics Predict Visually Guided Motor Imagery of Grasp Shaping

2017-02-21

Identification of intended movement type and movement phase of hand grasp shaping are critical features for the control of volitional neuroprosthetics. We demonstrate that neural dynamics during visually-guided imagined …

Motor Imagery