paper-with-me

홈 › Papers

Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation

2026-06-07 · Xuanyi Liu, Deyi Ji, Junyu Lu, Jing Wang, Lanyun Zhu, Qianxiong Xu, Xuhang Chen, Tianrun Chen, Siwei Ma arxiv

Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due to the brevity and ambiguity of user prompts. Existing approaches primarily polish the prompt for fluency and readability. However, the enhancement process still lacks visual grounding. As a result, the rewriter may over-infer missing details, causing an intent-generation gap. To address this limitation, we propose FaithRewriter, a novel prompt-enhancement framework for T2I generation. Specifically, FaithRewriter first leverages a multimodal MLLM to generate an image from the original prompt as an intermediate visual cue. This cue is then combined with the prompt and fed into a large-scale LLM to produce visually grounded augmentations that better reflect how the intended content should appear in images. Finally, these augmentations are distilled into a small-scale LLM for efficient deployment, enhancing its ability to generate effective T2I prompts. Experiments show that FaithRewriter yields prompts that are more faithful to the user intent and more visually plausible than strong baselines, helping narrow the intent-generation gap.

📄 PDF Abstract BibTeX arXiv:2606.08492

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image GenerationVisual Grounding

Similar Papers 제목 키워드 기반

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

2025-10-20 · Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo 외 arxiv

Vision-Language Models (VLMs) achieve strong results on multimodal tasks such as visual question answering, yet they can still fail even when the correct visual evidence is present. In this work, we systematically invest…

Visual Question Answering

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

2025-08-22 · Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello 외 arxiv

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion…

Emotion Recognition

Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise

2026-04-10 · Zibin Geng, Xuefeng Jiang, Jia Li, Zheng Li 외 arxiv

Prompt learning is a parameter-efficient approach for vision-language models, yet its robustness under label noise is less investigated. Visual content contains richer and more reliable semantic information, which remain…

Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding

2025-11-15 · Pinxue Guo, Chongruo Wu, Xinyu Zhou, Lingyi Hong 외 arxiv

Multimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensur…

Visual Grounding

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

2025-05-26 · Weihao Xuan, Qingcheng Zeng, Heli Qi, Junjue Wang 외

Uncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems. Among existing approaches, verbalized uncertainty, where models express their confidence through natural lan…

Uncertainty QuantificationVisual Reasoning