paper-with-me

홈 › Papers

Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models

2025-12-22 · Zixuan Ye, Quande Liu, Cong Wei, Yuanxing Zhang, Xintao Wang, Pengfei Wan, Kun Gai, Wenhan Luo arxiv

Recently, the introduction of Chain-of-Thought (CoT) has largely improved the generation ability of unified models. However, it is observed that the current thinking process during generation mainly focuses on the text consistency with the text prompt, ignoring the \textbf{visual context consistency} with the visual reference images during the multi-modal generation, e.g., multi-reference generation. The lack of such consistency results in the failure in maintaining key visual features (like human ID, object attribute, style). To this end, we integrate the visual context consistency into the reasoning of unified models, explicitly motivating the model to sustain such consistency by 1) Adaptive Visual Planning: generating structured visual check list to figure out the visual element of needed consistency keeping, and 2) Iterative Visual Correction: performing self-reflection with the guidance of check lists and refining the generated result in an iterative manner. To achieve this, we use supervised finetuning to teach the model how to plan the visual checking, conduct self-reflection and self-refinement, and use flow-GRPO to further enhance the visual consistency through a customized visual checking reward. The experiments show that our method outperforms both zero-shot unified models and those with text CoTs in multi-modal generation, demonstrating higher visual context consistency.

📄 PDF Abstract BibTeX arXiv:2512.19686

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis

2025-04-08 · Tri Ton, Ji Woo Hong, Chang D. Yoo

This paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transform…

Audio SynthesisFAD

Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans

2025-11-16 · Hongbin Huang, Junwei Li, Tianxin Xie, Zhuang Li 외 arxiv

High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversation…

Dialogue GenerationResponse GenerationSpeech Synthesis

Depth-Aware Endoscopic Video Inpainting

2024-07-02 · Francis Xiatian Zhang, Shuang Chen, Xianghua Xie, Hubert P. H. Shum

Video inpainting fills in corrupted video content with plausible replacements. While recent advances in endoscopic video inpainting have shown potential for enhancing the quality of endoscopic videos, they mainly repair …

Depth EstimationVideo Inpainting

FontCrafter: High-Fidelity Element-Driven Artistic Font Creation with Visual In-Context Generation

2026-03-23 · Wuyang Luo, Chengkai Tan, Chang Ge, Binye Hong 외 arxiv

Artistic font generation aims to synthesize stylized glyphs based on a reference style. However, existing approaches suffer from limited style diversity and coarse control. In this work, we explore the potential of eleme…

Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation

2026-05-07 · Yixin Zhu, Zixiong Wang, Jian Yang, Jin Xie 외 arxiv

Reliable simulation evaluation of robot manipulation policies serves as a high-fidelity proxy for real-world performance. Although existing benchmarks cover a wide range of task categories, they lack visual realism, crea…

Robot Manipulation