paper-with-me

홈 › Papers

CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation

2025-07-11 · Zhanxin Gao, Beier Zhu, Liang Yao, Jian Yang, Ying Tai arxiv

Subject-consistent generation (SCG)-aiming to maintain a consistent subject identity across diverse scenes-remains a challenge for text-to-image (T2I) models. Existing training-free SCG methods often achieve consistency at the cost of layout and pose diversity, hindering expressive visual storytelling. To address the limitation, we propose subject-Consistent and pose-Diverse T2I framework, dubbed as CoDi, that enables consistent subject generation with diverse pose and layout. Motivated by the progressive nature of diffusion, where coarse structures emerge early and fine details are refined later, CoDi adopts a two-stage strategy: Identity Transport (IT) and Identity Refinement (IR). IT operates in the early denoising steps, using optimal transport to transfer identity features to each target image in a pose-aware manner. This promotes subject consistency while preserving pose diversity. IR is applied in the later denoising steps, selecting the most salient identity features to further refine subject details. Extensive qualitative and quantitative results on subject consistency, pose diversity, and prompt fidelity demonstrate that CoDi achieves both better visual perception and stronger performance across all metrics. The code is provided in https://github.com/NJU-PCALab/CoDi.

📄 PDF Abstract BibTeX arXiv:2507.08396

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image GenerationVisual Storytelling

Similar Papers 제목 키워드 기반

Enhanced Generative Machine Listener

2025-09-25 · Vishnu Raj, Gouthaman KV, Shiv Gehlot, Lars Villemoes 외 arxiv

We present GMLv2, a reference-based model designed for the prediction of subjective audio quality as measured by MUSHRA scores. GMLv2 introduces a Beta distribution-based loss to model the listener ratings and incorporat…

EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding

2026-08-13 · Shuailei Zhang, Muyun Jiang, Wei Zhang, Jinbo Chen 외 arxiv

Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG founda…

Representation LearningEmotion RecognitionEeg Decoding

MindLLM: A Subject-Agnostic and Versatile Model for fMRI-to-Text Decoding

2025-02-18 · Weikang Qiu, Zheng Huang, Haoyu Hu, Aosong Feng 외

Decoding functional magnetic resonance imaging (fMRI) signals into text has been a key challenge in the neuroscience community, with the potential to advance brain-computer interfaces and uncover deeper insights into bra…

Decision Making

Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding

2026-04-09 · Mu Nan, Muquan Yu, Weijian Mai, Jacob S. Prince 외 arxiv

Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural representations and computational models of vision. A field-wide goal is…

Brain Decoding

MindFormer: Semantic Alignment of Multi-Subject fMRI for Brain Decoding

2024-05-28 · Inhwa Han, Jaayeon Lee, Jong Chul Ye

Research efforts for visual decoding from fMRI signals have attracted considerable attention in research community. Still multi-subject fMRI decoding with one model has been considered intractable due to the drastic vari…

Brain DecodingImage GenerationLanguage ModellingLarge Language Model+1