paper-with-me

홈 › Papers

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation

2025-11-25 · Zhoujie Fu, Xianfang Zeng, Jinghong Lan, Xinyao Liao, Cheng Chen, Junyi Chen, Jiacheng Wei, Wei Cheng, Shiyu Liu, Yunuo Chen, Gang Yu, Guosheng Lin arxiv

Pre-trained video models learn powerful priors for generating high-quality, temporally coherent content. While these models excel at temporal coherence, their dynamics are often constrained by the continuous nature of their training data. We hypothesize that by injecting the rich and unconstrained content diversity from image data into this coherent temporal framework, we can generate image sets that feature both natural transitions and a far more expansive dynamic range. To this end, we introduce iMontage, a unified framework designed to repurpose a powerful video model into an all-in-one image generator. The framework consumes and produces variable-length image sets, unifying a wide array of image generation and editing tasks. To achieve this, we propose an elegant and minimally invasive adaptation strategy, complemented by a tailored data curation process and training paradigm. This approach allows the model to acquire broad image manipulation capabilities without corrupting its invaluable original motion priors. iMontage excels across several mainstream many-in-many-out tasks, not only maintaining strong cross-image contextual consistency but also generating scenes with extraordinary dynamics that surpass conventional scopes. Find our homepage at: https://kr1sjfu.github.io/iMontage-web/.

📄 PDF Abstract BibTeX arXiv:2511.20635

Code (0)

등록된 구현이 없습니다.

Tasks

Image ManipulationImage Generation

Similar Papers 제목 키워드 기반

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

2026-05-21 · Haokun Wen, Xuemeng Song, Xinghao Xie, Xiaolin Chen 외 arxiv

Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is highly desired in practice. However, existing approaches focus on na…

Image Retrieval

A Unified Model for Compressed Sensing MRI Across Undersampling Patterns

2024-10-05 · CVPR 2025 1 · Armeet Singh Jatyani, Jiayun Wang, Aditi Chandrashekar, Zihui Wu 외

Compressed Sensing MRI reconstructs images of the body's internal anatomy from undersampled measurements, thereby reducing scan time. Recently, deep learning has shown great potential for reconstructing high-fidelity ima…

Anatomycompressed sensingDiagnosticMRI Reconstruction+2

$\mathcal{P}^3$: Toward Versatile Embodied Agents

2025-08-09 · Shengli Zhou, Xiangchen Wang, Jinrui Zhang, Ruozai Tian 외 arxiv

Embodied agents have shown promising generalization capabilities across diverse physical environments, making them essential for a wide range of real-world applications. However, building versatile embodied agents poses …

Learning Versatile Skills with Curriculum Masking

2024-10-23 · Yao Tang, Zhihui Xie, Zichuan Lin, Deheng Ye 외

Masked prediction has emerged as a promising pretraining paradigm in offline reinforcement learning (RL) due to its versatile masking schemes, enabling flexible inference across various downstream tasks with a unified mo…

Decision MakingOffline RLReinforcement Learning (RL)Sequential Decision Making

GuideWalk: Learning Unified Autonomous Navigation and Locomotion for Humanoid Robots across Versatile Terrains

2026-06-09 · Haoxuan Han, Chen Chen, Linao Gong, Xin Yang 외 arxiv

Humanoid robots have achieved strong locomotion capabilities, but reliable navigation on versatile terrains remains challenging because obstacle avoidance must be coordinated with dynamically feasible motion. In this wor…

Reinforcement Learning