paper-with-me

홈 › Papers

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

2025-10-29 · Nicolas Dufour, Lucas Degeorge, Arijit Ghosh, Vicky Kalogeiton, David Picard arxiv

The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2).

📄 PDF Abstract BibTeX arXiv:2510.25897

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification

2026-03-16 · MiroMind Team, S. Bai, L. Bing, L. Lei 외 arxiv

We present MiroThinker-1.7, a new research agent designed for complex long-horizon reasoning tasks. Building on this foundation, we further introduce MiroThinker-H1, which extends the agent with heavy-duty reasoning capa…

Domain Generalization by Mutual-Information Regularization with Pre-trained Models

2022-03-21 · Junbum Cha, Kyungjae Lee, Sungrae Park, Sanghyuk Chun

Domain generalization (DG) aims to learn a generalized model to an unseen target domain using only limited source domains. Previous attempts to DG fail to learn domain-invariant representations only from the source domai…

Domain Generalization

SAMIRO: Spatial Attention Mutual Information Regularization with a Pre-trained Model as Oracle for Lane Detection

2025-11-13 · Hyunjong Lee, Jangho Lee, Jaekoo Lee arxiv

Lane detection is an important topic in the future mobility solutions. Real-world environmental challenges such as background clutter, varying illumination, and occlusions pose significant obstacles to effective lane det…

Lane Detection

Language Reward Modulation for Pretraining Reinforcement Learning

2023-08-23 · Ademi Adeniji, Amber Xie, Carmelo Sferrazza, Younggyo Seo 외

Using learned reward functions (LRFs) as a means to solve sparse-reward reinforcement learning (RL) tasks has yielded some steady progress in task-complexity through the years. In this work, we question whether today's L…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Robot Manipulation

Future-conditioned Unsupervised Pretraining for Decision Transformer

2023-05-26 · Zhihui Xie, Zichuan Lin, Deheng Ye, Qiang Fu 외

Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promising, return conditioning is limited to …

Decision MakingReinforcement Learning (RL)