paper-with-me

홈 › Papers

Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models

2025-08-17 · Xun Su, Jianming Huang, Yang Yusen, Zhongxi Fang, Hiroyuki Kasai arxiv

Inference-time scaling has achieved remarkable success in language models, yet its adaptation to diffusion models remains underexplored. We observe that the efficacy of recent Sequential Monte Carlo (SMC)-based methods largely stems from globally fitting the The reward-tilted distribution, which inherently preserves diversity during multi-modal search. However, current applications of SMC to diffusion models face a fundamental dilemma: early-stage noise samples offer high potential for improvement but are difficult to evaluate accurately, whereas late-stage samples can be reliably assessed but are largely irreversible. To address this exploration-exploitation trade-off, we approach the problem from the perspective of the search algorithm and propose two strategies: Funnel Schedule and Adaptive Temperature. These simple yet effective methods are tailored to the unique generation dynamics and phase-transition behavior of diffusion models. By progressively reducing the number of maintained particles and down-weighting the influence of early-stage rewards, our methods significantly enhance sample quality without increasing the total number of Noise Function Evaluations. Experimental results on multiple benchmarks and state-of-the-art text-to-image diffusion models demonstrate that our approach outperforms previous baselines.

📄 PDF Abstract BibTeX arXiv:2508.12361

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Recommandation mobile, sensible au contexte de contenus évolutifs: Contextuel-E-Greedy

2014-02-09 · Djallel Bouneffouf

We introduce in this paper an algorithm named Contextuel-E-Greedy that tackles the dynamicity of the user's content. It is based on dynamic exploration/exploitation tradeoff and can adaptively balance the two aspects by …

Fiduciary Bandits

2019-05-16 · ICML 2020 1 · Gal Bahar, Omer Ben-Porat, Kevin Leyton-Brown, Moshe Tennenholtz

Recommendation systems often face exploration-exploitation tradeoffs: the system can only learn about the desirability of new options by recommending them to some user. Such systems can thus be modeled as multi-armed ban…

Recommendation Systems

Thompson Sampling in Dynamic Systems for Contextual Bandit Problems

2013-10-17 · Tianbing Xu, Yaming Yu, John Turner, Amelia Regan

We consider the multiarm bandit problems in the timevarying dynamic system for rich structural features. For the nonlinear dynamic model, we propose the approximate inference for the posterior distributions based on Lapl…

Thompson Sampling

Risk and Ambiguity in Information Seeking: Eye Gaze Patterns Reveal Contextual Behaviour in Dealing with Uncertainty

2016-06-27 · Wittek Peter, Liu Ying-Hsang, Darányi Sándor, Gedeon Tom 외

Information foraging connects optimal foraging theory in ecology with how humans search for information. The theory suggests that, following an information scent, the information seeker must optimize the tradeoff between…

Generative Exploration and Exploitation

2019-04-21 · Jiechuan Jiang, Zongqing Lu

Sparse reward is one of the biggest challenges in reinforcement learning (RL). In this paper, we propose a novel method called Generative Exploration and Exploitation (GENE) to overcome sparse reward. GENE automatically …

Reinforcement LearningReinforcement Learning (RL)