paper-with-me

홈 › Papers

Three Creates All: You Only Sample 3 Steps

2026-03-23 · Yuren Cai, Guangyi Wang, Zongqing Li, Li Li, Zhihui Liu, Songzhi Su arxiv

Diffusion models deliver high-fidelity generation but remain slow at inference time due to many sequential network evaluations. We find that standard timestep conditioning becomes a key bottleneck for few-step sampling. Motivated by layer-dependent denoising dynamics, we propose Multi-layer Time Embedding Optimization (MTEO), which freeze the pretrained diffusion backbone and distill a small set of step-wise, layer-wise time embeddings from reference trajectories. MTEO is plug-and-play with existing ODE solvers, adds no inference-time overhead, and trains only a tiny fraction of parameters. Extensive experiments across diverse datasets and backbones show state-of-the-art performance in the few-step sampling and substantially narrow the gap between distillation-based and lightweight methods. Code will be available.

📄 PDF Abstract BibTeX arXiv:2603.22375

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ALLoRA: Adaptive Learning Rate Mitigates LoRA Fatal Flaws

2024-10-13 · Hai Huang, Randall Balestriero

Low-Rank Adaptation (LoRA) is the bread and butter of Large Language Model (LLM) finetuning. LoRA learns an additive low-rank perturbation, $AB$, of a pretrained matrix parameter $W$ to align the model to a new task or d…

Large Language Model

VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation

2026-01-05 · Shikun Sun, Liao Qu, Huichao Zhang, Yiheng Liu 외 arxiv

Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion, and Visual AutoRegressive (VAR) models. Unlike AR and diffusion, VARs operate on heterogeneous input structures across their generation s…

Reinforcement Learning

Learning to Select Views for Efficient Multi-View Understanding

2024-01-01 · CVPR 2024 1 · Yunzhong Hou, Stephen Gould, Liang Zheng

Multiple camera view (multi-view) setups have proven useful in many computer vision applications. However the high computational cost associated with multiple views creates a significant challenge for end devices wit…

CPUGPU

P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning

2026-07-28 · Liyun Yan, Jianming Ma, Yang Zhang, Shengcheng Fu 외 arxiv

Variational Autoencoders are widely used to encode high-dimensional and noisy observations in robotics. However, their stochastic latent creates a mismatch with Proximal Policy Optimization (PPO): an effective policy mar…

The Value of Covariance Matching in Gaussian DDPMs and the Lanczos Sampler

2026-05-21 · Md Sahil Akhtar, Aymane El Gadarri, Vivek F. Farias, Adam D. Jozefiak arxiv

A central error measure in Gaussian DDPMs is the path-space KL divergence between the exact reverse chain and the learned Gaussian reverse process. This quantity is especially relevant for procedures such as classifier g…