paper-with-me

Papers

Support Before Frequency in Discrete Diffusion

2026-05-13 · Adrian Müller, Antoine Gonon, Zebang Shen, Ya-Ping Hsieh, Niao He arxiv

Discrete diffusion models are increasingly competitive for language modeling, yet it remains unclear how their denoising objectives organize learning. Although these objectives target the full data distribution, we show that the exact reverse process induces a hierarchy between coarse support information and finer frequency information. For uniform and absorbing (a.k.a. masking) diffusion, we prove that, in the small-noise regime of the final denoising steps, each single-token reverse edit decomposes into a leading scale, determined by whether it moves toward the data support (e.g., grammatically valid sentences), and a finer coefficient, determining relative probabilities within the same scale. Thus, recovering validity structure only requires learning the correct order of magnitude of reverse probabilities, whereas recovering data frequencies requires coefficient-level estimation. The separation is mechanism-dependent: uniform diffusion exhibits a trichotomy into validity-improving, validity-preserving, and validity-worsening edits, while absorbing diffusion places its leading-order mass on validity-improving moves. Experiments on a masked language diffusion model and synthetic regular-language tasks support these predictions: support-localization emerges earlier than within-support frequency ranking, and the contrast between uniform and absorbing diffusion matches the predicted rate separation. Together, our results suggest that discrete diffusion models learn data support before data frequencies.

📄 PDF Abstract BibTeX arXiv:2605.13999

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control

2026-05-02 · Jinhao Zhang, Zhexuan Zhou, Huizhe Li, Yichen Lai 외 arxiv

Diffusion-based visuomotor policies perform well in robotic manipulation, yet current methods still inherit image-generation-style decoders and multi-step sampling. We revisit this design from a frequency-domain perspect…

Image Generation

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

2025-08-27 · Zhixuan Liang, Yizhuo Li, Tianshuo Yang, Chengyue Wu 외 arxiv

Vision-Language-Action (VLA) models adapt large vision-language backbones to map images and instructions into robot actions. However, prevailing VLAs either generate actions autoregressively in a fixed left-to-right orde…

Improving Discrete Diffusion Models via Structured Preferential Generation

2024-05-28 · Severi Rissanen, Markus Heinonen, Arno Solin

In the domains of image and audio, diffusion models have shown impressive performance. However, their application to discrete data types, such as language, has often been suboptimal compared to autoregressive generative …

Computing the Discrete Fourier Transform of signals with spectral frequency support

2021-02-24 · P Charantej Reddy, V S S Prabhu Tej, Aditya Siripuram, Brad Osgood

We consider the problem of finding the Discrete Fourier Transform (DFT) of $N-$ length signals with known frequency support of size $k$. When $N$ is a power of 2 and the frequency support is a spectral set, we provide an…

Diffusion Approximations for Expert Opinions in a Financial Market with Gaussian Drift

2018-07-02 · Jörn Sass, Dorothee Westphal, Ralf Wunderlich

This paper investigates a financial market where returns depend on an unobservable Gaussian drift process. While the observation of returns yields information about the underlying drift, we also incorporate discrete-time…