paper-with-me

Papers

The Design Space of Tri-Modal Masked Diffusion Models

2026-02-25 · Louis Bethune, Victor Turrisi, Bruno Kacper Mlodozeniec, Pau Rodriguez Lopez, Lokesh Boominathan, Nikhil Bhendawade, Amitis Shidani, Joris Pelemans, Theo X. Olausson, Devon Hjelm, Paul Dixon, Joao Monteiro, Pierre Ablin, Vishnu Banna, Arno Blaas, Nick Henderson, Kari Noriy, Dan Busbridge, Josh Susskind, Marco Cuturi, Irina Belousova, Luca Zappella, Russ Webb, Jason Ramapuram arxiv

Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent work initializing and fine-tuning a base unimodal model for bimodal generation. Diverging from previous approaches, we introduce the first tri-modal masked diffusion model pretrained from scratch on text, image-text, and audio-text data. We systematically analyze multimodal scaling laws, modality mixing ratios, noise schedules, and batch-size effects, and we provide optimized inference sampling defaults. Our batch-size analysis yields a novel stochastic differential equation (SDE)-based reparameterization that eliminates the need for tuning the optimal batch size as reported in recent work. This reparameterization decouples the physical batch size, often chosen based on compute constraints (GPU saturation, FLOP efficiency, wall-clock time), from the logical batch size, chosen to balance gradient variance during stochastic optimization. Finally, we pretrain a preliminary 3B-parameter tri-modal model on 6.4T tokens, demonstrating the capabilities of a unified design and achieving strong results in text generation, text-to-image tasks, and text-to-speech tasks. Our work represents the largest-scale systematic open study of multimodal discrete diffusion models conducted to date, providing insights into scaling behaviors across multiple modalities.

📄 PDF Abstract BibTeX arXiv:2602.21472

Code (0)

등록된 구현이 없습니다.

Tasks

Stochastic OptimizationText Generation

Similar Papers 제목 키워드 기반

Text-driven Human Motion Generation with Motion Masked Diffusion Model

2024-09-29 · Xingyu Chen

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, wh…

DiversityMotion Generation

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

2026-03-09 · Jaeik Kim, Woojin Kim, Jihwan Hong, Yejoon Lee 외 arxiv

We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlik…

Cross-Modal RetrievalSpeech RecognitionImage Generation

Training-Free Self-Correction for Multimodal Masked Diffusion Models

2026-02-02 · Yidong Ouyang, Panwen Hu, Zhengyan Wan, Zhe Wang 외 arxiv

Masked diffusion models have emerged as a powerful framework for text and multimodal generation. However, their sampling procedure updates multiple tokens simultaneously and treats generated tokens as immutable, which ma…

Text-to-Image Generationmultimodal generation

A Cheaper and Better Diffusion Language Model with Soft-Masked Noise

2023-04-10 · Jiaao Chen, Aston Zhang, Mu Li, Alex Smola 외

Diffusion models that are based on iterative denoising have been recently proposed and leveraged in various generation tasks like image generation. Whereas, as a way inherently built for continuous data, existing diffusi…

DenoisingImage GenerationLanguage ModelingLanguage Modelling

DGMR: Diffusion Guided Masked Reconstruction Framework for Multimodal Cloud Removal

2025-05-01 · IEEE Geoscience and Remote Sensing Letters 2025 5 · Coupled and decoupled learning, diffusion guidance, masked reconstruction, noncloudy difference similarity (NDS)

Cloudy conditions affect the quality of captured data by optical satellites. Multimodal techniques rely on synthetic aperture radar (SAR) images to recover cloudy pixels in optical images. These techniques face challenge…

Cloud Removal