LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models
Diffusion language models (dLLMs) recently emerged as a promising alternative to auto-regressive LLMs. The latest works further extended it to multimodal understanding and generation tasks. In this work, we propose LaViDa-R1, a multimodal, general-purpose reasoning dLLM. Unlike existing works that build reasoning dLLMs through task-specific reinforcement learning, LaViDa-R1 incorporates diverse multimodal understanding and generation tasks in a unified manner. In particular, LaViDa-R1 is built with a novel unified post-training framework that seamlessly integrates supervised finetuning (SFT) and multi-task reinforcement learning (RL). It employs several novel training techniques, including answer-forcing, tree search, and complementary likelihood estimation, to enhance effectiveness and scalability. Extensive experiments demonstrate LaViDa-R1's strong performance on a wide range of multimodal tasks, including visual math reasoning, reason-intensive grounding, and image editing.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningImage EditingSimilar Papers 제목 키워드 기반
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks …
Text-to-Image GenerationMultimodal ReasoningImage EditingSparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
Masked Discrete Diffusion Models (MDMs) have achieved strong performance across a wide range of multimodal tasks, including image understanding, generation, and editing. However, their inference speed remains suboptimal …
Text-to-Image GenerationMathematical ReasoningImage EditingLaViDa: A Large Diffusion Language Model for Multimodal Understanding
Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraini…
Instruction FollowingLanguage ModelingLanguage ModellingText Infilling+1LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant represent…
BenchmarkingHuman-Object Interaction DetectionRepresentation LearningVideo Description+1No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
The collection and detection of video anomaly data has long been a challenging problem due to its rare occurrence and spatio-temporal scarcity. Existing video anomaly detection (VAD) methods under perform in open-world s…
Video Anomaly Detection