paper-with-me

Papers

LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models

2026-02-15 · Shufan Li, Yuchen Zhu, Jiuxiang Gu, Kangning Liu, Zhe Lin, Yongxin Chen, Molei Tao, Aditya Grover, Jason Kuen arxiv

Diffusion language models (dLLMs) recently emerged as a promising alternative to auto-regressive LLMs. The latest works further extended it to multimodal understanding and generation tasks. In this work, we propose LaViDa-R1, a multimodal, general-purpose reasoning dLLM. Unlike existing works that build reasoning dLLMs through task-specific reinforcement learning, LaViDa-R1 incorporates diverse multimodal understanding and generation tasks in a unified manner. In particular, LaViDa-R1 is built with a novel unified post-training framework that seamlessly integrates supervised finetuning (SFT) and multi-task reinforcement learning (RL). It employs several novel training techniques, including answer-forcing, tree search, and complementary likelihood estimation, to enhance effectiveness and scalability. Extensive experiments demonstrate LaViDa-R1's strong performance on a wide range of multimodal tasks, including visual math reasoning, reason-intensive grounding, and image editing.

📄 PDF Abstract BibTeX arXiv:2602.14147

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage Editing

Similar Papers 제목 키워드 기반

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

2025-09-23 · Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin 외 arxiv

We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks …

Text-to-Image GenerationMultimodal ReasoningImage Editing

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

2025-12-16 · Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin 외 arxiv

Masked Discrete Diffusion Models (MDMs) have achieved strong performance across a wide range of multimodal tasks, including image understanding, generation, and editing. However, their inference speed remains suboptimal …

Text-to-Image GenerationMathematical ReasoningImage Editing

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

2025-05-22 · Shufan Li, Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul 외

Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraini…

Instruction FollowingLanguage ModelingLanguage ModellingText Infilling+1

LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living

2024-06-13 · CVPR 2025 1 · Dominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind 외

Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant represent…

BenchmarkingHuman-Object Interaction DetectionRepresentation LearningVideo Description+1

No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

2026-02-22 · Zunkai Dai, Ke Li, Jiajia Liu, Jie Yang 외 arxiv

The collection and detection of video anomaly data has long been a challenging problem due to its rare occurrence and spatio-temporal scarcity. Existing video anomaly detection (VAD) methods under perform in open-world s…

Video Anomaly Detection