paper-with-me

홈 › Papers

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

2025-08-07 · Linghao Zhu, Yiran Guan, Dingkang Liang, Jianzhong Ju, Zhenbo Luo, Bin Qin, Jian Luan, Yuliang Liu, Xiang Bai arxiv

Reinforcement learning (RL) has emerged as an effective post-training paradigm for enhancing the reasoning capabilities of multimodal large language model (MLLM). However, current RL pipelines often suffer from training inefficiencies caused by two underexplored issues: Advantage Collapsing, where most advantages in a batch concentrate near zero, and Rollout Silencing, where the proportion of rollouts contributing non-zero gradients diminishes over time. These issues lead to suboptimal gradient updates and hinder long-term learning efficiency. To address these issues, we propose Shuffle-R1, a simple yet principled framework that improves RL fine-tuning efficiency by dynamically restructuring trajectory sampling and batch composition. It introduces (1) Pairwise Trajectory Sampling, which selects high-contrast trajectories with large advantages to improve gradient signal quality, and (2) Advantage-based Trajectory Shuffle, which increases exposure of valuable rollouts through informed batch reshuffling. Experiments across multiple reasoning benchmarks show that our framework consistently outperforms strong RL baselines with minimal overhead. These results highlight the importance of data-centric adaptations for more efficient RL training in MLLM.

📄 PDF Abstract BibTeX arXiv:2508.05612

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

2025-01-09 · Shiji Zhao, Ranjie Duan, Fengxiang Wang, Chi Chen 외

Multimodal Large Language Models (MLLMs) have achieved impressive performance and have been put into practical use in commercial applications, but they still have potential safety mechanism vulnerabilities. Jailbreak att…

Red Teaming

Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack

2025-11-02 · Xin Liu, Aoyang Zhou, Aoyang Zhou arxiv

Visual-Language Pre-training (VLP) models have achieved significant performance across various downstream tasks. However, they remain vulnerable to adversarial examples. While prior efforts focus on improving the adversa…

Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models

2025-04-24 · Xu Ma, Peize Sun, Haoyu Ma, Hao Tang 외

Autoregressive (AR) models, long dominant in language generation, are increasingly applied to image synthesis but are often considered less competitive than Diffusion-based models. A primary limitation is the substantial…

Image GenerationText GenerationText to Image GenerationText-to-Image Generation

HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

2020-05-01 · EMNLP 2020 11 · Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan 외

We present HERO, a novel framework for large-scale video+language omni-representation learning. HERO encodes multimodal inputs in a hierarchical structure, where local context of a video frame is captured by a Cross-moda…

Language ModelingLanguage ModellingMasked Language ModelingMoment Retrieval+7

On the Self Shuffle Language

2022-02-16 · Pamela Fleischmann, Tero Harju, Lukas Haschke, Jonas Höfer 외

The shuffle product \(u\shuffle v\) of two words \(u\) and \(v\) is the set of all words which can be obtained by interleaving \(u\) and \(v\). Motivated by the paper \emph{The Shuffle Product: New Research Directions} b…