paper-with-me

홈 › Papers

Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View

2025-11-10 · Jianyu Qi, Ding Zou, Wenrui Yan, Rui Ma, Jiaxu Li, Zhijie Zheng, Zhiguo Yang, Rongchang Zhao arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have spurred significant progress in Chain-of-Thought (CoT) reasoning. Building on the success of Deepseek-R1, researchers extended multimodal reasoning to post-training paradigms based on reinforcement learning (RL), focusing predominantly on mathematical datasets. However, existing post-training paradigms tend to neglect two critical aspects: (1) The lack of quantifiable difficulty metrics capable of strategically screening samples for post-training optimization. (2) Suboptimal post-training paradigms that fail to jointly optimize perception and reasoning capabilities. To address this gap, we propose two novel difficulty-aware sampling strategies: Progressive Image Semantic Masking (PISM) quantifies sample hardness through systematic image degradation, while Cross-Modality Attention Balance (CMAB) assesses cross-modal interaction complexity via attention distribution analysis. Leveraging these metrics, we design a hierarchical training framework that incorporates both GRPO-only and SFT+GRPO hybrid training paradigms, and evaluate them across six benchmark datasets. Experiments demonstrate consistent superiority of GRPO applied to difficulty-stratified samples compared to conventional SFT+GRPO pipelines, indicating that strategic data sampling can obviate the need for supervised fine-tuning while improving model accuracy. Our code will be released at https://github.com/qijianyu277/DifficultySampling.

📄 PDF Abstract BibTeX arXiv:2511.06722

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMultimodal Reasoning

Similar Papers 제목 키워드 기반

Scalable Nonparametric Sampling from Multimodal Posteriors with the Posterior Bootstrap

2019-02-08 · Edwin Fong, Simon Lyddon, Chris Holmes

Increasingly complex datasets pose a number of challenges for Bayesian inference. Conventional posterior sampling based on Markov chain Monte Carlo can be too computationally intensive, is serial in nature and mixes poor…

Bayesian Inferenceregression

When, why, and how do diffusion posterior samplers fail? A finite-sample lens

2026-05-28 · Benjamin A. Burns, Sara Fridovich-Keil arxiv

Diffusion models have excellent capacity to model complex distributions of natural data, which has made them a popular and effective choice for posterior sampling in imaging inverse problems. Existing methods can incorpo…

Sequential sampling without comparison to boundary through model-free reinforcement learning

2024-08-12 · Jamal Esmaily, Rani Moran, Yasser Roudi, Bahador Bahrami

Although evidence integration to the boundary model has successfully explained a wide range of behavioral and neural data in decision making under uncertainty, how animals learn and optimize the boundary remains unresolv…

Decision MakingDecision Making Under Uncertainty

Revisiting Greedy Decoding for Visual Question Answering: A Calibration Perspective

2026-04-25 · Boqi Chen, Xudong Liu, Yunke Ao, Jianing Qiu arxiv

Stochastic sampling strategies are widely adopted in large language models (LLMs) to balance output coherence and diversity. These heuristics are often inherited in Multimodal LLMs (MLLMs) without task-specific justifica…

Visual Question AnsweringMultimodal Reasoning

Jump-Diffusion Langevin Dynamics for Multimodal Posterior Sampling

2022-11-02 · Jacopo Guidolin, Vyacheslav Kungurtsev, Ondřej Kuželka

Bayesian methods of sampling from a posterior distribution are becoming increasingly popular due to their ability to precisely display the uncertainty of a model fit. Classical methods based on iterative random sampling …