paper-with-me

홈 › Papers

Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper

2026-08-27 · Rongjin Li, Yuanxin Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun arxiv

Multimodal large language models (MLLMs) are increasingly capable scientific assistants, yet they remain far from fully autonomous research. This transition requires models to actively inspect academic papers, build global evidence views, and make traceable judgments without prespecified issues or evidence. However, existing work provides limited task paradigms or training studies for such issue- and evidence-absent verification. We study this challenge through scientific error detection, where models must determine whether errors exist and justify them with evidence-based reasoning. To fill this gap, we present VERA-RL, a reinforcement-learning formulation for scientific error detection over academic papers. Following a Reason--Verify--Scan progression, we construct VERA-13K, a 12,900-sample dataset organized into 4,300 matched chains, covering 6 scientific-error categories across the research workflow and broad natural-science domains. We further introduce fine-grained rewards for reasoning completeness, evidence alignment, and error precision. Training Qwen3-VL-8B with VERA-RL substantially improves verifiable reasoning, approaching flagship MLLMs such as Gemini 3 Pro and Qwen3-VL-235B-A22B on Scan.

📄 PDF Abstract BibTeX arXiv:2608.26596

Code (3)

Aaron617/agent-arXiv-daily ★ 9
Tavish9/awesome-daily-AI-arxiv ★ 113
arxivsub/arXivSub_daily_arxiv ★ 4

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

2025-07-20 · Yi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng 외 arxiv

Large language models (LLMs), despite possessing latent safety understanding from their vast pretraining data, remain vulnerable to generating harmful content and exhibit issues such as over-refusal and utility degradati…

Reinforcement Learning

PICTS: A Novel Deep Reinforcement Learning Approach for Dynamic P-I Control in Scanning Probe Microscopy

2025-02-11 · Ziwei Wei, Shuming Wei, Qibin Zeng, Wanheng Lu 외

We have developed a Parallel Integrated Control and Training System, leveraging the deep reinforcement learning to dynamically adjust the control strategies in real time for scanning probe microscopy techniques.

Deep Reinforcement Learningreinforcement-learningReinforcement Learning

What-If Analysis of Large Language Models: Explore the Game World Using Proactive Thinking

2025-09-05 · Yuan Sui, Yanming Zhang, Yi Liao, Yu Gu 외 arxiv

LLMs struggle with decision-making in high-stakes environments like MOBA games, primarily due to a lack of proactive reasoning and limited understanding of complex game dynamics. To address this, we propose What-if Analy…

Reinforcement Learning

ProGuard: Towards Proactive Multimodal Safeguard

2025-12-29 · Shaohan Yu, Lijun Li, Chenyang Si, Lu Sheng 외 arxiv

The rapid evolution of generative models has led to a continuous emergence of multimodal safety risks, exposing the limitations of existing defense methods. To address these challenges, we propose ProGuard, a vision-lang…

Reinforcement Learning

Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking

2025-09-27 · Jinyi Han, Ying Huang, Ying Liao, Zishang Jiang 외 arxiv

Large Reasoning Models (LRMs) have achieved impressive performance on challenging tasks, yet their deep reasoning often incurs substantial computational costs. To achieve efficient reasoning, existing reinforcement learn…

Reinforcement Learning