paper-with-me

홈 › Papers

IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning

2026-04-27 · Navya Gupta, Rishitej Reddy Vyalla, Avinash Anand, Chhavi Kirtani, Erik Cambria, Zhengchen Zhang, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah arxiv

Curriculum learning helps language models tackle complex reasoning by gradually increasing task difficulty. However, it often fails to generate consistent step-by-step reasoning, especially in multilingual and low-resource settings where cross-lingual transfer from English to Indian languages remains limited. We propose IRIS: Interleaved Reinforcement with Incremental Staged Curriculum, a two-axis framework that combines Supervised Fine-Tuning on progressively harder problems (vertical axis) with Reverse Curriculum Reinforcement Learning to reduce reliance on step-by-step guidance (horizontal axis). We design a composite reward combining correctness, step-wise alignment, continuity, and numeric incentives, optimized via Group Relative Policy Optimization (GRPO). We release CL-Math, a dataset of 29k problems with step-level annotations in English, Hindi, and Marathi. Across standard benchmarks and curated multilingual test sets, IRIS consistently improves performance, with strong results on math reasoning tasks and substantial gains in low-resource and bilingual settings, alongside modest improvements in high-resource languages.

📄 PDF Abstract BibTeX arXiv:2604.24114

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningCross-Lingual TransferMathematical Reasoning

Similar Papers 제목 키워드 기반

Automating Staged Rollout with Reinforcement Learning

2022-04-01 · Shadow Pritchard, Vidhyashree Nagaraju, Lance Fiondella

Staged rollout is a strategy of incrementally releasing software updates to portions of the user population in order to accelerate defect discovery without incurring catastrophic outcomes such as system wide outages. Som…

Multi-Objective Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

IRIS-GAN: Staged Specialist Detection of Deepfake Faces

2026-06-03 · Jaume M. Trenchs, Veronica Sanz arxiv

We introduce IRIS-GAN, a specialist forensic detector for synthetic face images under cross-generator shift. Rather than addressing universal synthetic-image detection, we focus on faces generated by generative adversari…

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control

2025-08-27 · Quanfeng Lu, Zhantao Ma, Shuai Zhong, Jin Wang 외 arxiv

The rapid advancement of large vision language models (LVLMs) and agent systems has heightened interest in mobile GUI agents that can reliably translate natural language into interface operations. Existing single-agent a…

Multi-agent Reinforcement LearningMathematical Reasoning

Curriculum Learning for Safety Alignment

2026-05-25 · Sandeep Kumar, Virginia Smith, Chhavi Yadav arxiv

Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and exhibits poor out-of-distribution (OOD) generalisation. In this paper, we inv…

CLRecogEye : Curriculum Learning towards exploiting convolution features for Dynamic Iris Recognition

2025-11-26 · Geetanjali Sharma, Gaurav Jaswal, Aditya Nigam, Raghavendra Ramachandra arxiv

Iris authentication algorithms have achieved impressive recognition performance, making them highly promising for real-world applications such as border control, citizen identification, and both criminal investigations a…