paper-with-me

홈 › Papers

Let's Be Self-generated via Step by Step: A Curriculum Learning Approach to Automated Reasoning with Large Language Models

2024-10-29 · Kangyang Luo, Zichen Ding, Zhenmin Weng, Lingfeng Qiao, Meng Zhao, Xiang Li, Di Yin, Jinlong Shu

While Chain of Thought (CoT) prompting approaches have significantly consolidated the reasoning capabilities of large language models (LLMs), they still face limitations that require extensive human effort or have performance needs to be improved. Existing endeavors have focused on bridging these gaps; however, these approaches either hinge on external data and cannot completely eliminate manual effort, or they fall short in effectively directing LLMs to generate high-quality exemplary prompts. To address the said pitfalls, we propose a novel prompt approach for automatic reasoning named \textbf{LBS3}, inspired by curriculum learning which better reflects human learning habits. Specifically, LBS3 initially steers LLMs to recall easy-to-hard proxy queries that are pertinent to the target query. Following this, it invokes a progressive strategy that utilizes exemplary prompts stemmed from easy-proxy queries to direct LLMs in solving hard-proxy queries, enabling the high-quality of the proxy solutions. Finally, our extensive experiments in various reasoning-intensive tasks with varying open- and closed-source LLMs show that LBS3 achieves strongly competitive performance compared to the SOTA baselines.

📄 PDF Abstract BibTeX arXiv:2410.21728

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning

2024-10-09 · Xiyao Wang, Linfeng Song, Ye Tian, Dian Yu 외

Monte Carlo Tree Search (MCTS) has recently emerged as a powerful technique for enhancing the reasoning capabilities of LLMs. Techniques such as SFT or DPO have enabled LLMs to distill high-quality behaviors from MCTS, i…

Mathematical Reasoning

See Further When Clear: Curriculum Consistency Model

2024-12-09 · CVPR 2025 1 · Yunpeng Liu, Boxiao Liu, Yi Zhang, Xingzhong Hou 외

Significant advances have been made in the sampling efficiency of diffusion models and flow matching models, driven by Consistency Distillation (CD), which trains a student model to mimic the output of a teacher model at…

model

Curriculum Learning Meets Weakly Supervised Modality Correlation Learning

2022-12-15 · Sijie Mai, Ya Sun, Haifeng Hu

In the field of multimodal sentiment analysis (MSA), a few studies have leveraged the inherent modality correlation information stored in samples for self-supervised learning. However, they feed the training pairs in a r…

Multimodal Sentiment AnalysisSelf-Supervised LearningSentiment Analysis

Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs

2025-09-22 · Mariam Mahran, Katharina Simbeck arxiv

Large Language Models (LLMs) are increasingly used for educational support, yet their response quality varies depending on the language of interaction. This paper presents an automated multilingual pipeline for generatin…

SAGE: Multi-Agent Self-Evolution for LLM Reasoning

2026-03-16 · Yulin Peng, Xinxin Zhu, Chenxing Wei, Nianbo Zeng 외 arxiv

Reinforcement learning with verifiable rewards improves reasoning in large language models (LLMs), but many methods still rely on large human-labeled datasets. While self-play reduces this dependency, it often lacks expl…

Reinforcement Learning