paper-with-me

홈 › Papers

GFlowNet Fine-tuning for Diverse Correct Solutions in Mathematical Reasoning Tasks

2024-10-26 · Ryoichi Takase, Masaya Tsunokake, Yuta Tsuchiya, Shota Inuzuka

Mathematical reasoning problems are among the most challenging, as they typically require an understanding of fundamental laws to solve. The laws are universal, but the derivation of the final answer changes depending on how a problem is approached. When training large language models (LLMs), learning the capability of generating such multiple solutions is essential to accelerate their use in mathematical education. To this end, we train LLMs using generative flow network (GFlowNet). Different from reward-maximizing reinforcement learning (RL), GFlowNet fine-tuning seeks to find diverse solutions by training the LLM whose distribution is proportional to a reward function. In numerical experiments, we evaluate GFlowNet fine-tuning and reward-maximizing RL in terms of accuracy and diversity. The results show that GFlowNet fine-tuning derives correct final answers from diverse intermediate reasoning steps, indicating the improvement of the capability of alternative solution generation.

📄 PDF Abstract BibTeX arXiv:2410.20147

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityMathematical ReasoningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Routing by Reaching: Composition of Pre-trained GFlowNets for Multi-Objective Generation

2026-02-25 · Seokwon Yoon, Youngbin Choi, Seunghyuk Cho, Seungbeom Lee 외 arxiv

Generative Flow Networks (GFlowNets) learn to sample diverse candidates in proportion to a reward function, making them well-suited for scientific discovery, where exploring multiple promising solutions is crucial. Furth…

Multi-Objective GFlowNets

2022-10-23 · Moksh Jain, Sharath Chandra Raparthy, Alex Hernandez-Garcia, Jarrid Rector-Brooks 외

We study the problem of generating diverse candidates in the context of Multi-Objective Optimization. In many applications of machine learning such as drug discovery and material design, the goal is to generate candidate…

Active LearningDiversityDrug Discovery

Enhancing Solution Efficiency in Reinforcement Learning: Leveraging Sub-GFlowNet and Entropy Integration

2024-10-01 · Siyi He

Traditional reinforcement learning often struggles to generate diverse, high-reward solutions, especially in domains like drug design and black-box function optimization. Markov Chain Monte Carlo (MCMC) methods provide a…

Computational EfficiencyDiversityDrug Design

GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networks

2025-03-09 · CVPR 2025 1 · Haoqiang Kang, Enna Sachdeva, Piyush Gupta, Sangjae Bae 외

Vision-Language Models (VLMs) have recently shown promising advancements in sequential decision-making tasks through task-specific fine-tuning. However, common fine-tuning methods, such as Supervised Fine-Tuning (SFT) an…

Card GamesDiversityReinforcement Learning (RL)Sequential Decision Making

Flow of Reasoning:Training LLMs for Divergent Problem Solving with Minimal Examples

2024-06-09 · Fangxu Yu, Lai Jiang, Haoqiang Kang, Shibo Hao 외

The ability to generate diverse solutions to a given problem is a hallmark of human creativity. This divergent reasoning is also crucial for machines, enhancing their robustness and enabling them to assist humans in many…

ARCDiversityLogical ReasoningMathematical Reasoning+3