paper-with-me

홈 › Papers

Flow of Reasoning:Training LLMs for Divergent Problem Solving with Minimal Examples

2024-06-09 · Fangxu Yu, Lai Jiang, Haoqiang Kang, Shibo Hao, Lianhui Qin

The ability to generate diverse solutions to a given problem is a hallmark of human creativity. This divergent reasoning is also crucial for machines, enhancing their robustness and enabling them to assist humans in many applications such as scientific discovery. However, existing approaches to multi-step reasoning with large language models (LLMs) have mostly focused only on reasoning accuracy, without further discovering more diverse valid solutions. For example, supervised fine-tuning can improve LLM reasoning quality, but requires extensive supervised data to capture the full range of possible solutions. Reinforcement learning aims to find limited highest-reward solutions while neglecting the solution diversity. To fill this gap, we propose Flow of Reasoning (FoR), an efficient diversity-seeking LLM finetuning method aimed at improving reasoning quality and diversity with minimal data. FoR formulates multi-step LLM reasoning as a Markovian flow on a DAG-structured reasoning graph. This formulation allows us to incorporate and adapt principled GFlowNet approaches, for finetuning LLMs to sample diverse reasoning paths with probabilities proportional to the (unnormalized) reward of target problems. Extensive experiments show that, with limited training examples (e.g., 15 examples), FoR enables the discovery of diverse, creative, high-quality solutions, greatly outperforming a wide range of existing inference and training methods across five challenging puzzle-solving tasks, including BlocksWorld (embodied reasoning), Game24 (math puzzle solving), Rubik's Cube (spatial reasoning), 1D-ARC (abstraction reasoning), and PrOntoQA (logical reasoning). Code is available at https://github.com/Yu-Fangxu/FoR.

📄 PDF Abstract BibTeX arXiv:2406.05673

Code (1)

yu-fangxu/for 공식 구현 pytorch

Tasks

ARCDiversityLogical ReasoningMathematical ReasoningRubik's Cubescientific discoverySpatial Reasoning

Methods 이 논문이 사용한 방법론

Entropy Regularization 설명 없음
PPO Proximal Policy Optimization, or PPO, is a policy gradient method for reinforcement learning. The motivation was to have an algorithm with the data efficiency and reliable…

Similar Papers 제목 키워드 기반

From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations

2025-05-18 · Yuzhi Li, Haojun Xu, Fang Tian

Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable reasoning and generalization capabilities in video understanding; however, their application in video editing remains largely un…

Video EditingVideo Understanding

Reasoning Beyond the Obvious: Evaluating Divergent and Convergent Thinking in LLMs for Financial Scenarios

2025-07-24 · Zhuang Qiang Bok, Watson Wei Khong Chua arxiv

Most reasoning benchmarks for LLMs emphasize factual accuracy or step-by-step logic. In finance, however, professionals must not only converge on optimal decisions but also generate creative, plausible futures under unce…

DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning

2025-06-05 · Tanmay Parekh, Kartik Mehta, Ninareh Mehrabi, Kai-Wei Chang 외

Zero-shot Event Detection (ED), the task of identifying event mentions in natural language text without any training data, is critical for document understanding in specialized domains. Understanding the complex event on…

document understandingEvent DetectionTransfer Learning

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

2026-05-27 · Jihyeong Park, Ingeol Baek, Jeonghyun Park, Hwanhee Lee arxiv

Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generations, failing to capture how an agent reasons through iterative interac…

Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models

2024-07-03 · Haritz Puerto, Tilek Chubakov, Xiaodan Zhu, Harish Tayyar Madabushi 외

Requiring a Large Language Model to generate intermediary reasoning steps has been shown to be an effective way of boosting performance. In fact, it has been found that instruction tuning on these intermediary reasoning …

Language ModellingLarge Language Model