paper-with-me

홈 › Papers

Generalizable LLM Learning of Graph Synthetic Data with Reinforcement Learning

2025-06-01 · Yizhuo Zhang, Heng Wang, Shangbin Feng, Zhaoxuan Tan, Xinyun Liu, Yulia Tsvetkov

Previous research has sought to enhance the graph reasoning capabilities of LLMs by supervised fine-tuning on synthetic graph data. While these led to specialized LLMs better at solving graph algorithm problems, we don't need LLMs for shortest path: we need generalization from synthetic graph data to real-world tasks with implicit graph structures. In this work, we propose to unlock generalizable learning of graph synthetic data with reinforcement learning. We first design solution-based and process-based rewards for synthetic graph problems: instead of rigid memorizing response patterns in direct fine-tuning, we posit that RL would help LLMs grasp the essentials underlying graph reasoning and alleviate overfitting. We employ RL algorithms such as GRPO and DPO, aligning both off-the-shelf LLMs and LLMs fine-tuned on synthetic graph data. We then compare them against existing settings on both in-domain synthetic tasks and out-of-domain real-world tasks with implicit graph structures such as multi-hop QA, structured planning, and more. Extensive experiments demonstrate that our RL recipe leads to statistically significant improvement on 5 datasets, with an average gain of 12.9\% over baseline settings. Further analysis reveals that process-based rewards consistently outperform solution-based rewards, mixing synthetic and real-world task data yields potential gains, while compositionality and explainable intermediate steps remains a critical challenge even after RL.

📄 PDF Abstract BibTeX arXiv:2506.00845

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

DPO 설명 없음

Similar Papers 제목 키워드 기반

Learning Generalizable Hand-Object Tracking from Synthetic Demonstrations

2025-12-22 · Yinhuai Wang, Runyi Yu, Hok Wai Tsui, Xiaoyi Lin 외 arxiv

We present a system for learning generalizable hand-object tracking controllers purely from synthetic data, without requiring any human demonstrations. Our approach makes two key contributions: (1) HOP, a Hand-Object Pla…

Reinforcement LearningObject Tracking

Adversarial Attack on Graph Structured Data

2018-06-06 · ICML 2018 7 · Hanjun Dai, Hui Li, Tian Tian, Xin Huang 외

Deep learning on graph structures has shown exciting results in various applications. However, few attentions have been paid to the robustness of such models, in contrast to numerous research work for image or text adver…

Adversarial AttackGraph Neural NetworkReinforcement Learning

Distilling Reinforcement Learning into Single-Batch Datasets

2025-08-12 · Connor Wilhelm, Dan Ventura arxiv

Dataset distillation compresses a large dataset into a small synthetic dataset such that learning on the synthetic dataset approximates learning on the original. Training on the distilled dataset can be performed in as l…

Reinforcement LearningAtari Games

Generalizable Resource Allocation in Stream Processing via Deep Reinforcement Learning

2019-11-19 · Xiang Ni, Jing Li, Mo Yu, Wang Zhou 외

This paper considers the problem of resource allocation in stream processing, where continuous data flows must be processed in real time in a large distributed system. To maximize system throughput, the resource allocati…

DecoderDeep Reinforcement LearningGraph Embeddinggraph partitioning+3

Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks

2025-02-06 · Niccolò Grillo, Andrea Toccaceli, Joël Mathys, Benjamin Estermann 외

Despite incredible progress, many neural architectures fail to properly generalize beyond their training distribution. As such, learning to reason in a correct and generalizable way is one of the current fundamental chal…

Inductive Bias