paper-with-me

홈 › Papers

GFlowNets with Human Feedback

2023-05-11 · Yinchuan Li, Shuang Luo, Yunfeng Shao, Jianye Hao

We propose the GFlowNets with Human Feedback (GFlowHF) framework to improve the exploration ability when training AI models. For tasks where the reward is unknown, we fit the reward function through human evaluations on different trajectories. The goal of GFlowHF is to learn a policy that is strictly proportional to human ratings, instead of only focusing on human favorite ratings like RLHF. Experiments show that GFlowHF can achieve better exploration ability than RLHF.

📄 PDF Abstract BibTeX arXiv:2305.07036

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets

2024-10-19 · Oh Joon Kwon, Daiki E. Matsunaga, Kee-Eung Kim

A critical component of the current generation of language models is preference alignment, which aims to precisely control the model's behavior to meet human needs and values. The most notable among such methods is Reinf…

Diversity

Stochastic Generative Flow Networks

2023-02-19 · Ling Pan, Dinghuai Zhang, Moksh Jain, Longbo Huang 외

Generative Flow Networks (or GFlowNets for short) are a family of probabilistic agents that learn to sample complex combinatorial structures through the lens of "inference as control". They have shown great potential in …

Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets

2025-04-28 · Adam Younsi, Abdalgader Abubaker, Mohamed El Amine Seddik, Hakim Hacid 외

Achieving both accuracy and diverse reasoning remains challenging for Large Language Models (LLMs) in complex domains like mathematics. A key bottleneck is evaluating intermediate reasoning steps to guide generation with…

Data AugmentationDiversityMathMathematical Reasoning

Start Small: Training Controllable Game Level Generators without Training Data by Learning at Multiple Sizes

2022-09-29 · Yahia Zakaria, Magda Fayek, Mayada Hadhoud

A level generator is a tool that generates game levels from noise. Training a generator without a dataset suffers from feedback sparsity, since it is unlikely to generate a playable level via random exploration. A common…

DiversitySokoban

Learning to Scale Logits for Temperature-Conditional GFlowNets

2023-10-04 · Minsu Kim, Joohwan Ko, Taeyoung Yun, Dinghuai Zhang 외

GFlowNets are probabilistic models that sequentially generate compositional structures through a stochastic policy. Among GFlowNets, temperature-conditional GFlowNets can introduce temperature-based controllability for e…