paper-with-me

홈 › Papers

WildReward: Learning Reward Models from In-the-Wild Human Interactions

2026-02-09 · Hao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao, Lei Hou, Juanzi Li arxiv

Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs. With the widespread deployment of LLMs, in-the-wild interactions have emerged as a rich source of implicit reward signals. This raises the question: Can we develop reward models directly from in-the-wild interactions? In this work, we explore this possibility by adopting WildChat as an interaction source and proposing a pipeline to extract reliable human feedback, yielding 186k high-quality instances for training WildReward via ordinal regression directly on user feedback without preference pairs. Extensive experiments demonstrate that WildReward achieves comparable or even superior performance compared to conventional reward models, with improved calibration and cross-sample consistency. We also observe that WildReward benefits directly from user diversity, where more users yield stronger reward models. Finally, we apply WildReward to online DPO training and observe significant improvements across various tasks. Code and data are released at https://github.com/THU-KEG/WildReward.

📄 PDF Abstract BibTeX arXiv:2602.08829

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

2024-06-07 · Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Faeze Brahman 외

We introduce WildBench, an automated evaluation framework designed to benchmark large language models (LLMs) using challenging, real-world user queries. WildBench consists of 1,024 tasks carefully selected from over one …

BenchmarkingChatbot

DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies

2025-05-12 · Tony Tao, Mohan Kumar Srirama, Jason Jingzhou Liu, Kenneth Shaw 외

Large-scale, diverse robot datasets have emerged as a promising path toward enabling dexterous manipulation policies to generalize to novel environments, but acquiring such datasets presents many challenges. While teleop…

Shaping embodied agent behavior with activity-context priors from egocentric video

2021-10-14 · NeurIPS 2021 12 · Tushar Nagarajan, Kristen Grauman

Complex physical tasks entail a sequence of object interactions, each with its own preconditions -- which can be difficult for robotic agents to learn efficiently solely through their own experience. We introduce an appr…

WildQA: In-the-Wild Video Question Answering

2022-09-14 · Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo 외

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the "in the wild" settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding d…

Evidence SelectionQuestion AnsweringVideo Question AnsweringVideo Understanding

Facilitating human-wildlife cohabitation through conflict prediction

2021-09-22 · Susobhan Ghosh, Pradeep Varakantham, Aniket Bhatkhande, Tamanna Ahmad 외

With increasing world population and expanded use of forests as cohabited regions, interactions and conflicts with wildlife are increasing, leading to large-scale loss of lives (animal and human) and livelihoods (economi…

Prediction