paper-with-me

Papers

Enhancing RLHF with Human Gaze Modeling

2025-07-11 · Karim Galliamov, Ivan Titov, Ilya Pershin arxiv

Reinforcement Learning from Human Feedback (RLHF) aligns language models with human preferences but is computationally expensive. We explore two approaches that leverage human gaze modeling to enhance RLHF: (1) gaze-aware reward models and (2) gaze-based distribution of sparse rewards at token level. Our experiments demonstate that gaze-informed RLHF achieves faster convergence while maintaining or slightly improving performance, thus, reducing computational costs during policy optimization. These results show that human gaze provides a valuable and underused signal for policy optimization, pointing to a promising direction for improving RLHF efficiency.

📄 PDF Abstract BibTeX arXiv:2507.09016

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Gaze patterns predict preference and confidence in pairwise AI image evaluation

2026-03-25 · Nikolas Papadopoulos, Shreenithi Navaneethan, Sheng Bai, Ankur Samanta 외 arxiv

Preference learning methods, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on pairwise human judgments, yet little is known about the cognitive processes underly…

Reinforcement Learning

TransGOP: Transformer-Based Gaze Object Prediction

2024-02-21 · Binglu Wang, Chenxi Guo, Yang Jin, Haisheng Xia 외

Gaze object prediction aims to predict the location and category of the object that is watched by a human. Previous gaze object prediction works use CNN-based object detectors to predict the object's location. However, w…

Gaze EstimationObjectobject-detectionObject Detection+1

Multimodal Learning and Cognitive Processes in Radiology: MedGaze for Chest X-ray Scanpath Prediction

2024-06-28 · Akash Awasthi, Ngan Le, Zhigang Deng, Rishi Agrawal 외

Predicting human gaze behavior within computer vision is integral for developing interactive systems that can anticipate user attention, address fundamental questions in cognitive science, and hold implications for field…

Scanpath prediction

FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF

2024-12-20 · Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong, Roger Wattenhofer 외

In the era of increasing privacy concerns and demand for personalized experiences, traditional Reinforcement Learning with Human Feedback (RLHF) frameworks face significant challenges due to their reliance on centralized…

Privacy Preservingreinforcement-learningReinforcement Learning

Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot Interaction

2026-05-06 · Berk Sezer, Ali Görkem Küçük, Erol Şahin, Sinan Kalkan arxiv

While zero-shot appearance-based 3D gaze estimation offers significant cost-efficiency by directly mapping RGB images to gaze vectors, its reliability in Human-Robot Interaction (HRI) settings remains uncertain. Existing…

Gaze Estimation