paper-with-me

Papers

PAWS: Preference Learning with Advantage-Weighted Segments

2026-06-10 · Aleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li, Serge Thilges, Huy Le, Tai Hoang, Rania Rayyes, Gerhard Neumann arxiv

Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods typically train utility functions on trajectory or segment-level preferences while relying on per-step utility estimates during policy optimization. This training and inference mismatch induces a distribution shift that severely degrades temporal credit assignment and limits policy learning. We analyze this issue and propose PAWS, a segment-based preference learning method that performs policy updates directly using segment-level advantage functions. By aligning utility training with policy optimization, PAWS preserves trajectory-level preference information and avoids unreliable per-step learning signals. Experiments on simulated robotic manipulation and locomotion tasks demonstrate that PAWS consistently outperforms existing PbRL approaches, highlighting the importance of distribution-consistent preference learning.

📄 PDF Abstract BibTeX arXiv:2606.11982

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data

2023-02-28 · Sangwoo Mo, Jong-Chyi Su, Chih-Yao Ma, Mido Assran 외

Semi-supervised learning aims to train a model using limited labels. State-of-the-art semi-supervised methods for image classification such as PAWS rely on self-supervised representations learned with large-scale unlabel…

Density Estimationimage-classificationImage ClassificationRepresentation Learning

RuPAWS: A Russian Adversarial Dataset for Paraphrase Identification

2022-06-01 · LREC 2022 6 · Nikita Martynov, Irina Krotova, Varvara Logacheva, Alexander Panchenko 외

Paraphrase identification task can be easily challenged by changing word order, e.g. as in “Can a good person become bad?”. While for English this problem was tackled by the PAWS dataset (Zhang et al., 2019), datasets fo…

Paraphrase Identification

PAWS: A Multi-lingual Parallel Treebank with Anaphoric Relations

2018-06-01 · WS 2018 6 · Anna Nedoluzhko, Michal Nov{\'a}k, Maciej Ogrodniczuk

We present PAWS, a multi-lingual parallel treebank with coreference annotation. It consists of English texts from the Wall Street Journal translated into Czech, Russian and Polish. In addition, the texts are syntacticall…

Coreference ResolutionMachine Translation

Learning Optimal Advantage from Preferences and Mistaking it for Reward

2023-10-03 · W. Bradley Knox, Stephane Hatgis-Kessell, Sigurdur Orn Adalgeirsson, Serena Booth 외

We consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most recent work assumes that human preferences…

PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification

2019-08-30 · IJCNLP 2019 11 · Yinfei Yang, Yuan Zhang, Chris Tar, Jason Baldridge

Most existing work on adversarial data generation focuses on English. For example, PAWS (Paraphrase Adversaries from Word Scrambling) consists of challenging English paraphrase identification pairs from Wikipedia and Quo…

Paraphrase IdentificationSentence