paper-with-me

Papers

Wait, That Feels Familiar: Learning to Extrapolate Human Preferences for Preference Aligned Path Planning

2023-09-18 · Haresh Karnan, Elvin Yang, Garrett Warnell, Joydeep Biswas, Peter Stone

Autonomous mobility tasks such as lastmile delivery require reasoning about operator indicated preferences over terrains on which the robot should navigate to ensure both robot safety and mission success. However, coping with out of distribution data from novel terrains or appearance changes due to lighting variations remains a fundamental problem in visual terrain adaptive navigation. Existing solutions either require labor intensive manual data recollection and labeling or use handcoded reward functions that may not align with operator preferences. In this work, we posit that operator preferences for visually novel terrains, which the robot should adhere to, can often be extrapolated from established terrain references within the inertial, proprioceptive, and tactile domain. Leveraging this insight, we introduce Preference extrApolation for Terrain awarE Robot Navigation, PATERN, a novel framework for extrapolating operator terrain preferences for visual navigation. PATERN learns to map inertial, proprioceptive, tactile measurements from the robots observations to a representation space and performs nearest neighbor search in this space to estimate operator preferences over novel terrains. Through physical robot experiments in outdoor environments, we assess PATERNs capability to extrapolate preferences and generalize to novel terrains and challenging lighting conditions. Compared to baseline approaches, our findings indicate that PATERN robustly generalizes to diverse terrains and varied lighting conditions, while navigating in a preference aligned manner.

📄 PDF Abstract BibTeX arXiv:2309.09912

Code (0)

등록된 구현이 없습니다.

Tasks

NavigateRobot NavigationVisual Navigation

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Should Machine Learning Models Report to Us When They Are Clueless?

2022-03-23 · Roozbeh Yousefzadeh, Xuenan Cao

The right to AI explainability has consolidated as a consensus in the research community and policy-making. However, a key component of explainability has been missing: extrapolation, which describes the extent to which …

BIG-bench Machine Learning

Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction

2026-07-16 · Ryuichi Sumida, Mao Saeki, Masaki Eguchi, Sadahiro Yoshikawa 외 arxiv

As conversational AI systems are designed for repeated use, a central question is how a series of interactions becomes a relationship. We present a longitudinal multimodal study of a memory-augmented conversational agent…

Treats or Affection? Understanding Reward Preferences in Indian Free-ranging Dogs for Bonding with Humans

2025-05-05 · Srijaya Nandi, Aesha Lahiri, Tuhin Subhra Pal, Anamitra Roy 외

Free-ranging dogs constitute approximately 80% of the global dog population. These dogs are freely breeding and live without direct human supervision, making them an ideal model system for studying dog-human relationship…

Accounting for Human Learning when Inferring Human Preferences

2020-11-11 · Harry Giles, Lawrence Chan

Inverse reinforcement learning (IRL) is a common technique for inferring human preferences from data. Standard IRL techniques tend to assume that the human demonstrator is stationary, that is that their policy $\pi$ does…

Accounting for Human Learning when Inferring Human Preferences

2020-10-15 · NeurIPS Workshop HAMLETS 2020 12 · Anonymous

Inverse reinforcement learning (IRL) is a common technique for inferring human preferences from data. Standard IRL techniques tend to assume that the human demonstrator is stationary, that is that their policy $\pi$ does…