Time-Efficient Reward Learning via Visually Assisted Cluster Ranking
One of the most successful paradigms for reward learning uses human feedback in the form of comparisons. Although these methods hold promise, human comparison labeling is expensive and time consuming, constituting a major bottleneck to their broader applicability. Our insight is that we can greatly improve how effectively human time is used in these approaches by batching comparisons together, rather than having the human label each comparison individually. To do so, we leverage data dimensionality-reduction and visualization techniques to provide the human with a interactive GUI displaying the state space, in which the user can label subportions of the state space. Across some simple Mujoco tasks, we show that this high-level approach holds promise and is able to greatly increase the performance of the resulting agents, provided the same amount of human labeling time.
Code (0)
등록된 구현이 없습니다.
Tasks
Dimensionality ReductionMuJoCoSimilar Papers 제목 키워드 기반
Clustered Policy Decision Ranking
Policies trained via reinforcement learning (RL) are often very complex even for simple tasks. In an episode with n time steps, a policy will make n decisions on actions to take, many of which may appear non-intuitive to…
Fault localizationReinforcement Learning (RL)State Drug Policy Effectiveness: Comparative Policy Analysis of Drug Overdose Mortality
Opioid overdose rates have reached an epidemic level and state-level policy innovations have followed suit in an effort to prevent overdose deaths. State-level drug law is a set of policies that may reinforce or undermin…
ClusteringTime SeriesTime Series AnalysisTime Series RegressioniVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy during the inference phase remains underexpl…
Reinforcement LearningVisual LocalizationVisual GroundingWelfare estimations from imagery. A test of domain experts ability to rate poverty from visual inspection of satellite imagery
The present study uses domain experts to estimate welfare levels and indicators from high-resolution satellite imagery. We use the wealth quintiles from the 2015 Tanzania DHS dataset as ground truth data. We analyse the …
MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG
Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic knowledge bases and answer related questions. Existing multimodal Retr…
Visual Question AnsweringAnswer Generation