paper-with-me

Papers

Off-Policy Selection for Initiating Human-Centric Experimental Design

2024-10-26 · Ge Gao, Xi Yang, Qitong Gao, Song Ju, Miroslav Pajic, Min Chi

In human-centric tasks such as healthcare and education, the heterogeneity among patients and students necessitates personalized treatments and instructional interventions. While reinforcement learning (RL) has been utilized in those tasks, off-policy selection (OPS) is pivotal to close the loop by offline evaluating and selecting policies without online interactions, yet current OPS methods often overlook the heterogeneity among participants. Our work is centered on resolving a pivotal challenge in human-centric systems (HCSs): how to select a policy to deploy when a new participant joining the cohort, without having access to any prior offline data collected over the participant? We introduce First-Glance Off-Policy Selection (FPS), a novel approach that systematically addresses participant heterogeneity through sub-group segmentation and tailored OPS criteria to each sub-group. By grouping individuals with similar traits, FPS facilitates personalized policy selection aligned with unique characteristics of each participant or group of participants. FPS is evaluated via two important but challenging applications, intelligent tutoring systems and a healthcare application for sepsis treatment and intervention. FPS presents significant advancement in enhancing learning outcomes of students and in-hospital care outcomes.

📄 PDF Abstract BibTeX arXiv:2410.20017

Code (0)

등록된 구현이 없습니다.

Tasks

Experimental DesignReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Handoff Design in User-Centric Cell-Free Massive MIMO Networks Using DRL

2025-07-28 · Hussein A. Ammar, Raviraj Adve, Shahram Shahbazpanahi, Gary Boudreau 외 arxiv

In the user-centric cell-free massive MIMO (UC-mMIMO) network scheme, user mobility necessitates updating the set of serving access points to maintain the user-centric clustering. Such updates are typically performed thr…

Reinforcement Learning

Towards Monotonic Improvement in In-Context Reinforcement Learning

2025-09-27 · Wenhao Zhang, Shao Zhang, Xihuai Wang, Yang Li 외 arxiv

In-Context Reinforcement Learning (ICRL) has emerged as a promising paradigm for developing agents that can rapidly adapt to new tasks by leveraging past experiences as context, without updating their parameters. Recent …

Reinforcement Learning

EgoVLM: Policy Optimization for Egocentric Video Understanding

2025-06-03 · Ashwin Vinod, Shrey Pandit, Aditya Vavre, Linshen Liu

Emerging embodied AI applications, such as wearable cameras and autonomous agents, have underscored the need for robust reasoning from first person video streams. We introduce EgoVLM, a vision-language model specifically…

EgoSchemaQuestion Answeringreinforcement-learningReinforcement Learning+2

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

2026-07-01 · Jihyeok Jung, Jeewu Lee, Sanghyeop Kim, Chanhee Han 외 arxiv

Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation. However, understanding first-…

Scene Understanding

HOPE: Human-Centric Off-Policy Evaluation for E-Learning and Healthcare

2023-02-18 · Ge Gao, Song Ju, Markel Sanz Ausin, Min Chi

Reinforcement learning (RL) has been extensively researched for enhancing human-environment interactions in various human-centric tasks, including e-learning and healthcare. Since deploying and evaluating policies online…

Off-policy evaluationReinforcement Learning (RL)