paper-with-me

홈 › Papers

OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings

2026-05-27 · Shadmehr Zaregarizi, Khashayar Yavari arxiv

Large language models (LLMs) have demonstrated promising capability in generating reward functions for deep reinforcement learning (DRL)-based building energy management. However, their potential to exhibit or exacerbate disparities in occupant comfort across heterogeneous demographic populations remains unexplored. We present OccuReward, a framework investigating how LLM-mediated reward design affects demographic equity. Our contribution is three-fold: the introduction of the Comfort Equity Index (CEI) as a novel feedback signal; a methodology for iterative, equity-aware LLM reward shaping; and a performance analysis of DRL agents under these refined objectives. Utilizing four empirically grounded occupant profiles from the ASHRAE Global Thermal Comfort Database II (13,440 votes), we deploy a Soft Actor-Critic agent in CityLearn v2. Our approach employs the Gemini API to generate reward function logic and weights--rather than performing per-step inference--across three refinement rounds. Results across 15 experimental runs reveal that elderly female occupants consistently experience the lowest satisfaction in initial rounds. By Round 3, equity-aware LLM refinement activates specific reward components that improve satisfaction for Young Males (+17.6%), Mid-aged Females (+28.2%), Health Sensitive (+53.8%), and Elderly Females (+567%), while simultaneously reducing energy costs by 3.2%. Our findings highlight that while reward-level intervention significantly improves equity, demographic disparities in AI-driven controllers persist, necessitating further research into algorithmic fairness in building systems.

📄 PDF Abstract BibTeX arXiv:2605.28168

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge

2024-10-02 · Xiefeng Wu

Q-shaping is an extension of Q-value initialization and serves as an alternative to reward shaping for incorporating domain knowledge to accelerate agent training, thereby improving sample efficiency by directly shaping …

Language ModelingLanguage ModellingLarge Language Model

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization

2025-07-09 · Xinjie Chen, Minpeng Liao, Guoxin Chen, Chengxi Li 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has recently advanced the reasoning capabilities of large language models (LLMs). While prior work has emphasized algorithmic design, data curation, and reward shapin…

Reinforcement LearningData Augmentation

PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management

2026-05-27 · Shadmehr Zaregarizi, Khashayar Yavari arxiv

Occupant comfort and grid-aware energy efficiency are competing objectives whose joint optimization depends critically on how reward functions are specified in deep reinforcement learning (DRL) controllers for buildings.…

Reinforcement Learning

Search-P1: Path-Centric Reward Shaping for Stable and Efficient Agentic RAG Training

2026-02-26 · Tianle Xia, Ming Xu, Lingxiang Hu, Yiding Sun 외 arxiv

Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by incorporating external knowledge, yet traditional single-round retrieval struggles with complex multi-step reasoning. Agentic RAG addresses th…

Segmentation Analysis in Human Centric Cyber-Physical Systems using Graphical Lasso

2018-10-24 · Hari Prasanna Das, Ioannis C. Konstantakopoulos, Aummul Baneen Manasawala, Tanya Veeravalli 외

A generalized gamification framework is introduced as a form of smart infrastructure with potential to improve sustainability and energy efficiency by leveraging humans-in-the-loop strategy. The proposed framework enable…

Decision MakingSegmentation