paper-with-me

홈 › Papers

Can the Environment Speak for Itself? $T^{2}$-GRPO: A Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents

2026-06-07 · Yutong Song, Jiang Wu, Pengfei Zhang, Wenjun Huang, Honghui Xu, Nikil Dutt, Amir M. Rahmani arxiv

Optimizing large language models (LLMs) for long-horizon caregiver agents requires balancing delayed task objectives with immediate environment dynamics, such as patient distress and resistance. In dementia care, this balance is especially difficult: trajectory level rewards are too sparse for turn level credit assignment, while external LLM-based evaluators are costly and can misread fragmented or indirect patient responses. To address this issue, we propose \textbf{T}urn-\textbf{T}rajectory \textbf{G}roup \textbf{R}elative \textbf{P}olicy \textbf{O}ptimization (\textbf{T$^{2}$-GRPO}), a framework that decouples caregiver RL into two normalized reward horizons and enforces safety through a binary hard veto. $T^2$-GRPO derives dense turn-level rewards directly from environment state transitions, measuring changes in patient distress and resistance from a frozen dementia patient simulator. These environment-grounded rewards are combined with trajectory-level evaluations through independent centered-rank normalization, which preserves heterogeneous reward signals and mitigates reward collapse. Extensive experiments on dementia caregivers show that T $^{2}$-GRPO outperforms competitive baselines, indicating a substantial improvement for emotionally sensitive caregiver scenarios that effectively handles immediate patient feedback, long-term care outcomes, and safety constraints.

📄 PDF Abstract BibTeX arXiv:2606.08875

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization

2026-01-23 · Peiji Li, Linyang Li, Handa Sun, Wenjin Mai 외 arxiv

Large language models have demonstrated strong reasoning capabilities in complex tasks through tool integration, which is typically framed as a Markov Decision Process and optimized with trajectory-level RL algorithms su…

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

2026-02-06 · Yunze Tong, Mushui Liu, Canyu Zhao, Didi Zhu 외 arxiv

Deploying GRPO on Flow Matching models has proven effective for text-to-image generation. However, existing paradigms typically propagate an outcome-based reward to all preceding denoising steps without distinguishing th…

Text-to-Image Generation

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

2026-07-30 · Xilin Jiang, Riki Shimizu, Sukru Samet Dindar, Junkai Wu 외 arxiv

Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speake…

Reinforcement Learning

Drowning in Routine: Signal Dilution in Multi-Turn Agent Training

2026-06-20 · Yann Pernot, Vi Retault arxiv

Multi-turn agents interleave consequential decisions with routine execution: some actions change the downstream return distribution, while others are necessary but reward-equivalent. The cost of trajectory-level credit a…

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

2026-05-31 · Liyang Li, Muzhi Zhu, Zhiyue Zhao, Hengyu Zhao 외 arxiv

Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has largely been studied as passive understanding of pre-collected observa…