paper-with-me

Papers

Scaffolding Reflection in Reinforcement Learning Framework for Confinement Escape Problem

2020-11-13 · Nishant Mohanty, Suresh Sundaram

In this paper, a novel Scaffolding Reflection in Reinforcement Learning (SR2L) is proposed for solving the confinement escape problem (CEP). In CEP, an evader's objective is to attempt escaping a confinement region patrolled by multiple pursuers. Meanwhile, the pursuers aim to reach and capture the evader. The inverse solution for pursuers to try and capture has been extensively studied in the literature. However, the problem of evaders escaping from the region is still an open issue. The SR2L employs an actor-critic framework to enable the evader to escape the confinement region. A time-varying state representation and reward function have been developed for proper convergence. The formulation uses the sensor information about the observable environment and prior knowledge of the confinement boundary. The conventional Independent Actor-Critic (IAC) method fails to converge due to sparseness in the reward. The effect becomes evident when operating in such a dynamic environment with a large area. In SR2L, along with the developed reward function, we use the scaffolding reflection method to improve the convergence significantly while increasing its efficiency. In SR2L, a motion planner is used as a scaffold for the actor-critic network to observe, compare and learn the action-reward pair. It enables the evader to achieve the required objective while using lesser resources and time. Convergence studies show that SR2L learns faster and converges to higher rewards as compared to IAC. Extensive Monte-Carlo simulations show that a SR2L consistently outperforms conventional IAC and the motion planner itself as the baselines.

📄 PDF Abstract BibTeX arXiv:2011.06764

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Quantifying the Potential to Escape Filter Bubbles: A Behavior-Aware Measure via Contrastive Simulation

2025-11-27 · Difu Feng, Qianqian Xu, Zitai Wang, Cong Hua 외 arxiv

Nowadays, recommendation systems have become crucial to online platforms, shaping user exposure by accurate preference modeling. However, such an exposure strategy can also reinforce users' existing preferences, leading …

Recommendation Systems

Hybrid LLM-Embedded Dialogue Agents for Learner Reflection: Designing Responsive and Theory-Driven Interactions

2026-02-24 · Paras Sharma, YuePing Sha, Janet Shufor Bih Epse Fofang, Brayden Yan 외 arxiv

Dialogue systems have long supported learner reflections, with theoretically grounded, rule-based designs offering structured scaffolding but often struggling to respond to shifts in engagement. Large Language Models (LL…

EscapeBench: Pushing Language Models to Think Outside the Box

2024-12-18 · Cheng Qian, Peixuan Han, Qinyu Luo, Bingxiang He 외

Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative adaptation in unfamiliar environments. To a…

Language ModelingLanguage Modelling

Self-evolving expertise in complex non-verifiable subject domains: dialogue as implicit meta-RL

2025-10-17 · Richard M. Bailey arxiv

So-called `wicked problems', those involving complex multi-dimensional settings, non-verifiable outcomes, heterogeneous impacts and a lack of single objectively correct answers, have plagued humans throughout history. Mo…

Reinforcement Learning

Pedagogical Reflections on the Holistic Cognitive Development (HCD) Framework and AI-Augmented Learning in Creative Computing

2025-11-10 · Anand Bhojan arxiv

This paper presents an expanded account of the Holistic Cognitive Development (HCD) framework for reflective and creative learning in computing education. The HCD framework integrates design thinking, experiential learni…