paper-with-me

Papers

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence

2025-05-30 · Guiyang Hou, Xing Gao, Yuchuan Wu, Xiang Huang, Wenqi Zhang, Zhe Zheng, Yongliang Shen, Jialu Du, Fei Huang, Yongbin Li, Weiming Lu

Recently, Large Language Models (LLMs) have made significant progress in IQ-related domains that require careful thinking, such as mathematics and coding. However, enhancing LLMs' cognitive development in social domains, particularly from a post-training perspective, remains underexplored. Recognizing that the social world follows a distinct timeline and requires a richer blend of cognitive modes (from intuitive reactions (System 1) and surface-level thinking to deliberate thinking (System 2)) than mathematics, which primarily relies on System 2 cognition (careful, step-by-step reasoning), we introduce Temporal-aware Hierarchical Cognitive Reinforcement Learning (TimeHC-RL) for enhancing LLMs' social intelligence. In our experiments, we systematically explore improving LLMs' social intelligence and validate the effectiveness of the TimeHC-RL method, through five other post-training paradigms and two test-time intervention paradigms on eight datasets with diverse data patterns. Experimental results reveal the superiority of our proposed TimeHC-RL method compared to the widely adopted System 2 RL method. It gives the 7B backbone model wings, enabling it to rival the performance of advanced models like DeepSeek-R1 and OpenAI-O3. Additionally, the systematic exploration from post-training and test-time interventions perspectives to improve LLMs' social intelligence has uncovered several valuable insights.

📄 PDF Abstract BibTeX arXiv:2505.24500

Code (1)

zju-real/timehc-rl 공식 구현

Similar Papers 제목 키워드 기반

Discounting and Drug Seeking in Biological Hierarchical Reinforcement Learning

2025-06-05 · Vardhan Palod, Pranav Mahajan, Veeky Baths, Boris S. Gutkin

Despite a strong desire to quit, individuals with long-term substance use disorder (SUD) often struggle to resist drug use, even when aware of its harmful consequences. This disconnect between knowledge and compulsive be…

Hierarchical Reinforcement Learningreinforcement-learningReinforcement Learning

Delay-Empowered Causal Hierarchical Reinforcement Learning

2026-05-12 · Chenran Zhao, Dianxi Shi, Haotian Wang, Mengzhu Wang 외 arxiv

Many real-world tasks involve delayed effects, where the outcomes of actions emerge after varying time lags. Existing delay-aware reinforcement learning methods often rely on state augmentation, prior knowledge of delay …

Hierarchical Reinforcement Learning

Semantic RL with Action Grammars: Data-Efficient Learning of Hierarchical Task Abstractions

2019-07-29 · Robert Tjarko Lange, Aldo Faisal

Hierarchical Reinforcement Learning algorithms have successfully been applied to temporal credit assignment problems with sparse reward signals. However, state-of-the-art algorithms require manual specification of sub-ta…

Hierarchical Reinforcement LearningLogical Reasoningreinforcement-learningReinforcement Learning+2

Intelligent problem-solving as integrated hierarchical reinforcement learning

2022-08-18 · Manfred Eppe, Christian Gumbsch, Matthias Kerzel, Phuong D. H. Nguyen 외

According to cognitive psychology and related disciplines, the development of complex problem-solving behaviour in biological agents depends on hierarchical cognitive mechanisms. Hierarchical reinforcement learning is a …

Hierarchical Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

MotionHiFlow: Text-to-motion via hierarchical flow matching

2026-04-25 · Heng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang 외 arxiv

Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce comple…