paper-with-me

홈 › Papers

Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective

2026-03-09 · Liyuan Mao, Le Yu, Jing Zhou, Chujie Zheng, Bowen Yu, Chang Gao, Shixuan Liu, An Yang, Weinan Zhang, JunYang Lin arxiv

In this work, we reveal that Large Language Models (LLMs) possess intrinsic behavioral plasticity-akin to chameleons adapting their coloration to environmental cues-that can be exposed through token-conditional generation and stabilized via reinforcement learning. Specifically, by conditioning generation on carefully selected token prefixes sampled from responses exhibiting desired behaviors, LLMs seamlessly adapt their behavioral modes at inference time (e.g., switching from step-by-step reasoning to direct answering) without retraining. Based on this insight, we propose Token-Conditioned Reinforcement Learning (ToCoRL), a principled framework that leverages RL to internalize this chameleon-like plasticity, transforming transient inference-time adaptations into stable and learnable behavioral patterns. ToCoRL guides exploration with token-conditional generation and keep enhancing exploitation, enabling emergence of appropriate behaviors. Extensive experiments show that ToCoRL enables precise behavioral control without capability degradation. Notably, we show that large reasoning models, while performing strongly on complex mathematics, can be effectively adapted to excel at factual question answering, which was a capability previously hindered by their step-by-step reasoning patterns.

📄 PDF Abstract BibTeX arXiv:2603.08398

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningQuestion Answering

Similar Papers 제목 키워드 기반

A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models

2026-07-06 · Nima Eshraghi, Lovedeep Gondara, Yuqing Huang, Sagarika Suresh 외 arxiv

Activation steering via sparse autoencoders (SAEs) enables behavioral control of large language models without task-specific fine-tuning, but standard methods apply the steering signal at every generated token, incurring…

One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers

2025-06-12 · Diana Abagyan, Alejandro R. Salamanca, Andres Felipe Cruz-Salinas, Kris Cao 외

Pretraining massively multilingual Large Language Models (LLMs) for many languages at once is challenging due to limited model capacity, scarce high-quality data, and compute constraints. Moreover, the lack of language c…

All

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff

2026-06-07 · Runze Liu, Jiashun Liu, Xu Wan, Yuqian Fu 외 arxiv

Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training. SFT is expected to provide a useful behavioral prior for RL to further enh…

Reinforcement Learning

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

2026-08-24 · Bin Dou, Junru Zhang, Zhaoyi Yuan, Wuliang Huang 외 arxiv

User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for r…

Representation Learning

M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation

2026-03-27 · Jihoon Jeong arxiv

We introduce M-CARE (Model Clinical Assessment and Reporting for Evaluation), a clinical case report framework for AI model behavioral disorders adapted from human medicine. M-CARE provides a 13-section report format, a …