paper-with-me

홈 › Papers

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

2026-08-17 · Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu arxiv

Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Trained only on forward perturbation-response prediction, PertMind improved response inference in unseen cellular contexts while retaining general language capabilities. It also transferred without task-specific post-training to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generated biological profiles that supported competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.

📄 PDF Abstract BibTeX arXiv:2608.16419

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

2025-03-07 · Hengguang Zhou, Xirui Li, Ruochen Wang, Minhao Cheng 외

Recently DeepSeek R1 demonstrated how reinforcement learning with simple rule-based incentives can enable autonomous development of complex reasoning in large language models, characterized by the "aha moment", in which …

Multimodal Reasoningreinforcement-learningReinforcement LearningVisual Reasoning

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

2026-07-14 · Xinyu Tang, Gangqiang Cao, Yurou Liu, Yuliang Zhan 외 arxiv

Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constr…

Reinforcement Learning

Emergent time-keeping mechanisms in a deep reinforcement learning agent performing an interval timing task

2025-08-06 · Amrapali Pednekar, Alvaro Garrido, Pieter Simoens, Yara Khaluf arxiv

Drawing parallels between Deep Artificial Neural Networks (DNNs) and biological systems can aid in understanding complex biological mechanisms that are difficult to disentangle. Temporal processing, an extensively resear…

Reinforcement Learning

First Return, Entropy-Eliciting Explore

2025-07-09 · Tianyu Zheng, Tianshun Xing, Qingshui Gu, Taoran Liang 외 arxiv

Reinforcement Learning from Verifiable Rewards (RLVR) improves the reasoning abilities of Large Language Models (LLMs) but it struggles with unstable exploration. We propose FR3E (First Return, Entropy-Eliciting Explore)…

Reinforcement LearningMathematical Reasoning

Incorporating Pragmatic Reasoning Communication into Emergent Language

2020-06-07 · NeurIPS 2020 12 · Yipeng Kang, Tonghan Wang, Gerard de Melo

Emergentism and pragmatics are two research fields that study the dynamics of linguistic communication along substantially different timescales and intelligence levels. From the perspective of multi-agent reinforcement l…

Multi-agent Reinforcement LearningReinforcement Learning (RL)StarcraftStarcraft II