Escaping Local Optima in the Waddington Landscape: A Two-Stage TRPO-PPO Approach for Single-Cell Perturbation Analysis
Modeling cellular responses to genetic and chemical perturbations remains a central challenge in single-cell biology. Existing data-driven frameworks have advanced perturbation prediction through variational autoencoders, chemically conditioned autoencoders, and large-scale transformer pretraining. However, most existing models rely exclusively on either in silico perturbation data or experimental perturbation data but rarely integrate both, limiting their ability to generalize and validate predictions across simulated and real biological contexts in a digital twin system. Moreover, the models are prone to local optima in the nonconvex Waddington landscape of cell fate decisions, where poor initialization can trap trajectories in spurious lineages. In this work, we introduce a two-stage reinforcement learning algorithm for modeling single-cell perturbation. We first compute an explicit natural gradient update using Fisher-vector products and a conjugate gradient solver, scaled by a KL trust-region constraint to provide a safe, curvature-aware first step for the policy. Starting with these preconditioned parameters, we then apply a second phase of proximal policy optimization (PPO) with a KL penalty, exploiting minibatch efficiency to refine the policy. We demonstrate that this initialization strategy substantially improves generalization on Single-cell RNA sequencing (scRNA-seq) perturbation analysis in a digital twin system.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningSimilar Papers 제목 키워드 기반
Incorporating stochastic gene expression, signaling-mediated intercellular interactions, and regulated cell proliferation in models of coordinated tissue development
Formulating quantitative and predictive models for tissue development requires consideration of the complex, stochastic gene expression dynamics, its regulation via cell-to-cell interactions, and cell proliferation. Incl…
Drawing a Waddington landscape to capture dynamic epigenetics
Epigenetics is most often reduced to chromatin marking in the current literature, whereas this notion was initially defined in a more general context. This restricted view ignores that epigenetic memories are in fact mor…
Homeorhesis in Waddington's Landscape by Epigenetic Feedback Regulation
In multicellular organisms, cells differentiate into several distinct types during early development. Determination of each cellular state, along with the ratio of each cell type, as well as the developmental course duri…
A Waddington landscape for prototype learning in generalized Hopfield networks
Networks in machine learning offer examples of complex high-dimensional dynamical systems reminiscent of biological systems. Here, we study the learning dynamics of Generalized Hopfield networks, which permit a visualiza…
Homotopic Convex Transformation: A New Landscape Smoothing Method for the Traveling Salesman Problem
This paper proposes a novel landscape smoothing method for the symmetric Traveling Salesman Problem (TSP). We first define the Homotopic Convex (HC) transformation of a TSP as a convex combination of a well-constructed s…
Traveling Salesman Problem