paper-with-me

Papers

PerturbCellRL: Verifier-Guided Reinforcement Learning for Single-Cell Perturbation Prediction

2026-06-26 · Dongxia Wu, Mingyu Li, Yuhui Zhang, Anurendra Kumar, Emma Lundberg, Serena Yeung-Levy, Emily B. Fox arxiv

Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions. While recent generative models improve population-level prediction, individual generated cells are not explicitly checked for biological consistency. We introduce PerturbCellRL, a reinforcement learning (RL) framework that post-trains a pretrained single-cell transcriptomic generator using a suite of cell-level verifiers as rewards. These verifiers define four rewards: Pearson top-k similarity, RMSE top-k proximity, DE Spearman, and Pathway activity. The Pathway activity verifier rewards cells whose pathway responses match known perturbation biology. We evaluate PerturbCellRL on multiple genetic and chemical perturbation benchmarks. Across these benchmarks, PerturbCellRL improves over the pretrained flow-matching generator on reward-aligned evaluation metrics and a held-out evaluation metric. Moreover, PerturbCellRL remains competitive with state-of-the-art methods on population-level metrics. Together, these results frame trustworthy single-cell prediction as verifier-guided generative alignment, moving beyond matching expression distributions toward predictions whose single-cell perturbation effects are explicitly checked for biological consistency.

📄 PDF Abstract BibTeX arXiv:2606.27752

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification

2026-08-26 · Zebei Zhao, Zhihao Shi, Minqi Shi arxiv

Reference-based verifiers are important for evaluating reasoning models and providing accurate outcome rewards in reinforcement learning with verifiable rewards. To improve verification accuracy, prior work has explored …

Reinforcement Learning

Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering

2026-05-28 · Shicheng Fan, Haochang Hao, Dehai Min, Weihao Liu 외 arxiv

Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewards provide only coarse supervision and cannot distinguish correct f…

Reinforcement LearningQuestion Answering

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction

2026-02-13 · Xin-Qiang Cai, Masashi Sugiyama arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a dominant paradigm for enhancing Large Language Models (LLMs) reasoning, yet its reliance on external verifiers limits its scalability. Recent finding…

Reinforcement Learning

Guided Verifier: Collaborative Multimodal Reasoning via Dynamic Process Supervision

2026-02-04 · Lingzhuang Sun, Ruitong Liu, Yuxia Zhu, Xiaohan Xu 외 arxiv

Reinforcement Learning (RL) has emerged as a pivotal mechanism for enhancing the complex reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevailing paradigms typically rely on solitary rollou…

Reinforcement LearningMultimodal Reasoning

RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning

2025-05-21 · Kaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong 외

Reinforcement learning (RL) has recently emerged as a compelling approach for enhancing the reasoning capabilities of large language models (LLMs), where an LLM generator serves as a policy guided by a verifier (reward m…

MathMathematical ReasoningReinforcement Learning (RL)