paper-with-me

홈 › Papers

TEDDY: A Family Of Foundation Models For Understanding Single Cell Biology

2025-03-05 · Alexis Chevalier, Soumya Ghosh, Urvi Awasthi, James Watkins, Julia Bieniewska, Nichita Mitrea, Olga Kotova, Kirill Shkura, Andrew Noble, Michael Steinbaugh, Julien Delile, Christoph Meier, Leonid Zhukov, Iya Khalil, Srayanta Mukherjee, Judith Mueller

Understanding the biological mechanism of disease is critical for medicine, and in particular drug discovery. AI-powered analysis of genome-scale biological data hold great potential in this regard. The increasing availability of single-cell RNA sequencing data has enabled the development of large foundation models for disease biology. However, existing foundation models either do not improve or only modestly improve over task-specific models in downstream applications. Here, we explored two avenues for improving the state-of-the-art. First, we scaled the pre-training dataset to 116 million cells, which is larger than those used by previous models. Second, we leveraged the availability of large-scale biological annotations as a form of supervision during pre-training. We trained the TEDDY family of models comprising six transformer-based state-of-the-art single-cell foundation models with 70 million, 160 million, and 400 million parameters. We vetted our models on two downstream evaluation tasks -- identifying the underlying disease state of held-out donors not seen during training and distinguishing healthy cells from diseased ones for disease conditions and donors not seen during training. Scaling experiments showed that performance improved predictably with both data volume and parameter count. Our models showed substantial improvement over existing work on the first task and more muted improvements on the second.

📄 PDF Abstract BibTeX arXiv:2503.03485

Code (0)

등록된 구현이 없습니다.

Tasks

Drug Discovery

Similar Papers 제목 키워드 기반

CellVerse: Do Large Language Models Really Understand Cell Biology?

2025-05-09 · Fan Zhang, Tianyu Liu, Zhihong Zhu, Hao Wu 외

Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehe…

Drug Response PredictionQuestion Answering

TEDDY: Trimming Edges with Degree-based Discrimination strategY

2024-02-02 · Hyunjin Seo, Jihun Yun, Eunho Yang

Since the pioneering work on the lottery ticket hypothesis for graph neural networks (GNNs) was proposed in Chen et al. (2021), the study on finding graph lottery tickets (GLT) has become one of the pivotal focus in the …

Towards Applying Large Language Models to Complement Single-Cell Foundation Models

2025-07-14 · Steven Palayew, Bo Wang, Gary Bader arxiv

Single-cell foundation models such as scGPT represent a significant advancement in single-cell omics, with an ability to achieve state-of-the-art performance on various downstream biological tasks. However, these models …

Teddy: Efficient Large-Scale Dataset Distillation via Taylor-Approximated Matching

2024-10-10 · Ruonan Yu, Songhua Liu, Jingwen Ye, Xinchao Wang

Dataset distillation or condensation refers to compressing a large-scale dataset into a much smaller one, enabling models trained on this synthetic dataset to generalize effectively on real data. Tackling this challenge,…

Dataset Distillation

Robofriend: An Adpative Storytelling Robotic Teddy Bear - Technical Report

2023-01-04 · Ido Glanz, Matan Weksler, Erez Karpas, Tzipi Horowitz-Kraus

In this paper we describe Robofriend, a robotic teddy bear for telling stories to young children. Robofriend adapts its behavior to keep the childrens' attention using reinforcement learning.

reinforcement-learningReinforcement LearningReinforcement Learning (RL)