paper-with-me

홈 › Papers

Entropy, Disagreement, and the Limits of Foundation Models in Genomics

2026-04-05 · Maxime Rochkoulets, Lovro Vrček, Mile Šikić arxiv

Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing. Yet, the reasons for their limited effectiveness remain poorly understood. In this work, we investigate the role of entropy as a fundamental factor limiting the capacities of such models to learn from their training data and develop foundational capabilities. We train ensembles of models on text and DNA sequences and analyze their predictions, static embeddings, and empirical Fisher information flow. We show that the high entropy of genomic sequences -- from the point of view of unseen token prediction -- leads to near-uniform output distributions, disagreement across models, and unstable static embeddings, even for models that are matched in architecture, training and data. We then demonstrate that models trained on DNA concentrate Fisher information in embedding layers, seemingly failing to exploit inter-token relationships. Our results suggest that self-supervised training from sequences alone may not be applicable to genomic data, calling into question the assumptions underlying current methodologies for training genomic foundation models.

📄 PDF Abstract BibTeX arXiv:2604.04287

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement

2025-08-06 · Karrtik Iyer, Manikandan Ravikiran, Prasanna Pendse, Shayan Mohanty arxiv

Automated grading systems can efficiently score short-answer responses, yet they often fail to indicate when a grading decision is uncertain or potentially contentious. We introduce semantic entropy, a measure of variabi…

GQVis: A Dataset of Genomics Data Questions and Visualizations for Generative AI

2025-09-19 · Skylar Sargent Walters, Arthea Valderrama, Thomas C. Smits, David Kouřil 외 arxiv

Data visualization is a fundamental tool in genomics research, enabling the exploration, interpretation, and communication of complex genomic features. While machine learning models show promise for transforming data int…

Theoretical Foundation of Co-Training and Disagreement-Based Algorithms

2017-08-15 · Wei Wang, Zhi-Hua Zhou

Disagreement-based approaches generate multiple classifiers and exploit the disagreement among them with unlabeled data to improve learning performance. Co-training is a representative paradigm of them, which trains two …

Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation

2026-05-26 · Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang 외 arxiv

On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uniformity of OPD signals by prioritizing high-entropy or high-disagree…

Interpretable Multimodal Cancer Prototyping with Whole Slide Images and Incompletely Paired Genomics

2025-11-26 · Yupei Zhang, Yating Huang, Wanming Hu, Lequan Yu 외 arxiv

Multimodal approaches that integrate histology and genomics hold strong potential for precision oncology. However, phenotypic and genotypic heterogeneity limits the quality of intra-modal representations and hinders effe…