paper-with-me

홈 › Papers

Going From Molecules to Genomic Variations to Scientific Discovery: Intelligent Algorithms and Architectures for Intelligent Genome Analysis

2022-05-16 · Mohammed Alser, Joel Lindegger, Can Firtina, Nour Almadhoun, Haiyu Mao, Gagandeep Singh, Juan Gomez-Luna, Onur Mutlu

We now need more than ever to make genome analysis more intelligent. We need to read, analyze, and interpret our genomes not only quickly, but also accurately and efficiently enough to scale the analysis to population level. There currently exist major computational bottlenecks and inefficiencies throughout the entire genome analysis pipeline, because state-of-the-art genome sequencing technologies are still not able to read a genome in its entirety. We describe the ongoing journey in significantly improving the performance, accuracy, and efficiency of genome analysis using intelligent algorithms and hardware architectures. We explain state-of-the-art algorithmic methods and hardware-based acceleration approaches for each step of the genome analysis pipeline and provide experimental evaluations. Algorithmic approaches exploit the structure of the genome as well as the structure of the underlying hardware. Hardware-based acceleration approaches exploit specialized microarchitectures or various execution paradigms (e.g., processing inside or near memory) along with algorithmic changes, leading to new hardware/software co-designed systems. We conclude with a foreshadowing of future challenges, benefits, and research directions triggered by the development of both very low cost yet highly error prone new sequencing technologies and specialized hardware chips for genomics. We hope that these efforts and the challenges we discuss provide a foundation for future work in making genome analysis more intelligent. The analysis script and data used in our experimental evaluation are available at: https://github.com/CMU-SAFARI/Molecules2Variations

📄 PDF Abstract BibTeX arXiv:2205.07957

Code (1)

cmu-safari/molecules2variations 공식 구현

Tasks

scientific discovery

Similar Papers 제목 키워드 기반

Textomics: A Dataset for Genomics Data Summary Generation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Summarizing biomedical discovery from genomics data using natural languages is an essential step in biomedical research but is mostly done manually. Here, we introduce Textomics, a novel dataset of genomics data descript…

Textomics: A Dataset for Genomics Data Summary Generation

2022-05-01 · ACL 2022 5 · Mu-Chun Wang, Zixuan Liu, Sheng Wang

Summarizing biomedical discovery from genomics data using natural languages is an essential step in biomedical research but is mostly done manually. Here, we introduce Textomics, a novel dataset of genomics data descript…

Genomic reproducibility in the bioinformatics era

2023-08-18 · Pelin Icer Baykal, Paweł P. Łabaj, Florian Markowetz, Lynn M. Schriml 외

In biomedical research, validation of a new scientific discovery is tied to the reproducibility of its experimental results. However, in genomics, the definition and implementation of reproducibility still remain impreci…

scientific discovery

Scientific Large Language Models: A Survey on Biological & Chemical Domains

2024-01-26 · Qiang Zhang, Keyang Ding, Tianwen Lyv, Xinda Wang 외

Large Language Models (LLMs) have emerged as a transformative power in enhancing natural language comprehension, representing a significant stride toward artificial general intelligence. The application of LLMs extends b…

scientific discoverySurvey

Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders

2025-11-21 · Samuel Stevens, Jacob Beattie, Tanya Berger-Wolf, Yu Su arxiv

Scientific archives now contain hundreds of petabytes of data across genomics, ecology, climate, and molecular biology that could reveal undiscovered patterns if systematically analyzed at scale. Large-scale, weakly-supe…