paper-with-me

홈 › Papers

Scientific Large Language Models: A Survey on Biological & Chemical Domains

2024-01-26 · Qiang Zhang, Keyang Ding, Tianwen Lyv, Xinda Wang, Qingyu Yin, Yiwen Zhang, Jing Yu, Yuhao Wang, Xiaotong Li, Zhuoyi Xiang, Kehua Feng, Xiang Zhuang, Zeyuan Wang, Ming Qin, Mengyao Zhang, Jinlu Zhang, Jiyu Cui, Tao Huang, Pengju Yan, Renjun Xu, Hongyang Chen, Xiaolin Li, Xiaohui Fan, Huabin Xing, Huajun Chen

Large Language Models (LLMs) have emerged as a transformative power in enhancing natural language comprehension, representing a significant stride toward artificial general intelligence. The application of LLMs extends beyond conventional linguistic boundaries, encompassing specialized linguistic systems developed within various scientific disciplines. This growing interest has led to the advent of scientific LLMs, a novel subclass specifically engineered for facilitating scientific discovery. As a burgeoning area in the community of AI for Science, scientific LLMs warrant comprehensive exploration. However, a systematic and up-to-date survey introducing them is currently lacking. In this paper, we endeavor to methodically delineate the concept of "scientific language", whilst providing a thorough review of the latest advancements in scientific LLMs. Given the expansive realm of scientific disciplines, our analysis adopts a focused lens, concentrating on the biological and chemical domains. This includes an in-depth examination of LLMs for textual knowledge, small molecules, macromolecular proteins, genomic sequences, and their combinations, analyzing them in terms of model architectures, capabilities, datasets, and evaluation. Finally, we critically examine the prevailing challenges and point out promising research directions along with the advances of LLMs. By offering a comprehensive overview of technical developments in this field, this survey aspires to be an invaluable resource for researchers navigating the intricate landscape of scientific LLMs.

📄 PDF Abstract BibTeX arXiv:2401.14656

Code (1)

hicai-zju/scientific-llm-survey 공식 구현

Tasks

scientific discoverySurvey

Similar Papers 제목 키워드 기반

nach0: Multimodal Natural and Chemical Languages Foundation Model

2023-11-21 · Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov 외

Large Language Models (LLMs) have substantially driven scientific progress in various domains, and many papers have demonstrated their ability to tackle complex problems with creative solutions. Our paper introduces a ne…

Decodermodelnamed-entity-recognitionNamed Entity Recognition+1

Integrating Chemistry Knowledge in Large Language Models via Prompt Engineering

2024-04-22 · Hongxuan Liu, Haoyu Yin, Zhiyao Luo, Xiaonan Wang

This paper presents a study on the integration of domain-specific knowledge in prompt engineering to enhance the performance of large language models (LLMs) in scientific domains. A benchmark dataset is curated to encaps…

HallucinationPrompt Engineeringscientific discovery

EnzChemRED, a rich enzyme chemistry relation extraction dataset

2024-04-22 · Po-Ting Lai, Elisabeth Coudert, Lucila Aimo, Kristian Axelsen 외

Expert curation is essential to capture knowledge of enzyme functions from the scientific literature in FAIR open knowledgebases but cannot keep pace with the rate of new discoveries and new publications. In this work we…

Benchmarkingnamed-entity-recognitionNamed Entity RecognitionNER+2

AI for Biomedicine in the Era of Large Language Models

2024-03-23 · Zhenyu Bi, Sajib Acharjee Dip, Daniel Hajialigol, Sindhura Kommu 외

The capabilities of AI for biomedicine span a wide spectrum, from the atomic level, where it solves partial differential equations for quantum systems, to the molecular level, predicting chemical or protein structures, a…

Language ModelingLanguage ModellingLarge Language ModelTime Series

Bio-xLSTM: Generative modeling, representation and in-context learning of biological and chemical sequences

2024-11-06 · Niklas Schmidinger, Lisa Schneckenreiter, Philipp Seidl, Johannes Schimunek 외

Language models for biological and chemical sequences enable crucial applications such as drug discovery, protein engineering, and precision medicine. Currently, these language models are predominantly based on Transform…

Drug DiscoveryIn-Context Learning