paper-with-me

홈 › Papers

Omics-scale polymer computational database transferable to real-world artificial intelligence applications

2025-11-07 · Ryo Yoshida, Yoshihiro Hayashi, Hidemine Furuya, Ryohei Hosoya, Kazuyoshi Kaneko, Hiroki Sugisawa, Yu Kaneko, Aiko Takahashi, Yoh Noguchi, Shun Nanjo, Keiko Shinoda, Tomu Hamakawa, Mitsuru Ohno, Takuya Kitamura, Misaki Yonekawa, Stephen Wu, Masato Ohnishi, Chang Liu, Teruki Tsurimoto, Arifin, Araki Wakiuchi, Kohei Noda, Junko Morikawa, Teruaki Hayakawa, Junichiro Shiomi, Masanobu Naito, Kazuya Shiratori, Tomoki Nagai, Norio Tomotsu, Hiroto Inoue, Ryuichi Sakashita, Masashi Ishii, Isao Kuwajima, Kenji Furuichi, Norihiko Hiroi, Yuki Takemoto, Takahiro Ohkuma, Keita Yamamoto, Naoya Kowatari, Masato Suzuki, Naoya Matsumoto, Seiryu Umetani, Hisaki Ikebata, Yasuyuki Shudo, Mayu Nagao, Shinya Kamada, Kazunori Kamio, Taichi Shomura, Kensaku Nakamura, Yudai Iwamizu, Atsutoshi Abe, Koki Yoshitomi, Yuki Horie, Katsuhiko Koike, Koichi Iwakabe, Shinya Gima, Kota Usui, Gikyo Usuki, Takuro Tsutsumi, Keitaro Matsuoka, Kazuki Sada, Masahiro Kitabata, Takuma Kikutsuji, Akitaka Kamauchi, Yusuke Iijima, Tsubasa Suzuki, Takenori Goda, Yuki Takabayashi, Kazuko Imai, Yuji Mochizuki, Hideo Doi, Koji Okuwaki, Hiroya Nitta, Taku Ozawa, Hitoshi Kamijima, Toshiaki Shintani, Takuma Mitamura, Massimiliano Zamengo, Yuitsu Sugami, Seiji Akiyama, Yoshinari Murakami, Atsushi Betto, Naoya Matsuo, Satoru Kagao, Tetsuya Kobayashi, Norie Matsubara, Shosei Kubo, Yuki Ishiyama, Yuri Ichioka, Mamoru Usami, Satoru Yoshizaki, Seigo Mizutani, Yosuke Hanawa, Shogo Kunieda, Mitsuru Yambe, Takeru Nakamura, Hiromori Murashima, Kenji Takahashi, Naoki Wada, Masahiro Kawano, Yosuke Harada, Takehiro Fujita, Erina Fujita, Ryoji Himeno, Hiori Kino, Kenji Fukumizu arxiv

Developing large-scale foundational datasets is a critical milestone in advancing artificial intelligence (AI)-driven scientific innovation. However, unlike AI-mature fields such as natural language processing, materials science, particularly polymer research, has significantly lagged in developing extensive open datasets. This lag is primarily due to the high costs of polymer synthesis and property measurements, along with the vastness and complexity of the chemical space. This study presents PolyOmics, an omics-scale computational database generated through fully automated molecular dynamics simulation pipelines that provide diverse physical properties for over $10^5$ polymeric materials. The PolyOmics database is collaboratively developed by approximately 260 researchers from 48 institutions to bridge the gap between academia and industry. Machine learning models pretrained on PolyOmics can be efficiently fine-tuned for a wide range of real-world downstream tasks, even when only limited experimental data are available. Notably, the generalisation capability of these simulation-to-real transfer models improve significantly as the size of the PolyOmics database increases, exhibiting power-law scaling. The emergence of scaling laws supports the "more is better" principle, highlighting the significance of ultralarge-scale computational materials data for improving real-world prediction performance. This unprecedented omics-scale database reveals vast unexplored regions of polymer materials, providing a foundation for AI-driven polymer science.

📄 PDF Abstract BibTeX arXiv:2511.11626

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hyperactive Learning (HAL) for Data-Driven Interatomic Potentials

2022-10-09 · Cas van der Oord, Matthias Sachs, Dávid Péter Kovács, Christoph Ortner 외

Data-driven interatomic potentials have emerged as a powerful class of surrogate models for {\it ab initio} potential energy surfaces that are able to reliably predict macroscopic properties with experimental accuracy. I…

CPU

PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design

2026-05-26 · Manpreet Kaur, Xingying Zhang, Qian Liu arxiv

Polymer discovery is central to fields ranging from energy storage to biomedicine, but it is hindered by an astronomically large chemical design space and fragmented representations of structure, properties, and prior kn…

Representation Learning

POINT$^{2}$: A Polymer Informatics Training and Testing Database

2025-03-30 · Jiaxin Xu, Gang Liu, Ruilan Guo, Meng Jiang 외

The advancement of polymer informatics has been significantly propelled by the integration of machine learning (ML) techniques, enabling the rapid prediction of polymer properties and expediting the discovery of high-per…

Uncertainty Quantification

FAPS: A Fast Platform for Protein Structureomics Analysis

2025-06-11 · Lucas Wilken, Nihjum Paul, Troy Timmerman, Sara A. Tolba 외

Protein quantification and analysis are well-accepted approaches for biomarker discovery but are limited to identification without structural information. High-throughput omics data (i.e., genomics, transcriptomics, and …

A Systematic Overview of Single-Cell Transcriptomics Databases, their Use cases, and Limitations

2024-04-15 · Mahnoor N. Gondal, Saad Ur Rehman Shah, Arul M. Chinnaiyan, Marcin Cieslik

Rapid advancements in high-throughput single-cell RNA-seq (scRNA-seq) technologies and experimental protocols have led to the generation of vast amounts of genomic data that populates several online databases and reposit…