Search arXivSearch

arXiv · 2511.11626

Omics-scale polymer computational database transferable to real-world artificial intelligence applications

Ryo Yoshida·Yoshihiro Hayashi·Hidemine Furuya·Ryohei Hosoya·Kazuyoshi Kaneko·Hiroki Sugisawa·Yu Kaneko·Aiko Takahashi·Yoh Noguchi·Shun Nanjo·Keiko Shinoda·Tomu Hamakawa·Mitsuru Ohno·Takuya Kitamura·Misaki Yonekawa·Stephen Wu·Masato Ohnishi·Chang Liu·Teruki Tsurimoto·Arifin·Araki Wakiuchi·Kohei Noda·Junko Morikawa·Teruaki Hayakawa·Junichiro Shiomi·Masanobu Naito·Kazuya Shiratori·Tomoki Nagai·Norio Tomotsu·Hiroto Inoue·Ryuichi Sakashita·Masashi Ishii·Isao Kuwajima·Kenji Furuichi·Norihiko Hiroi·Yuki Takemoto·Takahiro Ohkuma·Keita Yamamoto·Naoya Kowatari·Masato Suzuki·Naoya Matsumoto·Seiryu Umetani·Hisaki Ikebata·Yasuyuki Shudo·Mayu Nagao·Shinya Kamada·Kazunori Kamio·Taichi Shomura·Kensaku Nakamura·Yudai Iwamizu·Atsutoshi Abe·Koki Yoshitomi·Yuki Horie·Katsuhiko Koike·Koichi Iwakabe·Shinya Gima·Kota Usui·Gikyo Usuki·Takuro Tsutsumi·Keitaro Matsuoka·Kazuki Sada·Masahiro Kitabata·Takuma Kikutsuji·Akitaka Kamauchi·Yusuke Iijima·Tsubasa Suzuki·Takenori Goda·Yuki Takabayashi·Kazuko Imai·Yuji Mochizuki·Hideo Doi·Koji Okuwaki·Hiroya Nitta·Taku Ozawa·Hitoshi Kamijima·Toshiaki Shintani·Takuma Mitamura·Massimiliano Zamengo·Yuitsu Sugami·Seiji Akiyama·Yoshinari Murakami·Atsushi Betto·Naoya Matsuo·Satoru Kagao·Tetsuya Kobayashi·Norie Matsubara·Shosei Kubo·Yuki Ishiyama·Yuri Ichioka·Mamoru Usami·Satoru Yoshizaki·Seigo Mizutani·Yosuke Hanawa·Shogo Kunieda·Mitsuru Yambe·Takeru Nakamura·Hiromori Murashima·Kenji Takahashi·Naoki Wada·Masahiro Kawano

Abstract

Developing large-scale foundational datasets is a critical milestone in advancing artificial intelligence (AI)-driven scientific innovation. However, unlike AI-mature fields such as natural language processing, materials science, particularly polymer research, has significantly lagged in developing extensive open datasets. This lag is primarily due to the high costs of polymer synthesis and property measurements, along with the vastness and complexity of the chemical space. This study presents PolyOmics, an omics-scale computational database generated through fully automated molecular dynamics simulation pipelines that provide diverse physical properties for over $10^5$ polymeric materials. The PolyOmics database is collaboratively developed by approximately 260 researchers from 48 institutions to bridge the gap between academia and industry. Machine learning models pretrained on PolyOmics can be efficiently fine-tuned for a wide range of real-world downstream tasks, even when only limited experimental data are available. Notably, the generalisation capability of these simulation-to-real transfer models improve significantly as the size of the PolyOmics database increases, exhibiting power-law scaling. The emergence of scaling laws supports the "more is better" principle, highlighting the significance of ultralarge-scale computational materials data for improving real-world prediction performance. This unprecedented omics-scale database reveals vast unexplored regions of polymer materials, providing a foundation for AI-driven polymer science.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ryo Yoshida, Yoshihiro Hayashi, Hidemine Furuya, Ryohei Hosoya, Kazuyoshi Kaneko, Hiroki Sugisawa, Yu Kaneko, Aiko Takahashi, Yoh Noguchi, Shun Nanjo, Keiko Shinoda, Tomu Hamakawa, Mitsuru Ohno, Takuya Kitamura, Misaki Yonekawa, Stephen Wu, Masato Ohnishi, Chang Liu, Teruki Tsurimoto, Arifin, Araki Wakiuchi, Kohei Noda, Junko Morikawa, Teruaki Hayakawa, Junichiro Shiomi, Masanobu Naito, Kazuya Shiratori, Tomoki Nagai, Norio Tomotsu, Hiroto Inoue, Ryuichi Sakashita, Masashi Ishii, Isao Kuwajima, Kenji Furuichi, Norihiko Hiroi, Yuki Takemoto, Takahiro Ohkuma, Keita Yamamoto, Naoya Kowatari, Masato Suzuki, Naoya Matsumoto, Seiryu Umetani, Hisaki Ikebata, Yasuyuki Shudo, Mayu Nagao, Shinya Kamada, Kazunori Kamio, Taichi Shomura, Kensaku Nakamura, Yudai Iwamizu, Atsutoshi Abe, Koki Yoshitomi, Yuki Horie, Katsuhiko Koike, Koichi Iwakabe, Shinya Gima, Kota Usui, Gikyo Usuki, Takuro Tsutsumi, Keitaro Matsuoka, Kazuki Sada, Masahiro Kitabata, Takuma Kikutsuji, Akitaka Kamauchi, Yusuke Iijima, Tsubasa Suzuki, Takenori Goda, Yuki Takabayashi, Kazuko Imai, Yuji Mochizuki, Hideo Doi, Koji Okuwaki, Hiroya Nitta, Taku Ozawa, Hitoshi Kamijima, Toshiaki Shintani, Takuma Mitamura, Massimiliano Zamengo, Yuitsu Sugami, Seiji Akiyama, Yoshinari Murakami, Atsushi Betto, Naoya Matsuo, Satoru Kagao, Tetsuya Kobayashi, Norie Matsubara, Shosei Kubo, Yuki Ishiyama, Yuri Ichioka, Mamoru Usami, Satoru Yoshizaki, Seigo Mizutani, Yosuke Hanawa, Shogo Kunieda, Mitsuru Yambe, Takeru Nakamura, Hiromori Murashima, Kenji Takahashi, Naoki Wada, Masahiro Kawano. 2025-11-07. Omics-scale polymer computational database transferable to real-world artificial intelligence applications. https://arxiv.org/abs/2511.11626

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Global approximations to the error function of real argument for vectorized computation

The error function of real argument can be approximated to a given uniform relative accuracy by a single closed-form expression for the whole variable range either in terms of addition, multiplication, division, and square root operations only, or also using the exponential function. The coefficients have been tabulated for up to 128-bit precision. Tests of a computer code implementation using the standard single- and double-precision floating-point arithmetic show good performance and vectorizability. Approximations to the complementary error function have also been found, including those where only additions, multiplications, and one division are needed to reach a uniform absolute accuracy.

physics.chem-ph

From Heuristics to Machine Learning: The Performance Ceiling for Single-Ion Magnets and Its Electronic Origin

Machine learning (ML) is expected to speed up the discovery of single-ion magnets (SIMs), but does the structural information available before synthesis allow such predictions? For 1215 lanthanide complexes from the SIMDAVIS 1.2.1 database we compared three increasing levels of structural description: tabular features of the coordination site, continuous symmetry measures of the coordination polyhedron, and the complete 3D arrangement of atoms. All three converge to an accuracy near 76%, only slightly above the 71% of the single rule "predict SIM for Dy3+". To explain the failures, we combined multireference ab initio calculations with an inspection of the structures behind the high-confidence errors. The SIMs missed by the geometric models are field-induced relaxers whose ground Kramers doublets are prone to tunnelling, a property invisible to geometric descriptors. Many false positives contain several lanthanide centers or radicals, so their relaxation is collective and outside the single-ion picture. The electronic-structure and connectivity information needed to identify SIMs is therefore not accessible to geometric methods alone. Geometric models remain useful: restricting the screening to compounds with high prediction confidence raises the accuracy to 88% while retaining 48% of the dataset. Building on the analysis of the failures, we propose a strategy that combines simple filters for nuclearity and for radicals with ligand-field descriptors from ab initio calculations.

physics.chem-ph

Composition-Dependent Self-Diffusion Coefficients in Liquid Mixtures from Hybrid Machine Learning

Self-diffusion coefficients are key descriptors of molecular mobility, yet experimental data remain scarce, highlighting the need for reliable prediction methods. In previous work, we introduced the hybrid Enhanced Stokes-Einstein (ESE) model, which advanced the state of the art in the physically consistent prediction of self-diffusion coefficients of solutes at infinite dilution in pure solvents by integrating the Stokes-Einstein equation with machine learning (ML). Here, we extend this approach to concentration-dependent self-diffusion coefficients and multicomponent solvents with HADES. This hybrid architecture leverages a deep-set neural network to connect pure-component and mixture prediction within a single framework. HADES predicts self-diffusion coefficients in liquid mixtures with any number of components at any composition and temperature. The only required inputs are SMILES-encoded molecular structures of the components and the pure-component viscosities, making the method broadly applicable. Trained and evaluated on a comprehensive dataset of 2526 data points for 600 systems, HADES significantly outperforms benchmark prediction methods. The trained model and its source code are fully disclosed, and the application is available via an interactive website https://ml-prop.mv.rptu.de/.

physics.chem-ph