Molecular quantum chemical data sets and databases for machine learning potentials

被引:1
|
作者
Ullah, Arif [1 ]
Chen, Yuxinxin [2 ]
Dral, Pavlo O. [2 ,3 ]
机构
[1] Anhui Univ, Sch Phys & Optoelect Engn, Hefei 230601, Anhui, Peoples R China
[2] Xiamen Univ, Coll Chem & Chem Engn, State Key Lab Phys Chem Solid Surfaces, Fujian Prov Key Lab Theoret & Computat Chem, Xiamen 361005, Fujian, Peoples R China
[3] Nicolaus Copernicus Univ Torun, Inst Phys, Fac Phys Astron & Informat, Ul Grudzdzka 5, PL-87100 Torun, Poland
来源
基金
中国国家自然科学基金;
关键词
database; quantum chemistry; electronic properties; data set; machine learning; FORCE-FIELD; NONCOVALENT INTERACTIONS; DENSITY FUNCTIONALS; VIRTUAL EXPLORATION; ORBITAL METHODS; CHEMISTRY; UNIVERSE; THERMOCHEMISTRY; RESOLUTION; BENCHMARK;
D O I
10.1088/2632-2153/ad8f13
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The field of computational chemistry is increasingly leveraging machine learning (ML) potentials to predict molecular properties with high accuracy and efficiency, providing a viable alternative to traditional quantum mechanical (QM) methods, which are often computationally intensive. Central to the success of ML models is the quality and comprehensiveness of the data sets on which they are trained. Quantum chemistry data sets and databases, comprising extensive information on molecular structures, energies, forces, and other properties derived from QM calculations, are crucial for developing robust and generalizable ML potentials. In this review, we provide an overview of the current landscape of quantum chemical data sets and databases. We examine key characteristics and functionalities of prominent resources, including the types of information they store, the level of electronic structure theory employed, the diversity of chemical space covered, and the methodologies used for data creation. Additionally, an updatable resource is provided to track new data sets and databases at https://github.com/Arif-PhyChem/datasets_and_databases_4_MLPs. This resource also has the overview in a machine-readable database format with the Jupyter notebook example for analysis. Looking forward, we discuss the challenges associated with the rapid growth of quantum chemical data sets and databases, emphasizing the need for updatable and accessible resources to ensure the long-term utility of them. We also address the importance of data format standardization and the ongoing efforts to align with the FAIR principles to enhance data interoperability and reusability. Drawing inspiration from established materials databases, we advocate for the development of user-friendly and sustainable platforms for these data sets and databases.
引用
收藏
页数:29
相关论文
共 50 条
  • [1] Impact of the Characteristics of Quantum Chemical Databases on Machine Learning Prediction of Tautomerization Energies
    Vazquez-Salazar, Luis Itza
    Boittier, Eric D.
    Unke, Oliver T.
    Meuwly, Markus
    JOURNAL OF CHEMICAL THEORY AND COMPUTATION, 2021, 17 (08) : 4769 - 4785
  • [2] Applications and training sets of machine learning potentials
    Hong, Changho
    Kim, Jaehoon
    Kim, Jaesun
    Jung, Jisu
    Ju, Suyeon
    Choi, Jeong Min
    Han, Seungwu
    SCIENCE AND TECHNOLOGY OF ADVANCED MATERIALS-METHODS, 2023, 3 (01):
  • [3] Robust estimation of the intrinsic dimension of data sets with quantum cognition machine learning
    Candelori, Luca
    Abanov, Alexander G.
    Berger, Jeffrey
    Hogan, Cameron J.
    Kirakosyan, Vahagn
    Musaelian, Kharen
    Samson, Ryan
    Smith, James E. T.
    Villani, Dario
    Wells, Martin T.
    Xu, Mengjia
    SCIENTIFIC REPORTS, 2025, 15 (01):
  • [4] Maximizing information from chemical engineering data sets: Applications to machine learning
    Thebelt, Alexander
    Wiebe, Johannes
    Kronqvist, Jan
    Tsay, Calvin
    Misener, Ruth
    CHEMICAL ENGINEERING SCIENCE, 2022, 252
  • [5] Data mining: Machine learning, statistics, and databases
    Mannila, H
    EIGHTH INTERNATIONAL CONFERENCE ON SCIENTIFIC AND STATISTICAL DATABASE SYSTEMS, PROCEEDINGS, 1996, : 2 - 9
  • [6] Quantum Chemical Roots of Machine-Learning Molecular Similarity Descriptors
    Gugler, Stefan
    Reiher, Markus
    JOURNAL OF CHEMICAL THEORY AND COMPUTATION, 2022, : 6670 - 6689
  • [7] Correlation and redundancy on machine learning performance for chemical databases
    Li, Hongzhi
    Li, Wenze
    Pan, Xuefeng
    Huang, Jiaqi
    Gao, Ting
    Hu, LiHong
    Li, Hui
    Lu, Yinghua
    JOURNAL OF CHEMOMETRICS, 2018, 32 (07)
  • [8] Applications of machine learning methods for chemical reaction databases
    Tkachenko, Valery
    Sattarov, Boris
    Korotcov, Alexandru
    Lowe, Daniel
    Nugmanov, Ramil
    Madzhidov, Timur
    Varnek, Alexandre
    ABSTRACTS OF PAPERS OF THE AMERICAN CHEMICAL SOCIETY, 2017, 254
  • [9] Analysis of Data Sets With Learning Conflicts for Machine Learning
    Ledesma, Sergio
    Ibarra-Manzano, Mario-Alberto
    Cabal-Yepez, Eduardo
    Almanza-Ojeda, Dora-Luz
    Avina-Cervantes, Juan-Gabriel
    IEEE ACCESS, 2018, 6 : 45062 - 45070
  • [10] Negative Data in Data Sets for Machine Learning Training
    Maloney, Michael P.
    Coley, Connor W.
    Genheden, Samuel
    Carson, Nessa
    Helquist, Paul
    Norrby, Per-Ola
    Wiest, Olaf
    ORGANIC LETTERS, 2023, 25 (17) : 2945 - 2947