Sequence-based prediction model of protein crystallization propensity using machine learning and two-level feature selection

被引:20
|
作者
Le, Nguyen Quoc Khanh [1 ]
Li, Wanru [2 ]
Cao, Yanshuang [2 ]
机构
[1] Taipei Med Univ, Coll Med, Profess Master Program Artificial Intelligence Med, Taipei 110, Taiwan
[2] Natl Univ Singapore, Inst Syst Sci, Singapore, Singapore
关键词
crystallization; feature selection; machine learning; protein sequence; prediction model; support vector machine; NETWORK;
D O I
10.1093/bib/bbad319
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Protein crystallization is crucial for biology, but the steps involved are complex and demanding in terms of external factors and internal structure. To save on experimental costs and time, the tendency of proteins to crystallize can be initially determined and screened by modeling. As a result, this study created a new pipeline aimed at using protein sequence to predict protein crystallization propensity in the protein material production stage, purification stage and production of crystal stage. The newly created pipeline proposed a new feature selection method, which involves combining Chi-square (chi(2)) and recursive feature elimination together with the 12 selected features, followed by a linear discriminant analysisfor dimensionality reduction and finally, a support vector machine algorithm with hyperparameter tuning and 10-fold cross-validation is used to train the model and test the results. This new pipeline has been tested on three different datasets, and the accuracy rates are higher than the existing pipelines. In conclusion, our model provides a new solution to predict multistage protein crystallization propensity which is a big challenge in computational biology.
引用
收藏
页数:8
相关论文
共 50 条
  • [1] Sequence-Based Prediction of Transmembrane Protein Crystallization Propensity
    Qizhi Zhu
    Lihua Wang
    Ruyu Dai
    Wei Zhang
    Wending Tang
    Yannan Bin
    Zeliang Wang
    Junfeng Xia
    Interdisciplinary Sciences: Computational Life Sciences, 2021, 13 : 693 - 702
  • [2] Sequence-Based Prediction of Transmembrane Protein Crystallization Propensity
    Zhu, Qizhi
    Wang, Lihua
    Dai, Ruyu
    Zhang, Wei
    Tang, Wending
    Bin, Yannan
    Wang, Zeliang
    Xia, Junfeng
    INTERDISCIPLINARY SCIENCES-COMPUTATIONAL LIFE SCIENCES, 2021, 13 (04) : 693 - 702
  • [3] Sequence-based prediction of protein crystallization, purification and production propensity
    Mizianty, Marcin J.
    Kurgan, Lukasz
    BIOINFORMATICS, 2011, 27 (13) : I24 - I33
  • [4] CRYSTALP2: sequence-based protein crystallization propensity prediction
    Kurgan, Lukasz
    Razib, Ali A.
    Aghakhani, Sara
    Dick, Scott
    Mizianty, Marcin
    Jahandideh, Samad
    BMC STRUCTURAL BIOLOGY, 2009, 9
  • [5] RFCRYS: Sequence-based protein crystallization propensity prediction by means of random forest
    Jahandideh, Samad
    Mahdavi, Abbas
    JOURNAL OF THEORETICAL BIOLOGY, 2012, 306 : 115 - 119
  • [6] CRYSpred: Accurate Sequence-Based Protein Crystallization Propensity Prediction Using Sequence-Derived Structural Characteristics
    Mizianty, Marcin J.
    Kurgan, Lukasz A.
    PROTEIN AND PEPTIDE LETTERS, 2012, 19 (01): : 40 - 49
  • [7] Fast Prediction of Protein Methylation Sites Using a Sequence-Based Feature Selection Technique
    Wei, Leyi
    Xing, Pengwei
    Shi, Gaotao
    Ji, Zhiliang
    Zou, Quan
    IEEE-ACM TRANSACTIONS ON COMPUTATIONAL BIOLOGY AND BIOINFORMATICS, 2019, 16 (04) : 1264 - 1273
  • [8] DeepCrystal: a deep learning framework for sequence-based protein crystallization prediction
    Elbasir, Abdurrahman
    Moovarkumudalvan, Balasubramanian
    Kunji, Khalid
    Kolatkar, Prasanna R.
    Mall, Raghvendra
    Bensmail, Halima
    BIOINFORMATICS, 2019, 35 (13) : 2216 - 2225
  • [9] DeepCrystal: A Deep Learning Framework for Sequence-based Protein Crystallization Prediction
    Elbasir, Abdurrahman
    Moovarkumudalvan, Balasubramanian
    Kunji, Khalid
    Kolatkar, Prasanna R.
    Bensmail, Halima
    Mall, Raghvendra
    PROCEEDINGS 2018 IEEE INTERNATIONAL CONFERENCE ON BIOINFORMATICS AND BIOMEDICINE (BIBM), 2018, : 2747 - 2749
  • [10] Sequence-Based Prediction of Cysteine Reactivity Using Machine Learning
    Wang, Haobo
    Chen, Xuemin
    Li, Can
    Liu, Yuan
    Yang, Fan
    Wang, Chu
    BIOCHEMISTRY, 2018, 57 (04) : 451 - 460