Assessing the limits of genomic data integration for predicting protein networks

被引：148

作者：

Lu, LJ

Xia, Y

Paccanaro, A

Yu, HY

Gerstein, M

机构：

[1] Yale Univ, Dept Mol Biophys & Biochem, New Haven, CT 06520 USA

[2] Yale Univ, Dept Comp Sci, New Haven, CT 06520 USA

[3] Yale Univ, Program Computat Biol & Bioinformat, New Haven, CT 06520 USA

来源：

GENOME RESEARCH | 2005年 / 15卷 / 07期

关键词：

D O I：

10.1101/gr.3610305

中图分类号：

Q5 [生物化学]; Q7 [分子生物学];

学科分类号：

071010 ; 081704 ;

摘要：

Genomic data integration-the process of statistically combining diverse Sources of information from functional genomics experiments to make large-scale predictions-is becoming increasingly prevalent. One might expect that this process should become progressively more powerful With the integration of more evidence. Here, we explore the limits of genomic data integration, assessing the degree to which predictive power increases with the addition of more features. We focus oil a predictive context that has been extensively investigated and benchmarked in the past-the prediction of protein-protein interactions in yeast. We start by using a simple Naive Bayes classifier for integrating diverse Sources of genomic evidence, ranging from coexpression relationships to similar phylogenetic profiles. We expand the number of features considered for prediction to 16, significantly more than previous Studies. Overall, we observe a small, but measurable improvement in prediction performance over previous benchmarks, based on four strong features. This allows us to identify new yeast interactions with high confidence. It also allows us to quantitatively assess the inter-relations amongst different genomic features. It is known that subtle correlations and dependencies between features call confound the strength of interaction predictions. We investigate this issue in detail through calculating mutual information. To Our Surprise, we find no appreciable statistical dependence between the many possible pairs of features. We further explore feature dependencies by comparing the performance Of Our simple Naive Bayes classifier with a boosted version of the same classifier, which is fairly resistant to feature dependence. We find that boosting does not improve performance, indicating that, at least for prediction purposes, Our genomic features are essentially independent. In Summary, by integrating a few (i.e., four) good features, we approach the maximal predictive power of current genomic data integration; moreover, this limitation does not reflect (potentially removable) inter-relationships between the features.

引用

页码：945 / 953

页数：9

共 50 条

[1] A Bayesian networks approach for predicting protein-protein interactions from genomic data
Jansen, R
Yu, HY
Greenbaum, D
Kluger, Y
Krogan, NJ
Chung, SB
Emili, A
Snyder, M
Greenblatt, JF
Gerstein, M
SCIENCE, 2003, 302 (5644) : 449 - 453
[2] Predicting co-complexed protein pairs using genomic and proteomic data integration
Zhang, LV
Wong, SL
King, OD
Roth, FP
BMC BIOINFORMATICS, 2004, 5 (1)
[3] Predicting co-complexed protein pairs using genomic and proteomic data integration
Lan V Zhang
Sharyl L Wong
Oliver D King
Frederick P Roth
BMC Bioinformatics, 5
[4] Bayesian Inference for Genomic Data Integration Reduces Misclassification Rate in Predicting Protein-Protein Interactions
Xing, Chuanhua
Dunson, David B.
PLOS COMPUTATIONAL BIOLOGY, 2011, 7 (07)
[5] Integration of genomic data for inferring protein complexes from global protein-protein interaction networks
Zheng, Huiru
Wang, Haiying
Glass, David H.
IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART B-CYBERNETICS, 2008, 38 (01): : 5 - 16
[6] Predicting biological networks from genomic data
Harrington, Eoghan D.
Jensen, Lars J.
Bork, Peer
FEBS LETTERS, 2008, 582 (08) : 1251 - 1258
[7] Predicting clinical outcomes in neuroblastoma with genomic data integration
Baali, Ilyes
Acar, D. Alp Emre
Aderinwale, Tunde W.
HafezQorani, Saber
Kazan, Hilal
BIOLOGY DIRECT, 2018, 13
[8] Predicting clinical outcomes in neuroblastoma with genomic data integration
Ilyes Baali
D Alp Emre Acar
Tunde W. Aderinwale
Saber HafezQorani
Hilal Kazan
Biology Direct, 13
[9] Predicting Protein Function Based on Genomic Data Mining
Song, Chang-xin
Wu, Xiaoming
Wang, Bo
Cheng, Jingzhi
2008 INTERNATIONAL CONFERENCE ON ADVANCED COMPUTER THEORY AND ENGINEERING, 2008, : 759 - 762
[10] Predicting Protein Function by Genomic Data-Mining
Song, Changxin
Ma, Ke
ADVANCED INTELLIGENT COMPUTING THEORIES AND APPLICATIONS, PROCEEDINGS: WITH ASPECTS OF CONTEMPORARY INTELLIGENT COMPUTING TECHNIQUES, 2008, 15 : 229 - +

← 1 2 3 4 5 →