Large-scale machine learning for metagenomics sequence classification

被引：52

作者：

Vervier, Kevin ^{[1
,2
,3
,4
]}

Mahe, Pierre ^{[1
]}

Tournoud, Maud ^{[1
]}

Veyrieras, Jean-Baptiste ^{[1
]}

Vert, Jean-Philippe ^{[2
,3
,4
]}

机构：

[1] bioMerieux, Bioinformat Res Dept, F-69280 Marcy Letoile, France

[2] PSL Res Univ, CBIO Ctr Computat Biol, MINES ParisTech, F-77300 Fontainebleau, France

[3] Inst Curie, F-75248 Paris, France

[4] INSERM U900, F-75248 Paris, France

来源：

BIOINFORMATICS | 2016年 / 32卷 / 07期

基金：

欧洲研究理事会;

关键词：

D O I：

10.1093/bioinformatics/btv683

中图分类号：

Q5 [生物化学];

学科分类号：

071010 ; 081704 ;

摘要：

Motivation: Metagenomics characterizes the taxonomic diversity of microbial communities by sequencing DNA directly from an environmental sample. One of the main challenges in metagenomics data analysis is the binning step, where each sequenced read is assigned to a taxonomic clade. Because of the large volume of metagenomics datasets, binning methods need fast and accurate algorithms that can operate with reasonable computing requirements. While standard alignment-based methods provide state-of-the-art performance, compositional approaches that assign a taxonomic class to a DNA read based on the k-mers it contains have the potential to provide faster solutions. Results: We propose a new rank-flexible machine learning-based compositional approach for taxonomic assignment of metagenomics reads and show that it benefits from increasing the number of fragments sampled from reference genome to tune its parameters, up to a coverage of about 10, and from increasing the k-mer size to about 12. Tuning the method involves training machine learning models on about 10(8) samples in 10(7) dimensions, which is out of reach of standard softwares but can be done efficiently with modern implementations for large-scale machine learning. The resulting method is competitive in terms of accuracy with well-established alignment and composition-based tools for problems involving a small to moderate number of candidate species and for reasonable amounts of sequencing errors. We show, however, that machine learning-based compositional approaches are still limited in their ability to deal with problems involving a greater number of species and more sensitive to sequencing errors. We finally show that the new method outperforms the state-of-the-art in its ability to classify reads from species of lineage absent from the reference database and confirm that compositional approaches achieve faster prediction times, with a gain of 2-17 times with respect to the BWA-MEM short read mapper, depending on the number of candidate species and the level of sequencing noise.

引用

页码：1023 / 1032

页数：10

共 50 条

[1] Quick extreme learning machine for large-scale classification
Audi Albtoush
Manuel Fernández-Delgado
Eva Cernadas
Senén Barro
Neural Computing and Applications, 2022, 34 : 5923 - 5938
[2] Quick extreme learning machine for large-scale classification
Albtoush, Audi
Fernandez-Delgado, Manuel
Cernadas, Eva
Barro, Senen
NEURAL COMPUTING & APPLICATIONS, 2022, 34 (08): : 5923 - 5938
[3] Extreme Learning Machine for Large-Scale Graph Classification Based on MapReduce
Wang, Zhanghui
Zhao, Yuhai
Wang, Guoren
PROCEEDINGS OF ELM-2015, VOL 1: THEORY, ALGORITHMS AND APPLICATIONS (I), 2016, 6 : 93 - 105
[4] Chimera: Large-Scale Classification using Machine Learning, Rules, and Crowdsourcing
Sun, Chong
Rampalli, Narasimhan
Yang, Frank
Doan, Anhai
PROCEEDINGS OF THE VLDB ENDOWMENT, 2014, 7 (13): : 1529 - 1540
[5] Extreme Learning Machine for large-scale graph classification based on MapReduce
Wang, Zhanghui
Zhao, Yuhai
Yuan, Ye
Wang, Guoren
Chen, Lei
NEUROCOMPUTING, 2017, 261 : 106 - 114
[6] Large-scale data classification method based on machine learning model
Department of Electrical Engineering, Dalian Institute of Science and Technology, Dalian, China
Int. J. Database Theory Appl., 2 (71-80):
[7] A Survey on Large-Scale Machine Learning
Wang, Meng
Fu, Weijie
He, Xiangnan
Hao, Shijie
Wu, Xindong
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2022, 34 (06) : 2574 - 2594
[8] Hierarchical Classification for Large-Scale Learning
Wang, Boshi
Barbu, Adrian
ELECTRONICS, 2023, 12 (22)
[9] Supervised machine learning for diagnostic classification from large-scale neuroimaging datasets
Lanka, Pradyumna
Rangaprakash, D.
Dretsch, Michael N.
Katz, Jeffrey S.
Denney, Thomas S., Jr.
Deshpande, Gopikrishna
BRAIN IMAGING AND BEHAVIOR, 2020, 14 (06) : 2378 - 2416
[10] Supervised machine learning for diagnostic classification from large-scale neuroimaging datasets
Pradyumna Lanka
D Rangaprakash
Michael N. Dretsch
Jeffrey S. Katz
Thomas S. Denney
Gopikrishna Deshpande
Brain Imaging and Behavior, 2020, 14 : 2378 - 2416

← 1 2 3 4 5 →