Comparison of Read Mapping and Variant Calling Tools for the Analysis of Plant NGS Data

被引:32
|
作者
Schilbert, Hanna Marie [1 ,2 ]
Rempel, Andreas [1 ,2 ,3 ]
Pucker, Boas [1 ,2 ,4 ]
机构
[1] Bielefeld Univ, Genet & Genom Plants, CeBiTec, D-33615 Bielefeld, Germany
[2] Bielefeld Univ, Fac Biol, D-33615 Bielefeld, Germany
[3] Bielefeld Univ, Grad Sch DILS, Bielefeld Inst Bioinformat Infrastruct BIBI, Fac Technol, D-33615 Bielefeld, Germany
[4] Ruhr Univ Bochum, Mol Genet & Physiol Plants, Fac Biol & Biotechnol, D-44801 Bochum, Germany
来源
PLANTS-BASEL | 2020年 / 9卷 / 04期
关键词
Single Nucleotide Variants (SNVs); Single Nucleotide Polymorphisms (SNPs); Insertions/Deletions (InDels); population genomics; re-sequencing; mapper; benchmarking; Next Generation Sequencing (NGS); bioinformatics; plant genomics; GENOME ANALYSIS; ALIGNMENT; SEQUENCE; DISCOVERY; FRAMEWORK; ACCURATE;
D O I
10.3390/plants9040439
中图分类号
Q94 [植物学];
学科分类号
071001 ;
摘要
High-throughput sequencing technologies have rapidly developed during the past years and have become an essential tool in plant sciences. However, the analysis of genomic data remains challenging and relies mostly on the performance of automatic pipelines. Frequently applied pipelines involve the alignment of sequence reads against a reference sequence and the identification of sequence variants. Since most benchmarking studies of bioinformatics tools for this purpose have been conducted on human datasets, there is a lack of benchmarking studies in plant sciences. In this study, we evaluated the performance of 50 different variant calling pipelines, including five read mappers and ten variant callers, on six real plant datasets of the model organism Arabidopsis thaliana. Sets of variants were evaluated based on various parameters including sensitivity and specificity. We found that all investigated tools are suitable for analysis of NGS data in plant research. When looking at different performance metrics, BWA-MEM and Novoalign were the best mappers and GATK returned the best results in the variant calling step.
引用
收藏
页数:14
相关论文
共 50 条
  • [1] Comparison of variant calling algorithms in NGS analysis
    Barquin, M.
    Perez-Barrios, C.
    Sanchez, E.
    Auglyte, M.
    Sanz, S.
    Ortiz, N.
    Rodriguez, A.
    Gutierrez, L.
    Sanchez Ruiz, A. C.
    Provencio, M.
    Maynou, J.
    Romero, A.
    CLINICA CHIMICA ACTA, 2019, 493 : S113 - S114
  • [2] Comparison of INDEL Calling Tools with Simulation Data and Real Short-Read Data
    Li, Donghe
    Kim, Wonji
    Wang, Longfei
    Yoon, Kyong-Ah
    Park, Boyoung
    Park, Charny
    Kong, Sun-Young
    Hwang, Yongdeuk
    Baek, Daehyun
    Lee, Eun Sook
    Won, Sungho
    IEEE-ACM TRANSACTIONS ON COMPUTATIONAL BIOLOGY AND BIOINFORMATICS, 2019, 16 (05) : 1635 - 1644
  • [3] GNATY: Optimized NGS Variant Calling and Coverage Analysis
    Wolf, Beat
    Kuonen, Pierre
    Dandekar, Thomas
    BIOINFORMATICS AND BIOMEDICAL ENGINEERING (IWBBIO 2016), 2016, 9656 : 446 - 454
  • [4] Performance evaluation of pipelines for mapping, variant calling and interval padding, for the analysis of NGS germline panels
    Zanti, Maria
    Michailidou, Kyriaki
    Loizidou, Maria A.
    Machattou, Christina
    Pirpa, Panagiota
    Christodoulou, Kyproula
    Spyrou, George M.
    Kyriacou, Kyriacos
    Hadjisavvas, Andreas
    BMC BIOINFORMATICS, 2021, 22 (01)
  • [5] Performance evaluation of pipelines for mapping, variant calling and interval padding, for the analysis of NGS germline panels
    Maria Zanti
    Kyriaki Michailidou
    Maria A. Loizidou
    Christina Machattou
    Panagiota Pirpa
    Kyproula Christodoulou
    George M. Spyrou
    Kyriacos Kyriacou
    Andreas Hadjisavvas
    BMC Bioinformatics, 22
  • [6] IMPROVING NGS HLA DATA ANALYSIS BY USE OF SURROGATE BASES FOR READ MAPPING ACCURACY
    Shi, Joel
    Agostini, Tina
    Dinauer, David
    Radick, Marie
    Bialozynski, Carolyn
    Zhao, Bin
    Veldre, Inta
    Gifford, Benjamin D.
    Conradson, Scott
    TISSUE ANTIGENS, 2014, 84 (01): : 123 - 123
  • [7] Porting the Variant Calling Pipeline for NGS data in cloud-HPC environment
    Mulone, Alberto
    Awad, Sherine
    Chiarugi, Davide
    Aldinucci, Marco
    2023 IEEE 47TH ANNUAL COMPUTERS, SOFTWARE, AND APPLICATIONS CONFERENCE, COMPSAC, 2023, : 1858 - 1863
  • [8] ngs_backbone: a pipeline for read cleaning, mapping and SNP calling using Next Generation Sequence
    Blanca, Jose M.
    Pascual, Laura
    Ziarsolo, Peio
    Nuez, Fernando
    Canizares, Joaquin
    BMC GENOMICS, 2011, 12
  • [9] ngs_backbone: a pipeline for read cleaning, mapping and SNP calling using Next Generation Sequence
    Jose M Blanca
    Laura Pascual
    Peio Ziarsolo
    Fernando Nuez
    Joaquin Cañizares
    BMC Genomics, 12
  • [10] Teaser: Individualized benchmarking and optimization of read mapping results for NGS data
    Moritz Smolka
    Philipp Rescheneder
    Michael C. Schatz
    Arndt von Haeseler
    Fritz J. Sedlazeck
    Genome Biology, 16