GemSIM: general, error-model based simulator of next-generation sequencing data

被引:114
|
作者
McElroy, Kerensa E. [1 ,2 ,3 ]
Luciani, Fabio [3 ]
Thomas, Torsten [1 ,2 ]
机构
[1] UNSW, Ctr Marine Bioinnovat, Sydney, NSW 2052, Australia
[2] UNSW, Sch Biotechnol & Biomol Sci, Sydney, NSW 2052, Australia
[3] Univ New S Wales, Sch Med Sci, Inflammat & Infect Res Grp, Sydney, NSW 2052, Australia
来源
BMC GENOMICS | 2012年 / 13卷
基金
英国医学研究理事会; 澳大利亚国家健康与医学研究理事会;
关键词
QUALITY; ACCURACY; FORMAT;
D O I
10.1186/1471-2164-13-74
中图分类号
Q81 [生物工程学(生物技术)]; Q93 [微生物学];
学科分类号
071005 ; 0836 ; 090102 ; 100705 ;
摘要
Background: GemSIM, or General Error-Model based SIMulator, is a next-generation sequencing simulator capable of generating single or paired-end reads for any sequencing technology compatible with the generic formats SAM and FASTQ (including Illumina and Roche/454). GemSIM creates and uses empirically derived, sequence-context based error models to realistically emulate individual sequencing runs and/or technologies. Empirical fragment length and quality score distributions are also used. Reads may be drawn from one or more genomes or haplotype sets, facilitating simulation of deep sequencing, metagenomic, and resequencing projects. Results: We demonstrate GemSIM's value by deriving error models from two different Illumina sequencing runs and one Roche/454 run, and comparing and contrasting the resulting error profiles of each run. Overall error rates varied dramatically, both between individual Illumina runs, between the first and second reads in each pair, and between datasets from Illumina and Roche/454 technologies. Indels were markedly more frequent in Roche/454 than Illumina and both technologies suffered from an increase in error rates near the end of each read. The effects of these different profiles on low-frequency SNP-calling accuracy were investigated by analysing simulated sequencing data for a mixture of bacterial haplotypes. In general, SNP-calling using VarScan was only accurate for SNPs with frequency > 3%, independent of which error model was used to simulate the data. Variation between error profiles interacted strongly with VarScan's 'minumum average quality' parameter, resulting in different optimal settings for different sequencing runs. Conclusions: Next-generation sequencing has unprecedented potential for assessing genetic diversity, however analysis is complicated as error profiles can vary noticeably even between different runs of the same technology. Simulation with GemSIM can help overcome this problem, by providing insights into the error profiles of individual sequencing runs and allowing researchers to assess the effects of these errors on downstream data analysis.
引用
收藏
页数:9
相关论文
共 50 条
  • [31] Pathway analysis with next-generation sequencing data
    Jinying Zhao
    Yun Zhu
    Eric Boerwinkle
    Momiao Xiong
    European Journal of Human Genetics, 2015, 23 : 507 - 515
  • [32] Identification of indels in next-generation sequencing data
    Ratan, Aakrosh
    Olson, Thomas L.
    Loughran, Thomas P., Jr.
    Miller, Webb
    BMC BIOINFORMATICS, 2015, 16
  • [33] Visualizing next-generation sequencing data with JBrowse
    Westesson, Oscar
    Skinner, Mitchell
    Holmes, Ian
    BRIEFINGS IN BIOINFORMATICS, 2013, 14 (02) : 172 - 177
  • [34] Focus on next-generation sequencing data analysis
    Rusk N.
    Nature Methods, 2009, 6 (Suppl 11) : S1 - S1
  • [35] Next-generation sequencing: adjusting to data overload
    Monya Baker
    Nature Methods, 2010, 7 : 495 - 499
  • [36] Next-generation sequencing: adjusting to data overload
    Baker, Monya
    NATURE METHODS, 2010, 7 (07) : 495 - 499
  • [37] Assembly algorithms for next-generation sequencing data
    Miller, Jason R.
    Koren, Sergey
    Sutton, Granger
    GENOMICS, 2010, 95 (06) : 315 - 327
  • [38] Applications and data analysis of next-generation sequencing
    Vogl, Ina
    Benet-Pages, Anna
    Eck, Sebastian H.
    Kuhn, Marius
    Vosberg, Sebastian
    Greif, Philipp A.
    Metzeler, Klaus H.
    Biskup, Saskia
    Mueller-Reible, Clemens
    Klein, Hanns-Georg
    LABORATORIUMSMEDIZIN-JOURNAL OF LABORATORY MEDICINE, 2013, 37 (06): : 305 - 315
  • [39] Identification of indels in next-generation sequencing data
    Aakrosh Ratan
    Thomas L Olson
    Thomas P Loughran
    Webb Miller
    BMC Bioinformatics, 16
  • [40] Next-generation sequencing and the evolution of data sharing
    de Macena Sobreira, Nara Lygia
    Hamosh, Ada
    AMERICAN JOURNAL OF MEDICAL GENETICS PART A, 2021, 185 (09) : 2633 - 2635