Aggressive assembly of pyrosequencing reads with mates

被引:384
作者
Miller, Jason R. [1 ]
Delcher, Arthur L. [2 ]
Koren, Sergey [1 ]
Venter, Eli [1 ]
Walenz, Brian P. [1 ]
Brownley, Anushka [1 ]
Johnson, Justin [1 ]
Li, Kelvin [1 ]
Mobarry, Clark [3 ]
Sutton, Granger [1 ]
机构
[1] J Craig Venter Inst, Rockville, MD 20850 USA
[2] Univ Maryland, Ctr Bioinformat & Computat Biol, College Pk, MD 20742 USA
[3] White Oak Technol Inc, Silver Spring, MD 20910 USA
关键词
D O I
10.1093/bioinformatics/btn548
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Motivation: DNA sequence reads from Sanger and pyrosequencing platforms differ in cost, accuracy, typical coverage, average read length and the variety of available paired-end protocols. Both read types can complement one another in a 'hybrid' approach to whole-genome shotgun sequencing projects, but assembly software must be modified to accommodate their different characteristics. This is true even of pyrosequencing mated and unmated read combinations. Without special modifications, assemblers tuned for homogeneous sequence data may perform poorly on hybrid data. Results: Celera Assembler was modified for combinations of ABI 3730 and 454 FLX reads. The revised pipeline called CABOG (Celera Assembler with the Best Overlap Graph) is robust to homopolymer run length uncertainty, high read coverage and heterogeneous read lengths. In tests on four genomes, it generated the longest contigs among all assemblers tested. It exploited the mate constraints provided by paired-end reads from either platform to build larger contigs and scaffolds, which were validated by comparison to a finished reference sequence. A low rate of contig mis-assembly was detected in some CABOG assemblies, but this was reduced in the presence of sufficient mate pair data.
引用
收藏
页码:2818 / 2824
页数:7
相关论文
共 29 条
[11]   Whole-genome shotgun assembly and comparison of human genome assemblies [J].
Istrail, S ;
Sutton, GG ;
Florea, L ;
Halpern, AL ;
Mobarry, CM ;
Lippert, R ;
Walenz, B ;
Shatkay, H ;
Dew, I ;
Miller, JR ;
Flanigan, MJ ;
Edwards, NJ ;
Bolanos, R ;
Fasulo, D ;
Halldorsson, BV ;
Hannenhalli, S ;
Turner, R ;
Yooseph, S ;
Lu, F ;
Nusskern, DR ;
Shue, BC ;
Zheng, XQH ;
Zhong, F ;
Delcher, AL ;
Huson, DH ;
Kravitz, SA ;
Mouchard, L ;
Reinert, K ;
Remington, KA ;
Clark, AG ;
Waterman, MS ;
Eichler, EE ;
Adams, MD ;
Hunkapiller, MW ;
Myers, EW ;
Venter, JC .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 2004, 101 (07) :1916-1921
[12]   Whole-genome sequence assembly for mammalian genomes: Arachne 2 [J].
Jaffe, DB ;
Butler, J ;
Gnerre, S ;
Mauceli, E ;
Lindblad-Toh, K ;
Mesirov, JP ;
Zody, MC ;
Lander, ES .
GENOME RESEARCH, 2003, 13 (01) :91-96
[13]  
Jarvie Thomas, 2008, Biotechniques, V44, P829, DOI 10.2144/000112894
[14]  
Korbel JO, 2007, SCIENCE, V318, P420, DOI 10.1126/science.1149504
[15]   Versatile and open software for comparing large genomes [J].
Kurtz, S ;
Phillippy, A ;
Delcher, AL ;
Smoot, M ;
Shumway, M ;
Antonescu, C ;
Salzberg, SL .
GENOME BIOLOGY, 2004, 5 (02)
[16]   REPuter: the manifold applications of repeat analysis on a genomic scale [J].
Kurtz, S ;
Choudhuri, JV ;
Ohlebusch, E ;
Schleiermacher, C ;
Stoye, J ;
Giegerich, R .
NUCLEIC ACIDS RESEARCH, 2001, 29 (22) :4633-4642
[17]   The diploid genome sequence of an individual human [J].
Levy, Samuel ;
Sutton, Granger ;
Ng, Pauline C. ;
Feuk, Lars ;
Halpern, Aaron L. ;
Walenz, Brian P. ;
Axelrod, Nelson ;
Huang, Jiaqi ;
Kirkness, Ewen F. ;
Denisov, Gennady ;
Lin, Yuan ;
MacDonald, Jeffrey R. ;
Pang, Andy Wing Chun ;
Shago, Mary ;
Stockwell, Timothy B. ;
Tsiamouri, Alexia ;
Bafna, Vineet ;
Bansal, Vikas ;
Kravitz, Saul A. ;
Busam, Dana A. ;
Beeson, Karen Y. ;
Mclntosh, Tina C. ;
Remington, Karin A. ;
Abril, Josep F. ;
Gill, John ;
Borman, Jon ;
Rogers, Yu-Hui ;
Frazier, Marvin E. ;
Scherer, Stephen W. ;
Strausberg, Robert L. ;
Venter, J. Craig .
PLOS BIOLOGY, 2007, 5 (10) :2113-2144
[18]   Genome sequencing in microfabricated high-density picolitre reactors [J].
Margulies, M ;
Egholm, M ;
Altman, WE ;
Attiya, S ;
Bader, JS ;
Bemben, LA ;
Berka, J ;
Braverman, MS ;
Chen, YJ ;
Chen, ZT ;
Dewell, SB ;
Du, L ;
Fierro, JM ;
Gomes, XV ;
Godwin, BC ;
He, W ;
Helgesen, S ;
Ho, CH ;
Irzyk, GP ;
Jando, SC ;
Alenquer, MLI ;
Jarvie, TP ;
Jirage, KB ;
Kim, JB ;
Knight, JR ;
Lanza, JR ;
Leamon, JH ;
Lefkowitz, SM ;
Lei, M ;
Li, J ;
Lohman, KL ;
Lu, H ;
Makhijani, VB ;
McDade, KE ;
McKenna, MP ;
Myers, EW ;
Nickerson, E ;
Nobile, JR ;
Plant, R ;
Puc, BP ;
Ronan, MT ;
Roth, GT ;
Sarkis, GJ ;
Simons, JF ;
Simpson, JW ;
Srinivasan, M ;
Tartaro, KR ;
Tomasz, A ;
Vogt, KA ;
Volkmer, GA .
NATURE, 2005, 437 (7057) :376-380
[19]   A whole-genome assembly of Drosophila [J].
Myers, EW ;
Sutton, GG ;
Delcher, AL ;
Dew, IM ;
Fasulo, DP ;
Flanigan, MJ ;
Kravitz, SA ;
Mobarry, CM ;
Reinert, KHJ ;
Remington, KA ;
Anson, EL ;
Bolanos, RA ;
Chou, HH ;
Jordan, CM ;
Halpern, AL ;
Lonardi, S ;
Beasley, EM ;
Brandon, RC ;
Chen, L ;
Dunn, PJ ;
Lai, ZW ;
Liang, Y ;
Nusskern, DR ;
Zhan, M ;
Zhang, Q ;
Zheng, XQ ;
Rubin, GM ;
Adams, MD ;
Venter, JC .
SCIENCE, 2000, 287 (5461) :2196-2204
[20]   Complete genome sequence of the oral pathogenic bacterium Porphyromonas gingivalis strain W83 [J].
Nelson, KE ;
Fleischmann, RD ;
DeBoy, RT ;
Paulsen, IT ;
Fouts, DE ;
Eisen, JA ;
Daugherty, SC ;
Dodson, RJ ;
Durkin, AS ;
Gwinn, M ;
Haft, DH ;
Kolonay, JF ;
Nelson, WC ;
Mason, T ;
Tallon, L ;
Gray, J ;
Granger, D ;
Tettelin, H ;
Dong, H ;
Galvin, JL ;
Duncan, MJ ;
Dewhirst, FE ;
Fraser, CM .
JOURNAL OF BACTERIOLOGY, 2003, 185 (18) :5591-5601