MEGAN analysis of metagenomic data

被引:2243
作者
Huson, Daniel H.
Auch, Alexander F.
Qi, Ji
Schuster, Stephan C.
机构
[1] Univ Tubingen, Ctr Bioinformat, D-72076 Tubingen, Germany
[2] Penn State Univ, Ctr Comparat Genom & Bioinformat, Ctr Infect Dis Dynam, University Pk, PA 16802 USA
关键词
D O I
10.1101/gr.5969107
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
Metagenomics is the study of the genomic content of a sample of organisms obtained from a common habitat using targeted or random sequencing. Goals include understanding the extent and role of microbial diversity. The taxonomical content of such a sample is usually estimated by comparison against sequence databases of known sequences. Most published studies use the analysis of paired-end reads, complete sequences of environmental fosmid and BAC clones, or environmental assemblies. Emerging sequencing- by-synthesis technologies with very high throughput are paving the way to low-cost random "shotgun" approaches. This paper introduces MEGAN, a new computer program that allows laptop analysis of large metagenomic data sets. In a preprocessing step, the set of DNA sequences is compared against databases of known sequences using BLAST or another comparison tool. MEGAN is then used to compute and explore the taxonomical content of the data set, employing the NCBI taxonomy to summarize and order the results. A simple lowest common ancestor algorithm assigns reads to taxa such that the taxonomical level of the assigned taxon reflects the level of conservation of the sequence. The software allows large data sets to be dissected without the need for assembly or the targeting of specific phylogenetic markers. It provides graphical and statistical output for comparing different data sets. The approach is applied to several data sets, including the Sargasso Sea data set, a recently published metagenomic data set sampled from a mammoth bone, and several complete microbial genomes. Also, simulations that evaluate the performance of the approach for different read lengths are presented.
引用
收藏
页码:377 / 386
页数:10
相关论文
共 27 条
[1]   BASIC LOCAL ALIGNMENT SEARCH TOOL [J].
ALTSCHUL, SF ;
GISH, W ;
MILLER, W ;
MYERS, EW ;
LIPMAN, DJ .
JOURNAL OF MOLECULAR BIOLOGY, 1990, 215 (03) :403-410
[2]   Bacterial rhodopsin:: Evidence for a new type of phototrophy in the sea [J].
Béjà, O ;
Aravind, L ;
Koonin, EV ;
Suzuki, MT ;
Hadd, A ;
Nguyen, LP ;
Jovanovich, S ;
Gates, CM ;
Feldman, RA ;
Spudich, JL ;
Spudich, EN ;
DeLong, EF .
SCIENCE, 2000, 289 (5486) :1902-1906
[3]   Proteorhodopsin phototrophy in the ocean [J].
Béjà, O ;
Spudich, EN ;
Spudich, JL ;
Leclerc, M ;
DeLong, EF .
NATURE, 2001, 411 (6839) :786-789
[4]   GenBank [J].
Benson, Dennis A. ;
Karsch-Mizrachi, Ilene ;
Lipman, David J. ;
Ostell, James ;
Wheeler, David L. .
NUCLEIC ACIDS RESEARCH, 2006, 34 :D16-D20
[5]   The complete genome sequence of Escherichia coli K-12 [J].
Blattner, FR ;
Plunkett, G ;
Bloch, CA ;
Perna, NT ;
Burland, V ;
Riley, M ;
ColladoVides, J ;
Glasner, JD ;
Rode, CK ;
Mayhew, GF ;
Gregor, J ;
Davis, NW ;
Kirkpatrick, HA ;
Goeden, MA ;
Rose, DJ ;
Mau, B ;
Shao, Y .
SCIENCE, 1997, 277 (5331) :1453-+
[6]   Microbial community genomics in the ocean [J].
DeLong, EE .
NATURE REVIEWS MICROBIOLOGY, 2005, 3 (06) :459-469
[7]   Community genomics among stratified microbial assemblages in the ocean's interior [J].
DeLong, EF ;
Preston, CM ;
Mincer, T ;
Rich, V ;
Hallam, SJ ;
Frigaard, NU ;
Martinez, A ;
Sullivan, MB ;
Edwards, R ;
Brito, BR ;
Chisholm, SW ;
Karl, DM .
SCIENCE, 2006, 311 (5760) :496-503
[8]   A review of DNA sequencing techniques [J].
França, LTC ;
Carrilho, E ;
Kist, TBL .
QUARTERLY REVIEWS OF BIOPHYSICS, 2002, 35 (02) :169-200
[9]   Reverse methanogenesis: Testing the hypothesis with environmental genomics [J].
Hallam, SJ ;
Putnam, N ;
Preston, CM ;
Detter, JC ;
Rokhsar, D ;
Richardson, PM ;
DeLong, EF .
SCIENCE, 2004, 305 (5689) :1457-1462
[10]   Metagenomics: Application of genomics to uncultured microorganisms [J].
Handelsman, J .
MICROBIOLOGY AND MOLECULAR BIOLOGY REVIEWS, 2004, 68 (04) :669-+