Identifying metabolic enzymes with multiple types of association evidence

被引:70
作者
Kharchenko, P
Chen, LF
Freund, Y
Vitkup, D
Church, GM
机构
[1] Department of Genetics, Harvard Medical School, Boston, MA 02115
[2] Center for Computational Biology and Bioinformatics, Department of Biomedical Informatics, Columbia University, New York, NY 10032
[3] Department of Computer Science and Engineering, University of California San Diego, La Jolla, CA 92093
关键词
D O I
10.1186/1471-2105-7-177
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background: Existing large-scale metabolic models of sequenced organisms commonly include enzymatic functions which can not be attributed to any gene in that organism. Existing computational strategies for identifying such missing genes rely primarily on sequence homology to known enzyme-encoding genes. Results: We present a novel method for identifying genes encoding for a specific metabolic function based on a local structure of metabolic network and multiple types of functional association evidence, including clustering of genes on the chromosome, similarity of phylogenetic profiles, gene expression, protein fusion events and others. Using E. coli and S. cerevisiae metabolic networks, we illustrate predictive ability of each individual type of association evidence and show that significantly better predictions can be obtained based on the combination of all data. In this way our method is able to predict 60% of enzyme-encoding genes of E. coli metabolism within the top 10 (out of 3551) candidates for their enzymatic function, and as a top candidate within 43% of the cases. Conclusion: We illustrate that a combination of genome context and other functional association evidence is effective in predicting genes encoding metabolic enzymes. Our approach does not rely on direct sequence homology to known enzyme-encoding genes, and can be used in conjunction with traditional homology-based metabolic reconstruction methods. The method can also be used to target orphan metabolic activities.
引用
收藏
页数:16
相关论文
共 61 条
[11]   Protein interaction maps for complete genomes based on gene fusion events [J].
Enright, AJ ;
Iliopoulos, I ;
Kyrpides, NC ;
Ouzounis, CA .
NATURE, 1999, 402 (6757) :86-90
[12]   Genome-scale reconstruction of the Saccharomyces cerevisiae metabolic network [J].
Förster, J ;
Famili, I ;
Fu, P ;
Palsson, BO ;
Nielsen, J .
GENOME RESEARCH, 2003, 13 (02) :244-253
[13]  
Freund Y, 1999, MACHINE LEARNING, PROCEEDINGS, P124
[14]   A decision-theoretic generalization of on-line learning and an application to boosting [J].
Freund, Y ;
Schapire, RE .
JOURNAL OF COMPUTER AND SYSTEM SCIENCES, 1997, 55 (01) :119-139
[15]   Functional organization of the yeast proteome by systematic analysis of protein complexes [J].
Gavin, AC ;
Bösche, M ;
Krause, R ;
Grandi, P ;
Marzioch, M ;
Bauer, A ;
Schultz, J ;
Rick, JM ;
Michon, AM ;
Cruciat, CM ;
Remor, M ;
Höfert, C ;
Schelder, M ;
Brajenovic, M ;
Ruffner, H ;
Merino, A ;
Klein, K ;
Hudak, M ;
Dickson, D ;
Rudi, T ;
Gnau, V ;
Bauch, A ;
Bastuck, S ;
Huhse, B ;
Leutwein, C ;
Heurtier, MA ;
Copley, RR ;
Edelmann, A ;
Querfurth, E ;
Rybin, V ;
Drewes, G ;
Raida, M ;
Bouwmeester, T ;
Bork, P ;
Seraphin, B ;
Kuster, B ;
Neubauer, G ;
Superti-Furga, G .
NATURE, 2002, 415 (6868) :141-147
[16]   The psychosocial and health care needs of HIV-positive people in the United Kingdom following HAART: a review [J].
Green, G ;
Smith, R .
HIV MEDICINE, 2004, 5 :1-46
[17]   PROPERTIES OF THE EXTENDED HYPERGEOMETRIC DISTRIBUTION [J].
HARKNESS, WL .
ANNALS OF MATHEMATICAL STATISTICS, 1965, 36 (03) :938-945
[18]   Systematic identification of protein complexes in Saccharomyces cerevisiae by mass spectrometry [J].
Ho, Y ;
Gruhler, A ;
Heilbut, A ;
Bader, GD ;
Moore, L ;
Adams, SL ;
Millar, A ;
Taylor, P ;
Bennett, K ;
Boutilier, K ;
Yang, LY ;
Wolting, C ;
Donaldson, I ;
Schandorff, S ;
Shewnarane, J ;
Vo, M ;
Taggart, J ;
Goudreault, M ;
Muskat, B ;
Alfarano, C ;
Dewar, D ;
Lin, Z ;
Michalickova, K ;
Willems, AR ;
Sassi, H ;
Nielsen, PA ;
Rasmussen, KJ ;
Andersen, JR ;
Johansen, LE ;
Hansen, LH ;
Jespersen, H ;
Podtelejnikov, A ;
Nielsen, E ;
Crawford, J ;
Poulsen, V ;
Sorensen, BD ;
Matthiesen, J ;
Hendrickson, RC ;
Gleeson, F ;
Pawson, T ;
Moran, MF ;
Durocher, D ;
Mann, M ;
Hogue, CWV ;
Figeys, D ;
Tyers, M .
NATURE, 2002, 415 (6868) :180-183
[19]   Functional discovery via a compendium of expression profiles [J].
Hughes, TR ;
Marton, MJ ;
Jones, AR ;
Roberts, CJ ;
Stoughton, R ;
Armour, CD ;
Bennett, HA ;
Coffey, E ;
Dai, HY ;
He, YDD ;
Kidd, MJ ;
King, AM ;
Meyer, MR ;
Slade, D ;
Lum, PY ;
Stepaniants, SB ;
Shoemaker, DD ;
Gachotte, D ;
Chakraburtty, K ;
Simon, J ;
Bard, M ;
Friend, SH .
CELL, 2000, 102 (01) :109-126
[20]   Predicting protein function by genomic context: Quantitative evaluation and qualitative inferences [J].
Huynen, M ;
Snel, B ;
Lathe, W ;
Bork, P .
GENOME RESEARCH, 2000, 10 (08) :1204-1210