The relationship between protein structure and function: a comprehensive survey with application to the yeast genome

被引:294
作者
Hegyi, H [1 ]
Gerstein, M [1 ]
机构
[1] Yale Univ, Dept Mol Biophys & Biochem, New Haven, CT 06520 USA
关键词
structure-function; fold classification; structural convergence; functional divergence; yeast genomics;
D O I
10.1006/jmbi.1999.2661
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
For most proteins in the genome databases, function is predicted via sequence comparison. In spite of the popularity of this approach, the extent to which it can be reliably applied is unknown. We address this issue by systematically investigating the relationship between protein function and structure. We focus initially on enzymes functionally classified by the Enzyme Commission (EC) and relate these to by structurally classified domains the SCOP database. We find that the major SCOP fold classes have different propensities to carry out certain broad categories of functions. For instance, alpha/beta folds are disproportionately associated with enzymes, especially transferases and hydrolases, and all-alpha and small folds with non-enzymes, while alpha +beta folds have an equal tendency either way. These observations for the database overall are largely true for specific genomes. We focus, in particular, on yeast, analyzing it with many classifications in addition to SCOP and EC (i.e. COGs, CATH, MIPS), and find clear tendencies for fold-function association, across a broad spectrum of functions. Analysis with the COGs scheme also suggests that the func tions of the most ancient proteins are more evenly distributed among different structural classes than those of more modern ones. For the data base overall, we identify the most versatile functions, i.e. those that are associated with the most folds, and the most versatile folds, associated with the most functions. The two most versatile enzymatic functions (hydro-lyases and O-glycosyl glucosidases) are associated with seven folds each. The five most versatile folds (TIM-barrel, Rossmann, ferredoxin, alpha-beta hydrolase, and P-loop NTP hydrolase) are all mixed alpha-beta structures. They stand out as generic scaffolds, accommodating from six to as many as 16 functions (for the exceptional TIM-barrel). At the conclusion of our analysis we are able to construct a graph giving the chance that a functional annotation can be reliably transferred at different degrees of sequence and structural similarity. Supplemental information is available from http://bioinfo.mbb.yale.edu/genome/foldfunc. (C) 1999 Academic Press.
引用
收藏
页码:147 / 164
页数:18
相关论文
共 64 条
[1]   Gapped BLAST and PSI-BLAST: a new generation of protein database search programs [J].
Altschul, SF ;
Madden, TL ;
Schaffer, AA ;
Zhang, JH ;
Zhang, Z ;
Miller, W ;
Lipman, DJ .
NUCLEIC ACIDS RESEARCH, 1997, 25 (17) :3389-3402
[2]   BASIC LOCAL ALIGNMENT SEARCH TOOL [J].
ALTSCHUL, SF ;
GISH, W ;
MILLER, W ;
MYERS, EW ;
LIPMAN, DJ .
JOURNAL OF MOLECULAR BIOLOGY, 1990, 215 (03) :403-410
[3]   The PRINTS protein fingerprint database in its fifth year [J].
Attwood, TK ;
Beck, ME ;
Flower, DR ;
Scordis, P ;
Selley, JN .
NUCLEIC ACIDS RESEARCH, 1998, 26 (01) :304-308
[4]   The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998 [J].
Bairoch, A ;
Apweiler, R .
NUCLEIC ACIDS RESEARCH, 1998, 26 (01) :38-42
[5]   The PROSITE database, its status in 1997 [J].
Bairoch, A ;
Bucher, P ;
Hofmann, K .
NUCLEIC ACIDS RESEARCH, 1997, 25 (01) :217-221
[6]   The ENZYME data bank in 1995 [J].
Bairoch, A .
NUCLEIC ACIDS RESEARCH, 1996, 24 (01) :221-222
[7]  
Barrett Alan J., 1997, European Journal of Biochemistry, V250, P1
[8]   Sequences and topology - Deriving biological knowledge from genomic sequences [J].
Bork, P ;
Eisenberg, D .
CURRENT OPINION IN STRUCTURAL BIOLOGY, 1998, 8 (03) :331-332
[9]  
BORK P, 1993, PROTEIN SCI, V2, P31
[10]   Predicting functions from protein sequences - where are the bottlenecks? [J].
Bork, P ;
Koonin, EV .
NATURE GENETICS, 1998, 18 (04) :313-318