Analysis of the yeast transcriptome with structural and functional categories: characterizing highly expressed proteins

被引:98
作者
Jansen, R [1 ]
Gerstein, M [1 ]
机构
[1] Yale Univ, Dept Mol Biophys & Biochem, New Haven, CT 06520 USA
关键词
D O I
10.1093/nar/28.6.1481
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
We analyzed 10 genome expression data sets by large-scale cross-referencing against broad structural and functional categories. The data sets, generated by different techniques (e.g. SAGE and gene chips), provide various representations of the yeast transcriptome (the set of all yeast genes, weighted by transcript abundance). Our analysis enabled us to determine features more prevalent in the transcriptome than the genome: i.e. those that are common to highly expressed proteins. Starting with simplest categories, we find that, relative to the genome, the transcriptome is enriched in Ala and Gly and depleted in Asn and very long proteins. We find, furthermore, that protein length and maximum expression level have a roughly inverse relationship. To relate expression level and protein structure, we assigned transmembrane helices and known folds (using PSI-blast) to each protein in the genome; this allowed us to determine that the transcriptome is enriched in mixed alpha-beta structures and depleted in membrane proteins relative to the genome. In particular, some enzymatic folds, such as the TIM barrel and the G3P dehydrogenase fold, are much more prevalent in the transcriptome than the genome, whereas others, such as the protein-kinase and leucine-zipper folds, are depleted. The TIM barrel, in fact, is overwhelmingly the 'top fold' in the transcriptome, while it only ranks fifth in the genome. The most highly enriched functional categories in the transcriptome (based on the MIPS system) are energy production and protein synthesis, while categories such as transcription, transport and signaling ave depleted. Furthermore, for a given functional category, transcriptome enrichment varies quite substantially between the different expression data sets, with a variation an order of magnitude larger than for the other categories cross-referenced (e.g. amino acids). One can readily see how the enrichment and depletion of the various functional categories relates directly to that of particular folds. Further information can be found at http://bioinfo.mbb.yale.edu/genome/expression.
引用
收藏
页码:1481 / 1488
页数:8
相关论文
共 47 条
  • [11] Data management and analysis for gene expression arrays
    Ermolaeva, O
    Rastogi, M
    Pruitt, KD
    Schuler, GD
    Bittner, ML
    Chen, YD
    Simon, R
    Meltzer, P
    Trent, JM
    Boguski, MS
    [J]. NATURE GENETICS, 1998, 20 (01) : 19 - 23
  • [12] Comprehensive, comprehensible, distributed and intelligent databases: current status
    Frishman, D
    Heumann, K
    Lesk, A
    Mewes, HW
    [J]. BIOINFORMATICS, 1998, 14 (07) : 551 - 561
  • [13] How representative are the known structures of the proteins in a complete genome? A comprehensive structural census
    Gerstein, M
    [J]. FOLDING & DESIGN, 1998, 3 (06): : 497 - 512
  • [15] Gerstein M, 1998, PROTEINS, V33, P518, DOI 10.1002/(SICI)1097-0134(19981201)33:4<518::AID-PROT5>3.0.CO
  • [16] 2-J
  • [17] Life with 6000 genes
    Goffeau, A
    Barrell, BG
    Bussey, H
    Davis, RW
    Dujon, B
    Feldmann, H
    Galibert, F
    Hoheisel, JD
    Jacq, C
    Johnston, M
    Louis, EJ
    Mewes, HW
    Murakami, Y
    Philippsen, P
    Tettelin, H
    Oliver, SG
    [J]. SCIENCE, 1996, 274 (5287) : 546 - &
  • [18] Goffeau A., 1996, SCIENCE, V274, p[546, 563]
  • [19] Correlation between protein and mRNA abundance in yeast
    Gygi, SP
    Rochon, Y
    Franza, BR
    Aebersold, R
    [J]. MOLECULAR AND CELLULAR BIOLOGY, 1999, 19 (03) : 1720 - 1730
  • [20] The relationship between protein structure and function: a comprehensive survey with application to the yeast genome
    Hegyi, H
    Gerstein, M
    [J]. JOURNAL OF MOLECULAR BIOLOGY, 1999, 288 (01) : 147 - 164