Identification and analysis of over 2000 ribosomal protein pseudogenes in the human genome

被引:145
作者
Zhang, ZL [1 ]
Harrison, P [1 ]
Gerstein, M [1 ]
机构
[1] Yale Univ, Dept Mol Biophys & Biochem, New Haven, CT 06520 USA
关键词
D O I
10.1101/gr.331902
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
Mammals have 79 ribosomal proteins (RP). Using a systematic procedure based on sequence-homology, we have comprehensively identified pseudogenes of these proteins in the human genome. Our assignments are available at http://www.pseudogene.org or http://bioinfo.mbb.yale.edu/genome/pseudogene. In total, we found 2090 processed pseudogenes and 16 duplications of RP genes. In relation to the matching parent protein, each of the processed pseudogenes has an average relative sequence length of 97% and an average sequence identity of 76%. A small number (258) of them do not contain obvious disablements (stop codons or frameshifts) and, therefore, could be mistaken as functional genes, and 178 are disrupted by one or more repetitive elements. On average, processed pseudogenes have a longer truncation at the S' end than the 3' end, consistent with the target-primed-reverse-transcription (TPRT) mechanism. Interestingly, on chromosome 16, an RPL26 processed pseudogene was found in the intron region of a functional RPS2 gene. The large-scale distribution of RP pseudogenes throughout the genome appears to result, chiefly, from random insertions with the numbers on each chromosome, consequently, proportional to its size. In contrast to RP genes, the RP pseudogenes have the highest density in GC-intermediate regions (41%-46%) of the genome, with the density pattern being between that of LINEs and Alus. This can be explained by a negative selection theory as we observed that GC-rich RP pseudogenes decay faster in GC-poor regions. Also, we observed a correlation between the number of processed pseudogenes and the GC content of the associated functional gene, i.e., relatively GC-poor RPs have more processed pseudogenes. This ranges from 145 pseudogenes for RPL21 down to 3 pseudogenes for RPL14. We were able to date the RP pseudogenes based on their sequence divergence from present-day RP genes, finding an age distribution similar to that for Alus. The distribution is consistent with a decline in retrotransposition activity in the hominid lineage during the last 40 Myr. We discuss the implications for retrotransposon stability and genome dynamics based on these new findings.
引用
收藏
页码:1466 / 1482
页数:17
相关论文
共 66 条
[21]   Ribonuclease and high salt sensitivity of the ribonucleoprotein complex formed by the human LINE-1 retrotransposon [J].
Hohjoh, H ;
Singer, MF .
JOURNAL OF MOLECULAR BIOLOGY, 1997, 271 (01) :7-12
[22]   Sequence-specific single-strand RNA binding protein encoded by the human LINE-1 retrotransposon [J].
Hohjoh, H ;
Singer, MF .
EMBO JOURNAL, 1997, 16 (19) :6034-6043
[23]   Cytoplasmic ribonucleoprotein complexes containing human LINE-1 protein and RNA [J].
Hohjoh, H ;
Singer, MF .
EMBO JOURNAL, 1996, 15 (03) :630-639
[24]   The Ensembl genome database project [J].
Hubbard, T ;
Barker, D ;
Birney, E ;
Cameron, G ;
Chen, Y ;
Clark, L ;
Cox, T ;
Cuff, J ;
Curwen, V ;
Down, T ;
Durbin, R ;
Eyras, E ;
Gilbert, J ;
Hammond, M ;
Huminiecki, L ;
Kasprzyk, A ;
Lehvaslaiho, H ;
Lijnzaad, P ;
Melsopp, C ;
Mongin, E ;
Pettett, R ;
Pocock, M ;
Potter, S ;
Rust, A ;
Schmidt, E ;
Searle, S ;
Slater, G ;
Smith, J ;
Spooner, W ;
Stabenau, A ;
Stalker, J ;
Stupka, E ;
Ureta-Vidal, A ;
Vastrik, I ;
Clamp, M .
NUCLEIC ACIDS RESEARCH, 2002, 30 (01) :38-41
[25]   Sequence patterns indicate an enzymatic involvement in integration of mammalian retroposons [J].
Jurka, J .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 1997, 94 (05) :1872-1877
[26]   The impact of L1 retrotransposons on the human genome [J].
Kazazian, HH ;
Moran, JV .
NATURE GENETICS, 1998, 19 (01) :19-24
[27]   The human ribosomal protein L6 gene in a critical region for Noonan syndrome [J].
Kenmochi, N ;
Yoshihama, M ;
Higa, S ;
Tanaka, T .
JOURNAL OF HUMAN GENETICS, 2000, 45 (05) :290-293
[28]   A map of 75 human ribosomal protein genes [J].
Kenmochi, N ;
Kawaguchi, T ;
Rozen, S ;
Davis, E ;
Goodman, N ;
Hudson, TJ ;
Tanaka, T ;
Page, DC .
GENOME RESEARCH, 1998, 8 (05) :509-523
[30]   MEGA2: molecular evolutionary genetics analysis software [J].
Kumar, S ;
Tamura, K ;
Jakobsen, IB ;
Nei, M .
BIOINFORMATICS, 2001, 17 (12) :1244-1245