Shotgun proteomics aids discovery of novel protein-coding genes, alternative splicing, and "resurrected" pseudogenes in the mouse genome

被引:94
作者
Brosch, Markus [1 ]
Saunders, Gary I. [1 ]
Frankish, Adam [1 ]
Collins, Mark O. [1 ]
Yu, Lu [1 ]
Wright, James [1 ]
Verstraten, Ruth [1 ]
Adams, David J. [1 ]
Harrow, Jennifer [1 ]
Choudhary, Jyoti S. [1 ]
Hubbard, Tim [1 ]
机构
[1] Wellcome Trust Sanger Inst, Cambridge CB10 1SA, England
基金
英国惠康基金;
关键词
POSTERIOR ERROR PROBABILITIES; MASS-SPECTROMETRY; PEPTIDE IDENTIFICATION; DROSOPHILA-MELANOGASTER; ANNOTATION; PREDICTION; DATABASE; SPECTRA; DUPLICATION; VALIDATION;
D O I
10.1101/gr.114272.110
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
Recent advances in proteomic mass spectrometry (MS) offer the chance to marry high-throughput peptide sequencing to transcript models, allowing the validation, refinement, and identification of new protein-coding loci. We present a novel pipeline that integrates highly sensitive and statistically robust peptide spectrum matching with genome-wide protein-coding predictions to perform large-scale gene validation and discovery in the mouse genome for the first time. In searching an excess of 10 million spectra, we have been able to validate 32%, 17%, and 7% of all protein-coding genes, exons, and splice boundaries, respectively. Moreover, we present strong evidence for the identification of multiple alternatively spliced translations from 53 genes and have uncovered 10 entirely novel protein-coding genes, which are not covered in any mouse annotation data sources. One such novel protein-coding gene is a fusion protein that spans the Ins2 and Igf2 loci to produce a transcript encoding the insulin II and the insulin-like growth factor 2-derived peptides. We also report nine processed pseudogenes that have unique peptide hits, demonstrating, for the first time, that they are not just transcribed but are translated and are therefore resurrected into new coding loci. This work not only highlights an important utility for MS data in genome annotation but also provides unique insights into the gene structure and propagation in the mouse genome. All these data have been subsequently used to improve the publicly available mouse annotation available in both the Vega and Ensembl genome browsers (http://vega.sanger.ac.uk).
引用
收藏
页码:756 / 767
页数:12
相关论文
共 75 条
[1]   Mouse project to find each gene's role [J].
Abbott, Alison .
NATURE, 2010, 465 (7297) :410-410
[2]   Manual annotation and analysis of the defensin gene cluster in the C57BL/6J mouse reference genome [J].
Amid, Clara ;
Rehaume, Linda M. ;
Brown, Kelly L. ;
Gilbert, James G. R. ;
Dougan, Gordon ;
Hancock, Robert E. W. ;
Harrow, Jennifer L. .
BMC GENOMICS, 2009, 10
[3]  
[Anonymous], 2006, GENOME BIOL S1
[4]   The Vertebrate Genome Annotation (Vega) database [J].
Ashurst, JL ;
Chen, CK ;
Gilbert, JGR ;
Jekosch, K ;
Keenan, S ;
Meidl, P ;
Searle, SM ;
Stalker, J ;
Storey, R ;
Trevanion, S ;
Wilming, L ;
Hubbard, T .
NUCLEIC ACIDS RESEARCH, 2005, 33 :D459-D465
[5]   Current topics in genome evolution: Molecular mechanisms of new gene formation [J].
Babushok, D. V. ;
Ostertag, E. M. ;
Kazazian, H. H., Jr. .
CELLULAR AND MOLECULAR LIFE SCIENCES, 2007, 64 (05) :542-554
[6]   Retrocopy contributions to the evolution of the human genome [J].
Baertsch, Robert ;
Diekhans, Mark ;
Kent, W. James ;
Haussler, David ;
Brosius, Juergen .
BMC GENOMICS, 2008, 9 (1)
[7]   CONTROLLING THE FALSE DISCOVERY RATE - A PRACTICAL AND POWERFUL APPROACH TO MULTIPLE TESTING [J].
BENJAMINI, Y ;
HOCHBERG, Y .
JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY, 1995, 57 (01) :289-300
[8]   CONTRIBUTIONS OF MASS-SPECTROMETRY TO PEPTIDE AND PROTEIN-STRUCTURE [J].
BIEMANN, K .
BIOMEDICAL AND ENVIRONMENTAL MASS SPECTROMETRY, 1988, 16 (1-12) :99-111
[9]  
Birney E, 1997, ISMB-97 - FIFTH INTERNATIONAL CONFERENCE ON INTELLIGENT SYSTEMS FOR MOLECULAR BIOLOGY, PROCEEDINGS, P56
[10]   Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project [J].
Birney, Ewan ;
Stamatoyannopoulos, John A. ;
Dutta, Anindya ;
Guigo, Roderic ;
Gingeras, Thomas R. ;
Margulies, Elliott H. ;
Weng, Zhiping ;
Snyder, Michael ;
Dermitzakis, Emmanouil T. ;
Stamatoyannopoulos, John A. ;
Thurman, Robert E. ;
Kuehn, Michael S. ;
Taylor, Christopher M. ;
Neph, Shane ;
Koch, Christoph M. ;
Asthana, Saurabh ;
Malhotra, Ankit ;
Adzhubei, Ivan ;
Greenbaum, Jason A. ;
Andrews, Robert M. ;
Flicek, Paul ;
Boyle, Patrick J. ;
Cao, Hua ;
Carter, Nigel P. ;
Clelland, Gayle K. ;
Davis, Sean ;
Day, Nathan ;
Dhami, Pawandeep ;
Dillon, Shane C. ;
Dorschner, Michael O. ;
Fiegler, Heike ;
Giresi, Paul G. ;
Goldy, Jeff ;
Hawrylycz, Michael ;
Haydock, Andrew ;
Humbert, Richard ;
James, Keith D. ;
Johnson, Brett E. ;
Johnson, Ericka M. ;
Frum, Tristan T. ;
Rosenzweig, Elizabeth R. ;
Karnani, Neerja ;
Lee, Kirsten ;
Lefebvre, Gregory C. ;
Navas, Patrick A. ;
Neri, Fidencio ;
Parker, Stephen C. J. ;
Sabo, Peter J. ;
Sandstrom, Richard ;
Shafer, Anthony .
NATURE, 2007, 447 (7146) :799-816