Characterization of the human ESC transcriptome by hybrid sequencing

被引:235
作者
Au, Kin Fai [1 ,2 ]
Sebastiano, Vittorio [3 ]
Afshar, Pegah Tootoonchi [4 ]
Durruthy, Jens Durruthy [3 ]
Lee, Lawrence [5 ,6 ]
Williams, Brian A. [7 ,8 ]
van Bakel, Harm [9 ]
Schadt, Eric E. [9 ]
Reijo-Pera, Renee A. [3 ]
Underwood, Jason G. [5 ,10 ]
Wong, Wing Hung [1 ,2 ]
机构
[1] Stanford Univ, Dept Stat, Stanford, CA 94305 USA
[2] Stanford Univ, Dept Hlth Res & Policy, Stanford, CA 94305 USA
[3] Stanford Univ, Dept Obstet & Gynecol, Inst Stem Cell Biol & Regenerat Med, Ctr Human Pluripotent Stem Cell Res & Educ, Stanford, CA 94305 USA
[4] Stanford Univ, Sch Engn, Dept Elect Engn, Stanford, CA 94305 USA
[5] Pacific Biosci Calif, Menlo Pk, CA 94025 USA
[6] Invitae Inc, San Francisco, CA 94107 USA
[7] CALTECH, Div Biol, Pasadena, CA 91125 USA
[8] CALTECH, Beckman Inst, Pasadena, CA 91125 USA
[9] Mt Sinai Sch Med, Dept Genet & Genom Sci, New York, NY 10029 USA
[10] Univ Washington, Dept Genome Sci, Seattle, WA 98105 USA
关键词
isoform discovery; PacBio; hESC transcriptome; alternative splicing; lncNRA; RNA-SEQ DATA; ISOFORM DISCOVERY; SPLICE JUNCTIONS; QUANTIFICATION; ANNOTATION; GENERATION;
D O I
10.1073/pnas.1320101110
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
Although transcriptional and posttranscriptional events are detected in RNA-Seq data from second-generation sequencing, fulllength mRNA isoforms are not captured. On the other hand, third-generation sequencing, which yields much longer reads, has current limitations of lower raw accuracy and throughput. Here, we combine second-generation sequencing and third-generation sequencing with a custom-designed method for isoform identification and quantification to generate a high-confidence isoform dataset for human embryonic stem cells (hESCs). We report 8,084 RefSeq-annotated isoforms detected as full-length and an additional 5,459 isoforms predicted through statistical inference. Over one-third of these are novel isoforms, including 273 RNAs from gene loci that have not previously been identified. Further characterization of the novel loci indicates that a subset is expressed in pluripotent cells but not in diverse fetal and adult tissues; moreover, their reduced expression perturbs the network of pluripotency-associated genes. Results suggest that gene identification, even in well-characterized human cell lines and tissues, is likely far from complete.
引用
收藏
页码:E4821 / E4830
页数:10
相关论文
共 32 条
[1]   RAPID CDNA SEQUENCING (EXPRESSED SEQUENCE TAGS) FROM A DIRECTIONALLY CLONED HUMAN INFANT BRAIN CDNA LIBRARY [J].
ADAMS, MD ;
SOARES, MB ;
KERLAVAGE, AR ;
FIELDS, C ;
VENTER, JC .
NATURE GENETICS, 1993, 4 (04) :373-386
[2]   Improving PacBio Long Read Accuracy by Short Read Alignment [J].
Au, Kin Fai ;
Underwood, Jason G. ;
Lee, Lawrence ;
Wong, Wing Hung .
PLOS ONE, 2012, 7 (10)
[3]   Detection of splice junctions from paired-end RNA-seq data by SpliceMap [J].
Au, Kin Fai ;
Jiang, Hui ;
Lin, Lan ;
Xing, Yi ;
Wong, Wing Hung .
NUCLEIC ACIDS RESEARCH, 2010, 38 (14) :4570-4578
[4]   Integrative annotation of human large intergenic noncoding RNAs reveals global properties and specific subclasses [J].
Cabili, Moran N. ;
Trapnell, Cole ;
Goff, Loyal ;
Koziol, Magdalena ;
Tazon-Vega, Barbara ;
Regev, Aviv ;
Rinn, John L. .
GENES & DEVELOPMENT, 2011, 25 (18) :1915-1927
[5]   Argonaute HITS-CLIP decodes microRNA-mRNA interaction maps [J].
Chi, Sung Wook ;
Zang, Julie B. ;
Mele, Aldo ;
Darnell, Robert B. .
NATURE, 2009, 460 (7254) :479-486
[6]   The GENCODE v7 catalog of human long noncoding RNAs: Analysis of their gene structure, evolution, and expression [J].
Derrien, Thomas ;
Johnson, Rory ;
Bussotti, Giovanni ;
Tanzer, Andrea ;
Djebali, Sarah ;
Tilgner, Hagen ;
Guernec, Gregory ;
Martin, David ;
Merkel, Angelika ;
Knowles, David G. ;
Lagarde, Julien ;
Veeravalli, Lavanya ;
Ruan, Xiaoan ;
Ruan, Yijun ;
Lassmann, Timo ;
Carninci, Piero ;
Brown, James B. ;
Lipovich, Leonard ;
Gonzalez, Jose M. ;
Thomas, Mark ;
Davis, Carrie A. ;
Shiekhattar, Ramin ;
Gingeras, Thomas R. ;
Hubbard, Tim J. ;
Notredame, Cedric ;
Harrow, Jennifer ;
Guigo, Roderic .
GENOME RESEARCH, 2012, 22 (09) :1775-1789
[7]   Landscape of transcription in human cells [J].
Djebali, Sarah ;
Davis, Carrie A. ;
Merkel, Angelika ;
Dobin, Alex ;
Lassmann, Timo ;
Mortazavi, Ali ;
Tanzer, Andrea ;
Lagarde, Julien ;
Lin, Wei ;
Schlesinger, Felix ;
Xue, Chenghai ;
Marinov, Georgi K. ;
Khatun, Jainab ;
Williams, Brian A. ;
Zaleski, Chris ;
Rozowsky, Joel ;
Roeder, Maik ;
Kokocinski, Felix ;
Abdelhamid, Rehab F. ;
Alioto, Tyler ;
Antoshechkin, Igor ;
Baer, Michael T. ;
Bar, Nadav S. ;
Batut, Philippe ;
Bell, Kimberly ;
Bell, Ian ;
Chakrabortty, Sudipto ;
Chen, Xian ;
Chrast, Jacqueline ;
Curado, Joao ;
Derrien, Thomas ;
Drenkow, Jorg ;
Dumais, Erica ;
Dumais, Jacqueline ;
Duttagupta, Radha ;
Falconnet, Emilie ;
Fastuca, Meagan ;
Fejes-Toth, Kata ;
Ferreira, Pedro ;
Foissac, Sylvain ;
Fullwood, Melissa J. ;
Gao, Hui ;
Gonzalez, David ;
Gordon, Assaf ;
Gunawardena, Harsha ;
Howald, Cedric ;
Jha, Sonali ;
Johnson, Rory ;
Kapranov, Philipp ;
King, Brandon .
NATURE, 2012, 489 (7414) :101-108
[8]   An integrated encyclopedia of DNA elements in the human genome [J].
Dunham, Ian ;
Kundaje, Anshul ;
Aldred, Shelley F. ;
Collins, Patrick J. ;
Davis, CarrieA. ;
Doyle, Francis ;
Epstein, Charles B. ;
Frietze, Seth ;
Harrow, Jennifer ;
Kaul, Rajinder ;
Khatun, Jainab ;
Lajoie, Bryan R. ;
Landt, Stephen G. ;
Lee, Bum-Kyu ;
Pauli, Florencia ;
Rosenbloom, Kate R. ;
Sabo, Peter ;
Safi, Alexias ;
Sanyal, Amartya ;
Shoresh, Noam ;
Simon, Jeremy M. ;
Song, Lingyun ;
Trinklein, Nathan D. ;
Altshuler, Robert C. ;
Birney, Ewan ;
Brown, James B. ;
Cheng, Chao ;
Djebali, Sarah ;
Dong, Xianjun ;
Dunham, Ian ;
Ernst, Jason ;
Furey, Terrence S. ;
Gerstein, Mark ;
Giardine, Belinda ;
Greven, Melissa ;
Hardison, Ross C. ;
Harris, Robert S. ;
Herrero, Javier ;
Hoffman, Michael M. ;
Iyer, Sowmya ;
Kellis, Manolis ;
Khatun, Jainab ;
Kheradpour, Pouya ;
Kundaje, Anshul ;
Lassmann, Timo ;
Li, Qunhua ;
Lin, Xinying ;
Marinov, Georgi K. ;
Merkel, Angelika ;
Mortazavi, Ali .
NATURE, 2012, 489 (7414) :57-74
[9]   Reassessing the Determinants of Breeding Synchrony in Ungulates [J].
English, Annie K. ;
Chauvenet, Alienor L. M. ;
Safi, Kamran ;
Pettorelli, Nathalie .
PLOS ONE, 2012, 7 (07)
[10]   Ensembl 2011 [J].
Flicek, Paul ;
Amode, M. Ridwan ;
Barrell, Daniel ;
Beal, Kathryn ;
Brent, Simon ;
Chen, Yuan ;
Clapham, Peter ;
Coates, Guy ;
Fairley, Susan ;
Fitzgerald, Stephen ;
Gordon, Leo ;
Hendrix, Maurice ;
Hourlier, Thibaut ;
Johnson, Nathan ;
Kaehaeri, Andreas ;
Keefe, Damian ;
Keenan, Stephen ;
Kinsella, Rhoda ;
Kokocinski, Felix ;
Kulesha, Eugene ;
Larsson, Pontus ;
Longden, Ian ;
McLaren, William ;
Overduin, Bert ;
Pritchard, Bethan ;
Riat, Harpreet Singh ;
Rios, Daniel ;
Ritchie, Graham R. S. ;
Ruffier, Magali ;
Schuster, Michael ;
Sobral, Daniel ;
Spudich, Giulietta ;
Tang, Y. Amy ;
Trevanion, Stephen ;
Vandrovcova, Jana ;
Vilella, Albert J. ;
White, Simon ;
Wilder, Steven P. ;
Zadissa, Amonida ;
Zamora, Jorge ;
Aken, Bronwen L. ;
Birney, Ewan ;
Cunningham, Fiona ;
Dunham, Ian ;
Durbin, Richard ;
Fernandez-Suarez, Xose M. ;
Herrero, Javier ;
Hubbard, Tim J. P. ;
Parker, Anne ;
Proctor, Glenn .
NUCLEIC ACIDS RESEARCH, 2011, 39 :D800-D806