A comprehensive transcript index of the human genome generated using microarrays and computational approaches

被引:76
作者
Schadt, EE
Edwards, SW
GuhaThakurta, D
Holder, D
Ying, L
Svetnik, V
Leonardson, A
Hart, KW
Russell, A
Li, GY
Cavet, G
Castle, J
McDonagh, P
Kan, ZY
Chen, RH
Kasarskis, A
Margarint, M
Caceres, RM
Johnson, JM
Armour, CD
Garrett-Engele, PW
Tsinoremas, NF
Shoemaker, DD
机构
[1] Rosetta Inpharmat LLC, Kirkland, WA 98034 USA
[2] Merck Res Labs, Westpoint, PA USA
[3] Rally Sci, Watertown, MA 02472 USA
[4] Amgen Inc, Seattle, WA 98119 USA
[5] Scripps Res Inst, Jupiter, FL 33458 USA
关键词
D O I
10.1186/gb-2004-5-10-r73
中图分类号
Q81 [生物工程学(生物技术)]; Q93 [微生物学];
学科分类号
071005 ; 0836 ; 090102 ; 100705 ;
摘要
Background: Computational and microarray-based experimental approaches were used to generate a comprehensive transcript index for the human genome. Oligonucleotide probes designed from approximately 50,000 known and predicted transcript sequences from the human genome were used to survey transcription from a diverse set of 60 tissues and cell lines using ink-jet microarrays. Further, expression activity over at least six conditions was more generally assessed using genomic tiling arrays consisting of probes tiled through a repeat-masked version of the genomic sequence making up chromosomes 20 and 22. Results: The combination of microarray data with extensive genome annotations resulted in a set of 28,456 experimentally supported transcripts. This set of high-confidence transcripts represents the first experimentally driven annotation of the human genome. In addition, the results from genomic tiling suggest that a large amount of transcription exists outside of annotated regions of the genome and serves as an example of how this activity could be measured on a genome-wide scale. Conclusions: These data represent one of the most comprehensive assessments of transcriptional activity in the human genome and provide an atlas of human gene expression over a unique set of gene predictions. Before the annotation of the human genome is considered complete, however, the previously unannotated transcriptional activity throughout the genome must be fully characterized.
引用
收藏
页数:17
相关论文
共 48 条
[1]  
ADAMS MD, 1995, NATURE, V377, P3
[2]   Gapped BLAST and PSI-BLAST: a new generation of protein database search programs [J].
Altschul, SF ;
Madden, TL ;
Schaffer, AA ;
Zhang, JH ;
Zhang, Z ;
Miller, W ;
Lipman, DJ .
NUCLEIC ACIDS RESEARCH, 1997, 25 (17) :3389-3402
[3]  
Bateman A, 2004, NUCLEIC ACIDS RES, V32, pD138, DOI [10.1093/nar/gkp985, 10.1093/nar/gkr1065, 10.1093/nar/gkh121]
[4]   AVID: A global alignment program [J].
Bray, N ;
Dubchak, I ;
Pachter, L .
GENOME RESEARCH, 2003, 13 (01) :97-102
[5]   Prediction of complete gene structures in human genomic DNA [J].
Burge, C ;
Karlin, S .
JOURNAL OF MOLECULAR BIOLOGY, 1997, 268 (01) :78-94
[6]   d2_cluster: A validated method for clustering EST and full-length cDNA sequences [J].
Burke, J ;
Davison, D ;
Hide, W .
GENOME RESEARCH, 1999, 9 (11) :1135-1142
[7]   The contribution of 700,000 ORF sequence tags to the definition of the human transcriptome [J].
Camargo, AA ;
Samaia, HPB ;
Dias-Neto, E ;
Simao, DF ;
Migotto, IA ;
Briones, MRS ;
Costa, FF ;
Nagai, MA ;
Verjovski-Almeida, S ;
Zago, MA ;
Andrade, LEC ;
Carrer, H ;
El-Dorry, HFA ;
Espreafico, EM ;
Habr-Gama, A ;
Giannella-Neto, D ;
Goldman, GH ;
Gruber, A ;
Hackel, C ;
Kimura, ET ;
Maciel, RMB ;
Marie, SKN ;
Martins, EAL ;
Nóbrega, MP ;
Paçó-Larson, ML ;
Pardini, MIMC ;
Pereira, GG ;
Pesquero, JB ;
Rodrigues, V ;
Rogatto, SR ;
da Silva, IDCG ;
Sogayar, MC ;
Sonati, MDF ;
Tajara, EH ;
Valentini, SR ;
Alberto, FL ;
Amaral, MEJ ;
Aneas, I ;
Arnaldi, LAT ;
de Assis, AM ;
Bengtson, MH ;
Bergamo, NA ;
Bombonato, V ;
de Camargo, MER ;
Canevari, RA ;
Carraro, DM ;
Cerutti, JM ;
Corrêa, MLC ;
Corrêa, RFR ;
Costa, MCR .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 2001, 98 (21) :12103-12108
[8]   Optimization of oligonucleotide arrays and RNA amplification protocols for analysis of transcript structure and alternative splicing [J].
Castle, J ;
Garrett-Engele, P ;
Armour, CD ;
Duenwald, SJ ;
Loerch, PM ;
Meyer, MR ;
Schadt, EE ;
Stoughton, R ;
Parrish, ML ;
Shoemaker, DD ;
Johnson, JM .
GENOME BIOLOGY, 2003, 4 (10)
[9]   Unbiased mapping of transcription factor binding sites along human chromosomes 21 and 22 points to widespread regulation of noncoding RNAs [J].
Cawley, S ;
Bekiranov, S ;
Ng, HH ;
Kapranov, P ;
Sekinger, EA ;
Kampa, D ;
Piccolboni, A ;
Sementchenko, V ;
Cheng, J ;
Williams, AJ ;
Wheeler, R ;
Wong, B ;
Drenkow, J ;
Yamanaka, M ;
Patel, S ;
Brubaker, S ;
Tammana, H ;
Helt, G ;
Struhl, K ;
Gingeras, TR .
CELL, 2004, 116 (04) :499-509
[10]   Computational methods for the identification of genes in vertebrate genomic sequences [J].
Claverie, JM .
HUMAN MOLECULAR GENETICS, 1997, 6 (10) :1735-1744