Gene structure prediction and alternative splicing analysis using genomically aligned ESTs

被引:289
作者
Kan, ZY
Rouchka, EC
Gish, WR
States, DJ [1 ]
机构
[1] Washington Univ, Ctr Computat Biol, St Louis, MO 63110 USA
[2] Washington Univ, Dept Genet, St Louis, MO 63110 USA
关键词
D O I
10.1101/gr.155001
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
With the availability of a nearly complete sequence of the human genome, aligning expressed sequence tags (EST) to the genomic sequence has become a practical and powerful strategy for gene prediction. Elucidating gene structure is a complex problem requiring the identification of splice junctions, gene boundaries, and alternative splicing variants. We have developed a software tool, Transcript Assembly Program (TAP), to delineate gene structures using genomically aligned EST sequences. TAP assembles the joint gene structure of the entire genomic region from individual splice junction pairs, using a novel algorithm that uses the EST-encoded connectivity and redundancy information to sort out the complex alternative splicing patterns. A method called polyadenylation site scan (PASS) has been developed to detect poly-A sites in the genome. TAP uses these predictions to identify gene boundaries by segmenting the joint gene structure at polyadenylated terminal exons. Reconstructing 1007 known transcripts, TAP scored a sensitivity (Sn) of 60% and a specificity (Sp) of 92% at the exon level. The gene boundary identification process was found to be accurate 78% of the time. TAP also reports alternative splicing patterns in EST alignments. An analysis of alternative splicing in 1124 genic regions suggested that more than half of human genes undergo alternative splicing. Surprisingly, we saw an absolute majority of the detected alternative splicing events affect the coding region. Furthermore, the evolutionary conservation of alternative splicing between human and mouse was analyzed using an EST-based approach. (See http://stl.wustl.edu/-zkan/TAP/).
引用
收藏
页码:889 / 900
页数:12
相关论文
共 30 条
  • [1] BAFNA V, 2000, INTELL SYST MOL BIOL, V8, P3
  • [2] Human and mouse gene structure: Comparative analysis and application to exon prediction
    Batzoglou, S
    Pachter, L
    Mesirov, JP
    Berger, B
    Lander, ES
    [J]. GENOME RESEARCH, 2000, 10 (07) : 950 - 958
  • [3] MaskerAid:: a performance enhancement to RepeatMasker
    Bedell, JA
    Korf, I
    Gish, W
    [J]. BIOINFORMATICS, 2000, 16 (11) : 1040 - 1041
  • [4] Comparison of gene indexing databases
    Bouck, J
    Yu, W
    Gibbs, R
    Worley, K
    [J]. TRENDS IN GENETICS, 1999, 15 (04) : 159 - 162
  • [5] EST comparison indicates 38% of human mRNAs contain possible alternative splice forms
    Brett, D
    Hanke, J
    Lehmann, G
    Haase, S
    Delbrück, S
    Krueger, S
    Reich, J
    Bork, P
    [J]. FEBS LETTERS, 2000, 474 (01) : 83 - 86
  • [6] Prediction of complete gene structures in human genomic DNA
    Burge, C
    Karlin, S
    [J]. JOURNAL OF MOLECULAR BIOLOGY, 1997, 268 (01) : 78 - 94
  • [7] Alternative gene form discovery and candidate gene selection from gene indexing projects
    Burke, J
    Wang, H
    Hide, W
    Davison, DB
    [J]. GENOME RESEARCH, 1998, 8 (03): : 276 - 290
  • [8] Evaluation of gene structure prediction programs
    Burset, M
    Guigo, R
    [J]. GENOMICS, 1996, 34 (03) : 353 - 367
  • [9] Computational methods for the identification of genes in vertebrate genomic sequences
    Claverie, JM
    [J]. HUMAN MOLECULAR GENETICS, 1997, 6 (10) : 1735 - 1744
  • [10] Analysis of expressed sequence tags indicates 35,000 human genes
    Ewing, B
    Green, P
    [J]. NATURE GENETICS, 2000, 25 (02) : 232 - 234