IDENTIFYING POTENTIAL TRANSFER-RNA GENES IN GENOMIC DNA-SEQUENCES

被引:104
作者
FICHANT, GA
BURKS, C
机构
[1] Theoretical Biology and Biophysics Group T-10, MS K710 Los Alamos National Laboratory Los Alamos
关键词
DNA SEQUENCE DATABASES; GENE STRUCTURE; PATTERN RECOGNITION; MACHINE LEARNING; TRANSFER RNA;
D O I
10.1016/0022-2836(91)90108-I
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
We have developed an algorithm that automatically and reproducibly identifies potential tRNA genes in genomic DNA sequences, and we present a general strategy for testing the sensitivity of such algorithms. This algorithm is useful for the flagging and characterization of long genomic sequences that have not been experimentally analyzed for identification of functional regions, and for the scanning of nucleotide sequence databases for errors in the sequences and the functional assignments associated with them. In an exhaustive scan of the GenBank database, 97·5% of the 744 known tRNA genes were correctly identified (true-positives), and 42 previously unidentified sequences were predicted to be tRNAs. A detailed analysis of these latter predictions reveals that 16 of the 42 are very similar to known tRNA genes, and we predict that they do, in fact, code for tRNA, yielding a false-positive rate for the algorithm of 0·003%. The new algorithm and testing strategy are a considerable improvement over any previously described strategies for recognizing tRNA genes, and they allow detections of genes (including introns) embedded in long genomic sequences. © 1991.
引用
收藏
页码:659 / 671
页数:13
相关论文
共 48 条
[1]   NUCLEOTIDE DISTRIBUTION AND THE RECOGNITION OF CODING REGIONS IN DNA-SEQUENCES - AN INFORMATION-THEORY APPROACH [J].
ALMAGOR, H .
JOURNAL OF THEORETICAL BIOLOGY, 1985, 117 (01) :127-136
[2]   CODON RECOGNITION PATTERNS AS DEDUCED FROM SEQUENCES OF THE COMPLETE SET OF TRANSFER-RNA SPECIES IN MYCOPLASMA-CAPRICOLUM - RESEMBLANCE TO MITOCHONDRIA [J].
ANDACHI, Y ;
YAMAO, F ;
MUTO, A ;
OSAWA, S .
JOURNAL OF MOLECULAR BIOLOGY, 1989, 209 (01) :37-54
[4]  
BURKS C, 1990, UNPUB WORKSHOP DETEC
[5]  
BURKS C, 1989, BIOMOLECULAR DATA RE, P17
[6]  
BURKS C, 1990, METHOD ENZYMOL, V183, P3
[7]  
Chamberlin J.M., 1976, RNA POLYM, V17, P17
[8]   HEURISTIC INFORMATIONAL ANALYSIS OF SEQUENCES [J].
CLAVERIE, JM ;
BOUGUELERET, L .
NUCLEIC ACIDS RESEARCH, 1986, 14 (01) :179-196
[9]  
FICHANT G, 1987, COMPUT APPL BIOSCI, V3, P287
[10]  
FICHANT G, 1989, BIOMETRIE DONNEES DI, P65