NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

被引:2210
作者
Pruitt, Kim D. [1 ]
Tatusova, Tatiana [1 ]
Maglott, Donna R. [1 ]
机构
[1] NIH, Natl Ctr Biotechnol Informat, Natl Lib Med, Bethesda, MD 20892 USA
基金
美国国家卫生研究院;
关键词
D O I
10.1093/nar/gkl842
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2 879 860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.
引用
收藏
页码:D61 / D65
页数:5
相关论文
共 12 条
[1]   Gapped BLAST and PSI-BLAST: a new generation of protein database search programs [J].
Altschul, SF ;
Madden, TL ;
Schaffer, AA ;
Zhang, JH ;
Zhang, Z ;
Miller, W ;
Lipman, DJ .
NUCLEIC ACIDS RESEARCH, 1997, 25 (17) :3389-3402
[2]   BASIC LOCAL ALIGNMENT SEARCH TOOL [J].
ALTSCHUL, SF ;
GISH, W ;
MILLER, W ;
MYERS, EW ;
LIPMAN, DJ .
JOURNAL OF MOLECULAR BIOLOGY, 1990, 215 (03) :403-410
[3]  
BENSON DA, 2007, IN PRESS NUCL ACIDS
[4]   The Mouse Genome Database (MGD): updates and enhancements [J].
Blake, Judith A. ;
Eppig, Janan T. ;
Bult, Carol J. ;
Kadin, James A. ;
Richardson, Joel E. .
NUCLEIC ACIDS RESEARCH, 2006, 34 :D562-D567
[5]   Regulation of gene expression by stop codon recoding: selenocysteine [J].
Copeland, PR .
GENE, 2003, 312 :17-25
[6]   FlyBase: genes and gene models [J].
Drysdale, RA ;
Crosby, MA .
NUCLEIC ACIDS RESEARCH, 2005, 33 :D390-D395
[7]  
MAGLOTT D, 2007, IN PRESS NUCL ACIDS
[8]   The Arabidopsis Information Resource (TAIR):: a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community [J].
Rhee, SY ;
Beavis, W ;
Berardini, TZ ;
Chen, GH ;
Dixon, D ;
Doyle, A ;
Garcia-Hernandez, M ;
Huala, E ;
Lander, G ;
Montoya, M ;
Miller, N ;
Mueller, LA ;
Mundodi, S ;
Reiser, L ;
Tacklind, J ;
Weems, DC ;
Wu, YH ;
Xu, I ;
Yoo, D ;
Yoon, J ;
Zhang, PF .
NUCLEIC ACIDS RESEARCH, 2003, 31 (01) :224-228
[9]  
Schuler GD, 1996, METHOD ENZYMOL, V266, P141
[10]   WormBase:: better software, richer content [J].
Schwarz, Erich M. ;
Antoshechkin, Igor ;
Bastiani, Carol ;
Bieri, Tamberlyn ;
Blasiar, Darin ;
Canaran, Payan ;
Chan, Juancarlos ;
Chen, Nansheng ;
Chen, Wen J. ;
Davis, Paul ;
Fiedler, Tristan J. ;
Girard, Lisa ;
Harris, Todd W. ;
Kenny, Eimear E. ;
Kishore, Ranjana ;
Lawson, Dan ;
Lee, Raymond ;
Mueller, Hans-Michael ;
Nakamura, Cecilia ;
Ozersky, Phil ;
Petcherski, Andrei ;
Rogers, Anthony ;
Spooner, Will ;
Tuli, Mary Ann ;
Van Auken, Kimberly ;
Wang, Daniel ;
Durbin, Richard ;
Spieth, John ;
Stein, Lincoln D. ;
Sternberg, Paul W. .
NUCLEIC ACIDS RESEARCH, 2006, 34 :D475-D478