BLASTing small molecules - statistics and extreme statistics of chemical similarity scores

被引:14
作者
Baldi, Pierre [1 ,2 ,3 ]
Benz, Ryan W. [1 ,2 ]
机构
[1] Univ Calif Irvine, Dept Comp Sci, Irvine, CA 92697 USA
[2] Univ Calif Irvine, Inst Genom & Bioinformat, Irvine, CA 92697 USA
[3] Univ Calif Irvine, Dept Biol Chem, Irvine, CA 92697 USA
关键词
D O I
10.1093/bioinformatics/btn187
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Motivation: Small organic molecules, from nucleotides and amino acids to metabolites and drugs, play a fundamental role in chemistry, biology and medicine. As databases of small molecules continue to grow and become more open, it is important to develop the tools to search them efficiently. In order to develop a BLAST-like tool for small molecules, one must first understand the statistical behavior of molecular similarity scores. Results: We develop a new detailed theory of molecular similarity scores that can be applied to a variety of molecular representations and similarity measures. For concreteness, we focus on the most widely used measure - the Tanimoto measure applied to chem-ical fingerprints. In both the case of empirical fingerprints and fingerprints generated by several stochastic models, we derive accurate approximations for both the distribution and extreme value distribution of similarity scores. These approximation are derived using a ratio of correlated Gaussians approach. The theory enables the calculation of significance scores, such as Z-scores and P-values, and the estimation of the top hits list size. Empirical results obtained using both the random models and real data from the ChemDB database are given to corroborate the theory and show how it can be applied to mine chemical space.
引用
收藏
页码:I357 / I365
页数:9
相关论文
共 29 条
[1]  
ACKLEY DH, 1985, COGNITIVE SCI, V9, P147
[2]  
ALTSCHUL SF, 1997, NUCLEIC ACIDS RES, V25, P3402
[3]  
[Anonymous], GRAPHICAL MODELS MAC
[4]   Lossless compression of chemical fingerprints using integer entropy codes improves storage and retrieval [J].
Baldi, Pierre ;
Benz, Ryan W. ;
Hirschberg, Daniel S. ;
Swamidass, S. Joshua .
JOURNAL OF CHEMICAL INFORMATION AND MODELING, 2007, 47 (06) :2098-2109
[5]   Similarity searching of chemical databases using atom environment descriptors (MOLPRINT 2D): Evaluation of performance [J].
Bender, A ;
Mussa, HY ;
Glen, RC ;
Reiling, S .
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES, 2004, 44 (05) :1708-1718
[6]  
Bohacek RS, 1996, MED RES REV, V16, P3, DOI 10.1002/(SICI)1098-1128(199601)16:1<3::AID-MED1>3.3.CO
[7]  
2-D
[8]  
Cedilnik A., 2004, METODOLOSKI ZVEZKI, V1, P99, DOI 10.51936/zweu3253
[9]   ChemDB: a public database of small molecules and related chemoinformatics resources [J].
Chen, J ;
Swamidass, SJ ;
Bruand, J ;
Baldi, P .
BIOINFORMATICS, 2005, 21 (22) :4133-4139
[10]   ChemDB update - full-text search and virtual chemical space [J].
Chen, Jonathan H. ;
Linstead, Erik ;
Swamidass, S. Joshua ;
Wang, Dennis ;
Baldi, Pierre .
BIOINFORMATICS, 2007, 23 (17) :2348-2351