Similarity searching of chemical databases using atom environment descriptors (MOLPRINT 2D): Evaluation of performance

被引:283
作者
Bender, A
Mussa, HY
Glen, RC
Reiling, S
机构
[1] Univ Cambridge, Dept Chem, Unilever Ctr Mol Sci Informat, Cambridge CB2 1EW, England
[2] Aventis, DI&A, Bridgewater, NJ 08807 USA
来源
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES | 2004年 / 44卷 / 05期
关键词
D O I
10.1021/ci0498719
中图分类号
O6 [化学];
学科分类号
0703 ;
摘要
A molecular similarity searching technique based on atom environments, information-gain-based feature selection, and the naive Bayesian classifier has been applied to a series of diverse datasets and its performance compared to those of alternative searching methods. Atom environments are count vectors of heavy atoms present at a topological distance from each heavy atom of a molecular structure. In this application, using a recently published dataset of more than 100000 molecules from the MDL Drug Data Report database, the atom environment approach appears to outperform fusion of ranking scores as well as binary kernel discrimination, which are both used in combination with Unity fingerprints. Overall retrieval rates among the top 5% of the sorted library are nearly 10% better (more than 14% better in relative numbers) than those of the second best method, Unity fingerprints and binary kernel discrimination. In 10 out of 11 sets of active compounds the combination of atom environments and the naive Bayesian classifier appears to be the superior method, while in the remaining dataset, data fusion and binary kernel discrimination in combination with Unity fingerprints is the method of choice. Binary kernel discrimination in combination with Unity fingerprints generally comes second in performance overall. The difference in performance can largely be attributed to the different molecular descriptors used. Atom environments outperform Unity fingerprints by a large margin if the combination of these descriptors with the Tanimoto coefficient is compared. The naive Bayesian classifier in combination with information-gain-based feature selection and selection of a sensible number of features performs about as well as binary kernel discrimination in experiments where these classification methods are compared. When used on a monoaminooxidase dataset, atom environments and the naive Bayesian classifier perform as well as binary kernel discrimination in the case of a 50/50 split of training and test compounds. In the case of sparse training data, binary kernel discrimination is found to be superior on this particular dataset. On a third dataset, the atom environment descriptor shows higher retrieval rates than other 2D fingerprints tested here when used in combination with the Tanimoto similarity coefficient. Feature selection is shown to be a crucial step in determining the performance of the algorithm. The representation of molecules by atom environments is found to be more effective than Unity fingerprints for the type of biological receptor similarity calculations examined here. Combining information prior to scoring and including information about inactive compounds, as in the Bayesian classifier and binary kernel discrimination, is found to be superior to posterior data fusion (in the datasets tested here).
引用
收藏
页码:1708 / 1718
页数:11
相关论文
共 48 条
[1]   STRATEGIC CONSIDERATIONS IN DESIGN OF A SCREENING SYSTEM FOR SUBSTRUCTURE SEARCHES OF CHEMICAL STRUCTURE FILES [J].
ADAMSON, GW ;
COWELL, J ;
MCLURE, AHW ;
TOWN, WG ;
YAPP, AM ;
LYNCH, MF .
JOURNAL OF CHEMICAL DOCUMENTATION, 1973, 13 (03) :153-157
[2]   Uniform-length molecular descriptors for quantitative structure-property relationships (QSPR) and quantitative structure-activity relationships (QSAR): classification studies and similarity searching [J].
Baumann, K .
TRAC-TRENDS IN ANALYTICAL CHEMISTRY, 1999, 18 (01) :36-46
[3]   Molecular similarity searching using atom environments, information-based feature selection, and a naive Bayesian classifier [J].
Bender, A ;
Mussa, HY ;
Glen, RC ;
Reiling, S .
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES, 2004, 44 (01) :170-178
[4]   Molecular similarity based on DOCK-generated fingerprints [J].
Briem, H ;
Kuntz, ID .
JOURNAL OF MEDICINAL CHEMISTRY, 1996, 39 (17) :3401-3408
[5]   In vitro and in silico affinity fingerprints:: Finding similarities beyond structural classes [J].
Briem, H ;
Lessel, UF .
PERSPECTIVES IN DRUG DISCOVERY AND DESIGN, 2000, 20 (01) :231-244
[6]   Use of structure Activity data to compare structure-based clustering methods and descriptors for use in compound selection [J].
Brown, RD ;
Martin, YC .
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES, 1996, 36 (03) :572-584
[7]   VALIDATION OF THE GENERAL-PURPOSE TRIPOS 5.2 FORCE-FIELD [J].
CLARK, M ;
CRAMER, RD ;
VANOPDENBOSCH, N .
JOURNAL OF COMPUTATIONAL CHEMISTRY, 1989, 10 (08) :982-1012
[8]   COMPARATIVE MOLECULAR-FIELD ANALYSIS (COMFA) .1. EFFECT OF SHAPE ON BINDING OF STEROIDS TO CARRIER PROTEINS [J].
CRAMER, RD ;
PATTERSON, DE ;
BUNCE, JD .
JOURNAL OF THE AMERICAN CHEMICAL SOCIETY, 1988, 110 (18) :5959-5967
[9]  
*DAYL INC, DAYL VERS 4 62
[10]   The hidden component of size in two-dimensional fragment descriptors: Side effects on sampling in bioactive libraries [J].
Dixon, SL ;
Koehler, RT .
JOURNAL OF MEDICINAL CHEMISTRY, 1999, 42 (15) :2887-2900