Large-Scale Similarity Search Profiling of ChEMBL Compound Data Sets

被引:64
作者
Heikamp, Kathrin [1 ]
Bajorath, Juergen [1 ]
机构
[1] Univ Bonn, Dept Life Sci Informat, B IT, LIMES Program Unit Chem Biol & Med Chem, D-53113 Bonn, Germany
关键词
FINGERPRINTS; RECOMBINATION;
D O I
10.1021/ci200199u
中图分类号
R914 [药物化学];
学科分类号
100701 ;
摘要
A large-scale similarity search investigation has been carried out on 266 well-defined compound activity classes extracted from the ChEMBL database. The analysis was performed using two widely applied two-dimensional (2D) fingerprints that mark opposite ends of the current performance spectrum of these types of fingerprints, i.e., MACCS structural keys and the extended connectivity fingerprint with bond diameter four (ECFP4). For each fingerprint, three nearest neighbor search strategies were applied. On the basis of these search calculations, a similarity search profile of the ChEMBL database was generated. Overall, the fingerprint search campaign was surprisingly successful. In 203 of 266 test cases (similar to 76%), a compound recovery rate of at least 50% was observed with at least the better performing fingerprint and one search strategy. The similarity search profile also revealed several general trends. For example, fingerprint searching was often characterized by an early enrichment of active compounds in database selection sets. In addition, compound activity classes have been categorized according to different similarity search performance levels, which helps to put the results of benchmark calculations into perspective. Therefore, a compendium of activity classes falling into different search performance categories is provided. On the basis of our large-scale investigation, the performance range of state-of-the-art 2D fingerprinting has been delineated for compound data sets directed against a wide spectrum of pharmaceutical targets.
引用
收藏
页码:1831 / 1839
页数:9
相关论文
共 30 条
[1]  
*ACC, 2011, MDL DRUG DAT REP
[2]  
[Anonymous], 2010, SCIT PIP PIL
[3]  
[Anonymous], MACCS STRUCT KEYS
[4]  
[Anonymous], 2010, MOL OP ENV
[5]   The properties of known drugs .1. Molecular frameworks [J].
Bemis, GW ;
Murcko, MA .
JOURNAL OF MEDICINAL CHEMISTRY, 1996, 39 (15) :2887-2893
[6]   The use of the area under the roc curve in the evaluation of machine learning algorithms [J].
Bradley, AP .
PATTERN RECOGNITION, 1997, 30 (07) :1145-1159
[7]  
ChEMBL, 2011, CHEMBL
[8]   Molecular similarity analysis in virtual screening: foundations, limitations and novel approaches [J].
Eckert, Hanna ;
Bojorath, Juergen .
DRUG DISCOVERY TODAY, 2007, 12 (5-6) :225-233
[9]  
Gardiner EJ, 2011, FUTURE MED CHEM, V3, P405, DOI [10.4155/FMC.11.4, 10.4155/fmc.11.4]
[10]   Current Trends in Ligand-Based Virtual Screening: Molecular Representations, Data Mining Methods, New Application Areas, and Performance Evaluation [J].
Geppert, Hanna ;
Vogt, Martin ;
Bajorath, Juergen .
JOURNAL OF CHEMICAL INFORMATION AND MODELING, 2010, 50 (02) :205-216