Comparison of support vector machine and artificial neural network systems for drug/nondrug classification

被引:439
作者
Byvatov, E
Fechner, U
Sadowski, J
Schneider, G
机构
[1] Goethe Univ Frankfurt, Inst Organ Chem & Chem Biol, D-60439 Frankfurt, Germany
[2] AstraZeneca R&D, S-43183 Molndal, Sweden
来源
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES | 2003年 / 43卷 / 06期
关键词
D O I
10.1021/ci0341161
中图分类号
O6 [化学];
学科分类号
0703 ;
摘要
Support vector machine (SVM) and artificial neural network (ANN) systems were applied to a drug/nondrug classification problem as an example of binary decision problems in early-phase virtual compound filtering and screening. The results indicate that solutions obtained by SVM training seem to be more robust with a smaller standard error compared to ANN training. Generally, the SVM classifier yielded slightly higher prediction accuracy than ANN, irrespective of the type of descriptors used for molecule encoding, the size of the training data sets, and the algorithm employed for neural network training. The performance was compared using various different descriptor sets and descriptor combinations based on the 120 standard Ghose-Crippen fragment descriptors, a wide range of 180 different properties and physicochemical descriptors from the Molecular Operating Environment (MOE) package, and 225 topological pharmacophore (CATS) descriptors. For the complete set of 525 descriptors cross-validated classification by SVM yielded 82% correct predictions (Matthews cc = 0.63), whereas ANN reached 80% correct predictions (Matthews cc = 0.58). Although SVM outperformed the ANN classifiers with regard to overall prediction accuracy, both methods were shown to complement each other, as the sets of true positives, false positives (overprediction), true negatives, and false negatives (underprediction) produced by the two classifiers were not identical. The theory of SVM and ANN training is briefly reviewed.
引用
收藏
页码:1882 / 1889
页数:8
相关论文
共 42 条
[1]   Can we learn to distinguish between "drug-like" and "nondrug-like" molecules? [J].
Ajay ;
Walters, WP ;
Murcko, MA .
JOURNAL OF MEDICINAL CHEMISTRY, 1998, 41 (18) :3314-3324
[2]  
Baldi P., 1998, Bioinformatics: The machine learning approach
[3]  
Bishop C. M., 1995, NEURAL NETWORKS PATT
[4]   Drug design by machine learning: support vector machines for pharmaceutical data analysis [J].
Burbidge, R ;
Trotter, M ;
Buxton, B ;
Holden, S .
COMPUTERS & CHEMISTRY, 2001, 26 (01) :5-14
[5]   A tutorial on Support Vector Machines for pattern recognition [J].
Burges, CJC .
DATA MINING AND KNOWLEDGE DISCOVERY, 1998, 2 (02) :121-167
[6]  
Byvatov Evgeny, 2003, Appl Bioinformatics, V2, P67
[7]   Computational methods for the prediction of 'drug-likeness' [J].
Clark, DE ;
Pickett, SD .
DRUG DISCOVERY TODAY, 2000, 5 (02) :49-58
[8]   A reflective Newton method for minimizing a quadratic function subject to bounds on some of the variables [J].
Coleman, TF ;
Li, YY .
SIAM JOURNAL ON OPTIMIZATION, 1996, 6 (04) :1040-1058
[9]  
CORTES C, 1995, MACH LEARN, V20, P273, DOI 10.1023/A:1022627411411
[10]  
Cristianini N, 2000, Intelligent Data Analysis: An Introduction