Improving the prediction of disease-related variants using protein three-dimensional structure

被引:100
作者
Capriotti, Emidio [1 ,3 ]
Altman, Russ B. [1 ,2 ]
机构
[1] Stanford Univ, Dept Bioengn, Stanford, CA 94305 USA
[2] Stanford Univ, Dept Genet, Stanford, CA 94305 USA
[3] Univ Balearic Isl, Dept Math & Comp Sci, Palma De Mallorca, Spain
来源
BMC BIOINFORMATICS | 2011年 / 12卷
关键词
SINGLE-NUCLEOTIDE POLYMORPHISMS; AMINO-ACID SUBSTITUTIONS; SUPPORT VECTOR MACHINES; NON-SYNONYMOUS SNPS; STABILITY CHANGES; EVOLUTIONARY INFORMATION; POINT MUTATIONS; HUMAN GENOME; SEQUENCE; DATABASE;
D O I
10.1186/1471-2105-12-S4-S3
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background: Single Nucleotide Polymorphisms (SNPs) are an important source of human genome variability. Non-synonymous SNPs occurring in coding regions result in single amino acid polymorphisms (SAPs) that may affect protein function and lead to pathology. Several methods attempt to estimate the impact of SAPs using different sources of information. Although sequence-based predictors have shown good performance, the quality of these predictions can be further improved by introducing new features derived from three-dimensional protein structures. Results: In this paper, we present a structure-based machine learning approach for predicting disease-related SAPs. We have trained a Support Vector Machine (SVM) on a set of 3,342 disease-related mutations and 1,644 neutral polymorphisms from 784 protein chains. We use SVM input features derived from the protein's sequence, structure, and function. After dataset balancing, the structure-based method (SVM-3D) reaches an overall accuracy of 85%, a correlation coefficient of 0.70, and an area under the receiving operating characteristic curve (AUC) of 0.92. When compared with a similar sequence-based predictor, SVM-3D results in an increase of the overall accuracy and AUC by 3%, and correlation coefficient by 0.06. The robustness of this improvement has been tested on different datasets and in all the cases SVM-3D performs better than previously developed methods even when compared with PolyPhen2, which explicitly considers in input protein structure information. Conclusion: This work demonstrates that structural information can increase the accuracy of disease-related SAPs identification. Our results also quantify the magnitude of improvement on a large dataset. This improvement is in agreement with previously observed results, where structure information enhanced the prediction of protein stability changes upon mutation. Although the structural information contained in the Protein Data Bank is limiting the application and the performance of our structure-based method, we expect that SVM-3D will result in higher accuracy when more structural date become available.
引用
收藏
页数:11
相关论文
共 43 条
[31]   CUPSAT: prediction of protein stability upon point mutations [J].
Parthiban, Vijaya ;
Gromiha, M. Michael ;
Schomburg, Dietmar .
NUCLEIC ACIDS RESEARCH, 2006, 34 :W239-W242
[32]   AL2CO: calculation of positional conservation in a protein sequence alignment [J].
Pei, JM ;
Grishin, NV .
BIOINFORMATICS, 2001, 17 (08) :700-712
[33]   Human non-synonymous SNPs: server and survey [J].
Ramensky, V ;
Bork, P ;
Sunyaev, S .
NUCLEIC ACIDS RESEARCH, 2002, 30 (17) :3894-3900
[34]   dbSNP: the NCBI database of genetic variation [J].
Sherry, ST ;
Ward, MH ;
Kholodov, M ;
Baker, J ;
Phan, L ;
Smigielski, EM ;
Sirotkin, K .
NUCLEIC ACIDS RESEARCH, 2001, 29 (01) :308-311
[35]   Coding sing le-nucleotide polymorphisms associated with complex vs. Mendelian disease: Evolutionary evidence for differences in molecular effects [J].
Thomas, PD ;
Kejariwal, A .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 2004, 101 (43) :15398-15403
[36]   PANTHER: a browsable database of gene products organized by biological function, using curated protein family and subfamily classification [J].
Thomas, PD ;
Kejariwal, A ;
Campbell, MJ ;
Mi, HY ;
Diemer, K ;
Guo, N ;
Ladunga, I ;
Ulitsky-Lazareva, B ;
Muruganujan, A ;
Rabkin, S ;
Vandergriff, JA ;
Doremieux, O .
NUCLEIC ACIDS RESEARCH, 2003, 31 (01) :334-341
[37]   MuD: an interactive web server for the prediction of non-neutral substitutions using protein structural data [J].
Wainreb, Gilad ;
Ashkenazy, Haim ;
Bromberg, Yana ;
Starovolsky-Shitrit, Alina ;
Haliloglu, Turkan ;
Ruppin, Eytan ;
Avraham, Karen B. ;
Rost, Burkhard ;
Ben-Tal, Nir .
NUCLEIC ACIDS RESEARCH, 2010, 38 :W523-W528
[38]   SNPs, protein structure, and disease [J].
Wang, Z ;
Moult, J .
HUMAN MUTATION, 2001, 17 (04) :263-270
[39]   Aromatic interactions in peptides: Impact on structure and function [J].
Waters, ML .
BIOPOLYMERS, 2004, 76 (05) :435-445
[40]   Finding new structural and sequence attributes to predict possible disease association of single amino acid lpolymorphism (SAP) [J].
Ye, Zhi-Qiang ;
Zhao, Shu-Qi ;
Gao, Ge ;
Liu, Xiao-Qiao ;
Langlois, Robert E. ;
Lu, Hui ;
Wei, Liping .
BIOINFORMATICS, 2007, 23 (12) :1444-1450