Assignment of protein sequences to existing domain and family classification systems: Pfam and the PDB

被引:47
作者
Xu, Qifang [1 ]
Dunbrack, Roland L., Jr. [1 ]
机构
[1] Fox Chase Canc Ctr, Inst Canc Res, Philadelphia, PA 19111 USA
关键词
COMPREHENSIVE DATABASE; STRUCTURAL GENOMICS; PSI-BLAST;
D O I
10.1093/bioinformatics/bts533
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Motivation: Automating the assignment of existing domain and protein family classifications to new sets of sequences is an important task. Current methods often miss assignments because remote relationships fail to achieve statistical significance. Some assignments are not as long as the actual domain definitions because local alignment methods often cut alignments short. Long insertions in query sequences often erroneously result in two copies of the domain assigned to the query. Divergent repeat sequences in proteins are often missed. Results: We have developed a multilevel procedure to produce nearly complete assignments of protein families of an existing classification system to a large set of sequences. We apply this to the task of assigning Pfam domains to sequences and structures in the Protein Data Bank (PDB). We found that HHsearch alignments frequently scored more remotely related Pfams in Pfam clans higher than closely related Pfams, thus, leading to erroneous assignment at the Pfam family level. A greedy algorithm allowing for partial overlaps was, thus, applied first to sequence/HMM alignments, then HMM-HMM alignments and then structure alignments, taking care to join partial alignments split by large insertions into single-domain assignments. Additional assignment of repeat Pfams with weaker E-values was allowed after stronger assignments of the repeat HMM. Our database of assignments, presented in a database called PDBfam, contains Pfams for 99.4% of chains > 50 residues.
引用
收藏
页码:2763 / 2772
页数:10
相关论文
共 34 条
[31]   ProtBuD: a database of biological unit structures of protein families and superfamilies [J].
Xu, Qifang ;
Canutescu, Adrian ;
Obradovic, Zoran ;
Dunbrack, Roland L., Jr. .
BIOINFORMATICS, 2006, 22 (23) :2876-2882
[32]   The protein common interface database (ProtCID)-a comprehensive database of interactions of homologous proteins in multiple crystal forms [J].
Xu, Qifang ;
Dunbrack, Roland L., Jr. .
NUCLEIC ACIDS RESEARCH, 2011, 39 :D761-D770
[33]   Flexible structure alignment by chaining aligned fragment pairs allowing twists [J].
Ye, Yuzhen ;
Godzik, Adam .
BIOINFORMATICS, 2003, 19 :II246-II255
[34]   Glutamate-induced apoptosis in primary cortical neurons is inhibited by equine estrogens via down-regulation of caspase-3 and prevention of mitochondrial cytochrome c release [J].
Zhang, YM ;
Bhavnani, BR .
BMC NEUROSCIENCE, 2005, 6 (1)