AL2CO: calculation of positional conservation in a protein sequence alignment

被引:350
作者
Pei, JM
Grishin, NV
机构
[1] Univ Texas, SW Med Ctr, Howard Hughes Med Inst, Dallas, TX 75390 USA
[2] Univ Texas, SW Med Ctr, Dept Biochem, Dallas, TX 75390 USA
关键词
D O I
10.1093/bioinformatics/17.8.700
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Motivation: Amino acid sequence alignments are widely used in the analysis of protein structure, function and evolutionary relationships. Proteins within a superfamily usually share the same fold and possess related functions. These structural and functional constraints are reflected in the alignment conservation patterns. Positions of functional and/or structural importance tend to be more conserved. Conserved positions are usually clustered in distinct motifs surrounded by sequence segments of low conservation. Poorly conserved regions might also arise from the imperfections in multiple alignment algorithms and thus indicate possible alignment errors. Quantification of conservation by attributing a conservation index to each aligned position makes motif detection more convenient. Mapping these conservation indices onto a protein spatial structure helps to visualize spatial conservation features of the molecule and to predict functionally and/or structurally important sites. Analysis of conservation indices could be a useful tool in detection of potentially misaligned regions and will aid in improvement of multiple alignments. Results: We developed a program to calculate a conservation index at each position in a multiple sequence alignment using several methods. Namely, amino acid frequencies at each position are estimated and the conservation index is calculated from these frequencies. We utilize both unweighted frequencies and frequencies weighted using two different strategies. Three conceptually different approaches (entropy-based, variance-based and matrix score-based) are implemented in the algorithm to define the conservation index. Calculating conservation indices for 35 522 positions in 284 alignments from SMART database we demonstrate that different methods result in highly correlated (correlation coefficient more than 0.85) conservation indices. Conservation indices show statistically significant correlation between sequentially adjacent positions i and i + j, where j < 13, and averaging of the indices over the window of three positions is optimal for motif detection. Positions with gaps display substantially lower conservation properties. We compare conservation properties of the SMART alignments or FSSP structural alignments to those of the ClustalW alignments. The results suggest that conservation indices should be a valuable tool of alignment quality assessment and might be used as an objective function for refinement of multiple alignments.
引用
收藏
页码:700 / 712
页数:13
相关论文
共 55 条
  • [1] Gapped BLAST and PSI-BLAST: a new generation of protein database search programs
    Altschul, SF
    Madden, TL
    Schaffer, AA
    Zhang, JH
    Zhang, Z
    Miller, W
    Lipman, DJ
    [J]. NUCLEIC ACIDS RESEARCH, 1997, 25 (17) : 3389 - 3402
  • [2] WEIGHTS FOR DATA RELATED BY A TREE
    ALTSCHUL, SF
    CARROLL, RJ
    LIPMAN, DJ
    [J]. JOURNAL OF MOLECULAR BIOLOGY, 1989, 207 (04) : 647 - 653
  • [3] Positional dependence, cliques, and predictive motifs in the bHLH protein domain
    Atchley, WR
    Terhalle, W
    Dress, A
    [J]. JOURNAL OF MOLECULAR EVOLUTION, 1999, 48 (05) : 501 - 516
  • [4] A STRATEGY FOR THE RAPID MULTIPLE ALIGNMENT OF PROTEIN SEQUENCES - CONFIDENCE LEVELS FROM TERTIARY STRUCTURE COMPARISONS
    BARTON, GJ
    STERNBERG, MJE
    [J]. JOURNAL OF MOLECULAR BIOLOGY, 1987, 198 (02) : 327 - 337
  • [5] A METHOD TO PREDICT FUNCTIONAL RESIDUES IN PROTEINS
    CASARI, G
    SANDER, C
    VALENCIA, A
    [J]. NATURE STRUCTURAL BIOLOGY, 1995, 2 (02): : 171 - 178
  • [6] Dayhoff M.O., 1978, ATLAS PROTEIN SEQ ST, V5
  • [7] Structure-based evaluation of sequence comparison and fold recognition alignment accuracy
    Domingues, FS
    Lackner, P
    Andreeva, A
    Sippl, MJ
    [J]. JOURNAL OF MOLECULAR BIOLOGY, 2000, 297 (04) : 1003 - 1013
  • [8] Structural basis of activation and GTP hydrolysis in Rab proteins
    Dumas, JJ
    Zhu, ZY
    Connolly, JL
    Lambright, DG
    [J]. STRUCTURE, 1999, 7 (04) : 413 - 423
  • [9] Eddy S R, 1995, J Comput Biol, V2, P9, DOI 10.1089/cmb.1995.2.9
  • [10] Eddy S R, 1995, Proc Int Conf Intell Syst Mol Biol, V3, P114