AL2CO: calculation of positional conservation in a protein sequence alignment

被引:350
作者
Pei, JM
Grishin, NV
机构
[1] Univ Texas, SW Med Ctr, Howard Hughes Med Inst, Dallas, TX 75390 USA
[2] Univ Texas, SW Med Ctr, Dept Biochem, Dallas, TX 75390 USA
关键词
D O I
10.1093/bioinformatics/17.8.700
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Motivation: Amino acid sequence alignments are widely used in the analysis of protein structure, function and evolutionary relationships. Proteins within a superfamily usually share the same fold and possess related functions. These structural and functional constraints are reflected in the alignment conservation patterns. Positions of functional and/or structural importance tend to be more conserved. Conserved positions are usually clustered in distinct motifs surrounded by sequence segments of low conservation. Poorly conserved regions might also arise from the imperfections in multiple alignment algorithms and thus indicate possible alignment errors. Quantification of conservation by attributing a conservation index to each aligned position makes motif detection more convenient. Mapping these conservation indices onto a protein spatial structure helps to visualize spatial conservation features of the molecule and to predict functionally and/or structurally important sites. Analysis of conservation indices could be a useful tool in detection of potentially misaligned regions and will aid in improvement of multiple alignments. Results: We developed a program to calculate a conservation index at each position in a multiple sequence alignment using several methods. Namely, amino acid frequencies at each position are estimated and the conservation index is calculated from these frequencies. We utilize both unweighted frequencies and frequencies weighted using two different strategies. Three conceptually different approaches (entropy-based, variance-based and matrix score-based) are implemented in the algorithm to define the conservation index. Calculating conservation indices for 35 522 positions in 284 alignments from SMART database we demonstrate that different methods result in highly correlated (correlation coefficient more than 0.85) conservation indices. Conservation indices show statistically significant correlation between sequentially adjacent positions i and i + j, where j < 13, and averaging of the indices over the window of three positions is optimal for motif detection. Positions with gaps display substantially lower conservation properties. We compare conservation properties of the SMART alignments or FSSP structural alignments to those of the ClustalW alignments. The results suggest that conservation indices should be a valuable tool of alignment quality assessment and might be used as an objective function for refinement of multiple alignments.
引用
收藏
页码:700 / 712
页数:13
相关论文
共 55 条
  • [41] Structure-derived substitution matrices for alignment of distantly related sequences
    Prlic, A
    Domingues, FS
    Sippl, MJ
    [J]. PROTEIN ENGINEERING, 2000, 13 (08): : 545 - 550
  • [42] DATABASE OF HOMOLOGY-DERIVED PROTEIN STRUCTURES AND THE STRUCTURAL MEANING OF SEQUENCE ALIGNMENT
    SANDER, C
    SCHNEIDER, R
    [J]. PROTEINS-STRUCTURE FUNCTION AND GENETICS, 1991, 9 (01): : 56 - 68
  • [43] SMART, a simple modular architecture research tool: Identification of signaling domains
    Schultz, J
    Milpetz, F
    Bork, P
    Ponting, CP
    [J]. PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 1998, 95 (11) : 5857 - 5864
  • [44] INFORMATION-THEORETICAL ENTROPY AS A MEASURE OF SEQUENCE VARIABILITY
    SHENKIN, PS
    ERMAN, B
    MASTRANDREA, LD
    [J]. PROTEINS-STRUCTURE FUNCTION AND GENETICS, 1991, 11 (04): : 297 - 313
  • [45] PSIC: profile extraction from sequence alignments with position-specific counts of independent observations
    Sunyaev, SR
    Eisenhaber, F
    Rodchenkov, IV
    Eisenhaber, B
    Tumanyan, VG
    Kuznetsov, EN
    [J]. PROTEIN ENGINEERING, 1999, 12 (05): : 387 - 394
  • [46] A FLEXIBLE METHOD TO ALIGN LARGE NUMBERS OF BIOLOGICAL SEQUENCES
    TAYLOR, WR
    [J]. JOURNAL OF MOLECULAR EVOLUTION, 1988, 28 (1-2) : 161 - 169
  • [47] THOMPSON JD, 1994, COMPUT APPL BIOSCI, V10, P19
  • [48] CLUSTAL-W - IMPROVING THE SENSITIVITY OF PROGRESSIVE MULTIPLE SEQUENCE ALIGNMENT THROUGH SEQUENCE WEIGHTING, POSITION-SPECIFIC GAP PENALTIES AND WEIGHT MATRIX CHOICE
    THOMPSON, JD
    HIGGINS, DG
    GIBSON, TJ
    [J]. NUCLEIC ACIDS RESEARCH, 1994, 22 (22) : 4673 - 4680
  • [49] A comprehensive comparison of multiple sequence alignment programs
    Thompson, JD
    Plewniak, F
    Poch, O
    [J]. NUCLEIC ACIDS RESEARCH, 1999, 27 (13) : 2682 - 2690
  • [50] AMINO-ACID PREFERENCES AT PROTEIN-BINDING SITES
    VILLAR, HO
    KAUVAR, LM
    [J]. FEBS LETTERS, 1994, 349 (01) : 125 - 130