False-positive selection identified by ML-based methods:: Examples from the Sig1 gene of the diatom Thalassiosira weissflogii and the tax gene of a human T-cell lymphotropic virus

被引:97
作者
Suzuki, Y [1 ]
Nei, M
机构
[1] Natl Inst Genet, Ctr Informat Biol, Mishima, Shizuoka, Japan
[2] Natl Inst Genet, DNA Data Bank Japan, Mishima, Shizuoka, Japan
[3] Penn State Univ, Inst Mol Evolutionary Genet, University Pk, PA 16802 USA
[4] Penn State Univ, Dept Biol, University Pk, PA 16802 USA
关键词
positive selection; parsimony; likelihood; Thalassiosira weissflogii; sexually induced gene 1; human T-cell lymphotropic virus type 1; tax;
D O I
10.1093/molbev/msh098
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
Sexually induced gene 1 (Sig1) in the centric diatom Thalassiosira weissflogii is considered to encode a gamete recognition protein. Sorhannus (2003) analyzed nucleotide sequences of Sig1 using parsimony analysis and the maximum-likelihood (ML)-based Bayesian method for inferring positive selection at single amino acid sites and reported that positively selected sites were detected by the latter method but not by the former. He then concluded that for this type of study, the ML-based method is more reliable than parsimony analysis. Here we show that his results apparently represent false-positive cases of the ML-based method and that there is no solid evidence that this gene contains positively selected sites. We further demonstrate that in the tax gene of human T-cell lymphotropic virus type I (HTLV-I), all codon sites, including invariable sites, can be inferred as positively selected sites by the ML-based method. These observations indicate that the ML-based method may produce many false-positive sites. One of the main reasons for the occurrence of false positives is that in the ML-based method, codon sites are grouped into several categories, with different nonsynonymous/synonymous rate ratios (omegas), on a purely statistical basis, and positive selection is inferred indirectly by examining whether the average omega for each category is greater than 1. In parsimony analysis, however, the evolutionary change of nucleotides at each codon site is examined. For this reason, parsimony-based methods rarely produce false positives and are safer than ML-based methods for detecting positive selection at individual codon sites, although a large number of sequences are necessary.
引用
收藏
页码:914 / 921
页数:8
相关论文
共 34 条