Impacts of Terraces on Phylogenetic Inference

被引:34
作者
Sanderson, Michael J. [1 ]
McMahon, Michelle M. [2 ]
Stamatakis, Alexandros [1 ,3 ,4 ]
Zwickl, Derrick J. [1 ]
Steel, Mike [5 ]
机构
[1] Univ Arizona, Dept Ecol & Evolutionary Biol, Tucson, AZ 85721 USA
[2] Univ Arizona, Sch Plant Sci, Tucson, AZ 85721 USA
[3] Heidelberg Inst Theoret Studies, Sci Comp Grp, D-69118 Heidelberg, Germany
[4] Karlsruhe Inst Technol, Inst Theoret Informat, D-76131 Karlsruhe, Germany
[5] Univ Canterbury, Biomath Res Ctr, Christchurch 1, New Zealand
基金
美国国家科学基金会;
关键词
Bootstrap; partitioned model; phylogenetics; posterior probability; terrace; MISSING DATA; BOOTSTRAP SUPPORT; TREE; LIKELIHOOD; EVOLUTION; COMPLEXITY; CHOICE; RECONSTRUCTION; SUPERMATRICES; PHYLOGENOMICS;
D O I
10.1093/sysbio/syv024
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
Terraces are sets of trees with precisely the same likelihood or parsimony score, which can be induced by missing sequences in partitioned multi-locus phylogenetic data matrices. The potentially large set of trees on a terrace can be characterized by enumeration algorithms or consensus methods that exploit the pattern of partial taxon coverage in the data, independent of the sequence data themselves. Terraces can add ambiguity and complexity to phylogenetic inference, particularly in settings where inference is already challenging: data sets with many taxa and relatively few loci. In this article we present five new findings about terraces and their impacts on phylogenetic inference. First, we clarify assumptions about partitioning scheme model parameters that are necessary for the existence of terraces. Second, we explore the dependence of terrace size on partitioning scheme and indicate how to find the partitioning scheme associated with the largest terrace containing a given tree. Third, we highlight the impact of terrace size on bootstrap estimates of confidence limits in clades, and characterize the surprising result that the bootstrap proportion for a clade, as it is usually calculated, can be entirely determined by the frequency of bipartitions on a terrace, with some bipartitions receiving high support even when incorrect. Fourth, we dissect some effects of prior distributions of edge lengths on the computed posterior probabilities of clades on terraces, to understand an example in which long edges "attract" each other in Bayesian inference. Fifth, we describe how assuming relationships between edge-lengths of different loci, as an attempt to avoid terraces, can also be problematic when taxon coverage is partial, specifically when heterotachy is present. Finally, we discuss strategies for remediation of some of these problems. One promising approach finds a minimal set of taxa which, when deleted from the data matrix, reduces the size of a terrace to a single tree.
引用
收藏
页码:709 / 726
页数:18
相关论文
共 85 条
[21]  
FELSENSTEIN J, 1985, EVOLUTION, V39, P783, DOI 10.1111/j.1558-5646.1985.tb00420.x
[22]   TEMPO AND MODE IN PLANT BREEDING SYSTEM EVOLUTION [J].
Goldberg, Emma E. ;
Igic, Boris .
EVOLUTION, 2012, 66 (12) :3701-3709
[23]   HOMOPLASY AND THE CHOICE AMONG CLADOGRAMS [J].
GOLOBOFF, PA .
CLADISTICS-THE INTERNATIONAL JOURNAL OF THE WILLI HENNIG SOCIETY, 1991, 7 (03) :215-232
[24]   Hide and vanish: Data sets where the most parsimonious tree is known but hard to find, and their implications for tree search methods [J].
Goloboff, Pablo A. .
MOLECULAR PHYLOGENETICS AND EVOLUTION, 2014, 79 :118-131
[26]   Phylogenomic Resolution of Paleozoic Divergences in Harvestmen (Arachnida, Opiliones) via Analysis of Next-Generation Transcriptome Data [J].
Hedin, Marshal ;
Starrett, James ;
Akhter, Sajia ;
Schoenhofer, Axel L. ;
Shultz, Jeffrey W. .
PLOS ONE, 2012, 7 (08)
[27]   Assessing the root of bilaterian animals with scalable phylogenomic methods [J].
Hejnol, Andreas ;
Obst, Matthias ;
Stamatakis, Alexandros ;
Ott, Michael ;
Rouse, Greg W. ;
Edgecombe, Gregory D. ;
Martinez, Pedro ;
Baguna, Jaume ;
Bailly, Xavier ;
Jondelius, Ulf ;
Wiens, Matthias ;
Mueller, Werner E. G. ;
Seaver, Elaine ;
Wheeler, Ward C. ;
Martindale, Mark Q. ;
Giribet, Gonzalo ;
Dunn, Casey W. .
PROCEEDINGS OF THE ROYAL SOCIETY B-BIOLOGICAL SCIENCES, 2009, 276 (1677) :4261-4270
[28]   Addressing Inter-Gene Heterogeneity in Maximum Likelihood Phylogenomic Analysis: Yeasts Revisited [J].
Hess, Jaqueline ;
Goldman, Nick .
PLOS ONE, 2011, 6 (08)
[29]   Using Supermatrices for Phylogenetic Inquiry: An Example Using the Sedges [J].
Hinchliff, Cody E. ;
Roalson, Eric H. .
SYSTEMATIC BIOLOGY, 2013, 62 (02) :205-219
[30]   Algorithms, data structures, and numerics for likelihood-based phylogenetic inference of huge trees [J].
Izquierdo-Carrasco, Fernando ;
Smith, Stephen A. ;
Stamatakis, Alexandros .
BMC BIOINFORMATICS, 2011, 12