Searcharxiv⌕ Search

arXiv subjects

Peter J. Waddell

Publications and source records attributed to Peter J. Waddell.

14 recordsLinked to original sources

BME-like Quartet Weights for Phylogenetic Trees

Like pairwise distances, quartets can be highly redundant and correlated on a phylogenetic tree, and their number grows on the order of n^4 rather than n^2. I explore BME-like weights for reweighting quartet scores before summing them to score a full tree. Three weights are considered on an unrooted binary tree: w_ext(q)=2^(-I_ext(q)), w_int(q)=2^(-I_int(q)), and w_tot(q)=2^(-I_tot(q))=w_ext(q)w_int(q), where the exponents count specified internal nodes in the minimal connecting subtree of a quartet. Exact tree-shape counts, total quartet-weight sums, and internal-edge crossing sums are calculated for all unlabeled unrooted binary tree shapes on 6-10 taxa. For w_ext, the total quartet weight is tree-shape-invariant and the edge-crossing sum depends only on split size. For any n-leaf tree, we prove sum_q w_ext(q)=(n-2)(n-3)/8, and the sum over quartets crossing an internal edge with split a|b equals (a-1)(b-1)/4. Exact tree-shape-specific normalizers are also derived for w_int and w_tot. A degree-corrected hard-polytomy extension is given for multifurcating trees, and a conditional consistency result shows that these positive weights preserve consistency when the underlying quartet estimates are themselves consistent for the true induced quartet states. These results provide a mathematical foundation for evaluating and applying BME-like quartet weights to reduce redundancy with the particular aim of improving statistical efficiency with finite data.

q-bio.PE↗

Complimentary Phylogenetic Signals for Morphological Characters and Quantitative 3D Shape Data within genus Homo

Estimating the phylogeny of the genus Homo is entering a new phase of vastly improved data and methodology. There is increasing evidence of 6 to 10 competing species/lineages at any point in the last half million years, making the elucidation of the relationships of individual specimens particularly important. Recent estimates of the phylogeny of key specimens include Waddell (2013, 2014, 2015, 2016), and Mounier et al. (2016). These are made with quite different data (3D skull shapes and discrete morphological characters, respectively) and methods of analysis (unweighted least squares fitting of distances, OLS+, and reweighted maximum parsimony, respectively). Initial inspection of the trees in these articles might leave the impression of a great deal of disagreement and confused results. Here it is shown this need not be the case, and that these two types of data and analysis may be indicating a very similar tree, one that is in good agreement also with subjective current wisdom/expert opinions on particular parts of the phylogeny. The precise location of the African LH18 specimen arises as key to a better understanding of the likely form of the last common ancestors of H. sapiens and Neanderthals. A diverse approach seems to bring forth much more agreement of trees than otherwise perceived, and argues against being dogmatic about methods of phylogenetic analysis particularly when working with difficult problems.

q-bio.PE↗

Expanded Distance-based Phylogenetic Analyses of Fossil Homo Skull Shape Evolution

Analyses of a set of 47 fossil and 4 modern skulls using phylogenetic geometric morphometric methods corroborate and refine earlier results. These include evidence that the African Iwo Eleru skull, only about 12,000 years old, indeed represents a new species of near human. In contrast, the earliest known anatomically modern human skull, Qafzeh 9, the skull of Eve from Israel/Palestine, is validated as fully modern in form. Analyses clearly show evidence of archaic introgression into Gravettian, pre_Gravettian, Qafzeh, and Upper Cave (China) populations of near modern humans, and in about that order of increasing archaic content. The enigmatic Saldahna (Elandsfontein) skull emerges as a probable first representative of that lineage, which exclusive of Neanderthals that, eventually lead to modern humans. There is also evidence that the poorly dated Kabwe (Broken Hill) skull represents a much earlier distinct lineage. The clarity of the results bode well for quantitative statistical phylogenetic methods making significant inroads in the stalemates of paleoanthropology.

q-bio.PE↗

Extended Distance-based Phylogenetic Analyses Applied to 3D Homo Fossil Skull Evolution

This article shows how 3D geometric morphometric data can be analyzed using newly developed distance-based evolutionary tree inference methods, with extensions to planar graphs. Application of these methods to 3D representations of the skullcap (calvaria) of 13 diverse skulls in the genus Homo, ranging from Homo erectus (ergaster) at about 1.6 mya, all the way forward to modern humans, yields a remarkably clear phylogenetic tree. Various evolutionary hypotheses are tested. Results of these tests include rejection of the monophyly of Homo heidelbergensis, the Multi-Regional hypothesis, and the hypothesis that the unusual 12,000 year old (12kya) Iwo Eleru skull represents a modern human. Rather, by quantitative phylogenetic analyses the latter is seen to be an old (200-400kya) lineage that probably represents a novel African species, Homo iwoelerueensis. It diverged after the lineage leading to Neanderthals, and may have been driven to extinction in the last 10kya by modern humans, Homo sapiens, another African species of Homo that appeared about 100kya. Another enigmatic skull, Qafzeh 6 from the Middle East about 90kya, appears to be a hybrid of two thirds near, but not, anatomically modern human and one third of an archaic lineage diverging close to classic European Neanderthals. Overall, the tree clearly implies an accelerating rate of skullcap shape change, and by extension, change of the underlying brain, over the last 400kya in Africa. This acceleration may have extended right up to the origin of modern humans. Methods of distance-based evolutionary tree inference are refined and extended, with particular attention to diagnosing the model and achieving a better fit. This includes power transformations of the input data which favor root Procrustes distances.

q-bio.PE↗

Happy New Year Homo erectus? More evidence for interbreeding with archaics predating the modern human/Neanderthal split

A range of a priori hypotheses about the evolution of modern and archaic genomes are further evaluated and tested. In addition to the well-known splits/introgressions involving Neanderthal genes into out-of- Africa people, or Denisovan genes into Oceanians, a further series of archaic splits and hypotheses proposed in Waddell et al. (2011) are considered in detail. These include signals of Denisovans with something markedly more archaic and possibly something more archaic into Papuans as well. These are compared and contrasted with some well-advertised introgressions such as Denisovan genes across East Asia, archaic genes into San or non-tree mixing between Oceanians, East Asians and Europeans. The general result is that these less appreciated and surprising archaic splits have just as much or more support in genome sequence data. Further, evaluation confirms the hypothesis that archaic genes are much rarer on modern X chromosomes, and may even be near totally absent, suggesting strong selection against their introgression. Modeling of relative split weights allows an inference of the proportion of the genome the Denisovan seems to have gotten from an older archaic, and the best estimate is around 2%. Using a mix of quantitative and qualitative morphological data and novel phylogenetic methods, robust support is found for multiple distinct middle Pleistocene lineages. Of these, fossil hominids such as SH5, Petralona, and Dali, in particular, look like prime candidates for contributing pre-Neanderthal/Modern archaic genes to Denisovans, while the Jinniu-Shan fossil looks like the best candidate for a close relative of the Denisovan. That the Papuans might have received some truly archaic genes appears a good possibility and they might even be from Homo erectus.

q-bio.PE↗

New g%AIC, g%AICc, g%BIC, and Power Divergence Fit Statistics Expose Mating between Modern Humans, Neanderthals and other Archaics

The purpose of this article is to look at how information criteria, such as AIC and BIC, relate to the g%SD fit criterion derived in Waddell et al. (2007, 2010a). The g%SD criterion measures the fit of data to model based on a normalized weighted root mean square percentage deviation between the observed data and model estimates of the data, with g%SD = 0 being a perfectly fitting model. However, this criterion may not be adjusting for the number of parameters in the model comprehensively. Thus, its relationship to more traditional measures for maximizing useful information in a model, including AIC and BIC, are examined. This results in an extended set of fit criteria including g%AIC and g%BIC. Further, a broader range of asymptotically most powerful fit criteria of the power divergence family, which includes maximum likelihood (or minimum G^2) and minimum X^2 modeling as special cases, are used to replace the sum of squares fit criterion within the g%SD criterion. Results are illustrated with a set of genetic distances looking particularly at a range of Jewish populations, plus a genomic data set that looks at how Neanderthals and Denisovans are related to each other and modern humans. Evidence that Homo erectus may have left a significant fraction of its genome within the Denisovan is shown to persist with the new modeling criteria.

q-bio.GN↗

Homo denisova, Correspondence Spectral Analysis, Finite Sites Reticulate Hierarchical Coalescent Models and the Ron Jeremy Hypothesis

This article shows how to fit reticulate finite and infinite sites sequence spectra to aligned data from five modern human genomes (San, Yoruba, French, Han and Papuan) plus two archaic humans (Denisovan and Neanderthal), to better infer demographic parameters. These include interbreeding between distinct lineages. Major improvements in the fit of the sequence spectrum are made with successively more complicated models. Findings include some evidence of a male biased gene flow from the Denisova lineage to Papuan ancestors and possibly even more archaic gene flow. It is unclear if there is evidence for more than one Neanderthal interbreeding, as the evidence suggesting this largely disappears when a finite sites model is fitted.

q-bio.PE↗

What use are Exponential Weights for flexi-Weighted Least Squares Phylogenetic Trees?

The method of flexi-Weighted Least Squares on evolutionary trees uses simple polynomial or exponential functions of the evolutionary distance in place of model-based variances. This has the advantage that unexpected deviations from additivity can be modeled in a more flexible way. At present, only polynomial weights have been used. However, a general family of exponential weights is desirable to compare with polynomial weights and to potentially exploit recent insights into fast least squares edge length estimation on trees. Here describe families of weights that are multiplicative on trees, along with measures of fit of data to tree. It is shown that polynomial, but also multiplicative weights can approximate model-based variance of evolutionary distances well. Both models are fitted to evolutionary data from yeast genomes and while the polynomial weights model fits better, the exponential weights model can fit a lot better than ordinary least squares. Iterated least squares is evaluated and is seen to converge quickly and with minimal change in the fit statistics when the data are in the range expected for the useful evolutionary distances and simple Markov models of character change. In summary, both polynomial and exponential weighted least squares work well and justify further investment into developing the fastest possible algorithms for evaluating evolutionary trees.

q-bio.PE↗

A Unified Framework for Trees, Multi-Dimensional Scaling and Planar Graphs

Least squares trees, multi-dimensional scaling and Neighbor Nets are all different and popular ways of visualizing multi-dimensional data. The method of flexi-Weighted Least Squares (fWLS) is a powerful method of fitting phylogenetic trees, when the exact form of errors is unknown. Here, both polynomial and exponential weights are used to model errors. The exact same models are implemented for multi-dimensional scaling to yield flexi-Weighted MDS, including as special cases methods such as the Sammon Stress function. Here we apply all these methods to population genetic data looking at the relationships of "Abrahams Children" encompassing Arabs and now widely dispersed populations of Jews, in relation to an African outgroup and a variety of European populations. Trees, MDS and Neighbor Nets of this data are compared within a common likelihood framework and the strengths and weaknesses of each method are explored. Because the errors in this type of data can be complex, for example, due to unexpected genetic transfer, we use a residual resampling method to assess the robustness of trees and the Neighbor Net. Despite the Neighbor Net fitting best by all criteria except BIC, its structure is ill defined following residual resampling. In contrast, fWLS trees are favored by BIC and retain considerable strong internal structure following residual resampling. This structure clearly separates various European and Middle Eastern populations, yet it is clear all of the models have errors much larger than expected by sampling variance alone.

q-bio.PE↗

Resampling Residuals on Phylogenetic Trees: Extended Results

In this article the results of Waddell and Azad (2009) are extended. In particular, the geometric percentage mean standard deviation measure of the fit of distances to a phylogenetic tree is adjusted for the number of parameters fitted to the model. The formulae are also presented in their general form for any weight that is a function of the distance. The cell line gene expression data set of Ross et al. (2000) is reanalyzed. It is shown that ordinary least squares (OLS) is a much better fit to the data than a Neighbor Joining or BME tree. Residual resampling shows that cancer cell lines do indeed fit a tree fairly well and that the tree does have strong internal structure. Simulations show that least squares tree building methods, including OLS, are strong competitors with BME type methods for fitting model data, while real world examples often suggest the same conclusion.

q-bio.PE↗

Resampling Residuals: Robust Estimators of Error and Fit for Evolutionary Trees and Phylogenomics

Phylogenomics, even more so than traditional phylogenetics, needs to represent the uncertainty in evolutionary trees due to systematic error. Here we illustrate the analysis of genome-scale alignments of yeast, using robust measures of the additivity of the fit of distances to tree when using flexi Weighted Least Squares. A variety of DNA and protein distances are used. We explore the nature of the residuals, standardize them, and then create replicate data sets by resampling these residuals. Under the model, the results are shown to be very similar to the conventional sequence bootstrap. With real data they show up uncertainty in the tree that is either due to underestimating the stochastic error (hence massively overestimating the effective sequence length) and/or systematic error. The methods are extended to the very fast BME criterion with similarly promising results.

q-bio.PE↗

Expectation-Maximization (EM) Algorithms for Mapping Short Reads Illustrated with FAIRE data and the TP53-WRAP53 Gene Region

Huge numbers of short reads are being generated for mapping back to the genome to discover the frequency of transcripts, miRNAs, DNAase hypersensitive sites, FAIRE regions, nucleosome occupancy, etc. Since these reads are typically short (e.g., 36 base pairs) and since many eukaryotic genomes, including humans, have highly repetitive sequences then many of these reads map to two or more locations in the genome. Current mapping of these reads, grading them according to 0, 1 or 2 mismatches wastes a great deal of information. These short sequences are typically mapped with no account of the accuracy of the sequence, even in company software when per base error rates are being reported by another part of the machine. Further, multiply mapping locations are frequently discarded altogether or allocated with no regard to where other reads are accumulating. Here we show how to combine probabilistic mapping of reads with an EM algorithm to iteratively improve the empirical likelihood of the allocation of short reads. Mapping using LAST takes into account the per base accuracy of the read, plus insertions and deletions, plus anticipated occasional errors or SNPs with respect to the parent genome. The probabilistic EM algorithm iteratively allocates reads based on the proportion of reads mapping within windows on the previous cycle, along with any prior information on where the read best maps. The methods are illustrated with FAIRE ENCODE data looking at the very important head-to-head gene combination of TP53 and WRAP 53.

q-bio.GN↗

Measuring Fit of Sequence Data to Phylogenetic Model: Gain of Power using Marginal Tests

Testing fit of data to model is fundamentally important to any science, but publications in the field of phylogenetics rarely do this. Such analyses discard fundamental aspects of science as prescribed by Karl Popper. Indeed, not without cause, Popper (1978) once argued that evolutionary biology was unscientific as its hypotheses were untestable. Here we trace developments in assessing fit from Penny et al. (1982) to the present. We compare the general log-likelihood ratio (the G or G2 statistic) statistic between the evolutionary tree model and the multinomial model with that of marginalized tests applied to an alignment (using placental mammal coding sequence data). It is seen that the most general test does not reject the fit of data to model (p~0.5), but the marginalized tests do. Tests on pair-wise frequency (F) matrices, strongly (p < 0.001) reject the most general phylogenetic (GTR) models commonly in use. It is also clear (p < 0.01) that the sequences are not stationary in their nucleotide composition. Deviations from stationarity and homogeneity seem to be unevenly distributed amongst taxa; not necessarily those expected from examining other regions of the genome. By marginalizing the 4t patterns of the i.i.d. model to observed and expected parsimony counts, that is, from constant sites, to singletons, to parsimony informative characters of a minimum possible length, then the likelihood ratio test regains power, and it too rejects the evolutionary model with p << 0.001. Given such behavior over relatively recent evolutionary time, readers in general should maintain a healthy skepticism of results, as the scale of the systematic errors in published analyses may really be far larger than the analytical methods (e.g., bootstrap) report.

q-bio.PE↗

Combined Sum of Squares Penalties for Molecular Divergence Time Estimation

Estimates of molecular divergence times when rates of evolution vary require the assumption of a model of rate change. Brownian motion is one such model, and since rates cannot become negative, a log Brownian model seems appropriate. Divergence time estimates can then be made using weighted least squares penalties. As sequences become long, this approach effectively becomes equivalent to penalized likelihood or Bayesian approaches. Different forms of the least squares penalty are considered to take into account correlation due to shared ancestors. It is shown that a scale parameter is also needed since the sum of squares changes with the scale of time. Errors or uncertainty on fossil calibrations, may be folded in with errors due to the stochastic nature of Brownian motion and ancestral polymorphism, giving a total sum of squares to be minimized. Applying these methods to placental mammal data the estimated age of the root decreases from 125 to about 94 mybp. However, multiple fossil calibration points and relative molecular divergence times inflate the sum of squares more than expected. If fossil data are also bootstrapped, then the confidence interval for the root of placental mammals varies widely from ~70 to 130 mybp. Such a wide interval suggests that more and better fossil calibration data is needed and/or better models of rate evolution are needed and/or better molecular data are needed. Until these issues are thoroughly investigated, it is premature to declare either the old molecular dates frequently obtained (e.g. > 110 mybp) or the lack of identified placental fossils in the Cretaceous, more indicative of when crown-group placental mammals evolved.

q-bio.PE↗