SearcharxivSearch

arXiv subjects

Simon Gravel

Publications and source records attributed to Simon Gravel.

15 recordsLinked to original sources

Population-scale Ancestral Recombination Graphs with tskit 1.0

Ancestral recombination graphs (ARGs) are an increasingly important component of population and statistical genetics. The tskit library has become key infrastructure for the field, providing an expressive and general representation of ARGs together with a suite of efficient fundamental operations. In this note, we announce tskit version 1.0, describe its underlying rationale, and document its stability guarantees. These guarantees provide a foundation for durable computational artefacts and support long-term reproducibility of code and analyses.

q-bio.PE

The existence and abundance of ghost ancestors in biparental populations

In a randomly-mating biparental population of size $N$ there are, with high probability, individuals who are genealogical ancestors of every extant individual within approximately $\log_2(N)$ generations into the past. We use this result of J. Chang to prove a curious corollary under standard models of recombination: there exist, with high probability, individuals within a constant multiple of $ \log_2(N)$ generations into the past who are simultaneously (i) genealogical ancestors of {\em each} of the individuals at the present, and (ii) genetic ancestors to {\em none} of the individuals at the present. Such ancestral individuals - ancestors of everyone today that left no genetic trace -- represent `ghost' ancestors in a strong sense. In this short note, we use simple analytical argument and simulations to estimate how many such individuals exist in finite Wright-Fisher populations.

q-bio.PE

Predicting discovery rates of genomic features

Successful sequencing experiments require judicious sample selection. However, this selection must often be performed on the basis of limited preliminary data. Predicting the statistical properties of the final sample based on preliminary data can be challenging, because numerous uncertain model assumptions may be involved. Here, we ask whether we can predict ``omics" variation across many samples by sequencing only a fraction of them. In the infinite-genome limit, we find that a pilot study sequencing $5\%$ of a population is sufficient to predict the number of genetic variants in the entire population within $6\%$ of the correct value, using an estimator agnostic to demography, selection, or population structure. To reach similar accuracy in a finite genome with millions of polymorphisms, the pilot study would require about $15\%$ of the population. We present computationally efficient jackknife and linear programming methods that exhibit substantially less bias than the state of the art when applied to simulated data and sub-sampled 1000 Genomes Project data. Extrapolating based on the NHLBI Exome Sequencing Project data, we predict that $7.2\%$ of sites in the capture region would be variable in a sample of $50,000$ African-Americans, and $8.8\%$ in a European sample of equal size. Finally, we show how the linear programming method can also predict discovery rates of various genomic features, such as the number of transcription factor binding sites across different cell types.

q-bio.GN

Reconstructing Native American Migrations from Whole-genome and Whole-exome Data

There is great scientific and popular interest in understanding the genetic history of populations in the Americas. We wish to understand when different regions of the continent were inhabited, where settlers came from, and how current inhabitants relate genetically to earlier populations. Recent studies unraveled parts of the genetic history of the continent using genotyping arrays and uniparental markers. The 1000 Genomes Project provides a unique opportunity for improving our understanding of population genetic history by providing over a hundred sequenced low coverage genomes and exomes from Colombian (CLM), Mexican-American (MXL), and Puerto Rican (PUR) populations. Here, we explore the genomic contributions of African, European, and Native American ancestry to these populations. Estimated Native American ancestry is 48% in MXL, 25% in CLM, and 13% in PUR. Native American ancestry in PUR is most closely related to populations surrounding the Orinoco River basin, confirming the Southern America ancestry of the Taíno people of the Caribbean. We present new methods to estimate the allele frequencies in the Native American fraction of the populations, and model their distribution using a demographic model for three ancestral Native American populations. These ancestral populations likely split in close succession: the most likely scenario, based on a peopling of the Americas 16 thousand years ago (kya), supports that the MXL Ancestors split 12.2kya, with a subsequent split of the ancestors to CLM and PUR 11.7kya. The model also features effective populations of 62,000 in Mexico, 8,700 in Colombia, and 1,900 in Puerto Rico. Modeling Identity-by-descent and ancestry tract length, we show that post-contact populations differ markedly in their effective sizes and migration patterns, with Puerto Rico showing the smallest effective size and the earlier migration from Europe.

q-bio.PE

Reconstructing the Population Genetic History of the Caribbean

The Caribbean basin is home to some of the most complex interactions in recent history among previously diverged human populations. Here, by making use of genome-wide SNP array data, we characterize ancestral components of Caribbean populations on a sub-continental level and unveil fine-scale patterns of population structure distinguishing insular from mainland Caribbean populations as well as from other Hispanic/Latino groups. We provide genetic evidence for an inland South American origin of the Native American component in island populations and for extensive pre-Columbian gene flow across the Caribbean basin. The Caribbean-derived European component shows significant differentiation from parental Iberian populations, presumably as a result of founder effects during the colonization of the New World. Based on demographic models, we reconstruct the complex population history of the Caribbean since the onset of continental admixture. We find that insular populations are best modeled as mixtures absorbing two pulses of African migrants, coinciding with early and maximum activity stages of the transatlantic slave trade. These two pulses appear to have originated in different regions within West Africa, imprinting two distinguishable signatures in present day Afro-Caribbean genomes and shedding light on the genetic impact of the dynamics occurring during the slave trade in the Caribbean.

q-bio.PE

Population genetics models of local ancestry

Migrations have played an important role in shaping the genetic diversity of human populations. Understanding genomic data thus requires careful modeling of historical gene flow. Here we consider the effect of relatively recent population structure and gene flow, and interpret genomes of individuals that have ancestry from multiple source populations as mosaics of segments originating from each population. We propose general and tractable models for describing the evolution of these patterns of local ancestry and their impact on genetic diversity. We focus on the length distribution of continuous ancestry tracts, and the variance in total ancestry proportions among individuals. The proposed models offer improved agreement with Wright-Fisher simulation data when compared to state-of-the art models, and can be used to infer various demographic parameters in gene flow models. Considering HapMap African-American (ASW) data, we find that a model with two distinct phases of `European' gene flow significantly improves the modeling of both tract lengths and ancestry variances.

q-bio.PE

Upper bound on the packing density of regular tetrahedra and octahedra

We obtain an upper bound to the packing density of regular tetrahedra. The bound is obtained by showing the existence, in any packing of regular tetrahedra, of a set of disjoint spheres centered on tetrahedron edges, so that each sphere is not fully covered by the packing. The bound on the amount of space that is not covered in each sphere is obtained in a recursive way by building on the observation that non-overlapping regular tetrahedra cannot subtend a solid angle of $4π$ around a point if this point lies on a tetrahedron edge. The proof can be readily modified to apply to other polyhedra with the same property. The resulting lower bound on the fraction of empty space in a packing of regular tetrahedra is $2.6\ldots\times 10^{-25}$ and reaches $1.4\ldots\times 10^{-12}$ for regular octahedra.

math.MG

A method for dense packing discovery

The problem of packing a system of particles as densely as possible is foundational in the field of discrete geometry and is a powerful model in the material and biological sciences. As packing problems retreat from the reach of solution by analytic constructions, the importance of an efficient numerical method for conducting \textit{de novo} (from-scratch) searches for dense packings becomes crucial. In this paper, we use the \textit{divide and concur} framework to develop a general search method for the solution of periodic constraint problems, and we apply it to the discovery of dense periodic packings. An important feature of the method is the integration of the unit cell parameters with the other packing variables in the definition of the configuration space. The method we present led to improvements in the densest-known tetrahedron packing which are reported in [arXiv:0910.5226]. Here, we use the method to reproduce the densest known lattice sphere packings and the best known lattice kissing arrangements in up to 14 and 11 dimensions respectively (the first such numerical evidence for their optimality in some of these dimensions). For non-spherical particles, we report a new dense packing of regular four-dimensional simplices with density $ϕ=128/219\approx0.5845$ and with a similar structure to the densest known tetrahedron packing.

math.MG

Dense periodic packings of tetrahedra with small repeating units

We present a one-parameter family of periodic packings of regular tetrahedra, with the packing fraction $100/117\approx0.8547$, that are simple in the sense that they are transitive and their repeating units involve only four tetrahedra. The construction of the packings was inspired from results of a numerical search that yielded a similar packing. We present an analytic construction of the packings and a description of their properties. We also present a transitive packing with a repeating unit of two tetrahedra and a packing fraction $\frac{139+40\sqrt{10}}{369}\approx0.7194$.

math.MG

Laminating lattices with symmetrical glue

We use the automorphism group $Aut(H)$, of holes in the lattice $L_8=A_2\oplus A_2\oplus D_4$, as the starting point in the construction of sphere packings in 10 and 12 dimensions. A second lattice, $L_4=A_2\oplus A_2$, enters the construction because a subgroup of $Aut(L_4)$ is isomorphic to $Aut(H)$. The lattices $L_8$ and $L_4$, when glued together through this relationship, provide an alternative construction of the laminated lattice in twelve dimensions with kissing number 648. More interestingly, the action of $Aut(H)$ on $L_4$ defines a pair of invariant planes through which dense, non-lattice packings in 10 dimensions can be constructed. The most symmetric of these is aperiodic with center density 1/32. These constructions were prompted by an unexpected arrangement of 378 kissing spheres discovered by a search algorithm.

math.MG

Divide and concur: A general approach to constraint satisfaction

Many difficult computational problems involve the simultaneous satisfaction of multiple constraints which are individually easy to satisfy. Such problems occur in diffractive imaging, protein folding, constrained optimization (e.g., spin glasses), and satisfiability testing. We present a simple geometric framework to express and solve such problems and apply it to two benchmarks. In the first application (3SAT, a boolean satisfaction problem), the resulting method exhibits similar performance scaling as a leading context-specific algorithm (walksat). In the second application (sphere packing), the method allowed us to find improved solutions to some old and well-studied optimization problems. Based upon its simplicity and observed efficiency, we argue that this framework provides a competitive alternative to stochastic methods such as simulated annealing.

physics.comp-ph

Nonlinear response theories and effective pair potentials

We present a general method based on nonlinear response theory to obtain effective interactions between ions in an electron gas which can also be applied to other systems where an adiabatic separation of time-scales is possible. Nonlinear contributions to the interatomic potential are expressed in terms of physically meaningful quantities, giving insight in the physical properties of the system. The method is applied to various test cases and is found to improve the standard linear and quadratic response approaches. It also reduces the discrepancies previously observed between perturbation theory and density-functional theory results for the proton-proton pair potentials in metallic environments.

cond-mat.mtrl-sci

Hamiltonians separable in cartesian coordinates and third-order integrals of motion

We present in this article all Hamiltonian systems in E(2) that are separable in cartesian coordinates and that admit a third-order integral, both in quantum and in classical mechanics. Many of these superintegrable systems are new, and it is seen that there exists a relation between quantum superintegrable potentials, invariant solutions of the Korteweg-De Vries equation and the Painlevé transcendents.

math-ph

Superintegrability, isochronicity, and quantum harmonic behavior

We discuss the properties of superintegrable Hamiltonian systems, in particular those that admit separation of variables in cartesian coordinates. We show that the superintegrability of such potentials is equivalent to the isochronicity of the separated potentials. We use this fact to get a new insight into an old question about the relation between quantum and classical harmonic behavior.

math-ph

Superintegrability with third order invariants in quantum and classical mechanics

We consider here the coexistence of first- and third-order integrals of motion in two dimensional classical and quantum mechanics. We find explicitly all potentials that admit such integrals, and all their integrals. Quantum superintegrable systems are found that have no classical analog, i.e. the potentials are proportional to \hbar^2, so their classical limit is free motion.

math-ph