SearcharxivSearch

arXiv subjects

Phillip Endicott

Publications and source records attributed to Phillip Endicott.

2 recordsLinked to original sources

The Human Genomic Landscape of Oceania

Oceania and Island Southeast Asia have a rich, yet understudied, human genomic landscape. This region encompasses some of the first areas inhabited by humans following the out-of-Africa expansion, includes populations with the highest levels of archaic hominin introgression, and contains Pacific islands that are among the most remote continuously inhabited locations in the world. Here, we describe the first region-wide analysis of individuals from population groups spanning Oceania and its broad perimeter. In total we generate and analyze genome-wide data from 92 different populations, 58 separate islands, and 30 countries, covering one third of the planet. Leveraging this diverse dataset, we resolve genetic connections among islands, providing a detailed view of regional population structure and identifying the island groups involved in the settlement of several Polynesian Outliers. Ancestry-specific analyses allow us to deconvolve different layers of history, from tracing groups deriving their Austronesian ancestry via the Lapita expansion to quantifying variable archaic introgression across the basal Papuan component of Oceanians and Southeast Asians. Finally, we map biomedically relevant variants across Oceania and Southeast Asia, observing pronounced allele-frequency differences between population groups. Together, these findings refine models of oceanic settlement and admixture and establish a comprehensive reference that will advance global efforts to ensure broad and equitable representation in human genomics.

q-bio.PE

CLARITY -- Comparing heterogeneous data using dissimiLARITY

Integrating datasets from different disciplines is hard because the data are often qualitatively different in meaning, scale, and reliability. When two datasets describe the same entities, many scientific questions can be phrased around whether the (dis)similarities between entities are conserved across such different data. Our method, CLARITY, quantifies consistency across datasets, identifies where inconsistencies arise, and aids in their interpretation. We illustrate this using three diverse comparisons: gene methylation vs expression, evolution of language sounds vs word use, and country-level economic metrics vs cultural beliefs. The non-parametric approach is robust to noise and differences in scaling, and makes only weak assumptions about how the data were generated. It operates by decomposing similarities into two components: a `structural' component analogous to a clustering, and an underlying `relationship' between those structures. This allows a `structural comparison' between two similarity matrices using their predictability from `structure'. Significance is assessed with the help of re-sampling appropriate for each dataset. The software, CLARITY, is available as an R package from https://github.com/danjlawson/CLARITY.

stat.ME