bioRxiv doi: http://dx.doi.org/10.1101/008367
Modeling human population separation history using physically phased genomes
Shiya Song, Elzbieta Sliwerska, Sarah Emery, Jeffrey M Kidd
Phased haplotype sequences are a key component in many population genetic analyses since variation in haplotypes reflects the action of recombination, selection, and changes in population size. In humans, haplotypes are typically estimated from unphased sequence or genotyping data using statistical models applied to large reference panels. To assess the importance of correct haplotype phase on population history inference, we performed fosmid pool sequencing and resolved phased haplotypes of five individuals from diverse African populations (including Yoruba, Esan, Gambia, Massai and Mende). We physically phased 98% of heterozygous SNPs into haplotype-resolved blocks, obtaining a block N50 of 1 Mbp. We combined these data with additional phased genomes from San, Mbuti, Gujarati and CEPH European populations and analyzed population size and separation history using the Pairwise Sequentially Markovian Coalescent (PSMC) and Multiple Sequentially Markovian Coalescent (MSMC) models. We find that statistically phased haplotypes yield an earlier split-time estimation compared with experimentally phased haplotypes. To better interpret patterns of cross-population coalescence, we implemented an approximate Bayesian computation (ABC) approach to estimate population split times and migration rates by fitting the distribution of coalescent times inferred between two haplotypes, one from each population, to a standard Isolation-with-Migration model. We inferred that the separation between hunter-gather populations and other populations happened around 120,000 to 140,000 years ago with gene flow continuing until 30,000 to 40,000 years ago; separation between west African and out of African populations happened around 70,000 to 80,000 years ago, while the separation between Massai and out of African populations happened around 50,000 years ago.
Link
Showing posts with label Phasing. Show all posts
Showing posts with label Phasing. Show all posts
June 24, 2016
September 10, 2014
ASHG 2014 titles and abstracts
Some interesting titles from the ASHG 2014 conference.
UPDATE: I have added the abstracts.
The human X chromosome is the target of megabase wide selective sweeps associated with multi-copy genes expressed in male meiosis and involved in reproductive isolation. M. H. Schierup, K. Munch, K. Nam, T. Mailund, J. Y. Dutheil.
UPDATE: I have added the abstracts.
The human X chromosome is the target of megabase wide selective sweeps associated with multi-copy genes expressed in male meiosis and involved in reproductive isolation. M. H. Schierup, K. Munch, K. Nam, T. Mailund, J. Y. Dutheil.
The X chromosome differs from the autosomes in its hemizogosity in males and in its intimate relationship with the very different Y chromosome. It has a different gene content than autosomes and undergo specific processes such as meiotic sex chromosome inactivation (MSCI) and XY body formation. Previous studies have shown that natural selection is more efficient against deleterious mutations and, in chimpanzee, that positive selection is prevalent. We show that in all great apes species, megabase wide regions of the X chromosome has severely reduced diversity (by more than 80%). These regions are partly shared among species and indicate a large number of strong selective sweeps that have occurred independently on the same set of targets in different great apes species. We use simulations and deterministic calculations to show that background selection or soft selective sweeps are unlikely to be responsible. The regions also bear all the hallmarks of selective sweeps such as an increased proportion of singletons and higher divergence among closely related populations. Human populations are differently affected, suggesting that a large fraction of sweeps are private to specific human populations. The regions of reduced diversity correlates strongly with the position of X-ampliconic regions, which are 100-500 kb regions containing multiple copies of genes that are solely expressed during male meiosis. We propose that the genes in these regions escape MSCI and participate in an intragenomic conflict with regions of similar function on the Y chromosome for transmission of sex chromosomes to the next generation, i.e. sex chromosome meiotic drive. Recent results from Neanderthal introgression into humans point to the same regions as showing no introgression, consistent with the above process leading to reproductive isolation. Strikingly, the same regions of the X also shows much reduced divergence between human and chimpanzee, suggesting either that this speciation process was indeed complex or that the same regions were under strong selection in the human chimpanzee ancestor.
New insights on human de novo mutation rate and parental age. W. S. W. Wong, B. Solomon, D. Bodian, D. Thach, R. Iyer, J. Vockley, J. Niederhuber.
Germline mutations have a major role to play in evolution. Much attention has been given to studying the pattern and rate of human mutations using biochemical or phylogenetic methods based on closely related species. Massively parallel sequencing technologies have given scientists the opportunity to study directly measured de novo mutations (DNMs) at an unprecedented scale. Here we report the largest study (to our knowledge) of de novo point mutations in humans, in which we used whole genome deep sequencing (~60x) data from 605 family trios (father, mother and newborn). These trios represent the first group of approximately 2,700 trios who have undergone whole-genome sequencing (WGS) through our pediatric-based WGS research studies. The fathers ages range from 17 to 63 years and the mothers ages range from 17 to 43 years. We identified over 23000 DNMs (~40 per newborn) in the autosomal chromosomes using a customized pipeline and infer that the mutation rate per basepair is around 1.2x10-8 per generation, well within the reported range in previous studies. We were also able to confirm that the total number of DNMs in the newborn was directly proportional to the paternal age (P less than 2x10-16). Maternal age is shown to have a small but significant positive effect on the number of DNMs passed onto the offspring, (P =0.003) , even after accounting for the paternal age. This contradicts the prior dogma that maternal age only has an effect on chromosomal abnormalities related to nondisjunction events. Furthermore, 5% (22 total) of newborns in the analyzed group were conceived with assisted reproductive technologies (ARTs), and these infants have on average 5 more DNMs (Bias corrected and accelerated bootstrap 95% Confidence Interval, 1.24 to 8.00) than those conceived naturally, after controlling for both parents ages. Both parents ages remain significant as independently correlated with DNMs even after the families that used ARTs were removed from the analysis. Our study enhances current knowledge related to the human germline mutational rates.
Alignment to an ancestry specific reference genome discovers additional variants among 1000 Genomes ASW Cohort. R. A. Neff, J. Vargas, G. H. Gibbons, A. R. Davis.
Whole genome sequencing studies across certain populations, such as those with African ancestry, are often underpowered due to a larger divergence between the common reference genome and the true genetic sequence of the population. However, a common reference genome is not designed to account for this divergence in population-specific studies. Strong signals from common (MAF>50%) single nucleotide polymorphisms (SNPs), insertion-deletions (indels), and structural variants (SVs) can make alignment and variant calling difficult by masking nearby variants with weaker genetic signals. We present the results generated from alignment to an African descent population-specific reference genome by applying variants present in a majority of individuals with African descent from all phases of the 1000 Genomes Project and the International HapMap Consortium. We identified 882,826 single nucleotide polymorphisms, short insertion-deletion events, and large structural variations present at MAF>50%; in the population, representing 2.39 MB of genetic variation changed from hg19. We demonstrate that utilization of a population-specific reference improves variant call quality, coverage level, and imputation accuracy. We compared alignment of 27 African-American SW population (ASW) samples from the 1000 Genomes Phase 1 project between the population-specific and the hg19 reference. We discovered an additional 443,036 SNPs by alignment to the population specific reference in union across all samples, including thousands of exonic variants that are non-synonymous and are clinically relevant to the study of disease.
Using compressed data structures to capture variation in thousands of human genomes. S. A. McCarthy, Z. Lui, J. T. Simpson, Z. Iqbal, T. M. Keane, R. Durbin.
Currently the most widely used approach to catalogue variation amongst a set of samples is to align the sequencing reads to a single linear reference genome. This principle has been at the core of the 1000 Genomes data processing pipeline since the pilot phase of the project. However, there is now an increased awareness of the limitations of this approach, such as alignment artefacts, reference bias and unobserved variation on non-reference haplotypes. The Burrows-Wheeler transform and FM-index are compact data structures that have been successfully used in sequence alignment and assembly. One of the key features of these structures is that they are a searchable and reference-free representation of the raw sequencing reads. Our project aims to build a web server based on BWT data structures containing all the reads from many thousands of samples so as to efficiently retrieve matching reads and information about samples and populations. Enticingly, it is expected that data storage for this system would plateau as we collect more data since most new sequencing reads will have already been observed. We expect this to enable powerful new ways to query variation data from thousands of individuals. For the first phase of this project, we include all 87 Tbp of the low-coverage and exome data from the 2,535 samples in 1000 Genomes Phase 3. We envisage this would provide a means for researchers to easily check the prevalence of any human sequence in a control set of thousands of putatively healthy samples. We present our approaches and initial benchmarks on variant sensitivity and specificity against truth datasets and explore several applications for these structures such as validation of short insertion/deletion and structural variant calls, and rapid searching for traces of viral DNA.
Second-generation PLINK: Rising to the challenge of larger and richer datasets. C. C. Chang, C. C. Chow, L. C. A. M. Tellier, S. Vattikuti, S. M. Purcell, J. J. Lee.
PLINK 1 is a widely used open-source C/C++ toolset for genome-wide association studies (GWAS) and research in population genetics. However, the steady accumulation of data from imputation and whole-genome sequencing studies has exposed a strong need for even faster and more scalable implementations of key functions. In addition, GWAS and population-genetic data now frequently contain probabilistic calls, phase information, and/or multiallelic variants, none of which can be represented by PLINK 1's primary data format. To address these issues, we are developing a second-generation codebase for PLINK. The first major release from this codebase, PLINK 1.9, introduces extensive use of bit-level parallelism, O(sqrt(n))-time/constant-space Hardy-Weinberg equilibrium and Fisher's exact tests, and many other algorithmic improvements. In combination, these changes accelerate most operations by 1-4 orders of magnitude, and allow the program to handle datasets too large to fit in RAM. This will be followed by PLINK 2.0, which will introduce (a) a new data format capable of efficiently representing probabilities, phase, and multiallelic variants, and (b) extensions of many functions to account for the new types of information. The second-generation versions of PLINK will offer dramatic improvements in performance and compatibility. For the first time, users without access to high-end computing resources can perform several essential analyses of the feature-rich and very large genetic datasets coming into use.
Exploring genetic variation and genotypes among millions of genomes. R. M. Layer, A. R. Quinlann.
Integrated analysis of protein-coding variation in over 90,000 individuals from exome sequencing data. D. G. MacArthur, M. Lek, E. Banks, R. Poplin, T. Fennell, K. Samocha, B. Thomas, K. Karczewski, S. Purcell, P. Sullivan, S. Kathiresan, M. I. McCarthy, M. Boehnke, S. Gabriel, D. M. Altshuler, G. Getz, M. J. Daly, Exome Aggregation Consortium.
Rare, and thus largely unknown, variants are a major reason that, typically, less than 10% of the heritability of complex diseases currently can be explained by known genetic variation. While increasing the number of sequenced genomes may improve our ability to reveal this “hidden heritability,” the scale of the resulting dataset poses substantial storage and computational demands. Current efforts to sequence 100,000 genomes, and combined efforts that are likely to surpass 1 million genomes will identify hundreds of millions to billions of polymorphic loci. The minimum storage requirement for directly representing the variability found by these projects (1 bit per individual per variant, ignoring the necessary metadata) will range from terabytes to petabytes. Like most big-data problems, a balance must be found between optimizing storage and computational efficiency. For example, while compression can minimize storage by reducing file size, it can also cause inefficient computation since data must be decompressed before it can be analyzed. Conversely, highly structured data can reduce analysis times but typically require extra metadata that increase file size. Current variation storage schemes were not designed to quickly analyze massive datasets and fail to balance these competing goals. We present GENOTQ, an open source API and toolkit that reduces file size and data access time through use of a succinct data structure, a class of data structures that compress data such that operations can be performed without requiring the full decompression. Word aligned hybrid (WAH) bitmap compression is one such data structure that was developed to improve query times for relational databases. Binary values are encoded such that logical operations (AND, OR, NOT) can be performed on the compressed data. This encoding results in file sizes that are 20X smaller than uncompressed versions, and only 50% larger than the compressed version. Queries, such as finding shared variants among a subpopulation, are also 21X faster. Furthermore, representing the genotypes in this manner makes our method well suited to both distributed architectures like BigQuery and parallel processors like GPUs. We stress that this method is only part of a larger solution that would incorporate genomic annotations, medical histories, and pedigrees. Incorporating fast genotype queries with this web of metadata will provide a rich information source to both clinicians and researchers.
Capture of 390,000 SNPs in dozens of ancient central Europeans reveals a population turnover in Europe thousands of years after the advent of farming. I. Lazaridis, W. Haak, N. Patterson, N. Rohland, S. Mallick, B. Llamas, S. Nordenfelt, E. Harney, A. Cooper, K. W. Alt, D. Reich.
To understand the population transformations that took place in Europe since the early Neolithic, we used a DNA capture technique to obtain reads covering ~390 thousand single nucleotide polymorphisms (SNPs) from a number of different archaeological cultures of central Europe (Germany and Hungary). The samples spanned the time period from 7,500 BP to 3,500 BP (Early Neolithic to Early Bronze Age periods) and most of them were previously studied using mtDNA (Brandt, Haak et al., Science, 2013). The captured SNPs include about 360,000 SNPs from the Affymetrix Human Origins Array that were discovered in African individuals, as well as about 30,000 SNPs chosen for other reasons (that are thought to have been affected by natural selection, or to have phenotypic effects, or are useful in determining Y-chromosome haplogroups). By analyzing this data together with a dataset of 2,345 present-day humans and other published ancient genomes, we show that late Neolithic inhabitants of central Europe belonging to the Corded Ware culture were not a continuation of the earlier occupants of the region. Our results highlight the importance of migration and major population turnover in Europe long after the arrival of farming. * Contributed equally to this work.
Insights into British and European population history from ancient DNA sequencing of Iron Age and Anglo-Saxon samples from Hinxton, England. S. Schiffels, W. Haak, B. Llamas, E. Popescu, L. Loe, R. Clarke, A. Lyons, P. Paajanen, D. Sayer, R. Mortimer, C. Tyler-Smith, A. Cooper, R. Durbin.
British population history is shaped by a complex series of repeated immigration periods and associated changes in population structure. It is an open question however, to what extent each of these changes is reflected in the genetic ancestry of the current British population. Here we use ancient DNA sequencing to help address that question. We present whole genome sequences generated from five individuals that were found in archaeological excavations at the Wellcome Trust Genome Campus near Cambridge (UK), two of which are dated to around 2,000 years before present (Iron Age), and three to around 1,300 years before present (Anglo-Saxon period). Good preservation status allowed us to generate one high coverage sequence (12x) from an Iron Age individual, and four low coverage sequences (1x-4x) from the other samples. By providing the first ancient whole genome sequences from Britain, we get a unique picture of the ancestral populations in Britain before and after the Anglo-Saxon immigrations. We use modern genetic reference panels such as the 1000 Genomes Project to examine the relationship of these ancient samples with present day population genetic data. Results from principal component analysis suggest that all samples fall consistently within the broader Northern European context, which is also consistent with mtDNA haplogroups. In addition, we obtain a finer structural genetic classification from rare genetic variants and haplotype based methods such as FineStructure. Reflecting more recent genetic ancestry, results from these methods suggest significant differences between the Iron Age and the Anglo-Saxon period samples when compared to other European samples. We find in particular that while the Anglo-Saxon samples resemble more closely the modern British population than the earlier samples, the Iron Age samples share more low frequency variation than the later ones with present day samples from southern Europe, in particular Spain (1000GP IBS). In addition the Anglo-Saxon period samples appear to share a stronger older component with Finnish (1000GP FIN) individuals. Our findings help characterize the ancestral European populations involved in major European migration movements into Britain in the last 2,000 years and thus provide more insights into the genetic history of people in northern Europe.
Fine-scale population structure in Europe. S. Leslie, G. Hellenthal, S. Myers, P. Donnelly, International Multiple Sclerosis Genetics Consortium.
There is considerable interest in detecting and interpreting fine-scale population structure in Europe: as a signature of major events in the history of the populations of Europe, and because of the effect undetected population structure may have on disease association studies. Population structure appears to have been a minor concern for most of the recent generation of genome-wide association studies, but is likely to be important for the next generation of studies seeking associations to rare variants. Thus far, genetic studies across Europe have been limited to a small number of markers, or to methods that do not specifically account for the correlation structure in the genome due to linkage disequilibrium. Consequently, these studies were unable to group samples into clusters of similar ancestry on a fine (within country) scale with any confidence. We describe an analysis of fine-scale population structure using genome-wide SNP data on 6,209 individuals, sampled mostly from Western Europe. Using a recently published clustering algorithm (fineSTRUCTURE), adapted for specific aspects of our analysis, the samples were clustered purely as a function of genetic similarity, without reference to their known sampling locations. When plotted on a map of Europe one observes a striking association between the inferred clusters and geography. Interestingly, for the most part modern country boundaries are significant i.e. we see clear evidence of clusters that exclusively contain samples from a single country. At a high level we see: the Finns are the most differentiated from the rest of Europe (as might be expected); a clear divide between Sweden/Norway and the rest of Europe (including Denmark); and an obvious distinction between southern and northern Europe. We also observe considerable structure within countries on a hitherto unseen fine-scale - for example genetically distinct groups are detected along the coast of Norway. Using novel techniques we perform further analyses to examine the genetic relationships between the inferred clusters. We interpret our results with respect to geographic and linguistic divisions, as well as the historical and archaeological record. We believe this is the largest detailed analysis of very fine-scale human genetic structure and its origin within Europe. Crucial to these findings has been an approach to analysis that accounts for linkage disequilibrium.
The population structure and demographic history of Sardinia in relationship to neighboring populations. J. Novembre, C. Chiang, J. Marcus, C. Sidore, M. Zoledziewska, M. Steri, H. Al-asadi, G. Abecasis, D. Schlessinger, F. Cucca.
Numerous studies have made clear that Sardinian populations are relatively isolated genetically from other populations of the Mediterranean, and more recently, intriguing connections between Sardinian ancestry and early Neolithic ancient DNA samples have been made. In this study, we analyze a whole-genome low-coverage sequencing dataset from 2120 Sardinians to more fully characterize patterns of genetic diversity in Sardinia. The study contains one subsample that contains individuals from across Sardinia and a second subsample that samples 4 villages from the more isolated Ogliastra region. We also merge the data with published reference data from Europe and North Africa. Overall Fst values of Sardinia to other European populations are low (less than 0.015); however using a novel method for visualizing genetic differentiation on a geographic map, we formally show how Sardinia is more differentiated than would be expected given its geographic distance from the mainland, consistent with periods of isolation. Applications of the software Admixture show how Sardinia populations differ in the levels of recent admixture with mainland European populations and that there are only minor contributions from North African populations to Sardinian ancestry. Notably the Sardinians from Ogliastra contain a distinct genetic cluster with minimal evidence of recent admixture with mainland Europe. We found frequency-based f3 tests and the tree-based algorithm Treemix both also show minimal evidence of recent admixture. Given the relative isolation, one might expect to see a unique demographic history from neighboring populations. Using coalescent-based approaches, we find Sardinian populations have had more constant effective sizes over the past several thousand years than mainland European populations, which typically show evidence for rapid growth trajectories in the recent past. This unique demographic history has consequences for the abundance of putatively damaging and deleterious variants, and we use our data to address the prediction that the genetic architecture of disease traits is expected to involve fewer loci with a greater proportion of variants at common frequencies in Sardinia.
Population structure in African-Americans. S. Gravel, M. Barakatt, B. Maples, M. Aldrich, E. E. Kenny, C. D. Bustamante, S. Baharian.
We present a detailed population genetic study of 4 African-American cohorts comprising over 6000 genotyped individuals across US urban and rural communities: two nation-wide longitudinal cohorts, one biobank cohort, and the 1000 genomes ASW cohort. Ancestry analysis reveals a uniform breakdown of continental ancestry proportions across regions and urban/rural status, with 79% African, 19% European, and 1.5% Native American/Asian ancestries, with substantial between-individual variation. The Native Ancestry proportion is higher than previous estimates and is maintained after self-identified hispanics and individuals with substantial inferred Spanish ancestry are removed. This strongly supports direct admixture between Native Americans and African Americans on US territory, and linkage patterns suggest contact early after African-American arrival to the Americas. Local ancestry patterns and variation in ancestry proportions across individuals are broadly consistent with a single African-American population model with early Native American admixture and ongoing European gene flow in the South. The size and broad geographic sampling of our cohorts enables detailed analysis the geographic and cultural determinants of finer-scale population structure. Recent Identity-by-descent analysis reveals fine-scale structure consistent with the routes used during slavery and in the great African-American migrations of the twentieth century: east-to-west migrations in the south, and distinct south-to-north migrations into New England and the Midwest. These migrations follow transit routes available at the time, and are in stark contrast with European-American relatedness patterns.
Genetic testing of 400,000 individuals reveals the geography of ancestry in the United States. Y. Wang, J. M. Granka, J. K. Byrnes, M. J. Barber, K. Noto, R. E. Curtis, N. M. Natalie, C. A. Ball, K. G. Chahine.
The population of the United States is formed by the interplay of immigration, migration and admixture. Recent research (R. Sebro et al., ASHG 2013) has shed light on the U.S. demography by studying the self-reported ethnicity from the 2010 U.S. Census. However, self-reported ethnicity may not accurately represent true genetic ancestry and may therefore introduce unknown biases. Since launching its DNA service in May 2012, AncestryDNA has genotyped over 400, 000 individuals from the United States. Leveraging this huge volume of DNA data, we conducted a large-scale survey of the ancestry of the United States. We predicted genetic ethnicity for each individual, relying on a rigorously curated reference panel of 3,000 single-origin individuals. Combining that with birth locations, we explored how various ethnicities are distributed across the United States Our results reveal a distinct spatial distribution for each ethnicity. For example, we found that individuals from Massachusetts have the highest proportion of Irish genetic ancestry and individuals from New York have the highest proportion of Southern European genetic ancestry, indicating their unique immigration and migration histories. We also performed pairwise IBD analysis on the entire sample set and identified over 300 million shared genomic segments among all 400,000 individuals. From this data, we calculated the average amount of sharing for pairs of individuals born within the same state or from two different states. In general, we found the genetic sharing decreases as the geographic distance between two states increases. However, the pattern also varies substantially among the 50 states. In summary, our analysis has provided significant insight on the biogeographic patterns of the ancestry in the United States.
Statistical inference of archaic introgression and natural selection in Central African Pygmies. P. Hsieh, J. D. Wall, J. Lachance, S. A. Tishkoff, R. N. Gutenkunst, M. F. Hammer.
Recent evidence from ancient DNA studies suggests that genetic material introgressed from archaic forms of Homo, such as Neanderthals and Denisovans, into the ancestors of contemporary non-African populations. These findings also imply that hybridization may have given rise to some of adaptive novelties in anatomically modern humans (AMH) as they expanded from Africa into various ecological niches in Eurasia. Within Africa, fossil evidence suggests that AMH and a variety of archaic forms coexisted for much of the last 200,000 years. Here we present preliminary results leveraging high quality whole-genome data (>60X coverage) for three contemporary sub-Saharan African populations (Biaka, Baka, and Yoruba) from Central and West Africa to test for archaic admixture. With the current lack of African ancient DNA, especially in Central Africa due to its rainforest environment, our statistical inference approach provides an alternative means to understand the complex evolutionary dynamics among groups of the genus Homo. To identify candidate introgressive loci, we scan the genomes of 16 individuals and calculate S*, a summary statistic that was specifically designed by one of us (JDW) to detect archaic admixture. The significance of each candidate is assessed through extensive whole-genome level simulations using demographic parameters estimated by ∂a∂i to obtain a parametric distribution of S* values under the null hypothesis of no archaic introgression. As a complementary approach, top candidates are also examined by an approximate-likelihood computation method. The admixture time for each individual introgressive variant is inferred by estimating the decay of the genetic length of the diverged haplotype as a function of its underlying recombination rate. A neutrality test that controls for demography is performed for each candidate to test the hypothesis that introgressive variants rose to high frequency due to positive directional selection. Several genomic regions were identified by both selection and introgression scans, and we will discuss the possible genetic and functional properties of these “double-hits”. The present study represents one of the most comprehensive genomic surveys to date for evidence of archaic introgression to anatomically modern humans in Africa.
Inferences about human history and natural selection from 280 complete genome sequences from 135 diverse populations. S. Mallick, D. Reich, Simons Genome Diversity Project Consortium.
The most powerful way to study population history and natural selection is to analyze whole genome sequences, which contain all the variation that exists in each individual. To date, genome-wide studies of history and selection have primarily analyzed data from single nucleotide polymorphism (SNP) arrays which are biased by the choice of which SNPs to include. Alternatively they have analyzed sequence data that have been generated as part of medical genetic studies from populations with large census sizes, and thus do not capture the full scope of human genetic variation. Here we report high quality genome sequences (~40x average) from 280 individuals from 135 worldwide populations, including 45 Africans, 26 Native Americans, 27 Central Asians or Siberians, 46 East Asians, 25 Oceanians, 46 South Asians, and 71 West Eurasians. All samples were sequenced using an identical protocol at the same facility (Illumina Ltd.). We modified standard pipelines to eliminate biases that might confound population genetic studies. We report novel inferences, as well as a high resolution map that shows where archaic ancestry (Neanderthal and Denisovan) is distributed throughout the world. We compare and contrast the genomic landscape of the Denisovan introgression into mainland Eurasians to that in island Southeast Asians. We are making this dataset fully available on Amazon Web Services as a resource to the community, coincident with the American Society of Human Genetics meeting.
Improved haplotype phasing using identity by descent. B. L. Browning, S. R. Browning.
We present a new haplotype phasing method that achieves higher accuracy than existing methods. The method is based on the Beagle haplotype frequency model, but unlike the original Beagle phasing method, the new method incorporates genetic recombination, genotype error, and segments of identity by descent. We compared the new haplotype phasing method to Beagle (r1230) and to SHAPEIT version 2 (r778) using Illumina Human 1M SNP data for chromosome 20. We phased 44 HapMap3 CEU trio offspring together with subsets of Wellcome Trust Case Control Consortium 2 controls (n=650, 1300, 2600, 5200). Phase error was measured at trio offspring genotypes on chromosome 20 that have phase determined by parental genotypes. The SHAPEIT “states” parameter was set at 6400 in order to increase its phasing accuracy. The new haplotype phasing method produced haplotype switch error rates that were 20-25% lower than the error rates for the existing Beagle method and 1-7% lower than the error rates for SHAPEIT. The difference in switch error rates between the new method and SHAPEIT increased with increasing sample size. The new haplotype phasing method will be incorporated into version 4 of the Beagle software package (http://faculty.washington.edu/browning/beagle/beagle.html).
Reducing pervasive false positive identical-by-descent segments detected by large-scale pedigree analysis. E. Y. Durand, N. Eriksson, C. Y. McLean.
Analysis of genomic segments shared identical-by-descent (IBD) between individuals is fundamental to many genetic applications, from demographic inference to estimating the heritability of diseases. A large number of methods to detect IBD segments have been developed recently. However, IBD detection accuracy in non-simulated data is largely unknown. In principle, it can be evaluated using known pedigrees, as IBD segments are by definition inherited without recombination down a family tree. We extracted 25,432 genotyped European individuals containing 2,952 father-mother-child trios from the 23andMe, Inc. dataset. We then used GERMLINE, a widely used IBD detection method, to detect IBD segments within this cohort. Exploiting known familial relationships, we identified a false positive rate over 67% for 2-4 centiMorgan (cM) segments, in sharp contrast with accuracies reported in simulated data at these sizes. We show that nearly all false positives arise due to allowing switch errors between haplotypes when detecting IBD, a necessity for retrieving long (> 6 cM) segments in the presence of imperfect phasing. We introduce HaploScore, a novel, computationally efficient metric that enables detection and filtering of false positive IBD segments on population-scale datasets. HaploScore scores IBD segments proportional to the number of switch errors they contain. Thus, it enables filtering of spurious segments reported due to GERMLINE being overly permissive to imperfect phasing. We replicate the false IBD findings and demonstrate the generalizability of HaploScore to alternative genotyping arrays using an independent cohort of 555 European individuals from the 1000 Genomes project. HaploScore can be readily adapted to improve the accuracy of segments reported by any IBD detection method, provided that estimates of the genotyping error rate and switch error rate are available.
Parente2: A fast and accurate method for detecting identity by descent. S. Bercovici, J. M. Rodriguez, L. Huang, S. Batzoglou.
Identity-by-descent (IBD) inference is the problem of establishing a direct and explicit genetic connection between two individuals through a genomic segment that is inherited by both individuals from a recent common ancestor. IBD inference is key to a variety of population genomic studies, ranging from demographic studies to linking genomic variation with phenotype and disease. The problem of both accurate and efficient IBD detection has become increasingly challenging with the availability of large collections of human genotypes and genomes: given a cohort’s size, as quadratic number of pairwise genome comparisons must be performed, in principle. Therefore, computation time and the false discovery rate can also scale quadratically. To enable practical large-scale IBD detection, we developed Parente2, a novel method for detecting IBD segments. Parente2 is based on an embedded log-likelihood ratio and uses an ensemble windowing approach to model complex linkage disequilibrium in the underlying studied population. Parente2 is applied directly on genotype data without the need to phase data prior to IBD inference. Through extensive simulations using real data, we evaluate Parente2’s performance. We show that Parente2 is superior to previous state-of-the-art methods, detecting pairs of related individuals sharing a 4 cM IBD segment with 99.9%; sensitivity at a 0.1%; false positive rate, and achieving 79.2%; sensitivity at a 1%; false positive rate for the more challenging case of pairs sharing a 2 cM IBD segment. Additionally, Parente2 is efficient, providing one to two orders of magnitude speedup compared to previous state of the art methods. Parente2 is freely available at http://parente.stanford.edu/.
Fast PCA of very large samples in linear time. K. J. Galinsky, P. Loh, G. Bhatia, S. Georgiev, S. Mukherjee, N. J. Patterson, A. L. Price.
Principal components analysis (PCA) is an effective tool for inferring population structure and correcting for population stratification in genetic data. Traditionally, PCA runs in O(MN2+N3 ) time, where M is the number of variants and N is the number of samples. Here, we describe a new algorithm, fastpca, for approximating the top K PCs that runs in time O(MNK), making use of recent advances in random low-rank matrix approximation algorithms (Rokhlin et al. 2009). fastpca avoids computing the GRM and associated computational and memory storage costs, enabling PCA of very large datasets on standard hardware. We estimated the top 10 PCs of the WTCCC dataset (16k samples, 101k variants) in roughly 7 minutes while consuming 1GB of RAM, compared to 1 hour and 2.5GB for PLINK2. The fastpca approximation was extremely accurate (r2>99% between all fastpca and PLINK2 PCs). The improvement in running time becomes even larger at larger samples sizes; for example, fastpca estimated the top 10 PCs of a simulated data set with 100k samples and 300k variants in 135 minutes 8.5GB of RAM, vs. an estimated 350 hours and 85GB of RAM using PLINK2. A recently published O(MN2) time method, flashpca, did not complete on this data set due to exceeding 40GB memory requirement. All of these analyses were based on LD-pruning SNPs with r2>0.2, which leads to much more accurate PCs in simulations as compared to retaining all SNPs; more complex LD-adjustment strategies provide only a small further improvement.
Fast detection of IBD segments associated with quantitative traits in genome-wide association studies. Z. Wang, E. Kang, B. Han, S. Snir, E. Eskin.
Recently, many methods have been developed to detect the identity-by-descent (IBD) segments between a pair of individuals. These methods are able to detect very small shared IBD segments between a pair of individuals up to 2 centimorgans in length. This IBD information can be used to identify recent rare mutations associated with phenotype of interest. Previous approaches for IBD association were applicable to case/control phenotypes. In this work, we propose a novel and natural statistic for the IBD association testing, which can be applied to quantitative traits. A drawback of the statistic is that it requires a large number of permutations to assess the significance of the association, which can be a great computational challenge. We make a connection between the proposed statistic and linear models so that it does not require permutations to assess the significance of an association. In addition, our method can control population structure by utilizing linear mixed models.
Long-range haplotype mapping in Hispanic/Latinos reveals loci for short stature. G. Belbin, D. Ruderfer, K. Slivinski, M.C. Yee, J. Jeff, O. Gottesman, E.A. Stahl, R.J.F. Loos, E.P. Bottinger, E.E. Kenny.
The Hispanic/Latino (HL) population of Northern Manhattan represents a diverse recent diaspora population, with 95% of the individuals reporting having grandparents born outside of the United States. Of these 43% report grandparents born in Puerto Rico, 23% the Dominican Republic, 13% Central America, and 5%, 4%, and 2% from Mexico, South America, and Europe respectively. Despite complex patterns of migration, admixture, and diversity, strong signatures of cryptic relatedness persist amongst HLs. We have detected long-range genomic tract sharing (>3cM), or identity-by-descent (IBD), across 5,194 HL in the Mount Sinai BioMe Biobank. We observed an average population level IBD sharing of 0.0025 in HL, which is 2.5- and 5-fold higher than that observed in BioMe European- and African-American populations, respectively. We hypothesize that these patterns of recent migration and genetic drift may drive some otherwise rare functional alleles to detectable frequency. We clustered groups of homologous IBD tracts (n=112,250) segregating in this HL population. We observed that IBD clusters represent a class of low frequency alleles (median minor allele frequency =0.0077, s.d.=0.0015). We performed a genome-wide association of the IBD clusters, or ‘population-based linkage’, to detect loci implicated in height, a highly heritable polygenic trait. 15 independent loci surpassed our empirically derived genome-wide significance threshold of less than 4.4710-4, 11 of which replicated in an independent cohort of BioMe HLs. Strikingly, two regions confer strong recessive effects. In the case of the top hit on 9q32 (MAF less than 0.005; p less than8x10-6), homozygous non-referent individuals were shorter by 6” or 10”, for men or women, respectively, compared to the population mean (5’ 7” and 5’ 2” for men and women, respectively). In addition, IBD haplotypes in the 9q32 cluster harbored a significant enrichment of Native American ancestry (p less than 1x10-16). Finally, this interval contains a number of biologically compelling candidate genes, including COL27A1 and PALM2. This study demonstrates that rich population structure, rather than being a confounding factor in biomedical discovery efforts, may be leveraged to reveal novel genetic associations with complex human traits.
A haplotype reference panel of over 31,000 individuals and next-generation imputation methods. S. Das, on behalf of Haplotype Reference Consortium.
Genotype imputation is now a key tool in the analysis of human genetic studies, enabling array-based genetic association studies to examine the millions of variants that are being discovered by advances in whole genome sequencing. Examining these variants increases power and resolution of genetic association studies and makes it easier to compare the results of studies conducted using different arrays. Genotype imputation improves in accuracy with increasing numbers of sequenced samples, particularly for low frequency variants. The goal of the Haplotype Reference Consortium is to combine haplotype information from ongoing whole genome sequencing studies to create a large imputation resource. To date, we have collected information on >31,500 sequenced whole genomes, aggregated over 20 studies of predominantly European ancestry, to create a very large reference panel of human haplotypes where ~50M genetic variants are observed 5 or more times. These haplotypes can be used to guide genotype imputation and haplotype estimation. In preliminary empirical evaluations, our panel provides substantial increases in accuracy relative to the 1000 Genomes Project Phase 1 reference panel and other smaller panels, particularly for variants with frequency less than
5%. I will describe our evaluation of strategies for merging haplotypes and variant lists across studies and advances in methods for genotype likelihood-based haplotype estimation that can be applied to 10,000s of samples. I will also summarize new methods for next generation imputation that perform faster and require less memory than contemporary methods while attaining similar levels of imputation accuracy. Our full resource is available to the community through imputation servers that enable scientists to impute missing variants in any study and respect the privacy of subjects contributing to the studies that constitute the Haplotype Reference Consortium. The majority of haplotypes will also be deposited in the European Genotype Archive.
A rare variant local haplotype sharing method with application to admixed populations. S. Hooker, G. T. Wang, B. Li, Y. Guan, S. M. Leal.
With the advent of next generation sequencing there is great interest in studying the involvement of rare variants in complex trait etiology. For many complex traits sequence data is being generated on DNA samples from African Americans and Hispanics to elucidate rare variant associations. Analyses of admixed populations present special challenges due to spurious associations which can occur because of confounding. However using information on admixture and local ancestry can also be highly beneficial and increase the power to detect associations in these populations. Here a local haplotype sharing (LHS) method (Xu and Guan 2014) was extended to test for rare variant (RV) associations in admixed populations. Previously the Weighted Haplotype and Imputation-based Test (WHAIT) (Li et al. 2010) was proposed to test for rare variant associations using haplotype data. The RV-LHS method unlike WHAIT, does not require reconstruction of haplotypes which can be both computationally intensive and error prone. Additionally the RV-LHS uses information on local ancestry which is particularly advantageous when analyzing admixed populations. Results will be shown from simulation studies performed for rare variant data from an admixed population. Both Type I and II errors are evaluated for the RV-LHS method. Additionally the power of the RV-LHS method is compared to WHAIT as well as several other non-haplotype-based rare variant association methods including the combined multivariate collapsing (CMC) (Li and Leal, 2008), Variable Threshold (VT) (Price et al. 2010) and Sequence Kernel Association Test (SKAT) (Wu et al. 2010). Several heart, lung and blood phenotypes were analyzed using sequence data on African-Americans from the NHLBI-Exome Sequencing Project to better evaluate the performance of the RV-LHS compared to other rare variant association methods.
September 06, 2013
ASHG 2013 abstracts
Feel free to point me to more interesting abstracts than the ones I noticed during my "first pass".
Morphometric and ancient DNA study of human skeletal remanants in Indian Subcontinent.
N. Rai et al.
LL. Kang et al.
F. L. Mendez, M. F. Hammer University of Arizona, Tucson, AZ., USA.
J. C. Martinez-Cruzado et al.
M. Zhabagin et al.
M. Nothnagel et al.
T. Blauwkamp et al.
P. H. Hsieh et al.
A. T. Duggan et al.
Reconstructing Austronesian population history.
M. Lipson et al.
R. Do et al.
J. L. Rodriguez-Flores et al.
C. Lewis et al.
Morphometric and ancient DNA study of human skeletal remanants in Indian Subcontinent.
N. Rai et al.
Recovery and sequencing of mtDNA from ancient human remnants is a daunting task but provides valuable information about human migrations and evolution. Our present study is the first to recover, amplify and sequence (HVR and coding regions of mtDNA) inadequately preserved and highly degraded (1.5 Ky to ≤1.0 Ky ago) hominids mitochondrial DNA of three most intriguing and indigenous ancient population of South and South-East Asia (Myanmar=20 Buried individuals, Nicobar Islands=15 and Andaman Island=6). Following all parameters and to avoid the chance of contamination we independently extracted and sequenced the DNA in two different labs and measured the cranial variability in all hominid skulls using 128 cranial landmarks, compiled 3D morphometrics, genetic data of ancient DNA samples and analyzed the admixture and genetic affinities of above three populations. Results showed the predominant frequency of F1a1 and complete absence of 9bp deletion in ancient Nicobarese. Unlike in previous reports on modern Nicobarese, the high frequency of F1a1 haplogroup in ancient Nicobarese show the probable migration of Nicobarese from South East Asia and the complete absence of 9bp deletion suggests the different events of settlement. This study failed to detect genetic affinities of Burmese with Nicolbarese even though their phenotype and language appears to be same. We first time report any kind of population study on Burmese populations and with the genetic affinity of Burmese with East Asian, East Indian (Including Gadhwal region of Himalaya) and Bangladeshi populations, we found significant admixture with West Eurasians. Our study strongly supports the West Eurasian and East Asian route of migration and settlement of early Burmese population. The three populations in the present study are quite different in their genetic structure but 3D morphometric study using huge number of landmarks explains a close homology among these populations and this can be explained by the role of climatic signature on these populations.Y chromosomes of ancient Hunnu people and its implication on the phylogeny of East Asian linguistic families.
LL. Kang et al.
The Hunnu (Xiongnu) people, also called Huns in Europe, were the largest ethnic group to the north of Han Chinese until the 5th century. The ethno-linguistic affiliation of the Hunnu is controversial among Yeniseian, Altaic, Uralic, and Indo-European. Ancient DNA analyses on the remains of the Hunnu people had shown some clues to this problem. Y chromosome haplogroups of Hunnu remains included Q-M242, N-Tat, C-M130, and R1a1. Recently, we analyzed three samples of Hunnu from Barköl, Xinjiang, China, and determined Q-M3 haplogroup. Therefore, most Y chromosomes of the Hunnu samples examined by multiple studies are belonging to the Q haplogroup. Q-M3 is mostly found in Yeniseian and American Indian peoples, suggesting that Hunnu should be in the Yeniseian family. The Y chromosome diversity is well associated with linguistic families in East Asia. According to the similarity in the Y chromosome profiles, there are four pairs of congenetic families, i.e., Austronesian and Tai-Kadai, Mon-Khmer and Hmong-Mien, Sino-Tibetan and Uralic, Yeniseian and Palaesiberian. Between 4,000-2,000 years before present, Tai-Kadai, Hmong-Mien, Sino-Tibetan, and Yeniseian languages transformed into toned analytic languages, becoming quite different from the rest four. Since Hunnu was in the Yeniseian family, all these four toned families were distributed in the inland of China during the transformations. There must be some social or biological factors induced the transformations at that time, which is worth doing more linguistic and genetic researches.Genomic scans for haplotypes of Denisova and Neanderthal ancestry in modern human populations.
F. L. Mendez, M. F. Hammer University of Arizona, Tucson, AZ., USA.
Evidence of archaic introgression into modern humans has accumulated in recent years. While most efforts to characterize the introgression process have relied on genome averages, only a small number of introgressive haplotypes have been shown to have an archaic origin after rejection of the alternative hypothesis of incomplete lineage sorting. Accurate identification of introgressive haplotypes is crucial both to characterize potentially functional consequences of archaic admixture and to quantify more precisely the genomic impact of archaic introgression. We perform two independent genomic scans for haplotypes of Denisova and of Neanderthal origin in a geographically diverse sample of complete genome sequences. These scans are based on the local sharing of polymorphisms and linkage disequilibrium, respectively. The analysis of concordance between the methods is then used to estimate the power and to compare demographic inference when performed using either all the data or just the genomic regions with no evidence of introgression. Moreover, we evaluate the extent to which Denisova haplotypes are observed in non-Melanesian populations, and investigate whether the presence of such haplotypes is better explained by their persistence in the population since introgression or by more recent gene flow from Melanesians.
Admixture Estimation in a Founder Population.
Y. Banda1 et al.
Admixture between previously diverged populations yields patterns of genetic variation that can aid in understanding migrations and natural selection. An understanding of individual admixture (IA) is also important when conducting association studies in admixed populations. However, genetic drift, in combination with shallow allele frequency differences between ancestral populations, can make admixture estimation by the usual methods challenging. We have, therefore, developed a simple but robust method for ancestry estimation using a linear model to estimate allele frequencies in the admixed individual or sample as a function of ancestral allele frequencies. The model works well because it allows for random fluctuation in the observed allele frequencies from the expected frequencies based on the admixture estimation. We present results involving 3,366 Ashkenazi Jews (AJ) who are part of the Kaiser Permanente Genetic Epidemiology Research on Adult Health and Aging (GERA) cohort and genotyped at 674,000 SNPs, and compare them to the results of identical analyses for 2,768 GERA African Americans (AA). For the analysis of the AJ, we included surrogate Middle Eastern, Italian, French, Russian, and Caucasus subgroups to represent the ancestral populations. For the African Americans, we used surrogate Africans and Northern Europeans as ancestors. For the AJ, we estimated mean ancestral proportions of 0.380, 0.305, 0.113, 0.041 and 0.148 for Middle Eastern, Italian, French, Russian and Caucasus ancestry, respectively. For the African Americans, we obtained estimated means of 0.745 and 0.248 for African and European ancestry, respectively. We also noted considerably less variation in the individual admixture proportions for the AJ (s.d. = .02 to .05) compared to the AA (s.d.= .15), consistent with an older age of admixture for the former. From the linear model regression analysis on the entire population, we also obtain estimates of goodness of fit by r2. For the analysis of AJ, the r2 was 0.977; for the analysis of the AA, the r2 was 0.994, suggesting that genetic drift has played a more prominent role in determining the AJ allele frequencies. This was confirmed by examination of the distribution of differences for the observed versus predicted allele frequencies. As compared to the African Americans, the AJ differences were significantly larger, and presented some outliers which may have been the target of selection (e.g. in the HLA region on chromosome 6p).Admixture in the Pre-Columbian Caribbean.
J. C. Martinez-Cruzado et al.
The biological origin of the Caribbean aborigines that greeted Columbus is one of the most controversial issues regarding the population history of this region. Genome studies suggest an Equatorial-Tucanoan origin, consistent with the Arawakan language spoken by most natives of the region. However, the archaeological evidence suggests an early arrival from Mesoamerica, and their admixture with the more recent Arawak-speaking group stemming from the Amazon remains a possibility. The lineages comprehending most Puerto Rican samples belonging to haplogroups B1 and C1, which in turn encompass 44% of all Native American mtDNAs in the island, have an unambiguous South American origin. However, none of those belonging to haplogroup A2, encompassing 52% of all Native American mtDNAs, have been related to South America or any other continental region. To augment the scarce data from Mesoamerican countries other than Mexico, we present the complete mtDNA sequence of 6 Honduran samples belonging to distinct control region lineages in addition to 3 from the Dominican Republic and 3 from Puerto Rico. Interestingly, maximum likelihood phylogenetic reconstruction including 40 published haplogroup A2 sequence haplotypes from Mesoamerica, Central America and South America clusters 8 out of 10 Mesoamerican and Andean haplotypes in a deep rooted group, separate from, and excluding all Costa Rican, Panamian and Brasilian haplotypes, suggesting a relatively recent origin for Chibchan-Paezan and Amazonian groups. Furthermore, 4 of the 5 Greater Antillean A2 haplotypes are included in the deeply rooted Mesoamerican-Andean cluster. Moreover, the only Cuban haplotype in the literature and the remaining A2 haplotype from the Dominican Republic form even more deeply rooted private branches. Similarly, the only haplogroup C1d sample sequenced from the Dominican Republic forms a private branch with the deepest root in a maximum likelihood tree containing 19 additional C1d haplotypes from Mexico to Brasil plus the CRS. In conclusion, our preliminary results suggest that a substantial proportion of the Native American mtDNA lineages from the Greater Antilles do not share an Amazonian origin with the language their people spoke in 1492. Furthermore, the position of two Dominican lineages at the earliest split in both their respective trees suggests an early origin that could be explained by extensive lineage extinctions in Mesoamerica and the Andes or an origin in North America.The possible role of social selection in the distribution of the "Proto-Mongolian" haplotype in Kazakhs, Kyrgyz, Mongols and other Eurasian populations.
M. Zhabagin et al.
Social factors may be important contributors to reproductive success and determination of the selective survival of individuals. Therefore, social selection and other social factors are important for understanding population structure and its formation. The role of social selection on the distribution and formation of Y-chromosomal gene pool has been studied. There is a strong connection between social selection and birth rate of the descendants, whose fathers had achieved high social status during the expansion of the Mongol Empire and associated historical events. A total of 783 haplotypes, including 687 newly obtained and 96 retrieved from the literature were assigned to the haplogroup C3*-M217 (xM48) based on genotyping 17 Y-chromosomal STR markers. These haplotypes represent 11 populations of Eurasia: Kazakhs, Mongols, Kyrgyz, Telengits, Circassians, Balkar, Temirgoys, Karachai, Evenki, Kizhi and the Pashtuns. As the result, a major haplotype 13-16-25-15-16-18-14-10-22-11-10-11-13-10-21 (DYS389a-DYS389b-DYS390-DYS456-DYS19-DYS458-DYS437-DYS438-DYS448-GATA4-DYS391-DYS392-DYS393-DYS439-DYS635, N=94) was found to have 12.00% frequency within haplogroup C3*. This haplotype includes and extends the previously described “star-cluster” haplotype. Noteworthy, the frequency of this major haplotype within haplogroup C3* was 16.80% in Kazakhs, 10.13% in Mongols and 2.63% in Kirgiz who are not considered as direct descendants of Genghis Khan. 35.10% of the major haplotype was represented by Kazakh tribe Ashamayly-Kerey, 17.02% by the Khalkh Mongols and 7.44% by the Barguts. Therefore, we suppose this major ancestral haplotype to be the "proto-Mongolian haplotype", inherited by Genghis Khan and his descendants. It is important to mention that Temujin belongs to Kiyat-Borjigin tribe that in turn is a branch of the bigger Borjigin tribe, part of the Khalkh Mongols. Thus, Genghis Khan might be considered as a carrier rather than founder of the star-cluster haplotype. He and his descendants are the ones who contributed to a positive effect of social selection in the distribution of this haplotype. Other examples are the Barguts, who had Genghis Khan’s credit and were granted with a number of privileges, or the Kerey, based on the fact that Temujin had been brought up at the court of the Togrul Khan, belonging to the Kerey tribe.Y-chromosomal variation in native South Americans: bright dots on a gray canvas.
M. Nothnagel et al.
While human populations in Europe and Asia have often been reported to reveal a concordance between their extant genetic structure and the prevailing regional pattern of geography and language, such evidence is lacking for native South Americans. In the largest study of South American natives to date, we examined the relationship between Y-chromosomal genotype on the one hand, and male geographic origin and linguistic affiliation on the other. We observed virtually no structure for the extant Y-chromosomal genetic variation of South American males that could sensibly be related to their inter-tribal geographic and linguistic relationships, augmented by locally confined Y-STR autocorrelation. Analysis of repeatedly taken random subsamples from Europe adhering to the same sampling scheme excluded the possibility that this finding was due to our specific scheme. Furthermore, for the first time, we identified a distinct geographical cluster of Y-SNP lineages C-M217 (C3*) in South America, which are virtually absent from North and Central America, but occur at high frequency in Asia. Our data suggest a late introduction of C3* into South America no more than 6,000 years ago and low levels of migration between the ancestor populations of C3* carrier and non-carriers. Our findings are consistent with a rapid peopling of the continent, followed by long periods of isolation in small groups, and highlight the fact that a pronounced correlation between genetic and geographic/cultural structure can only be expected under very specific conditions.
The timing and history of Neandertal gene flow into modern humans.
S. Sankararaman et al.
Previous analyses of modern human variation in conjunction with the Neandertal genome have revealed that Neandertals contributed 1-4% of the genes of non-Africans with the time of last gene flow dated to 37,000-86,000 years before present. Nevertheless, many aspects of the joint demographic history of modern humans and Neandertals are unclear. We present multiple analyses that reveal details of the early history of modern humans since their dispersal out of Africa.
1.We analyze the difference between two allele frequency spectra in non-Africans: the spectrum conditioned on Neandertals carrying a derived allele while Denisovans carry the ancestral allele and the spectrum conditioned on Denisovans carrying a derived allele while Neandertals carry the ancestral allele. This difference spectrum allows us to study the drift since Neandertal gene flow under a simple model of neutral evolution in a panmictic population even when other details of the history before gene flow are unknown. Applying this procedure to the genotypes called in the 1000 Genomes Project data, we estimate the drift since admixture in Europeans of about 0.065 and about 0.105 in East Asians. These estimates are quite close to those in the European and East Asian populations since they diverged, implying that the Neandertal gene flow occurred close to the time of split of the ancestral populations.
2.Assuming only one Neandertal gene flow event in the common ancestry of Europeans and East Asians, we estimate the drift since gene flow in the common ancestral population. We show that an upper bound on this shared drift is 0.018. Because this is far less than the drift associated with the out-of-Africa bottleneck of all non-African populations, this shows that the Neandertal gene flow occurred after the out-of-Africa bottleneck.
3.We use the genetic drift shared between Europeans and East Asians, in conjunction with the observation of large regions deficient in Neandertal ancestry obtained from a map of Neandertal ancestry in Eurasians, to estimate the number of generations and effective population size in the period immediately after gene flow. These analyses suggest that only a few dozen Neandertals may have contributed to the majority of Neandertal ancestry in non-Africans today.
Genetic characterisation of two Greek population isolates.
K. Hatzikotoulas et al.
Genetic association studies of low-frequency and rare variants can be empowered by focusing on isolated populations. It is important to genetically characterize population isolates for substructure and recent admixture events as these may give rise to spurious associations. Under the auspices of the HELlenic Isolated Cohorts study (HELIC; www.helic.org) we have collected >3,000 samples from two isolated populations in Greece: the Pomak villages (HELIC Pomak), a set of religiously-isolated mountainous villages in the North of Greece; and Anogia and surrounding mountainous villages on Crete (HELIC MANOLIS). All samples have information on anthropometric, cardiometabolic, biochemical, haematological and diet-related traits. 1,500 individuals from each population isolate have been typed on the Illumina OmniExpress and Human Exome Beadchip platforms. Multidimensional scaling analysis with the 1000 Genomes Project data shows similarities of the two population isolates with Mediterranean populations such as the Tuscans from Italy and Iberians from Spain. We also observe evidence for structure within the isolates, with the Kentavros village in the Pomak strand demonstrating high levels of differentiation. To characterise the degree of isolatedness in these populations we estimated the proportion of individuals with at least one “surrogate parent” (using only the subset of samples with pairwise pi-hat<0 .2="" 707="" adolescents="" an="" and="" at="" attica="" compared="" comprises="" district.="" find="" for="" from="" genome="" greek="" in="" individuals="" is="" isolate="" least="" manolis="" of="" one="" outbred="" parent="" population="" proportion="" random="" regions="" study="" surrogate="" teenage="" that="" the="" this="" to="" unrelated="" we="" which="" with="">60% and in the Pomak isolate is >65% compared to ~1% in the outbred Greek population. Our results establish these populations as isolates and provide some insights into the genomic architecture of Greek populations, which have not been previously characterised.0>Efficient and Accurate Whole-Genome Human Phasing.
T. Blauwkamp et al.
High throughput DNA sequencing allows whole human genomes to be resequenced rapidly and inexpensively producing a comprehensive list of variants relative to the reference genome. However, short read sequencing technologies are limited in their ability to determine phasing information, thus resulting in heterozygous calls being represented as the average of the maternal and paternal chromosomes. Phasing information is of critical importance to personal medicine as it provides a better linkage between genotype and phenotype, permitting new advances in our understanding of compound heterozygote linked diseases, pharmacogenomics, HLA typing, and prenatal genome sequencing. Here, we describe a new sample prep method that enables whole human genome haplotyping at high accuracy using only 30Gb of sequence data. Genomic DNA was fragmented into ~10Kb fragments, end repaired, and ligated to adapters. Hundreds of aliquots with approximately 50MB of DNA in each were amplified, fragmented and converted into individual shotgun libraries. The pooled libraries were sequenced in a single lane of a HiSeq2500 at 2x100bp to generate ~30Gb of sequence. The resulting sequence information was analyzed to obtain a set of long blocks of ~10Kb, covering multiple heterozygous SNPs, allowing phasing of these SNPs relative to each other. An HMM-based phasing algorithm was used to compute the most likely phase and confidence intervals based on the observed coverage and sequencer quality scores. Phasing of those blocks relative to each other was done by another HMM-based algorithm which uses a panel of previously phased genomes. Comparing our results with phase information inferred by transmission from the parents, we found that over 98% of heterozygous SNPs were phased within long blocks (N50=500kb) at a switch error rate below 1 switch per megabase of phased sequence. We present results obtained from multiple cell lines and human samples. This new library prep method and data analysis pipeline enables whole human genome phasing with only 30Gb of raw sequence, which represents only ~30% more sequencing than current 30x baseline run for human sequencing. Compared to other published reports, this method is capable of phasing a greater fraction of SNPS with ~75% less sequencing. Coupling our higher percentage of SNPs phased with high accuracy and the lowest sequencing requirement, this new technology is the most affordable approach to generating completely phased whole human genomes.Inference of Natural Selection and Demographic History for African Pygmy Hunter-Gatherers.
P. H. Hsieh et al.
African Pygmies are hunter-gatherers primarily inhabiting the Central African rainforests, where they are exposed to high temperatures, high humidity, and a pathogen and parasite-enriched woody habitat. These factors undoubtedly influenced their evolutionary history as they adapted to this environment. Many Pygmy populations have historically been in socio-economic contact with neighboring Niger-Kordofanian speaking farmer populations, particularly since the agriculture expansion in sub-Saharan Africa that began five thousand years ago (kya). To look for the true signatures of adaptation to the rainforest habitat of pygmies we must control for this complex demographic history. We sequenced and combined 40x whole genome sequence data from 3 Baka pygmies from Cameroon, 4 Biaka pygmies from the Central African Republic, and 9 Niger-Kordofanian speaking Yoruba farmers from Nigeria. We used ?a?i, a model-based demographic inference tool, to infer the history of these populations. Our best-fit model suggests that the ancestors of the farmer and pygmy populations diverged 150 kya and remained isolated from each other until 40 kya. This divergence is more ancient than estimated by previous studies that included fewer loci, but is consistent with a PSMC analysis, a separate inference tool that uses different aspects of the genomic data than ?a?i. Interestingly, our analysis shows that models with bi-directional asymmetric gene flow between farmers and pygmies are statistically better supported than previously suggested models with a single wave of uni-directional migration from farmers to pygmies. To identify possible targets of positive selection, we conducted a genomic scan using complementary methods, including the frequency-spectrum based G2D test, the population differentiation based XP-CLR test, and the haplotype based iHS test. We performed 10,000 simulations based on the above best-fit demographic model in order to assign statistical significance to each reported target of natural selection. Our results reveal that genes involved in cell adhesion, cellular signaling, olfactory perception, and immunity were likely targeted by natural selection in the pygmies or their recent ancestors. Our analysis also shows that genes involved in the function of lipid binding are enriched in highly differentiated non-synonymous mutations, suggesting that this function may have acted differently on the Pygmies and farmers after their divergence from their common ancestor.Population demography and maternal history of Oceania.
A. T. Duggan et al.
We present a large-scale study of mtDNA diversity across Near and Remote Oceania with whole-genome mtDNA sequencing and a sample collection of more than 1,300 individuals spanning from the Bismarck Archipelago in the west to the Cook Islands in the east. As the location of at least two major migration events (initial colonization over 40,000 years ago, followed by an expansion of Austronesian-speaking migrants around 3,500 years ago), Oceania provides a unique opportunity to study the effects of population admixture. Our results support the idea of sex-biased admixture between the resident populations and the migrants of the Austronesian expansion. We find that haplogroups of putative Asian origin which are thought to have spread with the Austronesian expansion are found at high frequency in all but two populations and, in general, we see little evidence of distinction between Papuan and Austronesian speaking populations. Santa Cruz, which is part of the Solomon Islands but geographically distinct from the main island chain and considered part of Remote Oceania, has long been considered a linguistic oddity and is now accepted to represent a very deep branch in the Oceanic language family. We find that it is also a genetic outlier, with potential direct connections to the Bismarck Archipelago not evident in the main Solomon Islands chain. In this expanded dataset, we find additional evidence of instability and increased heteroplasmy at the ‘Polynesian motif’ position 16247, further confirming previous findings restricted to the Solomon Islands.
Reconstructing Austronesian population history.
M. Lipson et al.
Present-day populations that speak Austronesian languages are spread across half the globe, from Easter Island in the Pacific Ocean to Madagascar in the Indian Ocean. Evidence from linguistics and archaeology suggests that the "Austronesian expansion," a vast cultural and linguistic dispersal that began 4--5 thousand years ago, had its origin in Taiwan. However, genetic studies of Austronesian ancestry have been inconclusive, with some finding affinities with aboriginal Taiwanese, others advancing an autochthonous origin within Island Southeast Asia, and others proposing a model involving multiple waves of migration from Asia. Here, we analyze genome-wide data from a diverse set of 31 Austronesian-speaking and 25 other groups typed at 18,412 overlapping single nucleotide polymorphisms (SNPs) to trace the genetic origins of Austronesians. We use a recently developed computational tool for building phylogenetic models of population relationships incorporating the possibility of admixture, which allows us to infer ancestry proportions and sources of genetic material for 26 admixed Austronesian-speaking populations. Our analysis provides strong confirmation of widespread ancestry of Taiwanese origin: at least a quarter of the genetic material in all Austronesian-speaking populations that we studied---including all of the Asian ancestry in populations from eastern Indonesia and Oceania---is more closely related to aboriginal Taiwanese than to any populations we sampled from the mainland. Surprisingly, we also show that western Austronesian-speaking populations have inherited substantial proportions of their Asian ancestry from a source that falls within the variation of present-day Austro-Asiatic populations in Southeast Asia. No Austro-Asiatic languages are spoken in Island Southeast Asia today, although there are some linguistic and archaeological suggestions of an early connection between mainland and island populations. The most plausible explanation for these findings, in light of the historical evidence, is that western Island Southeast Asia was settled by Austronesian groups who had previously mixed with Austro-Asiatic speakers on the mainland.No significant differences in the accumulation of deleterious mutations across diverse human populations.
R. Do et al.
Differences in demographic history across populations are expected to cause differences in the accumulation of deleterious mutations because natural selection works less efficiently when population sizes are small. Surprisingly, however, the relative burden of deleterious mutations has never been directly measured across human populations on a per-haploid genome basis, despite the fact that this is what matters biologically in the absence of dominance and epistasis. Here we empirically measure the relative accumulation of deleterious mutations in 13 diverse populations (Yoruba, Mandenka, San, Mbuti, Dinka, Australian, French, Sardinian, Han, Dai, Mixe, Karitiana and Papuan) along with one archaic population (Denisova). All the present-day populations have statistically indistinguishable accumulations of coding mutations. We highlight two examples. First, we find no evidence for a lower mutational load in West Africans than in Europeans despite the approximately 30% higher genetic diversity in West Africans: the accumulation of nonsynonymous mutations in West Africans is 1.01±0.02 times that in Europeans, and for “probably damaging” mutations, the ratio is 1.03±0.04. Second, we find no evidence for a lower mutational load in populations that have experienced agriculture-related expansions over the last 10,000 years and those that have not: the ratio in Chinese to Karitiana hunter gatherers from Brazil is 0.99±0.07. We determined that these null results are not an artifact of insensitivity of our method to differences in demographic history. As a positive control, we also analyzed archaic Denisovans who are known to have had a small population size for hundreds of thousands of years since separation from modern humans. We show that the Denisovan lineage has accumulated “probably damaging” mutations 1.33±0.06 times more rapidly than modern humans since they split. These analyses are important because of the new constraints they place on the distribution of selection coefficients in humans. Given the currently estimated demographic histories of West Africans and Europeans, combined with the fact that we do not detect a lower accumulation of deleterious mutations in West Africans than Europeans, we can conclude that only a small proportion of nonsynonymous mutations have selection coefficients in the range s=-0.01 to -0.001, which is the range of selection coefficients which would be expected to show a lower accumulation in West Africans than in Africans.Deep coverage Bedouin genomes reveal Bedouin haplotypes shared among worldwide populations in the 1000 Genomes Project.
J. L. Rodriguez-Flores et al.
The 1000 Genomes Project (1000G) has sampled and sequenced over 2500 genomes that are representative of the genetic diversity in populations worldwide. The Arabian Peninsula has not been previously included in 1000G, hence the connections between genetic variation in the indigenous Bedouin people and worldwide populations is unknown. We have sampled genomes from Bedouin individuals in the nation of Qatar as a window into the genetic variation in this understudied region. Our goal was to use this sample to assess the hypothesis that there is detectable shared ancestry between Bedouin and Southern European populations resulting from the history of empires that spanned both the Mediterranean and Arabian regions and the hypothesis that there is shared ancestry between Bedouin and contemporary Latin American populations, since the majority of European settlers in Latin America from the past half millennia are primarily from Southern European countries. We selected 60 Qataris with over 95% Bedouin ancestry and at least 3 generations of ancestry in Qatar for deep coverage genome sequencing. Genomes were sequenced by the Illumina Genome Network using TruSeq DNA PCR-free sample preparation, generating over 120 gigabases of paired-end 100 base pair reads per genome on a HiSeq 2500, yielding over 30x depth and genotypes for >96% of the genome using both the ELAND/CASAVA and BWA/GATK pipelines. Using these genotypes, we inferred haplotypes using SHAPEIT for Bedouin Qataris and for 1000G populations on a set of sites polymorphic in both 1000G and Bedouins. We used admixture analysis to assess shared ancestry between our Bedouin sample and 1000G populations using the ancestry deconvolution method SUPPORTMIX. Given the lack of appropriate ancestral populations, we conducted a leave-one-out approach, where for each population (1000G + Bedouin = n), we removed the population and used the remaining n-1 populations as an ancestral reference panel. Using this approach, we observed up to 15% Bedouin ancestry in European, South Asian, and American populations. Likewise, we observed ancestry from Europe, South Asia, and America in the Bedouin. For individuals from the Americas, the analysis identified a considerable number of segments shared with Bedouins previously classified as European ancestry.Using a haplotype-based model to infer Native American colonization history.
C. Lewis et al.
We apply a powerful haplotype-based model (described in Lawson et al. 2012) to infer the population history of 410 individuals from ~50 Native American groups, using data interrogated at >470,000 genome-wide autosomal Single-Nucleotide-Polymorphisms (SNPs). The model matches haplotype patterns among individuals' chromosomes to infer which individuals share recent common ancestry at each location of the genome, an approach that has previously been demonstrated to increase power substantially over widely-used alternative approaches that consider SNPs independently. We apply this methodology to 1861 samples described in Reich et al. (2012), incorporating 263 additional samples from 32 relevant world-wide regions collated from other publicly available resources and currently unavailable data. We utilize these methodology and data in two ways. First, we infer intermixing (i.e. "admixture") events among different Native American groups by identifying the groups that share the most haplotype segments. Using additional unpublished techniques, we determine the dates of these intermixing events, the proportions of DNA contributed, and the precise genetic make-up of the groups involved. These unique characteristics set this methodology apart from all presently available software, allowing us to place these mixing events into a clear historical context and thus identify the factors (e.g. the rise or fall of various Native American empires) that have contributed most to the genetic architecture of present-day Native American groups. Second, we match DNA patterns from each Native American group to a set of over 30 populations from Siberia and East Asia, describing each Native American group as a mixture of DNA from these regions. This enables us to shed light on the widely debated number of distinct migrations into the Americas during the initial colonization across the Bering Strait, comparing our results to previous inference from the literature. Our application demonstrates the power gained by using rich haplotype information relative to approaches that ignore this information.
Using Ancient Genomes to Detect Positive Selection on the Human Lineage.
K. Prüfer et al.
At least two distinct groups of archaic hominins inhabited Eurasia before the arrival of modern humans: Neandertals and Denisovans. The analysis of the genomes of these archaic humans revealed that they are more closely related to one another than they are to modern humans. However, since modern and archaic humans are so closely related, only about 10% of the archaic DNA sequences fall outside the present-day human variation whereas for 90% of the genome, Neandertal or Denisova DNA sequences are more closely related to some humans than to others. The fact that the archaic sequence often falls within the diversity of modern humans can be used to detect selective sweeps that affected all modern humans after their split from archaic humans since such sweeps will result in genomic regions where both the Neandertal and Denisova genomes fall outside the modern human variation. The genetic lengths of such external regions are proportional to the strength of selection, since stronger selection will lead to faster sweeps allowing less time for recombination to decrease their size. We have implemented a test for such external regions as a hidden Markov model. At each polymorphic position the model emits ancestral or derived based on whether the tested archaic genome carries the ancestral or derived variant of SNPs observed in present-day humans. The model was applied to 185 African genomes from the 1000 genomes phase 1 data. We identified thousands of external regions using the Neandertal and Denisova genomes, separately. Approximately one third of the regions are overlapping between the two genomes. These regions are significantly longer than regions only identified in only one of the archaic genomes. Based on this excess of overlap for long regions, we devise a measure to identify a set of regions that are candidates for selective sweeps on the human lineage since the split from Neandertal and Denisova.
Pulling out the 1%: Whole-Genome In-Solution (WISC) capture for the targeted enrichment of ancient DNA sequencing libraries.
C. D. Bustamante et al.
The very low levels of endogenous DNA remaining in most ancient specimens has precluded the shotgun sequencing of many interesting samples due to cost. For example, ancient DNA (aDNA) libraries derived from bones and teeth often contain <1 b="" by="" capacity="" dna.="" dna="" endogenous="" environmental="" is="" majority="" meaning="" of="" sequencing="" taken="" that="" the="" up=""> We will present a method for the targeted enrichment of the endogenous component of human aDNA sequencing libraries. Using biotinylated RNA baits transcribed from genomic DNA libraries, we are able to significantly enrich for human-derived DNA fragments. This approach, which we call whole-genome in-solution capture (WISC), allows us to obtain genome-wide ancestral information from ancient samples with very low endogenous DNA contents. We demonstrate WISC on libraries created from four Iron Age and Bronze Age human teeth from Bulgaria, as well as bone samples from seven Peruvian mummies and a Bronze Age hair sample from Denmark. Prior to capture, shotgun sequencing of these libraries yielded an average of 1.2% of reads mapping to the human genome (including duplicates). After capture, this fraction increased dramatically, with up to 59% of reads mapped to human and folds enrichment ranging from 5X to 139X. Furthermore, we maintained coverage of the majority of fragments present in the pre-capture library. Intersection with the 1000 Genomes Project reference panel yielded an average of 50,723 SNPs (range 3,062-147,243) for the post-capture libraries sequenced with 1 million reads, compared with 13,280 SNPs (range 217-73,266) for the pre-capture libraries, increasing resolution in population genetic analyses. We will also present the results of performing WISC on other aDNA libraries from both archaic human and non-human samples, including ancient domestic dog samples. Our capture approach is flexible and cost-effective, allowing researchers to access aDNA from many specimens that were previously unsuitable for sequencing. Furthermore, this method has applications in other contexts, such as the enrichment of target human DNA in forensic samples.1>
Insights into population history from a high coverage Neandertal genome.
D. Reich1, for.the. Neandertal Genome Consortium2
We have sequenced to about 50-fold coverage a genome sequence from about 40 mg of a bone found in Denisova Cave in Southern Siberia. The genome of this female is much more closely related to the low-coverage Neandertal genomes from Croatia, Spain, Germany and the Caucasus than to the genome of archaic Denisovans, a sister group of Neandertals, and provides unambiguous evidence that both Neandertals and Denisovans inhabited the Altai Mountains in Siberia. The high-coverage Neandertal genome, combined with our earlier sequencing of a high quality Denisova genome, allows novel insights about the population history of archaic humans:
•We document recent inbreeding in this Altai Neandertal. The inbreeding coefficient of about 1/8 corresponds to about the homozygosity that would be expected from a mating of half siblings.
•The Altai Neandertal genome shares almost seven percent more derived alleles with present-day Africans than does the Denisova genome. This means that the Denisovans derived a proportion of their ancestry from a very archaic human lineage, and the amount of this ancestry they inherit is larger than in Neandertals.
• The Denisovan genome is affected by major recent gene flow from an Altai-related Neandertal.
• To further characterize the variation among Neandertals we sequenced the genome of a Neandertal from the Caucasus to about 0.5-fold coverage. Comparisons to present-day genomes show that the Neandertals who contributed genes to present-day non-Africans were more closely related to this Caucasian Neandertal than to the Neandertals we sequenced from the Altai.
•We built a map of Neandertal ancestry in modern humans, using data from all non-Africans in the 1000 Genomes Project. We show that the average Neandertal ancestry on chromosome X of present-day non-Africans is about a fifth of the genome average. It is known that hybrid incompatibility loci concentrate on chromosome X. Thus, this observation is consistent with a model of hybrid incompatibility in which Neandertal variants that introgressed into modern humans were rapidly selected away due to epistatic interactions with the modern human genetic background.
Inferring complex demographies from PSMC coalescent rate estimates: African substructure and the Out-of-Africa event.
S. Gopalakrishnan et al.
S. Gopalakrishnan et al.
Human population history is an intriguing and complex story with many events like population growth, bottlenecks, time-dependent/non-homogeneous migration, population splits and mixtures. Estimating complete demographies with population sizes, rates of gene flow and population split times has proven to be a challenging endeavor. We propose a framework for jointly estimating the demography parameters, especially gene-flow rates and split times, for a large number of populations. We use coalescent rate estimates obtained from Pairwise Sequentially Markovian Coalescent (PSMC) as the starting point for our analysis. Since PSMC works on only two chromosomes at a time, we apply PSMC to all pairs of individuals to obtain the pairwise coalescent rates for lineages from every pair of sampled populations. Using a mathematical model for calculating coalescent probabilites given population parameters, we estimate demography using the parameters that best fit the observed coalesecent rates.
In this study, we focus on two aspects of African population genetics, 1. the nature of population structure in Africa going back in time and 2. the timing of the Out-of-Africa event. To address these questions, we assembled a dataset with whole genome sequences from 162 individuals using both in-house sequencing and publicly available sources. These samples span 22 populations worldwide. These include eleven African populations which we use to dissect the population substructure in Africa. In addition, we also have 2 Middle Eastern, 5 European and 4 East/Central Asian populations which inform the population split time estimates for the Out-of-Africa event and the European-Asian split.
We find extensive population structure in Africa extending back to before the Out-of-Africa event. The Ethiopian populations, Amhara and Oromo, show evidence of mixing beyond 15 kya. The Maasai and Luhye merge with the Ethiopian populations to form a panmictic East African population ~40kya. We find evidence for extensive mixing between east and west African populations before 50kya. Among the pygmy populations, we see recent gene flow between the Batwa and Mbuti. All African populations except the San merge into a single population around 110 kya. The San exchange migrants with the other African populations beginning ~120 kya. We estimate the Out-of-Africa event to have occurred ~75kya and the European-Asian split to ~25kya.
Out of Africa, which way?
L. Pagani et al.
E. Elhaik et al.
L. Pagani et al.
While the African origin of all modern human populations is well-established, the dynamics of the diaspora that led anatomically modern humans to colonize the lands outside Africa are still under debate. Understanding the demographic parameters as well as the route (or routes) followed by the ancestors of all non-Africans could help to refine our understanding of the selection processes that occurred subsequently, as well as shedding light on a landmark process in our evolutionary history. Of the three possible gateways out of Africa (via Morocco across the Gibraltar strait, via Egypt through the Suez isthmus or via the Horn of Africa across Bab el Mandeb strait) only the latter two are supported by paleoclimatic and archaeological evidence. Furthermore, recent studies (Pagani et al. 2012) showed that, although the modern Ethiopian populations might be good candidates for the descendants of the source population of such a migration, modern Egyptians could be an even better candidate. Unfortunately, however, only a few Egyptian samples have been genotyped and, as yet, none have been fully sequenced. Here, we have generated 125 Ethiopian and 100 Egyptian whole genome sequences (Illumina HiSeq, 8x average depth). The genomes were partitioned using PCAdmix (Brisbin et al. 2012) to account for the confounding effects of recent introgression from neighboring non-African populations. To explore the genetic legacy of migration routes out of Africa, and in particular to test whether the observed genetic data support one route over another, the African components of Egyptians and Ethiopians were then compared to a panel of available non-African populations from the 1000 Genomes Project (1000 Genomes Project Consortium, 2012). The high resolution provided by whole genome sequencing allows us to shed new light on the paths followed by our ancestors as they left Africa, as well as refining the current knowledge of the demographic history of the populations analyzed.
The Saudi Arabian Genome Reveals a Two Step Out-of-Africa Migration.
J. J. Farrell et al.
Here we present the first high-coverage whole genome sequences from a Middle Eastern population consisting of 14 Eastern Province Saudi Arabians. Genomes from this region are of interest to further answer questions regarding “Out-of-Africa” human migration. Applying a pairwise sequentially Markovian coalescent model (PSMC), we inferred the history of population sizes between 10,000 years and 1,000,000 years before present (YBP) for the Saudi genomes and an additional 11 high-coverage whole genome sequences from Africa, Asia and Europe.Geographic Population Structure (GPS) of worldwide human populations infers biogeographical origin down to home village
The model estimated the initial separation from Africans at approximately 110,000 YBP. This intermediate population then underwent a long period of decreasing population size culminating in a bottleneck 50,000 YBP followed by an expansion into Asia and Europe. The split and subsequent bottleneck were thus two distinct events separated by a long intermediate period of genetic drift in the Middle East. The two most frequent mitochondria haplogroups (30% each) were the Middle Eastern U7a and the African L. The presence of the L haplogroup common in Africa was unexpected given the clustering of the Saudis with Europeans in the phylogenetic tree and suggests some recent African admixture. To examine this further, we performed formal tests for a history of admixture and found no evidence of African admixture in the Saudi after the split. Taken together, these analyses suggest that the L3 haplogroup found in the Saudi were present before the bottleneck 50,000 YBP. Given the TMRCA estimates for the L3 haplogroup of approximately 70,000 YBP and the timing of the Out-of-Africa split, these analyses suggest that L3 haplogroup arose in the Middle East with a subsequent back migration and expansion into Africa over the Horn-of-Africa during the lower sea levels found during the glacial period bottleneck.
These results are consistent with the hypothesis that modern humans populated the Middle East before a split 110,000 YBP, underwent genetic drift for 60,000 years before expanding to Asia and Europe as well as back-migration into Africa. Examination of genetic variants discovered by Saudi whole genome sequencing in ancestral African populations and European/Asian populations will contribute to the understanding human migration patterns and the origin of genetic variation in modern humans.
E. Elhaik et al.
The search for a method that utilizes biological information to predict human’s place of origin has occupied scientists for millennia. Modern biogeography methods are accurate to 700 km in Europe but are highly inaccurate elsewhere, particularly in Southeast Asia and Oceania. The accuracy of these methods is bound by the choice of genotyping arrays, the size and quality of the reference dataset, and principal component (PC)-based algorithms. To overcome the first two obstacles, we designed GenoChip, a dedicated genotyping array for genetic anthropology with an unprecedented number of ~12,000 Y-chromosomal and ~3,300 mtDNA SNPs and over 130,000 autosomal and X-chromosomal SNPs carefully chosen to study ancestry without any known health, medical, or phenotypic relevance. We also 615 individuals from 54 worldwide populations collected as part of the Genographic Project and the 1000 Genomes Project. To overcome the last impediment, we developed an admixture-based Geographic Population Structure (GPS) method that infers the biogeography of worldwide individuals down to their village of origin. GPS’s accuracy was demonstrated on three data sets: worldwide populations, Southeast Asians and Oceanians, and Sardinians (Italy) using 40,000-130,000 GenoChip markers. GPS correctly placed 80%; of worldwide individuals within their country of origin with an accuracy of 87%; for Asians and Oceanians. Applied to over 200 Sardinians villagers of both sexes, GPS placed a quarter of them within their villages and most of the remaining within 50 km of their villages, allowing us to identify the demographic processes that shaped the Sardinian society. These findings are significantly more accurate than PCA-based approaches. We further demonstrate two GPS applications in tracing the poorly understood biogeographical origin of the Druze and North American (CEU) populations. Our findings demonstrate the potential of the GenoChip array for genetic anthropology. Moreover, the accuracy and power of GPS underscore the promise of admixture-based methods to biogeography and has important ramifications for genetic ancestry testing, forensic and medical sciences, and genetic privacy.
Subscribe to:
Posts (Atom)