PLoS ONE 9(8): e102645. doi:10.1371/journal.pone.0102645
The South Asian Genome
John C. Chambers et al.
The genetic sequence variation of people from the Indian subcontinent who comprise one-quarter of the world's population, is not well described. We carried out whole genome sequencing of 168 South Asians, along with whole-exome sequencing of 147 South Asians to provide deeper characterisation of coding regions. We identify 12,962,155 autosomal sequence variants, including 2,946,861 new SNPs and 312,738 novel indels. This catalogue of SNPs and indels amongst South Asians provides the first comprehensive map of genetic variation in this major human population, and reveals evidence for selective pressures on genes involved in skin biology, metabolism, infection and immunity. Our results will accelerate the search for the genetic variants underlying susceptibility to disorders such as type-2 diabetes and cardiovascular disease which are highly prevalent amongst South Asians.
Link
Showing posts with label Genomics. Show all posts
Showing posts with label Genomics. Show all posts
August 12, 2014
July 29, 2014
Lethal mutations quantified
A very interesting new preprint on the arXiv (so it can be freely read). The founder population is the Hutterites. The key sentence:
arXiv:1407.7518 [q-bio.PE]
An estimate of the average number of recessive lethal mutations carried by humans
Ziyue Gao, Darrel Waggoner, Matthew Stephens, Carole Ober, Molly Przeworski
The effects of inbreeding on human health depend critically on the number and severity of recessive, deleterious mutations carried by individuals. In humans, existing estimates of these quantities are based on comparisons between consanguineous and non-consanguineous couples, an approach that confounds socioeconomic and genetic effects of inbreeding. To circumvent this limitation, we focused on a founder population with almost complete Mendelian disease ascertainment and a known pedigree. By considering all recessive lethal diseases reported in the pedigree and simulating allele transmissions, we estimated that each haploid set of human autosomes carries on average 0.29 (95% credible interval [0.10, 0.83]) autosomal, recessive alleles that lead to complete sterility or severe disorders at birth or before reproductive age when homozygous. Comparison to existing estimates of the deleterious effects of all recessive alleles suggests that a substantial fraction of the burden of autosomal, recessive variants is due to single mutations that lead to death between birth and reproductive age. In turn, the comparison to estimates from other eukaryotes points to a surprising constancy of the average number of recessive lethal mutations across organisms with markedly different genome sizes.
Link
Our approach indicates that on average, one in every two humans carries a recessive lethal allele on the autosomes that lead to lethality after birth and before reproductive age or to complete sterility.
arXiv:1407.7518 [q-bio.PE]
An estimate of the average number of recessive lethal mutations carried by humans
Ziyue Gao, Darrel Waggoner, Matthew Stephens, Carole Ober, Molly Przeworski
The effects of inbreeding on human health depend critically on the number and severity of recessive, deleterious mutations carried by individuals. In humans, existing estimates of these quantities are based on comparisons between consanguineous and non-consanguineous couples, an approach that confounds socioeconomic and genetic effects of inbreeding. To circumvent this limitation, we focused on a founder population with almost complete Mendelian disease ascertainment and a known pedigree. By considering all recessive lethal diseases reported in the pedigree and simulating allele transmissions, we estimated that each haploid set of human autosomes carries on average 0.29 (95% credible interval [0.10, 0.83]) autosomal, recessive alleles that lead to complete sterility or severe disorders at birth or before reproductive age when homozygous. Comparison to existing estimates of the deleterious effects of all recessive alleles suggests that a substantial fraction of the burden of autosomal, recessive variants is due to single mutations that lead to death between birth and reproductive age. In turn, the comparison to estimates from other eukaryotes points to a surprising constancy of the average number of recessive lethal mutations across organisms with markedly different genome sizes.
Link
June 15, 2014
Chimp mutation rate is equal to human mutation rate but driven more by males
This is important because (a) it shows evidence for the "slow" mutation rate in a species related to humans, (b) it shows that chimp and human mutation rates are equal and so using the human mutation rate in studies of divergence with chimps is justified, and (c) it is driven differently by males/females than in humans.
Science 13 June 2014: Vol. 344 no. 6189 pp. 1272-1275
DOI: 10.1126/science.344.6189.1272
Strong male bias drives germline mutation in chimpanzees
Oliver Venn
ABSTRACT
Germline mutation determines rates of molecular evolution, genetic diversity, and fitness load. In humans, the average point mutation rate is 1.2 × 10−8 per base pair per generation, with every additional year of father’s age contributing two mutations across the genome and males contributing three to four times as many mutations as females. To assess whether such patterns are shared with our closest living relatives, we sequenced the genomes of a nine-member pedigree of Western chimpanzees, Pan troglodytes verus. Our results indicate a mutation rate of 1.2 × 10−8 per base pair per generation, but a male contribution seven to eight times that of females and a paternal age effect of three mutations per year of father’s age. Thus, mutation rates and patterns differ between closely related species.
Link
Science 13 June 2014: Vol. 344 no. 6189 pp. 1272-1275
DOI: 10.1126/science.344.6189.1272
Strong male bias drives germline mutation in chimpanzees
Oliver Venn
ABSTRACT
Germline mutation determines rates of molecular evolution, genetic diversity, and fitness load. In humans, the average point mutation rate is 1.2 × 10−8 per base pair per generation, with every additional year of father’s age contributing two mutations across the genome and males contributing three to four times as many mutations as females. To assess whether such patterns are shared with our closest living relatives, we sequenced the genomes of a nine-member pedigree of Western chimpanzees, Pan troglodytes verus. Our results indicate a mutation rate of 1.2 × 10−8 per base pair per generation, but a male contribution seven to eight times that of females and a paternal age effect of three mutations per year of father’s age. Thus, mutation rates and patterns differ between closely related species.
Link
February 12, 2014
Ancient Clovis genome from Montana yields no surprises (Rasmussen et al. 2014)
Ancient DNA has consistently managed to surprise us, with pretty much no direct genetic continuity revealed between Pleistocene and modern populations anywhere in the world. So, it is refreshing to see that at least in the case of the Americas the people who lived there ~13 thousand years ago are clearly related to the people who lived there in pre-Columbian times, with no real evidence of subsequent gene flows from Eurasia (at least in the case of Central/South Americans).
Many people suspected this because of the difficulty to access the Americas from Eurasia: this must have limited gene flow between the two regions to a handful of migrants and a restricted set of time periods where geological and climatic conditions were advantageous. The much reduced genetic diversity of Native Americans also argues in favor of them being a relatively simple population, with low heterozygosity and a handful of unique "founder lineages" in both the Y-chromosome and mtDNA.
Nonetheless, there are also several theories in the realm of alernative history, involving Solutreans from Europe, trans-Pacific boat riders, bearded "White Gods", Minoans/Phoenicians/Atlanteans/Ancient Egyptians, "African" Olmecs, "Caucasoid" Paleo-Indians, lost Israelite tribes, to mention only a few of the most well-known ones.
The new study does not, of course, disprove any of the proposals in the preceding paragraph: one can still claim that diverse groups once inhabited the Americas and Rasmussen et al. (2014) just happened to chance upon one that looked just like modern native Americans. But, this certainly improves the odds of early "Native American simplicity", offering no evidence for the complexity postulated by many of the alternative theories.
Moreover, while the existence of other human groups in the Americas cannot be disproved by the study of a single ancient individual, what can be proved is the antiquity of the ancestors of Native Americans. Rather than being late arrivals arriving from Asia after the initial colonization, perhaps with derived Mongoloid physical morphology, we now know that they were already there as early as ~13 thousand years ago. It is remarkable that a single ancient DNA sample can sweep away much of the nonsense that has been written on the topic in the past.
A piece in Nature News addresses some of the "ethics" debate that seems ever-present in studies involving Native American remains. I don't know how this study will be perceived by living Native Americans: a possibility is that they'll be more receptive to ancient DNA research now that a team of scientists have stretched the time depth of their ancestry in the Americas to the earliest studied sample, revealing themselves not to be the evil-doers that western scientists are generally assumed to be according to a certain kind of mentality. A different -and more alarming- possibility, is that radical anti-science elements will be emboldened by these findings to claim that continuity with the earliest Americans (which in itself seems true enough) adds support to claims of ownership to pretty much all archaeological samples whose relationship to living Amerindians was hitherto uncertain in light of the many alternative theories.
In any case, it is remarkable that this ~13 thousand year old genome now exists while the genomes of modern native Americans that can be had for a fraction of the cost and technical difficulty do not. Indeed, not even genotype data exist from most Amerindian groups from the USA, which creates the rather bizarre state of affairs that the Anzick-1 genome had to be compared with native groups from several countries in the Western hemisphere except the one in which it was found.
Nature 506, 225–229 (13 February 2014) doi:10.1038/nature13025
The genome of a Late Pleistocene human from a Clovis burial site in western Montana
Morten Rasmussen et al.
Clovis, with its distinctive biface, blade and osseous technologies, is the oldest widespread archaeological complex defined in North America, dating from 11,100 to 10,700 14C years before present (BP) (13,000 to 12,600 calendar years BP)1, 2. Nearly 50 years of archaeological research point to the Clovis complex as having developed south of the North American ice sheets from an ancestral technology3. However, both the origins and the genetic legacy of the people who manufactured Clovis tools remain under debate. It is generally believed that these people ultimately derived from Asia and were directly related to contemporary Native Americans2. An alternative, Solutrean, hypothesis posits that the Clovis predecessors emigrated from southwestern Europe during the Last Glacial Maximum4. Here we report the genome sequence of a male infant (Anzick-1) recovered from the Anzick burial site in western Montana. The human bones date to 10,705 ± 35 14C years BP (approximately 12,707–12,556 calendar years BP) and were directly associated with Clovis tools. We sequenced the genome to an average depth of 14.4× and show that the gene flow from the Siberian Upper Palaeolithic Mal’ta population5 into Native American ancestors is also shared by the Anzick-1 individual and thus happened before 12,600 years BP. We also show that the Anzick-1 individual is more closely related to all indigenous American populations than to any other group. Our data are compatible with the hypothesis that Anzick-1 belonged to a population directly ancestral to many contemporary Native Americans. Finally, we find evidence of a deep divergence in Native American populations that predates the Anzick-1 individual.
Link
Many people suspected this because of the difficulty to access the Americas from Eurasia: this must have limited gene flow between the two regions to a handful of migrants and a restricted set of time periods where geological and climatic conditions were advantageous. The much reduced genetic diversity of Native Americans also argues in favor of them being a relatively simple population, with low heterozygosity and a handful of unique "founder lineages" in both the Y-chromosome and mtDNA.
Nonetheless, there are also several theories in the realm of alernative history, involving Solutreans from Europe, trans-Pacific boat riders, bearded "White Gods", Minoans/Phoenicians/Atlanteans/Ancient Egyptians, "African" Olmecs, "Caucasoid" Paleo-Indians, lost Israelite tribes, to mention only a few of the most well-known ones.
The new study does not, of course, disprove any of the proposals in the preceding paragraph: one can still claim that diverse groups once inhabited the Americas and Rasmussen et al. (2014) just happened to chance upon one that looked just like modern native Americans. But, this certainly improves the odds of early "Native American simplicity", offering no evidence for the complexity postulated by many of the alternative theories.
Moreover, while the existence of other human groups in the Americas cannot be disproved by the study of a single ancient individual, what can be proved is the antiquity of the ancestors of Native Americans. Rather than being late arrivals arriving from Asia after the initial colonization, perhaps with derived Mongoloid physical morphology, we now know that they were already there as early as ~13 thousand years ago. It is remarkable that a single ancient DNA sample can sweep away much of the nonsense that has been written on the topic in the past.
A piece in Nature News addresses some of the "ethics" debate that seems ever-present in studies involving Native American remains. I don't know how this study will be perceived by living Native Americans: a possibility is that they'll be more receptive to ancient DNA research now that a team of scientists have stretched the time depth of their ancestry in the Americas to the earliest studied sample, revealing themselves not to be the evil-doers that western scientists are generally assumed to be according to a certain kind of mentality. A different -and more alarming- possibility, is that radical anti-science elements will be emboldened by these findings to claim that continuity with the earliest Americans (which in itself seems true enough) adds support to claims of ownership to pretty much all archaeological samples whose relationship to living Amerindians was hitherto uncertain in light of the many alternative theories.
In any case, it is remarkable that this ~13 thousand year old genome now exists while the genomes of modern native Americans that can be had for a fraction of the cost and technical difficulty do not. Indeed, not even genotype data exist from most Amerindian groups from the USA, which creates the rather bizarre state of affairs that the Anzick-1 genome had to be compared with native groups from several countries in the Western hemisphere except the one in which it was found.
Nature 506, 225–229 (13 February 2014) doi:10.1038/nature13025
The genome of a Late Pleistocene human from a Clovis burial site in western Montana
Morten Rasmussen et al.
Clovis, with its distinctive biface, blade and osseous technologies, is the oldest widespread archaeological complex defined in North America, dating from 11,100 to 10,700 14C years before present (BP) (13,000 to 12,600 calendar years BP)1, 2. Nearly 50 years of archaeological research point to the Clovis complex as having developed south of the North American ice sheets from an ancestral technology3. However, both the origins and the genetic legacy of the people who manufactured Clovis tools remain under debate. It is generally believed that these people ultimately derived from Asia and were directly related to contemporary Native Americans2. An alternative, Solutrean, hypothesis posits that the Clovis predecessors emigrated from southwestern Europe during the Last Glacial Maximum4. Here we report the genome sequence of a male infant (Anzick-1) recovered from the Anzick burial site in western Montana. The human bones date to 10,705 ± 35 14C years BP (approximately 12,707–12,556 calendar years BP) and were directly associated with Clovis tools. We sequenced the genome to an average depth of 14.4× and show that the gene flow from the Siberian Upper Palaeolithic Mal’ta population5 into Native American ancestors is also shared by the Anzick-1 individual and thus happened before 12,600 years BP. We also show that the Anzick-1 individual is more closely related to all indigenous American populations than to any other group. Our data are compatible with the hypothesis that Anzick-1 belonged to a population directly ancestral to many contemporary Native Americans. Finally, we find evidence of a deep divergence in Native American populations that predates the Anzick-1 individual.
Link
December 27, 2013
Reconstructing Native American migrations
Of wider interest might be the authors' estimation of the autosomal mutation rate as 1.44x10-8 mutations/bp/generation. Of course, this might depend on the archaeological calibration used (where/when did the bottleneck in the ancestry of Native Americans occur?). It might also depend on recent evidence that Native Americans are of mixed origin and thus did not really split from CHB/JPT; only part of their ancestry did. Nonetheless, this is another fairly "low" autosomal mutation rate.
(This was previously released as a preprint to the arXiv).
PLoS Genet 9(12): e1004023. doi:10.1371/journal.pgen.1004023
Reconstructing Native American Migrations from Whole-Genome and Whole-Exome Data
Simon Gravel et al.
Link
(This was previously released as a preprint to the arXiv).
PLoS Genet 9(12): e1004023. doi:10.1371/journal.pgen.1004023
Reconstructing Native American Migrations from Whole-Genome and Whole-Exome Data
Simon Gravel et al.
Link
Site frequency spectrum from reads is unbiased (from genotype calls, biased at low coverage)
Mol Biol Evol (2013)
doi: 10.1093/molbev/mst229
Characterizing Bias in Population Genetic Inferences from Low-Coverage Sequencing Data
Eunjung Han et al.
The site frequency spectrum (SFS) is of primary interest in population genetic studies, because the SFS compresses variation data into a simple summary from which many population genetic inferences can proceed. However, inferring the SFS from sequencing data is challenging because genotype calls from sequencing data are often inaccurate due to high error rates and if not accounted for, this genotype uncertainty can lead to serious bias in downstream analysis based on the inferred SFS. Here, we compare two approaches to estimate the SFS from sequencing data: one approach infers individual genotypes from aligned sequencing reads and then estimates the SFS based on the inferred genotypes (call-based approach) and the other approach directly estimates the SFS from aligned sequencing reads by maximum likelihood (direct estimation approach). We find that the SFS estimated by the direct estimation approach is unbiased even at low coverage, whereas the SFS by the call-based approach becomes biased as coverage decreases. The direction of the bias in the call-based approach depends on the pipeline to infer genotypes. Estimating genotypes by pooling individuals in a sample (multisample calling) results in underestimation of the number of rare variants, whereas estimating genotypes in each individual and merging them later (single-sample calling) leads to overestimation of rare variants. We characterize the impact of these biases on downstream analyses, such as demographic parameter estimation and genome-wide selection scans. Our work highlights that depending on the pipeline used to infer the SFS, one can reach different conclusions in population genetic inference with the same data set. Thus, careful attention to the analysis pipeline and SFS estimation procedures is vital for population genetic inferences.
Link
Characterizing Bias in Population Genetic Inferences from Low-Coverage Sequencing Data
Eunjung Han et al.
The site frequency spectrum (SFS) is of primary interest in population genetic studies, because the SFS compresses variation data into a simple summary from which many population genetic inferences can proceed. However, inferring the SFS from sequencing data is challenging because genotype calls from sequencing data are often inaccurate due to high error rates and if not accounted for, this genotype uncertainty can lead to serious bias in downstream analysis based on the inferred SFS. Here, we compare two approaches to estimate the SFS from sequencing data: one approach infers individual genotypes from aligned sequencing reads and then estimates the SFS based on the inferred genotypes (call-based approach) and the other approach directly estimates the SFS from aligned sequencing reads by maximum likelihood (direct estimation approach). We find that the SFS estimated by the direct estimation approach is unbiased even at low coverage, whereas the SFS by the call-based approach becomes biased as coverage decreases. The direction of the bias in the call-based approach depends on the pipeline to infer genotypes. Estimating genotypes by pooling individuals in a sample (multisample calling) results in underestimation of the number of rare variants, whereas estimating genotypes in each individual and merging them later (single-sample calling) leads to overestimation of rare variants. We characterize the impact of these biases on downstream analyses, such as demographic parameter estimation and genome-wide selection scans. Our work highlights that depending on the pipeline used to infer the SFS, one can reach different conclusions in population genetic inference with the same data set. Thus, careful attention to the analysis pipeline and SFS estimation procedures is vital for population genetic inferences.
Link
June 20, 2013
Genetic load accumulation during range expansions
arXiv:1306.1652 [q-bio.PE]
On the accumulation of deleterious mutations during range expansions
Stephan Peischl et al.
We investigate the effect of spatial range expansions on the evolution of fitness when beneficial and deleterious mutations co-segregate. We perform individual-based simulations of a uniform linear habitat and complement them with analytical approximations for the evolution of mean fitness at the edge of the expansion. We find that deleterious mutations accumulate steadily on the wave front during range expansions, thus creating an expansion load. Reduced fitness due to the expansion load is not restricted to the wave front but occurs over a large proportion of newly colonized habitats. The expansion load can persist and represent a major fraction of the total mutation load thousands of generations after the expansion. Our results extend qualitatively and quantitatively to two-dimensional expansions. The phenomenon of expansion load may explain growing evidence that populations that have recently expanded, including humans, show an excess of deleterious mutations. To test the predictions of our model, we analyze patterns of neutral and non-neutral genetic diversity in humans and find an excellent fit between theory and data.
Link
On the accumulation of deleterious mutations during range expansions
Stephan Peischl et al.
We investigate the effect of spatial range expansions on the evolution of fitness when beneficial and deleterious mutations co-segregate. We perform individual-based simulations of a uniform linear habitat and complement them with analytical approximations for the evolution of mean fitness at the edge of the expansion. We find that deleterious mutations accumulate steadily on the wave front during range expansions, thus creating an expansion load. Reduced fitness due to the expansion load is not restricted to the wave front but occurs over a large proportion of newly colonized habitats. The expansion load can persist and represent a major fraction of the total mutation load thousands of generations after the expansion. Our results extend qualitatively and quantitatively to two-dimensional expansions. The phenomenon of expansion load may explain growing evidence that populations that have recently expanded, including humans, show an excess of deleterious mutations. To test the predictions of our model, we analyze patterns of neutral and non-neutral genetic diversity in humans and find an excellent fit between theory and data.
Link
June 19, 2013
Native American origins from whole-genome and exome data (Gravel et al. 2013)
arXiv:1306.4021 [q-bio.PE]
Reconstructing Native American Migrations from Whole-genome and Whole-exome Data
Simon Gravel et al.
There is great scientific and popular interest in understanding the genetic history of populations in the Americas. We wish to understand when different regions of the continent were inhabited, where settlers came from, and how current inhabitants relate genetically to earlier populations. Recent studies unraveled parts of the genetic history of the continent using genotyping arrays and uniparental markers. The 1000 Genomes Project provides a unique opportunity for improving our understanding of population genetic history by providing over a hundred sequenced low coverage genomes and exomes from Colombian (CLM), Mexican-American (MXL), and Puerto Rican (PUR) populations. Here, we explore the genomic contributions of African, European, and especially Native American ancestry to these populations. Estimated Native American ancestry is 48% in MXL, 25% in CLM, and 13% in PUR. Native American ancestry in PUR appears most closely related to Equatorial-Tucanoan-speaking populations, supporting a Southern America ancestry of the Taino people of the Caribbean. We present new methods to estimate the allele frequencies in the Native American fraction of the populations, and model their distribution using a three-population demographic model. The ancestral populations to the three groups likely split in close succession: the most likely scenario, based on a peopling of the Americas 16 thousand years ago (kya), supports that the MXL Ancestors split 12.2kya, with a subsequent split of the ancestors to CLM and PUR 11.7kya. The model also features a Mexican population of 62,000, a Colombian population of 8,700, and a Puerto Rican population of 1,900. Modeling Identity-by-descent (IBD) and ancestry tract length, we show that post-contact populations also differ markedly in their effective sizes and migration patterns, with Puerto Rico showing the smallest size and the earlier migration from Europe.
Link
Reconstructing Native American Migrations from Whole-genome and Whole-exome Data
Simon Gravel et al.
There is great scientific and popular interest in understanding the genetic history of populations in the Americas. We wish to understand when different regions of the continent were inhabited, where settlers came from, and how current inhabitants relate genetically to earlier populations. Recent studies unraveled parts of the genetic history of the continent using genotyping arrays and uniparental markers. The 1000 Genomes Project provides a unique opportunity for improving our understanding of population genetic history by providing over a hundred sequenced low coverage genomes and exomes from Colombian (CLM), Mexican-American (MXL), and Puerto Rican (PUR) populations. Here, we explore the genomic contributions of African, European, and especially Native American ancestry to these populations. Estimated Native American ancestry is 48% in MXL, 25% in CLM, and 13% in PUR. Native American ancestry in PUR appears most closely related to Equatorial-Tucanoan-speaking populations, supporting a Southern America ancestry of the Taino people of the Caribbean. We present new methods to estimate the allele frequencies in the Native American fraction of the populations, and model their distribution using a three-population demographic model. The ancestral populations to the three groups likely split in close succession: the most likely scenario, based on a peopling of the Americas 16 thousand years ago (kya), supports that the MXL Ancestors split 12.2kya, with a subsequent split of the ancestors to CLM and PUR 11.7kya. The model also features a Mexican population of 62,000, a Colombian population of 8,700, and a Puerto Rican population of 1,900. Modeling Identity-by-descent (IBD) and ancestry tract length, we show that post-contact populations also differ markedly in their effective sizes and migration patterns, with Puerto Rico showing the smallest size and the earlier migration from Europe.
Link
June 07, 2013
Demographic history from distribution of shared IBS lengths
An interesting new paper has appeared in PLoS Genetics, with what appears to be a nice new method for inferring demographic history from genome-scale data. The authors observe that segments inherited from common ancestors are "broken up" by mutation as time goes by: initially there are long identical tracts, but these are "split" whenever a new mutation appears, so they study the distribution of lengths of the pieces between mutations that remain identical by state.
A practical application of the new technique is applied to European-African history:
The authors estimate that perhaps a 1.75x increase in their estimates will be effected if the slower rates are used; this is not 2x as one might expect from a 2x slower rate, because their age estimates depend on both the mutation rate (for which there is controversy) and the recombination rate. By applying the 1.75x correction factor, we may obtain a time for European-African split at 96 thousand years and a continuation of gene flow between Europeans and Africans down to 23 thousand years.
I suppose that things might be complicated by the occurrence of Amerindian-like admixture in some West Eurasians in the past, as well as the occurrence of intra-African admixture (which I've called "Palaeoafrican") in the ancestry of Yoruba, both of which do not appear to be modeled here: the former might have infused an "African-less" component of ancestry at a time when the authors suggest that there was continuing gene flow between West Eurasia and African; the latter would inflate the effective population size of the Yoruba and make the appear earlier diverged from non-Africans.
In any case, this is a useful addition to our understanding of human history and may tie in to some of my arguments about Eurasian back-migration into Africa (although the authors consider bidrectional gene flow in their model). The lack of non-M,N mitochondria in non-Africans makes the post-OoA gene flow from Africa->Eurasia difficult to stomach, while the opposing migration of Y-haplogroup E bearers into Africa (as I have suggested) seems too instantaneous to account for the authors' evidence for protracted gene flow.
PLoS Genet 9(6): e1003521. doi:10.1371/journal.pgen.1003521
Inferring Demographic History from a Spectrum of Shared Haplotype Lengths
Kelley Harris, Rasmus Nielsen
There has been much recent excitement about the use of genetics to elucidate ancestral history and demography. Whole genome data from humans and other species are revealing complex stories of divergence and admixture that were left undiscovered by previous smaller data sets. A central challenge is to estimate the timing of past admixture and divergence events, for example the time at which Neanderthals exchanged genetic material with humans and the time at which modern humans left Africa. Here, we present a method for using sequence data to jointly estimate the timing and magnitude of past admixture events, along with population divergence times and changes in effective population size. We infer demography from a collection of pairwise sequence alignments by summarizing their length distribution of tracts of identity by state (IBS) and maximizing an analytic composite likelihood derived from a Markovian coalescent approximation. Recent gene flow between populations leaves behind long tracts of identity by descent (IBD), and these tracts give our method power by influencing the distribution of shared IBS tracts. In simulated data, we accurately infer the timing and strength of admixture events, population size changes, and divergence times over a variety of ancient and recent time scales. Using the same technique, we analyze deeply sequenced trio parents from the 1000 Genomes project. The data show evidence of extensive gene flow between Africa and Europe after the time of divergence as well as substructure and gene flow among ancestral hominids. In particular, we infer that recent African-European gene flow and ancient ghost admixture into Europe are both necessary to explain the spectrum of IBS sharing in the trios, rejecting simpler models that contain less population structure.
Link
A practical application of the new technique is applied to European-African history:
We estimate that the European-African divergence occurred 55 kya and that gene flow continued until 13 kya. About 5.8% of European genetic material is derived from a ghost population that diverged 420 kya from the ancestors of modern humans. The out-of-Africa bottleneck period, where the European effective population size is only 1,530, lasts until 5.9 kya.The authors use the "old" 2.5x10-8 mutation derived from a paleontological calibration of the human-chimp split, which renders their calculations comparable to many past papers on human demographic history, but at odds with many of the newer rates that are approximately twice slower. There is lingering controversy about the appropriateness of different rates.
The authors estimate that perhaps a 1.75x increase in their estimates will be effected if the slower rates are used; this is not 2x as one might expect from a 2x slower rate, because their age estimates depend on both the mutation rate (for which there is controversy) and the recombination rate. By applying the 1.75x correction factor, we may obtain a time for European-African split at 96 thousand years and a continuation of gene flow between Europeans and Africans down to 23 thousand years.
I suppose that things might be complicated by the occurrence of Amerindian-like admixture in some West Eurasians in the past, as well as the occurrence of intra-African admixture (which I've called "Palaeoafrican") in the ancestry of Yoruba, both of which do not appear to be modeled here: the former might have infused an "African-less" component of ancestry at a time when the authors suggest that there was continuing gene flow between West Eurasia and African; the latter would inflate the effective population size of the Yoruba and make the appear earlier diverged from non-Africans.
In any case, this is a useful addition to our understanding of human history and may tie in to some of my arguments about Eurasian back-migration into Africa (although the authors consider bidrectional gene flow in their model). The lack of non-M,N mitochondria in non-Africans makes the post-OoA gene flow from Africa->Eurasia difficult to stomach, while the opposing migration of Y-haplogroup E bearers into Africa (as I have suggested) seems too instantaneous to account for the authors' evidence for protracted gene flow.
PLoS Genet 9(6): e1003521. doi:10.1371/journal.pgen.1003521
Inferring Demographic History from a Spectrum of Shared Haplotype Lengths
Kelley Harris, Rasmus Nielsen
There has been much recent excitement about the use of genetics to elucidate ancestral history and demography. Whole genome data from humans and other species are revealing complex stories of divergence and admixture that were left undiscovered by previous smaller data sets. A central challenge is to estimate the timing of past admixture and divergence events, for example the time at which Neanderthals exchanged genetic material with humans and the time at which modern humans left Africa. Here, we present a method for using sequence data to jointly estimate the timing and magnitude of past admixture events, along with population divergence times and changes in effective population size. We infer demography from a collection of pairwise sequence alignments by summarizing their length distribution of tracts of identity by state (IBS) and maximizing an analytic composite likelihood derived from a Markovian coalescent approximation. Recent gene flow between populations leaves behind long tracts of identity by descent (IBD), and these tracts give our method power by influencing the distribution of shared IBS tracts. In simulated data, we accurately infer the timing and strength of admixture events, population size changes, and divergence times over a variety of ancient and recent time scales. Using the same technique, we analyze deeply sequenced trio parents from the 1000 Genomes project. The data show evidence of extensive gene flow between Africa and Europe after the time of divergence as well as substructure and gene flow among ancestral hominids. In particular, we infer that recent African-European gene flow and ancient ghost admixture into Europe are both necessary to explain the spectrum of IBS sharing in the trios, rejecting simpler models that contain less population structure.
Link
June 04, 2013
IBD sharing between Iberians and North Africans (Botigué et al. 2013)
An interesting new paper documents an excess of IBD sharing between Iberians (excluding Basques) and North African (and particular NW African) populations.
It would have been nice if the authors had used techniques such as rolloff and ALDER or those of Jin et al. (2012) to say something about the time/nature of the admixture event detected via IBD sharing; insteady, they use variance in admixture proportions, which gives a probably much noisier estimate, with the basic idea being that in the first few generations post-admixture there are individuals with much varying admixture proportions, but these tend to be homogenized over time.
The occurrence of North African-specific admixture in SW Europe has long been suspected on the basis of Y-chromosome/mtDNA work (e.g., the presence of E-M81 which is probably the best North African marker in existence). It also makes sense, because of the limited occurrence of Sub-Saharan markers in Iberia: such elements did not, presumably, fly over North Africa, but landed in Iberia via people who were themselves admixed.
A couple notes of caution:
(i) the use of ADMIXTURE as a means of estimating admixture proportions is dangerous in this case, because of the hybridity of "North Africans" themselves, which according to published estimates experienced Sub-Saharan African admixture in the last few thousand years. In my own experiments it is clear that "North Africans" are a mixture of three basic components related to Europe, Sub-Saharan Africa, and the Near East. Nonetheless, in my own experiments I do also get an excess of the component I've labeled "Northwest African" in Iberia that is not shared by Basques or French.
(ii) as I've emphasized before, IBD sharing between populations does not indicate the direction of gene flow. One would have to look at the ancestry of the shared segments to determine their origin. To give a simple example, an IBD segment shared by a Spaniard and a Mexican could be European, African, or Native American, and -thanks to historical knowledge- we can be fairly sure that the number of such segments is also in the given order.
Given that Iberia is the neighbor of NW Africa one would not be surprised if there was gene flow in both directions, and while North Africa gene flow into Iberia is one possible explanation, some of the gene flow may have gone the other way, e.g., with contacts during the Pax Romana, fleeing Iberian Muslim in the post-reconquista period, Barbary pirates attacking Christian ships and the like. In any case, it would be interesting to catalogue IBD shared segments between Iberia and NW Africa in terms of their geographical origin.
(iii) the sources and timing of admixture could potentially be determined by ancient DNA work. The three most recent time periods are related to the slave trade (both European of Africans and vice versa), the Islamic period, and the Roman Empire. Presumably, with the sampling of enough individuals, that type of admixture ought to manifest in populations living before/after each of these three events.
In any case, this is an interesting paper which is also accompanied by publicly accessible data.
PNAS doi: 10.1073/pnas.1306223110
Gene flow from North Africa contributes to differential human genetic diversity in southern Europe
Laura R. Botigué et al.
Human genetic diversity in southern Europe is higher than in other regions of the continent. This difference has been attributed to postglacial expansions, the demic diffusion of agriculture from the Near East, and gene flow from Africa. Using SNP data from 2,099 individuals in 43 populations, we show that estimates of recent shared ancestry between Europe and Africa are substantially increased when gene flow from North Africans, rather than Sub-Saharan Africans, is considered. The gradient of North African ancestry accounts for previous observations of low levels of sharing with Sub-Saharan Africa and is independent of recent gene flow from the Near East. The source of genetic diversity in southern Europe has important biomedical implications; we find that most disease risk alleles from genome-wide association studies follow expected patterns of divergence between Europe and North Africa, with the principal exception of multiple sclerosis.
Link
It would have been nice if the authors had used techniques such as rolloff and ALDER or those of Jin et al. (2012) to say something about the time/nature of the admixture event detected via IBD sharing; insteady, they use variance in admixture proportions, which gives a probably much noisier estimate, with the basic idea being that in the first few generations post-admixture there are individuals with much varying admixture proportions, but these tend to be homogenized over time.
The occurrence of North African-specific admixture in SW Europe has long been suspected on the basis of Y-chromosome/mtDNA work (e.g., the presence of E-M81 which is probably the best North African marker in existence). It also makes sense, because of the limited occurrence of Sub-Saharan markers in Iberia: such elements did not, presumably, fly over North Africa, but landed in Iberia via people who were themselves admixed.
A couple notes of caution:
(i) the use of ADMIXTURE as a means of estimating admixture proportions is dangerous in this case, because of the hybridity of "North Africans" themselves, which according to published estimates experienced Sub-Saharan African admixture in the last few thousand years. In my own experiments it is clear that "North Africans" are a mixture of three basic components related to Europe, Sub-Saharan Africa, and the Near East. Nonetheless, in my own experiments I do also get an excess of the component I've labeled "Northwest African" in Iberia that is not shared by Basques or French.
(ii) as I've emphasized before, IBD sharing between populations does not indicate the direction of gene flow. One would have to look at the ancestry of the shared segments to determine their origin. To give a simple example, an IBD segment shared by a Spaniard and a Mexican could be European, African, or Native American, and -thanks to historical knowledge- we can be fairly sure that the number of such segments is also in the given order.
Given that Iberia is the neighbor of NW Africa one would not be surprised if there was gene flow in both directions, and while North Africa gene flow into Iberia is one possible explanation, some of the gene flow may have gone the other way, e.g., with contacts during the Pax Romana, fleeing Iberian Muslim in the post-reconquista period, Barbary pirates attacking Christian ships and the like. In any case, it would be interesting to catalogue IBD shared segments between Iberia and NW Africa in terms of their geographical origin.
(iii) the sources and timing of admixture could potentially be determined by ancient DNA work. The three most recent time periods are related to the slave trade (both European of Africans and vice versa), the Islamic period, and the Roman Empire. Presumably, with the sampling of enough individuals, that type of admixture ought to manifest in populations living before/after each of these three events.
In any case, this is an interesting paper which is also accompanied by publicly accessible data.
PNAS doi: 10.1073/pnas.1306223110
Gene flow from North Africa contributes to differential human genetic diversity in southern Europe
Laura R. Botigué et al.
Human genetic diversity in southern Europe is higher than in other regions of the continent. This difference has been attributed to postglacial expansions, the demic diffusion of agriculture from the Near East, and gene flow from Africa. Using SNP data from 2,099 individuals in 43 populations, we show that estimates of recent shared ancestry between Europe and Africa are substantially increased when gene flow from North Africans, rather than Sub-Saharan Africans, is considered. The gradient of North African ancestry accounts for previous observations of low levels of sharing with Sub-Saharan Africa and is independent of recent gene flow from the Near East. The source of genetic diversity in southern Europe has important biomedical implications; we find that most disease risk alleles from genome-wide association studies follow expected patterns of divergence between Europe and North Africa, with the principal exception of multiple sclerosis.
Link
May 31, 2013
LOCO-LD paper and software
Link to software.
AJHG doi: 10.1016/j.ajhg.2013.04.023
Enhanced Localization of Genetic Samples through Linkage-Disequilibrium Correction
Yael Baran et al.
Characterizing the spatial patterns of genetic diversity in human populations has a wide range of applications, from detecting genetic mutations associated with disease to inferring human history. Current approaches, including the widely used principal-component analysis, are not suited for the analysis of linked markers, and local and long-range linkage disequilibrium (LD) can dramatically reduce the accuracy of spatial localization when unaccounted for. To overcome this, we have introduced an approach that performs spatial localization of individuals on the basis of their genetic data and explicitly models LD among markers by using a multivariate normal distribution. By leveraging external reference panels, we derive closed-form solutions to the optimization procedure to achieve a computationally efficient method that can handle large data sets. We validate the method on empirical data from a large sample of European individuals from the POPRES data set, as well as on a large sample of individuals of Spanish ancestry. First, we show that by modeling LD, we achieve accuracy superior to that of existing methods. Importantly, whereas other methods show decreased performance when dense marker panels are used in the inference, our approach improves in accuracy as more markers become available. Second, we show that accurate localization of genetic data can be achieved with only a part of the genome, and this could potentially enable the spatial localization of admixed samples that have a fraction of their genome originating from a given continent. Finally, we demonstrate that our approach is resistant to distortions resulting from long-range LD regions; such distortions can dramatically bias the results when unaccounted for.
Link
AJHG doi: 10.1016/j.ajhg.2013.04.023
Enhanced Localization of Genetic Samples through Linkage-Disequilibrium Correction
Yael Baran et al.
Characterizing the spatial patterns of genetic diversity in human populations has a wide range of applications, from detecting genetic mutations associated with disease to inferring human history. Current approaches, including the widely used principal-component analysis, are not suited for the analysis of linked markers, and local and long-range linkage disequilibrium (LD) can dramatically reduce the accuracy of spatial localization when unaccounted for. To overcome this, we have introduced an approach that performs spatial localization of individuals on the basis of their genetic data and explicitly models LD among markers by using a multivariate normal distribution. By leveraging external reference panels, we derive closed-form solutions to the optimization procedure to achieve a computationally efficient method that can handle large data sets. We validate the method on empirical data from a large sample of European individuals from the POPRES data set, as well as on a large sample of individuals of Spanish ancestry. First, we show that by modeling LD, we achieve accuracy superior to that of existing methods. Importantly, whereas other methods show decreased performance when dense marker panels are used in the inference, our approach improves in accuracy as more markers become available. Second, we show that accurate localization of genetic data can be achieved with only a part of the genome, and this could potentially enable the spatial localization of admixed samples that have a fraction of their genome originating from a given continent. Finally, we demonstrate that our approach is resistant to distortions resulting from long-range LD regions; such distortions can dramatically bias the results when unaccounted for.
Link
May 20, 2013
More population structure in the Netherlands (Lao et al. 2013)
There was a recent article on the topic by Abdellaoui et al., and here is another one.
Investigative Genetics 2013, 4:9 doi:10.1186/2041-2223-4-9
Clinal distribution of human genomic diversity across the Netherlands despite archaeological evidence for genetic discontinuities in Dutch population history
Oscar Lao et al.
Abstract (provisional)
Background
The presence of a southeast to northwest gradient across Europe in human genetic diversity is a well-established observation and has recently been confirmed by genome-wide single nucleotide polymorphism (SNP) data. This pattern is traditionally explained by major prehistoric human migration events in Palaeolithic and Neolithic times. Here, we investigate whether (similar) spatial patterns in human genomic diversity also occur on a micro-geographic scale within Europe, such as in the Netherlands, and if so, whether these patterns could also be explained by more recent demographic events, such as those that occurred in Dutch population history.
Methods
We newly collected data on a total of 999 Dutch individuals sampled at 54 sites across the country at 443,816 autosomal SNPs using the Genome-Wide Human SNP Array 5.0 (Affymetrix). We studied the individual genetic relationships by means of classical multidimensional scaling (MDS) using different genetic distance matrices, spatial ancestry analysis (SPA), and ADMIXTURE software. We further performed dedicated analyses to search for spatial patterns in the genomic variation and conducted simulations (SPLATCHE2) to provide a historical interpretation of the observed spatial patterns.
Results
We detected a subtle but clearly noticeable genomic population substructure in the Dutch population, allowing differentiation of a north-eastern, central-western, central-northern and a southern group. Furthermore, we observed a statistically significant southeast to northwest cline in the distribution of genomic diversity across the Netherlands, similar to earlier findings from across Europe. Simulation analyses indicate that this genomic gradient could similarly be caused by ancient as well as by the more recent events in Dutch history.
Conclusions
Considering the strong archaeological evidence for genetic discontinuity in the Netherlands, we interpret the observed clinal pattern of genomic diversity as being caused by recent rather than ancient events in Dutch population history. We therefore suggest that future human population genetic studies pay more attention to recent demographic history in interpreting genetic clines. Furthermore, our study demonstrates that genetic population substructure is detectable on a small geographic scale in Europe despite recent demographic events, a finding we consider potentially relevant for future epidemiological and forensic studies.
Link
Investigative Genetics 2013, 4:9 doi:10.1186/2041-2223-4-9
Clinal distribution of human genomic diversity across the Netherlands despite archaeological evidence for genetic discontinuities in Dutch population history
Oscar Lao et al.
Abstract (provisional)
Background
The presence of a southeast to northwest gradient across Europe in human genetic diversity is a well-established observation and has recently been confirmed by genome-wide single nucleotide polymorphism (SNP) data. This pattern is traditionally explained by major prehistoric human migration events in Palaeolithic and Neolithic times. Here, we investigate whether (similar) spatial patterns in human genomic diversity also occur on a micro-geographic scale within Europe, such as in the Netherlands, and if so, whether these patterns could also be explained by more recent demographic events, such as those that occurred in Dutch population history.
Methods
We newly collected data on a total of 999 Dutch individuals sampled at 54 sites across the country at 443,816 autosomal SNPs using the Genome-Wide Human SNP Array 5.0 (Affymetrix). We studied the individual genetic relationships by means of classical multidimensional scaling (MDS) using different genetic distance matrices, spatial ancestry analysis (SPA), and ADMIXTURE software. We further performed dedicated analyses to search for spatial patterns in the genomic variation and conducted simulations (SPLATCHE2) to provide a historical interpretation of the observed spatial patterns.
Results
We detected a subtle but clearly noticeable genomic population substructure in the Dutch population, allowing differentiation of a north-eastern, central-western, central-northern and a southern group. Furthermore, we observed a statistically significant southeast to northwest cline in the distribution of genomic diversity across the Netherlands, similar to earlier findings from across Europe. Simulation analyses indicate that this genomic gradient could similarly be caused by ancient as well as by the more recent events in Dutch history.
Conclusions
Considering the strong archaeological evidence for genetic discontinuity in the Netherlands, we interpret the observed clinal pattern of genomic diversity as being caused by recent rather than ancient events in Dutch population history. We therefore suggest that future human population genetic studies pay more attention to recent demographic history in interpreting genetic clines. Furthermore, our study demonstrates that genetic population substructure is detectable on a small geographic scale in Europe despite recent demographic events, a finding we consider potentially relevant for future epidemiological and forensic studies.
Link
Review on germline mutation rate in humans (Campbell and Eichler 2013)
This is a nice little review of the state of the art in germline mutation rate estimation in humans. This was previously estimated using paleontological calibrations (especially the human/chimp split), but a slower mutation rate emerged on the basis of whole genome data from humans. There may be problems with the latter (because of false positive/negative mutations using whole genome sequencing), but the problem is an important one due to the use of the mutation rate to estimate time depth of common ancestry. In any case, the table on the left summarizes the results of several studies on the topic.Trends in Genetics, 17 May 2013 doi:10.1016/j.tig.2013.04.005
Properties and rates of germline mutations in humans
Catarina D. Campbell, Evan E. EichlerSee Affiliations
Summary
All genetic variation arises via new mutations; therefore, determining the rate and biases for different classes of mutation is essential for understanding the genetics of human disease and evolution. Decades of mutation rate analyses have focused on a relatively small number of loci because of technical limitations. However, advances in sequencing technology have allowed for empirical assessments of genome-wide rates of mutation. Recent studies have shown that 76% of new mutations originate in the paternal lineage and provide unequivocal evidence for an increase in mutation with paternal age. Although most analyses have focused on single nucleotide variants (SNVs), studies have begun to provide insight into the mutation rate for other classes of variation, including copy number variants (CNVs), microsatellites, and mobile element insertions (MEIs). Here, we review the genome-wide analyses for the mutation rate of several types of variants and suggest areas for future research.
Link
May 10, 2013
Deleterious mutational load and recent population history (Simons et al. 2013)
UPDATE (Feb 28, 2014): This has now appeared in Nature Genetics.
arXiv:1305.2061 [q-bio.PE]
The deleterious mutation load is insensitive to recent population history
Yuval B. Simons, Michael C. Turchin, Jonathan K. Pritchard, Guy Sella (Submitted on 9 May 2013)
Human populations have undergone dramatic changes in population size in the past 100,000 years, including a severe bottleneck of non-African populations and recent explosive population growth. There is currently great interest in how these demographic events may have affected the burden of deleterious mutations in individuals and the allele frequency spectrum of disease mutations in populations. Here we use population genetic models to show that--contrary to previous conjectures--recent human demography has likely had very little impact on the average burden of deleterious mutations carried by individuals. This prediction is supported by exome sequence data showing that African American and European American individuals carry very similar burdens of damaging mutations. We next consider whether recent population growth has increased the importance of very rare mutations in complex traits. Our analysis predicts that for most classes of disease variants, rare alleles are unlikely to contribute a large fraction of the total genetic variance, and that the impact of recent growth is likely to be modest. However, for diseases that have a direct impact on fitness, strongly deleterious rare mutations likely do play important roles, and the impact of very rare mutations will be far greater as a result of recent growth. In summary, demographic history has dramatically impacted patterns of variation in different human populations, but these changes have likely had little impact on either genetic load or on the importance of rare variants for most complex traits.
Link
arXiv:1305.2061 [q-bio.PE]
The deleterious mutation load is insensitive to recent population history
Yuval B. Simons, Michael C. Turchin, Jonathan K. Pritchard, Guy Sella (Submitted on 9 May 2013)
Human populations have undergone dramatic changes in population size in the past 100,000 years, including a severe bottleneck of non-African populations and recent explosive population growth. There is currently great interest in how these demographic events may have affected the burden of deleterious mutations in individuals and the allele frequency spectrum of disease mutations in populations. Here we use population genetic models to show that--contrary to previous conjectures--recent human demography has likely had very little impact on the average burden of deleterious mutations carried by individuals. This prediction is supported by exome sequence data showing that African American and European American individuals carry very similar burdens of damaging mutations. We next consider whether recent population growth has increased the importance of very rare mutations in complex traits. Our analysis predicts that for most classes of disease variants, rare alleles are unlikely to contribute a large fraction of the total genetic variance, and that the impact of recent growth is likely to be modest. However, for diseases that have a direct impact on fitness, strongly deleterious rare mutations likely do play important roles, and the impact of very rare mutations will be far greater as a result of recent growth. In summary, demographic history has dramatically impacted patterns of variation in different human populations, but these changes have likely had little impact on either genetic load or on the importance of rare variants for most complex traits.
Link
April 02, 2013
More asymmetric migration (Sundqvist et al. 2013)
A day after the paper by Peter and Slatkin, a new paper has appeared on the arXiv dealing with the problem of detecting directionaliy in human migration patterns. This seems to be purely methodological, so no new insights on human history to report.
arXiv:1304.0118 [q-bio.PE]
A new approach to estimate directional genetic differentiation and asymmetric migration patterns
Lisa Sundqvist, Martin Zackrisson, David Kleinhans
In the field of population genetics measures of genetic differentiation are widely used to gather information on the structure and the amount of gene flow between populations. These indirect measures are based on a number of simplifying assumptions, for instance equal population size and symmetric migration. Structured populations with asymmetric migration patterns, frequently occur in nature and information about directional gene flow would here be of great interest. Nevertheless current measures of genetic differentiation cannot be used in such systems without violating the assumptions. To get information on asymmetric migration patterns from genetic data rather complex models using maximum likelihood or Bayesian approaches generally need to be applied. In such models a large number of parameters are estimated simultaneously and this involves complex optimization algorithms. We here introduce a new approach that intends to fill the gap between the complex approaches and the symmetric measures of genetic differentiation. Our approach makes it possible to calculate a directional component of genetic differentiation at low computational effort using any of the classical measures of genetic differentiation. The approach is based on defining a pool of migrants for any pair of populations and calculating measures for genetic differentiation between the populations and the respective pools. The directional measures of genetic differentiation can further be used to calculate asymmetric migration. The procedure is demonstrated with a simulated data set with known migration pattern. A comparison of the estimation results with the migration pattern used for simulation suggests, that our method captures relevant properties of migration patterns even at low migration frequencies and with few marker loci.
Link
arXiv:1304.0118 [q-bio.PE]
A new approach to estimate directional genetic differentiation and asymmetric migration patterns
Lisa Sundqvist, Martin Zackrisson, David Kleinhans
In the field of population genetics measures of genetic differentiation are widely used to gather information on the structure and the amount of gene flow between populations. These indirect measures are based on a number of simplifying assumptions, for instance equal population size and symmetric migration. Structured populations with asymmetric migration patterns, frequently occur in nature and information about directional gene flow would here be of great interest. Nevertheless current measures of genetic differentiation cannot be used in such systems without violating the assumptions. To get information on asymmetric migration patterns from genetic data rather complex models using maximum likelihood or Bayesian approaches generally need to be applied. In such models a large number of parameters are estimated simultaneously and this involves complex optimization algorithms. We here introduce a new approach that intends to fill the gap between the complex approaches and the symmetric measures of genetic differentiation. Our approach makes it possible to calculate a directional component of genetic differentiation at low computational effort using any of the classical measures of genetic differentiation. The approach is based on defining a pool of migrants for any pair of populations and calculating measures for genetic differentiation between the populations and the respective pools. The directional measures of genetic differentiation can further be used to calculate asymmetric migration. The procedure is demonstrated with a simulated data set with known migration pattern. A comparison of the estimation results with the migration pattern used for simulation suggests, that our method captures relevant properties of migration patterns even at low migration frequencies and with few marker loci.
Link
April 01, 2013
Directionality index for detecting origin of range expansions
This appears to be an interesting methodology for detecting directionality in genetic datasets. I am not sure how it might perform in the presence of admixture, a topic that was not discussed. Interestingly, the San appear as the only human population that had positive directionality values with all others, suggesting -to the authors- that they are closest to the origin of humans. On the other hand, in Pakistan, they found the Makrani to be the most ancestral population, and I strongly suspect that this may be related to the African admixture found in that population and not in others from that country.
arXiv:1303.7475v1 [q-bio.PE]
Detecting range expansions from genetic data
Benjamin M Peter, Montgomery Slatkin
We propose a method that uses genetic data to test for the occurrence of a recent range expansion and to infer the location of the origin of the expansion. We introduce a statistic for pairs of populations $\psi$ (the directionality index) that detects asymmetries in the two-dimensional allele frequency spectrum caused by the series of founder events that happen during an expansion. Such asymmetry arises because low frequency alleles tend to be lost during founder events, thus creating clines in the frequencies of surviving low-frequency alleles. Using simulations, we further show that $\psi$ is more powerful for detecting range expansions than both $F_{ST}$ and clines in heterozygosity. We illustrate the utility of $\psi$ by applying it to a data set from modern humans and show how we can include more complicated scenarios such as multiple expansion origins or barriers to migration in the model.
Link
arXiv:1303.7475v1 [q-bio.PE]
Detecting range expansions from genetic data
Benjamin M Peter, Montgomery Slatkin
We propose a method that uses genetic data to test for the occurrence of a recent range expansion and to infer the location of the origin of the expansion. We introduce a statistic for pairs of populations $\psi$ (the directionality index) that detects asymmetries in the two-dimensional allele frequency spectrum caused by the series of founder events that happen during an expansion. Such asymmetry arises because low frequency alleles tend to be lost during founder events, thus creating clines in the frequencies of surviving low-frequency alleles. Using simulations, we further show that $\psi$ is more powerful for detecting range expansions than both $F_{ST}$ and clines in heterozygosity. We illustrate the utility of $\psi$ by applying it to a data set from modern humans and show how we can include more complicated scenarios such as multiple expansion origins or barriers to migration in the model.
Link
February 15, 2013
Whole genome Y-SNP calling
BMC Genomics 2013, 14:101 doi:10.1186/1471-2164-14-101
AMY-tree: an algorithm to use whole genome SNP calling for Y chromosomal phylogenetic applications
Anneleen Van Geystelen et al.
Abstract (provisional)
Background
Due to the rapid progress of next-generation sequencing (NGS) facilities, an explosion of human whole genome data will become available in the coming years. These data can be used to optimize and to increase the resolution of the phylogenetic Y chromosomal tree. Moreover, the exponential growth of known Y chromosomal lineages will require an automatic determination of the phylogenetic position of an individual based on whole genome SNP calling data and an up to date Y chromosomal tree.
Results
We present an automated approach, 'AMY-tree', which is able to determine the phylogenetic position of a Y chromosome using a whole genome SNP profile, independently from the NGS platform and SNP calling program, whereby mistakes in the SNP calling or phylogenetic Y chromosomal tree are taken into account. Moreover, AMY-tree indicates ambiguities within the present phylogenetic tree and points out new Y-SNPs which may be phylogenetically relevant. The AMY-tree software package was validated successfully on 118 whole genome SNP profiles of 109 males with different origins. Moreover, support was found for an unknown recurrent mutation, wrong reported mutation conversions and a large amount of new interesting Y-SNPs.
Conclusions
Therefore, AMY-tree is a useful tool to determine the Y lineage of a sample based on SNP calling, to identify Y-SNPs with yet unknown phylogenetic position and to optimize the Y chromosomal phylogenetic tree in the future. AMY-tree will not add lineages to the existing phylogenetic tree of the Y-chromosome but it is the first step to analyse whole genome SNP profiles in a phylogenetic framework.
Link
AMY-tree: an algorithm to use whole genome SNP calling for Y chromosomal phylogenetic applications
Anneleen Van Geystelen et al.
Abstract (provisional)
Background
Due to the rapid progress of next-generation sequencing (NGS) facilities, an explosion of human whole genome data will become available in the coming years. These data can be used to optimize and to increase the resolution of the phylogenetic Y chromosomal tree. Moreover, the exponential growth of known Y chromosomal lineages will require an automatic determination of the phylogenetic position of an individual based on whole genome SNP calling data and an up to date Y chromosomal tree.
Results
We present an automated approach, 'AMY-tree', which is able to determine the phylogenetic position of a Y chromosome using a whole genome SNP profile, independently from the NGS platform and SNP calling program, whereby mistakes in the SNP calling or phylogenetic Y chromosomal tree are taken into account. Moreover, AMY-tree indicates ambiguities within the present phylogenetic tree and points out new Y-SNPs which may be phylogenetically relevant. The AMY-tree software package was validated successfully on 118 whole genome SNP profiles of 109 males with different origins. Moreover, support was found for an unknown recurrent mutation, wrong reported mutation conversions and a large amount of new interesting Y-SNPs.
Conclusions
Therefore, AMY-tree is a useful tool to determine the Y lineage of a sample based on SNP calling, to identify Y-SNPs with yet unknown phylogenetic position and to optimize the Y chromosomal phylogenetic tree in the future. AMY-tree will not add lineages to the existing phylogenetic tree of the Y-chromosome but it is the first step to analyse whole genome SNP profiles in a phylogenetic framework.
Link
January 04, 2013
Deep whole-genome sequencing of 100 Malays
The 1000 Genomes Project is the largest collection of full human genomes currently available, but most of its 2.5k samples have been sequenced at low coverage. One downside of this is that infrequent variants are often missed. If an individual is polymorphic at some site, then the chance of detecting this polymorphism increases with the number of reads covering that site. If a number of individuals are sampled, then polymorphisms that are common in the population will probably be detected in a few individuals even if a low number of reads is used for each of them; but, if they are infrequent, then they are more likely to be missed. Hence, low-coverage sequencing of population samples will tend to find common variants and will tend to miss less common variants relative to high-coverage sequencing.
This idea is intuitively correct, but the question of the added power of high-coverage sequencing to detect variants can only be addressed by giving the same individuals both low- and high-coverage sequencing. This is the topic of a new paper in AJHG which creates a useful comparison benchmark for the performance of the two types of sequencing methods. High-coverage sequencing may be needed for things like disease studies (because deleterious alleles tend to be low-frequency), or the study of recent human demography (because recent population growth has resulted in an abundance of low-frequency SNPs that have not had enough time to reach a high population frequency yet).
AJHG dx.doi.org/10.1016/j.ajhg.2012.12.005
Deep Whole-Genome Sequencing of 100 Southeast Asian Malays
Lai-Ping Wong et al.
Whole-genome sequencing across multiple samples in a population provides an unprecedented opportunity for comprehensively characterizing the polymorphic variants in the population. Although the 1000 Genomes Project (1KGP) has offered brief insights into the value of population-level sequencing, the low coverage has compromised the ability to confidently detect rare and low-frequency variants. In addition, the composition of populations in the 1KGP is not complete, despite the fact that the study design has been extended to more than 2,500 samples from more than 20 population groups. The Malays are one of the Austronesian groups predominantly present in Southeast Asia and Oceania, and the Singapore Sequencing Malay Project (SSMP) aims to perform deep whole-genome sequencing of 100 healthy Malays. By sequencing at a minimum of 30? coverage, we have illustrated the higher sensitivity at detecting low-frequency and rare variants and the ability to investigate the presence of hotspots of functional mutations. Compared to the low-pass sequencing in the 1KGP, the deeper coverage allows more functional variants to be identified for each person. A comparison of the fidelity of genotype imputation of Malays indicated that a population-specific reference panel, such as the SSMP, outperforms a cosmopolitan panel with larger number of individuals for common SNPs. For lower-frequency (less than 5%) markers, a larger number of individuals might have to be whole-genome sequenced so that the accuracy currently afforded by the 1KGP can be achieved. The SSMP data are expected to be the benchmark for evaluating the value of deep population-level sequencing versus low-pass sequencing, especially in populations that are poorly represented in population-genetics studies.
Link
This idea is intuitively correct, but the question of the added power of high-coverage sequencing to detect variants can only be addressed by giving the same individuals both low- and high-coverage sequencing. This is the topic of a new paper in AJHG which creates a useful comparison benchmark for the performance of the two types of sequencing methods. High-coverage sequencing may be needed for things like disease studies (because deleterious alleles tend to be low-frequency), or the study of recent human demography (because recent population growth has resulted in an abundance of low-frequency SNPs that have not had enough time to reach a high population frequency yet).
AJHG dx.doi.org/10.1016/j.ajhg.2012.12.005
Deep Whole-Genome Sequencing of 100 Southeast Asian Malays
Lai-Ping Wong et al.
Whole-genome sequencing across multiple samples in a population provides an unprecedented opportunity for comprehensively characterizing the polymorphic variants in the population. Although the 1000 Genomes Project (1KGP) has offered brief insights into the value of population-level sequencing, the low coverage has compromised the ability to confidently detect rare and low-frequency variants. In addition, the composition of populations in the 1KGP is not complete, despite the fact that the study design has been extended to more than 2,500 samples from more than 20 population groups. The Malays are one of the Austronesian groups predominantly present in Southeast Asia and Oceania, and the Singapore Sequencing Malay Project (SSMP) aims to perform deep whole-genome sequencing of 100 healthy Malays. By sequencing at a minimum of 30? coverage, we have illustrated the higher sensitivity at detecting low-frequency and rare variants and the ability to investigate the presence of hotspots of functional mutations. Compared to the low-pass sequencing in the 1KGP, the deeper coverage allows more functional variants to be identified for each person. A comparison of the fidelity of genotype imputation of Malays indicated that a population-specific reference panel, such as the SSMP, outperforms a cosmopolitan panel with larger number of individuals for common SNPs. For lower-frequency (less than 5%) markers, a larger number of individuals might have to be whole-genome sequenced so that the accuracy currently afforded by the 1KGP can be achieved. The SSMP data are expected to be the benchmark for evaluating the value of deep population-level sequencing versus low-pass sequencing, especially in populations that are poorly represented in population-genetics studies.
Link
December 28, 2012
Estonian Biocentre public data
The Estonian Biocentre (EBC) have put up all their free data in one convenient page. I have used most of these in my own experiments, and I must say that their public availability has been instrumental in enabling the type of "genome blogging" that I and others have engaged in over the last few years.
I have hitherto used some EBC data from GEO and some downloaded from the EBC site itself, so this is a good opportunity to rebuild all my datasets from a single source; as a bonus, the data has been lifted to build37/hg19, so I finally gave in and started to lift all my other datasets as well, using liftOver and the appropriate chain file.
I have hitherto used some EBC data from GEO and some downloaded from the EBC site itself, so this is a good opportunity to rebuild all my datasets from a single source; as a bonus, the data has been lifted to build37/hg19, so I finally gave in and started to lift all my other datasets as well, using liftOver and the appropriate chain file.
December 21, 2012
Estimating heterozygosity from low coverage sequencing data
arXiv:1212.4125 [q-bio.PE]
Estimating heterozygosity from a low-coverage genome sequence, leveraging data from other individuals sequenced at the same sites
Katarzyna Bryc, Nick Patterson, David Reich
High-throughput shotgun sequence data makes it possible in principle to accurately estimate population genetic parameters without confounding by SNP ascertainment bias. One such statistic of interest is the proportion of heterozygous sites within an individual's genome, which is informative about inbreeding and effective population size. However, in many cases, the available sequence data of an individual is limited to low coverage, preventing the confident calling of genotypes necessary to directly count the proportion of heterozygous sites. Here, we present a method for estimating an individual's genome-wide rate of heterozygosity from low-coverage sequence data, without an intermediate step calling genotypes. Our method jointly learns the shared allele distribution between the individual and a panel of other individuals, together with the sequencing error distributions and the reference bias. We show our method works well, first by its performance on simulated sequence data, and secondly on real sequence data where we obtain estimates using low coverage data consistent with those from higher coverage. We apply our method to obtain estimates of the rate of heterozygosity for 11 humans from diverse world-wide populations, and through this analysis reveal the complex dependency of local sequencing coverage on the true underlying heterozygosity, which complicates the estimation of heterozygosity from sequence data. We show filters can correct for the confounding by sequencing depth. We find in practice that ratios of heterozygosity are more interpretable than absolute estimates, and show that we obtain excellent conformity of ratios of heterozygosity with previous estimates from higher coverage data.
Link
Estimating heterozygosity from a low-coverage genome sequence, leveraging data from other individuals sequenced at the same sites
Katarzyna Bryc, Nick Patterson, David Reich
High-throughput shotgun sequence data makes it possible in principle to accurately estimate population genetic parameters without confounding by SNP ascertainment bias. One such statistic of interest is the proportion of heterozygous sites within an individual's genome, which is informative about inbreeding and effective population size. However, in many cases, the available sequence data of an individual is limited to low coverage, preventing the confident calling of genotypes necessary to directly count the proportion of heterozygous sites. Here, we present a method for estimating an individual's genome-wide rate of heterozygosity from low-coverage sequence data, without an intermediate step calling genotypes. Our method jointly learns the shared allele distribution between the individual and a panel of other individuals, together with the sequencing error distributions and the reference bias. We show our method works well, first by its performance on simulated sequence data, and secondly on real sequence data where we obtain estimates using low coverage data consistent with those from higher coverage. We apply our method to obtain estimates of the rate of heterozygosity for 11 humans from diverse world-wide populations, and through this analysis reveal the complex dependency of local sequencing coverage on the true underlying heterozygosity, which complicates the estimation of heterozygosity from sequence data. We show filters can correct for the confounding by sequencing depth. We find in practice that ratios of heterozygosity are more interpretable than absolute estimates, and show that we obtain excellent conformity of ratios of heterozygosity with previous estimates from higher coverage data.
Link
Subscribe to:
Posts (Atom)






