Showing posts with label 1000Genomes. Show all posts
Showing posts with label 1000Genomes. Show all posts

May 24, 2014

High genetic differentiation between populations often driven by classic selective sweeps

bioRxiv doi: http://dx.doi.org/10.1101/005462

Human genomic regions with exceptionally high or low levels of population differentiation identified from 911 whole-genome sequences

Vincenza Colonna et al.

Background: Population differentiation has proved to be effective for identifying loci under geographically-localized positive selection, and has the potential to identify loci subject to balancing selection. We have previously investigated the pattern of genetic differentiation among human populations at 36.8 million genomic variants to identify sites in the genome showing high frequency differences. Here, we extend this dataset to include additional variants, survey sites with low levels of differentiation, and evaluate the extent to which highly differentiated sites are likely to result from selective or other processes. Results: We demonstrate that while sites of low differentiation represent sampling effects rather than balancing selection, sites showing extremely high population differentiation are enriched for positive selection events and that one half may be the result of classic selective sweeps. Among these, we rediscover known examples, where we actually identify the established functional SNP, and discover novel examples including the genes ABCA12, CALD1 and ZNF804, which we speculate may be linked to adaptations in skin, calcium metabolism and defense, respectively. Conclusions: We have identified known and many novel candidate regions for geographically restricted positive selection, and suggest several directions for further research.

Link

June 19, 2013

Native American origins from whole-genome and exome data (Gravel et al. 2013)

arXiv:1306.4021 [q-bio.PE]

Reconstructing Native American Migrations from Whole-genome and Whole-exome Data

Simon Gravel et al.

There is great scientific and popular interest in understanding the genetic history of populations in the Americas. We wish to understand when different regions of the continent were inhabited, where settlers came from, and how current inhabitants relate genetically to earlier populations. Recent studies unraveled parts of the genetic history of the continent using genotyping arrays and uniparental markers. The 1000 Genomes Project provides a unique opportunity for improving our understanding of population genetic history by providing over a hundred sequenced low coverage genomes and exomes from Colombian (CLM), Mexican-American (MXL), and Puerto Rican (PUR) populations. Here, we explore the genomic contributions of African, European, and especially Native American ancestry to these populations. Estimated Native American ancestry is 48% in MXL, 25% in CLM, and 13% in PUR. Native American ancestry in PUR appears most closely related to Equatorial-Tucanoan-speaking populations, supporting a Southern America ancestry of the Taino people of the Caribbean. We present new methods to estimate the allele frequencies in the Native American fraction of the populations, and model their distribution using a three-population demographic model. The ancestral populations to the three groups likely split in close succession: the most likely scenario, based on a peopling of the Americas 16 thousand years ago (kya), supports that the MXL Ancestors split 12.2kya, with a subsequent split of the ancestors to CLM and PUR 11.7kya. The model also features a Mexican population of 62,000, a Colombian population of 8,700, and a Puerto Rican population of 1,900. Modeling Identity-by-descent (IBD) and ancestry tract length, we show that post-contact populations also differ markedly in their effective sizes and migration patterns, with Puerto Rico showing the smallest size and the earlier migration from Europe.

Link

January 04, 2013

Deep whole-genome sequencing of 100 Malays

The 1000 Genomes Project is the largest collection of full human genomes currently available, but most of its 2.5k samples have been sequenced at low coverage. One downside of this is that infrequent variants are often missed. If an individual is polymorphic at some site, then the chance of detecting this polymorphism increases with the number of reads covering that site. If a number of individuals are sampled, then polymorphisms that are common in the population will probably be detected in a few individuals even if a low number of reads is used for each of them; but, if they are infrequent, then they are more likely to be missed. Hence, low-coverage sequencing of population samples will tend to find common variants and will tend to miss less common variants relative to high-coverage sequencing.

This idea is intuitively correct, but the question of the added power of high-coverage sequencing to detect variants can only be addressed by giving the same individuals both low- and high-coverage sequencing. This is the topic of a new paper in AJHG which creates a useful comparison benchmark for the performance of the two types of sequencing methods. High-coverage sequencing may be needed for things like disease studies (because deleterious alleles tend to be low-frequency), or the study of recent human demography (because recent population growth has resulted in an abundance of low-frequency SNPs that have not had enough time to reach a high population frequency yet).

AJHG dx.doi.org/10.1016/j.ajhg.2012.12.005

Deep Whole-Genome Sequencing of 100 Southeast Asian Malays

Lai-Ping Wong et al.


Whole-genome sequencing across multiple samples in a population provides an unprecedented opportunity for comprehensively characterizing the polymorphic variants in the population. Although the 1000 Genomes Project (1KGP) has offered brief insights into the value of population-level sequencing, the low coverage has compromised the ability to confidently detect rare and low-frequency variants. In addition, the composition of populations in the 1KGP is not complete, despite the fact that the study design has been extended to more than 2,500 samples from more than 20 population groups. The Malays are one of the Austronesian groups predominantly present in Southeast Asia and Oceania, and the Singapore Sequencing Malay Project (SSMP) aims to perform deep whole-genome sequencing of 100 healthy Malays. By sequencing at a minimum of 30? coverage, we have illustrated the higher sensitivity at detecting low-frequency and rare variants and the ability to investigate the presence of hotspots of functional mutations. Compared to the low-pass sequencing in the 1KGP, the deeper coverage allows more functional variants to be identified for each person. A comparison of the fidelity of genotype imputation of Malays indicated that a population-specific reference panel, such as the SSMP, outperforms a cosmopolitan panel with larger number of individuals for common SNPs. For lower-frequency (less than 5%) markers, a larger number of individuals might have to be whole-genome sequenced so that the accuracy currently afforded by the 1KGP can be achieved. The SSMP data are expected to be the benchmark for evaluating the value of deep population-level sequencing versus low-pass sequencing, especially in populations that are poorly represented in population-genetics studies.



Link

November 29, 2012

South Indian Y chromosomes (+ a little complaining about methods)

The table of haplogroup frequencies (left) may prove quite useful, but I am fairly disappointed with what appears to be the state of the art in recent published research on Y chromosome variation. This is not to belittle the tremendous amount of labor and money needed to collect and genotype large representative samples of individuals; only to express hope that better use of the collected samples could be achieved.

First of all, it is inconceivable to me how scientists can continue to use the 3x slower "evolutionary mutation rate" for their analyses of Y-chromosome ages on the basis of Y-STR markers. I have done my small part in my Y-STR series to show that this mutation rate is applicable only for a rather specific demographic history, and completely unsuitable to real growing human populations where Y-STR variance accumulates at close to the genealogical rate. And, my observations merely elaborated quantitatively what was already present in Zhivotovsky et al. (2006) but has been completely ignored since:
In simulations of a neutral process with average rate of increase m = 1, the number of surviving haplogroups rapidly decreased with time and corresponded well with the theory of mutant survival (Li 1955, p. 242), and the average size of the surviving haplogroups increased each generation by a value rapidly approaching 0.5 (data not shown), which agrees with asymptotic fraction of 2/t of haplotypes that survive at generation t (Athreya and Ney 1972, p. 19). The accumulated variance increased almost linearly (fig. 1), at a rate of increase about 0.00028 per generation; that is, the actual rate of accumulation microsatellite variation was about 3.6 times less than that predicted from the germ line mutation rate. This corresponds perfectly to the 3- to 4-fold difference observed between germ line and evolutionarily effective mutation rate.
The issue is all but resolved in the amateur "genetic genealogy" community, but even professional geneticists often use either genealogical or evolutionary rate, or take an agnostic stance by reporting results based on both rates. To arrive at strong conclusions about a topic on the basis of a mutation rate that is, to say the least, controversial, without even acknowledging the existence of a controversy is unsatisfactory. Y-chromosome researchers ought to copy the attitude of those working with autosomal DNA, where a corresponding mutation rate controversy was not swept under the carpet, but acknowledged (e.g., in the recent Meyer et al. high-coverage Denisova paper), with the implications of the uncertainty during the present "transitional" period quantified in the form of wider confidence intervals.

This "mutation rate" issue  notwithstanding, it was also recently shown that by Busby et al. that Y-STR based estimates have a dependence on the set of Y-STRs used, with markers exhibiting linear behavior across different time spans. This does not invalidate their use as molecular clocks, but highlights the need to not only select a bunch of Y-STRs, but also either (i) demonstrate that the selected set exhibits linear behavior for the time span of interest, or (ii) correct for deviations from linearity. Again, this type of modelling of microsatellite behavior was recently achieved for autosomal STRs by Sun et al.  Note that such deviations result in a slower rate than the genealogical one, but the mechanism whereby this is produced is completely different than the one proposed by Zhivotovsky et al.: it is not drift in a non-growing (m=1) population that reduces the effective rate, but rather "saturation" of the mutation process, whereby the variance at fast-mutating markers grows sub-linearly with time, because of physical constraints on their possible range of values.

I don't hope that Y-STR based age estimation will have much to offer in the coming years. But the third set of the 1000 Genomes Project is on its way, and this will include a variety of South Asian samples. Very soon we will be in a good position to study the time depth of common ancestry between e.g., European and South Asian Y-chromosomes within various haplogroups using point mutations, and these are not plagued by many of the problems associated with Y-STR variation and its interpretation.

Finally, I can't help but notice that this paper has not acknowledged the tremendous progress in resolving the Y chromosome phylogeny done by non-academic researchers. With the current state of our knowledge, the claim that haplogroup R1a1 is "autochthonous" in India is not tenable. Even if one discounts all the evidence made by SNP discoveries in the commercial testing world (and why should they?), finer-scale structure within this haplogroup has now been officially published and appears to be inconsistent with a South Asian origin of this haplogroup.

Certainly, not all is resolved; for example, the representation of tribal populations in commercial DNA testing is almost non-existent, and a sampling of their Y-SNP diversity is urgently needed. A very useful paradigm of research is that of recent work on the most basal clade of the Y-chromosome phylogeny (A00) in which the identification of very unique Y-chromosomes by genetic genealogists was combined with academic samples of "indigenous" peoples to produce new knowledge.

Much of population genetic research will benefit from such consilience between academics and amateurs. This is not an idle hope, but a recognition that this field is one in which the public not only has a substantial interest but can also do something about it. Many might be interested in Mars exploration, but without Elon Musk's bank account, most are consigned to being consumers of information about the Red Planet. Hopefully, better ways of combining the efforts of research scientists and the educated public can be identified and used in the near future.

PLoS ONE 7(11): e50269. doi:10.1371/journal.pone.0050269

Population Differentiation of Southern Indian Male Lineages Correlates with Agricultural Expansions Predating the Caste System

GaneshPrasad ArunKumar et al.

Previous studies that pooled Indian populations from a wide variety of geographical locations, have obtained contradictory conclusions about the processes of the establishment of the Varna caste system and its genetic impact on the origins and demographic histories of Indian populations. To further investigate these questions we took advantage that both Y chromosome and caste designation are paternally inherited, and genotyped 1,680 Y chromosomes representing 12 tribal and 19 non-tribal (caste) endogamous populations from the predominantly Dravidian-speaking Tamil Nadu state in the southernmost part of India. Tribes and castes were both characterized by an overwhelming proportion of putatively Indian autochthonous Y-chromosomal haplogroups (H-M69, F-M89, R1a1-M17, L1-M27, R2-M124, and C5-M356; 81% combined) with a shared genetic heritage dating back to the late Pleistocene (10–30 Kya), suggesting that more recent Holocene migrations from western Eurasia contributed less than 20% of the male lineages. We found strong evidence for genetic structure, associated primarily with the current mode of subsistence. Coalescence analysis suggested that the social stratification was established 4–6 Kya and there was little admixture during the last 3 Kya, implying a minimal genetic impact of the Varna (caste) system from the historically-documented Brahmin migrations into the area. In contrast, the overall Y-chromosomal patterns, the time depth of population diversifications and the period of differentiation were best explained by the emergence of agricultural technology in South Asia. These results highlight the utility of detailed local genetic studies within India, without prior assumptions about the importance of Varna rank status for population grouping, to obtain new insights into the relative influences of past demographic events for the population structure of the whole of modern India.

Link

November 26, 2012

Medieval signal of Swedish (?) admixture in Finland

I took the FIN (Finnish), GBR (British), and CDX (Chinese Dai) samples of the 1000 Genomes Project, each of which has a sample size of 100 in order to investigate the signal of East-West Eurasian admixture in Finns. While neither Britons nor Dai could be imagine of having contributed to Finns directly, they ought to make useful proxies of a NW European population lacking recent East Eurasian ancestry, and an East Eurasian population lacking recent West Eurasian ancestry respectively.

In the following, I will assume a generation length of 29 years and a sample birthyear of 1980 as in previous experiments.

First, the 1-reference analysis of FIN using GBR produced an admixture proportion lower bound of 37.4 +/- 5.1 percent.

The corresponding analysis of FIN using CDX produced an admixture proportion lower bound of 4.4 +/- 1.0 percent.

The 2-ref admixture test with {GBR,CDX} reported success:

Test SUCCEEDS (z=2.76, p=0.0057) for FIN with {GBR, CDX} weights
But, the decay rates were inconsistent, a situation which might occur when major admixture from different sources took place at different times. In particular, the one using CDX corresponded to 65.57 +/- 8.36 generations, and the one using GBR to 25.48 +/- 4.93 generations.

In calendar dates, Finns are estimated to have mixed with an East Eurasian CDX-like population between 170BC-320AD and with a NW European GBR-like population between 1100-1380AD.

The central date of the latter estimate is 1,240AD, which corresponds quite closely to the beginning of Swedish rule and is in the middle of the 13th. century, between the time when Finland was initially claimed for western Christendom (12th c.) and the time when the conflict between Sweden and Russia was settled (14th c.).

November 09, 2012

Multiway Admixture Deconvolution with MULTIMIX

The software will appear here. This has already been used in the recent 1000 Genomes paper. Below is the analysis of the MEX data:



I ran a small CEU/YRI/MEX K=3 analysis using ADMIXTURE on 30 random individuals from each population.
Notice that ADMIXTURE assigns 100% "American" ancestry to the most Amerindian-admixed individuals. MULTIMIX, on the other hand, has correctly not assigned any Mexicans 100% to the Amerindian component, because it makes use of LD to infer the ancestry of individual segments.

A couple advantages of the new method is that it is not limited to two ancestral populations, and does not require phased data as input, although phased data may provide some accuracy benefit, if available.

I'm eager to try the new software when it becomes available. I am not sure how it will scale (CPU/Memory-wise) with more individuals/components, so it'll be fun to experiment with.

Genet Epidemiol. 2012 Nov 7. doi: 10.1002/gepi.21692. [Epub ahead of print]

Multiway Admixture Deconvolution Using Phased or Unphased Ancestral Panels. 

Churchhouse C, Marchini J.

Abstract

We describe a novel method for inferring the local ancestry of admixed individuals from dense genome-wide single nucleotide polymorphism data. The method, called MULTIMIX, allows multiple source populations, models population linkage disequilibrium between markers and is applicable to datasets in which the sample and source populations are either phased or unphased. The model is based upon a hidden Markov model of switches in ancestry between consecutive windows of loci. We model the observed haplotypes within each window using a multivariate normal distribution with parameters estimated from the ancestral panels. We present three methods to fit the model-Markov chain Monte Carlo sampling, the Expectation Maximization algorithm, and a Classification Expectation Maximization algorithm. The performance of our method on individuals simulated to be admixed with European and West African ancestry shows it to be comparable to HAPMIX, the ancestry calls of the two methods agreeing at 99.26% of loci across the three parameter groups. In addition to it being faster than HAPMIX, it is also found to perform well over a range of extent of admixture in a simulation involving three ancestral populations. In an analysis of real data, we estimate the contribution of European, West African and Native American ancestry to each locus in the Mexican samples of HapMap, giving estimates of ancestral proportions that are consistent with those previously reported.

Link

October 31, 2012

A thousand (and ninety two) genomes

This is an open access paper describing the phase1 data of the 1000 Genomes Project. There is plenty of interest in the paper and supplement, but look at Figure S8 (left). This indicates the median shared haplotype length around f2 sites, i.e., sites where the variant exists twice in the sample and hence it makes sense to speak of shared length.

The maximum such length is for FIN (Finns) at 140kb, but it seems fairly obvious visually that the lowest sharing is found in multi-origin populations from the Americas (MXL, CLM, PUR, AS), in which segments are probably "interrupted" because of admixture. African populations (LWK/YRI) also tend to have low sharing, followed by Europeans, and East Asians.

There are little details in evidence: for example, IBS sharing with Luhya (LWK) seems higher than the European average, consistent with some level of African admixture in Spain, that has probably contributed some African haplotypes.

There seems to be a hint of an excess of sharing between Japanese (JPT) and Luhya (LWK). I have to wonder whether this might have something to do with Y-haplogroup D which links the Japanese with African Y-haplogroup E bearers. An excess of sharing between CHS (Singapore Chinese) and PUR (Puerto Rico) also seems to be suggested, for which I have no good hypothesis.

Nature 491, 56–65 (01 November 2012) doi:10.1038/nature11632

An integrated map of genetic variation from 1,092 human genomes

The 1000 Genomes Project Consortium

By characterizing the geographic and functional spectrum of human genetic variation, the 1000 Genomes Project aims to build a resource to help to understand the genetic contribution to disease. Here we describe the genomes of 1,092 individuals from 14 populations, constructed using a combination of low-coverage whole-genome and exome sequencing. By developing methods to integrate information across several algorithms and diverse data sources, we provide a validated haplotype map of 38|[thinsp]|million single nucleotide polymorphisms, 1.4|[thinsp]|million short insertions and deletions, and more than 14,000 larger deletions. We show that individuals from different populations carry different profiles of rare and common variants, and that low-frequency variants show substantial geographic differentiation, which is further increased by the action of purifying selection. We show that evolutionary conservation and coding consequence are key determinants of the strength of purifying selection, that rare-variant load varies substantially across biological pathways, and that each individual contains hundreds of rare non-coding variants at conserved sites, such as motif-disrupting changes in transcription-factor-binding sites. This resource, which captures up to 98% of accessible single nucleotide polymorphisms at a frequency of 1% in related populations, enables analysis of common and low-frequency variants in individuals from diverse, including admixed, populations.

Link

October 13, 2012

An estimate of the admixture time for Finns

Using a similar procedure as in my recent post on the Baltic (Update II), I used 15 FIN individuals from the 1000 Genomes together with 12 Nganasans from Rasmussen et al. (2010) as reference populations, and 15 other FIN individuals to estimate admixture LD in a rolloff analysis. Three outlier Nganasan individuals (GSM558800, GSM558802, GSM558807) were removed.
The estimated time of admixture is 86.095 +/- 10.187 generations, or 2500 +/- 300 years. It corresponds rather well to the beginning of the Iron Age in northern Europe.

As I mention in my previous post, there is evidence for intrusive cultures (Battle Axe and Seima Turbino) converging on the area from different directions during the preceding Bronze Age. If the above date is accurate, it will suggest a rather late admixture event between the Europeoid and Siberian elements of Finns. The former may have included both the descendants of Mesolithic European hunter-gatherers and intruders from Central Europe (Corded Ware/Battle Axe); the latter may have included both Comb Ceramic and the descendants of the Seima Turbino metallurgists.

September 11, 2012

West Asian and North European admixture in Basques and Indo-Europeans

In a previous post I showed that Basques are lacking in the West Asian admixture present in all their West European Indo-European neighbors, consistent with my theory of a late Indo-European invasion of Europe whose ultimate source was the highlands of West Asia.

But, there are alternative theories, one of which purports that the Proto-Indo-Europeans were northern Europeoid pastoralists from the eastern European steppe. Since the North_European ancestral component is lacking in the Tyrolean Iceman and Gok4, the TRB Swede, it is conceivable that North_European bearing populations introduced this component during the Indo-European invasion.

Of course, there is absolutely no archaeological evidence for a massive migration out of the steppe into Europe, as even the main proponents of the steppe hypothesis accept, and as physical anthropology makes clear. And, we don't have to invoke an eastern European invasion to explain the North_European component, since it was present among pre-Indo-European hunter-gatherers from both Gotland and Iberia, as two ancient DNA studies have shown.

In any case, I took the HGDP and 1000Genomes European populations, together with the West_Asian and North_European Dodecad components, and calculated f3 statistics of the form:

f3(IE; Basque, Dodecad Component)

where Basque is either HGDP French_Basque or 1000 Genomes Pais_Vasco_1KG, and Dodecad Component is either West_Asian or North_European.

All the results can be found in the spreadsheet.

Again, there is evidence of West Asian+Basque admixture in all Indo-Europeans (|Z| less than -3) except the islanders from Canarias and Orkney, and the Russians; in the latter case, Basques are probably a poor stand-in for their pre-Indo-European ancestry. So, 32 of 38 comparisons are significant.

One would expect such negative f3 statistics to also apply in the North European+Basque case. After all, there are historically known migrations of Northern Europeoids into Western Europe (both Celts as well as Germanics) which did not affect Basques linguistically; moreover, Basques are southern Europeans, and many of the tested populations are northern Europeans, who are expected to turn up as mixtures of North European+Basque. However, a total of 16 of 38 comparisons are significant, involving, as expected mostly northern European populations.

It thus appears that geography and recent history is sufficient to explain the excess of North_European in some populations. Despite having a dataset with an excess of Iberian and North European populations, not many significant f3 statistics appear, and these are mostly as expected.

In conclusion, by comparing Basques vs. Indo-Europeans there appears no good evidence for the theory that Indo-European languages were brought into western Europe by a massive migration of northern Europeoids from eastern Europe. Basques do not appear distinctive in terms of the North_European component, but they do appear distinctive in terms of the West_Asian one.

This confirms previous ADMIXTURE analyses that Basques occupy an "intermediate" position along the north-south axis of variation in Europe, and an absolutely terminal one in terms of the West Asian component.

It is very interesting that ancient DNA research has provided clues about a very "uneven" landscape of prehistoric Europe, with Sardinian-like farmers in Sweden and Northern European-like hunter-gatherers in Iberia and very little in-betweens. But, these two elements eventually did mix, and, with the addition of a new group of people emanating from the highlands of West Asia, acquired their Indo-European speech, and went on to become the living nations of Europe.

Much remains to be discovered: the first ancient DNA traces of the constituent elements must be identified in space and time, and the history of their intermixture must be tracked.

September 04, 2012

Progress in the Y-chromosome phylogeny

There is much of interest in the abstracts of the DNA in Forensics 2012 conference which will start very shortly, including news on the Tyrolean Iceman (he belonged to G-L91; and G2a "featured unexpectedly high densities within or near the Ötztal Alps."), speculations about a trans-Pacific spread of a lineage found in South Americans, and many other topics besides.

But, for me, the most interesting abstracts relate new developments in the Y-chromosome phylogeny world. The titles of the 3 abstracts are:

Y-chromosomal insights from large-scale resequencing

A calibrated human Y-chromosomal phylogeny based on resequencing
Insight into human Y chromosome variation from low-coverage whole-genome resequencing

and, they all seem to be from authors working at The Wellcome Trust Sanger Institute.

Researchers used 1000Genomes low coverage data (2x) and high coverage Complete Genomics data to untangle the Y chromosome phylogeny. As expected, the 1000Genomes data weer of poorer quality, and had a large number (~14-17%) of false negatives, i.e., SNPs that were actually present in the samples were not discovered. 

I list the main findings from the three abstracts:
The TMRCA of the entire tree was ~115 KYA (thousand years ago), and of the lineagesoutside Africa ~60 KYA, both as expected. Additional insights included a rapid expansion of hg F~40 KYA, and of R1b in Europe ~5-10 KYA. The archaeological counterpart of the former isunclear, but the latter is likely to represent a Neolithic expansion of this lineage

The GENETREE TMRCA for the complete set of chromosomes examined was 105-125KYA; times for the out-of-Africa movement were 62-79 KYA, a Paleolithic expansion 37-48KYA, and the expansion of R1b in Europe 7-10 KYA; rho times were broadly similar.

It confirmed Hg E (Bantu), O (China) and R1b (Europe) expansions associated with the Neolithic transitions in different parts of the world, and revealed that the expansion in Europe was the most extreme. One novel finding was a striking expansion of lineages F to R ~20 thousand years after the out-of-Africa movement, suggesting a previously unknown event of importance to male demography at this time.
I would say that these results are consistent with my "two deserts" theory and the climatic history of Africa and the Near East. Of course, I don't think there was a 60ky Out-of-Africa event, for a number of different reasons that I've written about to death. With respect to Y chromosome phylogeny, it is important to highlight one more time where I'm coming from:

The major African Y-haplogroup E belongs to the DE subclade of the CT clade:



I have color-coded the Eurasian lineages as "green", and the African ones as "red". Now, those who think that the age of CT corresponds to Out-of-Africa believe that this event was accompanied by a massive bottleneck which is responsible for the reduced genetic diversity of Eurasians compared to Africans.

But, the question is obvious: if there was such a massive bottleneck in Eurasian ancestors, then how come it is the Eurasians (the bottlenecked population) that ended up with most of the CT descendants?

There really is no archaeology to support a 62-79ky Out-of-Africa, the only archaeology (and anthropology) in support of Out-of-Africa relates to the pre-100ky period, with things like the Nubian complex, the Mt. Carmel hominins, Jebel Faya, and others links between Africa and the Near East.

There are no genetics to support it either: track every paper that has argued for ~60ky Out-of-Africa, and you will invariably find a 2.5x10-8/bp/generation or similar mutation rate and/or a recent human-chimp calibration hiding somewhere in the details. While the mutation wars rage, it is not certain how they will be resolved, but I would put money on the true mutation rate ending up much lower than the one dominating the literature, and, consequently, Out-of-Africa being much earlier.

Getting back to the topic, the "striking expansion of lineages F to R ~20 thousand years after the out-of-Africa movement" corresponds to the UP Revolution in west Eurasia. So, to recapitulate my thinking in bullet form:

  1. Pre-100ky Out-of-North Africa (Mt. Carmel, Nubian, Jebel Faya?)
  2. c. 70ky climate crisis in North Africa-Arabia. Reduction of Y-chromosome diversity: CT founder.
  3. 70-50ky. Modern human biocultural evolution accelerates as they (i) face climate crisis, (ii) face new environments as they move out of North Africa-Arabia, (iii) face archaic humans in Eurasia and Africa.   Haplogroup DE is group of "southern"  Out-of-Arabians heading east (D) or west (E); Haplogroup CF is group of "northern" Out-of-Arabians, some of which head east (C) or stay around (F).
  4. 50-40ky. Culmination of the process leads to UP/LSA Revolution: 
  • In East: some F descendants come to dominate over the early D and C settlers
  • In West Eurasia: other F descendants break through the Neandertal bottlecap and invade Europe with UP technologies
  • In Africa: E descendants (descended from DE back-migrants) kick-start the Lower Stone Age.
.

August 23, 2012

Or, maybe they speciated 3.7-6.6Ma ago? (Sun et al. 2012)

This has certainly been an eventful August in human origins research; if the Neandertal Wars weren't enough, a different issue that had simmered for a while now, the human autosomal sequence mutation rate, has now come to a full boil.

A couple of weeks ago, Langergraber et al. (2012) came out, and combined direct measurement of generation lengths in humans and other primates with the directly measured human autosomal sequence mutation rate to argue for an old 7-13Ma divergence between Pan and Homo.

Yesterday, Kong et al. (2012) independently derived a low direct mutation rate of 1.2x10^-8, and added the observation that older human fathers pass on more mutations to their offspring than younger ones. As I point out in my post on the topic, this has implications for the Homo-Pan divergence as well: if chimp dads are younger than human dads, they will tend to pass fewer mutations to their offspring. Thus, the chimp mutation rate (/generation) might be lower rather than equal to the human one, and this might push the speciation time even further back in time.

Today, a new paper has appeared in Nature Genetics which argues for an "intermediate" rate between  the direct ~1-1.3x10^-8 rate and the widely used 2.5x10^-8 one: their rate estimate is: 1.4–2.3x10^-8 and the corresponding Human-Chimp speciation time is 3.7-6.6 million years ago. Kari Stefansson is a co-author of the new paper, as he is of the Kong et al. one, which estimated the mutation rate at 1.2x10^-8.

The new paper builds what appears to be a very exhaustive model of microsatellite mutation:
Microsatellites have been widely used to make inferences about evolutionary history. However, the accuracy of these inferences has been limited by a poor understanding of the mutation process. We developed a new model of microsatellite evolution (Supplementary Note). This model can estimate the time to the most recent common ancestor (TMRCA) of two samples at a microsatellite by taking into account (i) the dependence of the mutation rate on allele length and parental age (Fig. 2a,c); (ii) the step size of mutations (Fig. 2b); (iii) the size constraints on allele length (Fig. 2d and Supplementary Figs. 8 and 9); and (iv) the variation in generation interval over history. In contrast to the generalized stepwise mutation model (GSMM), which predicts a linear increase of average squared distance (ASD) between microsatellite alleles over time, the new model predicts a sublinear increase (Fig. 3) and saturation of the molecular clock, due to the constraints on allele length. We also extended the model to estimate the sequence mutation rate, using the per-nucleotide diversity flanking each microsatellite as an additional datum. To implement the model, we used a Bayesian hierarchical approach, first generating global parameters common to all loci, followed by locus-specific parameters and finally the microsatellite alleles at each locus (Online Methods). We used Markov chain Monte Carlo to infer TMRCA and sequence mutation rate. 
I haven't delved deeply into the details of how the sequence mutation rate (per nucleotide/per generation) can be derived by exploiting the microsatellite rate. But, why would the rate estimated with the new method be different than the directly measured one? The authors propose some ideas:
We hypothesize that the lower mutation rate estimates from the whole-genome sequencing studies might be due to (i) the limited number of mutations detected in these studies, which explains why their confidence intervals overlap ours, (ii) possible underestimation of the false negative rate in the whole-genome sequencing studies or (iii) variability in the mutation rate across individuals, such that a few families cannot provide a reliable estimate of the population-wide rate.  
 Apparently, the team behind Sun et al. became aware of the new Kong et al. after the paper was accepted, so they attached the following note at the end of it, as well as a discussion in the supplement:
Note added in proof: After this paper was accepted, another study35 was published that independently estimates the human sequence mutation rate, using a direct measurement in contrast to the indirect measurement we report here. In spite of some key similarities between our results and those of Kong et al.35 (the male-to-female mutation rate ratio and the absence of an effect of mother's age), they estimate a considerably stronger effect of father's age and an overall sequence mutation rate below the range we infer. The discrepancies in the sequence mutation rate may be in part due to the fact that Kong et al. focus on a more intensively filtered subset of the human genome than we analyze here, but other factors are also likely to be at work (Supplementary Note). As an initial attempt to compare the two studies in terms of their implications for evolutionary history, we ran the same Bayesian inference procedure we developed in this paper (integrating over uncertainty in unknown parameters), now using the sequence-based estimates rather than the microsatellite-based estimates as input (Supplementary Note). Notably, the inferred dates based on the measurement of the sequence mutation rate are older and no longer in direct conflict with the inference that S. tchadensis is on the human lineage since the split from chimpanzees. The sequence- and microsatellite-based data sets are very different, and an important direction for future research will be to understand why the direct sequence–based mutation rate estimate is lower than the one inferred on the basis of microsatellites. 
All this leaves me rather perplexed. I guess one take-home lesson from the debate would be to avoid making strong statements about the past that are dependent on a particular mutation rate. The following table from the supplementary material pretty much says it all:


Notice that the two estimates are approximately double one of the other. Personally, I tend to favor the older dates, since they might "match" better with key developments: Out-of-Africa will become pre-100ka and consistent with the appearance of the Nubian technocomplex in Arabia, which seems to be the only real solid evidence of Out-of-Africa in the archaeological record. It would also be consistent with the appearance of modern humans in the Levant c. 100ca at Mt. Carmel, the first clear evidence of Homo sapiens in Eurasia. Moreover, it would explain the early appearance of Neandertaloid features in the Atapuerca hominins at c. 600ka, long before the inferred split of modern humans from Neandertals when the slowest rate is used.

But, my confidence in these correspondences is low until the controversy is resolved one way or another. If the 1.8x10^-8 rate of this paper is closer to the truth, then my money would be on the false negative rate, i.e., full genome sequencing is systematically overlooking SNPs that exist in the genomes.

Apparently, now, we have three rates to contend with: (i) the Icelandic 1.2x10^-8 rate (and other similar rates, such as the 1.36x10^-8 one); the 2.5x10^-8 one that has been very widely used in the literature, and (iii) the "1.82x10^-8 mutations per base pair per generation (90% CI 1.40–2.28 × 10-8; Table 2)" from this paper. This may be disheartening, but all setbacks represent opportunities to learn something new, and now that the issue is out in the open, I'm sure that many "top dogs" will try to figure out what is going on.

Nature Genetics doi:10.1038/ng.2398

A direct characterization of human mutation based on microsatellites

James X Sun et al.

Mutations are the raw material of evolution but have been difficult to study directly. We report the largest study of new mutations to date, comprising 2,058 germline changes discovered by analyzing 85,289 Icelanders at 2,477 microsatellites. The paternal-to-maternal mutation rate ratio is 3.3, and the rate in fathers doubles from age 20 to 58, whereas there is no association with age in mothers. Longer microsatellite alleles are more mutagenic and tend to decrease in length, whereas the opposite is seen for shorter alleles. We use these empirical observations to build a model that we apply to individuals for whom we have both genome sequence and microsatellite data, allowing us to estimate key parameters of evolution without calibration to the fossil record. We infer that the sequence mutation rate is 1.4–2.3-10^-8 mutations per base pair per generation (90% credible interval) and that humanchimpanzee speciation occurred 3.7–6.6 million years ago.

Link

August 05, 2012

1000 Genomes Project Community meeting video

can be found here.

I fast forwarded through a few of the talks. Some interesting tidbits:

1) John Novembre reports on the human autosomal mutation rate on the basis of a very large sample sequenced in a small number of regions; he is using an approach that simultaneously estimates effective size and mutation rate: he gets a median of 1.38x10^-8 per bp per gen. This tends to agree with previously published pedigree estimates, and seems to also be quite lower than a widely used value of 2.5x10^-8 that assumed an ancestral effective population size and an age for the human-chimp divergence. Adoption of the slower rates has implications: if people start using the lower rates that come out of sequencing, population splits will probably have to be redated.

2) Jay Shendure speaks about the future of sequencing; we are now at around $3,500; he expects costs to continue to decline, but we shouldn't be overly optmistic due to technology limits+market forces. I think that the latter may be significant: when a company can offer a better deal than any of its competitors, it has no incentive to drop prices further, even if its own costs are dropping; moreover, demand is going to surge as prices approach the magical $1,000 mark: I know I'll be very tempted to buy at that point myself.

3) Alexander Platt talks about inferring population history using haplotypes. Of interest: during the last 1,000 generations there are more coalescences between Beijing Chinese and Japanese rather than Beijing Chinese and southern Chinese; in more recent times, there are more coalescences between Chinese groups. This makes some sense, if we suppose that -as seems likely- Mongoloids spread north-to-south across China during prehistory; the Japanese are thus linked -in older times- with northern Chinese, both of which are mostly descended from the northern Mongoloids; in more recent times, especially after the emergence of a uniquely Chinese polity and culture, the Chinese tend to marry other Chinese, hence they share more recent common ancestors within the country itself.

August 01, 2012

Let's play ASHG 2012 title imputation! (+open science miscellanea)

It's that time of year again, and the titles for the ASHG 2012 presentations have just been posted online. Well, part of them anyway; the dreaded (...) have made another appearance. Still it is fun to try to guess what each contribution is about, and we'll only have to wait ~1 month for the abstract text.

In related news, Ewen Callaway reports on the trend (?) for biologists to put their unpublished work in arXiv. My own views are strictly for open science, so I applaud the people who are dragging their disciplines into the 21st century.

Finally, recent initiatives in the UK and the EU will mandate open access for work funded by research agencies. This is a good step in the right direction, but a very incomplete one: open access solves the problem of ensuring wide dissemination of new science, but merely shifts the flow of public  money rather than sever it. With open access Government->University Library->Journal is replaced by Government->Research Agency->Scholar->Journal.

Moreover, open access does not address the more fundamental issue of how journals impede scientific progress by imposing the antiquated pre-publication peer review process. The sky hasn't fallen over the heads of physicists who post their work on arXiv when they're done with it and carry out post-arXiv publication peer review. So, it probably won't fall on the heads of biologists who do the same either.

A good example of this is the recent work on ChromoPainter/fineSTRUCTURE that appeared months before publication: lots of people -including myself- started using their software right away, which spurred new insight, and they got their peer-reviewed publication too. More recently, a group of independent researchers co-ordinated their efforts in public to hack 1000 Genomes data, discovered and validated new SNPs, and they got their publication too. Open science works, so everyone should try it!

July 31, 2012

On the age of Y-chromosome haplogroup R1b-M343

Continuing my exploration of the human Y-chromosome, on the basis of 1000 Genomes data, I turned my attention to Y-haplogroup R1b-M343, one of the most populous lineages in extant Europeans. A total of 109 Y-chromosomes possessed the mutation diagnostic of this haplogroup.

At first, I calculated the histogram of pairwise TMRCA between all these 109 Y-chromosomes:
Most times are around 6-7 thousand years ago, but there is an outlier bump at around 15 thousand years ago. To further investigate this bump, I carried out multidimensional scaling of the collection of Y-chromosomes:

It is clear that the group of high pairwise TMRCAs correspond to the individual on the left of the figure that emerges as a clear outlier vis a vis the rest. The ID of that individual is HG00640 (from PUR population). One possibility is that this individual is M343+ due to sequencing error and belongs to a different lineage altogether. However, HG00640 is also R1-S1+ and R1b1-L278+ but R1b1a-P297-.

It will appear therefore that the HG00640 Puerto Rican belongs to the R1b1-L278 clade, but not to the R1b1a-P297 subclade. He thus represents an earlier split from the tree than the R1b1a2-M269 (frequent in West Eurasia), as well as the R1b1a1-M73 (frequent in Central Asia). It seems that I have chanced upon a real relic Y-chromosome!

The estimate of the age difference between HG00640 and the remaining M343+ chromosomes that cluster on the right is: 15,426 years. We now have direct evidence that haplogroup R1b1 is quite old, and R1b-M343 itself must have emerged sometime between 23,657 years (the TMRCA of R1a vs. R1b) and 15,426 years.

This little exercise reinforces my belief, first expressed in the outliers article, that there are real relic Y chromosomes in the world today, and we neglect them at our own peril.

Most European and European-derived men from the 1000 Genomes Project who belong to the R1b-M343 clade share patrilineal descent within the last 7,000 years or so. But, not all of them do, and outliers like HG00640 can only be caught with very large worldwide sample sizes and full genome sequencing.


* * *

Addendum: There appears to be a R1b1(xP297) DNA Project. There appear to be a quite rich collection of men with SNP results similar to HG00640, including R1b1c-V88+ (as suggested by Roy King in the comments), but also of V88- individuals. I see great utility in such projects, because if one can detect very aberrant Y-STR haplotypes (which can be done with a simple histrogram or MDS plot, as in this post), then one can identify candidate Y chromosomes for full sequencing.

* * *

I have also assessed all PUR individuals with the world9 calculator. This can be found here. HG00640 does not appear to be unusual in terms of his overall genomic ancestry.

July 30, 2012

Estimating the age of Y-chromosome Adam (again)

UPDATE (August 1): The dates in this post have been superseded by the ones in Dates of major clades of the Y-chromosome phylogeny.

I have used the official phase1 chrY SNP data instead of the working data used in my first experiment. The histogram of pairwise TMRCA values looks much sharper now; not sure what the difference between the two datasets was:

In any case, the divergence of the most basal African clade is very evident here on the right, corresponding to an age for the human Y-MRCA of 159,298 years.

Also of interest are the other peaks in the distribution of pairwise TMRCAs which correspond to 6.5, 40.0, 66.3 thousand years. I think we are getting some good signals corresponding to Out-of-Arabia (66.3ky?) where a hyper-arid phase in Arabia may very well have caused a bottleneck in the population of modern humans, and full behavioral modernity/UP revolution (40.0ky?) where modern humans start turning up all over Eurasia, and even some Africans look like UP Europeans.

Perhaps, I'll spend some more time assigning the Y-chromosomes to haplogroups so that I can give a more complete estimate of the major clades of the Y-chromsome phylogeny.

UPDATE: Node Ages


I will add various node ages as I calculate them. Thanks to ISOGG for a convenient correspondence between haplogroups and SNP genetic positions.



Clade
Comparison
Age (years)
Notes
BT
(B-M247 vs. CT-M294)
 71,188
C1-C3
(C1-P122 vs. C3-Z1453)
 25,022
CF
(C-M130 vs. F-P137)
 47,379
CT
(DE-M145 vs. CF-P143)
 62,439
DE
(D-Page3 vs. E-P169 )
 62,205 
(***)
E
(E1-P147 vs. E2-M75)
 57,703 
(****)
E1b1
(E1b1a-V38 vs. E1b1b-M215)
 43,587
E1b1b1a1a-E1b1b1a1b
(E1b1b1a1a-V12 vs. E1b1b1a1b-V13)
 13,817
E1b1b1a1a-E1b1b1a1c
(E1b1b1a1a-V12 vs. E1b1b1a1c-V22)
 19,482
E1b1b1a1b-E1b1b1a1c
(E1b1b1a1b-V13 vs. E1b1b1a1c-V22)
 16,210
I
(I1-L75 vs. I2-L68)
  26,885
I2a1
(I2a1a-M26 vs. I2a1b-S328)
 19,513 
(*****)
IJ
(I1-L75 vs. J-L60)
 35,589
IJK
(IJ-P125 vs. K-P131)
 41,910
J
(J1-M267 vs. J2-M172)
 24,497 
(*)
J2
(J2a-L212 vs. J2b-M102)
 23,420
K-M9  
36,389
(x)
NO
(N-M231 vs. O-P191)
 32,467
O1-O2
(O1a-M119 vs. O2-M268)
 26,145
O1-O3
(O1a-M119 vs. O3-P198)
 26,303
O2-O3
(O2-M268 vs. O3-P198)
 27,870
O3a1-O3a2
(O3a1-L465 vs. O3a2-P201)
 18,765
P
(Q-M242 vs. R-P224)
 33,043
R1
(R1a-M420 vs. R1b-M343)
 23,657 
(**)
R1b1a2a1a1a5
(R1b1a2a1a1a5a-Z156 vs. R1b1a2a1a1a5b-Z301)
 6,476



(*) This is, strictly speaking, the common ancestor of J1 and J2, since J*(xJ1, J2) chromosomes have also been observed.
(**) This is also the common ancestor of R1a and R1b, since R1* chromosomes have also been observed
(***) Paragroup DE* chromosomes have also been observed
(****)  Paragroup E* chromosomes have also been observed, so this is, strictly speaking, the common ancestor of E1 and E2; this is only a small underestimate, given that the DE node (62,205 years) is only marginally older
(*****) There is also the paragroup I2a1* and I2a1c

(x) Due to absence of published bifurcating structure within K-M9, I estimated this using the same model-free method used for Y-MRCA. This result in an "older peak" of pairwise K-M9 TMRCAs of 36,389 years. This seems appropriate as it lies between IJK (41,910 years) and P (33,043 years)

July 29, 2012

Estimating the age of Y chromosome Adam

I took chrY SNP data from the 1000Genomes Project, with the goal of estimating the age of the root of the Y chromosome phylogeny. I used the mutation rate of 3x10^-8 mutations/nucleotide/generation of Xue et al. (2009), and a generation length of 25 years.

I decided to use a rather unconventional model-free approach, not building a phylogeny, but simply observing the distribution of pairwise TMRCA between all 526 Y chromosomes. A histogram thereof can be seen below:

My assumption is that the most basal split in the tree will correspond to the local peak on the right. I used MCLUST to cluster the observations, and, as is visually apparent, the right-most local peak forms a cluster with a mean of ~184 thousand years. This is a bit older than the 142 thousand year old date of Cruciani et al. (2011) which was based on a 200kb region of the MSY.

In any case, a re-rooting of the Y chromosome phylogeny is apparently in the works, so we're bound to know much more in the near future; we are long overdue for a systematic revision of node split times using the SNP counting method.

It is also imperative that African outliers in Eurasia be better studied to determine whether they have shallow or deep divergences with their African cousins. Much will depend on this, since the argument for recent Out-of-Africa crucially depends on Eurasia harboring a subset of African variation, a proposition which relies on the dismissal of apparently African outliers in Eurasia as recent migrants rather than relics.

UPDATE (July 30): I have now updated my age estimate using a newer cleaner version of the 1000 Genomes chrY data.

July 25, 2012

R1b1a2 variants in 1000G data

It's nice to see a group of independent researchers documenting their work using 1000 Genomes data. I've been following on-and-off developments in this field, and I have to say that it requires deep commitment from the persons involved to keep a mental picture of the ever-deepening phylogeny.

But, in a sense, that is what's great about the efforts of citizen scientists tackling a scientific problem: they are deeply invested in understanding their little part of the human Y-chromosome phylogeny, because it's their part and every SNP discovery in it represents a small victory. Thus, they can expend the time and effort to push the discovery process to its technological limits.

It's only too sad that this can at present be done only using 1000 Genomes data, a.k.a. the global collection of full genome data that completely ignores the part of the world between  Italy and China/India. Hopefully, sometime in the future, the ever-better-understood twig of R1b1a2 will be placed within its wider Eurasian context.

PLoS ONE 7(7): e41634. doi:10.1371/journal.pone.0041634


Discovery of Western European  R1b1a2 Y Chromosome Variants in 1000 Genomes Project Data: An Online Community Approach

Richard A. Rocca1*, Gregory Magoon2, David F. Reynolds3, Thomas Krahn4, Vincent O. Tilroe5, Peter M. Op den Velde Boots6, Andrew J. Grierson7

The authors have used an online community approach, and tools that were readily available via the Internet, to discover genealogically and therefore phylogenetically relevant Y-chromosome polymorphisms within core haplogroup R1b1a2-L11/S127 (rs9786076). Presented here is the analysis of 135 unrelated L11 derived samples from the 1000 Genomes Project. We were able to discover new variants and build a much more complex phylogenetic relationship for L11 sub-clades. Many of the variants were further validated using PCR amplification and Sanger sequencing. The identification of these new variants will help further the understanding of population history including patrilineal migrations in Western and Central Europe where R1b1a2 is the most frequent haplogroup. The fine-grained phylogenetic tree we present here will also help to refine historical genetic dating studies. Our findings demonstrate the power of citizen science for analysis of whole genome sequence data.

Link

May 02, 2012

Drawing the human Y chromosome tree with SNPs

Terry, (tdrobb@gmail.com), a poster at GENEALOGY-DNA-L reported age estimates for various nodes of the Y-chromosome tree based on SNPs. These can be found in this PDF file and here (scroll down for UPDATE10). He used 1000 Genomes data and SNP counting to reach these estimates.

It will be nice to see others join in on the SNP bandwagon, because that is really the way forward in age estimation for Y-chromosome lineages. SNPs have an extremely low (=negligible) rate of back-mutation, but they occur at a much lower rate than Y-STR step mutations. On the other hand, there are at most a few hundred Y-STRs and only ~100 tested by commercial companies, while scientific datasets generally include at most a few dozen of them. The Y chromosome includes millions of mutable sites and these will be generally reported both by the 1000 Genomes Project, and the plethora of full genome sequences that is about to become available.

Y-SNP based age estimation has the potential of greatly improving estimates by tightening confidence intervals substantially; there will, of course, be lingering uncertainty of parameters such as generation length, but Y chromosome mutation rates are likely to become very secure once full genome sequencing becomes so cheap that it can be applied to a number of father-son pairs.

Looking at the inferred tree, what is striking is the great distance between haplogroup A1b and the rest of the tree, or about 100,000 years. Note that these are not "relative" estimates as were published by the 1000 Genomes Project (based on "archaeologically" calibrating a node and estimating ages of other nodes by counting the relative number of SNPs), but "absolute" ones (dividing SNPs with a mutation rate).

(UPDATE: There is apparently an even more basal clade than A1b currently investigated; I have removed the link to an announcement regarding this clade, since there are issues regarding the release of this information)

Going back to the age estimates, I cannot help but notice the concordance between Terry's age estimates for DE/CF split (55ky) with the mtDNA estimates for most mtDNA L3 subclades. Terry labels DE "African" and CF "Eurasian", but, in fact DE is Afrasian and "CF" Eurasian. Together with the absence of any evidence for a post-70ka Out-of-Africa, I'd say that it is becoming increasingly clear that while modern humans can be ultimately traced to the Middle Stone Age in Africa, their major expansion that went on to colonize the entire world originated in Asia, and included a major episode of back-migration into Africa.

I also earnestly hope that the next set of Y chromosome papers on recent populations will forego the cost of testing hundreds of samples on Y-STRs and invest in full Y-chromosome sequencing of a few samples after an initial Y-SNP screening.

February 09, 2012

John Hawks on Neandertal similarity

In Which population in the 1000 Genomes Project samples has the most Neandertal similarity? John Hawks publishes some figures showing how different populations from the 1000 Genomes Project are more/less similar to the Vindija Neandertal sequence. My most recent thoughts on the topic on Neandertal admixture can be found here.

The most important findings are the following:

First, Tuscans appear to have more shared derived variants with Vindija than Brits do. I can think of two explanations for this finding:
  • Higher genetic diversity in southern Europe compared to northern Europe; if these shared variants occurred at a low frequency in southern Europe vs. northern Europe to begin with, then genetic drift would have driven more of them to extinction in the north than in the south.
  • Possibility that the Vindija sequence has modern human admixture from a population of early Southern Europeans that would have contributed more to the gene pool of Tuscans (who are geographically quite close to Croatia) than to Brits


Second, North Chinese appear to be more similar to Neandertals than South Chinese are. 
  • I am not entirely sure whether genetic drift would work in this case, since presumably CHB is representative of a larger Chinese population than Chinese settlers in Singapore.
  • A different explanation is that south Chinese may possess more admixture from an Asian type of erectus-like population that predates the common ancestor of modern humans and Neandertals, who almost certainly lived in the western end of Eurafrasian region.

Third, the Luhya from Kenya appear to have less Neandertal similarity than the Yoruba from Nigeria. This, of course, makes very little sense if Neandertal similarity can only be attributed to Neandertal admixture, since there were no Neandertals in Africa at all. Moreover, while the Luhya live in Kenya, their ultimate origins are further west, since they are a Bantu group, although they show mixed West/East African affiliations.

There are two alternatives:
  • Modern humans did originate in North-West Africa as the Y-chromosome phylogeny suggests, and East Africans possess some distinctive Palaeo-African ancestry from before the sapiens-Neandertal common ancestor.
  • Back-migration from Eurasia that affected different populations to different degrees; it is my impression, for example, that the Yoruba are almost completely a Y-haplogroup E population, but if anyone has any comparative results at hand, feel free to comment.
All in all, it will be interesting to see how the Vindija (and now Denisova) sequence compares to modern humans. The original papers could only compare against a handful of modern humans, but the availability of many full genomes from around the world will add statistical power to the comparisons.

January 16, 2012

Phased Omni haplotypes with ShapeIT

The working directory of the 1000 Genomes ftp site contains phased haplotypes for 2,123 individuals from the 1000 Genomes Project (US/Europe). The data were phased with ShapeIT, which I've recently played with, and could recommend as a fairly user-friendly and high quality phasing software.

You can use vcftools to convert the data into PLINK format, which appears to be quite efficient (but use --plink-tped) compared to doing it on the single file I previously linked to. So, it's also a way to get 1000Genomes data into the more useful PLINK format, and it's pre-phased as a bonus.