Showing posts with label Genomes Unzipped. Show all posts
Showing posts with label Genomes Unzipped. Show all posts

December 13, 2012

Daniel MacArthur's chromosome 10

In a previous post I confirmed Razib's observation on the South Asian ancestry of part of Dan MacArthur's chromosome 10. Razib has a new post up in which he argues that this type of ancestry is not of Romani origin. I figured it was my turn to continue the series of gratis ancestry analysis, with two goals in mind: (i) to let people know that if you put your genome online, you may find interesting things about it, and (ii) to indeed figure out what is going on with this interesting case of unexpected admixture.

I used Beagle/fastIBD with default parameters and with the HapMap recombination map to figure out the mean IBD sharing between Dan's chr10 and a number of different populations, including most of my available South Asian references. I also included YRI as an appropriate outgroup, as well as CEU30 and the 1000 Genomes British populations, given that Dr. MacArthur is Australian and has a Scottish surname, so it's a good bet that he has plenty of British-type ancestry.

The mean chr10 sharing is plotted below:



Now, it's cool that the top-2 matches are Argyll and Orkney, both of which are part of Scotland. But, what is interesting, is that North_Kannadi squeezes in ahead of CEU, with a very respectable mean of 1.4cM, and a number of other Indian populations are not far behind, while most of the ones from Pakistan are not. I'd say this looks consistent with an "Indian" origin of this type of ancestry.

A useful control is to repeat this experiment with a different chromosome, the similar-sized chr9 which lacks evidence for South Asian ancestry:

It is now visually clear that the difference between the British populations and the South Asian ones is greatly diminished. And, while the North Kannadi were near the top of the order for chr10, they are near the bottom for chr9, even lower than the YRI outgroup at the "noise left end" of the spectrum. Even the highest ranked South Asian population has about 1/3 estimated IBD sharing as the British ones. 

Moreover, whereas in chr10 the top-ranked South Asian populations were often from the south, in chr9 the situation is reversed, with most of the top-ranked ones being from Pakistan. Again, this suggests that there is no real South Asian admixture here, but just some low-level sharing with the (more West Eurasian) populations of the northern part of South Asia.

So, to make a long story short, it does look to me like an excellent suggestion that there is some type of peninsular or even south Indian ancestry in the chr10 in question.

November 01, 2010

Joe Pickrell redux

Joe Pickrell discovers Jewish great-grandparent
Second, Dienekes followed up on his analysis of the ancestry of the GNZ participants with a much larger data set, including individuals of southwest European descent. As expected, when including more data, there was no evidence that Vincent has any Ashkenazi ancestry. Unexpectedly, this was not true for me—even in this larger analysis, the evidence for Ashkenazi ancestry didn’t disappear.

...

As I was mulling over these sorts of issues, I sent the link to my previous analysis to a family member. I didn’t really expect this person to find it that interesting, but hey, you never know. I then got a phone call. I’ll summarize a couple days worth of moderate confusion, second-hand reports of conversations with distant relatives, and family intrigue with this: as it turns out, one of my great-grandparents was indeed a Polish Ashkenazi Jew who immigrated to the United States around the turn of the century. I, obviously, was completely unaware of this.
So to conclude, a tip of my hat to Dienekes and everyone else who looked at these data—this has been the first genuinely unexpected thing to come out of my genetic data.
I've estimated Joe's ancestry here and here.

He is included in the Dodecad Project's spreadsheet as JKP001.

His "Southwest Asian" score of 6.7% is consistent with but not indicative of Jewish ancestry, as this component is found in West Asia and Europe, although it attains its maximum in Saudi Arabia and occurs at about 20% in European Jews.

So, while I wouldn't conclude that he had partial Jewish ancestry based on his data, the issue is no longer relevant due to the emergence of the new genealogical information.

This is an example of what I called "cryptic ancestry" as a possible explanation for people getting unexpected results.

October 23, 2010

Detailed admixture analysis of West Eurasian populations (+ GenomesUnzipped individuals)

Here is the result of my ancestry analysis of several West Eurasian populations from HapMap 3, HGDP-CEPH, and Behar et al. (2010) datasets using ADMIXTURE for K=10. Some non West Eurasian populations were also added, to squeeze out the partial admixtures from outside the region in some populations.
I have tried to find an informative label for each of the inferred components, corresponding to the common attribute of the populations where it appears the highest. I make no claim about the geographical/temporal/ethnic origin of any of them, or what they actually represent, so please don't take the labels to be more than mental aids for processing the visual information.

Some comments on the components:

The West African component is centered on the Yoruba, is represented among the Maasai, and North Africans, and also occurs in the Near East. As my previous analysis of African populations suggests, this component is not really limited to West African, as the Yoruba show clear ties with the Luhya of East Africa. However, I prefer to call it West African, rather than Sub-Saharan here, as the West African Yoruba are its only representative.

The East African component is centered on the Maasai. It is also represented in North Africans, being more important than the West African among Egyptians, but, as expected, the reverse is true for Moroccans and Mozabites. Also, as I've noted in my analysis of the African ancestry of Near Eastern populations, it is also more important in them than the West African.

The North European component reaches its maximum in Balto-Slavs. However, its substantial presence among Hungarians, French, and Basque, and indeed all European populations, and even Lezgins, suggest that it is a broader phenomenon.

The Druze component is centered on the Druze of the Near East, occurring at a lower frequency even in Europe, and in many populations of the Near East.

The East Eurasian component is centered on the HapMap Chinese, and occurs at a noticeable frequency among Turks, Iranians, Adygei, Russians, and Chuvash.

The Arab component reaches its highest frequency in Bedouins and Saudis but occurs widely in Semitic populations. A very interesting observation about it is that it occurs in Egyptians and Moroccans, but not in the Mozabite Berbers, perhaps reinforcing its Arab and/or Semitic associations. Its paucity among people from the Caucasus (Georgians, Adygei, Lezgins) and Armenians is a further argument for that association.

The NW African component has a clear association with Moroccans and Mozabites, also occurring in Egypt, Morocco and and Sephardic Jews and various other populations.

The Semitic component attains high frequency among Arabs and Jews, hence its name. It seems more widely distributed than the Arab component, which makes sense, as many Near Eastern and European populations have had a longer time in which to encounter different Semitic groups since their Bronze Age appearance in the historical scene, but the corresponding time for encountering Arabs has been shorter.

The SW European component attains its highest frequency among Sardinians and Basques, hence its name. It has an opposite cline of distribution compared to the next component:

The W Asian component dominates in people from the Caucasus and West Asia and is widely distributed across West Eurasia.

Revisiting the Genomes Unzipped individuals

I started the party of using the data of Genomes Unzipped volunteers for ancestry analysis and was soon joined by others. I took another jab at this data as an afterthought in my Near Eastern African analysis, and now it's time to revisit the topic using dense genotype data on a variety of populations, and the power of ADMIXTURE.

First of all, a note of explanation as to why I don't use PCA/MDS in my ancestry analysis. I have nothing against them per se, and the visual examination of a large number of principal components can give a very useful overview of the data. The main problem, however, that I perceive with these techniques is their treatment of individuals of different backgrounds:

The offspring of a Greek and a Norwegian, a Russian and a Spaniard, and two Hungarians may all fall in the same spot on a PCA map. By its very nature, PCA synthesizes across ancestries, representing individuals as singular dots, whereas ADMIXTURE tries to analyze them into their several underlying components.

If we are interested in how an individual compares against other living humans, by all means PCA/MDS are great tools. But you'll never know how the dot came to be, whether it is (a) descended from a long line of ancestors inhabiting the corresponding geographic space, or it is (b) the product of admixture between more distant ancestors whose genomic average projected on the first few principal components matches that of (a).

With that long introduction, here is the analysis of the Genomes Unzipped individuals. These were listed as the "People" group in the above plot, but now we are looking at their individual level components.
The preponderence of "N European" and "SW European" components in everyone but Dan Vorhaus immediately tells us that these are Europeans. Moreover, the relative importance of the "SW European" vs. the "N European" component tells us that they are not from eastern Europe (compare with Balto-Slavs of previous figure). Finally, the relative insignificance of "W Asian" and "Semitic" components is suggestive of NW EUropean ancestry.

Only three individuals stand out in the analysis:

VXP001 (Vincent Plagnol) shows an excess of "SW European" component, supporting an excess of SW over NW European ancestry, but also a tiny slice (1.5% to be precise) of the "Arab" component lacking in the other individuals. Such tiny slices occur in several southern populations, so these results reinforce the "SW" rather than "NW" impression for this sample.

JKP001 (Joe Pickrell) also has elevated levels of "SW European" vs "N European" component, also suggesting more southern ancestry. He also has small slices of "Semitic" (4.4%) "Arab" (1.1%) and "Druze" (1.1%) components. Inherited from his Italian grandparent, perhaps?

Finally, DBV001 (Dan Vorhaus) shows a relative importance of "SW European" (26.6%) "Semitic" (17.9%), "Arab" (2.7%), and "Druze" (1.8%) components, suggestive of a more south- and eastern- origin than the other individuals. It is useful to compare him with the Ashkenazi Jewish average, side by side, as can be seen on the left. Dan is a close match for his people, with the exception of a small green slice of "NW African" in the Ashkenazi sample, which is lacking in Dan.

But, how is that component distributed among Ashkenazi Jews? Here is the answer for the 21 individuals in the sample:


It is evident that this is detectible in some individuals (like #418), excessive in others (like #426). Thus, Dan falls perfectly within the continuum of his people.

It's quite interesting that the three individuals that stood out from the rest in my initial analysis are also the ones who stand out in this more comprehensive assessment, and some of the reasons for it were revealed.

UPDATE (Nov 1): Joe Pickrell discovers Jewish great-grandparent

October 20, 2010

The shape of things to come

(Last Update: Oct 22)

Here is a teaser from my ongoing ADMIXTURE experiments. I have assembled a nice set of Eurasian populations, and this time I will continue increasing K until the Bayes Information Criterion maxes out.

I had previously used BIC to successfully choose K in my craniometric analysis of world populations. I am now at K=7 and it keeps on rising! Either it will stop or my computer will burn out doing Quasi-Newton iterations.

Onto the teaser: I'll only say that "People" are the 12 Genomes Unzipped individuals; I will leave their individual proportions a mystery - for now.

In color:
Sampled populations, # of individuals, ancestral components:


UPDATE (Oct 21):

Here are the K=8 results. I've added some more populations. The next step is to integrate HapMap data with HGDP-CEPH and the Behar et al. dataset. Stay tuned.

I've decided to present the data as population averages, as this is much more readable, especially as the number of individuals grows.

In this run, one of the clusters (purple) became associated with north Mongoloids (Yakuts, and partially Mongols and Daur), whereas Han are mostly in the green component. Notice that Yakut, Uzbeks, Chuvash, and Turks (all Turkic populations) are predominantly in the north cluster as far as their Mongoloid component is concerned, as expected.

UPDATE II (Oct 22):

Here are results for K=9. The addition of Maasai (MKK) has revealed the east African component which they share with Ethiopians; the latter also have a significant West Asian component.

For lack of a better term, I decided to call the lime component "Sardinian", as it is dominant in that population, but it is clearly reflective of something much broader.

October 18, 2010

Joe Pickrell on his ancestry

Joe Pickrell has a nice post on his ancestry at Genomes Unzipped, prompted in part by my use of the EURO-DNA-CALC program on his data. He explores his data using PCA on a much larger set of markers than the 192 used by EURO-DNA-CALC (the overlap between the Price et al. paper on which it is based and the 23andMe set).

I did not report Joe's results in my re-analysis of the Genomes Unzipped data, as his confidence interval intersected 0. For the record, his sample, like that Vincent Plagnol, didn't show any of the "yellow" cluster that was shared by Arabs and Dan Vorhaus (an Ashkenazi Jew).

Joe's explanation that his anomalous results is due to similar allelic frequencies in some markers between Ashkenazi Jews and southern Europeans is quite interesting, and it is based on a dissection of the markers used by EURO-DNA-CALC.

As I stated in my reanalysis of Plagnol's data, his result might also be due to:
a European-origin component in the composite Ashkenazi Jewish gene pool that he happens to share.
Joe's discovery about similar allele frequencies in some markers between Ashkenazi Jews and Italians is quite interesting in terms of my theory about the origin of the European component in the Ashkenazi Jewish gene pool.

Some early studies on AJ, using Y-chromosomes (pdf) overestimated their Near Eastern component by considering them a mix between Levantine and north/central European populations. But, this made the assumption that Jewish ancestors took a Lufthansa flight to Germany rather than spend 1,000+ years in the territory of the Roman Empire and the Hellenistic world where there might also have been introgression of European elements into their gene pool.

Thus, it may very well be that Ashkenazi Jews are distinctive with respect to NW Europeans both in terms to an ancient Near Eastern ancestry (shared with Arabs and "discovered" in my re-analysis) and with respect to Southern European ancestry corresponding to particular populations they interacted with during their sojourn in the Roman Empire.

To conclude, the release of the Genomes Unzipped data on the web has been beneficial to all involved: a few bloggers like myself could run and test their tools on the data, and people like Joe could give feedback on these tools. This is a strong argument in favor of open source tools and public available data and against the use of proprietary databases/ancestry estimation methods/datasets barricaded behind various controls.

As I continue my various ADMIXTURE-experiments, I will be sure to revisit the Genomes Unzipped folk and all those who wish to join them.

UPDATE (Oct 23): A much more detailed analysis of Genomes Unzipped individuals.
UPDATE (Nov 1): Joe Pickrell discovers Jewish great-grandparent

October 14, 2010

African admixture in the Near East: where from?

Here is the result of running ADMIXTURE for K=5 using 275K SNPs on the combined HGDP + HapMap African and West Asian populations, also including Adygei and Tuscans. The populations are in order: Luhya, Maasai, Tuscans, Yoruba, Adygei, Bedouins, Druze, Mozabites, Palestinians.
At this level of detail, Africans are divided into three clusters which can be labeled Sub-Saharan (red), East African (blue), and "Mozabite" or North African (purple). Europeans and West Asians form the green cluster, while the Arab samples have a substantial contribution of the yellow cluster.

Here are the admixture proportions:

African admixture in the two European populations is probably in the limits of statistical noise and consists of "Mozabite" (0.4%) for Tuscans and "E African" for Adygei (0.6%).

Druze, an Arab population that was religiously isolated from Arab Muslims for about a thousand years seems to have correspondingly missed most African admixture, registering 0.6% "Mozabite" and 1.1% "E African".

Non-Druze Arabs have clear traces of African admixture both in the form of "Mozabite" North African (4.5% for Palestinians, 4.9% for Bedouins), E African (6% for Palestinians and 5.7% for Bedouins) and a little Sub-Saharan (1.3% for Palestinians and 2.1% for Bedouins).

I had pointed the mainly eastern African admixture in Near Eastern Arabs a year ago in my review of HAPMIX. Clearly Maasai are a better stand-in than the Yoruba for whatever African ancestry Arabs have.

It is quite interesting to note the genetic distance (expressed in Fst) between the five inferred clusters:

We can plainly see that proximity to Eurasians increases in the order of Sub-Saharan, East African, "Mozabite". I have little doubt that Somalis and Ethiopians from East Africa would occupy an intermediate position between Maasai from Kenya and "Mozabites" in that order.

An interesting observation is that the "Arab" cluster is slightly more distant to all African clusters than the European/W Asian cluster is. This might seem perplexing as geography might dictate that it should be closer to the African clusters.

However, this is not very surprising to me, as there was gene flow between West Asia and Europe and Africa in old times, evidenced by such things as the presence of Eurasian Y-haplogroup R-V88 in Africa and African haplogroup E1b in Europe and West Asia.

The original Arab ancestors, were probably haplogroup J1e-bearing Semites exploiting arid environments of West Asia. Present-day Levantine Arabs (especially Bedouins, in the available samples) maintain a strong signal of this component of their ancestry, admixed, however, principally with the original Tuscan- and Adygei- like West Asians, and secondarily with E and N Africans.

Revisiting GenomesUnzipped "Ashkenazi Jewish" admixture

There were two individuals in my recent post who showed some evidence of "Ashkenazi Jewish" admixture (DBV001: 100% and VXP001: 32%). I list in the comments of that post some possible explanatios for why VXP001 (who has no knowledge of Jewish ancestry) might get such a result. Naturally, using 275K SNPs is better than the 192 of EURO-DNA-CALC, so I did a separate run that included these two individuals.

The results are:

DBV001: 85.1% European/W Asian, 10.5% "Arab", 0.5% "E African", and 3.8% "Mozabite". This is entirely consistent with full known Jewish ancestry. The closest population to the Middle Eastern component of Jews are presumably the Druze, who have about 16.9% of the "Arab" (which should probably be relabeled "Semitic") cluster. Ashkenazi Jews are known to be intermediate between Levantine and European populations, and DBV001's result is entirely consistent with this.

As I've mentioned before, the exact percentage of Middle Eastern ancestry in modern European Jews is difficult to estimate, as this would depend on determining the exact percentage of "European/W Asian" and "Semitic" components was present in their gene pool before they settled in Europe. If, for example, they were 100% in the "Semitic" cluster, then DBV001 would be about 10% of Middle Eastern ancestry, but if they were like modern Druze, then this percentage would be 100*10.5/16.9 = 62.1%. The truth is probably somewhere in between.

VXP001: A shorter story, as VXP001 comes out 100% "European/W Asian". Thus, I am inclined to believe that VXP001's AJ score is either due to the small number of markers, or to a European-origin component in the composite Ashkenazi Jewish gene pool that he happens to share.

UPDATE (Oct 23): A much more detailed analysis of Genomes Unzipped individuals.

October 11, 2010

Running EURO-DNA-CALC on GenomesUnzipped

(Last Update Oct 23)

genomesunzipped is a new initiative to put data of personal genomics customers online. It's a great idea, and the data will be quite useful to many people.

I downloaded the available data and ran EURO-DNA-CALC on them. Of course it is meant to be used for European or West Eurasian people, which all of them seem to be.

Here are the results for the 12 people whose data were online as of this writing. In bold are components whose confidence intervals do not intersect 0.


Most of them seem to be of NW European descent as their names suggest, and a couple seem to be partly or significantly of Jewish descent.

If you are one of these people, feel free to write to me or leave a comment to tell me if I'm right or dead wrong!

UPDATE (Oct 14): A more detailed analysis of DBV001 and VXP001 in this post.
UPDATE (Oct 23): A much more detailed analysis of all Genomes Unzipped individuals in the context of western Eurasia.
UPDATE (Nov 1): Joe Pickrell discovers Jewish great-grandparent