Showing posts with label Ethnicity. Show all posts
Showing posts with label Ethnicity. Show all posts

March 01, 2013

Genomewide diversity in the Levant (Haber et al. 2013)

Razib points me to a new paper (and its associated data, consisting of Christian, Druze, and Muslim Lebanese).


Genome-Wide Diversity in the Levant Reveals Recent Structuring by Culture

Marc Haber et al.

The Levant is a region in the Near East with an impressive record of continuous human existence and major cultural developments since the Paleolithic period. Genetic and archeological studies present solid evidence placing the Middle East and the Arabian Peninsula as the first stepping-stone outside Africa. There is, however, little understanding of demographic changes in the Middle East, particularly the Levant, after the first Out-of-Africa expansion and how the Levantine peoples relate genetically to each other and to their neighbors. In this study we analyze more than 500,000 genome-wide SNPs in 1,341 new samples from the Levant and compare them to samples from 48 populations worldwide. Our results show recent genetic stratifications in the Levant are driven by the religious affiliations of the populations within the region. Cultural changes within the last two millennia appear to have facilitated/maintained admixture between culturally similar populations from the Levant, Arabian Peninsula, and Africa. The same cultural changes seem to have resulted in genetic isolation of other groups by limiting admixture with culturally different neighboring populations. Consequently, Levant populations today fall into two main groups: one sharing more genetic characteristics with modern-day Europeans and Central Asians, and the other with closer genetic affinities to other Middle Easterners and Africans. Finally, we identify a putative Levantine ancestral component that diverged from other Middle Easterners ~23,700–15,500 years ago during the last glacial period, and diverged from Europeans ~15,900–9,100 years ago between the last glacial warming and the start of the Neolithic.

Link

February 06, 2013

Clustering folk tales

Proc. R. Soc. B 7 April 2013 vol. 280 no. 1756
doi: 10.1098/rspb.2012.3065

Population structure and cultural geography of a folktale in Europe 

Robert M. Ross et al.

Despite a burgeoning science of cultural evolution, relatively little work has focused on the population structure of human cultural variation. By contrast, studies in human population genetics use a suite of tools to quantify and analyse spatial and temporal patterns of genetic variation within and between populations. Human genetic diversity can be explained largely as a result of migration and drift giving rise to gradual genetic clines, together with some discontinuities arising from geographical and cultural barriers to gene flow. Here, we adapt theory and methods from population genetics to quantify the influence of geography and ethnolinguistic boundaries on the distribution of 700 variants of a folktale in 31 European ethnolinguistic populations. We find that geographical distance and ethnolinguistic affiliation exert significant independent effects on folktale diversity and that variation between populations supports a clustering concordant with European geography. This pattern of geographical clines and clusters parallels the pattern of human genetic diversity in Europe, although the effects of geographical distance and ethnolinguistic boundaries are stronger for folktales than genes. Our findings highlight the importance of geography and population boundaries in models of human cultural variation and point to key similarities and differences between evolutionary processes operating on human genes and culture.

Link

May 30, 2012

Spatial Ancestry Analysis (Yang et al. 2012)

Link to SPA software.

Nature Genetics 44, 725–731 (2012) doi:10.1038/ng.2285

A model-based approach for analysis of spatial structure in genetic data

Wen-Yun Yang et al.

Characterizing genetic diversity within and between populations has broad applications in studies of human disease and evolution. We propose a new approach, spatial ancestry analysis, for the modeling of genotypes in two- or three-dimensional space. In spatial ancestry analysis (SPA), we explicitly model the spatial distribution of each SNP by assigning an allele frequency as a continuous function in geographic space. We show that the explicit modeling of the allele frequency allows individuals to be localized on the map on the basis of their genetic information alone. We apply our SPA method to a European and a worldwide population genetic variation data set and identify SNPs showing large gradients in allele frequency, and we suggest these as candidate regions under selection. These regions include SNPs in the well-characterized LCT region, as well as at loci including FOXP2, OCA2 and LRP1B.

Link

December 29, 2011

Chinese, Korean, Japanese (genetic edition)

My 2006 post on facial composites of Chinese, Korean, and Japanese women is, surprisingly, the most widely read single entry of this blog. People still occasionally guess "who is who" in that post, five years later.

As I was going through the list of the Dodecad populations, I realized that there are 5+ participants in each of the Korean, Japanese, and Chinese groups. So, it seemed like a simple exercise to see whether the relatively high success rate of people's guesses could be corroborated using the DNA data.

Below is the MDS plot; there are 9 Chinese, 5 Japanese, 5 Koreans in the Dodecad Project; I have also added 30 HapMap Chinese (CHB) and Japanese (JPT):
Only the first MDS dimension showed deviation from normality according to a Shapiro-Wilk test. Using MCLUST, that dimension was enough (as can be seen from the above figure) to infer the presence of 3 clusters which corresponded to the 3 groups, with 100% correct assignments.

Interestingly, when I did not use the extra HapMap individuals, MCLUST did not split Koreans from Chinese. This goes to show that the absence of apparent structure does not imply absence of structure. The extra Chinese and Japanese individuals helped flesh out the existing structure in these East Asian groups.

Below is the list of the Dodecad populations that are below the 5-individual limit:


Algerian_D 4 East_African_Various_D 3 Greek_Italian_D 2 Belgian_D 1
North_African_Jews_D 4 Danish_D 3 Swiss_German_D 2 Latvian_D 1
Slovenian_D 4 Tunisian_D 3 Szekler_D 2 Estonian_D 1
Mixed_Scandinavian_D 4 Austrian_D 3 Mandaean_D 2 Bangladesh_D 1
Moroccan_D 4 Saudi_D 3 Azeri_D 2 Yemenese_D 1
Serb_D 4 Pakistani_D 3 Czech_D 2 Sri_Lanka_D 1
Tatar_Various_D 3 Georgian_D 2 Hungarian_D 1
Palestinian_D 3 Kazakh_D 2 Basque_D 1
Romanian_D 3 Udmurt_D 1
Ukrainian_D 1
Egyptian_D 1

If you belong to one of the above groups (all 4 grandparents) and have tested with either 23andMe or Family Finder, you are especially invited to contact me at dodecad@gmail.com (but do not send data right away!), about possible inclusion in the project. 

For example, in the most recent Clusters Galore analysis, there was a generic "Balkan" cluster. Does this imply that Balkan ethnic groups cannot be distinguished from each other, or that sample sizes are simply not yet sufficient to make manifest the existing structure?

December 28, 2011

Genetic structure in China

After my experiment on Spain, I decided to carry out a similar experiment in China, for which there is a large number of regional/ethnic sub-populations.15 clusters were inferred with 22 MDS dimensions.

The Uygur are the clear outlier population, doubtlessly due to their substantial Caucasoid admixture and geographical position in Central Asia, a region that was traditionally at the outskirts of Chinese civilization. Other Altaic speakers (both Mongolic and Tungusic) are also divergent, as are the Dai/Lahu people from the China/Thailand/Laos area.

Interestingly, the Tujia people from Central China seem to be the ones most like the Han overall, with Hmongic Miaozu/She more like the southern Han.

August 12, 2011

Armenian population structure with autosomal STRs

It is unfortunate that what seems to be a good sized sample of regional Armenian sample was only analyzed with a small set of autosomal STRs. Hopefully, the sample may still be available in the future for more comprehensive tests.

Population codes can be found in a freely available supplementary table.

I had previously used a similar set of STRs with structure to estimate inter-continental admixture in a set of "pure", and widely differentiated individuals (Germans, Africans, Chinese). That suggested to me, that while the continent origin can indeed be guessed for most individuals, admixture proportions are extremely noisy. For much closer-related populations (Armenian groups and various other Caucasoid neighbors), I fear that we should be extremely cautious in over-interpreting the limited evidence that a small panel of STRs provides.

Here are also Dodecad v3 population portraits for the Behar et al. (2010) Armenian sample and the Armenian_D sample from the Dodecad project:


I do not have enough Armenian participants yet, to speak about possible substructure in that population, but, from the samples available so far, which are from several different regions, a picture of relative homogeneity emerges. I encourage individuals with 4 Armenian grandparents who've tested with either 23andMe or Family Finder to join the Project (send e-mail to dodecad@gmail.com).

American Journal of Physical Anthropology
DOI: 10.1002/ajpa.21558

Regionalized autosomal STR profiles among Armenian groups suggest disparate genetic influences

Robert K. Lowery

Abstract
The archeology and ethnology of Armenia suggest that this region has acted as a crossroads for human migrations from Europe and the Middle East since at least the Neolithic. Near continual foreign influx has, in turn, led to the supposition that the gene pools of geographically separated Armenian populations may have diverged as differing historical influences potentially left distinct genetic traces in the various regions of the Armenian plateau. In this study, we seek to address whether any evidence for such genetic regional partitioning in Armenians exists by analyzing, for the first time, 15 autosomal short tandem repeat (STR) loci in 404 Armenians from four geographically well-characterized collections (Ararat Valley, Gardman, Sasun, and Lake Van) that represent distinct communities from across Historical Armenia. In addition, to determine whether genetic differences among these four Armenian populations are the result of differential affinities to populations of known historical influence in Armenia, we utilize 27 biogeographically targeted reference populations for phylogenetic and admixture analyses. From these examinations, we find that while close genetic affiliations exist between the two easternmost Armenian groups analyzed, Ararat Valley and Gardman, the remaining two populations display substantial distinctions. In particular, Sasun is distinguished by evidence for genetic contributions from Turkey, while a stronger Balkan component is detected in Lake Van, potentially suggestive of remnant genetic influences from ancient Greek and Phrygian populations in this region.

Link

July 28, 2011

Slavonic origin of Sorbs (again)

Another recent related study on the origin of Sorbs.

BMC Genetics 2011, 12:67doi:10.1186/1471-2156-12-67

Population-genetic comparison of the Sorbian isolate population in Germany with the German KORA population using genome-wide SNP arrays

Arnd Gross et al.

Abstract (provisional)
Background
The Sorbs are an ethnic minority in Germany with putative genetic isolation, making the population interesting for disease mapping. A sample of N=977 Sorbs is currently analysed in several genome-wide meta-analyses. Since genetic differences between populations are a major confounding factor in genetic meta-analyses, we compare the Sorbs with the German outbred population of the KORA F3 study (N=1644) and other publically available European HapMap populations by population genetic means. We also aim to separate effects of over-sampling of families in the Sorbs sample from effects of genetic isolation and compare the power of genetic association studies between the samples.

Results
The degree of relatedness was significantly higher in the Sorbs. Principal components analysis revealed a west to east clustering of KORA individuals born in Germany, KORA individuals born in Poland or Czech Republic, Half-Sorbs (less than four Sorbian grandparents) and Full-Sorbs. The Sorbs cluster is nearest to the cluster of KORA individuals born in Poland. The number of rare SNPs is significantly higher in the Sorbs sample. FST between KORA and Sorbs is an order of magnitude higher than between different regions in Germany. Compared to the other populations, Sorbs show a higher proportion of individuals with runs of homozygosity between 2.5 Mb and 5 Mb. Linkage disequilibrium (LD) at longer range is also slightly increased but this has no effect on the power of association studies. Oversampling of families in the Sorbs sample causes detectable bias regarding higher FST values and higher LD but the effect is an order of magnitude smaller than the observed differences between KORA and Sorbs. Relatedness in the Sorbs also influenced the power of uncorrected association analyses.

Conclusions
Sorbs show signs of genetic isolation which cannot be explained by over-sampling of relatives, but the effects are moderate in size. The Slavonic origin of the Sorbs is still genetically detectable. Regarding LD structure, a clear advantage for genome-wide association studies cannot be deduced. The significant amount of cryptic relatedness in the Sorbs sample results in inflated variances of Beta-estimators which should be considered in genetic association analyses.

Link

May 02, 2011

Six months of the Dodecad Project

The Dodecad Ancestry Project has recently reached its half-year milestone, and to celebrate, I present a small analysis of the Dodecad populations with at least 10 participants. With 17 such populations, and 254 individuals, we have now reached a point where the Dodecad dataset is equivalent to many published datasets in the literature.

The total sample now exceeds 600 members, and that includes both individuals from single populations that have not yet reached either the 5+ or 10+ person mark, as well as individuals of mixed heritage.

The first PCA dimension distinguishes between Indians and West Eurasians. The second one is clinal between Finns and North Africans.
In the following ones, I present the first dimension with the 3rd, 4th, etc.

The third dimension distinguishes between North Africans and the rest:
The fourth dimension separates the Ashkenazi Jews:
The fifth dimension distinguishes Finns:

Here are the results of the Clusters Galore analysis. I usually perform this over MDS data, but it works just as well on PCA data as well, as I've mentioned in the post where I introduced the idea.

Most Greeks align themselves with Italians, but some do so with the Balkans and West Asia. I would not draw many conclusions from this, as it might be a consequence of the composition of the Greek sample, but it's not a coincidence, I think, that these three places are Greece's geographical neighbors, as well as the only ones where Greek has continued to be spoken until quite recently.

Another interesting statistic is the average intra-population identity-by-state (IBS), which is a good measure of a population's homogeneity:

Here is the full IBS matrix:


I have recently introduced the idea of the population concordance ratio γ(A, B), the value of which becomes 1 if two individuals of population in a row A are always more similar to each other than any individual from column B:

As expected, most populations are perfectly concordant with respect to Indians and North Africans, the two clear outgroups in this set. Another interesting observation is that the most heterogeneous populations are the least concordant, as they include individuals of quite varying ancestry (e.g., a "white Mexican" may be closer to an Englishman than to a very Amerindian-admixed Mexican).

I decided to test this idea by calculating the correlation between a population's average concordance ratio and its average IBS similarity, which turned out to be -.36, which is not significant, but in the right direction. Perhaps average IBS similarity is less than ideal for the purpose of gauging homogeneity as it applies to the concordance ratio.

Notice also the asymmetry between γ(A, B) and γ(B, A). For example, the sample from the Balkans consists of an assortment of non-Greek people from the Balkans, so it's not particularly concordant with respect to many North European populations: people from the Balkans differ from each other substantially in their North European-ness. North European populations, however, tend to be concordant with respect to people from the Balkans.

Submission to the Project is currently closed, but I do encourage all individuals with 4 grandparents from the same European, North/East African, West/Central/South Asian group to contact me at dodecad@gmail.com for possible inclusion in the Project. Send e-mail first, not data, to check whether I can process your data. Or, follow the project blog for future submission opportunities.

April 23, 2011

Genetic structure of West Eurasians

I have decided to generate a new major data dump of ADMIXTURE results. In comparison to previous such experiments:
  1. The focus is entirely on West Eurasians (Caucasoids).
  2. I have excluded all potential relatives from the source datasets, as well as several populations that tend to create uninformative clusters of their own (e.g., Druze or Ashkenazi Jews); exceptions are populations of great anthropological interest (e.g., Basques).
  3. I have included all relevant Dodecad Ancestry Project populations with 5+ participants.
  4. I have developed a new way of "framing" the region of interest by choosing appropriate sets of individuals from outside of it.
"Framing" populations

I have, since the beginning of my ADMIXTURE experiments, emphasized the importance of including appropriate population controls designed to squeeze out minor distant admixture in populations of interest, so that it does not confound the inference of region-specific components.

This leads to a problem: there are many possible sources of admixture. For example, we do not know a priori which set of African populations may have contributed to Caucasoid populations, or which set of East Asian ones. We could choose e.g., the Yoruba and the Chinese to represent Sub-Saharans and East Asians, but that might exclude possible sources of variation, and lead to Yoruba- and Chinese- specific clusters rather than more general Sub-Saharan and East Asian ones. If we included more population controls, we would cover more possible sources of variation, but ADMIXTURE would infer components of little interest (e.g., between Pygmies vs. Bushmen or Mongols vs. Chinese)

To avoid this, I propose to create meta-populations consisting of a single individual from many populations, i.e., a Yoruba, a Mandenka, a San, a Mbuti Pygmy, etc. for Sub-Saharan Africa, or a Miaozu, a Han, a Mongol, a She, a Hezhen, etc. for East Asia. That way we are both helping ADMIXTURE infer general components, while at the same time preventing it from inferring non-region specific ones.

Results

The entirety of the results presented here can be downloaded. They include:
  1. Population sources
  2. ADMIXTURE proportions for populations
  3. Fst divergences between components
  4. Population portraits showing individual level variation
See spreadsheet and associated bundle (or here).

At K=3, we observe the emergence West Eurasian, Sub-Saharan, and East/South Asian components.

The impact of the Sub-Saharan component is felt most distinctly in North Africa and the Near East, especially among Arabs; the impact of the East/South Asian one in West Asia and Northeastern Europe, especially among Finnic and Turkic speakers.

It is interesting to note that 39.8% of the Indian_D sample is assigned to the E/S Asian component. I had previously estimated in a roundabout way, and in a slightly smaller sample that the Ancestral South Indian component in Project participants was 33.3%, so ADMIXTURE has roughly managed to infer correctly that about 1/3 of this Indian sample's ancestry is more closely related to East Asians than to West Eurasians.

At K=4, the first split within the Caucasoid group appears: a component centered onn Europe, and one on West/South Asia.

Many populations possess both these components in clinal proportions.

The European component shrinks to insignificance in Arabians, such as Saudis and Yemenese.

The West/South Asian component shrinks to insignificance in Northeast Europeans, such as Finns, Lithuanians, north Russians, and Chuvash.


At K=5, a new Mediterranean component emerges. This is highly represented in populations to the North, South, and East of the Mediterranean sea.

This component is noteworthy for its absence in India and Northeastern Europe.

In Northeastern Europe, the Mediterranean component is hardly represented at all, whereas the West/South Asian component, freed of its K=4 Mediterranean associations now makes its appearance.

Conversely, in the West Mediterranean, among Basques, Sardinians, Moroccans, and Mozabites the West/South Asian component vanishes to non-existence.


At K=6, a North African component emerges.

Notice its presence in the Near East and parts of Southern Europe.

The two regions can be contrasted in terms of their African components, with very high North/Sub-Saharan African ratio in Europe vs. much lower in the Near East.

The explanation for this seems straightforward, as Europe was affected by North Africa in prehistoric and historic times, whereas the Near East also shares a border with more southern parts of the African continent, as well as the potential influence of the medieval slave trade that seems to have affected Muslim Near Eastern populations disproportionately.


At K=7, a Southwest Asian component emerges which is highest in Arabia and East Africa. I could've called this Red Sea, but I've reserved this name for a similar component that emerges at higher K.

It is clear that this is the main Caucasoid component present in East Africa.

It vanishes to non-existence in the Northern fringe of Europe, in the British Isles, Scandinavia, and among the Finns and Lithuanians.

Another interesting aspect of its distribution is its presence in Pakistan but not India. Perhaps, in this case, it reflects historical contacts between the Islamic Near East and parts of South Asia.


At K=8, we observe most of the familiar components from the K=10 analysis of the Dodecad Project. However, the use of the framing populations has meant that these components emerge before either Africans or East Eurasians split.

Now, the South Asian component appears, which swallows up most of the E/S Asian component that previously linked South with East Asians. This component extends a great way to the Near East and eastern parts of the Caucasus.

Quite interestingly, the remainder of the Caucasoid component in South Asia that is not absorbed by the new South Asian component seems to be split between the West Asian and North/Central European components, with an absence of the South European component.

It is among the Lezgins of the Caucasus that such a combination occurs, on the western shore of the Caspian Sea. The same combination of Caucasoid components also occurs in Uzbeks and Chuvash.

I conclude from this that the Caucasoids who entered South and Central Asia were probably derived from the eastern fringes of the Caucasoid world where only the West Asian (in the south) and North/Central European (in the north) are in existence. The area around the Caspian Sea seems like an excellent candidate for their origin, as I have speculated before, as that region has two important properties:
  1. It is transitional between predominantly N/C European populations to the north and predominantly W Asian populations to the south
  2. It is the border of the influence of the S European element, with Georgians possessing some of it, while Lezgins do not.

At K=9, we see the emergence of specific Sardinian and Basque components. Normally this is undesirable, but, I believe this breakup serves to divide the previously inferred South European component meaningfully.

What was South European in lower K seems to have an Atlantic vs. Mediterranean dimension, with the Basque/Sardinian ratio being particularly high in the Atlantic facade of Europe. Conversely, this ratio is low in the Mediterranean as we move eastwards: it is already low in Italy and the Balkans and becomes virtually zero in Cypriots, Armenians, and Levantine Arabs.

North Africa is also particularly interesting in having a low Basque/Sardinian rate, even in Morocco. It appears that Sardinians are a much better proxy of European influences in the region than Basques are.

K=10 is particularly exciting because, for the first time, there is clear evidence of structure in the North/Central European component that can now be split, for the first time, into Northwestern and Northeastern ones.

The NW European component is maximized in Orcadians, and people from the British Isles in general, as well as in Scandinavia. These populations have a low NE/NW ratio, as do the French, Iberians, and Italians.

Conversely, Balto-Slavs have a high NE/NW ratio.

Interestingly, Greeks have a balanced NE/NW ratio (1.2), intermediate between Italians and Balto-Slavs. Similar balanced ratios are also found among Lezgins (1.08), Turks, and Iranians. I conclude that Slavic or other Eastern European admixture cannot account for the totality of presence of this component in Greeks.

Indians have a 1.8 NE/NW ratio. In Pakistan this is 6.5, in Uzbeks it is 2.9, and in the North Eurasian_Ra it is 14.2. My conclusion is that a single migration of steppe people from eastern Europe cannot account for the presence of North European-like genes in Asia.

I propose that a palimpsest of population movements has brought such elements into the interior of Asia: the migration of the early Indo-Iranians from West Asia or the Balkans with a balanced NE/NW ratio, and, the migration of steppe people from Eastern Europe with a high NE/NW ratio. The latter, did affect much of Asia, but it is in India, where Iranian groups did not penetrate in great numbers the lower ratio of the Indo-Aryans has been best preserved.

The case of the Finns is also interesting, as there is a surplus of NE over NW European elements. Their position is intermediate between Scandinavians and Lithuanians/Russians but toward the latter. So, Finns appear to (i) have a substratum similar to Balto-Slavs, (ii) to be influenced by Scandinavians, and (iii) with a balance of East Eurasian elements (5.8% at this analysis) preserving the legacy of their linguistic ancestors from the east. At present it is difficult to determine how much of the NE European component in Finns is due to their eastern ancestors who were presumably mixed Caucasoid/Mongoloid long before they arrived in the Baltic, and how much was absorbed in situ.


At K=11 the Ethiopian/East African component emerges, absorbing some of the Red Sea and Sub-Saharan components from the previous K=10 run.

In comparison to the East African component of the Dodecad Project analysis, this component is closer to West Eurasians than to Sub-Saharan Africans, and a residual Sub-Saharan element remains in the two East African (Ethiopian and East_African_D) population samples. Presumably this is due to the more complete sampling of Sub-Saharan genetic diversity using the Sub_Saharan_H "framing" population.

Outside Africa, both E African/Sub-Saharan components are present in the Near East and North Africa with higher E African/Sub-Saharan ratios in the Near East and lower ones in North Africa.

In Europe, there are low such ratios in the few populations where African admixture is present, together with some N African. We can probably conclude that African admixture is mostly due to North Africans, and African-influenced Near Eastern populations, rather than directly from Sub-Saharan Africa.

At K=12 the first uninformative cluster emerged, centered on Iraqi Jews, hence I decided to stop the analysis at this point.

Population Portraits

There is a plethora of population portraits in the download bundle, showing how admixture proportions vary in individuals within populations, and how they vary between successive K.

Here is, for example, the K=11 portrait of Cypriots. A picture of overall homogeneity of this sample emerges, but notice how the NW European and NE European have disjoint presence in the Cypriot individuals, with 5 having some of the former, 6 having some of the latter, and only 1 of these having both.

Compare with Lezgins (right) where these two components occur in all individuals. Whatever this admixture represents, it must be old enough if it is so uniformly distributed in the population.



Here are the Georgians at K=10. Notice that their NE European component is unevenly distributed, and in every case where it occurs it is accompanied by a thin slice of East Asian. This may well indicate partial Russian or other Eastern European ancestry in these individuals.



Side-by-side comparisons are also quite useful. Consider Armenians vs. Lezgins vs. Iranians at K=7







Notice how Lezgins, who live north of the Caucasus mountains possess some of the N/C European component, which the Armenians, who live to the south of them lack. This should come as no surprise, as the Lezgins inhabit parts of the ancient Sarmatia Asiatica. Compare with Iranians, who are differentiated by their Indo-European Armenian neighbors by the presence of a "S Asian" component, which, in turn, ties them to their Indo-Aryan linguistic relatives.

Much more can be said, but I'll let readers explore the data on their own, and draw their own conclusions from them.

April 12, 2011

Population concordance ratio

An interesting problem in population genetics is the following: how often are two individuals from a population A more similar to each other than either of them is to an individual from another population B?

Suppose you have a similarity function for individuals, sim(a, b)

(In my experiments I will use identity-by-state (IBS) as calculated by PLINK (--cluster --matrix) as a measure of similarity, but any symmetrical similarity function (that is, sim(a,b)=sim(b,a)) will do.)

We want to calculate the rate at which the following condition occurs:


If this condition holds, then a and a' from population A are more similar to each other than either of them is to an individual b from population B. We then say that the trio of individuals is concordant.

I will use the indicator function I(a, a', b) = 1 in case of a concordant trio, and =0 otherwise.

I can then estimate the probability of concordance, if I have n individuals from A and m from B as follows:

The rationale behind this formula is straightforward: we are counting the number of all concordant trios, and dividing by n(n-1)m/2 since there are n(n-1)/2 pairs of individuals from A, and each pair is compared against all m individuals from B.

The expected value of this concordance ratio can vary between 0.25 and 1:
  • It is 1 if the two populations are so well-differentiated so that every trio is concordant.
  • On the other hand, if the two populations are genetically identical, then each similarity comparison is equivalent to a coin toss (probability = 0.5) and we are testing this condition for two different individuals from A: hence the probability of concordance for each trio is 0.5*0.5 = 0.25.

In a finite sample of individuals it is possible that the concordance ratio estimate may actually be lower than 0.25.

An interesting property of the concordance ratio is its asymmetry, that is:

We will see how this property gives some useful insight in some of the following examples.

#1. Han vs. Yoruba

This is based on the Stanford HGDP set, including only individuals as recommended by Rosenberg (2006) in his H952 set. For each experiment only SNPs with at least a 99% genotyping rate have been retained.

The first experiment is designed to showcase the concordance ratio in two well-differentiated human populations, 44 Han Chinese and 21 Yoruba Nigerians. The analysis is based on 617,602 SNPs.

That is, two Han Chinese are always more similar to each other than to a Yoruba and vice versa.

I repeated the experiment, sequentially thinning the market set randomly by a factor of 10 using PLINK's --thin 0.1 argument. Concordance remained 1 for ~61k, ~6k, and ~600 markers, and became different than 1 (but still greater than 0.95) with 50 SNPs.

Genome-wide, or, across a sizeable number of markers concordance of Han vs. Yoruba and vice versa is perfect, while for a few random SNPs it may not be.

#2. Britons vs. Mexicans

In the next experiment, I have used 90 Britons (GBR) and 69 Mexican Americans from Los Angeles (MXL) from the 1000 Genomes Project. I have included only SNPs with 99%+ genotyping rate that are also included in the HGDP Stanford data, for a total of 351,521 markers.

Notice the previously mentioned asymmetry: two Britons are virtually always closer to each other than they are to a Mexican, but a Mexican is sometimes closer to a Briton than he is to a fellow Mexican. This is due to the fact that Mexicans have variable European admixture, so a substantially European-admixed Mexican may be closer to a Briton than he is to one of his substantially Amerindian-admixed compatriots.

#3: Various Europeans

In the final experiment, I included HGDP European populations, together with 12 Dodecad Project Greeks. The analysis is based on 492,176 SNPs. Each row in the following table represents the first argument of the concordance ratio function, and each column the second one.

Here is a way to read the table, using Greek_D as an example:
  • The last row represents a test in which pairs of Greeks were compared against individuals from the population of each column. These comparisons were concordant 84.4% of the time against Russians (the most distant population), and 37.5% of the time against Tuscans (the closest one).
  • The last column represents in which Greeks were used as an "outgroup" for comparison against pairs of individuals from each row. These comparisons were concordant (~1) for most populations, except the Tuscan (47.9%), North Italian (71.5%), and Adygei (81.9%)
Let's do the same using French as an example:
  • With a pair of French individuals against individuals from other populations, concordance ranged between 36.2% (for North Italians) to 95.8% (for Adygei)
  • When French individuals were used as an outgroup to compare against pairs of individuals from other populations, concordance ranged between 63.1% (for North Italians) to 100% (for Sardinians). The latter means that a pair of Sardinians is always closer to each other than to a French sample (or, at least, an HGDP French one).

Conclusion

The study of concordance is an interesting thought experiment that illustrates how genome-wide comparisons between individuals show the following:
  1. Two individuals from a homogeneous population are virtually always more similar to each other than to an individual from a genetically differentiated population
  2. Two individuals from a population may, or may not, be more similar to each other than to an individual from a genetically related population
  3. More variable populations are usually more discordant with respect to other populations, whereas very homogeneous populations tend to be concordant
The concordance ratio is useful for personal genomics customers, as it puts their IBS similarities to various other individuals in perspective. For example, a Greek should not be surprised if he matches a particular Tuscan more than he does a fellow Greek, nor should he seek mysterious Italian ancestors because of it, as such a discordant result occurs frequently.

The concordance ratio is also useful because it provides a truly model-free test of population differentiation:

It is different from techniques such as PCA which allow the separability of individuals from different populations by projecting them on a number of dimensions, the first few of which are usually correlated with the inter-population fraction of genetic diversity. Hence, individuals that appear well-separated on a few PCA dimensions may in fact be overall more genetically similar to individuals from other populations across the full marker set. The concordance ratio avoids any accusation of privileging aspects of the genome (the ones that differentiate populations), as it is based on a single genome-wide similarity function for individuals.

It is also different from clustering algorithms such as ADMIXTURE that infer allele frequencies in putative ancestral populations, again implicitly using markers with high frequency differences to estimate admixture proportions. Hence, a "match" in a marker with low population differentiation is treated differently as a source of evidence than a match in a marker of strong population differentiation. The concordance ratio avoids this issue by using a single similarity function for individuals that does not privilege one marker over another.

R Code

Code for the calculation of the concordance ratio ratio can be downloaded from here as an R function. Two files are required:
  • A symmetrical similarity matrix, as output by plink --matrix --cluster command. Any similarity matrix file in the PLINK MIBS format will do.
  • A file in which each row has a population name and the number of individuals from that population.
An example of the latter file count.txt for Experiment #3 is:

North_Italian 12
Russian 25
Orcadian 15
Sardinian 28
Tuscan 8
French 28
French_Basque 24
Adygei 17
Greek_D 12

Of course, individuals must appear in that order in the plink file, i.e., first the 12 North Italians, then the 25 Russians, etc.

Assuming you have such a file, e.g., in binary BED/BIM/FAM format, you first calculate the IBS matrix:

plink --matrix --cluster --bfile datafile

This creates a plink.mibs file. Then, in R, after changing to the appropriate directory, where plink.mibs, count.txt and the source code gamma.r is, you enter:

source("gamma.r")
gamma(simfile="plink.mibs", popfile="count.txt")


PS: The concordance ratio should not be confused with Witherspoon's ω fraction. That is defined by comparing all pairs of between- and within- population distances, and ranges between 0 (highest concordance in my terminology) and 0.5 (lowest concordance). The concordance ratio, on the other hand, tests all possible trios of individuals, and it also has the asymmetrical property explained above.

March 29, 2011

The power of Clusters Galore: Iranians and Arabs

The full power of Clusters Galore depends on its ability to infer clusters of arbitrary size, shape, and orientation in a high-dimensional space. It achieves this by using MCLUST over an MDS or PCA representation of dense genomic data.

Nonetheless, we can still see get a sense of it even in a simple 2D representation as the following:
This was produced by applying MDS on 240 individuals (from Behar et al. 2010, HGDP, Xing et al. 2010, and the Dodecad Project).

One can see that the Behar et al. and Dodecad Iranians form a small cluster on the right, together with the Xing et al. Kurds and the single Dodecad Kurd. Arabs are quite more variable: Druze extend to the bottom of the figure, Bedouin form two groups: one similar to other Arabs, the other extending to the left of the figure. There are also a few Arabs stretching to the top.

The variability of the Arabs can be attributed to reproductive isolation, inbreeding, and variable amounts of African admixture. Let's apply MCLUST over these 240 2D points:
The above visual representation shows the centroids and shapes of the 5 inferred clusters. Here are the numbers of individuals from each population assigned to each cluster:

Notice cluster #5: it consists of all Kurds, most Behar et al. Iranians and all Dodecad ones, and the single Dodecad Kurd, plus a Lebanese and a Syrian. It is overall 96% Iranic in composition. It is quite tempting to think that the two Syrian and Lebanese members have some links to Iranian peoples either due to Kurdish ancestry or the Shia form of Islam.

The more variable Arabs are split into multiple clusters: the main, tight, cluster #3 which includes most of the Levantine Arabs, but also some Saudis and Yemenese, the extremely variable African-admixed cluster #1 dominated by some Yemenese but including a few others, the "Arabian" Saudi-Bedouin dominated cluster #2, and the Druze-specific cluster #4.

It seems that just as the distinction between Celto-Germans and Balto-Slavs is not only cultural, but also genetic, so is the distinction between Iranian and Arab. In the case of the Arabs though, religious distinctions (e.g., the Druze), variable African admixture, and quite possibly Arabization of Levantine populations has resulted in a non-homogeneous array of genetic clusters.

PS: Iranic groups are also not homogeneous if one includes some of those from South Asia, as evidenced by this previous genetic map of West Eurasians which analyzed Kurds and Iranians together with Pathans and Balochis.

March 18, 2011

Analysis of 1000 Genomes + HapMap 3 data

A reader tipped me on the availability of data from the 1000 Genomes Project genotyped on the Illumina Omni 2.5 chip. Out of 2.5 million or so SNPs, there are about 720,000 with rs-numbers in the working dataset. There are a few new populations in the data:
  • GBR (Great Britain)
  • FIN (Finland)
  • IBS (Iberian Spanish)
  • CLM (Colombians)
  • MXL (Mexican Americans from Los Angeles)
  • PUR (Puerto Ricans)
I've been rebuilding my various datasets to account for common markers, high quality SNPs, and linkage disequilibrium, so this is based on about 133,000 markers. I also limited the number of individuals at 25 per population.

I took the HapMap-3 data to make sure that the integration was correct and ran various analytical techniques over the joint dataset of 17 populations and 425 individuals.

Multidimensional Scaling

As expected the three poles correspond to West Eurasians (top left, GBR, CEU, TSI), East Eurasians (bottom left, CHB, CHD, JPT), and Sub-Saharan Africans (YRI).

Other populations fall in between the three poles: for example, FIN slightly removed from West Eurasians in an East Eurasian direction, Mexicans and Gujarati Indians (GIH) in-between West and East Eurasians, African Americans (ASW) and Maasai East Africans (MKK) in-between Sub-Saharan Africans and West Eurasians.

Clusters Galore Analysis

I then used the Clusters Galore approach to cluster individuals. As I've mentioned before, individuals with quite distinct origins may overlap in the MDS representation, and the Galore approach is able to discover distinct clusters by looking at several dimensions at the same time, and using a state-of-the art clustering algorithm, MCLUST.

As can be seen in the MDS plot, Mexicans and Gujarati Indians overlap, as well as African Americans and Maasai. Obviously these populations are completely different mixtures that happen to coincide in genomic space due to the relatedness of their ancestral components that intermixed at different times and in different continents.

Here are the results of the Galore analysis. With 20 MDS dimensions retained (the maximum I considered) there were 35 clusters in the MCLUST solution that maximized the Bayes Information Criterion.


This is quite instructive:
  • Some populations (FIN and YRI) form their own very specific clusters #2 and #35
  • Some clusters join 2 or more populations. For example White Americans (CEU) and Britons (GBR) form cluster #1
  • Latinos form several clusters, especially the Mexicans. This should've been anticipated from the MDS plot where they are shown to be widely dispersed (quite variable). In essence, Latinos are not homogeneous populations but sets of individuals possessing variable admixture proportions
Note also, that some populations that are folded into a single cluster in this analysis (e.g., Spanish and Tuscans in #3) can in fact be distinguished from each other although not so easily in the first 20 dimensions considered here, as these are dominated by more salient features of the global genetic landscape.

ADMIXTURE analysis

I then ran ADMIXTURE over the dataset for K=5.

Here are the admixture proportions corresponding to this plot:


This is quite instructive with respect to the absence of particular reference populations: Finns show East Eurasian influences in the form of "Native American" (1.5%) and "East Asian" (6.2%) elements. Clearly, we don't have to imagine Native Americans moving into Finland, and these two components are standins for the Siberian ancestors of the European Finns. Similarly, Spanish show African admixture (1.6%). This is also probably due to both North and Sub-Saharan African elements, but the absence of appropriate North African references makes the distinction impossible. Finally, the Maasai show European and African admixture. This may be due to the non-emergence of a specific East African component at this level of resolution, as well as the absence of appropriate West Asian Caucasoid groups that are more likely to have influenced them. The absence of West Asian reference populations also probably affects Tuscans as their West Asian admixture may be misinterpreted as South Asian.

Here are the Fst distances between components:


This is also instructive: the South Asian component, in the absence of relatively unadmixed South Asian references is closer to Europeans than to East Asians. In fact, it is a composite of West Eurasian and indigenous South Asian population elements, the latter being distantly related to East Asians. Similarly, in the absence of Amerindian references, the Native American component (a bit of a misnomer) is equidistant to Europeans and East Asians. In fact, it is also a composite of West Eurasian and pre-Columbian American populations.


Conclusion

The Omni 2.5 data seem to work fine, and genome bloggers can anticipate good things in the future from the 1000 Genomes Project, as many more populations are in the pipeline. Clearly, the full-sequence data will probably be too much to handle for most hobbyists at the moment, but for anthropological investigations the 2.5 million SNPs will be more than enough.

The few experiments I carried out here also served to highlight the problems associated with using a limited number of reference populations. But, thankfully, this was a contrived problem aimed to make a point: there are now publicly available data for most major human populations, so the field is wide open for anyone interested in the study of human variation.

March 02, 2011

Origin of Tibetans (Wang et al. 2011)

The correspondence of language-ethnic affiliation with genomic data is quite striking as can be seen in the neighbor-joining tree (bottom). From the paper:
The migration routes of the Chinese population as a single group have been outlined based on Y chromosome haplotype distributions. After the ancestors of Sino-Tibetans reached the upper and middle Yellow River basin, they divided into two subgroups: Proto-Tibeto-Burman and Proto-Chinese [2]. These two subgroups were similar to the two ancestral components of EA populations at K = 2 (Figure S1B). The ancestral component which was dominant in Tibetan and Yi arose from the Proto-Tibeto-Burman subgroup, which marched on to south-west China and later, through one of its branches, became the ancestor of modern Tibetans. Proto-Tibeto-Burmans also spread over the Hengduan Mountains where the Yi have lived for hundreds of generations [28]. Taking the optimal living condition and the easiest migration route into account, we favor the single-route hypothesis; it is more likely that their migration into the Tibetan Plateau through the Hengduan Mountain valleys occurred after Tibetan ancestors separated from the other Proto-Tibeto-Burman groups and diverged to form the modern Tibetan population.
I recently uncovered a genetic component specific to Altaic populations, and this paper shows that within the Sino-Tibetan group, a component centered on Tibeto-Burmans can be uncovered as well. Genomics is already contributing greatly to our understanding of ancient languages, and it is perhaps time for geneticists to use their tools in order to date the breakup of these language groups, providing important new data on which to base linguistic theories.

PLoS ONE 6(2): e17002. doi:10.1371/journal.pone.0017002

On the Origin of Tibetans and Their Genetic Basis in Adapting High-Altitude Environments

Binbin Wang et al.

Since their arrival in the Tibetan Plateau during the Neolithic Age, Tibetans have been well-adapted to extreme environmental conditions and possess genetic variation that reflect their living environment and migratory history. To investigate the origin of Tibetans and the genetic basis of adaptation in a rigorous environment, we genotyped 30 Tibetan individuals with more than one million SNP markers. Our findings suggested that Tibetans, together with the Yi people, were descendants of Tibeto-Burmans who diverged from ancient settlers of East Asia. The valleys of the Hengduan Mountain range may be a major migration route. We also identified a set of positively-selected genes that belong to functional classes of the embryonic, female gonad, and blood vessel developments, as well as response to hypoxia. Most of these genes were highly correlated with population-specific and beneficial phenotypes, such as high infant survival rate and the absence of chronic mountain sickness.

Link