Showing posts with label Uyghur. Show all posts
Showing posts with label Uyghur. Show all posts

November 03, 2012

Admixture in the Chuvash and the Uygur

I took the Behar et al. (2010) sample of Chuvash, excluding GSM536731 which has atypical ancestry and merged it with the Li et al. HGDP French_Basque and Dai. The latter two populations don't show evidence of admixture according to both the f3-statistic and ALDER (Loh et al. 2012). (I used a --geno 0.03 flag in PLINK and extracted a subset of SNPs including in the Rutgers recombination map for Illumina chips).

The f3-statistic f3(Chuvashs_16; French_Basque, Dai) was equal to -0.011311 (Z=-31.308), indicative of admixture.

I then ran an ALDER analysis:


Test SUCCEEDS (z=4.85, p=1.2e-06) for Chuvashs_16 with {French_Basque, Dai} weights

DATA: success (warning: decay rates inconsistent) 1.2e-06 Chuvashs_16 French_Basque Dai 4.85 3.78 5.18 50% 40.27 +/- 5.80 0.00032377 +/- 0.00006676 28.21 +/- 7.47 0.00004231 +/- 0.00000962 47.08 +/- 4.53 0.00016628 +/- 0.00003212

DATA: test status p-value test pop ref A ref B 2-ref z-score 1-ref z-score A 1-ref z-score B max decay diff % 2-ref decay 2-ref amp_exp 1-ref decay A 1-ref amp_exp A 1-ref decay B 1-ref amp_exp B

This indicates that the Chuvash can be seen as admixed, but with inconsistent decays: the one with the French Basque (=28.21) is younger than the one with the Dai (=47.08). I think this makes fairly good sense, because the Chuvash are descended from people who came to Europe during the 1st millennium AD and must have later mixed with Europeans, perhaps with eastern Slavs as these made their way eastward during the 2nd millennium AD.

I then carried out similar analyses on the HGDP Uygur. As expected f3(Uygur; French_Basque, Dai) = -0.023917 (Z = -60.362), indicative of admixture. The ALDER analysis:


Test SUCCEEDS (z=6.85, p=7.4e-12) for Uygur with {French_Basque, Dai} weights

DATA: success 7.4e-12 Uygur French_Basque Dai 6.85 4.47 7.39 15% 20.56 +/- 3.00 0.00036760 +/- 0.00003660 22.59 +/- 5.06 0.00010920 +/- 0.00002025 19.46 +/- 2.64 0.00007864 +/- 0.00000710

DATA: test status p-value test pop ref A ref B 2-ref z-score 1-ref z-score A 1-ref z-score B max decay diff % 2-ref decay 2-ref amp_exp 1-ref decay A 1-ref amp_exp A 1-ref decay B 1-ref amp_exp B

suggests a very recent admixture on both the European and East Asian side. It seems fairly clear that whatever admixture was taking place in Central Asia, perhaps for thousands of years, the present-day Ugyur were formed, at least in part, by a fairly recent, perhaps post-Mongol admixture event.

May 20, 2011

On Tocharian origins

Where did the Tocharians originate from? J.P. Mallory's recent talk has been somewhat of an eye-opener for me, as Prof. Mallory brought to my attention two important issues:
  1. The lack of a clear connection between Afanasyevo and the Tarim Basin.
  2. The existence (in Tocharian) of a rich agricultural IE terminology related to cereals, as well as the domesticated pig, which cannot be easily explained if Tocharians arrived in Xinjiang from the steppes to the north, and, ultimately from eastern Europe.
To begin with, I want to point out an important issue: we cannot assume that the earliest Caucasoids of Xinjiang, including some of the famous early Tarim mummies were Tocharian speaking. There are several arguments why this is so:
  1. Tocharian is first attested in the 8th c. AD, that is, about 3 thousand years after the earliest detected Caucasoids in the region
  2. There has been a shift in the region from Tocharian and eastern Iranian languages to Turkic over the last thousand years or so. Why assume linguistic continuity in the preceding three thousand?
  3. Indeed, there has been linguistic shift throughout other regions of Eurasia in shorter timespans, such as the spread of Slavic across most of eastern Europe, the virtual extinction of Celtic in most of western Europe, the replacement of multiple languages by Arabic in the Near East, and so on. Linguistic continuity does not seem to be an appropriate default position in the absence of direct evidence.
  4. The earliest Caucasoids of the Tarim were already substantially mixed with Mongoloids at least in their mtDNA. This reduces our confidence that they spoke an Indo-European language, as there is a pattern of Caucasoid patrilineages combined with Mongoloid mtDNA in present-day non-IE South Siberians
  5. Indeed, the current Turkic Uyghurs, who are closer (temporally) to the Tocharians than the early Bronze Age Caucasoids have a rich assortment of Caucasoid Y-chromosome haplogroups, whereas the early Bronze Age ones seem to have belonged uniformly to R1a1. What languages were spoken by the non-R1a1 Caucasoids who arrived in the Tarim prior to the Turkification of the region?
To summarize the first part of the argument: the early population of the Tarim does not have clear steppe connections, it may not have been Indo-European speaking, and even if it were, it did not necessarily speak the same language as the later Tocharians. Moreover, the Tocharian language has a vocabulary without clear steppe associations, but with rich agricultural ones.

In search of the Tocharians

We may discover the origin of the Tocharians by a careful sorting of Y-chromosome lineages in the present-day Uyghur population of Xinjiang that is assumed to have absorbed the pre-Turkic inhabitants of the region:
  1. Remove all east Eurasian lineages that are likely to be associated with the Xiongnu, Mongols, or Uyghur
  2. Remove all west Eurasian lineages that can be explained from a non-Tocharian source (such as Iranians, or various Silk Road outliers)
  3. See if anything is left
A recent paper by Zhong et al. provides rich data on Uyghurs that can be used to carry out this program.

The phylogeographic analysis of these lineages does leave some candidates:
  1. Haplogroup D can be excluded as Mongolian/Tibetan
  2. Haplogroup E can be excluded as Mediterranean/African
  3. Haplogroup C can be excluded as Altaic/South Asian (C5)
  4. Haplogroup G2a* (West Asian) does not seem to have an important presence (3 samples)
  5. Haplogroup H can be excluded as South Asian
  6. Haplogroup I can be excluded as a European outlier (1 sample)
  7. Haplogroup J*(xJ2) can be excluded as NE Caucasian/Semitic with small presence (2 samples)
  8. Haplogroup NO; haplogroup N has been founded in a Xiongnu context, so it is likely intrusive; O is East Eurasian
  9. Haplogroup Q is also associated with Xiongnu nomads from Pengyang
This analysis leaves four candidates: J2-M172, R1a1a-M17, R1b-M343, and L-M20.

We can exclude L-M20 because its overall low frequency in most populations makes it difficult, at present, to make a definitive pronouncement on its origin, except perhaps for its Indian L1 clade which is absent here.

J2, present in both its J2a and J2b subclades here at substantial frequencies has an origin in West Asia, as well as a substantial presence among Indo-Iranian speakers. While it is possible (indeed likely, in my opinion) to have been present among the Tocharians, we cannot exclude the possibility that it represents either a specifically Iranian influence, or even something earlier than both.

R1a1a is present in both the steppe, as well as South Asia and West Asia. Its high frequency among some Indo-Iranian populations also makes it difficult to ascribe a specifically Tocharian origin to it.

This leaves only R1b-M343 as a candidate. Have we found a genuine Tocharian genetic signature?

The West Asian roots of R-M343 (?)

R-M343 and its main R-M269 clade are in a sense exasperating: the combination of their widespread distribution from Africa, the Atlantic, to the depths of Inner Asia, combined with their apparent Y-STR-estimated youth make it nearly impossible to associate them with a specific archaeological or historical phenomenon.

Where could R-M269 have come from? It was not present, as far as we can tell, in early Bronze Age Xinjiang, and neither has it been detected in south Siberians. The steppe/"northern" route seems out.

A southern route, from the Indian subcontinent also seems out, as despite its ubiquity elsewhere in Eurasia, it seems to have (mostly) skipped both India and (to an extent) Pakistan.

An indigenous origin seems highly unparsimonious, as it would require that it trek all the way to the Atlantic, but make hardly an impact in either East Asia or South Asia.

As far as I can tell, the only explanation for the presence of R-M343 in Xinjiang is West Asia, or at least Central Asia west of the Tarim. There it can be found at a high frequency in Armenians, Turks, north Iranians, and Lezgins among others. And, unlike both J2 and R1a1a, R-M343 does not seem to be Indo-Iranian (due to its absence in India).

Gamkrelidze and Ivanov cited W. N. Henning to the effect that the ancestors of the Tocharians could be identified with the Gutians from the Zagros, a people that attacked the Sumerians and founded a dynasty. As usual, I don't presume to know the linguistic evidence for this, but this hypothesis would place the ancestors of the Tocharians in the "right spot": virtually all of their Caucasoid Y-chromosome gene pool could be explained with an origin in north Iran.


A model of Tocharian origins

The model of Tocharian origins I present is simplicity itself:

First, Tocharians are descended from a group of farmers that moved east of the PIE homeland and settled on the Zagros and beyond, south of the Caspian sea.

Second, their trek to the Tarim was a simple west-to-east movement along what would later become the Silk Road, beyond the Taklamakan desert and into the Tarim basin. There they must've mixed with the early pre-IE mixed Caucasoid/Mongoloid population of the early Bronze Age. The desert probably sheltered them, to an extent, from encroachments by the Iranians.

An open question remains: were the Tocharians late fugitives who were pushed out of their ancestral homelands by the emergence of the Iranians and entered the Tarim late? Or were they established there fairly early and were the historical Tocharians are the eastern relics of a once great people that was not Iranized unlike most of the people of Central Asia?

Autosomal evidence


The fine-scale analysis of the Dodecad project on a sample of 10 Uyghurs provides some additional evidence:

The Uyghurs seem to lack the Southwest Asian component that is ubuiquitous in most of West Asia today, and may have, in large part, expanded with the more recent spread of Semitic languages. They are similar, in that respect with South Asians, suggesting that neither the spread of Islam to the east nor the cosmopolitanism of the Silk Road were enough to bring this component to the region. Hence, the plethora of Caucasoid Y-haplogroups in the region cannot be attributed to recent arrivals.

The absence of specific South European components in them also suggests that the opinion of some linguists and archaeologists that would see the Tocharians related to Celts and moving from deep within Europe, or even Western Europe to the Tarim, are unlikely; the south European component is ubuiquitous in Europe, and the Uyghurs, like South Asians, seem to lack it entirely.

Their Caucasoid components are primarily West Asian and North European. Projecting them on the East Eurasian/West Asian/North European PCA plot (left), it is clear that they are more West Asian than North European, a result that is in agreement with their ADMIXTURE results.

Notice also how the North European/West Asian ratio is reversed for the more northern-latitude Uralic/Altaic speakers (Selkups, Dolgans, etc.).

Of course, the results should be interpreted with caution, but they seem perfectly in agreement with the model presented here:
  • the Uyghurs are partly Mongoloid both because they may carry the legacy of the ancient mixed population of the Tarim, and also because of their more recent Turkic/Xiongnu associations.
  • with respect to their Caucasoid components, they are mainly West Asian (with the West Asian component also being primary in South Asia), but somewhat shifted to the north due to their absorption of mixed Northern Caucasoid/Mongoloid peoples from the steppelands.

Conclusion

The mystery of the Tocharians may be that there is no mystery. The Tocharians are revealed to have been just another West Asian branch of the Indo-European family that, unlike most of its cousins, went east, absorbed Northern Caucasoid, Mongoloid, and South Asian population elements, emerged long enough in history to leave us a written record of their presence, before succumbing to the Xiongnu and the Mongols.

Thankfully, by combining the remnants of their language, and fragments of their DNA in their descendants, we are able to reconstruct the history of this, once forgotten people

November 07, 2010

Multidimensional scaling and ADMIXTURE across Northern Eurasia corresponds to geography and language

Here is a multi-dimensional scaling plot of a number of North Eurasian populations. In comparison to my previous post, I have excluded Americans and Greenlanders, and added several other populations from Central Asia and West Eurasia.

Population labels have been printed in the co-ordinates of the population averages; these largely correspond with identifiable blobs of colored points, but note that some populations have several outliers, so labels appear in white space. Most notable in that respect are the Koryak, Chukchi, and the Nganasan, all of whom have some apparently European-admixed individuals.


"Mongol" corresponds to Rasmussen et al. (2010) Mongol sample, while "Mongola" to the HGDP-CEPH one. The population codes on the left may not be clearly visible as they overlap with each other and are CEU, LT, HU (relatively unadmixed Caucasoids), FI/RU (Uralian-admixed northern Caucasoids), IR/TR (Altaic-admixed southern Caucasoids). The West Eurasian part of the plot can be seen blown up on the right.

The correspondence with geography and language is striking. Siberian isolates from the extreme north and east, Koryak and Chuckhi are on top; HapMap Chinese at the bottom. Between them are Uralians (Selkup, Yukagir, Nganassan) and Altaics (Mongol-Tungus-Turkic people).

Below is ADMIXTURE analysis for the same set of populations, for K=7:


Finns and Russians seem to have an excess of the "Nganasan" component over the Altaic, while Turks have the opposite. Below is a table of Fst distances between components:


The close relationship between the two Caucasoid components is apparent (Fst=0.033), but note fairly large Fst divergences between the morphologically Mongoloid groups. I attribute this mostly to the very low population sizes of these groups, which have probably affected them by drift. For the less demographically constrained Altaic and East Asian components, Fst=0.044.

If you are not familiar with these ethnic groups, the Red Book of the Peoples of the Russian Empire and the Ethnologue indexes on Altaic and Uralic are invaluable, as are the portraits of ethnic groups of China. On the right a picture of a Nganasan.

UPDATE: Also, a past post from the blog, collating Y-haplogroup N frequencies with anthropological descriptions. Nganasans apparently belong to haplogroup N at a frequency of 92.1%!

October 17, 2010

ADMIXTURE across Eurasia: from Anatolia to Siberia

(Last Update: Oct 17)

Here is a result of an ADMIXTURE run of a few populations from Eurasia (left to right: Turks, Armenians, Georgians, followed by a mix of Uygur, Mongolians, Yakut, Hezhen in no order), combining the HGDP dataset with that of Behar et al. (2010).

It's more of a test, rather than a final result, as I've just finished integrating the two datasets, but it's a nice comparison of a wide assortment of linguistic families.

Notice Turks and Armenians being quite similar to each other, (green+blue), although Turks are differentiated by the presence of an east Eurasian component (5.5%). On the basis of uniparental markers, five years ago, I estimated this component as 6.2% which seems to be right on the money. In the combined Armenian/Georgian sample this admixture is only 0.14% and as can be seen is limited to a handful of Georgian individuals.

It is interesting that Georgians belong semi-uniquely to the green cluster. Turks' non-Mongoloid ancestors were Indo-European speaking like the Armenians still are. It would be tempting to see in the blue-green contrast an Indo-European/Caucasian one, especially as the Caucasoid component further east seems to be mainly blue, in agreement with the idea that it was Indo-Europeans (in particular mainly Iranic speakers) who brought Caucasoid genes to the heartland of Asia.

UPDATE I (Oct 17):

Moving to the north, we see (left-to-right) Han (red), Hungarian/Belorussian (blue), Chuvash (first red "step"), Uzbek (second red "step"). Unlike the Turks, the Hungarians, who also speak a language that came from the east, seem to lack a noticeable east Eurasian component.

Their linguistic conversion was one of elite dominance, where a handful of Mongoloid and quasi-Mongoloid upper echelons left their language but not their genes:
According to his observations, the “overlords” were characterized by Turanid, Uralian and Pamir race elements and also by certain long-headed components. The “middle layer” or “warriors’ layer”, however, showed an anthropological profile distinctly different from that of the overlords. It was essentially constituted by Mediterraneans, Nordoids (who might also have been tall robust Mediterraneans) and Pamir component while the absence of Turanid and Uralian race characteristics was remarkable. As regards the third layer, the so-called “common folk”, they were dominated, just as the middle layer was, by Mediterranean and Nordoid elements but, in addition, the Cromagnoid ones were also significant.
The Chuvash are Turkic and live in Europe, while the Uzbeks, closer to the Altaic homeland in Asia are also Turkic, and have a predictable higher percentage of east Eurasian genes.

September 28, 2010

Some ADMIXTURE estimates in Eurasia

(Last Update: Sep 29)

Continuing my exploration of ADMIXTURE, I turned to the HGDP data, which has 660,918 SNPs for a wide assortment of worldwide populations. After pruning 12,086 SNPs with more than 1% missing genotypes, I was still left with ~650k SNPs.

Here are some experiments on this dataset. First, a clustering with K=2 of Han Chinese, Russians, and Orcadians (left to right)

The emergence of 2 clusters (red=Mongoloid, blue=Caucasoid) is as expected, with Russians showing a small participation in the red cluster (7.2%). These northern Russians are believed to have a substantial Finno-Ugric genetic origin, so this is inline with a recent estimate for the eastern component in the westernmost Finno-Ugric speakers being less than 10% (but see below).

Notice a couple of Chinese individuals with a small Caucasoid component: as I've mentioned before Mongolians, and presumably northern Han have a small Caucasoid component from early movements of Iranian speakers from the west. That's an advantage of doing your own admixture analysis, that you can look at the data at a fine detail, and not rely on the published figures.


Next, a clustering of Orcadians, Uygur, and Han Chinese:
The variable admixture in Uygurs is evident (47.2-63.7%, mean: 54.2%)

Next, a clustering of Druze, Bedouin, and Bantu from Kenya.

Druze appear complete Caucasoid (red), Bantu completely Negroid (save for a couple of individuals), while Bedouins show a quite variable minor Negroid component. This variable African contribution (0-17.6%) makes an elongated cluster out of Bedouins in a recent analysis, pulling them away from other Middle Eastern populations in a Sub-Saharan direction.

Finally, I clustered European populations together with Mandenka and Han Chinese:

The populations are in the following order: Han, Mandenka, Orcadian, French Basque, French, North Italian, Tuscan, Sardinian, Russian.

Here are the admixture proportions:


Notice how the eastern component in Russians is now estimated as 10.9%. This probably reflects the inclusion of French Basque and Sardinians, i.e., populations which have historically no opportunity for eastern Eurasian admixture, rather than only Orcadians. This underscores the importance of having appropriate poles in inter-continental admixture estimates (see Appendix I).

Note also that the 100% value for the Han Chinese is not incompatible with the presence of the two aforementioned Caucasoid-admixed individuals, who are present here with an estimated 1.9% and 0.5% such admixture. However, this contributes little to the sample average of 40+ individuals.

The minor (0.1%) Sub-Saharan admixture in Tuscans and Sardinians is also interesting. As you can guess from the figure, this stems from a handful of individuals (green specks) with less than 1% admixture, which is, however more than the numerical low of 0.001% inferred for most Europeans by the software.


UPDATE I: Eurasian Cline

Below is a run for the following populations (left-to-right: French Basque, Russians, Uygur, Mongolians, Daur, Han Chinese). Notice that the Mongolic-speakers (Mongolian and Daur from HGDP have a small Caucasoid admixture, as I have mentioned before.
APPENDIX I: The importance of choosing poles

The choice of appropriate poles in the estimation of inter-continental admixture is extremely important.

If there is a racial admixture continuum between two major races, such as we observe in Eurasia, then we can express each intermediate population as a weighted sum of populations that live to the east and west of it.

For example, I will use a variable in interval [0, 1] to represent the position in the continuum, with 0: pure western, and 1: pure eastern.

A population at 0.4 can be expressed as the following weighted sum:

0.4 = 0.6*0 + 0.4*1

i.e., as an admixture of 60% western, and 40% eastern.

But, it can also be expressed as e.g.,

0.4 = 0.612*0.02 + 0.388*1

Notice that the choice of a slightly eastward-tilted "western pole" (at position 0.02 in the continuum) has resulted in a reduction of the inferred eastern component (from 40% to 38.8%).

This is exactly what happened in our example: Russian eastern admixture reduced when we used Orcadians, rather than French Basque as the western pole.

Note also, that this is all done automatically: no one told ADMIXTURE to identify these two poles: it was the presence of unlabeled individuals from different ends of the spectrum that influenced the admixture estimates for the rest.

APPENDIX II: Latent populations

Another important point that needs to be remembered has to do with the possible existence of latent ancestral populations.

For example, it is true that Eurasia (minus South Asia) is economically described as a continuum from the Caucasoids of the Atlantic coast to the Mongoloids of the Pacific, with a transition zone in Central Asia and Siberia, and spillovers on either side. But, we cannot exclude the prehistoric existence of other races in the Eurasian landmass that do not exist today in a relatively unadmixed form.

In Eurasia, the Proto-Uralic race was postulated as such a "third race" with features of its own and not reducible to simple Caucasoid-Mongoloid admixture. It is difficult to see whether these features are ancestral peculiarites (prior to admixture with Caucasoids and Mongoloids), or if they have arisen in a mixed Caucasoid-Mongoloid population.

It is also important to understand how such latent populations affect genetic continua:

First, if the latent population is equidistant from the two major races, then its admixture has no effect on an individual's position in the continuum between the two races. However, it is possible that the latent population was more related to one of the two major races. In that case, admixture with it will move a population towards that race.

So while the jury is still out about the existence of a Proto-Uralic race in Eurasia, its effects on admixed populations indicates that if it had existed it was genetically closer to Mongoloids than to Caucasoids.

July 02, 2010

Admixture in Uyghurs (again)

Gene Expression points me to this letter which suggests that Western Eurasian admixture in Uyghurs has been overestimated in previous studies. They base their claim on the alleged low population coverage of previous studies. I don't buy their argument, mainly because of Behar et al. (2010) which has as good a coverage of Eurasia as one might hope for, and finds (expectedly) Uyghurs to be about a 50-50 Caucasoid/Mongoloid mix. The same is also true for Hazara, shown by the authors as majority Mongoloid, but revealed by Behar et al. (2010) to be about 50-50 as well.

The authors make another claim:
STRUCTURE cannot distinguish recent admixture from a cline of other origin, and these analyses cannot prove admixture in the Uyghurs; however, historical records indicate that the present Uyghurs were formed by admixture between Tocharians from the west and Orkhon Uyghurs (Wugusi-Huihu, according to present Chinese pronunciation) from the east in the 8th century CE.14 The Uyghur Empire was originally located in Mongolia and conquered the Tocharian tribes in Xinjiang. Tocharians such as Kroran have been shown by archaeological findings to appear phenotypically similar to northern Europeans,15 whereas the Orkhon Uyghur people were clearly Mongolians. The two groups of people subsequently mixed in Xinjiang to become one population, the present Uyghurs. We do not know the genetic constitution of the Tocharians, but if they were similar to western Siberians, such as the Khanty, admixture would already be biased toward similarity with East Asian populations.
First of all, the authors forget about the eastern Iranian peoples (Sakas) who most surely were present among the ancestors of the Uyghurs. Second, there is no reason to think that Tocharians were similar to the Khanty, a population with a substantial presence of northern Eurasian Y-haplogroup N does not make a link with Uyghurs likely.

Am J Hum Genet. 2009 December 11; 85(6): 934–937.

Genetic Landscape of Eurasia and “Admixture” in Uyghurs

Hui Li et al.

Link

November 23, 2009

Genetic Variation and Recent Positive Selection with 1 million SNPs

On the left Figure S4 shows PCA and frappe analysis for Eurasia. From the paper:
When just the Central/South Asia, Middle East, North Africa, and European groups are analyzed, PC1 (Fig. S4A) distinguishes the Mozabite (North Africa), Middle East, and Europe groups from the Central/South Asian groups, while PC2 separates the Mozabite and Middle East groups from the Europe groups, with no overlap among individuals from the different North Africa/Middle East/Europe groups. By contrast, there is overlap among individuals from the different Central/South Asia groups; in addition, the Makrani and Sindhi individuals identified in the worldwide analysis as having experienced recent sub-Saharan African admixture are clearly differentiated by PC2. The frappe analysis at K = 5 (Fig. S4B) indicates ancestry components corresponding to the Mozabite, Kalash, Hazara/Uygur, other Central/South Asia, and Europe groups. The three Middle East groups have varying amounts of the Europe, Mozabite, and Central/South Asia ancestry components. The three Italian groups are alone among European groups in having low amounts of the Mozabite ancestry component, possibly indicating gene flow across the Mediterranean. The Sardinians differ from continental European groups in lacking any Asian ancestry component, while the Russians and Adygei differ from other European groups in having appreciable amounts of the Hazara/Uygur and other Central/South Asia ancestry components, respectively, indicating more gene flow and/or ancestry with these groups (Fig. S4B).
Also from the paper, referring to Figure 3 and Figure S3:
Second, many of the subsequent statistically-significant PCs (Fig. 3 and Fig. S3) distinguish among various combinations of the sub-Saharan African groups (or among individuals within such groups), despite the fact that there are only six such groups in the analysis. This disproportionate impact of structure within sub-Saharan Africa on analyses of worldwide genetic diversity clearly emphasizes both the importance of such structure and the great need for further in-depth genetic characterization of sub-Saharan African populations [22]; we would hardly expect that these six groups encompass all of the genetic diversity in sub-Saharan Africa.

PLoS ONE doi:10.1371/journal.pone.0007888

Genetic Variation and Recent Positive Selection in Worldwide Human Populations: Evidence from Nearly 1 Million SNPs

David López Herráez et al.

Abstract

Background
Genome-wide scans of hundreds of thousands of single-nucleotide polymorphisms (SNPs) have resulted in the identification of new susceptibility variants to common diseases and are providing new insights into the genetic structure and relationships of human populations. Moreover, genome-wide data can be used to search for signals of recent positive selection, thereby providing new insights into the genetic adaptations that occurred as modern humans spread out of Africa and around the world.

Methodology
We genotyped approximately 500,000 SNPs in 255 individuals (5 individuals from each of 51 worldwide populations) from the Human Genome Diversity Panel (HGDP-CEPH). When merged with non-overlapping SNPs typed previously in 250 of these same individuals, the resulting data consist of over 950,000 SNPs. We then analyzed the genetic relationships and ancestry of individuals without assigning them to populations, and we also identified candidate regions of recent positive selection at both the population and regional (continental) level.

Conclusions
Our analyses both confirm and extend previous studies; in particular, we highlight the impact of various dispersals, and the role of substructure in Africa, on human genetic diversity. We also identified several novel candidate regions for recent positive selection, and a gene ontology (GO) analysis identified several GO groups that were significantly enriched for such candidate genes, including immunity and defense related genes, sensory perception genes, membrane proteins, signal receptors, lipid binding/metabolism genes, and genes involved in the nervous system. Among the novel candidate genes identified are two genes involved in the thyroid hormone pathway that show signals of selection in African Pygmies that may be related to their short stature.

Link

November 22, 2009

Ancestry-related assortative mating in latino populations (Risch et al. 2009)

When different races admix, then in the first few generations there is a spectrum of ancestry proportions, ranging from pure individuals of the constituent races to admixed individuals with varying proportions of ancestry.

If there is random mating, then over several generations all individuals tend to have similar ancestry proportions, determined by the number of founders from the two constituent races. Mix 30,000 Europeans with 70,000 Africans, randomly mate them for 10-20 generations, and pretty soon almost everyone will have 30:70 European/African ancestral proportions with a little variation.

However, if there is assortative mating, then this process takes much longer to complete, as matings of individuals with very different ancestry proportions are rare, and the spectrum of varying individual ancestry is maintained. In the above-mentioned example, if there is perfect assortative mating, then after 10-20 generations you will still have 30% of the population having 100% European genes, and 70% of them having 100% African ones.

Previously, I had argued that the fact that Latin Americans, unlike Central Asian Turkic populations (such as the Uyghurs), have such a wide spectrum of ancestry proportions is due to the more recent admixture in the Americas than in Central Asia (less time for homogenization to take place), and the continued importation of Europeans.

Assortative mating is a third factor that may be behind this phenomenon. A stronger parallel may be found in South Asia, where the two constituents have been co-existing for a much longer time, but under a rigid, formalized regime of assortative mating (the caste system), homogenization has not taken place.

Genome Biology doi:10.1186/gb-2009-10-11-r132

Ancestry-related assortative mating in latino populations

Neil Risch et al.

Abstract

Background

While spouse correlations have been documented for numerous traits, no prior studies have assessed assortative mating for genetic ancestry in admixed populations.

Results

Using 104 ancestry informative markers, we examined spouse correlations in genetic ancestry for Mexican spouse pairs recruited from Mexico City and the San Francisco Bay Area, and Puerto Rican spouse pairs recruited from Puerto Rico and New York City. In the Mexican pairs, we found strong spouse correlations for European and Native American ancestry, but no correlation in African ancestry. In the Puerto Rican pairs, we found significant spouse correlations for African ancestry and European ancestry but not Native American ancestry. Correlations were not attributable to variation in socioeconomic status or geographic heterogeneity. Past evidence of spouse correlation was also seen in the strong evidence of linkage disequilibrium between unlinked markers, which was accounted for in regression analysis by ancestral allele frequency difference at the pair of markers (European versus Native American for Mexicans, European versus African for Puerto Ricans). We also observed an excess of homozygosity at individual markers within the spouses, but this provided weaker evidence, as expected, of spouse correlation. Ancestry variance is predicted to decline in each generation, but less so under assortative mating. We used the current observed variances of ancestry to infer even stronger patterns of spouse ancestry correlation in previous generations.

Conclusions

Assortative mating related to genetic ancestry persists in Latino populations to the current day, and has impacted on the genomic structure in these populations.

Link

June 30, 2009

Uyghurs as an admixed not source population (Xu et al. 2009)

This paper is interesting not so much because it estimates admixture in Uyghurs (click on the post label for previous studies on the topic), but because it explicitly rejects the hypothesis that they are a source ("donor") population.

If a population has substantial genetic variation which overlaps with that of two other groups, then there are two possible interpretations:
  1. It represents the population from which the other two groups sprang, or at least contributed genes to both of them
  2. It represents a mixture of the two other groups
What this paper does, is to show that Uyghurs are best explained as a mixture of Caucasoids and Mongoloids (#2) rather than #1.

Molecular Biology and Evolution, doi:10.1093/molbev/msp130

Haplotype Sharing Analysis Showing Uyghurs Are Unlikely Genetic Donors

Shuhua Xu et al.

Abstract

The Uyghur are a group of people primarily residing in Xinjiang of China which is geographically located in Central Asia, from where modern humans were presumably spread in all directions reaching Europe, east and northeast Asia about 40 kya. A recent study suggested that the Uyghur are ancestry donors of the East Asian gene pool. However, an alternative hypothesis, i.e. the Uyghur is an admixture population with both East Asian (EAS) and European (EUR) ancestries is also supported by our previous studies. To test the two competing hypotheses, here we conducted a haplotype sharing analysis based on empirical and simulated data of high density single nucleotide polymorphisms (SNPs). Our results showed that more than 95% of Uyghur (UIG) haplotypes could be found in either East Asian (EAS) or European (EUR) populations, which contradicts the expectation of the null models assuming that UIG are donors. Simulation studies further indicated that the proportion of UIG private haplotypes observed in empirical data is only expected in alternative models assuming that UIG is an admixture population. Interestingly, the estimated ancestry contribution of 44%:56% (EAS:EUR) based on haplotype sharing analysis is consistent with our previous estimation with STRUCTURE analysis. Although the history of Uyghurs could be complex, our method is explicit and conservative in rejecting the null hypothesis. We concluded that the gene pool of modern Uyghurs is more likely a sole recipient with contribution from both EAS and EUR.

Link

May 14, 2009

Admixture in Mexican Mestizos

Gene Expression points me towards a new open access paper in PNAS about genetic diversity in Mexican Mestizos:
We analyzed data from 300 nonrelated self-identified Mestizo individuals from 6 states located in geographically distant regions in Mexico: Sonora (SON) and Zacatecas (ZAC) in the north, Guanajuato (GUA) in the center, Guerrero (GUE) in the center– Pacific, Veracruz (VER) in the center–Gulf, and Yucatan (YUC) in the southeast. Considering that Zapotecos have been shown as a good ancestral population for predicting Amerindian (AMI) ancestry in Mexican Mestizos (16), we included 30 Zapotecos (ZAP) from the southwestern state of Oaxaca (Fig. 1). For comparative purposes, we included similar data sets from HapMap populations: northern Europeans (CEU), Africans (YRI), and East Asians (EA), including Chinese (CHB) and Japanese (JPT).
As expected, Mestizo admixture is mainly between Caucasoids and Amerindians, with a very little Sub-Saharan African thrown in at the individual level. Moreover, as with many populations, such as the Uyghur, where admixture took place several generations ago, individual admixture levels are fairly uniform, with very few individuals deviating strongly towards either the Caucasoid or Amerindian end of the spectrum.

The variation in individual admixture appears only somewhat stronger than in the Uyghur, which may be explained either by the smaller number of markers used here, making the assessment of admixture "noisier", or alternatively might be the result of the fact that admixture in Mexican Mestizos happened more recently, and immigration into the Americas from Europe continued hence, hence the homogenization of the population is still ongoing.

Table S1 from the Supplementary material (pdf) shows the exact admixture proportions in the studied Mestizo populations and the HapMap populations.

Related:

PNAS doi:10.1073/pnas.0903045106

Analysis of genomic diversity in Mexican Mestizo populations to develop genomic medicine in Mexico

Irma Silva-Zolezzi et al.

Abstract

Mexico is developing the basis for genomic medicine to improve healthcare of its population. The extensive study of genetic diversity and linkage disequilibrium structure of different populations has made it possible to develop tagging and imputation strategies to comprehensively analyze common genetic variation in association studies of complex diseases. We assessed the benefit of a Mexican haplotype map to improve identification of genes related to common diseases in the Mexican population. We evaluated genetic diversity, linkage disequilibrium patterns, and extent of haplotype sharing using genomewide data from Mexican Mestizos from regions with different histories of admixture and particular population dynamics. Ancestry was evaluated by including 1 Mexican Amerindian group and data from the HapMap. Our results provide evidence of genetic differences between Mexican subpopulations that should be considered in the design and analysis of association studies of complex diseases. In addition, these results support the notion that a haplotype map of the Mexican Mestizo population can reduce the number of tag SNPs required to characterize common genetic variation in this population. This is one of the first genomewide genotyping efforts of a recently admixed population in Latin America.

Link

January 23, 2009

Another paper on Ashkenazi Jewish distinctiveness

Gene Expression points me to a new paper which demonstrates that Ashkenazi Jews can be distinguished perfectly from non-Jewish European Americans. This was previously seen, but it is nice to see it confirmed once more. Some previous studies:
  1. New paper on genomic differences between Ashkenazi Jews and Europeans
  2. 300K SNP paper on European genetic substructure
  3. European population substructure revealed by genetics
The authors also observed that heterozygosity among Ashkenazi Jews is higher than in European Americans. This is fairly conclusive evidence that Ashkenazi Jews are not so distinctive because they passed through a genetic bottleneck that shifted their allele frequencies in a direction specific to their group.

On the other hand, the conclusion that the genetic distinctness of Ashkenazi Jews is due to Middle Eastern ancestry is not demonstrated by this study. For example, the Uyghur of Central Asia are distinct from both East Asians and Caucasoids, but this is not due to any mysterious "Central Asian" component in them, but rather due to the fact that they are an admixed population of Caucasoids and Mongoloids.

To demonstrate the specific Middle Eastern background of Ashkenazi Jews, it would be a good idea to also study other Jewish groups. If it is shown that the various Jewish groups possess a common autosomal genetic component, then the simplest explanation would be that this component stems from the ancestral Jewish population of the 1st millennium AD, prior to the separation of the various Jewish groups from each other.

Moreover, the genetic distinctiveness of Ashkenazi Jews does not in itself say anything about the extent of Middle Eastern ancestry in this group. For example, in this paper, the Middle Eastern groups (mostly non-Jewish Semitic groups recruited in Israel) were different from other Caucasoids by the possession of a specific ancestral component (color-coded brown), but the extent of this component differed among them.

With that said, I do suspect that the distinctiveness of the Ashkenazi Jews is in part due to the possession of a Middle Eastern component of unspecified strength. I base this hypothesis on the results reported to me about the EURO-DNA-CALC test. This test distinguishes between NW, SE Europeans and Ashkenazi Jews; a few Arab individuals who have communicated their results to me have reported fairly high AJ components, indicating that part of what distinguishes an AJ from Europeans is related to the Middle Eastern Semitic background of that group.

The way forward is of course to perform a comprehensive admixture analysis where Europeans, various Jewish groups, and various non-Jewish Middle Eastern groups will be represented. That is the only way to ascertain the ancestral components of the various Jewish groups. Moreover, such an analysis would establish the extent of the common genomic element between the various Jewish groups, which -so far- has been established for a limited number of Y-chromosome and mtDNA lineages.

UPDATE: While a formal admixture analysis is not performed, the EIGENSOFT plot is suggestive of what common sense would dictate, namely that Ashkenazi Jews (reds) are intermediate between a native Near Eastern group (the Druze) and Europeans. Unfortunately the inclusion of the Mozabites (off the chart to the left) who have substantial Sub-Saharan ancestry, makes the resolution of the visible part of the chart less than desirable.



Genome Biology doi:10.1186/gb-2009-10-1-r7

A genome-wide genetic signature of Jewish ancestry perfectly separates individuals with and without full Jewish ancestry in a large random sample of European Americans

Anna C Need et al.

Abstract

Background

It was recently shown that the genetic distinction between self-identified Ashkenazi Jewish and non-Jewish individuals is a prominent component of genome-wide patterns of genetic variation in European Americans. No study however has yet assessed how accurately self-identified (Ashkenazi) Jewish ancestry can be inferred from genomic information, nor whether the degree of Jewish ancestry can be inferred among individuals with fewer than four Jewish grandparents.

Results

Using a principal components analysis, we found that the individuals with full Jewish ancestry formed a clearly distinct cluster from those individuals with no Jewish ancestry. Using the position on the first principal component axis, every single individual with self-reported full Jewish ancestry had a higher score than any individual with no Jewish ancestry.

Conclusions

Here we show that within Americans of European ancestry there is a perfect genetic corollary of Jewish ancestry which, in principle, would permit near perfect genetic inference of Ashkenazi Jewish ancestry. In fact, even subjects with a single Jewish grandparent can be statistically distinguished from those without Jewish ancestry. We also found that subjects with Jewish ancestry were slightly more heterozygous than the subjects with no Jewish ancestry, suggesting that the genetic distinction between Jews and non-Jews may be more attributable to a Near-Eastern origin for Jewish populations than to population bottlenecks.

Link (pdf)

December 08, 2008

A visual display of biological and social race

This is as clear display of the difference between biological and social race.

Populations from the three major human biological races (European Americans from Caucasoids, Yoruba from Negroids, Japanese/Chinese from Mongoloids) are clearly separable, with no overlap.

The "black race" to which African Americans are said to belong is seen as an almost perfect linear combination of Caucasoids and Negroids. It is not a biological race, but rather the result of admixture between the two races.

The same can be seen in other admixed groups such as the Uyghur, who are a combination of Caucasoids and Mongoloids. In that case, however, the admixture is more ancient, and the opportunity to further mix with representatives of the unadmixed groups is more limited. Therefore, the blend has been completed, and most individuals have similar admixture proportions from the ancestral groups. African Americans, on the other hand are much more variable in their individual ancestry components, from ~100% Negroid, to more Caucasoid than Negroid.

PLoS Genetics doi: 10.1371/journal.pgen.1000294

Effects of cis and trans Genetic Ancestry on Gene Expression in African Americans

Alkes L. Price et al.

Abstract

Variation in gene expression is a fundamental aspect of human phenotypic variation. Several recent studies have analyzed gene expression levels in populations of different continental ancestry and reported population differences at a large number of genes. However, these differences could largely be due to non-genetic (e.g., environmental) effects. Here, we analyze gene expression levels in African American cell lines, which differ from previously analyzed cell lines in that individuals from this population inherit variable proportions of two continental ancestries. We first relate gene expression levels in individual African Americans to their genome-wide proportion of European ancestry. The results provide strong evidence of a genetic contribution to expression differences between European and African populations, validating previous findings. Second, we infer local ancestry (0, 1, or 2 European chromosomes) at each location in the genome and investigate the effects of ancestry proximal to the expressed gene (cis) versus ancestry elsewhere in the genome (trans). Both effects are highly significant, and we estimate that 12±3% of all heritable variation in human gene expression is due to cis variants.

Link

November 02, 2008

23andme's advanced global similarity tool

UPDATE: I am told that this tool is currently in alpha version, so it's not clear when it will be fully ready for 23andme customers. As per my comments below, I think this is a great initiative to tie individual customers' genetic data to the many new genetic studies showing genomic-geographic correlations. I am sure that 23andme's blog, the Spittoon, will cover this when it is ready for public release, including any features that I may have overlooked. I will be following this story closely. [end update]

23andme has added a new advanced global similarity tool to their website (you need to register in order to play with it). This tool places a customer, as well as other customers he is "connected" with on the map of the first two principal components like the ones recently published in several papers.

The tools allows one to look at the PC map at the global, continental, or subcontinental level.


This is quite useful, and a right step in the direction I pointed out earlier. However, there are some points of criticism.
  • The axes are labeled North/South Migration and East/West Migration. While the pattern in the first two principal components does correspond roughly with longitude and latitude, it is erroneous to label these principal components as "North/South" and "East/West". It is even more erroneous to label them as "Migration", since a geographical cline is not necessarily produced by a migration event.
  • The "Take a Tour" feature presents a simplistic and misleading account of human prehistory in terms of "migrations". This account is a simple branching pattern, e.g., Africa -> Near East Europe, or Africa -> Near East -> Central Asia -> East Asia. The observed pattern did not emerge in this manner. For example, Central Asian people such as the Uyghur are intermediate between Western Eurasians (Caucasoids) and Eastern Eurasians (Mongoloids) because of a later admixture event; they can't be thought of as "ancestors" of the East Eurasians.
  • Partitioning human variation into this hierarchical set of groups is not the best way to satisfy customers' needs. For example, a Hispanic person may wish to see himself on a PC map which includes "Southern European" and "Native American" groups, an African American person may wish to see himself on a PC map which includes "Northern European" and "West African" groups, an Ethiopian, on a Sub-Saharan/Near Eastern map, while a European Jew on a European/Near Eastern map. Of course, there is a combinatorial number of possible combinations, but there is no reason why some of the more common ones (customer feedback may play a role here) many not be supported.
  • Why should this tool be limited to the first two principal components? Of course, additional components do not have such a strong geographical correspondence, but they -nonetheless- will separate populations in different ways, and allow individuals to place themselves more fully in context.
  • The tool could offer much more information. On mouse hover over an individual, a small label identifying it (e.g. origin and HGDP code), and listing its PC coordinates could appear. This is especially useful for power users. A pretty uncluttered picture is no substitute for as much information as possible.

August 30, 2008

Admixture mapping in Uyghurs


This is a followup to this earlier study. Note that the "European" label for the Caucasoid component in Uyghurs is inappropriate, since this is composed of href="http://dienekes.blogspot.com/2008/02/huge-paper-on-human-genetic.html">two distinct "European" and "Caucasoid Central Asian" elements.

From the paper:
Figure 3A shows summary plot of individual admixture proportions based on the highest-probability run of ten STRUCTURE runs. The results show that individuals from the same population often share membership coefficients in the inferred cluster, with the exception that one Japanese outlier shows obvious admixture. Mongola, Adygei, and Russian individuals show some degree of admixture as well.
Most of the EAS admixture in the Adygei from the Caucasus seems mostly spurious, as the Adygei have a substantial "Central Asian" Caucasoid component (38%) rather than Mongoloid admixture (2%).

Note that, as in the previous study, the Uyghur individuals seem to have similar proportions of "Western" and "Eastern" genes, due to the fact that the blend which produced them is fairly old and there are really no individuals in which either of the two components predominate.

The American Journal of Human Genetics, doi:10.1016/j.ajhg.2008.08.001

A Genome-wide Analysis of Admixture in Uyghurs and a High-Density Admixture Map for Disease-Gene Discovery

Shuhua Xu and Li Jin

Abstract



Following up on our previous study, we conducted a genome-wide analysis of admixture for two Uyghur population samples (HGDP-UG and PanAsia-UG), collected from the northern and southern regions of Xinjiang in China, respectively. Both HGDP-UG and PanAsia-UG showed a substantial admixture of East-Asian (EAS) and European (EUR) ancestries, with an empirical estimation of ancestry contribution of 53:47 (EAS:EUR) and 48:52 for HGDP-UG and PanAsia-UG, respectively. The effective admixture time under a model with a single pulse of admixture was estimated as 110 generations and 129 generations, or admixture events occurred about 2200 and 2580 years ago for HGDP-UG and PanAsia-UG, respectively, assuming an average of 20 yr per generation. Despite Uyghurs' earlier history compared to other admixture populations, admixture mapping, holds promise for this population, because of its large size and its mixture of ancestry from different continents. We screened multiple databases and identified a genome-wide single-nucleotide polymorphism panel that can distinguish EAS and EUR ancestry of chromosomal segments in Uyghurs. The panel contains 8150 ancestry-informative markers (AIMs) showing large frequency differences between EAS and EUR populations (FST > 0.25, mean FST = 0.43) but small frequency differences (7999 AIMs validated) within both populations (FST < 0.05, mean FST < 0.01). We evaluated the effectiveness of this admixture map for localizing disease genes in two Uyghur populations. To our knowledge, our map constitutes the first practical resource for admixture mapping in Uyghurs, and it will enable studies of diseases showing differences in genetic risk between EUR and EAS populations.

Link

April 23, 2008

Origin of Yugur subclans

Ann Hum Biol. 2008 Mar-Apr;35(2):198-211.

Origin and evolution of two Yugur sub-clans in Northwest China: a case study in paternal genetic landscape.

Zhou R, Yang D, Zhang H, Yu W, An L, Wang X, Li H, Xu J, Xie X.

Background: Yugur is an ethnic group that was officially identified by the Chinese Government in 1953. Within the population there are two sub-clans distinctly identified as the Eastern Yugur and Western Yugur, partly because they have different local languages. Aim: A parentage comparison was conducted between the two sub-clans to investigate their genetic relationship. Subjects and methods: Male subjects were chosen from the two clans to investigate their paternal genetic landscape through typing 14 single nucleotide polymorphisms (SNP) and 12 short tandem repeats (STR) of the Y chromosome. Results: Significant differences were revealed between the sub-clans at the haplogroup level. Genetic divergence was also observed by analyses of multidimensional scaling (MDS) and principal components (PC). Genetically, the Eastern Yugur are closer to the Han Chinese and Mongolian people than the Western Yugur. The Uygur people, who share a common ancestor (ancient Huihu) with the Yugur, were genetically separate from both sub-clans of Yugur. Moreover, the constructed phylogenetic network for haplogroup O provided further evidence that the two Yugur sub-groups present an underlying genetic difference. Conclusion: Overall, the diffusion of Mongolians during the Mongol Period has affected the Eastern Yugur more than the Western Yugur. The genetic contribution of the Han people to the Eastern Yugur seems to be more pronounced than to the Western Yugur. Besides the two different contributions referred to above, small population size and genetic drift have resulted in the genetic differentiation of the current sub-clans of Yugur.

Link

March 25, 2008

Origins of the Uighur

Interesting bit from the paper:
Notably, the distribution of admixture proportions among UIG individuals is relatively even, with 48.7% the lowest admixture from European ancestry and the highest 62.2%. The standard deviation is only 3.8%, which is much smaller than the estimation for the African-American (AfA) population,58 suggesting a much longer history of admixture events for the Uyghur population compared with the AfA population.


The American Journal of Human Genetics, doi:10.1016/j.ajhg.2008.01.017

Analysis of Genomic Admixture in Uyghur and Its Implication in Mapping Strategy

Shuhua Xu et al.

Abstract

The Uyghur (UIG) population, settled in Xinjiang, China, is a population presenting a typical admixture of Eastern and Western anthropometric traits. We dissected its genomic structure at population level, individual level, and chromosome level by using 20,177 SNPs spanning nearly the entire chromosome 21. Our results showed that UIG was formed by two-way admixture, with 60% European ancestry and 40% East Asian ancestry. Overall linkage disequilibrium (LD) in UIG was similar to that in its parental populations represented in East Asia and Europe with regard to common alleles, and UIG manifested elevation of LD only within 500 kb and at a level of 0.1 < style="font-weight: bold;">we estimated that the admixture event of UIG occurred about 126 [107∼146] generations ago, or 2520 [2140∼2920] years ago assuming 20 years per generation. In spite of the long history and short LD of Uyghur compared with recent admixture populations such as the African-American population, we suggest that mapping by admixture LD (MALD) is still applicable in the Uyghur population but ∼10-fold AIMs are necessary for a whole-genome scan.

Link

February 11, 2005

How Turkish are the Anatolians?

The Anatolians are the ethnic descendants of both the indigenous populations of Asia Minor who converted to Islam (and were thus spared from the genocidal campaign of the Ottomans and Kemalists during the early 20th century), and also of non-indigenous populations from the Balkans, the Middle East, and Central Asia. From Central Asia came the Turks, who were the main agent for the Islamization and during the last century Turkification of Asia Minor.

To what extent are the Anatolians descended from Central Asian Turks? The study of Cinnioglu et al. (2004) discovered an occurrence of 3.4% of Mongoloid Y-chromosomal haplogroups in Anatolia (haplogroups Q, O, and C).

According to Tambets et al. (2004) the occurrence of Mongoloid haplogroups in present-day Central Asian Turkic Altaic speakers (Altaians) is at least 40%, with an additional 10% which might belong to haplogroup O which was not tested in this study. According to Zerjal et al. (2002) this percentage is for various Turkic speakers: Kyrgyz (22%), Dungans (32%), Uyghurs (33%), Kazaks (86%), Uzbeks (18%).

It is clear that the percentage of Mongoloid ancestry among the Turkic speakers is very variable, yet it is clear that the Proto-Turks must have been partially Mongoloid in lieu of the fact that all current Turkic speakers possess some Mongoloid admixture. The average of the six Central Asian population samples listed above is 38.5% and may serve as a first-order estimate of the paternal contribution of early Turks, who (judging by their modern descendants in Central Asia) were more Caucasoid paternally and more Mongoloid maternally.

Using the figure of 38.5%, the paternal contribution of Turks to the Anatolian population is estimated to about 11%. In lieu of the approximation, allowing for 33% relative error in either direction for both the true frequency of Mongoloid lineages in Anatolia and in early Turks, we obtain a range of 6-22%. It would thus appear that the Turkish element is a minority one in the composition of the Anatolians, but it is by no means negligible.