Showing posts with label Tungus. Show all posts
Showing posts with label Tungus. Show all posts

December 23, 2013

mtDNA and Y chromosomes of Tungus

PLoS ONE 8(12): e83570. doi:10.1371/journal.pone.0083570

Investigating the Prehistory of Tungusic Peoples of Siberia and the Amur-Ussuri Region with Complete mtDNA Genome Sequences and Y-chromosomal Markers

Ana T. Duggan et al.

Evenks and Evens, Tungusic-speaking reindeer herders and hunter-gatherers, are spread over a wide area of northern Asia, whereas their linguistic relatives the Udegey, sedentary fishermen and hunter-gatherers, are settled to the south of the lower Amur River. The prehistory and relationships of these Tungusic peoples are as yet poorly investigated, especially with respect to their interactions with neighbouring populations. In this study, we analyse over 500 complete mtDNA genome sequences from nine different Evenk and even subgroups as well as their geographic neighbours from Siberia and their linguistic relatives the Udegey from the Amur-Ussuri region in order to investigate the prehistory of the Tungusic populations. These data are supplemented with analyses of Y-chromosomal haplogroups and STR haplotypes in the Evenks, Evens, and neighbouring Siberian populations. We demonstrate that whereas the North Tungusic Evenks and Evens show evidence of shared ancestry both in the maternal and in the paternal line, this signal has been attenuated by genetic drift and differential gene flow with neighbouring populations, with isolation by distance further shaping the maternal genepool of the Evens. The Udegey, in contrast, appear quite divergent from their linguistic relatives in the maternal line, with a mtDNA haplogroup composition characteristic of populations of the Amur-Ussuri region. Nevertheless, they show affinities with the Evenks, indicating that they might be the result of admixture between local Amur-Ussuri populations and Tungusic populations from the north.

Link

November 03, 2012

Recent admixture in Altaic populations: a legacy of Empire?

Continuing my experiments with ALDER, I took every single Altaic population publicly available, i.e., the following 25 populations:
Altai, Balkars_Y, Buryat, Chuvashs_16, Daur, Dolgan, Evenk_15, Hezhen, Kumyks_Y, Kyrgyz_Bishkek_Ho, Mongol, Mongola, Nogais_Y, Oroqen, Tu, Turkish_Aydin_Ho, Turkish_Istanbul_Ho, Turkish_Kayseri_Ho, Turkmens_Y, Turks, Tuva, Uygur, Uzbeks, Xibo, Yakut
I also took three West Eurasian populations unlikely to have historical East Asian admixture (French, French_Basque, and Sardinians), and three East Eurasian populations unlikely to have historical West Eurasian admixture (Dai, She, Miaozu). I merged all of the above in PLINK with a --geno 0.03 flag, and extracting SNPs present in the Rutgers recombination map for Illumina chips (a total of 524,822 SNPs).

I then ran ALDER for all 25 Altaic populations using any of the 3*3 West/East Eurasian reference pairs, or a total of 25*3*3= 225 runs. I retained only those 2-ref admixture analyses for which ALDER reported "success" with no warnings.

I then converted reported times to calendar dates: a generation of 29 years was assumed; lacking information about the age of the sampled individuals, I assumed that the "present" is 1980; finally, I report the earliest and latest -/+ limits of any confidence interval, as well as the median of all estimates.

The results can be seen below; for 11 of the 25 populations there was at least one test which was successful with no warnings. This does not mean that the other populations are unadmixed, but the following cases appear to be most "well-behaved":


Now, these appear to make excellent sense.

Of the Dolgans:
There also existed a group of Russian settlers on the River Heta, who, by the end of the 19th century, had become Dolganized and had gradually adopted the way of life of nomadic reindeer breeders. ... The tribes forming the nucleus of the Dolgans migrated from the banks of the River Lena at the end of the 17th century. One of the reasons for migration was the fact that Russian goods, flour, for instance, were coming to the Taimyr Peninsula by the boats on the Lena.
The 1770-1860AD range for the admixture appears to coincide with the period where the Dolgans came under Russian influence.

Of the Evenks:
The history of the Evenks' habitation can be traced in detail from the 17th century on. At that time the Evenks left several of their previous territories, for instance, the River Angara, when the Yakut, the Buryat and the Russians appeared in the province. The Evenks had especially bad relations with the Yakuts, who had settled in the river basin of the Lena in the 13th century. In the 18th and 19th centuries the Evenks living there adopted the Yakut language. In the Baikal area the Evenks began to speak the Buryat and the Mongolian languages, and even converted to lamaism. The southern Evenk -- the Manegir, the Birar, the Solon -- were influenced by the Manchu, Daur and Chinese cultures. The arable lands in Siberia were occupied by Russian settlers, migrating there in the 17th century, and those Evenks, living in the vicinity on the upper reaches of the Lena and near Baikal, were russified.
Again, the  1630-1800AD admixture range seems consistent with the time when Evenks came into contact with Russians.

Of the Nogais:
 In the first half of the 17th century a number of Nogay tribes were nomadic on the steppes between the Danube and the Caspian. The invasion of the warlike Kalmyks forced several of the Nogay tribes to leave their home steppes and withdraw to the foothills of the North Caucasus. By the River Kuban they met with the Cherkess.  In the Moscow chronicles from the 16th and 17th centuries there are several mentions of the Nogay, including the two Nogay Hordes, the Great and the Small. The former roamed beyond the River Volga, the latter somewhat to the west. Both had numerous military encounters with the Russians. In the 17th century some of the Nogay chiefs entered into an alliance with Moscow and fought at times together with the Russians against the Kabardians, the Kalmyks and peoples of Dagestan. 
 The 1610-1730AD range intersects the period when the Nogais settled in the North Caucasus and interacted with North Caucasians and Russians.

Not much needs to be said for the admixture signal in the Uygur, Uzbek, Kyrgyz, and Mongols which collectively ranges from 1260-1500AD. This was a period of Mongol power when Mongolian and Turkic speaking peoples assumed control over Central Asia and replaced to a great degree the previous inhabitants of the area.

The origin of the Balkars is less certain, because they are an old Turkic group that settled in the Caucasus, but the admixture (830-1220AD) date seems plausible. So does, of course, that of the Turks from Caesaria (990-1260AD) which parallels those of my recent experiment, and can be associated with the takeover of Anatolia following the Battle of Manzikert. Finally, I don't have a read explanation for the 11-12th century signal of admixture in the Siberian Altai and Buryat, but presumably it has something to do with the expansions of Altaic peoples around that time that were also felt in the west during this period; presumably, this involved some type of mixture with Caucasoid groups in Siberia.

The admixture dates are quite helpful in helping us better interpret other signals of admixture such as those of ADMIXTURE analyses (e.g., globe13). For example, the Dolgan have 13.1% North_European in that experiment, and the Altai have 13.2%, but apparently this occurred centuries apart and may have involved different groups of West Eurasian people.

In conclusion, ALDER seems to find some quite plausible dates for major admixture episodes in the history of Altaic populations that are compatible with fairly recent historical events.

July 22, 2012

Clarifying the phylogeny of Y-chromosome haplogroup C3c

A short and to the point paper that addresses the issue of classification within Y-haplogroup C3c and refines our knowledge about the distribution of both C3c* and C3c1. I wish more researchers would publish such short technical papers that refine the classification of their Y-chromosome samples as more phylogenetic information becomes available.

From the paper:
In our study, the highest frequencies of subhaplogroup C3c1-(M77, M86) were observed in Tungusic-speaking people of North-Eastern Asia, such as Evens and Evenks, as well as in Turkic-speaking Altaian Kazakhs and Mongolic-speaking Kalmyks. These results are in agreement with previous observations based on separate or joint genotyping of M77 and M86 markers.3,9,12,16

C3c* haplotypes were detected in aboriginal populations of North- Eastern Asia—Koryaks (28.2%) and Evens (1.6%) from the Sea of Okhotsk coast (Magadan region) and West Evenks (2.4%) from Central Siberia (Evenki Autonomous District) (Table 1). Earlier, two Evenk individuals from southern part of Yakutia, one Yakut-speaking Evenk and one Yukaghir were found to belong to C3c*.2,3 Therefore, the geographic distribution of subhaplogroup C3c* is limited to the eastern part of Siberia.
The authors apply the evolutionary mutation rate -although they acknowledge that molecular dating is controversial- to obtain ages of 9.9 (C3c), 6.5 (C3c1), and 4.5 (C3c*). While I don't trust the ability of Y-STR-based molecular dating to provide reasonably accurate age estimates, I would not be surprised if C3c1 was somehow implicated in the deeper origins of the Altaic language family, at least in the "narrow-sense" (Mongolian-Tungusic-Turkic).

J Hum Genet. 2012 Jul 19. doi: 10.1038/jhg.2012.93. [Epub ahead of print]

On the Y-chromosome haplogroup C3c classification.

Malyarchuk BA, Derenko M, Denisova G.

Abstract As there are ambiguities in classification of the Y-chromosome haplogroup C3c, relatively frequent in populations of Northern Asia, we analyzed all three haplogroup-defining markers M48, M77 and M86 in C3-M217-individuals from Siberia, Eastern Asia and Eastern Europe. We have found that haplogroup C3c is characterized by the derived state at M48, whereas mutations at both M77 and M86 define subhaplogroup C3c1. The branch defined by M48 alone would belong to subhaplogroup C3c*, characteristic for some populations of Central and Eastern Siberia, such as Koryaks, Evens, Evenks and Yukaghirs. Subhaplogroup C3c* individuals could be considered as remnants of the Neolithic population of Siberia, based on the age of C3c*-short tandem repeat variation amounting to 4.5±2.4 thousand years.

Link

March 16, 2012

TreeMix analysis of North Eurasians (and an African surprise)

I have used my K12b dataset to isolate a set of 537 individuals who had less than 10% membership in the South Asian, Northwest African, Southeast Asian, South Asian, East African, Gedrosia, South Asian, East African, Southwest_Asian, and Sub_Saharan components. Hence, the remaining 537 individuals had 90%+ membership in the remaining Atlantic_Med, North_European, Caucasus, Siberian, and East_Asian components.
  • The Atlantic_Med component is frequent in northwestern Europe
  • The North_European component is dominant in northeastern Europe and forays into Siberia
  • The Caucasus component is dominant in the Caucasus and forays into Central Asia
  • The Siberian component is dominant in North Asia and forays into Europe
  • The East_Asian component is frequent in East Asia and forays into North Asia
This pruning procedure may not be perfect, but it helps isolate a dataset consisting (mostly) of North Eurasian individuals. Furthermore, I removed all populations who had less than 5 remaining individuals after the first pruning step. Hence, in the end, I had a dataset of 38 populations/452 individuals. The remaining populations were:
Russian_D, Polish_D, German_D, Finnish_D, Swedish_D, Mixed_Slav_D, Norwegian_D, Lithuanian_D, Japanese_D, Daur, French, French_Basque, Hezhen, Japanese, Oroqen, Russian, Sardinian, Yakut, CEU30, JPT30, Belorussian, Chuvashs, Hungarians, Lithuanians, Romanians, Selkup, Evenk, Tuva, Yukagir, Nganassan, Dolgan, Buryat, Mongol, FIN30, Kent_1KG, Bulgarians_Y, Ukranians_Y, Mordovians_Y
Additionally, a sample of 30 Yoruba from the HapMap-3 was used as an outgroup.

TreeMix analysis

The TreeMix analysis was performed with default parameters, and allowing for a different number of migration edges.

Nomenclature: The direction of gene flow is best seen in the figure and/or associated treeout files.
For the text, I will put in (), the common ancestor of two populations, e.g., (French_Basque,Sardinian) and also as (X, *) the tree rooted at a particular node X, e.g., (Buryat, *)

0 migration edges:

The West and East Eurasian clusters are identified, with some populations with likely admixture being placed closer to the Eurasian root.

1 migration edge:
64% from (Sardinians/Basques) to Yoruba; this is difficult to interpret, but there has been evidence in the past that Africans and West Eurasians share more ancestry than Africans and East Asians do. In the linked post, I proposed a major episode of back-migration into Africa, and it is perhaps this that is being captured by this migration edge: Sardinians/Basques are the only two South-West Eurasian populations included, and any back-migration into Africa must have originated in the southern parts of West Eurasia.

Such a high level of back-migration may in fact be plausible, since Yoruba are a predominantly Y-haplogroup E bearing population, and the origin of the DE clade of the human Y-chromosome phylogeny is up in the air with both an African and Eurasian case having been advanced. Personally, I favor the Eurasian case, since within the CT clade, we have two subclades: CF (Eurasian) and DE (Eurasian/African).

Interestingly, John Hawks has recently discovered an unanticipated excess of "Neandertal ancestry" in Yoruba. This may also point to a back-migration into Africa and/or admixture of a group of Africans related to Eurasians (whom I've called Afrasians), with groups of Africans (Palaeoafricans) that split before the H. sapiens/H. neandertalensis common ancestor.

There is, however, another detail in the figure that may have escaped your notice: there is now about 0.5 worth of drift in the figure (left-to-right) as opposed to only 0.12 in the tree without migration edges. So, perhaps what we are seeing is indeed the first sign of admixture between modern and archaic humans in Africa, which has been made more likely by recent anthropological discoveries.

It's not clear to me whether TreeMix has stumbled onto something important or not, but it is certainly worth keeping in mind that the above model fits the data better than the simple tree model. Moreover, TreeMix attempts to reverse the polarity of migration edges, and -apparently- the (Sardinian, French_Basque)-to-Yoruba edge is preferable to the reverse.

So, we should keep our minds open to the possibility that the greater similarity of West Eurasians to Africans is not the result of multiple Out-of-Africa waves, one of which affected only West Eurasians, but of an Into-Africa back-migration from West Eurasia.

So far, tree-based models have focused on how diverse African groups are, and hence, the reduced diversity of Eurasians has been interpreted as an Out-of-Africa bottleneck that carried a subset of African variation into Eurasia.

But, there is an alternative interpretation of the evidence, namely that African groups are diverse because they carry a superset of ancient Into-Africa variation, with the African-specific part of their variation being the result of admixture with pre-existing African hominins. Such a scenario cannot be captured by tree models, but is apparently considered and not rejected by TreeMix which allows for lateral gene flow. Let's wait and see what new things come from full genome sequencing.

2 migration edges:

The (French_Basque/Sardinian)-to-Yoruba edge persists (64%) and a new edge was added from  (Buryat, *)-to-Mongol (85%). The "Mongol" sample consists of Siberian Mongols described by Rasmussen et al. (2010). An inspection of their K12b population portrait indicates that they do, in fact, have West Eurasian admixture, which according to the K12b spreadsheet amounts to about 18% in total. 

3 migration edges:
The aforementioned (French_Basque/Sardinian)-to-Yoruba (64%) and (Buryat,*)-toMongol (85%) edges persist, and now we have a 68% Nganasan-to-Selkup edge. 

These are the two Siberian Uralic populations in the dataset. This seems to parallel the K12b results, as Selkups have a North_European element which the Nganasans (Uralic speakers from the Arctic coast of Central Siberia lack), so we are seeing the hybridity of the Selkups here, who, like the Mongol sample are partly of West Eurasian ancestry.

4 migration edges:

The aforementioned (French_Basque,Sardinian)-to-Yoruba (64%), (Buryat,*)-to-Mongol (84%), and Nganasan-to-Selkup (68%) persist, and now we have a 89% (Buryat, *)-to-Tuva edge. According to the K12b the Tuva have 13.3% West Eurasian admixture, so again we have reasonably good agreement between TreeMix and ADMIXTURE. 

Interestingly, the non-"eastern" component of Selkups and Tuvans now forms a clade. It seems that a Nganasan-like and a (Buryat, *)-like population have converged into southern Siberia, absorbed a common local element and became the Selkup and Tuva respectively.

5 migration edges:

The aforementioned (French_Basque,Sardinian)-to-Yoruba (64%), (Buryat,*)-to-Mongol (85%),  Nganasan-to-Selkup (68%) persist, and 90% (Buryat, *)-to-Tuva persist, and now we have a new 18% Oroqen-to-(Yakut, Evenk) edge. The Oroqen and the Evenk are Tungusic speakers, whereas the Yakut are Turkic people from northeastern Siberia, having migrated there from the vicinity of Lake Baikal during the last millennium.

migration edges:

The aforementioned (French_Basque,Sardinian)-to-Yoruba (64%), (Buryat,*)-to-Mongol (85%),  Nganasan-to-Selkup (68%), 90(Buryat, *)-to-Tuva persist, 18% Oroqen-to-(Yakut, Evenk), persist, and a new 16% Nganasan-to-Oroqen edge appears. Interestingly, this has allowed the Oroqen and Hezhen to now form their own clade, which makes sense as these are both Tungusic speakers from northeastern China. The other Tungusic population, the Evenk group with the Turkic Yakut: what they share in common is that they both share origins close to Lake Baikal in Siberia.

migration edges:

The aforementioned (French_Basque,Sardinian)-to-Yoruba (64%), (Buryat,*)-to-Mongol (85%),  Nganasan-to-Selkup (68%), 90(Buryat, *)-to-Tuva persist, 18% Oroqen-to-(Yakut, Evenk),  16% Nganasan-to-Oroqen edges persist, and there is a new 81% Evenk-to-Yukagir edge. The remainder of the Yukagirs' ancestry is derived from the West Eurasian tree. The Yukagir language is rather mysterious, with some links to Uralic having been postulated. Here it pays off to look at the population portraits, since it is apparent that -unlike the Selkup- their West Eurasian ancestry is limited to a few individuals.

It is fairly interesting that Russian anthropologists placed the Yukagirs in the Baikal group of the Central Asian race, the same as the Evenks, who are their biggest donors. So, Yakuts, Evenks, and Yukagirs all seem to share the same Baikal-type of origin.


migration edges:
There is now a 64% Sardinian-to-Yoruba edge, a 16% Oroqen-to-Yukagir edge, 20% (Buryat, *)-to-(Yakut, Evenk), and a 24% Nganasan-to-Chuvash edge, 29% Oroqen-to-(Yakut,Evenk) edge, 88% (Buryat, *)-to-Tuva, 62% Nganasan-to-Selkup, 85% (Buryat, *)-to-Mongol. 

The tree has been rather re-organized, with two main Siberian groups identified: an eastern group (Hezhen, Daur, Oroqen, Buryat), and a central group (Yukagir, Dolgan, Nganasan, Yakut, Evenk, Selkup). The Chuvash, predominantly Europeoid Turkic speakers from Russia show evidence of gene flow from the central group as well, whereas the Selkup, Uralic speakers from Siberia, who belong to the central group, show evidence of gene flow from Europe.

migration edges:

64% (French_Basque,Sardinian)-to-Yoruba, 85% (Buryat, *)-to-Mongol, 68% Nganasan-to-Selkup, 92% (Buryat,*)-to-Tuva, 14% Oroqen-to-(Yakut,Evenk), 14% Nganasan-to-Oroqen, 82% Yakut-to-Yukagir, 90% Evenk-to-Dolgan, 13% Hezhen-to-(Nganasan, *).

10 migration edges:

64% (French_Basque, Sardinian)-to-Yoruba, (85% Nganasan, *)-to-Mongol, 68% Nganasan-to-Selkup, 92% (Nganasan,*)-to-Tuva, 15% Oroqen-to-(Yakut,Evenk), 15% Nganasan-to-Oroqen, 82% Yakut-to-Yukagir, 90% Evenk-to-Dolgan, 43% Hezhen-to-Buryat, 14% Sardinian-to-Bulgarian.

I will stop at this point. I may add more migration edges later to this post, but I'm tired of typing this stuff.

You can download all the plots and *.treeout files here.


UPDATE (March 20): I have repeated the experiment with HGDP San, rather than Yoruba as the outrgroup:

There is now a 63% migration edge from (Basque, Sardinian) to San.

May 24, 2011

The reality of the Altaic language family

Personally I'm not surprised by this; my own look at genomic data has identified an "Altaic" component which peaks at the Turkic Yakut and Tungusic Evenk, and is shared by every Turkic, Mongolic, and Tungusic population available to me. The same component also occurs to some extent among all the Japanese (5) and Korean (4) members of the Dodecad Project, while it is lacking in all the Chinese ones (8).


Of particular interest is the degree of CCM between Indo-European and Semitic languages (Tables 2 and 3). In many of the most geographically distant languages these are less than 10; by comparison, among Semitic languages the are all greater than 20. This seems to be quite in agreement with the idea that Semitic is a Bronze Age language family, Indo-European a Neolithic one.

This impression is strengthened by the fact that CCM between reconstructed proto-languages (e.g. Proto-Iranian and Proto-Slavic = 20) are much higher. Since these proto-languages are a few thousand years closer to the root of PIE than present-day languages, and differences between them are similar to those of Semitic languages, the notion that PIE is a few thousand years older than Proto-Semitic seems quite consistent with the evidence.

Journal of Language Relationship • Вопросы языкового родства • 3 (2010) • Pp. 117–126 • © Turchin P., Peiros I., Gell-Mann M., 2010

Analyzing genetic connections between languages by matching consonant classes

Peter Turchin (University of Connecticut)
Ilia Peiros (Santa Fe Institute)
Murray Gell-Mann (Santa Fe Institute)

The idea that the Turkic, Mongolian, Tungusic, Korean, and Japanese languages are genetically related (the “Altaic hypothesis”) remains controversial within the linguistic community. In an effort to resolve such controversies, we propose a simple approach to analyzing genetic connections between languages. The Consonant Class Matching (CCM) method uses strict phonological identification and permits no changes in meanings. This allows us to estimate the probability that the observed similarities between a pair (or more) of languages occurred by chance alone. The CCM procedure yields reliable statistical inferences about historical connections between languages: it classifies languages correctly for well-known families (Indo-European and Semitic) and does not appear to yield false positives. The quantitative patterns of similarity that we document for languages within the Altaic family are similar to those in the non-controversial Indo-European family. Thus, if the Indo-European family is accepted as real, the same conclusion should also apply to the Altaic family.

Link (pdf)

May 30, 2010

ESHG 2010 abstracts

Some excerpts from this year's European Society of Human Genetics meeting:

P10.31 - Pigmentation gene MC1R shows strong genetic patterning in Eurasia
We present a comprehensive analysis of allele/haplotype frequencies from five functional SNPs (rs1805005, rs2228479, rs1805007, rs1805008, and rs885479) in MC1R throughout Eurasia, including from 12,151 individuals from 141 regional populations, focussing on novel genotype data from 38 Central Asian populations.
P10.39 - Genetic variation in Bulgarians: a mitochondrial DNA perspective
The structure and diversity of the Bulgarian mitochondrial DNA (mtDNA) gene pool is still almost unknown. In the present study, we have evaluated the extent and nature of mtDNA variation in the largest Bulgarian sample to date, comprising 855 healthy unrelated subjects from across the country.
P10.41 - Mitochondrial genome diversity in Ulchi, the tungusic-speaking tribe of the Russian Far East
The present report is based on the study of mtDNA variation in Ulchi (n=74), a Tungusic-speaking tribe of hunters and fishermen dispersed along the lakes and reaches of the Lower Amur. MtDNA analysis revealed 39 distinct mtDNA haplotypes belonging to 21 Eurasian haplogroups C2-C3, D3-D8, D11, G1-G2, M7-M9, Z, B, F, N9, Y and U4, with overall N macrohaplogroup derivatives frequency 53%, M - 43%, and R - 4%.
P10.62 - Genetic structure of Western Caucasus populations on the base of uniparental polymorphisms
We have analyzed 52 markers in coding region of the mtDNA and 48 markers in the non-recombining part of the Y-chromosome in 592 individuals representing five populations from western Caucasus (Abkhazians, Adyghes, Abazins, Georgians, and Circassians). Y-chromosome haplogroups G-M201 and J2 (J-M172) account for more than 50% of all haplogroup diversity in the studied populations. Haplogroup G-M201 in the Western Caucasus populations is represented only by subclade G2a (G-P15) with the insignificantly low exception in the Adyghe population where G1a (G-P20) amounts to less than 1%. In contrast to high frequency of J2 haplogroup J1 exhibit moderate occurrence and vary from 2 to 6 %. Haplogroup R1a (R-SRY10831.2) is also present in all studied populations.