Showing posts with label J2. Show all posts
Showing posts with label J2. Show all posts

November 16, 2015

West_Asian in the flesh (hunter-gatherers from Georgia) (Jones et al. 2015)

Years ago, I detected the presence of a West_Asian genetic component (with dual modes in "Caucasus" and "Gedrosia") whose origins I placed in the "highlands of West Asia" and which I proposed spread into Europe post-5kya with Indo-European languages.

Earlier this year, the study by Haak et al. showed that steppe invaders after 5kya brought into Europe a 50/50 mix of "Eastern European Hunter-Gatherer" (EHG) ancestry/An unknown population from the Near East/Caucasus. The "unknown population" was most similar to Caucasians/Near Easterners like Armenians but did not correspond to any ancient sample.

A new paper in Nature Communications by Jones et al. finds this "missing link" in the flesh in Upper Paleolithic/Mesolithic hunter-gatherers from Georgia which they call "Caucasus Hunter-Gatherers" (CHG). From the paper:
The separation between CHG and both EF and WHG ended during the Early Bronze Age when a major ancestral component linked to CHG was carried west by migrating herders from the Eurasian Steppe. The foundation group for this seismic change was the Yamnaya, who we estimate to owe half of their ancestry to CHG-linked sources.
The authors also make the connection to South Asia:
In modern populations, the impact of CHG also stretches beyond Europe to the east. Central and South Asian populations received genetic influx from CHG (or a population close to them), as shown by a prominent CHG component in ADMIXTURE (Supplementary Fig. 5; Supplementary Note 9) and admixture f3-statistics, which show many samples as a mix of CHG and another South Asian population (Fig. 4b; Supplementary Table 9).
Also of interest:
Both Georgian hunter-gatherer samples were assigned to haplogroup J with Kotias belonging to the subhaplogroup J2a (see methods).
The paper is open access, so go ahead and read it for other details.

Nature Communications 6, Article number: 8912 doi:10.1038/ncomms9912

Upper Palaeolithic genomes reveal deep roots of modern Eurasians

Eppie R. Jones et al.

We extend the scope of European palaeogenomics by sequencing the genomes of Late Upper Palaeolithic (13,300 years old, 1.4-fold coverage) and Mesolithic (9,700 years old, 15.4-fold) males from western Georgia in the Caucasus and a Late Upper Palaeolithic (13,700 years old, 9.5-fold) male from Switzerland. While we detect Late Palaeolithic–Mesolithic genomic continuity in both regions, we find that Caucasus hunter-gatherers (CHG) belong to a distinct ancient clade that split from western hunter-gatherers ~45 kya, shortly after the expansion of anatomically modern humans into Europe and from the ancestors of Neolithic farmers ~25 kya, around the Last Glacial Maximum. CHG genomes significantly contributed to the Yamnaya steppe herders who migrated into Europe ~3,000 BC, supporting a formative Caucasus influence on this important Early Bronze age culture. CHG left their imprint on modern populations from the Caucasus and also central and south Asia possibly marking the arrival of Indo-Aryan languages.

Link

November 25, 2014

Paternal lineages and languages in the Caucasus

An interesting new study on Y chromosome and languages in the Caucasus. The distribution of haplogroups is on the left. The authors make some associations of haplogroups with language families:

  • R1b: Indo-European
  • R1a: Scytho-Sarmatian
  • J2: Hurro-Urartian
  • G2: Kartvelian

Hum Biol. 2014 May;86(2):113-30.

Human paternal lineages, languages, and environment in the caucasus.

Tarkhnishvili D1, Gavashelishvili A1, Murtskhvaladze M1, Gabelaia M1, Tevzadze G2.

Abstract

Publications that describe the composition of the human Y-DNA haplogroup in diffferent ethnic or linguistic groups and geographic regions provide no explicit explanation of the distribution of human paternal lineages in relation to specific ecological conditions. Our research attempts to address this topic for the Caucasus, a geographic region that encompasses a relatively small area but harbors high linguistic, ethnic, and Y-DNA haplogroup diversity. We genotyped 224 men that identified themselves as ethnic Georgian for 23 Y-chromosome short tandem-repeat markers and assigned them to their geographic places of origin. The genotyped data were supplemented with published data on haplogroup composition and location of other ethnic groups of the Caucasus. We used multivariate statistical methods to see if linguistics, climate, and landscape accounted for geographical diffferences in frequencies of the Y-DNA haplogroups G2, R1a, R1b, J1, and J2. The analysis showed significant associations of (1) G2 with wellforested mountains, (2) J2 with warm areas or poorly forested mountains, and (3) J1 with poorly forested mountains. R1b showed no association with environment. Haplogroups J1 and R1a were significantly associated with Daghestanian and Kipchak speakers, respectively, but the other haplogroups showed no such simple associations with languages. Climate and landscape in the context of competition over productive areas among diffferent paternal lineages, arriving in the Caucasus in diffferent times, have played an important role in shaping the present-day spatial distribution of patrilineages in the Caucasus. This spatial pattern had formed before linguistic subdivisions were finally shaped, probably in the Neolithic to Bronze Age. Later historical turmoil had little influence on the patrilineage composition and spatial distribution. Based on our results, the scenario of postglacial expansions of humans and their languages to the Caucasus from the Middle East, western Eurasia, and the East European Plain is plausible.

Link (pdf)

October 26, 2013

Afghan mega-paper (Di Cristofaro et al.)

The admixture results nicely presented on a map:


The authors note that none of the ancestral components peaks in Central Asia, concluding that this region has been a destination rather than a source of population movements. I certainly agree that Central Asia has a lot of recent history affecting it from virtually all directions. On the other hand, we should be cautious about interpreting geographical clines in terms of directionality of population movement; a good example is Sardinia which often emerges as a "focus" of Mediterranean ancestry, but this does not mean that it is the origin of such ancestry. It would certainly be interesting to remove the layers of more recent ancestry from Central Asia to see what was there before the last few thousand years.

The PCA based on autosomal data:


The Y-chromosome haplogroup data can be found in Figure S7. The authors comment:
94% of the chromosomes are distributed within the following 9 main haplogroups: R-M207 (34%), J-M304 (16%), C-M130 (15%), L-M20 (6%), G-M201 (6%), Q-M242 (6%), N-M231 (4%), O-M175 (4%) and E-M96 (3%). Within the core haplogroups observed in the Afghan populations, there are sub-haplogroups that provide more refined insights into the underlying structure of the Y-chromosome gene pool. One of the important sub-haplogroups includes the C3b2b1-M401 lineage that is amplified in Hazara, Kyrgyz and Mongol populations. Haplogroup G2c-M377 reaches 14.7% in Pashtun, consistent with previous results [31], whereas it is virtually absent from all other populations. J2a1-Page55 is found in 23% of Iranians, 13% of the Hazara from the Hindu Kush, 11% of the Tajik and Uzbek from the Hindu Kush, 10% of Pakistanis, 4% of the Turkmen from the Hindu Kush, 3% of the Pashtun and 2% of the Kyrgyz and Mongol populations. Concerning haplogroup L, L1c-M357 is significantly higher in Burusho and Kalash (15% and 25%) than in other populations. L1a-M76 is most frequent in Balochi (20%), and is found at lower levels in Kyrgyz, Pashtun, Tajik, Uzbek and Turkmen populations. Q1a2-M25 lineage is characteristic of Turkmen (31%), significantly higher than all other populations. Haplogroup R1a1a-M198/M17 is characterized by its absence or very low frequency in Iranian, Mongol and Hazara populations and its high frequency in Pashtun and Kyrgyz populations.


PLoS ONE 8(10): e76748. doi:10.1371/journal.pone.0076748

Afghan Hindu Kush: Where Eurasian Sub-Continent Gene Flows Converge

Julie Di Cristofaro et al.

Despite being located at the crossroads of Asia, genetics of the Afghanistan populations have been largely overlooked. It is currently inhabited by five major ethnic populations: Pashtun, Tajik, Hazara, Uzbek and Turkmen. Here we present autosomal from a subset of our samples, mitochondrial and Y- chromosome data from over 500 Afghan samples among these 5 ethnic groups. This Afghan data was supplemented with the same Y-chromosome analyses of samples from Iran, Kyrgyzstan, Mongolia and updated Pakistani samples (HGDP-CEPH). The data presented here was integrated into existing knowledge of pan-Eurasian genetic diversity. The pattern of genetic variation, revealed by structure-like and Principal Component analyses and Analysis of Molecular Variance indicates that the people of Afghanistan are made up of a mosaic of components representing various geographic regions of Eurasian ancestry. The absence of a major Central Asian-specific component indicates that the Hindu Kush, like the gene pool of Central Asian populations in general, is a confluence of gene flows rather than a source of distinctly autochthonous populations that have arisen in situ: a conclusion that is reinforced by the phylogeography of both haploid loci.

Link

March 22, 2013

Y chromosomes and mtDNA from the Maldives

Of interest from the paper:

The haplogroup J(M304) Y chromosomes are all in subgroup J2(M172). 
... 
However, Eaaswarkhanth et al. (2010) report that Muslims and non-Muslims in India largely have the same Y-haplogroup frequency distribution, except that in Muslims low frequencies of Y-E1b1b1a(M78), Y-J(M304)(xJ2(M172)), and Y-G(M201) are found that are absent in non-Muslims (Eaaswarkhanth et al., 2010). In our Maldivian sample, none of those Y-haplogroups were found.

AJPA DOI: 10.1002/ajpa.22256

Indian ocean crossroads: Human genetic origin and population structure in the maldives

Jeroen Pijpe et al.

The Maldives are an 850 km-long string of atolls located centrally in the northern Indian Ocean basin. Because of this geographic situation, the present-day Maldivian population has potential for uncovering genetic signatures of historic migration events in the region. We therefore studied autosomal DNA-, mitochondrial DNA-, and Y-chromosomal DNA markers in a representative sample of 141 unrelated Maldivians, with 119 from six major settlements. We found a total of 63 different mtDNA haplotypes that could be allocated to 29 mtDNA haplogroups, mostly within the M, R, and U clades. We found 66 different Y-STR haplotypes in 10 Y-chromosome haplogroups, predominantly H1, J2, L, R1a1a, and R2. Parental admixture analysis for mtDNA- and Y-haplogroup data indicates a strong genetic link between the Maldive Islands and mainland South Asia, and excludes significant gene flow from Southeast Asia. Paternal admixture from West Asia is detected, but cannot be distinguished from admixture from South Asia. Maternal admixture from West Asia is excluded. Within the Maldives, we find a subtle genetic substructure in all marker systems that is not directly related to geographic distance or linguistic dialect. We found reduced Y-STR diversity and reduced male-mediated gene flow between atolls, suggesting independent male founder effects for each atoll. Detected reduced female-mediated gene flow between atolls confirms a Maldives-specific history of matrilocality. In conclusion, our new genetic data agree with the commonly reported Maldivian ancestry in South Asia, but furthermore suggest multiple, independent immigration events and asymmetrical migration of females and males across the archipelago. Am J Phys Anthropol 000:000–000, 2013. © 2013 Wiley Periodicals, Inc.

Link

March 07, 2013

Y chromosomes of Bulgarians (Karachanak et al. 2013)

Bulgaria had been something of a blank area in studies of uniparental markers, so it's nice to finally see a comprehensive Y-chromosome study of the country.

The dates in the paper are based on the "evolutionary mutation rate". I suspect that ancient DNA will be the final arbiter in this issue, because, for example, a Mesolithic TMRCA of E-V13 in Bulgaria implies that we'll find a lot of it in Neolithic contexts, whereas a Bronze Age one implies that we'll find a little if any of it, and a discontinuity across time.

Of interest is the occurrence of some E*(xM35, M2) in this sample in Burgas, Varna, and Plovdiv. It would be interesting to trace the ancestry of the bearers of these Y-chromosomes. I know that there still exists a minority-within-a-minority of Black Muslims in Greek Thrace, and it's not inconceivable that these Y-chromosomes may represent the legacy of a similar population; in any case, their haplotypes can be found in Table S5 for anyone wanting to investigate.

SNP Diversity within R seems substantial, and as always, it is difficult to say much, since this may be a consequence of either (i) a plausible role of the Balkans as a staging point of the likely invasion of Europe in late prehistory, or (ii) back-migration of derived R-bearers into the Balkans, be them Slavs or Goths or "eastern" folks of various stripes during history. Once again, I suspect that ancient DNA might solve this riddle, or, alternatively, routine high-coverage sequencing of the Y chromosome that might inform us, e.g., about the TMRCA of a Bulgarian and a German R-U152 or a Bulgarian and Polish R-M458.

PLoS ONE 8(3): e56779. doi:10.1371/journal.pone.0056779

Y-Chromosome Diversity in Modern Bulgarians: New Clues about Their Ancestry

Sena Karachanak et al

To better define the structure and origin of the Bulgarian paternal gene pool, we have examined the Y-chromosome variation in 808 Bulgarian males. The analysis was performed by high-resolution genotyping of biallelic markers and by analyzing the STR variation within the most informative haplogroups. We found that the Y-chromosome gene pool in modern Bulgarians is primarily represented by Western Eurasian haplogroups with ~ 40% belonging to haplogroups E-V13 and I-M423, and 20% to R-M17. Haplogroups common in the Middle East (J and G) and in South Western Asia (R-L23*) occur at frequencies of 19% and 5%, respectively. Haplogroups C, N and Q, distinctive for Altaic and Central Asian Turkic-speaking populations, occur at the negligible frequency of only 1.5%. Principal Component analyses group Bulgarians with European populations, apart from Central Asian Turkic-speaking groups and South Western Asia populations. Within the country, the genetic variation is structured in Western, Central and Eastern Bulgaria indicating that the Balkan Mountains have been permeable to human movements. The lineage analysis provided the following interesting results: (i) R-L23* is present in Eastern Bulgaria since the post glacial period; (ii) haplogroup E-V13 has a Mesolithic age in Bulgaria from where it expanded after the arrival of farming; (iii) haplogroup J-M241 probably reflects the Neolithic westward expansion of farmers from the earliest sites along the Black Sea. On the whole, in light of the most recent historical studies, which indicate a substantial proto-Bulgarian input to the contemporary Bulgarian people, our data suggest that a common paternal ancestry between the proto-Bulgarians and the Altaic and Central Asian Turkic-speaking populations either did not exist or was negligible.

Link

January 31, 2013

Y chromosome and mtDNA study of modern Middle Eastern populations (Badro et al. 2013)

I will just briefly comment on the occurrence of L3* mtDNA in the Near East. This is a critical haplogroup because of its age of ~70ky. If all L3* in the Near East represents African migrants, then only the M and N macrogroups appeared in Eurasia, and a good case can be made for a "late" OoA event.

On the other hand, it is quite possible that some of the L3* in the Near East does not represent recent admixture, but rather native forms of L3 with deep ancestry in the region. If that is the case, then the Near East will emerge as the origin of L3, with M, N representing Out-of-Near East-into-Eurasia founders, and the various L3*(xM, N) representing Out-of-Near East-into-Africa founders.

It is difficult to say at present what will turn out to be the case. Ancient DNA has the potential of resolving this issue, because if L3*(xM, N) in Eurasia is really recent (e.g., associated with Islamic/Arab dispersals spanning Africa and Eurasia), then it ought to be missing from the earliest genetic layers.

Also of interest the geographical distribution of Y-haplogroups; nothing much new here, but still useful as a reference:





PLoS ONE 8(1): e54616. doi:10.1371/journal.pone.0054616

Y-Chromosome and mtDNA Genetics Reveal Significant Contrasts in Affinities of Modern Middle Eastern Populations with European and African Populations 

Danielle A. Badro et al.

The Middle East was a funnel of human expansion out of Africa, a staging area for the Neolithic Agricultural Revolution, and the home to some of the earliest world empires. Post LGM expansions into the region and subsequent population movements created a striking genetic mosaic with distinct sex-based genetic differentiation. While prior studies have examined the mtDNA and Y-chromosome contrast in focal populations in the Middle East, none have undertaken a broad-spectrum survey including North and sub-Saharan Africa, Europe, and Middle Eastern populations. In this study 5,174 mtDNA and 4,658 Y-chromosome samples were investigated using PCA, MDS, mean-linkage clustering, AMOVA, and Fisher exact tests of FST's, RST's, and haplogroup frequencies. Geographic differentiation in affinities of Middle Eastern populations with Africa and Europe showed distinct contrasts between mtDNA and Y-chromosome data. Specifically, Lebanon's mtDNA shows a very strong association to Europe, while Yemen shows very strong affinity with Egypt and North and East Africa. Previous Y-chromosome results showed a Levantine coastal-inland contrast marked by J1 and J2, and a very strong North African component was evident throughout the Middle East. Neither of these patterns were observed in the mtDNA. While J2 has penetrated into Europe, the pattern of Y-chromosome diversity in Lebanon does not show the widespread affinities with Europe indicated by the mtDNA data. Lastly, while each population shows evidence of connections with expansions that now define the Middle East, Africa, and Europe, many of the populations in the Middle East show distinctive mtDNA and Y-haplogroup characteristics that indicate long standing settlement with relatively little impact from and movement into other populations.

December 11, 2012

Y chromosome study of Italy (Brisighelli et al. 2012) incl. sample of Greek speakers from Salento

This is a wonderful new source of information on Y-chromosome variation in Italy, that also includes some samples of the linguistic minorities of Ladins and Griko speakers.

The latter is particularly interesting to me, because, these last Greeks of Magna Graecia are descended either from the ancient colonists or medieval Eastern Roman settlers, and as such may represent a group of Greek descendants that (i) may have admixed to some extent with local Italic speakers, but (ii) will not have had an opportunity to experience much post-medieval gene flow that may have affected Greeks from the Aegean.

There may be something wrong with the presentation of the haplogroup frequencies on the left; in particular, based on the text, I think that what appears as R1* is in fact R1*(xR1a1).

In any case, here are my observations on the Grecani Salentini sample:
  • They, as well as the Messapi, possess the highest frequencies of E-M78. This ties them to the Balkans in a very obvious way; this haplogroup was also interpreted as a signal of Greek colonization in Sicily and Massalia. This seems like the most obvious explanation; note that Salento is in Messapia, so the high frequency in the non-Greek denizens of the region may be simply the result of language shift, since the remaining Greek speakers are presumably the last remnant of a once much more numerous population that was linguistically Italicized as have most other Greek speaking populations of Italy and Sicily.
  • Their highest frequency haplogroups are R1*(xR1a1) and J2. Both are fairly common haplogroups in both Greece and Italy, so only a fine-scale analysis would be able to differentiate between what might be pre-Greek and what is Greek in origin. In any case, I have proposed that these two haplogroups were typical of (albeit not limited to) the Graeco-Phrygo-Armenian clade, so their occurrence in this sample is not surprising.
  • There is an occurrence of I*(xM26) chromosomes. This requires finer phylogenetic resolution, but certainly the absence of M26 -which has a SW European distribution- is interesting to note.
  • Haplogroup G-M201 again requires finer-scale resolution, and could be anything from a relative of the Neolithic Italians (having been found in the Tyrolean Iceman) to much more recent events.
  • Within haplogroup J, the majority of the chromosomes belong to clade J2, with about a tenth of the frequency made up of J*(xM62, M172). Note that these are not necessarily J*(xJ1,J2) as indicated in the figure, since M62 defines only a part of the J1 lineage.
  • The absence of haplogroup R1a1 in this sample is perhaps the most interesting finding. This occurs at a frequency of ~10% in Greek samples from Greece and is fairly variable. I have previously observed that it was absent in the south stream of Indo-European based on its paucity in Armenians, Albanians, and its uneven distribution in Greeks. Its absence in the Italian Griko sample reinforces this idea. A caveat, however, is that the origin of the Greek settlement of Italy can be traced to southern Greece and western Anatolia, so it's still possible that some R1a1 was present in other areas of the Aegean basin since pre-medieval times.
The authors of the paper use many conventional labels of what is "Neolithic" and what is not (e.g., R1*(xR1a1) is claimed as Mesolithic). But, certainly, both age estimation of modern chromosomes (e.g., Wei et al. 2012) and the ancient Y chromosome studies cast doubt on this association. I would say that rather than being predominantly pre-Neolithic, it might appear that the Y-chromosome gene pool of Italy may have been formed in late Neolithic to medieval times, with the only lineages that can convincingly trace their ancestry to the Neolithic or earlier epochs being G and I-M26.

As for the Ladins, the high frequency (67.7%) of R1*(xR1a1) is consistent with what I believe to have been the main Italo-Celtic lineage.

Finally, I should point out the occurrence of a couple of haplogroup L samples; this haplogroup is more typical of populations much to the east, being the "eastern" cousin of the more "western" haplogroup T within the LT clade. Certainly a finer-scale resolution of these two L samples might be informative about their potential origins and/or the ancient distribution of this rather mysterious haplogroup.


PLoS ONE 7(12): e50794. doi:10.1371/journal.pone.0050794

Uniparental Markers of Contemporary Italian Population Reveals Details on Its Pre-Roman Heritage

Francesca Brisighelli et al.

Abstract
Background

According to archaeological records and historical documentation, Italy has been a melting point for populations of different geographical and ethnic matrices. Although Italy has been a favorite subject for numerous population genetic studies, genetic patterns have never been analyzed comprehensively, including uniparental and autosomal markers throughout the country.

Methods/Principal Findings

A total of 583 individuals were sampled from across the Italian Peninsula, from ten distant (if homogeneous by language) ethnic communities — and from two linguistic isolates (Ladins, Grecani Salentini). All samples were first typed for the mitochondrial DNA (mtDNA) control region and selected coding region SNPs (mtSNPs). This data was pooled for analysis with 3,778 mtDNA control-region profiles collected from the literature. Secondly, a set of Y-chromosome SNPs and STRs were also analyzed in 479 individuals together with a panel of autosomal ancestry informative markers (AIMs) from 441 samples. The resulting genetic record reveals clines of genetic frequencies laid according to the latitude slant along continental Italy – probably generated by demographical events dating back to the Neolithic. The Ladins showed distinctive, if more recent structure. The Neolithic contribution was estimated for the Y-chromosome as 14.5% and for mtDNA as 10.5%. Y-chromosome data showed larger differentiation between North, Center and South than mtDNA. AIMs detected a minor sub-Saharan component; this is however higher than for other European non-Mediterranean populations. The same signal of sub-Saharan heritage was also evident in uniparental markers.

Conclusions/Significance

Italy shows patterns of molecular variation mirroring other European countries, although some heterogeneity exists based on different analysis and molecular markers. From North to South, Italy shows clinal patterns that were most likely modulated during Neolithic times.

Link

December 05, 2012

Y chromosomes in Iranians and Tajiks (Malyarchuk et al. 2013)

An interesting paper on Iranian and Tajik Y chromosomes. Iranian Y chromosomes were comprehensively studied by Grugni et al. but it is always good to have additional samples.


I have mentioned before the apparent distinction between west and east Iranians in terms of haplogroup J/R1a frequencies, with high ratios in Persians and Kurds, and low ones in Pathans, and this seems to be reinforced here; the Tajiks are speakers of Persian (hence "western") but trace their ancestry to the east of the modern country of Iran, and in-between Persians and eastern Iranians.

The absence of R1a in this Kurdish sample, coupled with high J frequency parallels the situation in the Kurdish Anatolian settlement studied by Gokcument et al., as well as the Georgian Kurmanji sample studied by Nasidze et al. On the other hand, R1a is present in the Kurmanji samples from Turkey and Turkmenistan in the latter study, as well as in the aforementioned Kurdish sample from Iran by Grugni et al. and the Kurdish sample from Turkmenistan studied by Wells et al. I'd say that there is potential variation of this haplogroup within Kurdish groups, which might be worth further exploration.

It would also be very interesting to study the haplogroup I chromosomes from this region. Do they represent historical introgression from Europe, or are they, perhaps, local basal clades that reinforce the idea of a relic distribution of I in West Asia, prior to the migration into Europe, that was recently suggested by the discovery of IJ* chromosomes in Iran by Grugni et al.?


Annals of Human Biology, 2013; Early Online: 1–7

Y-chromosome variation in Tajiks and Iranians

Boris Malyarchuk et al.

Aim: The purpose of this study was to characterize Y-chromosome diversity in Tajiks from Tajikistan and in Persians and Kurds from Iran.

Method: Y-chromosome haplotypes were identified in 40 Tajiks, 77 Persians and 25 Kurds, using 12 short tandem repeats (STR) and 18 binary markers.

Results: High genetic diversity was observed in the populations studied. Six of 12 haplogroups were common in Persians, Kurds and Tajiks, but only three haplogroups (G-M201, J-12f2 and L-M20) were the most frequent in all populations, comprising together 60% of the Y-chromosomes in the pooled data set. Analysis of genetic distances between Y-STR haplotypes revealed that the Kurds showed a great distance to the Iranian-speaking populations of Iran, Afghanistan and Tajikistan. The presence of Indian-specific haplogroups L-M20, H1-M52 and R2a-M124 in both Tajik samples from Afghanistan and Tajikistan demonstrates an apparent genetic affinity between Tajiks from these two regions.

Conclusions: Despite the marked similarities between Y-chromosome gene pools of Iranian-speaking populations, there are differences between them, defined by many factors, including geographic and linguistic relationships.

Link

November 29, 2012

South Indian Y chromosomes (+ a little complaining about methods)

The table of haplogroup frequencies (left) may prove quite useful, but I am fairly disappointed with what appears to be the state of the art in recent published research on Y chromosome variation. This is not to belittle the tremendous amount of labor and money needed to collect and genotype large representative samples of individuals; only to express hope that better use of the collected samples could be achieved.

First of all, it is inconceivable to me how scientists can continue to use the 3x slower "evolutionary mutation rate" for their analyses of Y-chromosome ages on the basis of Y-STR markers. I have done my small part in my Y-STR series to show that this mutation rate is applicable only for a rather specific demographic history, and completely unsuitable to real growing human populations where Y-STR variance accumulates at close to the genealogical rate. And, my observations merely elaborated quantitatively what was already present in Zhivotovsky et al. (2006) but has been completely ignored since:
In simulations of a neutral process with average rate of increase m = 1, the number of surviving haplogroups rapidly decreased with time and corresponded well with the theory of mutant survival (Li 1955, p. 242), and the average size of the surviving haplogroups increased each generation by a value rapidly approaching 0.5 (data not shown), which agrees with asymptotic fraction of 2/t of haplotypes that survive at generation t (Athreya and Ney 1972, p. 19). The accumulated variance increased almost linearly (fig. 1), at a rate of increase about 0.00028 per generation; that is, the actual rate of accumulation microsatellite variation was about 3.6 times less than that predicted from the germ line mutation rate. This corresponds perfectly to the 3- to 4-fold difference observed between germ line and evolutionarily effective mutation rate.
The issue is all but resolved in the amateur "genetic genealogy" community, but even professional geneticists often use either genealogical or evolutionary rate, or take an agnostic stance by reporting results based on both rates. To arrive at strong conclusions about a topic on the basis of a mutation rate that is, to say the least, controversial, without even acknowledging the existence of a controversy is unsatisfactory. Y-chromosome researchers ought to copy the attitude of those working with autosomal DNA, where a corresponding mutation rate controversy was not swept under the carpet, but acknowledged (e.g., in the recent Meyer et al. high-coverage Denisova paper), with the implications of the uncertainty during the present "transitional" period quantified in the form of wider confidence intervals.

This "mutation rate" issue  notwithstanding, it was also recently shown that by Busby et al. that Y-STR based estimates have a dependence on the set of Y-STRs used, with markers exhibiting linear behavior across different time spans. This does not invalidate their use as molecular clocks, but highlights the need to not only select a bunch of Y-STRs, but also either (i) demonstrate that the selected set exhibits linear behavior for the time span of interest, or (ii) correct for deviations from linearity. Again, this type of modelling of microsatellite behavior was recently achieved for autosomal STRs by Sun et al.  Note that such deviations result in a slower rate than the genealogical one, but the mechanism whereby this is produced is completely different than the one proposed by Zhivotovsky et al.: it is not drift in a non-growing (m=1) population that reduces the effective rate, but rather "saturation" of the mutation process, whereby the variance at fast-mutating markers grows sub-linearly with time, because of physical constraints on their possible range of values.

I don't hope that Y-STR based age estimation will have much to offer in the coming years. But the third set of the 1000 Genomes Project is on its way, and this will include a variety of South Asian samples. Very soon we will be in a good position to study the time depth of common ancestry between e.g., European and South Asian Y-chromosomes within various haplogroups using point mutations, and these are not plagued by many of the problems associated with Y-STR variation and its interpretation.

Finally, I can't help but notice that this paper has not acknowledged the tremendous progress in resolving the Y chromosome phylogeny done by non-academic researchers. With the current state of our knowledge, the claim that haplogroup R1a1 is "autochthonous" in India is not tenable. Even if one discounts all the evidence made by SNP discoveries in the commercial testing world (and why should they?), finer-scale structure within this haplogroup has now been officially published and appears to be inconsistent with a South Asian origin of this haplogroup.

Certainly, not all is resolved; for example, the representation of tribal populations in commercial DNA testing is almost non-existent, and a sampling of their Y-SNP diversity is urgently needed. A very useful paradigm of research is that of recent work on the most basal clade of the Y-chromosome phylogeny (A00) in which the identification of very unique Y-chromosomes by genetic genealogists was combined with academic samples of "indigenous" peoples to produce new knowledge.

Much of population genetic research will benefit from such consilience between academics and amateurs. This is not an idle hope, but a recognition that this field is one in which the public not only has a substantial interest but can also do something about it. Many might be interested in Mars exploration, but without Elon Musk's bank account, most are consigned to being consumers of information about the Red Planet. Hopefully, better ways of combining the efforts of research scientists and the educated public can be identified and used in the near future.

PLoS ONE 7(11): e50269. doi:10.1371/journal.pone.0050269

Population Differentiation of Southern Indian Male Lineages Correlates with Agricultural Expansions Predating the Caste System

GaneshPrasad ArunKumar et al.

Previous studies that pooled Indian populations from a wide variety of geographical locations, have obtained contradictory conclusions about the processes of the establishment of the Varna caste system and its genetic impact on the origins and demographic histories of Indian populations. To further investigate these questions we took advantage that both Y chromosome and caste designation are paternally inherited, and genotyped 1,680 Y chromosomes representing 12 tribal and 19 non-tribal (caste) endogamous populations from the predominantly Dravidian-speaking Tamil Nadu state in the southernmost part of India. Tribes and castes were both characterized by an overwhelming proportion of putatively Indian autochthonous Y-chromosomal haplogroups (H-M69, F-M89, R1a1-M17, L1-M27, R2-M124, and C5-M356; 81% combined) with a shared genetic heritage dating back to the late Pleistocene (10–30 Kya), suggesting that more recent Holocene migrations from western Eurasia contributed less than 20% of the male lineages. We found strong evidence for genetic structure, associated primarily with the current mode of subsistence. Coalescence analysis suggested that the social stratification was established 4–6 Kya and there was little admixture during the last 3 Kya, implying a minimal genetic impact of the Varna (caste) system from the historically-documented Brahmin migrations into the area. In contrast, the overall Y-chromosomal patterns, the time depth of population diversifications and the period of differentiation were best explained by the emergence of agricultural technology in South Asia. These results highlight the utility of detailed local genetic studies within India, without prior assumptions about the importance of Varna rank status for population grouping, to obtain new insights into the relative influences of past demographic events for the population structure of the whole of modern India.

Link

October 03, 2012

rolloff analysis of South Indian Brahmins as Armenian+Chamar

The first analysis of this population showed that there were negative f3(Brahmin; X, Y) signals when X were a variety of West European, Balkan, and West Asian population, and Y either the Chamar or North Kannadi. In the first analysis I used Orcadians and North Kannadi. I have now carried out a new rolloff analysis on 470,559 SNPs, using Armenians_Y and Chamar_M as the reference populations.

The exponential fit can be seen below.
The admixture date is 142.814 +/- 15.010 generations, or 4,140 +/- 440 years, which seems to correspond quite well with commonly accepted dates for the formation of Indo-Iranian.

I have previously observed that:

These patterns can be well-explained, I believe, if we accept that Indo-Iranians are partially descended not only from the early Proto-Indo-Europeans of the Near East, but also from a second element that had conceivable "South Asian" affiliations. The most likely candidate for the "second element" is the population of the Bactria Margiana Archaeological Complex (BMAC). The rise and demise of the BMAC fits well with the relative shallowness of the Indo-Iranian language family and its 2nd millennium BC breakup, and has been assigned an Indo-Iranian identity on other grounds by its excavator. As climate change led to the decline and abandonment of BMAC sites, its population must have spread outward: to the Iranian plateau, the steppe, and into South Asia, reinforcing the linguistic differentiation that must have already began over the extensive territory of the complex.
Quite possibly, as the West Asian element began mixing with the Sardinian-like population in Greece, another branch of the Indo-Europeans made its appearance east of the Caspian, in the territory of the BMAC, admixing with South Asian-like populations. Thus, it might seem that the Graeco-Aryan clade of Indo-European broke down during the Bronze Age, with one branch heading off to the Balkans, and another to the east. 

This scenario would also explain how the likely J2-bearing population associated with the earliest Proto-Indo-Europeans may have acquired the contrasting pattern I have previously described: the western (cis-Caspian) population would have admixed with R1b-bearers who occupy the "small arc" west and south of the Caspian, while the eastern (trans-Caspian) populations would have admixed with R1a-bearers who occupy the "large arc" in the flatlands north and east of the Caspian. It would also explain how the "western" branch (Graeco-Armenian) would have picked up Sardinian-like "Atlantic_Med" admixture, which is absent in the "eastern" Indo-Iranian branch.

At the same time, this scenario would explain the lack of "North European" admixture in the "western" branch (since this was shielded by the Caucasus and Black Sea from the northern Europeoids who may have lived north of these barriers), and explain it in the "eastern" branch (since the BMAC agriculturalists were in contact with presumably northern Europeoid groups inhabiting the steppelands, unhindered by any major physical barriers). (The relative absence of this admixture in the Graeco-Armenian branch may be advanced on the strength of its absence in Armenians, the evidence of a Sardinian-like Iron Age individual from Bulgaria, and the historical-era timing of admixture for the Greek population.)

It would be interesting to carry out similar experiments on Iranian groups, to see if they, too, present a similar pattern of admixture.

July 26, 2012

A look at Y chromosomes of Romania via Count Dracula

In short: researchers tried to see whether they could identify a specific Y chromosome lineage associated with the House of Basarab in Romania, the most famous member of which is Vlad the Impaler, an inspiration for the mythical Count Dracula. To do this, they tested Basarab-surnamed individuals, as well as the general Romanian population.

The whole exercise was, in a sense, a failure, since it neither disclosed a Basarab-specific lineage, nor resolved the historical question about the origin of the House of Basarab (Vlach or Cuman). But, it gave us some wonderful new data on Romania that is, of course, quite welcome.

This seems like a good candidate for a future ancient DNA study, assuming of course, that Vlad and his family are still in their final resting place, and there are brave enough researchers to disturb them (j/k).

On a more serious note, the authors correctly state that even if the Basarab house was originally Turkic, they could still have carried West Eurasian chromosomes, since incoming Turkic groups in Europe were not purely Mongoloid like their more remote ancestors. On the other hand, I note that most of the Basarab-surnamed individuals belonged to E-V13, I-P37.2, J-M241 all of which are almost certainly native Romanian. If one of them carries the original chromosome, then the odds are in favor of a Romanian origin, although nothing short of ancient DNA work can resolve the issue, assuming that's possible.

Table S1 contains the new Romanian data, and Table S2 data from surrounding populations (Hungary, Bulgaria, Ukraine).

PLoS ONE 7(7): e41803. doi:10.1371/journal.pone.0041803

Y-Chromosome Analysis in Individuals Bearing the Basarab Name of the First Dynasty of Wallachian Kings

Begoña Martinez-Cruz et al.

Vlad III The Impaler, also known as Dracula, descended from the dynasty of Basarab, the first rulers of independent Wallachia, in present Romania. Whether this dynasty is of Cuman (an admixed Turkic people that reached Wallachia from the East in the 11th century) or of local Romanian (Vlach) origin is debated among historians. Earlier studies have demonstrated the value of investigating the Y chromosome of men bearing a historical name, in order to identify their genetic origin. We sampled 29 Romanian men carrying the surname Basarab, in addition to four Romanian populations (from counties Dolj, N = 38; Mehedinti, N = 11; Cluj, N = 50; and Brasov, N = 50), and compared the data with the surrounding populations. We typed 131 SNPs and 19 STRs in the non-recombinant part of the Y-chromosome in all the individuals. We computed a PCA to situate the Basarab individuals in the context of Romania and its neighboring populations. Different Y-chromosome haplogroups were found within the individuals bearing the Basarab name. All haplogroups are common in Romania and other Central and Eastern European populations. In a PCA, the Basarab group clusters within other Romanian populations. We found several clusters of Basarab individuals having a common ancestor within the period of the last 600 years. The diversity of haplogroups found shows that not all individuals carrying the surname Basarab can be direct biological descendants of the Basarab dynasty. The absence of Eastern Asian lineages in the Basarab men can be interpreted as a lack of evidence for a Cuman origin of the Basarab dynasty, although it cannot be positively ruled out. It can be therefore concluded that the Basarab dynasty was successful in spreading its name beyond the spread of its genes.

July 19, 2012

Huge study on Y-chromosome variation in Iran (Grugni et al. 2012)

This is the equivalent of a box of candy for anyone interested in Eurasian (pre-)history. I will have digest all the goodies within, and post any of my comments as updates to this post.

UPDATE I: Here is the table of haplogroup frequencies for easy reference:

One of the most interesting finds is the presence of a few IJ-M429* chromosomes  in the sample. Haplogroup IJ encompasses the major European I subclade, and the major West Asian J subclade. The discovery of IJ* chromosomes is consistent with the origin of this haplogroup in West Asia; it is widely believed that haplogroup I represents a pre-Neolithic lineage in Europe, although at present there are no Y chromosome-tested pre-Neolithic remains.

There is also a wide assortment of Q and R in Iran. While some of these may be intrusive (e.g., the 42.6% of Q1a2 in Turkmen, likely a legacy of their Central Asian origins), the overall picture appears consistent with a deep presence of these lineages in Iran. This is especially true for haplogroup R where pretty much every paragroup and derived group is present, excepting those likely to have originated recently elsewhere.

UPDATE II: From the paper:
Although accounting only for 25% of the total variance, the first two components (Figure 3) separate populations according to their geographic and ethnic origin and define five main clusters: East-African, North-African and Near Eastern Arab, European, Near Eastern and South Asian. The 1stPC clearly distinguishes the East African groups (showing a high frequency of haplogroup E) from all the others which distribute longitudinally along the axis with a wide overlapping between European and Arab peoples and between Near Eastern and South Asian groups. The 2ndPC separates the North-African and Near Eastern Arabs (characterized by the highest frequency of haplogroup J1) from Europeans (characterized by haplogroups I, R1a and R1b) and the Near Easterners from the South Asians (due to the distribution of haplogroups G, R2 and L). Iranian groups do not cluster all together, occupying intermediate positions among Arab, Near Eastern and Asian clusters. In this scenario, it is worth of noticing the position of three Iranian groups: (i) Khuzestan Arabs (KHU-Ar) who, despite their Arabic origin, are close to the Iranian samples; (ii) Armenians from Tehran (THE-Ar), whose position, in the upper part of the Iranian distribution, indicates a close affinity with the Near Eastern cluster, while their position near Turkey and Caucasus groups, due to the high frequency R1b-M269 and other European markers (eg: I-M170), is in agreement with their Armenia origin; (iii) Sistan Baluchestan (SB-Ba) that clusters with its neighbouring Pakistan.
UPDATE III: There are lots of little details in the haplogroup distribution that make historical sense. For example, C3 exists in Assyrians from Azarbaijan, and both C*, C3, and O exists in Zoroastrians from Yazd. It is often forgotten that before the spread of Islam, and quite time thereafter, Inner Asia was teeming with Zoroastrians and Nestorian Christians. It seems quite likely that these outliers represent a legacy of these communities.

UPDATE IV: I have a feeling that Razib will take exception with this statement: "Ancient Persian people were firstly characterized by the Zoroastrianism. After the Islamization, Shi'a became the main doctrine of all Iranian people."


UPDATE V: This confirms my observation from the recent studies in Afghanistan, that there is an inverse relationship of J2a and R1a in Iranian-speaking groups, with an excess of the latter among the eastern Iranians, and of the former among the Persians. From the paper:
Among the different J2a haplogroups, J2a-M530 [46] is the most informative as for ancient dispersal events from the Iranian region. This lineage probably originated in Iran where it displays its highest frequency and variance in Yazd and Mazandaran (Figure 2). Taking into account its microsatellite variation and age estimates along its distribution area (Tables S3 and S7), it is likely that its diffusion could have been triggered by the Euroasiatic climatic amelioration after the Last Glacial Maximum and later increased by agriculture spread from Turkey and Caucasus towards southern Europe. The high variance observed in the Italian Peninsula is probably the result of stratifications of subsequent migrations and/or of the presence of sub-lineages not yet identified. Of interest in the M530 network (Figures 2 and S3) is the presence of a lateral branch that is characterized by a DYS391 repeat number equal to 9. Differently from previous observations [46], this branch is not restricted to Anatolian Greek samples being shared with different eastern Mediterranean coastal populations. The M530 diffusion pattern seems to be also shared by the paragroups J2a-M410* and J2a-PAGE55*. In addition, the variance distribution of the rare R1b-M269* Y chromosomes, displaying decreasing values from Iran, Anatolia and the western Black Sea coastal region, is also suggestive of a westward diffusion from the Iranian plateau, although more complex scenarios can be still envisioned because of its non-star like structure.
Of course, the idea that the diffusion of J2a related lineages ties in with early agricultural expansions has been with us for a long time, but it is time to abandon it. First of all, as we have seen, J2a diminishes greatly as we head towards South Asia; it certainly doesn't look like the lineage of the multitude of agricultural settlements that sprang up along the southeastern vector soon after the invention of agriculture. Second, it is lacking so far in all ancient Y chromosome data from Europe down to 5,000 years ago. It seems much more probably that J2 related lineages spread from the highlands of West Asia much later. 


The "age estimates" are the result of using the inappropriate "evolutionary mutation rate", and become even older because of the inclusion of the DYS388 marker that is very stable in many haplogroups but very mutable within haplogroup J. On the left you can see frequency, Y-STR variance, and haplotype network structures for various J-related groups.


It is unfortunate that there is no progress in the phylogeographic assessment of R1a in this paper. There have been substantial discoveries of SNPs within this haplogroup as a result of commercial testing; however there is clearly an ascertainment bias in the newer discoveries, as almost all these SNPs have been detected in Europeans. The new paper confirms the high levels of Y-STR variance in India, Pakistan, and Iran. Together with the cornucopia of related paragroups in Iran, there is little doubt that this haplogroup originated in the general area of Central/South Asia.


Personally, as I have stated before, I would relate this R1a with Neolithic peoples living east of the Caspian, in contrast to the R1b bearers who lived west and south of it. These two populations came under the influence of the Indo-Europeans and spread in different directions. The Indo-Iranians were then initially the mixed descendants of the Indo-Europeans and the R1a old agricultural population, and were formed in the territory of the Bactria-Margiana Archaeological Complex. 


This also explains the contrast between Iranian and Armenian groups: the latter mostly lack the R1a lineage, contrasting with all Iranian groups (even their Kurdish neighbors) who possess it. Conversely, Iranian groups, and especially eastern Iranians and Indo-Ayrans lack the R1b lineage. This is due to the fact that neither R1a nor R1b were originally part of the Indo-European community, but their geographical position was such that they came under the influence of the Indo-Europeans when the latter began their expansion.


UPDATE VI: I have created my own dendrogram using the Y-haplogroup frequencies and the hclust package of R (default parameters):


From top to bottom, one can identify some clusters:

  • Eastern Europe, further broken down into Balkans and Slavic+Hungary
  • West Asian/Caucasus
  • Iranian Proper
  • Arab

These correspond largely to the clusters identified by the authors, with India and the Turkmen sample emerging as the clear outliers. I omitted the Ethiopian samples, since E-M78 was not resolved phylogenetically, causing the Ethiopians to group with the likely E-V13 from the Balkans.

UPDATE VII: I have also run MCLUST over the haplogroup frequency data over the MDS representation of the distance matrix. The maximum number of 10 clusters occurred with 5 MDS dimensions retained. Population assignments in the 10 clusters can be found in the table below:


Iran/Azerbaijan_Gharbi+Tehran_(Assyrian) 1
Iran/Lorestan_(Lur) 1
Iran/Tehran_(Armenian) 1
Iran/Azerbaijan_Gharbi_(Azeri) 2
Iran/Hormozgan_(Bandari+Afro-Iranian) 2
Iran/Hormozgan/Qeshmi 2
Iran/Khorasan_(Persian) 2
Iran/Kurdistan_(Kurd) 2
Iran/Sistan_Baluchestan_(Baluch) 2
Pakistan 2
Iran/Fars+Isfahan_(Persian) 3
Iran/Gilan_(Gilak) 3
Iran/Yazd+Tehran_(Zoroastrian) 3
Turkey/Central 3
Turkey/East 3
Turkey/West_ 3
Iran/Golestan_(Turkmen) 4
India 4
Iran/Khuzestan_(Arab) 5
Egypt_(Arab) 5
Iraq/Baghdad 5
Oman 5
Saudi_Arabia 5
Tunisia 5
United_Arab_Emirates 5
Iran/Mazandaran_(Mazandarani) 6
Iran/Yazd_(Persian) 6
Balkarian 6
Georgia 6
Albania 7
Greece 7
Bosnia 8
Croatia 8
Slovenia 8
Czech_Republic 9
Hungary 9
Poland 9
Ukraine 9
Iraq_(Marsh_Arab) 10
Qatar 10
Yemen 10


We can ignore cluster #4 which consists of the two outliers (India + Turkmen). The rest of the clusters seem relatively coherent. Notice, for example, the Arabian cluster #10, Balkan cluster #8, Eastern European cluster #9, Greek-Albanian cluster #7, Mixed Arab cluster #5.

PLoS ONE 7(7): e41252. doi:10.1371/journal.pone.0041252

Ancient Migratory Events in the Middle East: New Clues from the Y-Chromosome Variation of Modern Iranians

Viola Grugni et al.


Knowledge of high resolution Y-chromosome haplogroup diversification within Iran provides important geographic context regarding the spread and compartmentalization of male lineages in the Middle East and southwestern Asia. At present, the Iranian population is characterized by an extraordinary mix of different ethnic groups speaking a variety of Indo-Iranian, Semitic and Turkic languages. Despite these features, only few studies have investigated the multiethnic components of the Iranian gene pool. In this survey 938 Iranian male DNAs belonging to 15 ethnic groups from 14 Iranian provinces were analyzed for 84 Y-chromosome biallelic markers and 10 STRs. The results show an autochthonous but non-homogeneous ancient background mainly composed by J2a sub-clades with different external contributions. The phylogeography of the main haplogroups allowed identifying post-glacial and Neolithic expansions toward western Eurasia but also recent movements towards the Iranian region from western Eurasia (R1b-L23), Central Asia (Q-M25), Asia Minor (J2a-M92) and southern Mesopotamia (J1-Page08). In spite of the presence of important geographic barriers (Zagros and Alborz mountain ranges, and the Dasht-e Kavir and Dash-e Lut deserts) which may have limited gene flow, AMOVA analysis revealed that language, in addition to geography, has played an important role in shaping the nowadays Iranian gene pool. Overall, this study provides a portrait of the Y-chromosomal variation in Iran, useful for depicting a more comprehensive history of the peoples of this area as well as for reconstructing ancient migration routes. In addition, our results evidence the important role of the Iranian plateau as source and recipient of gene flow between culturally and genetically distinct populations.

Link

February 29, 2012

Serbian Y-chromosomes

Gene. 2012 Jan 31. [Epub ahead of print]

High levels of Paleolithic Y-chromosome lineages characterize Serbia.

Regueiro M, Rivera L, Damnjanovic T, Lukovic L, Milasin J, Herrera RJ.

Abstract

Whether present-day European genetic variation and its distribution patterns can be attributed primarily to the initial peopling of Europe by anatomically modern humans during the Paleolithic, or to latter Near Eastern Neolithic input is still the subject of debate. Southeastern Europe has been a crossroads for several cultures since Paleolithic times and the Balkans, specifically, would have been part of the route used by Neolithic farmers to enter Europe. Given its geographic location in the heart of the Balkan Peninsula at the intersection of Central and Southeastern Europe, Serbia represents a key geographical location that may provide insight to elucidate the interactions between indigenous Paleolithic people and agricultural colonists from the Fertile Crescent. In this study, we examine, for the first time, the Y-chromosome constitution of the general Serbian population. A total of 103 individuals were sampled and their DNA analyzed for 104 Y-chromosome bi-allelic markers and 17 associated STR loci. Our results indicate that approximately 58% of Serbian Y-chromosomes (I1-M253, I2a-P37.2, R1a1a-M198) belong to lineages believed to be pre-Neolithic. On the other hand, the signature of putative Near Eastern Neolithic lineages, including E1b1b1a1-M78, G2a-P15, J1-M267 and J2-M172 and R1b1a2-M269 accounts for 39% of the Y-chromosome. Furthermore, an examination of the distribution of Y-chromosome filiations in Europe indicates extreme levels of Paleolithic lineages in a region encompassing Serbia, Bosnia-Herzegovina and Croatia, possibly the result of Neolithic migrations encroaching on Paleolithic populations against the Adriatic Sea.

Link

November 16, 2011

Armenian Y-chromosomes revisited (Herrera et al. 2011)

Armenian Y-chromosomes have been a largely ignored since the publication of the classic Weale et al. (2001) paper a decade ago. The Armenian DNA Project has largely covered the void during the intervening years, but it is nice that the topic is revisited by academics.

Armenia is sandwiched between Anatolia, the Fertile Crescent, the Iranian plateau, the Caucasus, and the Black and Caspian seas, making the study of Armenian Y-chromosomes extremely interesting for the student of Eurasian prehistory.

Gene flow from the surrounding regions may have affected the Armenian population over historical time, but the remoteness of the Armenian highlands, coupled with the national church -- which distinguished Armenians from both the Orthodoxy of the Roman Empire, the Zoroastrianism of the Persians, and, later the Islam of Arabs and Ottomans -- may have prevented it.

My comments on the paper will follow below once I read it.

UPDATE I: The paper spends a lot of time on analysis of Y-STR variance; my opinion of Y-STRs as a tool for inferring past population movements is, to put it mildly, low. When Bahamian Y-STR variance is higher than African one, and E-V13, one of the youngest European Y-haplogroups (in terms of Y-STR variance) turns up in Spain in one of the earliest ancient DNA samples, it goes without saying that the burden of proof is on those who wish to continue to talk about Neolithic or other population movements to make the assumptions of their models clearer. Nonetheless, there is still some utility in Y-STRs, so I reproduce some tree diagrams from the paper (top left), and link to the supplementary info that has a collection of haplotypes that may be useful to genealogists.

From the paper:
However, owing to the contentions associated with the current calibrations of the Y-STR mutation rates,32,34,35,41 as well as the limitations of the assumptions utilized by the methodologies for time estimations, the absolute dates generated in this study should only be taken as rough estimates of upper bounds.
Indeed. We are at the point where Y-STRs are at the end of their utility, but the replacement technology of extensive Y-chromosome sequencing has not quite arrived in an economical way yet.


UPDATE II:
I will have some additional thoughts on Y-chromosome distribution in the third update, but, for the time being, the two most important "nuggets" of information are: (i) the unusual haplogroup frequencies in Sasun (high R2 and T), which may be due to a founder effect, but it would be interesting if Armenian historians could find some explanation for their occurrence there, and (ii) the occurrence of R-M269*(xL23) in Ararat Valley. I invite more knowledgeable readers to comment on the issue; the haplotypes are in Table 2 of the supplement.

UPDATE III: The ubuiquity of haplogroup G2a in Neolithic Europe, coupled with the absence of other prominent present-day European haplogroups, has important implications about European discontinuity.

But, it also has implications about West Asian discontinuity. The Neolithic in Europe arrived by all accounts from either of two principal areas: Anatolia or the Levant. Today, in Anatolia and the Levant, we see a set of haplogroups of which haplogroup J is the most important and ubiquitous one. Haplogroup R1b is also quite frequent in Armenia, the east Caucasus, Anatolia, and Iran, but its frequency drops dramatically to the east and south. And, there is a whole assortment of other haplogroups with varying frequency.

Why didn't all these non-G2a haplogroups participate in the early Neolithic colonization of Europe? It could very well be that a very small founder population crossed the Aegean into Europe, one that happened to be G2a-dominated. But, that is ultimately not very satisfying: if there was plenty of J and R1b in West Asia at the time of the Neolithic expansion, why are these haplogroups so conspicuous in their absence -at least so far- from Neolithic Europe?

The case of haplogroup J is particularly problematic. If we had to guess, by looking at present-day distribution, which lineage tracks population movements from the Near East to Europe, there is simply no better candidate: every map of this haplogroup, and especially of its J2a sublineage shows an unambiguous pattern of radiation, with a core area consisting of Southern Italy, Greece, Anatolia, West Asia, Mesopotamia and the northern parts of the Levant. All these regions are crucial to the story of the Neolithic, so the absence of J in Neolithic Europe is perplexing.

And, the story has other complications. From the current paper:
The relative expansion times for haplogroup J2-M172 (Table 4) generally correspond with those yielded for R1b-M343, with the exception of Greece and Crete, which, unlike haplogroup R1b-M343, are slightly older than the dates yielded for several of the Near Eastern groups as well as the four Armenian populations.
As mentioned above, I don't give much weight on Y-STR evidence, but observations such as the above certainly add to the feeling of unease that something is not quite right with the default picture of prehistory.

Another observation on the Armenian population, is its very low frequency of haplogroup R1a1. Proponents of the Kurgan model of Indo-European dispersals sometimes associate this haplogroup with the Proto-Indo-European community, and it is strange why -if their ideas are right- Armenia is so lacking in this haplogroup, like its Caucasian neighbors. Why would these hypothetical migrants make such a huge impact in faraway India and barely a dent in nearby Armenia?

Finally, the occurrence of some I2, E-V13, and, perhaps, J2b in Armenia may point to Balkan contacts. But, when did these contacts occur? Are they traceable to the migration of Phrygians to Anatolia, according to the Herodotean account of Armenian origins, or can they be attributed to later contacts with Greeks or other Europeans?

The veil of mystery seems to be raised even higher by every new study: we may be less certain of what really happened today than in the days of happy ignorance, ten years ago. Ultimately it is new data, like the ones included in this paper, that will make every piece of evidence fit, and the grand puzzle of the history of Eurasia will be revealed in all its glory.

European Journal of Human Genetics , (16 November 2011) | doi:10.1038/ejhg.2011.192

Neolithic patrilineal signals indicate that the Armenian plateau was repopulated by agriculturalists

Kristian J Herrera, Robert K Lowery, Laura Hadden, Silvia Calderon, Carolina Chiou, Levon Yepiskoposyan, Maria Regueiro, Peter A Underhill and Rene J Herrera

Abstract
Armenia, situated between the Black and Caspian Seas, lies at the junction of Turkey, Iran, Georgia, Azerbaijan and former Mesopotamia. This geographic position made it a potential contact zone between Eastern and Western civilizations. In this investigation, we assess Y-chromosomal diversity in four geographically distinct populations that represent the extent of historical Armenia. We find a striking prominence of haplogroups previously implicated with the Agricultural Revolution in the Near East, including the J2a-M410-, R1b1b1*-L23-, G2a-P15- and J1-M267-derived lineages. Given that the Last Glacial Maximum event in the Armenian plateau occured a few millennia before the Neolithic era, we envision a scenario in which its repopulation was achieved mainly by the arrival of farmers from the Fertile Crescent temporally coincident with the initial inception of farming in Greece. However, we detect very restricted genetic affinities with Europe that suggest any later cultural diffusions from Armenia to Europe were not associated with substantial amounts of paternal gene flow, despite the presence of closely related Indo-European languages in both Armenia and Southeast Europe.

Link

September 14, 2011

The Caucasus revisited (Yunusbayev et al. 2011)


This is another treasure trove of a paper, and together with Balanovsky et al. (2011) we now have a very clear picture of genetic variation in this most interesting of world regions.

Here is the ADMIXTURE analysis:

The authors also post results up to K=10 in the supplementary material, which show Druze/Bedouin/Basque-centered component. It is actually possible to push the analysis higher than K=7 without such problem components appearing, by retaining non-closely related individuals (using --genome in PLINK and then iteratively removing individuals from pairs with PI_HAT greater than some value).

Nonetheless, the components emerging from this analysis will be familiar to followers of the Dodecad Project. In terms of Dodecad v3:
  • light yellow "North East Asian"
  • orange "South East Asian"
  • brown "Neo African" or "Sub_Saharan", as there are no African hunter-gatherers
  • dark blue "North European", as there is no split of east/west Europe at this level
  • middle blue "West Asian"
  • light blue "Southwest Asian"
  • green "South Asian", but anchored on Sindhi, a population from Pakistan, due to the lack of more southern populations from India
The labels of new populations sampled in this study can be seen in brown. I particularly hope that the substantial new autosomal data will become publicly available, so that I can use them in the Dodecad Project. It will be an invaluable new resource, filling some "holes" in the Eurasian landscape (e.g., east of the Caspian; Bulgarians; several new Caucasus populations) in the Li et al. (HGDP), and Behar et al. data.

(to be continued)

UPDATE I (Y-chromosomes):


Some observations:
  • C has a concentration in the Turkic Nogays
  • The presence of D this far west is very surprising, again in the Nogays. This haplogroup has a relic distribution, with particular concentrations in Tibet, Mongolia, Japan, and Andaman Islanders. In all likelihood its presence here is linked to the Nogays' eastern origin
  • E and its subclades occurs at a very low frequency here
  • G2a has a clear West Caucasus (both north and south) concentration
  • I seems to have a mainly West Caucasus distribution as well; this is a common European haplogroup; it has quite elevated frequencies among the Andis and Kara Nogays. It would be interesting to discover some historical correlate for the presence of I in Kara Nogays but not Kuban Nogays and in Andis but not in most of the NE Caucasus
  • J1 has the expected Northeast Caucasus nexus. This haplogroup is bimodal, with a mode in Arabians and a secondary mode in NE Caucasus. Note the paucity of J1e-P58, the reverse of the situation of Arabians; I've noted before the likely association of the P58 clade with Semitic languages.
  • The extreme concentration of J2 in Chechens and Ingush are probably associated with low variance. Apart from these atypical populations, a substantial presence of this haplogroup can be found in the NW/S Caucasus in different populations and in the form of different subclades.
  • The new LT mystery clade has its usual low-frequency wide distribution
  • N occurs in Nogays as expected, and, like C, also in the NW Caucasus. This probably also represents an eastern influence, probably associated not only with the Nogays but also with various Tatar influences on the Caucasus.
  • Q occurs widely in the NW Caucasus but only in 1 Nogay. Perhaps this is more of a Tatar marker, although a finer-scale resolution of this haplogroup is really necessary.
  • R1a-related lineages occur less frequently here among eastern Slavs, a main reason for the disconnect between the Eastern European plain and the Caucasus. There does, however, appear to be good diversity here, with the presence of R1a*, R1a1-M198*, Note again how the Iranic Ossetians (both North and South) have almost no R1a1 compared to both their NW Caucasian and S Caucasian neighbors, again, suggesting that this may not have been an important Alan or steppe Iranian lineage, at least during the late antique time horizon. The occurrence of R1a1f-M458 may represent Slavic influence in the NW Caucasus.
  • R1b-related lineages seem ubuiquitous in the Caucasus. R-M73 occurs substantially in Kara Nogays and Balkars, an apparent link with Central Asia where this haplogroup occurs frequently.
UPDATE II (Caucasus-Eastern Europe discontinuity)

The authors of this paper highlight the genetic discontinuity between the eastern European plain and the Caucasus. This was also apparent in the Balanovsky et al. (2011) paper, and was also a major conclusion of the Dodecad Project, with Caucasians exhibiting a high percentage of the "West Asian" component, while eastern Slavs low "West Asian" and high "East European".

The interpretation of this discontinuity is more difficult. There are surely parts of the Caucasus region that are mountainous and pose an ecological contrast to the flatlands of eastern Europe. That is consistent with a different type of population living in either region for a long time, despite the well-attested archaological contacts (e.g., Maikop or the settlement of steppe nomads such as Alans or Sarmatians).

On the other hand, the eastern Slavic population can, at least in part, have expanded more recently, in the medieval period, as part of the early Slavic dispersals, as well as the push to the north and east of the Russians. These appear to have partly displaced Turkic groups from the north Pontic region, with all of the above having displaced historical Scythian (Iranic) nomads, who, in turn, displaced the mysterious Cimmerians. If the discovery of east Eurasian mtDNA C in Neolithic and Bronze Age Ukraine stands up, there will be another layer of population replacement, as mtDNA C is quite rare in the broader region today. On the other hand, the Caucasus itself may have been affected from population movements from the Near East, as Balanovsky et al. suggest.

So, in conclusion, the discontinuity is a fact that emerges from different types of analyses, but its causes remain uncertain, and it is not clear when and how it was first established.

Mol Biol Evol (2011) doi: 10.1093/molbev/msr221

The Caucasus as an asymmetric semipermeable barrier to ancient human migrations

Bayazit Yunusbayev et al.

Abstract

The Caucasus, inhabited by modern humans since the Early Upper Paleolithic and known for its linguistic diversity, is considered to be important for understanding human dispersals and genetic diversity in Eurasia. We report a synthesis of autosomal, Y chromosome and mitochondrial DNA (mtDNA) variation in populations from all major subregions and linguistic phyla of the area. Autosomal genome variation in the Caucasus reveals significant genetic uniformity among its ethnically and linguistically diverse populations, and is consistent with predominantly Near/Middle Eastern origin of the Caucasians, with minor external impacts. In contrast to autosomal and mtDNA variation, signals of regional Y chromosome founder effects distinguish the eastern from western North Caucasians. Genetic discontinuity between the North Caucasus and the East European Plain contrasts with continuity through Anatolia and the Balkans, suggesting major routes of ancient gene flows and admixture.

Link