Showing posts with label N1c. Show all posts
Showing posts with label N1c. Show all posts

May 21, 2015

More Y-chromosome super-fathers

The time estimates are based on a mutation rate of 1x10-9 mutations/bp/year which is ~1/3 higher than mutation rate of Karmin et al.  So the values on the table may be a little lower.

There may be additional founders with recent time depths than shown in the table, e.g., a very shallow clusters within E-M35 (probably E-V13?) and a couple of shallow clusters within I-P215

Also of interest is the fact that Greeks and Anatolian Turks do not show evidence of the recent Y-chromosomal bottleneck:
The plots are consistent with patterns seen in the relative numbers of singletons, described above, in that the Saami and Palestinians show markedly different demographic histories compared with the rest, featuring very recent reductions, while the Turks and Greeks show evidence of general expansion, with increased growth rate around 14 KYA. A different pattern is seen in the remaining majority (13/17) of populations, which share remarkably similar histories featuring a minimum effective population size ~2.1–4.2 KYA (considering the 95% confidence intervals (CIs) reported in Supplementary Table 4), followed by expansion to the present.


Related:
Nature Communications 6, Article number: 7152 doi:10.1038/ncomms8152

Large-scale recent expansion of European patrilineages shown by population resequencing

Chiara Batini, Pille Hallast et al.

The proportion of Europeans descending from Neolithic farmers ~10 thousand years ago (KYA) or Palaeolithic hunter-gatherers has been much debated. The male-specific region of the Y chromosome (MSY) has been widely applied to this question, but unbiased estimates of diversity and time depth have been lacking. Here we show that European patrilineages underwent a recent continent-wide expansion. Resequencing of 3.7 Mb of MSY DNA in 334 males, comprising 17 European and Middle Eastern populations, defines a phylogeny containing 5,996 single-nucleotide polymorphisms. Dating indicates that three major lineages (I1, R1a and R1b), accounting for 64% of our sample, have very recent coalescent times, ranging between 3.5 and 7.3 KYA. A continuous swathe of 13/17 populations share similar histories featuring a demographic expansion starting ~2.1–4.2 KYA. Our results are compatible with ancient MSY DNA data, and contrast with data on mitochondrial DNA, indicating a widespread male-specific phenomenon that focuses interest on the social structure of Bronze Age Europe.

Link

May 03, 2015

Structure of Y-haplogroup N

arXiv:1504.06463 [q-bio.PE]

The dichotomy structure of Y chromosome Haplogroup N

Kang Hu et al.

Haplogroup N-M231 of human Y chromosome is a common clade from Eastern Asia to Northern Europe, being one of the most frequent haplogroups in Altaic and Uralic-speaking populations. Using newly discovered bi-allelic markers from high-throughput DNA sequencing, we largely improved the phylogeny of Haplogroup N, in which 16 subclades could be identified by 33 SNPs. More than 400 males belonging to Haplogroup N in 34 populations in China were successfully genotyped, and populations in Northern Asia and Eastern Europe were also compared together. We found that all the N samples were typed as inside either clade N1-F1206 (including former N1a-M128, N1b-P43 and N1c-M46 clades), most of which were found in Altaic, Uralic, Russian and Chinese-speaking populations, or N2-F2930, common in Tibeto-Burman and Chinese-speaking populations. Our detailed results suggest that Haplogroup N developed in the region of China since the final stage of late Paleolithic Era.

Link

June 21, 2013

Origins and dispersals of Y-chromosome haplogroup N

I will simply note that the authors use the effective mutation rate that is ~1/3 the genealogical mutation rate and hence their age estimates are inflated by ~3x. I have expressed reservations about using Y-STR based age estimates in general, but these concerns become more important for older lineages.

In particular, I would be very surprised if Y-haplogroup N turns up in Europe 8-10 thousand years ago, and I expect to see it make its first appearance in the 3rd millennium BC or thereabouts, perhaps together with the Seima-Turbino expansion across northern Eurasia. Thanks to the ancient DNA -preserving boreal cold, it may be possible to find out.

Irrespective of my disagreement on the mutation rate issue, I have to applaud the comprehensive survey carried out by these Chinese scientists: numbers invariably pay off.

PLoS ONE 8(6): e66102. doi:10.1371/journal.pone.0066102

Genetic Evidence of an East Asian Origin and Paleolithic Northward Migration of Y-chromosome Haplogroup N

Hong Shi et al.

The Y-chromosome haplogroup N-M231 (Hg N) is distributed widely in eastern and central Asia, Siberia, as well as in eastern and northern Europe. Previous studies suggested a counterclockwise prehistoric migration of Hg N from eastern Asia to eastern and northern Europe. However, the root of this Y chromosome lineage and its detailed dispersal pattern across eastern Asia are still unclear. We analyzed haplogroup profiles and phylogeographic patterns of 1,570 Hg N individuals from 20,826 males in 359 populations across Eurasia. We first genotyped 6,371 males from 169 populations in China and Cambodia, and generated data of 360 Hg N individuals, and then combined published data on 1,210 Hg N individuals from Japanese, Southeast Asian, Siberian, European and Central Asian populations. The results showed that the sub-haplogroups of Hg N have a distinct geographical distribution. The highest Y-STR diversity of the ancestral Hg N sub-haplogroups was observed in the southern part of mainland East Asia, and further phylogeographic analyses supports an origin of Hg N in southern China. Combined with previous data, we propose that the early northward dispersal of Hg N started from southern China about 21 thousand years ago (kya), expanding into northern China 12–18 kya, and reaching further north to Siberia about 12–14 kya before a population expansion and westward migration into Central Asia and eastern/northern Europe around 8.0–10.0 kya. This northward migration of Hg N likewise coincides with retreating ice sheets after the Last Glacial Maximum (22–18 kya) in mainland East Asia.

Link

March 18, 2013

Thesis of Oleg Balonovsky

is available here as pdf. Lots of interesting information, and a few striking maps. Hopefully, the fact that it's all in Russian won't be much of a problem in this day and age.

I will highlight a few pieces of information. First, a distribution of Y-chromosome haplogroups in Russian groups:

Notice:

  • N1c-Tat is a general feature of the Russians, but N1b-P43 is only really found at any significant frequency in the northern groups.
  • A strong contrast of E-M78 between central (present) and northern (absent) groups, consistent with a late introduction of this haplogroup in easternmost Europe.
  • South-Central-North decreasing frequency of R1a; now, it's not clear how R1a came to be in Russians: some of it may be legacy of its initial entry into Europe from the east, other could be of historical import, and may have even arrived during the Slavic expansion from Central Europe. The pattern probably is the reverse of the high frequency of N1, indicating increasing importance of Finno-Ugric substratum in the north.
  • Fairly interesting that of the two likely "Balkan" haplogroups E-M78 and I-P37, the former is modal in central region, the latter in southern one. The absence of both in "deep Asia" suggests a late introduction, as mentioned before, but when?
Also of interest a haplotype analysis within R1a1a-M198:


My most immediate observation is the set of mainly Indian highly divergent haplotypes on the left. There has been (well-deserved) excitement about recent Y-SNP progress within this haplogroup, but we should not neglect the occurrence of outliers/relics in our reconstruction of a haplogroup's history. I'd love to see those few Indian haplotypes SNP-tested using the currently available SNPs, or even used to develop new SNPs for this important Eurasian haplogroup.

September 05, 2012

East to West across Eurasia

A couple more interesting abstracts from the DNA in Forenscics 2012.


Genetic journey of the N1c haplogroup
Pamjav H, Nemeth E, Feher T, Volgyi A
Binary and Y-STR polymorphisms associated with the NRY region of the human Y chromosome preserve the paternal genetic legacy that has persisted to the present, permitting inference of human evolution, population migration and demographic history.The NRY region of the Y chromosome acts much like mtDNA to reveal the structure among human populations and possiblyto infer the order and timing of their descents. In the present study, we have investigated the originof haplogroup N1c-Tat phylogeographic structure and the genetic relationship of Eurasianpopulations by examining STR variation in a large number of individuals. We have identified 54samples as the haplogroup N1c-Tat from 5 population groups (N=632). To place the results into awider geographic context, we included 209 samples from published sources and 296 samples from the FTDNA public database into the phylogenetic analysis. According to previous studieshaplogroup N-M231 is of East Asian ancestry. Our results suggest that N1c-Tat mutation probably originated in South Siberia 8-9 thousand years ago and had spread through the Urals into the European part of present-day Russia. Its distribution is not fully correlated with the spread of Uralic languages. Turkic-speaking ethnic groups in South Siberia have high N1c-Tat presence and STR variance, while the N1c-L550 subgroup largely occurs among non-Uralic-speaking Europeanpopulations. Only the European N1c-Tat (xL550) subgroup can be linked to the spread of Finno-Ugric languages from the Kama-Urals area ~6,000 years ago. The subgroup N1c-L550 cannot be considered Finno-Ugric origin and its carriers might have been assimilated by Indo-European groups, resulting in their spread across Europe in historical times with Vikings and Balto-Slavs. Based on the present study Buryats were dominated by a young, about 800-years old N1c-Tat cluster, which suggest that this ethnic group could be a relatively recent admixture of Mongolian conquerors with a Paleo-Siberian population groups.
Of course these ages should be taken with a grain of salt because it is unclear how they were derived (i.e., whether the "evolutionary mutation rate" was used). Hopefully, someone will treat the  subject of N1c ages with Y-SNPs that do not have the problem of saturation that affects microsatellites. This is an interesting test case, because a ~3-fold change in ages will have important consequences for our understanding of the spread of Finno-Ugric languages into Europe: an earlier date would associate them with the Comb Ceramic, while a later, Bronze Age date would associate them with the Seima-Turbino phenomenon.


Huns in Bavaria? Genetic analyses of an artificially deformed skull from an early medieval cemetery in Burgweinting (Regensburg, Germany)

Schleuder R, Wilde S, Burger J, Grupe G, Forster P, Harbeck M
The morphological examination of an early medieval burial site in Burgweinting, which is dated to the end of the 5th century, revealed one female with an artificially, circularly deformed skull, a practice that is thought to be associated with the arrival of Nomads of the Eurasian steppe, particularly the Huns.    

Individuals with such artificial cranial deformations also can be found in other Late Roman and Early Medieval cemeteries in Europe mostly in the Carpathian basin but only as few isolated cases in Western Europe, where mostly women show such deformations.  
Regarding the artificial cranial deformations it is unclear whether a foreign custom was taken over by Germanic tribes or whether the individuals were members or descendants of Eurasian nomads.  
With the help of the find of Burgweinting, we exemplarily investigated this question.To identify the possible foreign origin of this female with alleged “Asian” skull deformation we sequenced the HVRI and HVRII region of the mitochondrial DNA.  
Our results show that the ancestry of a woman with artificially deformed skull can be linked to an at least partly Asian origin. So this indicates that at least some of the few individuals with skull deformation had not adopted the costume but can be seen as former members or descendants of the hunnish tribal community.   
It will be worthwhile if geneticists can co-operate with physical anthropologists and/or archaeologists more broadly in cases where morphology, or burial customs indicate that a possibly heterogeneous population exists at that site. The above is a good example of that synergy in action.

August 22, 2012

East Eurasian-like admixture in Northern Europe (part 2)

This is a continuation of my earlier post. Please refer to it for the methodology. A new part 3 can be found here.

I have repeated the experiment with a much larger set of populations:
English_D, British_D, Ukranians_Y,  Karitiana, Spaniards, Sardinian,  Serb_D, Mordovians_Y, Irish_D,  French, Finnish_D, Chuvashs_16,  Romanian_D, N_Italian_D, French_Basque,  Austrian_D, Russian_D, Hungarians_19,  Kent_1KG, German_D, Belorussian,  Tuscan, Lithuanian_D, Orkney_1KG,  Dutch_D, TSI30, Ukrainian_D,  Bulgarians_Y, Bulgarian_D, Russian,  Swedish_D, Pais_Vasco_1KG, French_D,  Castilla_Y_Leon_1KG, Lithuanians, San,  Polish_D, Romanians_14, Orcadian,  Cornwall_1KG, Valencia_1KG, North_Italian,  FIN30, Norwegian_D, CEU30
I used Sardinians as the Caucasoid reference population, Karitiana for Mongoloids, and San for Africans. The latter two were chosen because they live at maximally opposite corners of the Earth (South America vs. South Africa).

A first plot of the f4 statistics used for f4 regression ancestry estimation is seen below:

Clearly, some evidence of a cline is present, but several populations appear to deviate from it. In order to get the cleanest possible cline, I carried out the following greedy procedure: I calculate the correlation coefficient of this set, and iteratively remove one population that leads to the maximum improvement of the correlation, until no further improvement takes place. The following populations were removed with this procedure:

Spaniards, Serb_D, Romanian_D, N_Italian_D, Tuscan, TSI30, Bulgarians_Y, Bulgarian_D, Castilla_Y_Leon_1KG, Romanians_14, Valencia_1KG
This seems to make sense, as all these are southern European populations. Note that their removal does not mean that they do not partake in the same phenomenon as northern Europeans: they also exhibit Karitiana-shift relative to the Sardinians, but there are probably other confounding factors that make them fall "off-cline". Including them would diminish the clarity of the cline for Northern European populations. The regression of the remaining populations can be seen on the right:



f4 regression ancestry estimation results are shown on the left. These appear to be much higher than was the case with the Han and Dai in the previous experiment.

I can't say that I've made any obvious mistakes, but these admixture proportions are substantial, and call for an explanation. Whatever their true levels, I am fairly confident on at least a few points:

First, it is evident that northern Europeans have higher levels of this element than southern Europeans; the latter are not altogether deficient in it, but they fall "off-cline", making estimation of their admixture proportions more difficult.

Second, within northern Europe, there is a fairly clear east-west cline of diminishing Amerasian-like admixture. The minimum occurs in Sardinians and secondarily in Southwest Europe. Romance, Celtic, and Germanic populations all have less of it than Balto-Slavic and Uralic ones. And, some populations of northeastern Europe seem to have a noticeable excess of it.

The groups with the most Amerasian-like admixture possess Y-haplogroup N, a clear trace of eastern ancestry that is not shared by most Europeans. The arrival of this haplogroup, either with Comb Ceramic of the Baltic Neolithic or later with Seima Turbino Bronze Age expansions is probably responsible for the local excess in Northeastern Europe. The Chuvash are, of course, a Turkic population but of Finno-Ugrian genetic origin.

But, the presence of this element even in Western Europe cannot be explained on the basis of typically Mongoloid elements which are almost completely lacking there. If Mesolithic Europeans were themselves Asian-shifted, then this would account for the presence of the element, but not necessarily for its clinal manifestation. The double (north-south and east-west) cline indicates every sign of an intrusive element. So, for the time being, I will propose that this is associated with late (e.g., Copper and Bronze Age) phenomena, such as the northern stream of the Bronze Age Indo-European invasion of Europe.

This may be due to the

  • (i) northern Indo-European groups picking up some native east European or Siberian elements as they made their way into Europe, 
  • or (ii), more likely, in my opinion, that the Y-haplogroup R1 group of people, whose closest relatives are in Central/South Asia (R2) , and whose more distant relatives (Q) are in Siberia and the Americas, were from the beginning an "intermediate population" between West and East Eurasia. The R1 group of people in its R1b and R1a varieties first appear in Europe during the Copper Age, and they are lacking in early Neolithic sites.


Eight years ago, and in a totally different context, I wrote:

Similarly, 9 out of 10 Basques are descended from a man who has also fathered 9 out of 10 Kets from Siberia and 9 out of 10 Maya Indians from America. That man, founder of haplogroup P thus has descendants who belong to two of the major human races (or three, if Amerindians are considered as separate from Asian Mongoloids)   
... 
In conclusion, human continental populations form groups of genetic and phenotypic similarity, and these groups can be considered races in the phenetic sense. However, these groups are not monophyletic, hence in the cladistic sense they should not be considered as valid taxa. Since the principle of common descent is generally applied in modern systematics (or at least it should!), I think it's best not to recognize human subspecies. 

If these data pan out, it may be revealed that the European branch of the Caucasoids is actually a product of admixture too, with at least two of its constituent elements being the "Palaeo-West Eurasians" (Y-haplogroups G, IJ, possibly LT) and the "Neo-NW Eurasians" (Y-haplogroups N1 and R1), with the "Neo-Afrasians" (Y-haplogroup E1b1b) forming a third element.

(A raw dump of fourpop output can be found here).

September 14, 2011

The Caucasus revisited (Yunusbayev et al. 2011)


This is another treasure trove of a paper, and together with Balanovsky et al. (2011) we now have a very clear picture of genetic variation in this most interesting of world regions.

Here is the ADMIXTURE analysis:

The authors also post results up to K=10 in the supplementary material, which show Druze/Bedouin/Basque-centered component. It is actually possible to push the analysis higher than K=7 without such problem components appearing, by retaining non-closely related individuals (using --genome in PLINK and then iteratively removing individuals from pairs with PI_HAT greater than some value).

Nonetheless, the components emerging from this analysis will be familiar to followers of the Dodecad Project. In terms of Dodecad v3:
  • light yellow "North East Asian"
  • orange "South East Asian"
  • brown "Neo African" or "Sub_Saharan", as there are no African hunter-gatherers
  • dark blue "North European", as there is no split of east/west Europe at this level
  • middle blue "West Asian"
  • light blue "Southwest Asian"
  • green "South Asian", but anchored on Sindhi, a population from Pakistan, due to the lack of more southern populations from India
The labels of new populations sampled in this study can be seen in brown. I particularly hope that the substantial new autosomal data will become publicly available, so that I can use them in the Dodecad Project. It will be an invaluable new resource, filling some "holes" in the Eurasian landscape (e.g., east of the Caspian; Bulgarians; several new Caucasus populations) in the Li et al. (HGDP), and Behar et al. data.

(to be continued)

UPDATE I (Y-chromosomes):


Some observations:
  • C has a concentration in the Turkic Nogays
  • The presence of D this far west is very surprising, again in the Nogays. This haplogroup has a relic distribution, with particular concentrations in Tibet, Mongolia, Japan, and Andaman Islanders. In all likelihood its presence here is linked to the Nogays' eastern origin
  • E and its subclades occurs at a very low frequency here
  • G2a has a clear West Caucasus (both north and south) concentration
  • I seems to have a mainly West Caucasus distribution as well; this is a common European haplogroup; it has quite elevated frequencies among the Andis and Kara Nogays. It would be interesting to discover some historical correlate for the presence of I in Kara Nogays but not Kuban Nogays and in Andis but not in most of the NE Caucasus
  • J1 has the expected Northeast Caucasus nexus. This haplogroup is bimodal, with a mode in Arabians and a secondary mode in NE Caucasus. Note the paucity of J1e-P58, the reverse of the situation of Arabians; I've noted before the likely association of the P58 clade with Semitic languages.
  • The extreme concentration of J2 in Chechens and Ingush are probably associated with low variance. Apart from these atypical populations, a substantial presence of this haplogroup can be found in the NW/S Caucasus in different populations and in the form of different subclades.
  • The new LT mystery clade has its usual low-frequency wide distribution
  • N occurs in Nogays as expected, and, like C, also in the NW Caucasus. This probably also represents an eastern influence, probably associated not only with the Nogays but also with various Tatar influences on the Caucasus.
  • Q occurs widely in the NW Caucasus but only in 1 Nogay. Perhaps this is more of a Tatar marker, although a finer-scale resolution of this haplogroup is really necessary.
  • R1a-related lineages occur less frequently here among eastern Slavs, a main reason for the disconnect between the Eastern European plain and the Caucasus. There does, however, appear to be good diversity here, with the presence of R1a*, R1a1-M198*, Note again how the Iranic Ossetians (both North and South) have almost no R1a1 compared to both their NW Caucasian and S Caucasian neighbors, again, suggesting that this may not have been an important Alan or steppe Iranian lineage, at least during the late antique time horizon. The occurrence of R1a1f-M458 may represent Slavic influence in the NW Caucasus.
  • R1b-related lineages seem ubuiquitous in the Caucasus. R-M73 occurs substantially in Kara Nogays and Balkars, an apparent link with Central Asia where this haplogroup occurs frequently.
UPDATE II (Caucasus-Eastern Europe discontinuity)

The authors of this paper highlight the genetic discontinuity between the eastern European plain and the Caucasus. This was also apparent in the Balanovsky et al. (2011) paper, and was also a major conclusion of the Dodecad Project, with Caucasians exhibiting a high percentage of the "West Asian" component, while eastern Slavs low "West Asian" and high "East European".

The interpretation of this discontinuity is more difficult. There are surely parts of the Caucasus region that are mountainous and pose an ecological contrast to the flatlands of eastern Europe. That is consistent with a different type of population living in either region for a long time, despite the well-attested archaological contacts (e.g., Maikop or the settlement of steppe nomads such as Alans or Sarmatians).

On the other hand, the eastern Slavic population can, at least in part, have expanded more recently, in the medieval period, as part of the early Slavic dispersals, as well as the push to the north and east of the Russians. These appear to have partly displaced Turkic groups from the north Pontic region, with all of the above having displaced historical Scythian (Iranic) nomads, who, in turn, displaced the mysterious Cimmerians. If the discovery of east Eurasian mtDNA C in Neolithic and Bronze Age Ukraine stands up, there will be another layer of population replacement, as mtDNA C is quite rare in the broader region today. On the other hand, the Caucasus itself may have been affected from population movements from the Near East, as Balanovsky et al. suggest.

So, in conclusion, the discontinuity is a fact that emerges from different types of analyses, but its causes remain uncertain, and it is not clear when and how it was first established.

Mol Biol Evol (2011) doi: 10.1093/molbev/msr221

The Caucasus as an asymmetric semipermeable barrier to ancient human migrations

Bayazit Yunusbayev et al.

Abstract

The Caucasus, inhabited by modern humans since the Early Upper Paleolithic and known for its linguistic diversity, is considered to be important for understanding human dispersals and genetic diversity in Eurasia. We report a synthesis of autosomal, Y chromosome and mitochondrial DNA (mtDNA) variation in populations from all major subregions and linguistic phyla of the area. Autosomal genome variation in the Caucasus reveals significant genetic uniformity among its ethnically and linguistically diverse populations, and is consistent with predominantly Near/Middle Eastern origin of the Caucasians, with minor external impacts. In contrast to autosomal and mtDNA variation, signals of regional Y chromosome founder effects distinguish the eastern from western North Caucasians. Genetic discontinuity between the North Caucasus and the East European Plain contrasts with continuity through Anatolia and the Balkans, suggesting major routes of ancient gene flows and admixture.

Link

January 26, 2010

Ancient DNA from frozen Yakuts

From the paper:
Sixty one percent (8 out of 13) of the haplotypes (Ht1, Ht2, Yaka56, 65, 71, 80, 81, 86) were affiliated to the N1c (TAT-C) haplogroup on the basis of the SNP analyses. This haplogroup is considered as the most frequent in the Yakut population, and its frequency varies across studies from 75% [12] to 100% [13]. Sample YAKa26 was affiliated to haplogroups K . The SNP typing was inconclusive for 5 individuals (YAKa17, 19, 47, 49 and 57); nevertheless the affiliation to N1c was excluded on the basis of the absence of the TAT-C mutation.
and:
The origin of the most frequent Y-chromosomal haplotypes (Ht1 and Ht2) was difficult to establish on the basis of genetic information. Indeed, these two lineages belonging to haplogroup N1c seem to be restricted to Yakut populations, and were probably present since the period they were first located in Central Yakutia. Interestingly, the comparison with archaeological data revealed that the male individuals (YAKa34, 39, 40, 69, 78) at the beginning of the 18th century, identified as Clan Chiefs (or tojons) on the basis of their grave goods (weapons, jewelry, silk clothes, richly ornamented saddles and signet rings), belonged to these two haplotypes. Therefore, archaeological data could bring interesting information in tracing back the origin of these enigmatic male lineages. Indeed, the grave goods of the 15th/17th centuries (weapons and horse harnesses) and the construction of coffins with an empty trunk from the 18th century are similar to the burial customs of the Cis-Baïkal area [44] and of the Egyin Gol Necropolis during the 3rd century BC [45-47]. This suggests that the male ancestors of the Yakuts were probably formed of a small group of horse-riders originating from Northern Mongolia or the Baïkal Lake.
and:
Based on the analyses of the maternal and paternal lineages of ancient Yakuts, we were able to demonstrate that the formation of this population started before the 15th century, with a small group of settlers composed of horse-riders from the Cis-Baïkal region and a small number of women from different South Siberian origins.
BMC Evolutionary Biology doi:10.1186/1471-2148-10-25

Human evolution in Siberia: from frozen bodies to ancient DNA

Eric Crubezy et al.

Abstract (provisional)

Background
The Yakuts contrast strikingly with other populations from Siberia due to their cattle- and horse-breeding economy as well as their Turkic language. On the basis of ethnological and linguistic criteria as well as population genetic studies, it has been assumed that they originated from South Siberian populations. However, many questions regarding the origins of this intriguing population still need to be clarified (e.g. precise origin of paternal lineages and admixture rate with indigenous populations). This study attempts to better understand the origins of the Yakuts, by performing genetic analyses on 58 mummified frozen bodies dated from the 15th to the 19th century, excavated from Yakutia (Eastern Siberia).

Results
High quality data were obtained for the autosomal STRs, Y-chromosomal STRs and SNPs and mtDNA due to exceptional sample preservation. A comparison with the same markers on seven museum specimens excavated 3 to 15 years ago showed significant differences in DNA quantity and quality. Direct access to ancient genetic data from these molecular markers combined with the archaeological evidence, demographical studies and comparisons with 166 contemporary individuals from the same location as the frozen bodies, helped us to clarify the microevolution of this intriguing population.

Conclusion
We were able to trace the origins of the male lineages to a small group of horse-riders from the Cis-Baikal area. Furthermore, mtDNA data showed that intermarriages between the first settlers with Evenks women led to the establishment of genetic characteristics during the 15th century that are still observed today.

Link (pdf)

June 13, 2009

Y chromosomes of Turks from Antalya

Of interest the presence of R1*(R1a, R1b) at a frequency of 6.6%. Unfortunately the presence of haplogroups likely to have been introduced to Turkey from Central Asia was not directly measured, although BR*(xD2,E,F) is a good candidate for a haplogroup C stand-in.

Rom J Leg Med 17 (1) 59 – 68 (2009)

Y-SNP haplogroups in the Antalya population in Turkish Republic

Timur Serdar, Demircin Sema

Abstract

SNPs are known to be the most abundant source of sequence variation in the human
genome. The SNPs in the NRY (non-recombining Y-chromosome) region which passes from father to son as unchanged haplotype-blocks escaping recombination, provides important advantages in the investigations of sexual assault crimes, in the cases of parentage testing especially if the mother or alleged father is unavailable for testing and in the evolutionary studies. The aim of this study, was to determine the frequencies of Y SNP markers and the haplogroups, in order to define the Y-chromosome SNP markers which are polymorphic, have high discrimination power and can be used in forensic investigations in the Antalya population. For each of 75 unrelated males from Antalya, 35 different Y-SNP markers were amplified in a single reaction using multiplex minisequencing method. In the study, 18 markers of them were found to be polymorphic. The most frequent YSNP markers with mutations were M139 (100%), SRY10831/SRY1532 (92%), M89 (85.3%), M213 (85.3%), M9 (44%), 92R7 (30.6%), 12F2 (30.6%), M45 (29.3%), M172 (26.6%) and M173 (22.6%). The Y-chromosome haplogroups of Antalya population were defined by these 18 Y-SNP polymorphic loci and the frequencies and the distribution of haplogroups were determined. J2*(xJ2F2) (26.6%), K*(xN3,O,P) (13.3%), E3b (9.3%), F*(xH,I,J,K) (8%), R1a1*(xR1a1b) (8%), R1b*(xR1b1, R1b6, R1b8) (8%), P*(xQ3a,R1) (8%) haplogroups were identified as the most abundant in Antalya population. These haplogroups are reported as widespread also in European and neighboring Near Eastern populations.

Link (pdf)

May 07, 2009

Citation of my Y-STR mutation rate criticism

I was reading Tuuli Lappalainen Ph.D. dissertation, at the University of Helsinki: "Human genetic variation in the Baltic Sea region: Features of population history and natural selection," and, to my surprise, I noticed that my post on How Y-STR variance accumulates: a comment on Zhivotovsky, Underhill and Feldman (2006) was cited:
Estimating the age of haplogroups is important for connecting genetic patterns to historical phenomena. However, it is dependent on the correct estimation of the mutation rate, which has proven to be difficult. Rates calculated from pedigrees are 3-4 times higher than evolutionary rates (Parsons et al. 1997, Howell et al. 2003, Dupuy et al. 2004, Zhivotovsky et al. 2004, Zhivotovsky et al. 2006), and it is unclear which should be used for the calculation of the most recent common ancestor for major haplogroups in large geographic regions. It has recently been suggested (Pontikos 2008) that the widely used evolutionary rate of the Y chromosome (Zhivotovsky et al. 2004) is strongly underestimating the effective mutation rate due to not accounting for population growth and the bias of analyzing the biggest haplogroups that have grown at rates exceeding the general growth rate of the population. These analyses have not been published in a peer-reviewed journal, but they appear to correctly point out at least some problems of the commonly used models. Thus, the appropriate mutation rate to use for analyzing the temporal scale of the Y-chromosomal haplogroup variation may be a few times lower than was used in II – close to the pedigree rate. The same bias should apply to mitochondrial DNA, too. If the revised rates (Pontikos 2008) were used instead, TMRCAs for the main Y-chromosomal haplogroups I1a, N3 and R1a1 would be in the order of 3000-4000 years before present. These dates would imply that instead of the proposed Neolithic arrival of these haplogroups, their upper age limit would be in late Neolithic or early Bronze Age. Interestingly, the revised age of N3 variation in the Baltic Sea region would actually correspond nicely with the recently suggested Bronze Age arrival of the Finno-Ugric language (Häkkinen 2009). However, given the current uncertainty of the appropriate mutation rates, all time estimates should be used with great caution.
The cited post was the first one in the now extensive Y-STR series in which I have tried to dissect various aspects Y-STR based age estimation.

April 15, 2009

Genes of Finns revisited

This paper interprets the discrepancy between Y-chromosome and mtDNA results in Finland as the signature of Scandinavian gene flow into the western parts of the country, with the Y-chromosome gene pool of the east (typified by haplogroup N3) preserving the original inhabitants starting from the Holocene deglaciation.

It it is not at all clear, however "who got there first", and as far as I can see, the evidence just tells us there is a substantial east-west difference in Y-chromosomes in Finland, it doesn't really tell us which of the two elements represents the most ancient stratum.

In my opinion, the Finnish gene pool may contain traces of the aboriginal inhabitants, as well as the later eastern elements which brought the Finnish language, and the later still influences by Germanic Scandinavians. Hopefully the northern cold has been generous with DNA preservation and we may get some direct glimpses into the country's genetic history.

Some related posts:

European Journal of Human Genetics doi:

Genetic markers and population history: Finland revisited

Jukka U Palo et al.

Abstract

The Finnish population in Northern Europe has been a target of extensive genetic studies during the last decades. The population is considered as a homogeneous isolate, well suited for gene mapping studies because of its reduced diversity and homogeneity. However, several studies have shown substantial differences between the eastern and western parts of the country, especially in the male-mediated Y chromosome. This divergence is evident in non-neutral genetic variation also and it is usually explained to stem from founder effects occurring in the settlement of eastern Finland as late as in the 16th century. Here, we have reassessed this population historical scenario using Y-chromosomal, mitochondrial and autosomal markers and geographical sampling covering entire Finland. The obtained results suggest substantial Scandinavian gene flow into south-western, but not into the eastern, Finland. Male-biased Scandinavian gene flow into the south-western parts of the country would plausibly explain the large inter-regional differences observed in the Y-chromosome, and the relative homogeneity in the mitochondrial and autosomal data. On the basis of these results, we suggest that the expression of 'Finnish Disease Heritage' illnesses, more common in the eastern/north-eastern Finland, stems from long-term drift, rather than from relatively recent founder effects.

Link

March 07, 2009

Y chromosome distribution in northwestern Russia

As usual, the use of the "effective mutation rate" renders the age estimates in this paper useless (they need to be divided roughly by 3).

Indeed, in this paper they attempt to use Batwing to estimate ages using the effective rate. Batwing employs a Bayesian method with coalescent simulations, and thus takes into account "population history", the effects of which are supposedly encapsulated in the effective mutation rate. Thus, they are "correcting" (inappropriately of course) for population history twice.

This is clearly evident in the Table, where the Batwing age estimates exceed significantly those based on Y-STR variance, prompting the authors to reject the Batwing results as not "credible". Not credible indeed, if one blindly picks a "mutation rate" and a piece of software from the literature and combines them to arrive at an "estimate".

The R1a1 age estimates are properly all within a Neolithic time frame. Of course, the network topologies and associated Y-STR variance argue strongly against a simple Out-of-Eastern Europe scenario of the dispersal of R1a1, as non-star topologies with very high variance are found in India and Pakistan. This parallels another recent study in which a substantial subset of Indian R1a1 Y-chromosomes appeared to be distinctive from those of Europe.

Together, with the recent work on horse domestication, these results point to the fact that there is something wrong in the equation of R1a1 "PIE-speaking Bronze Age horse riders from the Pontic-Caspian steppe". Clearly, the picture is more complex, and will only be resolved when new SNPs resolve the phylogeny of this widespread haplogroup.

Also of interest:
As older ages are observed when grouping All Asians versus All Europeans (Table 5) for N1c, the available data suggest that the mutation may have originated in northern China as previously reported,14,15 but may have traversed through a different migratory route than has been postulated elsewhere,15 reaching northeastern European populations before the Urals


European Journal of Human Genetics doi:10.1038/ejhg.2009.6

Y-Chromosome distribution within the geo-linguistic landscape of northwestern Russia

Sheyla Mirabal et al.

Abstract

Populations of northeastern Europe and the Uralic mountain range are found in close geographic proximity, but they have been subject to different demographic histories. The current study attempts to better understand the genetic paternal relationships of ethnic groups residing in these regions. We have performed high-resolution haplotyping of 236 Y-chromosomes from populations in northwestern Russia and the Uralic mountains, and compared them to relevant previously published data. Haplotype variation and age estimation analyses using 15 Y-STR loci were conducted for samples within the N1b, N1c1 and R1a1 single-nucleotide polymorphism backgrounds. Our results suggest that although most genetic relationships throughout Eurasia are dependent on geographic proximity, members of the Uralic and Slavic linguistic families and subfamilies, yield significant correlations at both levels of comparison making it difficult to denote either linguistics or geographic proximity as the basis for their genetic substrata. Expansion times for haplogroup R1a1 date approximately to 18 000 YBP, and age estimates along with Network topology of populations found at opposite poles of its range (Eastern Europe and South Asia) indicate that two separate haplotypic foci exist within this haplogroup. Data based on haplogroup N1b challenge earlier findings and suggest that the mutation may have occurred in the Uralic range rather than in Siberia and much earlier than has been proposed (12.9plusminus4.1 instead of 5.2plusminus2.7 kya). In addition, age and variance estimates for haplogroup N1c1 suggest that populations from the western Urals may have been genetically influenced by a dispersal from northeastern Europe (eg, eastern Slavs) rather than the converse.

Link

October 23, 2008

Y chromosomes and mtDNA of Sweden

It's interesting to see a study of the current population of a country. If one is interested in deep pre-historical origins, one must ensure that his sample has no known foreign ancestry among the lines of interest. But, if one is interested in the current and future course of the population, then the totality of the population must be sampled.

This is a good opportunity to track shifting gene frequencies due to immigration and/or differential fertility. It would be a good idea for countries to start spending some money on a genetic census of their population. This would not need to involve all the inhabitants, and could be carried out for the fraction of cost that governments pay to collect all sorts of other statistics. Such a census would provide an important source of data to future scientists investigating the demography of Europe during this transitional era.

From the paper:
Several haplogroups had interesting frequency patterns, but also wide confidence intervals, necessitating caution in the interpretations. The mtDNA haplogroup with the strongest geographical cline, U5b, is known to have high frequencies among the northern Saami population, consistent with our results (Tambets et al. 2004). The high frequency of the Y-chromosomal R1b in the south, observed also by Karlsson et al. 2006; is consistent with its abundance in Central Europe and Denmark (Semino et al. 2000; Brion et al. 2005). Haplogroup R1a1 is more common in Norway than in Sweden (Dupuy et al. 2006), and its high frequency in Värmland/Dalarna and Halland supports the historically plausible connection to Norway (Lindqvist, 2006). Y-chromosomal haplogroup N3 (Lappalainen et al. 2006) and mtDNA haplogroup H1f (Loogväli et al. 2004; Lappalainen et al. 2008) are common in Finland, and had increased frequencies in several Swedish counties with historical ties to Finland: Eastern Svealand was the most important destination of the Finnish immigration wave in the 1970's; in Norrland the Finnish influences date back to ancient times and in Dalarna to the 17th century (Pitkänen 1994).

When compared to previous knowledge of ethnic Swedes without immigration in their familial background (Lappalainen et al. 2008), the frequencies of several haplogroups showed effects of 20th century immigration from more distant countries. The Y-chromosomal I1a had decreased frequencies in Malmö and Gothenburg most probably due to replacement by haplogroups that are common among immigrants. African immigration contributes to the frequencies of mtDNA haplogroups L3*(xN,M) and L* (xL3) (Chen et al. 2000), and Y-chromosomal haplogroup A (Underhill et al. 2001; Jobling & Tyler-Smith 2003), while Near Eastern influence can be seen in mtDNA haplogroup U7 and possibly J (Richards et al. 2000; Abu-Amero et al. 2007; Achilli et al. 2007). Asian and American immigration can be observed in the slightly elevated frequencies of mtDNA haplogroups M, A, C, D and G (Quintana-Murci et al. 2004; Hill et al. 2007) and the Y-chromosomal O, K* and P* (Underhill et al. 2001; Jobling & Tyler-Smith 2003). The frequency of the Y-chromosomal haplogroup I1b may associate to immigrants from Balkan and Eastern Europe (Rootsi et al. 2004). In Malmö and Gothenburg immigration was the main contributor to their isolated positions in the Y-chromosomal PCA plot and probably also to the higher diversities compared to the surrounding populations. These phenomena were not observed in Stockholm, where most of the immigrants come from Finland (Statistics Sweden, http://www.scb.se).

Annals of Human Genetics doi: 10.1111/j.1469-1809.2008.00487.x

Population Structure in Contemporary Sweden—A Y-Chromosomal and Mitochondrial DNA Analysis

T. Lappalainen et al.

Abstract

A population sample representing the current Swedish population was analysed for maternally and paternally inherited markers with the aim of characterizing genetic variation and population structure. The sample set of 820 females and 883 males were extracted and amplified from Guthrie cards of all the children born in Sweden during one week in 2003. 14 Y-chromosomal and 34 mitochondrial DNA SNPs were genotyped. The haplogroup frequencies of the counties closest to Finland, Norway, Denmark and the Saami region in the north exhibited similarities to the neighbouring populations, resulting from the formation of the Swedish nation during the past millennium. Moreover, the recent immigration waves of the 20th century are visible in haplogroup frequencies, and have led to increased diversity and divergence of the major cities. Signs of genetic drift can be detected in several counties in northern as well as in southern Sweden. With the exception of the most drifted subpopulations, the population structure in Sweden appears mostly clinal. In conclusion, our study yielded valuable information of the structure of the Swedish population, and demonstrated the usefulness of biobanks as a source of population genetic research. Our sampling strategy, nonselective on the current population rather than stratified according to ancestry, is informative for capturing the contemporary variation in the increasingly panmictic populations of the world.

Link

July 21, 2008

How Y-STR variance accumulates: a comment on Zhivotovsky, Underhill and Feldman (2006)

An important erratum for this post.

Additions to this entry at the bottom (last update July 29)


In recent years, in most population genetics papers, an evolutionary mutation rate for Y chromosome microsatellites (STRs) of 0.00069/locus/generation has been used. This rate was proposed by Zhivotovsky et al. (2004) (pdf), and defended in Zhivotovsky et al. (2005), and especially Zhivotovsky, Underhill and Feldman (2006) (henceforth Z.U.F.)

This mutation rate is smaller than the observed germline mutation rate by a factor of 3-4. The germline mutation rate is observed by counting mutations directly, e.g., in father-son pairs, or in known pedigrees. Zhivotovsky et al. have provided two pieces of evidence in favor of their evolutionary rate:
  • Study of accumalation of STR variation in populations with known founding events, namely Bulgarian Roma and Maori, in their 2004 paper.
  • Simulations indicating a 3.6x discrepancy between the two rates in their 2006 paper, which is due to multiple bottlenecks in a haplogroup's history.
I was always apprehensive about what the "right" mutation rate should be:
We need to obtain good estimates of the mutation rate in order to pinpoint in time the common ancestor of a set of Y chromosomes. A factor of 3, especially for relatively recent events may correspond to a difference between early historical and late Paleolithic events.
Thus, I decided to look into the matter myself to be convinced -one way or another- of what the evolutionary mutation rate must be.

Methodology

The following assumptions, following Z.U.F. are made:
  • A man has 0, 1, 2, ... sons according to a Poisson process with mean m=1.
  • A step mutation (increase or decrease by 1 repeat) occurs with a mutation rate of µ=0.00251
  • STR variance of the man's descendants is measured after g generations.
Results are averaged over N men who have descendants after g generations. I will call such men, "Patriarchs". Thus, I generate random family trees for men until I have harvested N=10,000 of them who have living descendants today.2

Patriarch vs. MRCA

A consequence of the time-forward methodology of simulation, is that a Patriarch may not be the Most Recent Common Ancestor (MRCA) of his descendants g generations into the future. Trivially, if a Patriarch has only one son, then, that son -not the Patriarch- is the MRCA of his descendants. But, even if the Patriarch has many sons, and his group of descendants grows, it is possible (due to randomness of the fathering process) that at some generation only 1 descendant will survive.

Suppose that the Patriarch has lived in generation 0, and the MRCA lived in generation i. Thus, STR variance in the descendants at generation g (today) has accumulated over a time span of g-i generations, since, of course, at the generation i (of the MRCA), STR variance is zero.

Now, if we use a time-forward methodology from known foundation events (e.g. the arrival of the Roma in Bulgaria, or the Maori in New Zealand), it is perfectly right to see how STR variance accumulates from the known foundational event. We would then divide the accumulated STR variance by the known time span to determine an effective evolutionary mutation rate, similar to Zhivotovsky et al. (2004).

But, when the foundational event is unknown, when we are trying to estimate its age, then we can only go as far back as the MRCA, since at his time variance is zero. Therefore, by dividing accumulated variance with the evolutionary mutation rate of Z.U.F., we are over-estimating the time to the MRCA.

For example, with g=100, the average STR variance for the descendants of N=10,000 Patriarchs is 0.0755. But, if we average only those Patriarchs who are also the MRCA of their descendants, we obtain a value of 0.0824, or about 9% higher.

In general, the over-estimate (as a percentage) decreases as g increases: as g increases, the average number of descendants of a Patriarch increases, making them much less susceptible to a variance-reset type of bottleneck described here.

Thus, while the age difference between the MRCA and the Patriarch is real, its effect in the age estimate is not very pronounced. There is, however, a second, and much more serious problem, with the Z.U.F. rates when applied to evolutionary studies.

Prolific vs. Non-Prolific Patriarchs: an Observation Selection effect

Patriarchs starting at generation 0 will have a very variable number of descendants at generation g. By averaging over all of them, we are estimating the average STR variance in the descendants of men who lived g generations ago.

Now, consider how this average changes if we average only over the k most "prolific" men (with the most descendants) out of all the N=10,000 Patriarchs:


k
Average Variance
100
0.1721
1000
0.1407
2500
0.1219
5000
0.1033
10000
0.0755


It is clear that the STR variance in the descendants of the most "prolific" Patriarchs is much higher than in the descendants of the least "prolific" ones. In fact, for the most prolific Patriarchs, variance accumulates near the germline mutation rate, and not at the lower evolutionary effective rate.

Below is the cumulative percentage of the descendants of the k most prolific Patriarchs, with k from 1 to N.

It can be seen that e.g., from the most prolific half of the Patriarchs stems 84% of the descendants. And this, assuming no social inequality in the number of progeny, i.e. each man having the exact same average probability (m=1) of fathering a son. Thus, in reality, the more prolific Patriarchs may have an even larger fraction of the descendants.

Why is this important? Because, in population studies, scientists are likely observe (in the finite samples they collect) multiple descendants only of the most prolific of the Patriarchs. Thus, for the vast majority of the Patriarchs with few descendants, we are likely to sample no, or few of their descendants.

This means that there is an inherent observation selection effect in the types of Patriarchs we are likely to study: they are much more likely to be among the prolific ones. Coupling this observation with the knowledge that STR variance in the descendants of prolific Patriarchs accumulates near the germline mutation rate (0.69µ for the 100 most prolific ones in my experiment), we, once again, conclude that the STR variance in haplogroups likely to be made the object of scientific study accumulates near the germline mutation rate, and at the very least, faster than the evolutionary rate of Z.U.F.

Closing Remarks

Z.U.F. have also proposed two additional demographic scenaria under which a higher effective mutation rate would be observed:
  • A sudden jump in the size of the haplogroup after it appears
  • An expanding population (m>1)
Both factors seem reasonable for post-Holocene human populations. It is well known that -whatever temporary setbacks there were- mankind has overall experienced a substantial population growth in recent millennia. Thus, an expanding population seems like a fair assumption.

Moreover, it is reasonable to assume that in stratified human societies, a few males, (leaders, or conquerors), or groups of closely related males may have generated a disproportionate number of descendants in the short-term.

In summary:
  • The age difference between the Patriarch and the MRCA indicates that Variance/0.00069 overestimates the age of the MRCA somewhat (but not very much).
  • A prolific Patriarch's descendants are more likely to be sampled by scientists, and tend to have a higher STR variance. Hence, Variance/0.00069 overestimates the age of the MRCA, perhaps substantially.
  • Demographic factors, such as population growth, or short-term success by related males indicates that Variance/0.00069 overestimates the age of the MRCA.
In view of the above, and keeping in mind both the stochastic factors that cause STR variance to fluctuate around its expected value, as well as uncertainties in demographic history, I do believe that ages calculated with the evolutionary mutation rate of 0.00069/locus/generation are significantly overestimated.

1 Z.U.F. used a germline mutation rate of µ=0.001. For the purposes of simulation, this is not an important difference, as they themselves note. I choose the rate of 0.0025 because it is closer to the actual human germline mutation rate for STRs.
2 Z.U.F. generated 50,000 men and then averaged over the men who had descendants. I, on the other hand, generate as many men as it takes to harvest at least N men with descendants, to ensure that I average a substantially large number of such men.

Editorial change (Jul 22): erroneously written "exceeds",in paragraph 2, changed to "is smaller than".

Update (July 23):

To further elucidate how the observation selection effect may make lineages seem older than they really are, I carried out another small experiment (g=110, N=10,000, m=1).

The age of each group is inferred by dividing the accumulated variance by the evolutionary rate of 0.0006944 (=μ/3.6).

The average variance over all N in this experiment is 0.0867, thus, the average inferred age is 125 generations, close to the truth (110 generations), allowing for the correction in age between the Patriarch and the TMRCA.

However, if we calculated the average variance over ten groups of 1,000 lineages (out of all N=10,000) according to the number of descendants, we see, as described above, that more "prolific" lineages have accumulated more variance, whereas less "prolific" ones have accumulated less variance than the overall average of 0.0867.

Thus, over the 10% most populous lineages (right of the figure), the average inferred age is 209 generations, or a 90% overestimate of the true age!

But, as I mentioned, it is precisely these populous lineages (which don't just have "some" descendants today, but thousands and millions of them) that are likely to be studied, because they are the only ones that have enough representatives in a sample of 100-1,000 men, typically seen in a population study, to allow for an age estimate via a variance calculation.

Update (July 24): Haplogroup sizes

The number of a Patriarch's descendants after g generations is a random variable which depends on the parameters m (the population growth constant), and g, the number of generations.

Scientists typically look at haplogroups with thousands or millions of existing members. Are such haplogroups produced in the types of simulations performed by Z.U.F.?

I estimate the average size of the haplogroups of the haplogroups produced by Z.U.F. for different g=10,20,...,700 and m=1.

It is evident that this number increases linearly with g at a rate estimated to be 0.5/generation [This was also noted by Z.U.F. who state: "the average size of the surviving haplogroups increased each generation by a value rapidly approaching 0.5"] However, this means, that the average haplogroup at 700 generations has a size of ~350 men.

Thus, not only is the average variance estimated by Z.U.F. inappropriate because of an observation selection effect (averaging over small and large haplogroups alike), but it seems to miss the relevant observations altogether, i.e. the really large haplogroups numbering in the hundreds of thousands or millions. Yet it is precise for such large haplogroups that it has often be used in the literature.

How can we produce "realistic" haplogroup sizes, close to those likely to become an object of scientific study in contemporary human populations? We can either:
  • increase the number of initial representatives, i.e. start with many related men with identical Y chromosomes rather than just 1, or we can
  • increase the population growth constant m to something higher than 1, i.e. a growing population.
Yet, both these changes have the same effect, namely the accumulation of variance at a higher rate than the Z.U.F. rate.

Indeed, Z.U.F. produce some such large haplogroups in some of their simulations (Fig. 1 asterisks, Fig. 2 squares/diamonds), all of which show -predictably- a higher effective rate than their 3.6x slower rate.

They caution against such large haplogroup sizes ["population size exceeds 1 million by generation 1000, which is not realistic for many local tribes."]. Granted, -- if one looks at local tribes never growing to large numbers.

And yet, some or all of the co-authors of Z.U.F. did not limit their use of the 3.6x slower rate to local tribes: Cinnioglu et al. 2004 (pdf), Sengupta et al. (2006), King et al. (2008) all apply the 0.00069 rate for populations (and haplogroups) that have grown to much more than 1 million in less time, thus overestimating severely their age.

Update (July 24): Variance of a large haplogroup

Following the previous observations, naturally, I wanted to see for myself what the STR variance of an ancient lineage with a large number of modern descendants actually looks like. My target size is 1,000,000, which is about 20% of modern Greek males.

I consider two cases:
  • Expansion commencing in the Late Bronze Age (g=120 or 1,600BC with a generation length of 30)
  • Expansion commencing in the early Neolithic (g=300 or 7,000BC)

I harvest N=1,000 haplogroups for each of these cases. I set the growth constant at m=1.100694 for the Bronze Age, and m=1.039122 for the Neolithic. This ensures that enough "large" haplogroups will be generated during simulation. Naturally, the overall population grows at a smaller rate, but the successful lineages will grow much faster than the population average.

Note that I harvest only haplogroups whose MRCA lived in the specified time span. Also, I harvest haplogroups whose final size is between 750,000 and 1,250,000 to match my target size of 1,000,000. Indeed, the average size of the harvested haplogroups is 964,327 for the Bronze Age, and 979693 for the Neolithic.

Here are the results:
  • ~1 million descendants of a Bronze Age (120 generations ago) ancestor have an STR variance of 0.269 +/ 0.087
  • ~1 million descendants of a Neolithic (300 generations ago) ancestor have an STR variance of 0.629 +/- 0.156
If we used the germline mutation rate (μ=0.0025) we would estimate the ages of these haplogroups as:
  • Bronze Age: 107.6 generations, or a 10% underestimate
  • Neolithic: 251.6 generations, or a 16% underestimate
On the other hand, if we used the evolutionary rate of 0.00069 of Z.U.F., our estimates would be:
  • Bronze Age: 389.9 generations, or a 225% overestimate
  • Neolithic 911.6 generations, or a 203% overestimate
It is clear that the Z.U.F. rate of 0.00069 substantially overestimates the ages of large recent haplogroups, whereas the germline rate underestimates them by a little.

Let's look at some concrete examples of age estimates in the literature, where I compare my own (first) estimates with the published ones. Here is how my estimates are derived:

For a Bronze Age ancestor (g=120) it is: 0.269 =(approx) 0.9 μg

For a Neolithic ancestor (g=300) it is: 0.629 =(approx) 0.84 μg

Thus, the correction multiplier, if the variance is between 0.269 and 0.629 is between 0.84 and 0.9; I will use the midpoint 0.87. If the variance is less than 0.269, then I use 0.9. If the variance is more than 0.629 then I use 0.84. Of course, the correction factor could be expressed more accurately as a function of the variance.

Note that the generation length preferred by these authors is 25, by me it is 30. All ages are ky BC.

Cinnioglu et al. (2004)

In this paper, an evolutionary rate of 0.0007 is used.



Variance
Cinnioglu
Dienekes
E-M78
0.18
4.4
0.4
G-P15
0.35
10.5
2.9
I-P37
0.23
6.2
1.1
J-M12
0.24
6.6
1.2
J-M67
0.33
9.8
2.6
R-M269
0.33
9.8
2.6

E-M78 is dated to 400BC, only a couple of centuries after the historical Greek colonization. E-M78 reaches its maximum in the Peloponese, a major source of Greek colonists.

I-P37 and J-M12 are dated to 1,100BC and 1,200BC, at around the time that e.g. the Phrygians from the Balkans are believed to have migrated to Asia Minor. I-P37 and J-M12 reach their maxima in areas north of Greece where the Phrygians are said to have originated.

Sengupta et al. (2006)



Variance
Sengupta
Dienekes
J2-M410
0.38
11.7
3.3
R-M17
0.39
12
3.4
R-M17 (upper caste)
0.26
7.3
1.5
G-P15
0.29
8.5
2
J-M241
0.38
11.8
3.3

Thus, all the exogenous West Asian lineages in India have post-Neolithic ages, with R-M17 having a suggestive age of 1,500BC coinciding with the suggested date for the Indo-Aryans.

King et al. (2008)



Variance
King
Dienekes
J-M12 (Nea Nikomedeia)
0.18
4.7
0.4
E-V13 (Sesklo/Dimini)
0.24
6.6
1.2
E-V13 (Lerna Franchthi)
0.25
7.2
1.3
J-M92 (Crete)
0.14
3.1
0.1 AD
J-M319 (Crete)
0.14
3.1
0.1 AD
E-V13 (Crete)
0.09
1.1
0.8 AD

These are very localized samples, so they should not be interpreted as reflecting expansion times in Greece itself, however, they do suggest a Bronze Age expansion of E-V13 and a much later arrival of E-V13 in Crete.

Note that for Crete, the 1,000,000-haplogroup size assumption is a substantial overestimate, so my age estimates are also substantial underestimates.

Update (July 25): R-M17 in South Siberia

Derenko et al. (2006) "Contrasting patterns of Y-chromosome variation in South Siberian populations from Baikal and Altai-Sayan regions" calculate the variance of R-M17 chromosomes in South Siberia, using the Z.U.F. rate, arriving at an age of 11.3kya corresponding to a value of 0.31. This corresponds to 2,300BC according to my estimate (see previous update).

Recently Bouakaze et al. (Int J Legal Med (2007) 121:493–499) reported the presence of R-M17 chromosomes in ancient inhabitants of South Siberia and the Andronovo culture (2,500BC-1,500BC).

The Andronovo culture is widely believed to be of Eastern European ultimate origin, reflecting the eastward movement of the Kurgan culture, and is associated by some with the ancestors of the Indo-Iranians.

In the Balkans, again in Z.U.F. years, the age of R-M17 is 15.8kya corresponding to variation of 0.44, corresponding to ~4,000BC according to my estimate.

Update (July 25): Baltic Y chromosomes

Lappalainen et al. (2008) use the Z.U.F. rate to estimate the antiquity of lineages in the Baltic region. Dates are ky BC.



Lappalainen Dienekes
I1a
5.7
1
N3
6.8
1.5
R1a1
8.7
1.9

1,000BC for I1a in the Baltic region is within the time frame of the emergence of the Germanic people who did experience a strong demographic growth.
1,500BC for N3 shows a rather late time for Finno-Ugrians. However, it must be noted that smaller demographic sizes would impose more drift, and hence a slower accumulation of variance. Therefore, this time is probably underestimated.
1,900BC for R1a1 is consistent with the northern edge of the expansion of R1a1. Once again, reduced variance may also be influenced by smaller population numbers, making this a possible underestimate.

Update (July 25): Southeastern Europe (the Balkans)

Pericic et al. (2005) use the Z.U.F. rate to estimate ages of Y-chromosome lineages in the Balkans. Dates are ky BC.



Pericic
Dienekes
I1b* (xM26)
8.1
2
E3b1α
5.3
0.9
R-M17
13.8
3.8
R-M269
9.6
2.3
J-M241 (without Kosovars) 1
0.8AD

Thus, Balkan haplogroup I seems related to a Bronze Age origin, with R-M17 being substantially older, and deriving perhaps from northern Balkan Neolithic or alternatively intrusive Kurgan populations. J-M241 seems to be quite young, similar to J-M12 in Nea Nikomedeia (see discussion of King et al. (2008) above).

The young ages of J-M12 and J-M241 also explain the striking inverse correlation between it and J-M410, which makes sense if it expanded later. A fairly late expansion also explains its under-representation in Southern Italy and Anatolia: it appears to be a rather young and "Epirotic" clade that was too late in coming to significantly participate in the historical Greek colonization.

Update (July 26): E3b in Cyprus and Southern Italy

Capelli et al. (2005) [Population Structure in the Mediterranean Basin: A Y Chromosome Perspective] study Y-chromosome variation in many Mediterranean populations including Cyprus. I use a mutation rate of 0.0018 for the six markers used in this study (Quintana-Murci et al. AJHG 68(2) pp. 537 - 542 ). Ages are in ky BC.

I come up with an age of 1.4ky BC for E3b in Cyprus, which is consistent with Mycenaean and later Greek settlements on the island.

I also looked at Southern Italian Y chromosomes. I removed those with values other than (13,12) in DYS19,DYS388), since these are universal in Greek E-V13, in order to remove possible contamination from non E-V13 chromosomes. The resulting age is 900BC, once again very close to the historical Greek colonization of Magna Graecia.

July (26): A more elaborate population growth model

Z.U.F. also propose (Fig. 2 triangles) a more elaborate population growth with:
  • m=1.002 before 400 generations
  • m=1.012 from 400 to to 14 generations ago
  • m=1.12 from 14 to 8 generations ago
  • m=1.25 from 8 generations ago to current time

I ran a simulation (g=1000, N=10,000) with this population growth model. The average size of the descent groups of the MRCAs is 692,982 men. Averaged all of them, variance is 1.37.
  • With the germline mutation rate, an estimate of 549 generations (45% underestimate)
  • With the Z.U.F. evolutionary rate, an estimate of 1,988 generations (99% overestimate)
If we limit ourselves only to the 10, 1000, 5000 most prolific MRCAs (out of the N=10,000), we obtain ages (respectively):
  • With the germline mutation rate: 776, 747, 668 generations
  • With the Z.U.F. evolutionary rate: 2,810, 2,707, 2,419 generations

Thus, one can estimate that STR variance since the time of the MRCA accumulates at a rate of ~0.75μ / generation.

And, yet, the 0.00069 rate has been used to date Paleolithic events, e.g., by Semino et al. (2004) [Am. J. Hum. Genet. 74:1023–1034, 2004], leading to general age overestimates.

Update (July 29)

My discussion is continued in Haplogroup sizes and observation selection effects (continued)