Showing posts with label Germanic. Show all posts
Showing posts with label Germanic. Show all posts

May 28, 2016

British Celts have more steppe ancestry than British English

An interesting tidbit in a preprint about blood pressure genes:
We consistently obtained significantly positive f4 statistics, implying that both the modern Celtic samples and the ancient Saxon samples have more Steppe ancestry than the modern Anglo-Saxon samples from southern and eastern England. This indicates that southern and eastern England is not exclusively a genetic mix of Celts and Saxons.
Southeastern England is genetically very homogeneous. If the people there were a mix of ancient Celts and Saxons you'd expect them to be intermediate between modern Celts (who should have more Celtic ancestry than the modern English) and ancient Saxons (who should have more Saxon ancestry than the modern English).

But, it seems that the English have less steppe ancestry than both modern Celts and ancient Saxons, so they're not really intermediate. My guess is that the English have Norman ancestry that the Celts don't. While the original Normans were Scandinavians with presumably lots of steppe ancestry, I'd be surprised if the post-1066 Normans that settled England were not already heavily admixed with the "French" and so had less steppe ancestry than the modern British Celts from Wales and Scotland.

bioRxiv http://dx.doi.org/10.1101/055855

Population structure of UK Biobank and ancient Eurasians reveals adaptation at genes influencing blood pressure

Kevin Galinsky et al.

Analyzing genetic differences between closely related populations can be a powerful way to detect recent adaptation. The very large sample size of the UK Biobank is ideal for detecting selection using population differentiation, and enables an analysis of UK population structure at fine resolution. In analyses of 113,851 UK Biobank samples, population structure in the UK is dominated by 5 principal components (PCs) spanning 6 clusters: Northern Ireland, Scotland, northern England, southern England, and two Welsh clusters. Analyses with ancient Eurasians show that populations in the northern UK have higher levels of Steppe ancestry, and that UK population structure cannot be explained as a simple mixture of Celts and Saxons. A scan for unusual population differentiation along top PCs identified a genome-wide significant signal of selection at the coding variant rs601338 in FUT2 (p=9.16×10-9). In addition, by combining evidence of unusual differentiation within the UK with evidence from ancient Eurasians, we identified new genome-wide significant (p less than 5×10-8) signals of recent selection at two additional loci: CYP1A2/CSK and F12. We detected strong associations to diastolic blood pressure in the UK Biobank for the variants with new selection signals at CYP1A2/CSK (p=1.10×10-19)) and for variants with ancient Eurasian selection signals in the ATXN2/SH2B3 locus (p=8.00×10-33), implicating recent adaptation related to blood pressure.

Link

September 13, 2012

Polish and German Y-chromosomes

I have often bemoaned the use of present-day populations as stand-ins for dealing with the subject of very old archaeological phenomena such as the Neolithic transition. Of course, I understand that until a few years ago, this was all we had to work with. But, this idea is now suspect, having been made so by a two-pronged attack. On the ancient DNA side, researchers have consistently discovered that for the better part of prehistory, ancient populations did not match modern ones: even if the constituent elements of later evolution could be identified, they were still in polarized non-admixed form as in the case of the Neolithic Swedes. On the recent side, researchers have used surnames or even toponyms to show that ethnic admixture in the recent historical past has shifted Y-chromosome frequencies around.

A new paper in EJHG follows on this tradition by comparing pre- and post-WWII patterns of Y chromosome variation in Germany and Poland. Y-haplogroup frequencies can be seen on top left. From the caption: "Phylogenetic relationship and frequencies of Y-chromosomal haplogroups in the studied populations. Ka Kaszuby; Ko Kociewie; Ku Kurpie; Lu Lusatia; Sl Slovakia; Me Mecklenburg; Ba Bavaria."

It is important to note that the researchers were able to study pre-war populations, because most everybody knows where their patrilineal ancestor lived less than 100 years ago. But, European history consists of many events whose effects on the current population are less known, because they occurred at an older time. In some cases, populations may have migrated (such as the Germans of eastern Europe following WWII), in others populations that have once existed there have almost disappeared, or become much less numerically significant (such as the Ashkenazi Jews due to persecution during WWII and after it through migration to Israel and elsewhere, or various Christian and Jewish communities that once flourished throughout the Middle East). Other, less known groups, such as the Old Prussians or the Jassic speakers of Hungary have been presumably absorbed by surrounding majorities, or through a process of elite dominance.

In many cases, the available information, in the form of linguistic, genealogical, or historical evidence, can be used to remove layers of admixture, migration, and extinction in history; but, the gap between the deep prehistoric past and the recent, historical one cannot be bridged by these methods alone. Ultimately, ancient DNA researchers must close in on the present by targeting more recent populations for analysis. As the realization of genetic change continues to amass on both sides of the divide, I suspect that this will come naturally, although I do expect some reticence to the findings as they begin to touch upon the most cherished origins traditions of the multitude of extant European nations.


European Journal of Human Genetics advance online publication 12 September 2012; doi: 10.1038/ejhg.2012.190

Contemporary paternal genetic landscape of Polish and German populations: from early medieval Slavic expansion to post-World War II resettlements

Krzysztof Rebala et al.

Abstract

Homogeneous Proto-Slavic genetic substrate and/or extensive mixing after World War II were suggested to explain homogeneity of contemporary Polish paternal lineages. Alternatively, Polish local populations might have displayed pre-war genetic heterogeneity owing to genetic drift and/or gene flow with neighbouring populations. Although sharp genetic discontinuity along the political border between Poland and Germany indisputably results from war-mediated resettlements and homogenisation, it remained unknown whether Y-chromosomal diversity in ethnically/linguistically defined populations was clinal or discontinuous before the war. In order to answer these questions and elucidate early Slavic migrations, 1156 individuals from several Slavic and German populations were analysed, including Polish pre-war regional populations and an autochthonous Slavic population from Germany. Y chromosomes were assigned to 39 haplogroups and genotyped for 19 STRs. Genetic distances revealed similar degree of differentiation of Slavic-speaking pre-war populations from German populations irrespective of duration and intensity of contacts with German speakers. Admixture estimates showed minor Slavic paternal ancestry (~20%) in modern eastern Germans and hardly detectable German paternal ancestry in Slavs neighbouring German populations for centuries. BATWING analysis of isolated Slavic populations revealed that their divergence was preceded by rapid demographic growth, undermining theory that Slavic expansion was primarily linguistic rather than population spread. Polish pre-war regional populations showed within-group heterogeneity and lower STR variation within R-M17 subclades compared with modern populations, which might have been homogenised by war resettlements. Our results suggest that genetic studies on early human history in the Vistula and Oder basins should rely on reconstructed pre-war rather than modern populations.

Link

August 22, 2012

East Eurasian-like ancestry in Northern Europe (part 3)

(This is the third part of the series. See part 1 and part 2.)

In the first two parts of the series, I showed that northern European populations show hints of East Eurasian ancestry when compared against Sardinians. I used Dai, Han, and Karitiana as reference populations for East Eurasia. In the current post, I extend this analysis by using HGDP Papuans and the Onge (Reich et al. 2009) from the Andaman Islands.

The f4 statistics using Karitiana, Papuan, and Onge populations can be found in this spreadsheet.

Below, you can see that they are all near perfectly correlated with each other.

The visual appraisal is confirmed when we calculate the correlation coefficients:


The fact that all three populations track the same signal is strong evidence for the direction of gene flow: from Asia into northern Europe. If the signal was present in only one of the three populations, then it could conceivably be an artefact of gene flow in the opposite direction (from northern Europeans to the affected population). But, the fact that all three populations show the same pattern would require northern European-like admixture in the Andaman Islands, Papuan New Guinea and South America, which does not appear very parsimonious.

While the signals from the three populations are correlated, their intensity varies. The Z-scores provide a measure of this intensity. The mean Z-scores using a Karitiana, Papuan, and Onge reference across all populations are respectively -17.7, -8.0, and -6.0.

While I did not include the Han reference of part 1 in this analysis, inspection of the f4 statistics (which can be obtained at the bottom of that part), suggests that the Z-scores become more significant when using an Onge, Papuan, Han, and Karitiana reference in that order. For example, for the Finnish_D population, they are: -10.037, -13.2949, -23.9305, and -27.764 respectively.

It thus appears that the element contributing East Eurasian-like ancestry in northern Europeans was derived from the northern spectrum of East Eurasians; the Karitiana may live in South America today, but they trace their ancestors to northern Eurasia, having entered the Americas c. 15ka.

In my opinion, the signal has been formed by a superposition of a few factors:

  1. The fact that Y-haplogroup R, the main lineage in modern northern Europeans has a common origin (Y-haplogroup P) with haplogroup Q, the main lineage in modern Amerindians, and many Siberians. We can hypothesize that the population that brought R into Europe was intermediate genetically across the Caucasoid-Mongoloid spectrum. In West Eurasia, this population admixed with the Palaeo-West Eurasians (Y-haplogroups IJ, G, and possibly LT), and contributed their DNA primarily to the northern Europeoids.
  2. Other population movements of more regional impact, such as Y-haplogroup N, which affected mainly Uralic, Baltic, and East Slavic populations, as well as elements from the mixed West/East Eurasian mtDNA contact zone that ancient DNA analysis has revealed in Eastern Europe and Siberia.
The raw dumps of fourpop output for Papuan and Onge reference can be found here.

East Eurasian-like admixture in Northern Europe (part 2)

This is a continuation of my earlier post. Please refer to it for the methodology. A new part 3 can be found here.

I have repeated the experiment with a much larger set of populations:
English_D, British_D, Ukranians_Y,  Karitiana, Spaniards, Sardinian,  Serb_D, Mordovians_Y, Irish_D,  French, Finnish_D, Chuvashs_16,  Romanian_D, N_Italian_D, French_Basque,  Austrian_D, Russian_D, Hungarians_19,  Kent_1KG, German_D, Belorussian,  Tuscan, Lithuanian_D, Orkney_1KG,  Dutch_D, TSI30, Ukrainian_D,  Bulgarians_Y, Bulgarian_D, Russian,  Swedish_D, Pais_Vasco_1KG, French_D,  Castilla_Y_Leon_1KG, Lithuanians, San,  Polish_D, Romanians_14, Orcadian,  Cornwall_1KG, Valencia_1KG, North_Italian,  FIN30, Norwegian_D, CEU30
I used Sardinians as the Caucasoid reference population, Karitiana for Mongoloids, and San for Africans. The latter two were chosen because they live at maximally opposite corners of the Earth (South America vs. South Africa).

A first plot of the f4 statistics used for f4 regression ancestry estimation is seen below:

Clearly, some evidence of a cline is present, but several populations appear to deviate from it. In order to get the cleanest possible cline, I carried out the following greedy procedure: I calculate the correlation coefficient of this set, and iteratively remove one population that leads to the maximum improvement of the correlation, until no further improvement takes place. The following populations were removed with this procedure:

Spaniards, Serb_D, Romanian_D, N_Italian_D, Tuscan, TSI30, Bulgarians_Y, Bulgarian_D, Castilla_Y_Leon_1KG, Romanians_14, Valencia_1KG
This seems to make sense, as all these are southern European populations. Note that their removal does not mean that they do not partake in the same phenomenon as northern Europeans: they also exhibit Karitiana-shift relative to the Sardinians, but there are probably other confounding factors that make them fall "off-cline". Including them would diminish the clarity of the cline for Northern European populations. The regression of the remaining populations can be seen on the right:



f4 regression ancestry estimation results are shown on the left. These appear to be much higher than was the case with the Han and Dai in the previous experiment.

I can't say that I've made any obvious mistakes, but these admixture proportions are substantial, and call for an explanation. Whatever their true levels, I am fairly confident on at least a few points:

First, it is evident that northern Europeans have higher levels of this element than southern Europeans; the latter are not altogether deficient in it, but they fall "off-cline", making estimation of their admixture proportions more difficult.

Second, within northern Europe, there is a fairly clear east-west cline of diminishing Amerasian-like admixture. The minimum occurs in Sardinians and secondarily in Southwest Europe. Romance, Celtic, and Germanic populations all have less of it than Balto-Slavic and Uralic ones. And, some populations of northeastern Europe seem to have a noticeable excess of it.

The groups with the most Amerasian-like admixture possess Y-haplogroup N, a clear trace of eastern ancestry that is not shared by most Europeans. The arrival of this haplogroup, either with Comb Ceramic of the Baltic Neolithic or later with Seima Turbino Bronze Age expansions is probably responsible for the local excess in Northeastern Europe. The Chuvash are, of course, a Turkic population but of Finno-Ugrian genetic origin.

But, the presence of this element even in Western Europe cannot be explained on the basis of typically Mongoloid elements which are almost completely lacking there. If Mesolithic Europeans were themselves Asian-shifted, then this would account for the presence of the element, but not necessarily for its clinal manifestation. The double (north-south and east-west) cline indicates every sign of an intrusive element. So, for the time being, I will propose that this is associated with late (e.g., Copper and Bronze Age) phenomena, such as the northern stream of the Bronze Age Indo-European invasion of Europe.

This may be due to the

  • (i) northern Indo-European groups picking up some native east European or Siberian elements as they made their way into Europe, 
  • or (ii), more likely, in my opinion, that the Y-haplogroup R1 group of people, whose closest relatives are in Central/South Asia (R2) , and whose more distant relatives (Q) are in Siberia and the Americas, were from the beginning an "intermediate population" between West and East Eurasia. The R1 group of people in its R1b and R1a varieties first appear in Europe during the Copper Age, and they are lacking in early Neolithic sites.


Eight years ago, and in a totally different context, I wrote:

Similarly, 9 out of 10 Basques are descended from a man who has also fathered 9 out of 10 Kets from Siberia and 9 out of 10 Maya Indians from America. That man, founder of haplogroup P thus has descendants who belong to two of the major human races (or three, if Amerindians are considered as separate from Asian Mongoloids)   
... 
In conclusion, human continental populations form groups of genetic and phenotypic similarity, and these groups can be considered races in the phenetic sense. However, these groups are not monophyletic, hence in the cladistic sense they should not be considered as valid taxa. Since the principle of common descent is generally applied in modern systematics (or at least it should!), I think it's best not to recognize human subspecies. 

If these data pan out, it may be revealed that the European branch of the Caucasoids is actually a product of admixture too, with at least two of its constituent elements being the "Palaeo-West Eurasians" (Y-haplogroups G, IJ, possibly LT) and the "Neo-NW Eurasians" (Y-haplogroups N1 and R1), with the "Neo-Afrasians" (Y-haplogroup E1b1b) forming a third element.

(A raw dump of fourpop output can be found here).

July 28, 2012

Complex Y chromosome structure in East Tyrol (and more IE thoughts)

Cultural diversity can disappear in a few generations, but genetic diversity -barring major genocides or disasters- usually persists.

The East Tyrol region in Austria has been Germanic-speaking since the Middle Ages, but historical evidence documents the presence of Romance, Germanic, and Slavic groups in its territory. How can we untangle the origin of the different groups when they are all jumbled up together now, and all Germanic-speaking? Previous research has shown that patrilineal groups can be isolated on the basis of surnames, but in the case of East Tyrol, the wide adoption of surnames happened after the region had become linguistically Germanic.

The authors of the new paper exploited the structure of local toponyms, in particular pasture names. The figure on the left shows the concentrations of Slavic (panel A), Romance (panel B), and Germanic (panel C) pasture names. While Germanic pasture names are evenly distributed, there is a contrast between those of Slavic and Romance origin. From the paper:
From the 853 analyzed pasture names in East Tyrol 71% were derived from Germanic (Bavarian) etymons, 17% from Slavic etymons, and 12% from Romance etymons. While pasture names with Germanic etymons were evenly distributed in high density within the whole study area the names with Slavic etymons were spatially focused in the east and north of East Tyrol (Fig. 2). Geographically, these are the lower Drau, Isel, Kals, Virgen and the Defereggen valleys (Fig. 1). No names with Slavic etymons were found in the southwestern Puster valley (Fig. 2). The pasture names with Romance etymons focus mainly in the southern part of East Tyrol (Gail, Puster, and Villgraten valley, Fig. 2). The slight northeastward trend observed in the distribution of Romance etymons is solely caused by the Kals valley, a medieval Romance linguistic enclave, which was separated from the Romance main territory in the 10th century [36]. On the basis of these results, East Tyrol was divided into two regions of former Romance (Puster, Gail, and Villgraten valley; region A) and Slavic (Isel, lower Drau, Defereggen, Virgen, and Kals valley; region B) main settlement (Fig. 2).
The authors dissected the occurrence of different haplogroups in the two contrasting regions (A: Romance, and B: Slavic) in some great detail:

Splitting the East Tyrolean population sample into regions A and B resulted in a partitioning of haplogroups E-M78, R-M17, R-M412/S167*, and R-S116*. E-M78, R-M17 and R-S116* Y chromosomes were exclusively found in region B whereas samples assigned to R-M412/S167*, R-U106/S21, and R-U152/S28 reached higher frequencies in region A (Fig. 3, Table S7). When attributing the samples to the fathers' and grandfathers' places of birth/residence, as reported by the participants, practically identical patterns were obtained for most of the haplogroups (Fig. 3). 
Y chromosomes belonging to haplogroups G-P15, I-M253, and J-M304 showed much lower regionalization in their frequencies (Fig. 3) at all three generation levels.

The non-localization of the G-P15, I-M253, and J-M304 seems reasonable as these may represent what is common in these populations (and one could indeed speculate -on the basis of current ancient DNA knowledge- that they correspond to Neolithic, Paleolithic, and Bronze Age processes respectively)

Two of the most interesting findings are:
Haplogroup R-M412/S167* was found at low frequencies in the combined East Tyrolean sample. However, the R-M412/S167* chromosomes were sorted by the subdivision of the study area and reached in region A levels of ~14% whereas their frequency in region B was well below the 5% threshold. At the probands and fathers level of analysis region A featured approximately fourfold higher frequencies of these chromosomes than region B. This ratio changed to about nine when placing the samples at the grandfathers' places of birth/residence. These contrasts remained statistically significant after correcting for multiple comparisons [22] at the fathers and grandfathers analysis level.
and:
The western border of the geographic expansion of haplogroup R-M17 Y chromosomes is to be found in Central Europe and largely follows the political border separating present-day Poland (57%) and Germany (East: ~30%, South: ~14%, West: ~10%) [42]. Frequencies of about 15% and 10% were also found for Austria [18] and North-East Italy [48], respectively. In South Italy and in West Europe R-M17 chromosomes are not present at informative frequencies. 
In this study, the proportion of Y chromosomes carrying the derived M17 allele was 14.1%, a value that nearly perfectly matched those reported for West Austria (North Tyrol, 15.4%) and South Germany (Munich; 14.3%) [18], [42]. However, haplogroup R-M17 was completely absent in the East Tyrolean sub-sample from region A, but made up to 16% in region B. This result remained practically unchanged when assigning the probands to their respective fathers' or grandfathers' places of birth/residence (Fig. 3).
The new study reinforces my belief that R-M17 was not originally present in the Italo-Celtic branch of Indo-European. Together with the paucity of the same lineage in Albanians (~5%), Armenians (less than 5%), and its quite uneven distribution in Greeks, it is becoming increasingly clear that R-M17 may represent a late entrant that affected minimally southern and western Europe.

The fountain of its spread was probably a trans-Caspian (?) Central Asian staging point that followed a counter-clockwise route into Europe that spawned the northern (Germanic and Balto-Slavic) groups of Europe and the Indo-Iranians, who remained longer in their BMAC homeland, finally breaking down during the 2nd millennium BC. This would also harmonize with the increasing evidence for complementary R-M17 distributions in Europe and Asia, associated with the Z93 marker. 


It might appear that Z93+ chromosomes may track the later expansion of the Indo-Iranian world. I have observed before that R-M17 seems distributed in a long arc north and east of the Caspian, and it is perhaps in different points along this arc that the dominant European (NW) and Asian (SE) types emerged out of the early Neolithic population.

Combining this insight with an analysis of Y chromosome variation within the Graeco-Armeno-Aryan group, it appears that Graeco-Armenian is characterized predominantly by J2+R1b related lineages, while Indo-Iranian by J2+R1a related lineages. The evidence for Tocharian would involve J2+R1b related lineages.  Overall, it would appear that the earliest J2 core of PIE affected two different groups of populations living on complementary sides of the Caspian:

  • The trans-Caspian R-M17 population followed an early (3rd, or late 4th millennium BC?) north-west trajectory into Europe (associated with northern European groups such as Balto-Slavic and Germanic) as well as a later expansion (2nd millennium BC? associated with climatic deterioration in BMAC) that brought Iranian speakers to the steppe, as well as to Iran, and Indo-Aryans to South Asia.
  • The cis-Caspian, trans-Caucasian R-M269 population followed an early (late 4th millennium, early 3rd millennium?) expansion into Europe, probably together with J2 in the Balkans (Graeco-Phrygian, perhaps Thracian), and arriving in the form of Bell Beakers in Western Europe (Italo-Celtic), as well as a later (2nd-1st millennium BC?) expansion to the east (Tocharians)
This long excursus was necessary as a preamble to an explanation on what happened in Europe itself, which brings us back to the topic of the current paper:

  • The lack of structure between regions A and B with respect to haplogroup J, together with the great difference in levels of this haplogroup between Italy and the Celtic world,  suggests that Italian J-related lineages  may have been inflated in proto-historical and historical times. There are candidates a-plenty: Greeks, Etruscans, Trojans to name but three. Excess of J in Italy, relative to the Celtic world, clearly relates to the abundant traditions of eastern origins for the historical groups of Italy.
  • It would appear that during proto-history, most of Europe was dominated by three sets of IE people (R-M269 in the west, who had transmitted Proto-Celto-Italic; R-M17 in the northeast of Proto-Balto-Slavic speech, and Proto-Germanic in-between, participating in both worlds, and --appropriately-- often linked with either Italo-Celtic or Balto-Slavic linguistically)
  • There were other (now-extinct) groups as well: the Illyrians vs. Thracians in the Balkans with complementary Y chromosome distributions, the former including an extra chunk of aboriginal legacy (haplogroup I), no doubt due to the much more difficult terrain of the western Balkans. These are contrasted with our final group, the Greeks who straddled three worlds (the Paleo-Mediterranean world of the first farmers, the Thraco-Phrygian world linked to the Indo-Iranians at a deeper level, and the Anatolian world)
The boundaries between these various groups were a little blurred in the course of history. But, apparently, they were still a little clearer during the Middle Ages, and probably much clearer before the Völkerwanderung of the Germans, and the expansion of the Slavs.


Geneticists are executing a remarkable pincer movement, zeroing in on the period of European ethnogenesis from both the remote past and the present: through a study of ancient DNA from the dawn of history, they are beginning to understand how Europe was peopled, layer after layer of settlement; and through the study of surnames and toponyms they are drilling ever deeper into the pre-genealogical past. Together with much anticipated technological progress related to full genome sequencing and ancient DNA extraction, it will not be long before the history of Europe will be laid bare in remarkable detail.

PLoS ONE 7(7): e41885. doi:10.1371/journal.pone.0041885

Pasture Names with Romance and Slavic Roots Facilitate Dissection of Y Chromosome Variation in an Exclusively German-Speaking Alpine Region

Harald Niederstatter et al.

The small alpine district of East Tyrol (Austria) has an exceptional demographic history. It was contemporaneously inhabited by members of the Romance, the Slavic and the Germanic language groups for centuries. Since the Late Middle Ages, however, the population of the principally agrarian-oriented area is solely Germanic speaking. Historic facts about East Tyrol's colonization are rare, but spatial density-distribution analysis based on the etymology of place-names has facilitated accurate spatial mapping of the various language groups' former settlement regions. To test for present-day Y chromosome population substructure, molecular genetic data were compared to the information attained by the linguistic analysis of pasture names. The linguistic data were used for subdividing East Tyrol into two regions of former Romance (A) and Slavic (B) settlement. Samples from 270 East Tyrolean men were genotyped for 17 Y-chromosomal microsatellites (Y-STRs) and 27 single nucleotide polymorphisms (Y-SNPs). Analysis of the probands' surnames revealed no evidence for spatial genetic structuring. Also, spatial autocorrelation analysis did not indicate significant correlation between genetic (Y-STR haplotypes) and geographic distance. Haplogroup R-M17 chromosomes, however, were absent in region A, but constituted one of the most frequent haplogroups in region B. The R-M343 (R1b) clade showed a marked and complementary frequency distribution pattern in these two regions. To further test East Tyrol's modern Y-chromosomal landscape for geographic patterning attributable to the early history of settlement in this alpine area, principal coordinates analysis was performed. The Y-STR haplotypes from region A clearly clustered with those of Romance reference populations and the samples from region B matched best with Germanic speaking reference populations. The combined use of onomastic and molecular genetic data revealed and mapped the marked structuring of the distribution of Y chromosomes in an alpine region that has been culturally homogeneous for centuries.

Link

July 18, 2012

fastIBD over 2,257 Europeans

Razib points me towards a very interesting new paper that applies fastIBD over the large POPRES dataset of Europeans. The most interesting thing about this is that the authors develop techniques for estimating the time depth of the pattern of common ancestry across Europe, and hence are able to conclude that the Slavic expansion has played a bigger role in European history than the Germanic one.

A worthwhile improvement would be to apply a clustering algorithm like I did back in January over the fastIBD output; that way, one does not have to arbitrarily partition Europe into regions, but have the partitions jump out of the data.

A different idea to confirm the scenario presented in this paper would be to drill into different European populations. For example, in the case of the Italians, it would be worthwhile to identify whether there are particular sub-populations with likely Greek or Albanian ancestry who share an excess of IBD with modern Greeks and Albanians.

Population averages may mask such interesting patterns lurking in the data. For example, sub-clusters within populations can be identified with both fineSTRUCTURE and fastIBD, and the corresponding clusters can be assessed with supervised ADMIXTURE to detect how they differ from each other. For example, using this technique, I was able to infer 3 sub-clusters within the ethnic Greek population:

  • pop8 (mainland Greek) with ~23% North_European
  • pop11 (Greek Cypriot) with ~5% North_European
  • pop14 (Cretan, islander, mainland+Asia Minor) with ~12% North_European
  • I have also a strong hunch based on a few half Pontic Greek+half mainland Greek data points that unmixed Pontic Greeks would be related to pop22 (Northeastern Anatolia) with ~5% North_European
Based on these results and the fastIBD analysis of Ralph and Coop (the POPRES Greek sample is from northern Greece), it might appear that a hefty portion of the North_European component in Greeks may date to the medieval period, since it is relatively smaller in eastern Greeks and Cypriots and also in the South Italian/Sicilian cluster pop16 of a different analysis, with Italians as a whole lacking the eastern European affiliations of some Greek groups.

Interestingly, ~5% North_European levels would be similar to those of Armenians who are the closest linguistic cousins of the Greeks within the Indo-European family, as well as the the Anatolian Turkish cluster pop13 at ~9%.

Overall, it would appear that some mainland Greek groups received some input as the result of the medieval Slavic intrusions, since the mainland North_European excess appears as a "wedge" within the South Italy/Sicily/Crete/Anatolia/Armenia arc and the fastIBD pattern of sharing suggests that this is due to fairly recent connections.

As I have pointed out before, one limitation of the method of counting shared blocks of ancestry is that it does not disclose the directionality of gene flow. For example, gene flow between Germans and Slavs is detected in this study, which could be ascribed to Germans living in eastern Europe and/or to Slavs becoming acculturated Germans as a result of living within Germanic states or intermarrying with them prior to the age of the nation state.

Finally -and most interestingly- I hope that similar haplotype-based methods can be applied to a wider dataset, because, as it is becoming clear, Europe has not been isolated from Asia or Africa during its long history. The authors mention "Slavic or Hunnic" as an explanation for the pattern of shared ancestry in eastern Europe, but it is only by including Asian groups that we can detect the existence of real Hunnic (or Avar, or Mongol, or Pecheneg, or, ...) ancestry.

Moreover, I am confident that the Bronze Age is well within the power of haplotype-based methods to detect IBD. For example, South Asian populations clearly show differential patterns of affiliation with modern West Eurasian groups, most of which can date to no later than the Bronze Age. Together with the gradual incorporation of the new ancient DNA genomes that are bound to be coming our way soon, it seems that our picture of not only recent history, but also of late prehistory is bound to become much sharper.

arXiv:1207.3815v1 [q-bio.PE]


The geography of recent genetic ancestry across Europe

Peter Ralph, Graham Coop
(Submitted on 16 Jul 2012)

The recent genealogical history of human populations is a complex mosaic formed by individual migration, large-scale population movements, and other demographic events. Population genomics datasets can provide a window into this recent history, as rare traces of recent shared genetic ancestry are detectable due to long segments of shared genomic material. We make use of genomic data for 2,257 Europeans (the POPRES dataset) to conduct one of the first surveys of recent genealogical ancestry over the past three thousand years at a continental scale. We detected 1.9 million shared genomic segments, and used the lengths of these to infer the distribution of shared ancestors across time and geography. We find that a pair of modern Europeans living in neighboring populations share around 10-50 genetic common ancestors from the last 1500 years, and upwards of 500 genetic ancestors from the previous 1000 years. These numbers drop off exponentially with geographic distance, but since genetic ancestry is rare, individuals from opposite ends of Europe are still expected to share millions of common genealogical ancestors over the last 1000 years. There is substantial regional variation in the number of shared genetic ancestors: especially high numbers of common ancestors between many eastern populations likely date to the Slavic and/or Hunnic expansions, while much lower levels of common ancestry in the Italian and Iberian peninsulas may indicate weaker demographic effects of Germanic expansions into these areas and/or more stably structured populations. Recent shared ancestry in modern Europeans is ubiquitous, and clearly shows the impact of both small-scale migration and large historical events. Population genomic datasets have considerable power to uncover recent demographic history, and will allow a much fuller picture of the close genealogical kinship of individuals across the world.

Link

November 14, 2011

Splits or Waves? Trees or Webs?

Tree models are used in both linguistics and genetics for inferring population history. The trouble with them is that human populations do not really evolve (either genetically or culturally, as in language), tree-like, but rather exchange both genes and words.

Linguistic evolution has been mostly described in terms of tree models, but languages are not insulated from each other, and they interact after their initial differentiation. This interaction is facilitated by geographic proximity, and also by linguistic proximity.

Geographic proximity makes it possible for speakers of different languages to talk to each other, learn each other's languages, or develop hybrid languages or a lingua franca. Linguistic proximity facilitates communication: it is fairly easy, for example, for speakers of Germanic languages to interact, and much more difficult for those of, say, English and Chinese.

If speakers of a language become separated by distance or geographical barriers, then lateral exchange between different groups becomes minimal, and language evolution can be well-described by a tree model. If, on the other hand, there exists a language continuum across a wide area, effected by a common process (say, the spread of agriculture), then there is room for substantial cross-interaction of different emergent languages at the stage when they can be still thought as dialects of the parent language.

While the current paper's focus is on Germanic languages, the endgame seems to be on the much harder and more vigorously contested field of Indo-European studies.

The author has put up a nice supplementary page online on a first attempt of using NeighborNet with an Indo-European dataset, pictured on the right. A publication on the topic is listed as being in preparation:
The utility of Germanic as a case-study is that it provides a (reasonably) known external history against which to assess our methodological approaches. On the strength of the findings here, a similar logic can now be extended to probing the unknown of how the early divergence history of Indo-European unfolded. In the full exploration in Heggarty (in preparation a), it transpires that even the data underlying figures 1 and 2 here suggest an early divergence along the lines of a dialect continuum. And for all the purported analytical elegance of binary branches, as a real-world demographic scenario it is this Indo-European continuum that offers the more straightforward and economical explanation. A splits-then-borrowing scenario has instead to invoke not just a complex series of divergent migrations, but then later movements to attenuate this by bringing certain groups back into contact again. This in turn entails consequences for which of the main rival hypotheses—the migratory Kurgan ‘horse culture’, or the progressive demic diffusion of agriculture—best fits as the driving force that shaped the pattern of the earliest Indo-European expansion.
Phil. Trans. R. Soc. B 12 December 2010 vol. 365 no. 1559 3829-3843

Splits or waves? Trees or webs? How divergence measures and network analysis can unravel language histories

Paul Heggarty et al.

Linguists have traditionally represented patterns of divergence within a language family in terms of either a ‘splits’ model, corresponding to a branching family tree structure, or the wave model, resulting in a (dialect) continuum. Recent phylogenetic analyses, however, have tended to assume the former as a viable idealization also for the latter. But the contrast matters, for it typically reflects different processes in the real world: speaker populations either separated by migrations, or expanding over continuous territory. Since history often leaves a complex of both patterns within the same language family, ideally we need a single model to capture both, and tease apart the respective contributions of each. The ‘network’ type of phylogenetic method offers this, so we review recent applications to language data. Most have used lexical data, encoded as binary or multi-state characters. We look instead at continuous distance measures of divergence in phonetics. Our output networks combine branch- and continuum-like signals in ways that correspond well to known histories (illustrated for Germanic, and particularly English). We thus challenge the traditional insistence on shared innovations, setting out a new, principled explanation for why complex language histories can emerge correctly from distance measures, despite shared retentions and parallel innovations.

Link

August 05, 2011

Genetic structure of Swedish population


PLoS ONE 6(8): e22547. doi:10.1371/journal.pone.0022547

The Genetic Structure of the Swedish Population

Keith Humphreys et al.

Patterns of genetic diversity have previously been shown to mirror geography on a global scale and within continents and individual countries. Using genome-wide SNP data on 5174 Swedes with extensive geographical coverage, we analyzed the genetic structure of the Swedish population. We observed strong differences between the far northern counties and the remaining counties. The population of Dalarna county, in north middle Sweden, which borders southern Norway, also appears to differ markedly from other counties, possibly due to this county having more individuals with remote Finnish or Norwegian ancestry than other counties. An analysis of genetic differentiation (based on pairwise Fst) indicated that the population of Sweden's southernmost counties are genetically closer to the HapMap CEU samples of Northern European ancestry than to the populations of Sweden's northernmost counties. In a comparison of extended homozygous segments, we detected a clear divide between southern and northern Sweden with small differences between the southern counties and considerably more segments in northern Sweden. Both the increased degree of homozygosity in the north and the large genetic differences between the south and the north may have arisen due to a small population in the north and the vast geographical distances between towns and villages in the north, in contrast to the more densely settled southern parts of Sweden. Our findings have implications for future genome-wide association studies (GWAS) with respect to the matching of cases and controls and the need for within-county matching. We have shown that genetic differences within a single country may be substantial, even when viewed on a European scale. Thus, population stratification needs to be accounted for, even within a country like Sweden, which is often perceived to be relatively homogenous and a favourable resource for genetic mapping, otherwise inferences based on genetic data may lead to false conclusions.

Link

March 05, 2011

Celto-Germans vs. Balto-Slavs

Here are the first two dimensions of a multidimensional scaling plot of the following samples:
  • Dodecad Ancestry Project: 6 Poles, 12 Russians (Russian_D), 11 Germans, 19 Scandinavians, 6 Mixed Slavs (various West and East Slav combinations), 17 Britons, 17 Irish
  • HGDP: 25 Russians from Vologda
  • Behar et al. (2010): 10 Lithuanians, 9 Belorussians
Applying MCLUST over these first two dimensions and with K=2, the following breakup of individuals ensues:

It's fascinating that Cluster #1 (which corresponds to the assortment of individuals on the left of the MDS plot) includes only Balto-Slavic individuals (65 in total), while Cluster #2 (on the right, includes all 64 Celto-Germanic individuals plus 2 Poles and a mixed Slav.

This surprising concordance is even more striking once we consider that one of the "mixed Slavs" in my sample may be of Prussian origin within present-day Poland. I will be happy to tell the Poles in my sample which cluster they belong to if they write to me at the Dodecad Project e-mail address.

A lot has transpired since the ancient ethnographers divided the little-known peoples of the far north into Keltoi and Skythai, or since the Franco-Russian anthropologist Deniker divided the light-pigmented Northern Europeans into a race nordique and a race orientale. So, it is a bit surprising to see that a basic division of northern Europeans into East and West has stood the test of time. (*)

(*) Minus the Finnic peoples of northeastern Europe who, as has become clear, owe their genetic distinctiveness to a Siberian element in their ancestry, tying them to their linguistic cousins in the east.

February 22, 2011

Medieval DNA from Usedom, Germany

With respect to the Slavic/Germanic origin of the studied samples, I would like to point to a 2005 study on the differentiation between Germans and Poles. R-M458 should probably be assigned to the Slavic side, while E1b1b on the German. The absence of R1b (in the albeit limited sample) is interesting, and should be interpreted as further evidence for the Slavic side of the argument, as R1b strongly differentiates Germans from Slavs in today's populations and less than a millennium ago is probably too short a timespan to expect dramatic changes in haplogroup frequencies.

Related:

Main title Die mittelalterlichen Skelette von Usedom
Subtitle Anthropologische Bearbeitung unter besonderer Berücksichtigung des ethnischen Hintergrundes
Title variations The mediaeval skeletons of Usedom
Subtitle for translated title Anthropological investigation in due consideration of the ethnical background
Author(s) Freder, Janine
Place of birth: Berlin
1. Referee Prof. Dr. Carsten Niemitz
Further Referee(s) Prof. Dr. Joachim Burger
Keywords Anthropology; osteometry; palaeodemography; Slavs; Danes; DNA; mitochondrial; Y chromosome
Classification (DDC) 570 Life sciences
Summary This study investigates 200 skeletons from an early Christian graveyard of the 12th to early 13th century in Usedom (Mecklenburg-Vorpommern, Germany). The city of Usedom was a notable maritime place of trade in a time of major political and social transformations. The Christianisation of the Slavic elite in 1128, the following raids of the Danes and the influx of German settlers starting in the 13th century were formative events.
The reconstruction of the living conditions of the Usedom population was achieved by means of well established anthropological and palaeodemographical methods. Age and sex distribution comply with other ordinary populations of that time frame: high proportion of children (32 %), comparatively few adolescents but many adults (59 %) as well as a slight surplus in men. Remarkably, a deficit in women in the mature age class is attended by an increased mortality of girls of the age class infans I. However, this may be due to a methodical error.
In order to clarify a possible Slavic, Danish or German background of the inhabitants of Usedom, eight skull measures, four skull indices and five measures of the long bones of the extremities were investigated typologically as well as statistically on the basis of their arithmetic means and compared to the measures of two series of Slavic or multiethnic/place of trade background (Sanzkow and Haithabu, respectively). The comparison of arithmetic means did yield statistically significant differences between the three populations. The men and women of Usedom seem to be more closely related to the Sanzkow population. However, they appear to take a position between the two other populations. Unfortunately, a comparison with Slavic and Germanic populations of the Neolithic till Early Middle Ages did not provide distinct results. The archaeologically based assumption of a mainly Slavic population cannot be rejected with anthropological means.
The analysis of mitochondrial and Y-chromosomal DNA, however, generated auspicious results despite adverse storage conditions. Results could be obtained from all four samples. Two individuals were of mtDNA haplogroup H and two of haplogroup K. Y-chromosome analysis yielded haplogroups E1b1b and R1a1a7, respectively, in two males. Future molecular research will see improved methods for the even more detailed reconstruction of human migration.

January 07, 2011

Of Cattle and Men (Edwards et al. 2010)

From the paper:
Apparently, the expansion of the dairy breeds have created, or largely maintained, a sharp genetic contrast of northern and southern Europe, which divides both France and Germany. It may be hypothesised that the northern landscapes, with large flat meadows, are suitable for large-scale farming with specialised dairy cattle (Niederungsvieh, lowland cattle), whilst the mixed-purpose or beef cattle (Höhenvieh, highland cattle) are better suited to the smaller farms and hilly regions of the south. However, it is also remarkable that in both France and Germany the bovine genetic boundary coincides with historic linguistic and cultural boundaries. In France, the Frankish invasion in the north created the difference between the northern langue d'oïl and the southern langue d'oc. The German language is still divided into the southern Hochdeutsch and northern Niederdeutsch dialects, which also correlates with the distribution of the Catholic and Protestant religions. On a larger scale, it is tempting to speculate that the difference between two types of European cattle reflects, and has even reinforced, the traditional and still visible contrast of Roman and Germanic Europe.
UPDATE: I wish there'd be some data points for the vast area between Eastern Europe and Yakutia. There might be a simple (and recent) expalanation for why Northeastern Europe is mostly "green" and Yakutia "red", but it would be nice to have actual datapoints in the quadrilater between NE Europe ("green"), SW Asia (mostly "red"), S Asia (zebu "black") and Yakutia.

PLoS ONE 6(1): e15922. doi:10.1371/journal.pone.0015922

Dual Origins of Dairy Cattle Farming – Evidence from a Comprehensive Survey of European Y-Chromosomal Variation

Ceiridwen J. Edwards et al.

Abstract
Background
Diversity patterns of livestock species are informative to the history of agriculture and indicate uniqueness of breeds as relevant for conservation. So far, most studies on cattle have focused on mitochondrial and autosomal DNA variation. Previous studies of Y-chromosomal variation, with limited breed panels, identified two Bos taurus (taurine) haplogroups (Y1 and Y2; both composed of several haplotypes) and one Bos indicus (indicine/zebu) haplogroup (Y3), as well as a strong phylogeographic structuring of paternal lineages.

Methodology and Principal Findings
Haplogroup data were collected for 2087 animals from 138 breeds. For 111 breeds, these were resolved further by genotyping microsatellites INRA189 (10 alleles) and BM861 (2 alleles). European cattle carry exclusively taurine haplotypes, with the zebu Y-chromosomes having appreciable frequencies in Southwest Asian populations. Y1 is predominant in northern and north-western Europe, but is also observed in several Iberian breeds, as well as in Southwest Asia. A single Y1 haplotype is predominant in north-central Europe and a single Y2 haplotype in central Europe. In contrast, we found both Y1 and Y2 haplotypes in Britain, the Nordic region and Russia, with the highest Y-chromosomal diversity seen in the Iberian Peninsula.

Conclusions
We propose that the homogeneous Y1 and Y2 regions reflect founder effects associated with the development and expansion of two groups of dairy cattle, the pied or red breeds from the North Sea and Baltic coasts and the spotted, yellow or brown breeds from Switzerland, respectively. The present Y1-Y2 contrast in central Europe coincides with historic, linguistic, religious and cultural boundaries.

Link

August 25, 2010

R1b founder effect in Central and Western Europe

Post will be updated after I read the paper. (Last Update: Aug. 29)

UPDATE I:

From the paper:
The ages of various haplogroups in populations were estimated using the
methodology described by Zhivotovsky et al,30 modified according to Sengupta
et al,10 using the evolutionary effective mutation rate of 6.9 x 10^-4 per 25 years.
The accuracy and appropriateness of this mutation rate has been independently
confirmed in several deep-rooted pedigrees of the Hutterites.
Of course readers of the blog are aware that I disagree with the use of the evolutionary rate. My comments on the Hutterites paper will be posted separately after I read it. I will simply say that there are numerous cases where the use of the genealogical rate makes better sense of the evidence than use of the "evolutionary" rate. Off the top of my head, the genealogical rate harmonizes with the Genghis Khan cluster, the expansion of Na-Dene speakers into the Americas, the expansion of Balto-Slavic, the Bronze Age spread of Semitic speakers, in accordance with the linguistic evidence, the expansion of Bantu in Angola, more recent British surnames, the formation of Arabian kingdoms, Greek colonization of Sicily, and the Bronze Age origin of Indo-Aryans and Finno-Ugrians (and I skipped a few).

UPDATE II (Aug 26):

Here is the phylogeny of R-M207 from the paper. For reference, the R-M207 page from ISOGG.


UPDATE III (Aug 26):

Going through the material in this paper in a systematic manner is not easy, so I will probably do a potpourri of updates covering various topics of interest.



As noted in the other recent paper, and shown in the above Figure from the current one, R-U106 peaks in northern Europe. Its frequency (including the R-U198 sublineage) is 36.8% in the Netherlands, 20.9% in Germany and Austria, 18.2% in Denmark, 18.2% in England, 12.6% in Switzerland, 7.5% in France, 6.1% in Ireland, 5.9% in Poland, 5.6% in north Italy 4.4% in Czech Republic and Slovakia, 3.5% in Hungary, 4.8% in Estonia, 4.3% in south Sweden, 2.5% in Spain and Portugal, 1.3% in eastern Slavs, 0.8% in south Italy, 0.6% in Balkan Slavs, 0.5% in Greeks (i.e. 2 of 193 Cretans, and no mainland Greeks), 0.4% in Turks, 0% in Middle East.

The age of R-U106 is estimated by the authors as 8.7ky BP, which translates to about 2.5ky BP with the germline rate. The existence of R-U106 as a major lineage within the Germanic group is self-evident, as Germanic populations have a higher frequency against all their neighbors (Romance, Irish, Slavs, Finns). Indeed, highest frequencies are attained in the Germanic countries, followed by countries where Germanic speakers are known to have settled in large numbers but to have ultimately been absorbed or fled (such as Ireland, north Italy, and the lands of the Austro-Hungarian empire). South Italy, the Balkans, and West Asia are areas of the world where no Germanic settlement of any importance is attested, and correspondingly R-U106 shrinks to near-zero.

UPDATE IV (Aug 26):

Another informative lineage, as noted in the other recent paper as well is R-U152:


Of interest is the fact that while
R-U152 has a clear French-Italian center of weight, the locations exhibiting highest STR variance are Germany and Slovakia, i.e., Central Europe. My guess is that R-U152 originated in Central Europe spreading to the west and south, perhaps with Italo-Celtic speakers or some subset thereof. In its home territory of Central Europe, its frequency decreased by the introduction of the Germanic and Slavic speaking elements which dominate the region.

Irrespective of what the ultimate origin of R-U152 is, it provides us with a good diagnostic marker for population movements out of the French-Italian area. In Italy for example it is noted at 26.6% for the north and 10.5% in the south. It would be extremely interesting to see its occurrence in Balkan Vlachs, as this would confirm/disprove the Italian component in their origin. However, R-U152 occurs in 7.3% of Cretans, suggesting introgression Y-chromosomes of North Italian (Venetian) origin, from the 4-century period of Venetian rule of the island. It also occurs in 4.1% of Greeks, where it might come from any period since the Roman annexation of the Hellenistic states to the Vlachs. However, its presence at only 1.8% of Romanians makes a large Italian contribution to the Romanian population unlikely. Balkan R-U152 chromosomes should be better resolved to determine when they arrived from the northwest.

The paucity of R-U152 in Turks (0.6%) make tales of wandering Galatians less likely to be true. There is no doubt that Galatians settled in Anatolia, but they were probably so few in numbers that they did not permanently alter the population. Knowledgeable readers should chime in about the Lebanese Christian R1b which was posited as a signature of the Crusades a couple of years ago, and its position in the phylogeny.

UPDATE V (Aug 26):

The most commong R1b subgroup in Europe is R-M269 and the most common subgroup is R-L23 which encompasses the vast majority of European R-M269 chromosomes. It is interesting to see where R-M269(xL23) is concentrated. In Europe I see cases in Germany, Switzerland, Slovenia, Poland, Hungary, Russia, the Ukraine. It is most prominent, however, in the Balkans, where every population except Croatia mainland (N=108) possesses it. In the Caucasus it does not exist except in the northeast. In Turkey and Iran there is some, albeit it is not clear in which regions.

UPDATE VI (Aug 27):

The authors write with respect to haplogroup R-V88:
With the exception of rareincidences of R1b-V88 in Corsica, Sardinia13 and Southern France(Supplementary Table S4), there is nearly mutually exclusive patterning of V88 across trans-Saharan Africa vs the prominence of P297-related varieties widespread across the Caucasus, Circum-Uralic regions, Anatolia and Europe. The detection of V88 in Iran, Palestine and especially the Dead Sea, Jordan (Supplementary Table S4) provides an insight into the back to Africa migration route.
Haplogroup R-V88 has been the subject of a recent study and was associated with the migration of Chadic speakers in Africa. It is difficult to say whether or not the authors' results really provide any insight into an alleged movement of this haplogroup from Asia to Africa, as it occurs in only a single Palestinian, and a single Iranian. Neither is the higher frequency (13.7%) observed in the Amman and Dead Sea area of Jordan really evidence of its antiquity there.

Neither the aforementioned paper nor the current one presents any evidence (e.g., Y-STR variance) for any great antiquity of the Asian R-V88 with respect to the African one. Indeed, with the exception of the aforementioned Jordanian sample, R-V88 is rare in Asia, while it is widespread in African Berbers. I see no clear reason at present to think that it migrated to Africa from Asia, and not to think of it as a relic of an older, widely dispersed R1b population leading to R-V88 in Africa itself.

UPDATE VII (Aug 28):

The paper repeats the standard claims about the origin of R1b and its main sublineage R-M269 in Asia, but presents no new information that would support this claim. With the state of the evidence, I see no real reason to prefer a West Asian to a Southeastern European origin for this haplogroup.

I don't give much credence to small differences in Y-STR variance, due to the large confidence intervals associated with such estimates, and it is interesting that the authors do not present an argument from Y-STR variation about the origin of R1b, preferring to make broad statements about Mesolithic-Neolithic movements into Europe.

A study of supplementary table S2 which gives coalescent times reveals that there is no clear pattern of greater Asian diversity within haplogroup R1b or its subclades. And, while Central-Western Europe does appear to be an outgrowth of R1b rather than a place of origin (with the dominance of derived R-M412 lineages) there is nothing in the paper that would make one prefer West Asia to Southeastern Europe as a place of origin.

Personally I think the issue cannot be settled yet, but there are reasons to prefer the latter option. An Asian origin of R1b has a major parsimony hurdle: it would require a seemingly directed drang nach westen for R1b, into Europe, and into North Africa, with a paucity of R1b in the opposite direction (among Arabians and to the south and in South Asia) and a scattering of very young R-M73 and R-M269 to the east of Europe.

UPDATE VIII (Aug 29):



R-S116 shows maximum Y-STR diversity in France and Germany but maximum frequency in Iberia and the British Isles. In the latter region it is represented mainly by R-M529 with the R-M222 subclade being particularly prominent in Ireland but also North England. It would be interesting to see data for Scotland, and I do not doubt that R-M222 would be prominent there as well. R-S116 also shows signs of being a Celtic, or Celtiberian-related lineage.

European Journal of Human Genetics doi: 10.1038/ejhg.2010.146

A major Y-chromosome haplogroup R1b Holocene era founder effect in Central and Western Europe

Natalie M Myres et al.

The phylogenetic relationships of numerous branches within the core Y-chromosome haplogroup R-M207 support a West Asian origin of haplogroup R1b, its initial differentiation there followed by a rapid spread of one of its sub-clades carrying the M269 mutation to Europe. Here, we present phylogeographically resolved data for 2043 M269-derived Y-chromosomes from 118 West Asian and European populations assessed for the M412 SNP that largely separates the majority of Central and West European R1b lineages from those observed in Eastern Europe, the Circum-Uralic region, the Near East, the Caucasus and Pakistan. Within the M412 dichotomy, the major S116 sub-clade shows a frequency peak in the upper Danube basin and Paris area with declining frequency toward Italy, Iberia, Southern France and British Isles. Although this frequency pattern closely approximates the spread of the Linearbandkeramik (LBK), Neolithic culture, an advent leading to a number of pre-historic cultural developments during the past ≤10 thousand years, more complex pre-Neolithic scenarios remain possible for the L23(xM412) components in Southeast Europe and elsewhere.

Link

September 20, 2009

History of the people of the Hungarian plain in the 1st millennium

Hum Biol. 2008 Dec;80(6):655-67

History of the peoples of the Great Hungarian Plain in the first millennium: a craniometric point of view

Holló G, Szathmáry L, Marcsik A, Barta Z.

We carried out an examination relying on six dimensions of 1,573 crania coming from the Great Hungarian Plain. The crania represent seven archeological periods: Sarmatian age (1-4th century), the period of transition (about 400-420), Hun and Gepidic epochs (about 420-455 and 455-567, respectively), early Avar age (about 568-670), late Avar period (about 670-895), the epoch of the Hungarian conquest and settlement (about 895-1000), and the Arpadian age (about 1000-1301). We were curious about the anatomical background behind cultural changes of the various populations that inhabited this area. After having noticed some discontinuities between the populations, as revealed by univariate analysis of single dimensions, we performed a principal-components analysis to see whether or not the diverse components showed eventual breaks in the sequence of the populations. Knowing that all the dominant populations had Asian roots, except for the Gepids of Germanic origin, we expected a considerable difference between the Gepidic population and all the other inhabitants. We also assumed that a conquest itself with a large-scale assimilation was unlikely to leave breaklike traits in anatomical patterns, except for aggressive conquests. We found that the second principal component (which correlated with cranial breadth and partly with height) showed a remarkable hiatus in both sexes between Gepids and early Avars. Having done a statistical proof (simultaneous tests for general linear hypotheses) of the observed phenomenon, we found that the gap referring to subsequent populations was significant only in males. A possible reason for this result is that the Avar conquest was much more radical than has been thought. In addition, considering that men were more likely to die in wars, women survived and were assimilated into the conquerors' populations with higher probability, so it is not surprising that the results of multicomparison tests are significant only in men.

Link

May 05, 2009

Supplement on "Geographical structure and differential natural selection amongst North European populations" (McEvoy et al. 2009)

From the supplemental material of a paper I covered in March, here are a couple of PCA plots of the first two principal components of the studied populations, with or without the Finns.


In the plot without the Finns, we see the expected British Isles -> Continental Europe differentiation in the order of Ireland, UK, Netherlands, along PC1. Swedes, and to a much lesser extent Danes deviate from this gradient in an orthogonal direction.
When Finns are included, PC1 now captures the major difference between them and the other Celto-Germanic populations which appear strikingly homogeneous along this component. The reason for the Swedes' divergence is now clear, as they are seemingly drawn towards the Finns, although the two clusters can be cleanly separated by a line at around PC1=-0.03.

It is fairly clear by now, that in northern Europe, there are two major distinctions (in that order): (i) between the Finns, and Finno-Ugrian influenced populations on the one hand, and the rest, and (b) a less important West-East gradient from Ireland to the Baltic.

The fact that factor (i) is the most important one pretty much vindicates the views of traditional physical anthropology since the time of Deniker at least. Despite the lack of data and statistical knowledge available at his time, Deniker, in the late 19th century, divided the light-pigmented northern European xanthochrooi of earlier classifications into two: the race nordique, associated primarily with the Germanic peoples, and the race orientale associated primarily with the eastern Slavs and Finns.

This classification scheme was continued by the better writers that followed, e.g., as razza nordica and razza baltica by Renato Biasutti, and as Атланто-балтийская раса (Atlanto-Baltic race) and Беломорско-балтийская раса (White Sea-Baltic race) in works written in Russian.

March 07, 2009

Genetic structure in northern Europe revisited

The results of the STRUCTURE analysis are quite interesting. When Finland is included (B), it is the first one to be separated from other Northern Europeans, confirming previous results. Sweden, and to a lesser degree Denmark seems to possess some admixture with the Finnish-related (red) element. The next split (blue vs. green) differentiates continental Germnics (esp. Scandinavians, and somewhat less Dutch) from Irish-British and Australians of largely British ancestry.

When Finland is not included (C) and for K=4, it becomes obvious that the UK population includes three components ("Dutch" green, "Scandinavian" red, and "British Isles" blue).
See also a previous study on the topic.

Genome Research doi:10.1101/gr.083394.108

Geographical structure and differential natural selection amongst North European populations

Brian P McEvoy et al.

Abstract

Population structure can provide novel insight into the human past and recognizing and correcting for such stratification is a practical concern in gene mapping by many association methodologies. We investigate these patterns, primarily through principal component (PC) analysis of whole genome SNP polymorphism, in 2099 individuals from populations of Northern European origin (Ireland, UK, Netherlands, Denmark, Sweden, Finland, Australia and HapMap European-American). The major trends (PC1 and PC2) demonstrate an ability to detect geographic substructure, even over a small area like the British Isles, and this information can then be applied to finely dissect the ancestry of the European-Australian and -American samples. They simultaneously point to the importance of considering population stratification in what might be considered a small homogenous region. There is evidence from FST based analysis of genic and non-genic SNPs that differential positive selection has operated across these populations despite their short divergence time and relatively similar geographic and environmental range. The pressure appears to have been focused on genes involved in immunity, perhaps reflecting response to infectious disease epidemic. Such an event may explain a striking selective sweep centered on the rs2508049-G allele, close to HLA-G gene on chromosome 6. Evidence of the sweep extends over 8Mb/3.5cM region. Overall the results illustrate the power of dense genotype and sample data to explore regional population variation, the events that have crafted it and their implications in both explaining disease prevalence and mapping these genes by association

Link

July 25, 2008

German origin of Transylvanian Saxons

Using Athey's haplogroup predictor, with equal priors and a threshold of 50 and probability of 90%, the following haplogroups were predicted in the 59 males:

5 E1b1b
1 G1
2 G2a
2 H
4 I1
3 I2a(xI2a2)
1 I2a2
1 I2b1
1 J2b
1 N
2 R1a
22 R1b

Rom J Leg Med 12 (4) 247 – 255 (2004)

A study on Y-STR haplotypes in the Saxon population from Transylvania (Siebenbürger Sachsen): is there an evidence for a German origin?

Ligia Barbarii et al.

ABSTRACT: A study on Y-STR haplotypes in the Saxon population from Transylvania
(Siebenbürger Sachsen): is there an evidence for a German origin? Y chromosome markers are increasingly used to investigate human population histories, being considered to be sensitive systems for detecting the population movements. In this study we present Y-STR data for a male population of Transylvanian Saxons in
comparison with Y-haplotypes from Romanians and other European populations. The Transylvanian Saxons, called like that since medieval times, are representing a western population with unknown origin, settled in the Arch of Romanian Carpathian Mountains in the earliest of the 12th century. Historical and dialectal studies strongly suggest that they do not originate from Saxony, but more probably from the Mosel riversides (Rhine affluent) and also from the Eifel Mountains Valley (present territory of Luxembourg). Living protected by fortified cities in compact communities, they still represent a quite distinct population in Transylvania. For this study, 59 male samples were collected from the Siebenburgen area, subjects being selected by their Saxon surnames and paternal grandfather birthplace. A set of nine STR polymorphic systems mapping on the male-specific region of the human Y chromosome (DYS19, DYS385, DYS389 I/II, DYS390, DYS391, DYS392, DYS393) were typed by means of
one or two two multiplex PCR reactions and capillary electrophoresis. The typing results reflect high Saxon population haplotype diversity. Furthermore, we present data on the haplotype sharing of the Saxon population with other European populations, especially with Germans as well as with the Romanians and the Transylvanian Szekely.

Link (pdf)

April 23, 2008

Criticism of Anglo-Saxon apartheid theory

See my entry on How the Anglo-Saxons outbred the British for details of the Anglo-Saxon apartheid theory. The New Scientist has details on the new study. Germanic invaders may not have ruled by apartheid:
Pattison reviewed existing archaeological and genetic evidence, and conducted a new analysis of British DNA. Then, starting in 2001 and working backwards to pre-Roman times, Pattison calculated for each generation the net population growth and the origins of immigrants.

He concludes that people with Germanic origins came to Britain well before and after the early Anglo-Saxon period, and this long period of immigration can explain a relatively strong Germanic genetic signal today.

He adds that about 60% of the current British population still has some native Briton DNA, arguing against the idea, put forward by Mark Thomas at University College London and colleagues that Saxon invaders ethnically purged the country.

The textual and archaeological evidence collected by Thomas's team is also controversial, says Pattison.
The home page of John Pattison.

Pattison, J.E., Is it Necessary to Assume an Apartheid-like Social Structure in Early Anglo-Saxon England; in press, Proceedings of the Royal Society B, April 2008.

UPDATE (Apr 23): The paper is now online. I like this kind of paper which tries to capture some of the complexity of human movements. Quite often in genetics one sees simplistic explanations of population origins. A prime example of this is the Paleolithic/Neolithic theory of European origins, that has been done to death, as if Europe and Asia weren't connected for tens of thousands of years before the emergence of agriculture and ten thousand years after it, allowing the movements of people both ways; all this complexity is shoved under the rug and origins are sought in a simple admixture event at the onset of the Neolithic.

In some cases, the nature of the migration makes a "repeat performance" unlikely, as in e.g., the arrival of the ancestors of Native Americans to the New World at a time when there was a land passage to it.

Quite often, the more dramatic events in a land's migration history tend to overshadow the more quiet and long-standing ones. The sudden arrival of a people in a short period of time, be them Anglo-Saxons in Great Britain, or European colonists in the New World, leaves a lasting impression, and is likely to be remembered well into the future, whereas the more limited and occasional movements of people from the same source areas but over longer periods of time do no attract attention: e.g., one tends to remember the Mayflower, the Puritans, etc. in the history of the colonization of North America, but the flow of British immigrants did not really cease for any substantial time ever since.

UPDATE 2:

Pattison's paper makes two unrelated arguments against the thesis of Thomas et al. (2006). The first one is that the disadvantages suffered by indigenous Britons were a kind of incentive for them to adopt a Germanic identity; the second, that pre- and post- Anglo-Saxon migration can account for the Germanic Y-DNA in England, so the effects of the social situation during the time of the Anglo-Saxon migration need not have been so dramatic or even at all present.

The first claim:
A similar strategy was employed by the Moorish Caliphate in Medieval Spain: Jews and Christians were subject to a special tax—the jizya, which Muslims did not pay—in an endeavour to encourage non-Muslims to convert to Islam. According to the Qur’an (1990), non-Muslims who refused to pay the tax, were required to either convert to Islam or face the death penalty. The ethnicities of the people involved were of no concern.
This does establish the essential difference between the regime of apartheid and the proposed situation in England at the time. In a South-African-style apartheid regime it was impossible for a member of the disadvantaged group (blacks) to become part of the advantaged one. In the Muslim case, it was possible for non-Muslims to become Muslims. Thus, in both cases there was discrimination, but in one case there were rigid boundaries between groups, while in the other there was not - at least not in the direction of Christian->Muslim conversion.

It should be noted, however, that while wherever Muslims and Christians co-existed, there was a steady discrimination against Christians, resulting in conversions, emigration, or massacres, all of which would have diminished the Christian element, at the same time, the Christians were a source of economic revenue for the regime, and the policy of the Muslim political authorities, e.g., the Sultan in the Ottoman Empire was not so much to eradicate the non-Muslim population, but rather to maintain it with enough freedom for it to be productive, but also with enough fear to be subservient to the dominant group.

The second issue is the real substance of the paper: was the social situation (whether one calls it apartheid or not) in early Anglo-Saxon England the cause of the significant Germanic element in modern Englishmen or not? That is an empirical question relying on the prevalence and arrival of Germanic elements in pre-Anglo-Saxon England and their subsequent arrival over the centuries. The paper does succeed in weakening the case for a substantial contribution of Anglo-Saxon/Briton social tensions in favor of the former's genetic success.

Is it necessary to assume an apartheid-like social structure in Early Anglo-Saxon England?

John E. Pattison

Abstract

It has recently been argued that there was an apartheid-like social structure operating in Early Anglo-Saxon England. This was proposed in order to explain the relatively high degree of similarity between Germanic-speaking areas of northwest Europe and England. Opinions vary as to whether there was a substantial Germanic invasion or only a relatively small number arrived in Britain during this period. Contrary to the assumption of limited intermarriage made in the apartheid simulation, there is evidence that significant mixing of the British and Germanic peoples occurred, and that the early law codes, such as that of King Ine of Wessex, could have deliberately encouraged such mixing. More importantly, the simulation did not take into account any northwest European immigration that arrived both before and after the Early Anglo-Saxon period. In view of the uncertainty of the places of origin of the various Germanic peoples, and their numbers and dates of arrival, the present study adopts an alternative approach to estimate the percentage of indigenous Britons in the current British population. It was found unnecessary to introduce any special social structure among the diverse Anglo-Saxon people in order to account for the estimates of northwest European intrusion into the British population.

Link