April 05, 2012

Meltwater Pulse 1A (MWP-1A)

Nature 483, 559–564 (29 March 2012) doi:10.1038/nature10902

Ice-sheet collapse and sea-level rise at the Bølling warming 14,600 years ago

Pierre Deschamps et al.

Past sea-level records provide invaluable information about the response of ice sheets to climate forcing. Some such records suggest that the last deglaciation was punctuated by a dramatic period of sea-level rise, of about 20 metres, in less than 500 years. Controversy about the amplitude and timing of this meltwater pulse (MWP-1A) has, however, led to uncertainty about the source of the melt water and its temporal and causal relationships with the abrupt climate changes of the deglaciation. Here we show that MWP-1A started no earlier than 14,650 years ago and ended before 14,310 years ago, making it coeval with the Bølling warming. Our results, based on corals drilled offshore from Tahiti during Integrated Ocean Drilling Project Expedition 310, reveal that the increase in sea level at Tahiti was between 12 and 22 metres, with a most probable value between 14 and 18 metres, establishing a significant meltwater contribution from the Southern Hemisphere. This implies that the rate of eustatic sea-level rise exceeded 40 millimetres per year during MWP-1A.

Link

Climate change and modern human dispersals

Stewart and Stringer have a paper in Science in which they argue for the role that climate change has played in the evolution of modern humans, and their expansion into Eurasia.


I have only a couple of observations with respect to their reconstruction of the migrations of modern humans. They support the idea of a fresh population movement Out-of-Africa c. 60,000 years ago, for which there really is no archaeological (or anthropological) evidence. The Origin of our Species is cited in support of this idea, so I am not sure exactly what it is based on, as I have not read that book.

There are a couple of new developments that change our understanding of how Homo sapiens arrived and dispersed in Eurasia; first, the discovery of the Nubian Complex in southern Arabia >100 thousand years ago. This establishes the presence of modern humans in a very large region of Eurasia (from the Levant to southern Arabia) >100 thousand years ago, and hence makes the idea that Skhul/Qafzeh represents an "Out-of-Africa that failed" less believable. It is more likely that the major human expansion during MIS4 and MIS3 originated in Arabia.


Second, in accordance with the Qafzeh/Skhul "OoA that failed" hypothesis, it is argued that Neandertals "reclaimed" the Levant during MIS4. But, the Kebara Neandertal sample from ~60 thousand years ago is more modern than its ~120 thousand year old Tabun predecessor. It seems to me that anatomically modern humans expanded Out-of-Africa before 100,000 years ago, but did not yet have the "tech" or behavioral adaptations to outcompete the Neandertals. It was in the Arabian refugium that the transition took place, followed by the gradual expansion of modern humans post-70ky which accelerated during the Upper Paleolithic/MIS 3.

This also explains why archaic humans persisted in West and Central Africa down to the Holocene boundary. I often note how strange it seems that this would be the case if modern behavior (as opposed to modern anatomy) originated in Africa (or even Southern Africa): why would it take much longer for humans to replace archaic Africans whereas it took them only a geological blink of an eye for them to do the same for archaic Eurasians? This puzzle is solved once we realize that an early pre-100ky expansion of AMH Out-of-Africa was followed by a late post-70ky expansion of fully modern humans Out-of-Arabia.

Science 16 March 2012:

Vol. 335 no. 6074 pp. 1317-1321

Human Evolution Out of Africa: The Role of Refugia and Climate Change

J. R. Stewart, C. B. Stringer

Although an African origin of the modern human species is generally accepted, the evolutionary processes involved in the speciation, geographical spread, and eventual extinction of archaic humans outside of Africa are much debated. An additional complexity has been the recent evidence of limited interbreeding between modern humans and the Neandertals and Denisovans. Modern human migrations and interactions began during the buildup to the Last Glacial Maximum, starting about 100,000 years ago. By examining the history of other organisms through glacial cycles, valuable models for evolutionary biogeography can be formulated. According to one such model, the adoption of a new refugium by a subgroup of a species may lead to important evolutionary changes.

Link

April 04, 2012

Modern humans responsible for early Aurignacian

The fact that modern humans created the Aurignacian has long been hypothesized, and was recently supported by dental analysis. This attribution is now supported by analysis of a couple of jawbones and associated dental remains from France.

Journal of Human Evolution

The Early Aurignacian human remains from La Quina-Aval (France)

Christine Verna et al.

There is a dearth of diagnostic human remains securely associated with the Early Aurignacian of western Europe, despite the presence of similarly aged early modern human remains from further east. One small and fragmentary sample of such remains consists of the two partial immature mandibles plus teeth from the Early Aurignacian of La Quina-Aval, Charente, France. The La Quina-Aval 4 mandible exhibits a prominent anterior symphyseal tuber symphyseos on a vertical symphysis and a narrow anterior dental arcade, both features of early modern humans. The dental remains from La Quina-Aval 1 to 4 (a dm1, 2 dm2, a P4 and a P4) are unexceptional in size and present occlusal configurations that combine early modern human features with a few retained ancestral ones. Securely dated to ∼33 ka 14C BP (∼38 ka cal BP), these remains serve to confirm the association of early modern humans with the Early Aurignacian in western Europe.

Link

Japanese population substructure (Nishiyama et al. 2012)

PLoS ONE 7(4): e35000. doi:10.1371/journal.pone.0035000

Detailed Analysis of Japanese Population Substructure with a Focus on the Southwest Islands of Japan

Takeshi Nishiyama et al.

Uncovering population structure is important for properly conducting association studies and for examining the demographic history of a population. Here, we examined the Japanese population substructure using data from the Japan Multi-Institutional Collaborative Cohort (J-MICC), which covers all but the northern region of Japan. Using 222 autosomal loci from 4502 subjects, we investigated population substructure by estimating FST among populations, testing population differentiation, and performing principal component analysis (PCA) and correspondence analysis (CA). All analyses revealed a low but significant differentiation between the Amami Islanders and the mainland Japanese population. Furthermore, we examined the genetic differentiation between the mainland population, Amami Islanders and Okinawa Islanders using six loci included in both the Pan-Asian SNP (PASNP) consortium data and the J-MICC data. This analysis revealed that the Amami and Okinawa Islanders were differentiated from the mainland population. In conclusion, we revealed a low but significant level of genetic differentiation between the mainland population and populations in or to the south of the Amami Islands, although genetic variation between both populations might be clinal. Therefore, the possibility of population stratification must be considered when enrolling the islander population of this area, such as in the J-MICC study.

Link

Cryptic distant relatives are common in genetic samples

Table S1 has some statistics on different populations.

PLoS ONE 7(4): e34267. doi:10.1371/journal.pone.0034267

Cryptic Distant Relatives Are Common in Both Isolated and Cosmopolitan Genetic Samples

Brenna M. Henn et al.

Although a few hundred single nucleotide polymorphisms (SNPs) suffice to infer close familial relationships, high density genome-wide SNP data make possible the inference of more distant relationships such as 2nd to 9th cousinships. In order to characterize the relationship between genetic similarity and degree of kinship given a timeframe of 100–300 years, we analyzed the sharing of DNA inferred to be identical by descent (IBD) in a subset of individuals from the 23andMe customer database (n = 22,757) and from the Human Genome Diversity Panel (HGDP-CEPH, n = 952). With data from 121 populations, we show that the average amount of DNA shared IBD in most ethnolinguistically-defined populations, for example Native American groups, Finns and Ashkenazi Jews, differs from continentally-defined populations by several orders of magnitude. Via extensive pedigree-based simulations, we determined bounds for predicted degrees of relationship given the amount of genomic IBD sharing in both endogamous and ‘unrelated’ population samples. Using these bounds as a guide, we detected tens of thousands of 2nd to 9th degree cousin pairs within a heterogenous set of 5,000 Europeans. The ubiquity of distant relatives, detected via IBD segments, in both ethnolinguistic populations and in large ‘unrelated’ populations samples has important implications for genetic genealogy, forensics and genotype/phenotype mapping studies.

Link

April 02, 2012

Craster's daughters

The second season of Game of Thrones has started and the series has become the topic of some fun discussion, I thought I'd give my €0.02 on the topic of genetic improbabilities.

(A warning: I have only read the first book, so please refrain from spoiling in the comments. Also, this post contains a minor spoiler about the first episode of season 2)

One of the good things about fiction is that it brings up interesting probabilities that would not often come up in the real world. In the first episode of Season 2, we are introduced to Craster, a character who lives north of the Wall. The interesting thing about Craster is that he practices a combination of incest and infanticide: he exposes his male offspring and weds his daughters, repeating the cycle. Let's examine whether or not this scenario is plausible:
  • In the first generation, Craster would have a 50% coefficient of relatedness with his daughters.
  • In the second generation, this would increase to 75%. In the third, this would be 87.5%, etc.

What is interesting is that their autozygosity would also increase across the generations.

In the first generation, at a random locus the probability that both copies have been inherited from the same ancestor is some random number corresponding to the probability of identity-by-descent in the general population.

In the second generation, at half the loci there is a 50-50 chance that the same allele will be inherited from Craster, hence the grand-daughters will be autozygous across at least 25% of the genome. This means that they will be homozygous for about a quarter of Craster's deleterious genetic load.

When we first see Craster, he seems to have quite a few "wives", so my guess is that this has been going on for 2-3 generations at least, which seems plausible given the early age of marriage in a quasi-medieval world.

However, given the severe abnormalities observed in children of father-daughter incest it will become increasingly difficult for Craster to obtain viable offspring through this practice. Many of his children would be aborted, die, or be severely incapacitated.

In order to increase his harem's size (as he seems to be doing) he would need each of his wives to produce 2 daughters for him. This translates to an average of 4 offspring per daughter (since 2 boys will also be born, on average, and exposed).

But, due to high levels of autozygosity, many of Craster's offspring would not be viable; hence, each wife must undergo a very large number of pregnancies. And, given that each wife is heavily inbred herself, she is less likely to survive many pregnancies, especially since there's no ob/gyn north of the Wall.

In conclusion, the tale of "Craster and his Wives" does seem to fit well in a work of fantasy...

March 31, 2012

Iceman's sheep belonged to mtDNA haplogroup B

This establishes that the main mtDNA haplogroup (B) of extant European sheep was already present in the ~5.3ky old sheep hair shafts of the Tyrolean Iceman's clothing. The fact that the precise sheep sequence (like that of its bearer's) has not been identified in modern sheep testifies to the importance that drift and/or selection has played in the recent evolution of the species.

PLoS ONE 7(3): e33792. doi:10.1371/journal.pone.0033792

Phylogenetic Position of a Copper Age Sheep (Ovis aries) Mitochondrial DNA



Abstract Top
Background
Sheep (Ovis aries) were domesticated in the Fertile Crescent region about 9,000-8,000 years ago. Currently, few mitochondrial (mt) DNA studies are available on archaeological sheep. In particular, no data on archaeological European sheep are available.

Methodology/Principal Findings
Here we describe the first portion of mtDNA sequence of a Copper Age European sheep. DNA was extracted from hair shafts which were part of the clothes of the so-called Tyrolean Iceman or Otzi (5,350 - 5,100 years before present). Mitochondrial DNA (a total of 2,429 base pairs, encompassing a portion of the control region, tRNAPhe, a portion of the 12S rRNA gene, and the whole cytochrome B gene) was sequenced using a mixed sequencing procedure based on PCR amplification and 454 sequencing of pooled amplification products. We have compared the sequence with the corresponding sequence of 334 extant lineages.

Conclusions/Significance
A phylogenetic network based on a new cladistic notation for the mitochondrial diversity of domestic sheep shows that the Otzi's sheep falls within haplogroup B, thus demonstrating that sheep belonging to this haplogroup were already present in the Alps more than 5,000 years ago. On the other hand, the lineage of the Otzi's sheep is defined by two transitions (16147, and 16440) which, assembled together, define a motif that has not yet been identified in modern sheep populations.


Link

Three quarters of Kerey clan men belong to Genghis Khan Y chromosome cluster

From the paper:
According to the historical data, the split between two sub-clans of the Kereys occurred about 20-22 generations ago (Khalidullin 2005). Estimation of divergence time (TD) of two groups of 15 STR haplotypes (except for DYS385a,b loci) found in the Kereys sub-clans demonstrates that TD value equal to 630 ± 190 years (or approximately 21 ± 6 generations) is resulted when a mean of per-locus, per-generation mutation rate of 0.0033 and a 30-year generation time are used. Note that similar value of mutation rate (0.00324) has been calculated as optimal for 15 STR haplotypes by Busby et al. (2011) who have investigated the question on how average squared distance (ASD) estimates change within haplotype sets when using different combinations of Y-chromosome STRs. This mutation rate belongs to a class of so called genealogical STR mutation rates revealed by direct observation in father/son pairs (Kayser et al. 2000; Goedbloed et al. 2009).
The correspondence between the split time of the Kerey sub-clans and the age estimate of their Y-STR divergence is quite interesting and provides an independent historical argument for the correspondence between the C3* star cluster and Genghis Khan (or at least his direct patrilineal kin). Note that the star cluster's age matches G. K. only using a genealogical mutation rate, and not the widely (mis)used "effective mutation rate. The timeframe is recent enough to render any saturation effects from non-linearity (as described by Busby et al.) relatively unimportant.

More:

The data reported above, taken together with the known arguments in favor of the
possible Genghis Khan‟s descent of Y-chromosome C3* star-cluster (Zerjal et al. 2003), allow us to suggest two hypotheses.
(1) The star-cluster is not directly related to the descendants of Genghis Khan, but rather is associated with the Kerait clan members. Mongol conquest with participation of the Keraits as special Khan‟s military forces allowed them to disseminate the Kerait-specific Y-chromosomes in the vast area inhabited by various peoples.
(2) Genghis Khan by himself belonged to the Keraits. This is supported by the following historical evidence (Man 2004; Khalidullin 2005). The Keraits inhabited the banks of the Onon River, where the camp of Genghis Khan‟s father Yesukhei was located. Yesukhei was declared as a blood brother of the Keraits‟ Khan Toghrul (Wang Khan). Toghrul then declared Genghis Khan his son-in-law. Fraternization of the Genghis Khan family with the Keraits‟ Khan suggests  that a real blood relationship, though probably not approved officially, existed between them.


Human Biology: Vol. 84: Iss. 1, Article 4.

The Y-chromosome C3* star-cluster attributed to Genghis Khan's descendants is present at high frequency in the Kerey clan from Kazakhstan

Serikbai Abilev et al.

In order to verify the possibility that the Y-chromosome C3* star-cluster attributed to Genghis Khan and his patrilineal descendants is relatively frequent in the Kereys, who are the dominant clan in Kazakhstan and in Central Asia as a whole, polymorphism of the Y-chromosome was studied in Kazakhs, represented mostly by members of the Kerey clan. The Kereys showed the highest frequency (76.5%) of individuals carrying the Y-chromosome variant known as C3* star-cluster ascribed to the descendants of Genghis Khan. C3* star-cluster haplotypes were found in two sub-clans, Abakh-Kereys and Ashmaily-Kereys, diverged about 20-22 generations ago according to the historical data. Median network of the Kerey star-cluster haplotypes at 17 STR loci displays a bipartite structure, with two subclusters defined by the only difference at DYS448 locus. It is noteworthy that there is a strong correspondence of these subclusters with the Kerey sub-clans affiliation. The data obtained suggest that the Kerey clan appears to be the largest known clan in the world descending from a common Y-chromosome ancestor. Possible ways of Genghis Khan‟s relation to the Kereys are discussed.

Link

March 28, 2012

A rare look at the Y chromosomes of Afghanistan

I often bemoan the fact that some of the regions of the world that are most interesting to the student of prehistory (e.g., Mesopotamia and the Iranian Plateau) seem to also be the ones with more than their fair share of political trouble, hindering efforts to study them with the newest set of tools. Afghanistan is certainly one case that hasn't been quite the most welcoming of places in recent decades.

The country is transitional between the Iranic speaking world of Iran and the Indo-Aryan speaking world of South Asia, as well as between the Indo-Iranian world and the (mostly) Turkic-speaking world of Central Asia. Hence, the absence of data for that country has been acutely felt for all those who are trying to understand "what happened" in Eurasia.

The appearance of a new paper by the Genographic Project is a welcome sight, and a good example of what is best about this Project. I haven't been exactly a fan of the Genographic's interpretation of their own data, but kudos to them for getting them in the first place.

From the paper:
Pashtuns are the largest ethnic group in Afghanistan, accounting for about 42 percent of the population, with Tajiks (27%), Hazaras (9%), Uzbeks (9%), Aimaqs (4%), Turkmen people (3%), Baluch (2%), and other groups (4%) making up the remainder [6]. In the present study, eight ethnic groups were examined, with a focus on the largest four groups: - The Pashtuns, traditionally lived a seminomadic lifestyle, they reside mainly in southern and eastern Afghanistan and in western Pakistan. They speak Pashto which is a member of the Eastern Iranian languages. - The Tajiks are a Persian-speaking ethnic group which are closely related to the Persians of Iran. In Afghanistan, they are the largest Tajik population outside their homeland to the north in Tajikistan. - The Hazara population speaks Persian with some Mongolian words. They believe they are descendants of Genghis Khan's army that invaded during the twelfth century. - The Uzbeks are a Turkic speaking group that have been living a sedentary farming lifestyle in Northern Afghanistan.
The main features of the Y-chromosome gene pool:
Genotyping revealed 32 halpogroups present in Afghanistan's ethnic groups among our samples. Haplogroups R1a1a-M17, C3-M217, J2-M172, and L-M20 were the most frequent when Afghan ethnic groups were pooled, together comprising >66% of the chromosomes. Absolute and relative haplogroup frequencies are tabulated in Table S4.
-The PCA analysis (left) showcases wonderfully the correspondence between different haplogroups and the three main regions of the Near East (green), South Asia (yellow), and Central Asia (purple).

It is a real shame that the newer markers available within the most prominent R-M17 haplogroup were not tested:
The prevailing Y-chromosome lineage in Pashtun and Tajik (R1a1a-M17), has the highest observed diversity among populations of the Indus Valley [46]. R1a1a-M17 diversity declines toward the Pontic-Caspian steppe where the mid-Holocene R1a1a7-M458 sublineage is dominant [46]. R1a1a7-M458 was absent in Afghanistan, suggesting that R1a1a-M17 does not support, as previously thought [47], expansions from the Pontic Steppe [3], bringing the Indo-European languages to Central Asia and India.
Nonetheless, I can't really disagree with the dismissal of the R-M17/Indo-European theory. R-M17 is simply too populous in South Asia to be the genetic legacy of "Indo-Europeans": (i) under an elite-dominance model, its frequency is way too high (compared to well-attested examples of elite dominance, e.g., Hungary or Turkey where the genetic legacy of the elite element is in the minority), (ii) under a folk migration model, it is difficult to understand why a hypothetical migrating Indo-European people would have such an overwhelming influence in the region while at the same time hardly influencing at all other densely occupied agricultural landscapes of the Eurasian steppe periphery; moreover, no autosomal signal corresponding to a migration from eastern Europe to South Asia really exists -the main cline of variation links South with West Asia, not Europe- and the small signal that does exist does not really correspond to observed levels of R-M17.

From the paper:
The E1b1b1-M35 lineages in some Pakistani Pashtun were previously traced to a Greek origin brought by Alexander's invasions [48]. However, RM network of E1b1b1-M35 found that Afghanistan's lineages are correlated with Middle Easterners and Iranians but not with populations from the Balkans.
Greek populations are not homogeneous in their haplogroup E frequencies, so it would be useful to consider the possibility that the lack of this frequent Southeastern European haplogroup in South Asia may not reflect a complete lack of Greek influence in this region, but rather, an influence from a structured ancient Greek population.

Looking at the Y-haplogroup composition:

A few points of interest:

  • The clear link between C/N/O with Central Asia
  • A clear difference between Persian and Pashto speakers in terms of inverse J2a/R1a frequences
  • The paucity of J1 chromosomes (only 1 Tajik) testifies to the absence of relatively recent Middle Eastern influences associated with the spread of Islam; consistent with the absence of the autosomal "Southwest Asian" component in South/Central Asia.
  • Paucity of R1b, except in a couple Uzbeks and a Tajik; I have argued before that R1a had an early distribution in the arc of flatlands north and east of the Caspian, while R1b a complementary distribution in the smaller arc of the highlands west and south of it, out of which the Tocharians may have originated.
  • The small Nurestani sample comprises of J2a, R1a, and R2; these are linguistic relatives of the Kalash of Pakistan who -unlike the latter- were converted to Islam in the 19th century.
I would say that the evidence is pretty clear that the earliest Iranians may have included haplogroups R1a and J2, although I would not wager on their relative proportions and overall contribution to modern Iranian-speaking populations. For whatever reason, it seems that Kurds and Persians ended up with a J2-over-R1a advantage, while Pathans and (plausibly) Turkified Central Asian former Iranian speakers with the reverse. Nonetheless, the occurrence of both haplogroups in most Iranian groups, as well as in most Indo-Aryan ones is quite telling. It is unfortunate that the relationships between these Y chromosomes (still J2a*! six years after Sengupta et al.) and their West Eurasian brethren was not further pursued.

Hopefully, the data can be re-used down the road once the phylogeny of different haplogroups (and R1a in particular) is better understood. As I've stated before on this blog, I take Y-STR based age estimates with a huge grain of salt, so I would not put much faith in any of the ones presented in this paper.

Related: Firasat et al. (2006), Y-chromosomes of Afghanistan, Lashgary et al. (2011), Regueiro et al. (2006).

PLoS ONE doi:10.1371/journal.pone.0034288


Afghanistan's Ethnic Groups Share a Y-Chromosomal Heritage Structured by Historical Events

Marc Haber et al.

Abstract


Afghanistan has held a strategic position throughout history. It has been inhabited since the Paleolithic and later became a crossroad for expanding civilizations and empires. Afghanistan's location, history, and diverse ethnic groups present a unique opportunity to explore how nations and ethnic groups emerged, and how major cultural evolutions and technological developments in human history have influenced modern population structures. In this study we have analyzed, for the first time, the four major ethnic groups in present-day Afghanistan: Hazara, Pashtun, Tajik, and Uzbek, using 52 binary markers and 19 short tandem repeats on the non-recombinant segment of the Y-chromosome. A total of 204 Afghan samples were investigated along with more than 8,500 samples from surrounding populations important to Afghanistan's history through migrations and conquests, including Iranians, Greeks, Indians, Middle Easterners, East Europeans, and East Asians. Our results suggest that all current Afghans largely share a heritage derived from a common unstructured ancestral population that could have emerged during the Neolithic revolution and the formation of the first farming communities. Our results also indicate that inter-Afghan differentiation started during the Bronze Age, probably driven by the formation of the first civilizations in the region. Later migrations and invasions into the region have been assimilated differentially among the ethnic groups, increasing inter-population genetic differences, and giving the Afghans a unique genetic diversity in Central Asia.

Link

High mtDNA mutation rate from deep-rooted Costa Rican pedigrees

American Journal of Physical Anthropology DOI: 10.1002/ajpa.22052

High mitochondrial mutation rates estimated from deep-rooting costa rican pedigrees

Lorena Madrigal et al.

Abstract

Estimates of mutation rates for the noncoding hypervariable Region I (HVR-I) of mitochondrial DNA vary widely, depending on whether they are inferred from phylogenies (assuming that molecular evolution is clock-like) or directly from pedigrees. All pedigree-based studies so far were conducted on populations of European origin. In this article, we analyzed 19 deep-rooting pedigrees in a population of mixed origin in Costa Rica. We calculated two estimates of the HVR-I mutation rate, one considering all apparent mutations, and one disregarding changes at sites known to be mutational hot spots and eliminating genealogy branches which might be suspected to include errors, or unrecognized adoptions along the female lines. At the end of this procedure, we still observed a mutation rate equal to 1.24 × 10−6, per site per year, i.e., at least threefold as high as estimates derived from phylogenies. Our results confirm that mutation rates observed in pedigrees are much higher than estimated assuming a neutral model of long-term HVRI evolution. We argue that until the cause of these discrepancies will be fully understood, both lower estimates (i.e., those derived from phylogenetic comparisons) and higher, direct estimates such as those obtained in this study, should be considered when modeling evolutionary and demographic processes.

Link

Improved eigenanalysis with Minimum Average Partial test (Shriner 2012)

I had mentioned a previous article by the same author on the topic of how many PCA dimensions to retain. I had identified this problem in the context of my "Clusters Galore" analysis, and the topic has recently re-surfaced in the recent Falush and Lawson pre-print, which, unfortunately, appeared at the same time as the publication of this new paper.

Personally, I have tried three methods for choosing the number of principal components to retain:

  1. Tracy Widom; which seems to retain more dimensions than are necessary, with a resulting reduction in clustering quality
  2. A test of normality (such as Shapiro-Wilk), which tends to identify a smaller number of dimensions where the data appear not normally distributed and hence may contain useful information about population structure.
  3. A more pragmatic approach of picking the number of components to retain that maximize the number of inferred clusters by MCLUST
It would be great if this test could be incorporated into future versions of EIGENSOFT.

Human Heredity Vol. 73, No. 2, 2012

Improved Eigenanalysis of Discrete Subpopulations and Admixture Using the Minimum Average Partial Test

Daniel Shriner

Abstract Principal components analysis of genetic data has benefited from advances in random matrix theory. The Tracy-Widom distribution has been identified as the limiting distribution of the lead eigenvalue, enabling formal hypothesis testing of population structure. Additionally, a phase change exists between small and large eigenvalues, such that population divergence below a threshold of FST is impossible to detect and above which it is always detectable. I show that the plug-in estimate of the effective number of markers in the EIGENSOFT software often exceeds the rank of the sample covariance matrix, leading to a systematic overestimation of the number of significant principal components. I describe an alternative plug-in estimate that eliminates the problem. This improvement is not just an asymptotic result but is directly applicable to finite samples. The minimum average partial test, based on minimizing the average squared partial correlation between individuals, can detect population structure at smaller FST values than the corrected test. The minimum average partial test is applicable to both unadmixed and admixed samples, with arbitrary numbers of discrete subpopulations or parental populations, respectively. Application of the minimum average partial test to the 11 HapMap Phase III samples, comprising 8 unadmixed samples and 3 admixed samples, revealed 13 significant principal components.

Link

mtDNA links between Africa and Europe, old and new (Cerezo et al. 2012)

Link to open access supplementary material.


Genome Research DOI: 10.1002/ajpa.22052

Reconstructing ancient mitochondrial DNA links between Africa and Europe

María Cerezo et al.

Abstract

Mitochondrial DNA (mtDNA) lineages of macro-haplogroup L (excluding the derived L3 branches M and N) represent the majority of the typical sub-Saharan mtDNA variability. In Europe, these mtDNAs account for less than 1% of the total but, when analyzed at the level of control region, they show no signals of having evolved within the European continent, an observation that is compatible with a recent arrival from the African continent. To further evaluate this issue, we analyzed 69 mitochondrial genomes belonging to various L sublineages from a wide range of European populations. Phylogeographic analyses showed that ∼65% of the European L lineages most likely arrived in rather recent historical times, including the Romanization period, the Arab conquest of the Iberian Peninsula and Sicily, and during the period of the Atlantic slave trade. However, the remaining 35% of L mtDNAs form European-specific subclades, revealing that there was gene flow from sub-Saharan Africa toward Europe as early as 11,000 yr ago.

Link

March 27, 2012

Cranial variation and the transition to agriculture in Europe

Students of physical anthropology won't be surprised that Pinhasi and von Cramon-Taubadel find that the Neolithic and pre-Neolithic populations in Europe were differentiated cranially as they were apparently genetically.

It has long been recognized that the ancient European population was different than the Upper Paleolithic population of the continent. Carleton Coon ascribed this differentiation to migration of narrow-faced Mediterraneans into the territory of the robust broad-faced Upper Paleolithics. Ilse Schwidetzky also viewed migration from the Southeast of gracile Mediterraneans who gradually replaced broad-faced Cro-Magnoids.

So, it is nice to read that the re-analysis of a wide assortment of skulls on 15 cranial variables has revealed that:
The major shape differences separating hunter-gatherer Mesolithic populations and farming Neolithic populations are coded by PC1 with Neolithic specimens having longer and taller vaults, and Mesolithic specimens having larger, and broader faces.
There are two (or three) puzzles in European prehistory:
  • How the robust, low-skulled, broad-faced hunter-gatherers became more high-skulled, narrow-faced and gracile
  • How the latter became brachycephalized until early modern times
  • Why they have become partially debrachycephalized in the most recent of times
Anthropologists have tended to favor either migration or adaptation to explain these trends, with some even suggesting simple phenotypic plasticity without any major genetic change. It is now clear that -whatever the role of adaptation or plasticity- the Upper Paleolithic population of Europe did not simply change to become more gracile on its own, but was affected by an already gracile population of foreign origin who set the ball rolling. There is already work on the genetic basis of facial structure, so, it is quite possible that eventually we'll be able to track directly the genetic changes underlying the phenotypic transformation of Europeans.

From the paper:
Nonetheless, the craniometric analysis allows us to discern certain patterns. For example, the ‘Forest Neolithic’ specimens are clearly much more similar to other Mesolithic hunter-gatherers than to Neolithic farmers in terms of their craniometric shape, suggesting a large degree of cultural diffusion in this region. However, it is also evident that the earliest potential colonisers of southeast and central Europe are very similar to the Anatolian Çatal Höyük population, congruent with an initial demic diffusion from the Near East/Anatolia.
The "Forest Neolithic" included pottery-using groups of eastern Europe (hence Neolithic, since pottery is one of the hallmarks of that period), but should not be confused with the early agriculturalists who apparently practiced farming without pottery early on in the Near East and Greece, and then acquired pottery and expanded with it into the rest of Europe, together with their full "package" of domesticated crops and animals.

Human Biology vol. 84

Cranial variation and the transition to agriculture in Europe

Ron Pinhasi, Noreen Von Cramon-Taubadel

Abstract

Debates surrounding the nature of the Neolithic demographic transition in Europe have historically centred on two opposing models; a 'demic' diffusion model whereby incoming farmers from the Near East and Anatolia effectively replaced or completely assimilated indigenous Mesolithic foraging communities and an 'indigenist' model resting on the assumption that ideas relating to agriculture and animal domestication diffused from the Near East, but with little or no gene flow. The extreme versions of these dichotomous models have been heavily contested primarily on the basis of archaeological and modern genetic data. However, in recent years there has been a growing acceptance of the likelihood that both processes were ongoing throughout the Neolithic transition and that a more complex, regional approach is required to fully understand the change from a foraging to a primarily agricultural mode of subsistence in Europe. Craniometric data have been particularly useful for testing these more complex scenarios, as they can reliably be employed as a proxy for the genetic relationships amongst Mesolithic and Neolithic populations. In contrast, modern genetic data assume that modern European populations accurately reflect the genetic structure of Europe at the time of the Neolithic transition, while ancient DNA data are still not geographically or temporally detailed enough to test continent-wide processes. Here, with particular emphasis on the role of craniometric analyses, we review the current state of knowledge regarding the cultural and biological nature of the Neolithic transition in Europe.

Link

March 26, 2012

Similarity matrices and clustering (Lawson and Falush)

Lawson and Falush have a new review paper on different clustering methods using haplotype data such as their own ChromoPainter/fineSTRUCTURE methodology, as well as the MCLUST/fastIBD methods that I started playing with a while back.

I won't have much time for the next few days to comprehensively review this new work, but I will add one data point to the discussion, by pointing to my ChromoPainter and fastIBD analyses over the same dataset. I will also add any further comments on this blog post, once I get the opportunity to read the paper.

Another point that needs to be made is how commendable the ChromoPainter folks' attitude towards the topic has been. Not only did they post their ChromoPainter preprint and software online months before their original paper was published, but they quickly jumped on my comments and suggestions on their paper to write their new review paper, making at available as a preprint itself. I'm guessing this saved about a year or two over what would have been possible if all the formalities of "traditional" publishing had been observed. It's also a very nice example of synergy between professional and amateur science, that the Internet and social media has made possible.

Similarity matrices and clustering algorithms for population identification using genetic data


Daniel John Lawson and Daniel Falush

Abstract

A large number of algorithms have been developed to identify population
structure from genetic data. Recent results show that the information used
by both model-based clustering methods and Principal Components Analysis
can be summarised by a matrix of pairwise similarity measures between
individuals. Similarity matrices have been constructed in a number of ways,
usually treating markers as independent but differing in the weighting given
to polymorphisms of different frequencies. Additionally, methods are now being
developed that better exploit the power of genome data by taking linkage
into account. We review several such matrices and evaluate their ‘information
content’. A two-stage approach for population identification is to first construct
a similarity matrix, and then perform clustering. We review a range
of common clustering algorithms, and evaluate their performance through a
simulation study. The clustering step can be performed either directly, or
after using a dimension reduction technique such as Principal Components
Analysis, which we find substantially improves the performance of most algorithms.
Based on these results, we describe the population structure signal
contained in each similarity matrix, finding that accounting for linkage leads
to significant improvements for sequence data. We also perform a comparison
on real data, where we find that population genetics models outperform
generic clustering approaches, particularly in regards to robustness against
features such as relatedness between individuals.


Link

March 24, 2012

Report on the symposium on Modern Human Genetic Variation

Joshua Akey summarizes the talks of a recent symposium at the Swedish Royal Academy of Sciences. Two bits of information stand out from his report. The first:

In another talk focused on demography, Mattias Jakobsson (Uppsala University, Sweden) presented novel data on the impact of the agricultural revolution on the genetics of contemporary European populations. Specifically, Jakobsson and colleagues obtained nearly 250 Mb of sequence from three 5,000-year-old remains of Neolithic hunter-gatherers and one Neolithic farmer excavated in Scandinavia. Analysis of these sequences in the context of the present day European gene pool suggests that the spread of agriculture involved the northward migrations of farmers. Thus, these data provide the most direct and compelling support for the demic diffusion model of agriculture (as opposed to cultural diffusion) described to date. 

It seems I have my answer to the what's next question. Jakobsson has been doing some interesting work on the demography of human emergence and dispersal, so it will be interesting to see not only the novel sequences from these Neolithic Scandinavians, but also how they fit into existing models of demic diffusion.

The second bit of information:

Similarly, Jeff Wall (University of California San Francisco, USA) described a novel method for inferring archaic admixture, which he applied to publicly available whole-genome sequence data generated by Complete Genomics. Provocatively, he finds higher rates of introgression in Asians compared to Europeans. An advantage of Wall’s method is that it does not require an archaic genome to infer introgression, and thus he was able to also test the hypothesis that contemporary African genomes have signatures of gene flow with archaic human ancestors. Strikingly, Wall indeed did find evidence of archaic admixture in African genomes, suggesting that modest amounts of gene flow were widespread throughout time and space during the evolution of anatomically modern humans.

I guess that I shouldn't throw explanation #1 out the window yet. Wall was involved in the recent paper on archaic African admixture, which only looked at a small subset of the genome, so it is nice to see that he is now working with full genomes, and that the race to data mine complete genomes for archaic admixture is afoot.

The book of abstracts is online at the symposium site. The Jakobsson paper does seem to agree with our emerging picture of a non-local origin of northern European farmers as well as greater survival of pre-farming populations in the northern periphery of Europe, but it will be interesting to see where exactly extant populations fall on the farmer-hunter/gatherer continuum.
Origins and genetic legacy of Neolithic farmers and hunter-gatherers in Northern Europe 
Mattias Jakobsson
Department of Evolutionary Biology, Evolutionary Biology Centre (EBC), Uppsala University, Sweden 
The prehistoric spread of farming in Europe has garnered intense interest for almost a century, and was one of the first questions to which population genetic data was used to investigate demographic hypotheses. However, the impact of the agricultural revolution on the European gene pool remains largely unknown. We obtained 249 million base pairs of quality-filtered human autosomal sequence data from some 5,000 year-old remains of three Neolithic hunter-gatherers and one Neolithic farmer excavated in Scandinavia, the northernmost fringe of agricultural practice at the time. Applying novel methods to study population structure based on low genome-coverage data, we find that Northern European Neolithic farmers are most similar to modern-day southern Europeans, contrasting sharply to Neolithic hunter-gatherers who are most similar to extant individuals from northern Europe. With most extant European populations appearing genetically intermediate between the two Neolithic groups, our results suggest that migration from the south by a genetically distinct group of humans accompanied the spread of agriculture to geographic regions where hunting and gathering was the mode of subsistence, but that admixture eventually shaped modern-day patterns of genomic variation.

Archaic admixture in the human genome 
Jeff D Wall
Department of Epidemiology & Biostatistics, University of California, San Francisco, USA 
We describe a method that uses patterns of linkage disequilibrium in extant human populations to identify regions of the genome that were inherited from ‘archaic’ human ancestors, such as Neandertals, Homo erectus or H. floresiensis. We validate this approach using two recently published archaic human genomes, and show that several ancient admixture events must have occurred, both within and outside of Africa. We also explore differences in the amount of archaic admixture across different contemporary human populations.


Finally, here is the meeting report:

Investigative Genetics 2012, 3:7 doi:10.1186/2041-2223-3-7

Understanding human evolutionary history: a meeting report of the Swedish Royal Academy of Sciences symposium of modern human genetic variation 

Joshua M Akey

Link (pdf)