Showing posts with label Semitic. Show all posts
Showing posts with label Semitic. Show all posts

January 07, 2013

mtDNA variation in East Africa (Boattini et al. 2013)

From the paper:
Language diversity in EA fits well with its complicated genetic history. In Fleming words, ‘‘Ethiopia by itself has more languages than all of Europe, even counting all the so-called dialects of the Romance family’’ (Fleming, 2006). All African linguistic phyla are found in EA: Afro-Asiatic (AA), Nilo-Saharan, Niger-Congo and Khoisan (however, the genealogical unit of Khoisan is no longer generally accepted). Among them, AA is the most differentiated, being represented by three (Omotic, Cushitic, Semitic) of its six major clades (the others being Chadic, Berber and Egyptian). Omotic and Cushitic are considered the deepest clades of AA, and both are found almost exclusively in the Horn of Africa, along with the linguistic relict Ongota that is traditionally assigned to the Cushitic family but whose classification is still widely debated (Fleming, 2006). These observations are in agreement with a North-Eastern African origin of the AA languages, most probably in pre-Neolithic times (Ehret, 1979, 1995; Kitchen et al., 2009).
and:

This study confirms the central role of EA and the Horn of Africa in the genetic and linguistic history of a wide area spanning from Central and Northern Africa to the Levant. Our results confirm high mtDNA diversity and strong genetic structuring in EA. We were indeed able to identify three population clusters (A, B1, B2) that are related both to geography and linguistics, and signaling different population events in the history of the region. The Horn of Africa (cluster A), in accordance with its role as a major gateway between sub-Saharan Africa and the Levant, shows widespread contacts with populations from CA (AA-Chadic speakers), the Arabian peninsula and the Nile Valley. Southwards, Kenya, and Tanzania (clusters B1 and B2), despite being both heavily involved in Bantu and Nilo-Saharan pastoralist expansions, reveal traces of a more ancient genetic stratum associated with Cushitic-speaking groups (cluster B2). Conversely, Berber- and Semitic-speaking populations of NA and the Levant show only marginal traces of admixture with sub-Saharan groups, as well as a different mtDNA genetic background, making the hypothesis of a Levantine origin of AA unlikely. In conclusion, EA genetic structure configures itself as a complicated palimpsest in which more ancient strata (AA-Cushiticspeaking groups) are largely overridden by recent different migration events. Further explorations of AA-Cushitic- speaking populations – both in terms of sampled groups and typed genetic markers – will be of great importance for the reconstruction of the genetic history of EA and AA-speakers. 

The African origin of Afroasiatic would agree with its linguistic separateness from Eurasian languages, and the fact that a single branch of the family (Semitic) is likely to have originated in Asia, and fairly recently at that.

Related:



Am J Phys Anthropol DOI: 10.1002/ajpa.22212

mtDNA variation in East Africa unravels the history of afro-asiatic groups

Alessio Boattini et al.

East Africa (EA) has witnessed pivotal steps in the history of human evolution. Due to its high environmental and cultural variability, and to the long-term human presence there, the genetic structure of modern EA populations is one of the most complicated puzzles in human diversity worldwide. Similarly, the widespread Afro-Asiatic (AA) linguistic phylum reaches its highest levels of internal differentiation in EA. To disentangle this complex ethno-linguistic pattern, we studied mtDNA variability in 1,671 individuals (452 of which were newly typed) from 30 EA populations and compared our data with those from 40 populations (2970 individuals) from Central and Northern Africa and the Levant, affiliated to the AA phylum. The genetic structure of the studied populations—explored using spatial Principal Component Analysis and Model-based clustering—turned out to be composed of four clusters, each with different geographic distribution and/or linguistic affiliation, and signaling different population events in the history of the region. One cluster is widespread in Ethiopia, where it is associated with different AA-speaking populations, and shows shared ancestry with Semitic-speaking groups from Yemen and Egypt and AA-Chadic-speaking groups from Central Africa. Two clusters included populations from Southern Ethiopia, Kenya and Tanzania. Despite high and recent gene-flow (Bantu, Nilo-Saharan pastoralists), one of them is associated with a more ancient AA-Cushitic stratum. Most North-African and Levantine populations (AA-Berber, AA-Semitic) were grouped in a fourth and more differentiated cluster. We therefore conclude that EA genetic variability, although heavily influenced by migration processes, conserves traces of more ancient strata.

Link

November 22, 2012

ALDER signal of admixture in Ashkenazi Jews

(You can skip the first part if you want, and head straight to the RESULTS section)

Previous studies on uniparental markers have indicated that Ashkenazi Jews (AJ) were formed by admixture between a Near Eastern population and European host populations; the evidence for the former element seems pretty clear on the basis of Y-chromosomes where Jews possess a relatively high frequency of Y-haplogroup J1 (and a few others) that are quite rare in non-Jewish north/east Europeans. As for the latter, it seems probable on the basis of the location of Ashkenazi Jews on PCA plots where they tend to occupy an intermediate position between extant populations of the Levant (including Near Eastern Jews) and non-Jewish Europeans.

Anyone who has played around with genetic data will know that while AJ may be positioned in the aforementioned "intermediate" location within the "West Eurasian continuum" between Europe and Near East, they tend to form their own cluster at higher dimensions. And, indeed, this is why it's fairly easy for a clustering algorithm, such as my "Clusters Galore" (MCLUST/MDS) approach to pick out a very specific AJ cluster (e.g., here, or here, using a fastIBD approach). An Ashkenazi Jewish-specific cluster also pops out at higher K in ADMIXTURE analyses. This cluster may reflect endogamy within the AJ community until quite recent times.

One way of detecting admixture in a group is through the use of f3-statistics. The statistic f3(AJ; European, Near_East) could be negative --which would indicate admixture-- but it is usually not -at least in the combinations of (European, Near_East) I've tried, and this is consistent with either the presence admixture or absence of admixture.

A simple and intuitive way to see why post-admixture drift might mask the presence of admixture can be seen by means of a simple calculation. Remember that the f3-statistic's +/- sign depends on the +/- sign of quantities (c-a)*(c-b) where c is an allele frequency in the admixed (?) population we are investigating, and a, b in the two reference populations. We can pick a to be less than b with no loss of generality.

In the absence of strong drift (e.g., if all populations have a very large number of individuals), then the allele frequency c=xa+(1-x)b where x is the amount of admixture --between 0 and 1-- from group A and (1-x) from group B, and this c will be maintained little changed in the post-admixture phase. With the aid of a little algebra, we get that:

(c-a)*(c-b) = (xa+(1-x)b-a)*(xa+(1-x)b-b)
= (xa+b-xb-a)*(xa+b-xb-b) =
= x(x-1)(a-b)^2

and this is of course negative because we assumed that x was less than 1.

In a large population, this c will remain near-constant, because of the lack of strong drift. As long as it remains within the interval (a,b), then (c-a)*(c-b) will also remain negative, and so will the f3 statistic.

But, what if strong drift affects the admixed population? Allele frequencies fluctuate more wildly in larger populations, so c might go outside the (a,b) interval. Without loss of generality, assume that c becomes greater than b in which case (c-a)*(c-b) will become positive.

The f3-statistic averages over many SNPs, so, depending on (i) the initial differentiation of the admixed populations, which could be seen as b-a, and (ii) the amount of drift, which causes c to jump outside the (a, b) interval as discussed above, it is possible that the evidence for admixture may disappear.

So, relying on allele frequency differences may help obliterate the signal of admixture. But, there is a different signal of admixture that uses the decay of admixture linkage-disequilibrium, most recently discussed in the ALDER paper. The admixture LD signal's evidence may also disappear in time, but only because the signal occurs at increasingly lower genetic distances over time due to recombination. Thankfully, it tends to occur at large enough --for the last few thousand years-- distances, for which the SNP density of existing genotyping platforms that measure a few hundred thousand SNPs per individual is sufficient.

METHODS

Naturally I was curious to see whether the admixture LD mechanism would produce the evidence of admixture that the f3-statistics did not. I combined three datasets in my possession (HGDP by Li et al. Behar et al. and Yunusbayev et al. ) and identified sets of European and Semitic populations. (Remember that these sets are non-exhaustive, but presumably usable surrogates for the true mixing populations exist within them):

Abhkasians_Y, Adygei, Belorussian, Bulgarians_Y, Chechens_Y, Chuvashs, French, French_Basque, Georgians, Hungarians, Lezgins, Lithuanians, Mordovians_Y, North_Italian, North_Ossetians_Y, Orcadian, Romanians, Russian, Sardinian, Spaniards, Tuscan, Ukranians_Y

and:

Bedouin, Druze, Egyptans, Ethiopian_Jews, Ethiopians, Iraq_Jews, Jordanians, Lebanese, Morocco_Jews, Palestinian, Saudis, Sephardic_Jews, Syrians, Yemenese, Yemen_Jews

I used my Dodecad Project sample of AJ which numbers 36 individuals and is larger than any other usable public sample available to me.

(ALDER was run with default parameters, using the Rutgets recombination map for Illumina chips, and with the merged dataset prepared with a --geno 0.03 flag. Note that the Ashkenazi_D sample consists of individuals typed on different Illumina platforms from 23andMe and FamilyTreeDNA. The total number of SNPs considered was 527,165.)

RESULTS

I report below the tests for which ALDER reported "success" for the test with no warnings:



The median of all these estimates is 36.78 generations or 1070 years which corresponds to a calendar date of 910CE, assuming the sample's birthday was 1980, and a generation length of 29 years.

Palamara et al. placed the beginning of demographic expansion of AJ in a similar timeframe (33 generations), following a severe founder effect reducing the population to ~270 individuals. Such a founder effect may have indeed served to produce positive f3-statistics, masking the presence of admixture, the occurrence of which appears to be substantiated on the basis of the ALDER test of admixture.

As for the levels of admixture, using a 1-ref analysis with the European populations, I get the following lower bounds:



I'd be interested in hearing people's opinions on the plausibility of these dates/proportions, as well as their potential historical associations; a lot of factors might affect these results, so perhaps this analysis could be improved in the future.

August 08, 2012

fastIBD analysis of several Jewish and non-Jewish groups

This is more of a "just the data" kind of post, inspired by the two recent papers on Jewish origins. A few quick points:
  • fastIBD was run with default parameters over a dataset of 512 individuals/264,539 SNPs
  • fastIBD identifies segments of relatively recent origin that are shared by individuals. These results should not be construed as measures of overall genetic similarity or origins. Rather, they suggest which populations have exchanged genes in the relative recent past, say, the last two thousand years or so.
  • I included all Ashkenazi_D and North_African_Jews_D samples; of the other Dodecad and reference populations, I took random samples of 10 each; running time of fastIBD increases with the square of the number of individuals, so doing this allowed me to run this in less than a day as opposed to about a week.
With that said, you can get:
  • Spreadsheet of numeric results, showing sharing (in centi-Morgans, cM)
  • Population-level graphical results, showing an ordering of other populations based on mean IBD sharing.
The following heat map allows for a quick appraisal of populations sharing an excess of IBD sharing (read row-by-row)



And, here are a couple of the visualizations for a few Jewish populations:

Note that all sources of data are listed on the bottom left of the Dodecad blog.

June 22, 2012

Assessing East Africans of Pagani et al. (2012) using 'weac2'

Thanks to the publication of new data from Pagani et al. (2012), we now have 235 more individuals from East Africa, mainly Ethiopians, but also Somalis and South Sudanese with dense genotype data.

Naturally, I wanted to make sure that everything was in order, so I applied the 'weac2' calculator on the new data. Here are the normalized median admixture proportions:



I have also created population portraits for the 12 different populations, which appear to show rather homogeneous samples.

Here are the descriptions of the data from the original paper:

The populations sampled (numbers) were the Semitic-speaking Amhara (26) and Tigray (21); the Cushitic-speaking Oromo (21), Ethiopian Somali (17), and Afar (12); the Omotic-speaking Ari Cultivators (24), Ari Blacksmiths (17), and Wolayta (8); and the Nilotic-speaking Gumuz (19) and Anuak (23). In addition to these groups, we also generated South Sudanese data from mixed populations (24) and Somali data from Somali populations (23).


Newer versions of the Dodecad tools will of course take into account the new samples, which ought to  help better define the "East_African" component that often arises at higher levels of detail.

And, of course kudos to all researchers who make their data publicly available and hence provide genome bloggers such as myself with much appreciated "fuel" for their inquiries.

June 21, 2012

Ethiopian origins (Pagani et al. 2012)

The study attempts to answer four questions:
Our current study is motivated by four questions. First, where do the Ethiopians stand in the African genetic landscape? Second, what is the extent of recent gene flow from outside Africa into Ethiopia, when did it occur, and is there evidence of selection effects? Third, do genomic data support a route for out-of-Africa migration of modern humans across the mouth of the Red Sea? Fourth, assuming temporal stability of current populations, what are the estimated ages of Ethiopian populations relative to other African groups?
Link to press release. Link the supplemental data.

The authors reiterate that modern humans left Africa 50-70kya, a hypothesis that seems to me pretty much dead in the light of recent archaeological evidence.

The lack of antiquity in the Ethiopian population, even in only the African component thereof argues against that population being ancestral to modern humans. Note that if the Out-of-East Africa hypothesis is correct, then skulls like Omo I represent ancestral modern humans and they are followed much later by modern humans anywhere else. So, while anatomical modernity may have emerged in East Africa --or maybe not; let's not forget that we have early modern skulls from the region in part because of the excellent preservation conditions and excess of scholarly interest-- there is no evidence that they spread from there.

I have little doubt that my own theory about substantial back-migration of Eurasians into Africa will eventually win the day. Of course, I am not referring to the recent (in the last 3,000 years) admixture with West Eurasians that the Ethiopian population has undergone, but rather to the more ancient migration that was probably associated with Y-haplogroup DE-YAP.

The fact that the African component of diverse African populations is more closely related to West than to East Eurasians is one piece of evidence among many for that scenario. Hopefully, it can be tested soon using whole genome data which may have enough density to detect much older admixture events.

UPDATE I: Since the dates in the paper are based on ROLLOFF, a piece of software that is not publicly available more than a year after its announcement, and which contradicts other software released by the same authors, I will take the Queen of Sheba stories circulated in the media with a huge grain of salt.

The American Journal of Human Genetics, 21 June 2012 doi:10.1016/j.ajhg.2012.05.015

Ethiopian Genetic Diversity Reveals Linguistic Stratification and Complex Influences on the Ethiopian Gene Pool

Luca Pagani et al.

Humans and their ancestors have traversed the Ethiopian landscape for millions of years, and present-day Ethiopians show great cultural, linguistic, and historical diversity, which makes them essential for understanding African variability and human origins. We genotyped 235 individuals from ten Ethiopian and two neighboring (South Sudanese and Somali) populations on an Illumina Omni 1M chip. Genotypes were compared with published data from several African and non-African populations. Principal-component and STRUCTURE-like analyses confirmed substantial genetic diversity both within and between populations, and revealed a match between genetic data and linguistic affiliation. Using comparisons with African and non-African reference samples in 40-SNP genomic windows, we identified “African” and “non-African” haplotypic components for each Ethiopian individual. The non-African component, which includes the SLC24A5 allele associated with light skin pigmentation in Europeans, may represent gene flow into Africa, which we estimate to have occurred ∼3 thousand years ago (kya). The African component was found to be more similar to populations inhabiting the Levant rather than the Arabian Peninsula, but the principal route for the expansion out of Africa ∼60 kya remains unresolved. Linkage-disequilibrium decay with genomic distance was less rapid in both the whole genome and the African component than in southern African samples, suggesting a less ancient history for Ethiopian populations.

Link

October 05, 2011

Y-chromosomes of Marsh Arabs

What do the Marsh Arabs have to do with ancient Sumer? Nothing that can be determined on the basis of this data. There are plenty of ancient Sumerian skulls, so how about we study them directly?

As far as I can see, the only link between Marsh Arabs and Sumerians presented in this paper comes from dating Y-STR variation of their major J1-Page08 group using the evolutionary mutation rate, with a divergence time of 4.5 +/- 2.6 ky. Even if that mutation rate was correct (it is not) and the assumptions on which the confidence interval are based were exhaustive (they are not), we still have +/- 2.6 ky leeway to deal with, which spans not only the Sumerians but plenty more besides.

Not to mention that the evolutionary mutation rate is wrongly applied to every case under the sun, and that Y-STR based age estimation in general has been conclusively shown to be a rather futile exercise.

Nonetheless, the paper does have value in demonstrating the paucity of J2 and R1 in the Marsh Arabs compared to the more cosmopolitan general Iraqi population:
Different from the Iraqi control sample, the Marsh Arab gene pool displays a very scarce input from the northern Middle East (Hgs J2-M172 and derivatives, G-M201 and E-M123), virtually lacks western Eurasian (Hgs R1-M17, R1-M412 and R1-L23) and sub-Saharan African (Hg E-M2) contributions.
Rather than "Sumerian", it seems that the Marsh Arabs have rather preserved a more pristine Semitic patrilineal gene pool compared to the cosmopolitan Iraqi samples that have absorbed pre-Arab and pre-Semitic population elements.


BMC Evolutionary Biology 2011, 11:288doi:10.1186/1471-2148-11-288

In search of the genetic footprints of Sumerians: a survey of Y-chromosome and mtDNA variation in the Marsh Arabs of Iraq.

Nadia Al-Zahery et al.

Abstract (provisional)

Background
For millennia, the southern part of the Mesopotamia has been a wetland region generated by the Tigris and Euphrates rivers before flowing into the Gulf. This area has been occupied by human communities since ancient times and the present-day inhabitants, the Marsh Arabs, are considered the population with the strongest link to ancient Sumerians. Popular tradition, however, considers the Marsh Arabs as a foreign group, of unknown origin, which arrived in the marshlands when the rearing of water buffalo was introduced to the region.

Results
To shed some light on the paternal and maternal origin of this population, Y chromosome and mitochondrial DNA (mtDNA) variation was surveyed in 143 Marsh Arabs and in a large sample of Iraqi controls. Analyses of the haplogroups and sub-haplogroups observed in the Marsh Arabs revealed a prevalent autochthonous Middle Eastern component for both male and female gene pools, with weak South-West Asian and African contributions, more evident in mtDNA. A higher male than female homogeneity is characteristic of the Marsh Arab gene pool, likely due to a strong male genetic drift determined by socio-cultural factors (patrilocality, polygamy, unequal male and female migration rates).

Conclusions
Evidence of genetic stratification ascribable to the Sumerian development was provided by the Y-chromosome data where the J1-Page08 branch reveals a local expansion, almost contemporary with the Sumerian City State period that characterized Southern Mesopotamia. On the other hand, a more ancient background shared with to Northern Mesopotamia is revealed by the less represented Y-chromosome lineage J1-M267*. Overall our results indicate that the introduction of water buffalo breeding and rice farming, most likely from the Indian sub-continent, only marginally affected the gene pool of autochthonous people of the region. Furthermore, a prevalent Middle Eastern ancestry of the modern population of the marshes of southern Iraq implies that if the Marsh Arabs are descendants of the ancient Sumerians, also the Sumerians were most likely autochthonous and not of Indian or South Asian ancestry.

Link

May 24, 2011

The reality of the Altaic language family

Personally I'm not surprised by this; my own look at genomic data has identified an "Altaic" component which peaks at the Turkic Yakut and Tungusic Evenk, and is shared by every Turkic, Mongolic, and Tungusic population available to me. The same component also occurs to some extent among all the Japanese (5) and Korean (4) members of the Dodecad Project, while it is lacking in all the Chinese ones (8).


Of particular interest is the degree of CCM between Indo-European and Semitic languages (Tables 2 and 3). In many of the most geographically distant languages these are less than 10; by comparison, among Semitic languages the are all greater than 20. This seems to be quite in agreement with the idea that Semitic is a Bronze Age language family, Indo-European a Neolithic one.

This impression is strengthened by the fact that CCM between reconstructed proto-languages (e.g. Proto-Iranian and Proto-Slavic = 20) are much higher. Since these proto-languages are a few thousand years closer to the root of PIE than present-day languages, and differences between them are similar to those of Semitic languages, the notion that PIE is a few thousand years older than Proto-Semitic seems quite consistent with the evidence.

Journal of Language Relationship • Вопросы языкового родства • 3 (2010) • Pp. 117–126 • © Turchin P., Peiros I., Gell-Mann M., 2010

Analyzing genetic connections between languages by matching consonant classes

Peter Turchin (University of Connecticut)
Ilia Peiros (Santa Fe Institute)
Murray Gell-Mann (Santa Fe Institute)

The idea that the Turkic, Mongolian, Tungusic, Korean, and Japanese languages are genetically related (the “Altaic hypothesis”) remains controversial within the linguistic community. In an effort to resolve such controversies, we propose a simple approach to analyzing genetic connections between languages. The Consonant Class Matching (CCM) method uses strict phonological identification and permits no changes in meanings. This allows us to estimate the probability that the observed similarities between a pair (or more) of languages occurred by chance alone. The CCM procedure yields reliable statistical inferences about historical connections between languages: it classifies languages correctly for well-known families (Indo-European and Semitic) and does not appear to yield false positives. The quantitative patterns of similarity that we document for languages within the Altaic family are similar to those in the non-controversial Indo-European family. Thus, if the Indo-European family is accepted as real, the same conclusion should also apply to the Altaic family.

Link (pdf)

May 19, 2011

Nicholls and Ryder: Semitic 4.4-5.1 thousand years before present

The same authors dated Proto-Indo-European at 8.4ky, in agreement with the work of Gray and Atkinson. In the current paper they re-analyze the data of Kitchen et al. (2009) for Semitic languages, and their estimate is somewhat younger than 5,750 years of that paper. All in all, it's good to see different researchers using different techniques but coming up with similar solutions.

It is increasingly clear that while the Proto-Indo-Europeans originated in the Neolithic Near East, the Proto-Semites followed them by about three thousand years. In the latter case there is also a Y-chromosome marker (J-P58) with an apparent age in impeccable agreement with the linguistic evidence, now that the genealogical-"evolutionary" mutation wars seem to have been won.

This also brings into focus the weakness of the argument that Anthony (2007) (p. 76) brings to the table by hypothesizing that the first farmers of northern Syria were Afro-Asiatic speakers like the Semites of the Near Eastern lowlands. Semites come into the picture 5,000 years after the onset of the Neolithic, and 3,000 years after the Proto-Indo-Europeans. Their relationship with Afroasiatic speakers of Africa make it quite likely that they lived in the south, probably in Arabia, and certainly not in eastern Anatolia or northern Syria.

Indeed, the recent discovery that haplogroup J1*(xP58) is associated with Northeast Caucasian languages, together with the absence or paucity of J1 in most African Afroasiatic speakers suggests to me that the J-P58 Proto-Semites may be the result of the transfer of an African language on a basically West Asian population. Such a scenario might also explain some of the -incorrectly quantified, but nonetheless existent- African genetic components in both Jews and Arabs, as well as the pastoralist/dry-climate J1 associations.

Proceedings of the 26th International Workshop on Statistical Modelling.

Phylogenetic models for Semitic vocabulary.
Geoff K Nicholls and Robin J. Ryder

Abstract: Kitchen et al. (2009) analyze a data set of lexical trait data for twenty five Semitic languages, including ancient languages Hebrew, Aramaic and Akkadian, modern South Arabian and Arabic languages and fifteen ethiosemitic languages. They estimate a phylogenetic tree for the diversification of lexical traits using tree and trait models and methods set up for genetic sequence data. We reanalyze the data in a homplasy-free model for lexical trait data. We use a prior on phylogenies which is non-informative with respect to some of the key scientific hypotheses (concerning topology and root time). Our results are in broad agreement with those of Kitchen et al. (2009), though our 95% HPD for the root of the Semitic tree (the branching of Akkadian) is [4400, 5100]BP and we place Moroccan and Ogaden Arabic in the Modern South Arabian Group.

January 22, 2011

Near Eastern Grape domestication

Kambiz links to an interesting paper on grape domestication. From the paper:
Archaeological evidence suggests that grape domestication took place in the South Caucasus between the Caspian and Black Seas and that cultivated vinifera then spread south to the western side of the Fertile Crescent, the Jordan Valley, and Egypt by 5,000 y ago (1, 21). Our analyses of relatedness between vinifera and sylvestris populations are consistent with archaeological data and support a geographical origin of grape domestication in the Near East (Fig. 4 and Table 1).
The genetic confirmation of the archaeological inference is particularly interesting, since "wine"is part of the Proto-Indo-European lexicon, and has related forms in both Kartvelian (South Caucasian) and Semitic languages. The Transcaucasus seems a quite good place to seek early contact between these three language families. Interestingly, the area between the Black Sea and Caspian is also where genetic analysis of Indo-Aryan origins has brought me.


PNAS doi: 10.1073/pnas.1009363108

Genetic structure and domestication history of the grape

Sean Myles et al.

The grape is one of the earliest domesticated fruit crops and, since antiquity, it has been widely cultivated and prized for its fruit and wine. Here, we characterize genome-wide patterns of genetic variation in over 1,000 samples of the domesticated grape, Vitis vinifera subsp. vinifera, and its wild relative, V. vinifera subsp. sylvestris from the US Department of Agriculture grape germplasm collection. We find support for a Near East origin of vinifera and present evidence of introgression from local sylvestris as the grape moved into Europe. High levels of genetic diversity and rapid linkage disequilibrium (LD) decay have been maintained in vinifera, which is consistent with a weak domestication bottleneck followed by thousands of years of widespread vegetative propagation. The considerable genetic diversity within vinifera, however, is contained within a complex network of close pedigree relationships that has been generated by crosses among elite cultivars. We show that first-degree relationships are rare between wine and table grapes and among grapes from geographically distant regions. Our results suggest that although substantial genetic diversity has been maintained in the grape subsequent to domestication, there has been a limited exploration of this diversity. We propose that the adoption of vegetative propagation was a double-edged sword: Although it provided a benefit by ensuring true breeding cultivars, it also discouraged the generation of unique cultivars through crosses. The grape currently faces severe pathogen pressures, and the long-term sustainability of the grape and wine industries will rely on the exploitation of the grape's tremendous natural genetic diversity.

Link

September 08, 2010

ASHG 2010 abstracts

The 2010 meeting of the American Society of Human Genetics is in November. Here are some interesting abstracts that caught my eye:

It's nice to finally see a genomic study on the Greek population.
P. Paschou et al. Evaluation of the HapMap dataset as reference for the Greek population.
The HapMap project has provided a unique tool for the analysis of human genetic variation, providing reference information for allele frequency and genotype distributions as well as linkage disequilibrium patterns of Single Nucleotide Polymorphisms (SNPs) across the entire genome. The latest release of HapMap phase 3 data provides genotypes for millions of SNPs in 11 populations from around the world, with Europe being represented by the CEU (originating from Northwestern Europe) and the TSI populations (Tuscan Italians from Southern Europe). Although initial studies support the fact that the CEU can be used as reference for the selection of tagging SNPs in other European populations, a critical step in the design of genetic association studies, this hypothesis has not been extensively studied across Europe and in particular in Southern Europe. We set out to explore the extent to which the HapMap populations can be used as reference for a previously unstudied population of South-Eastern Europe, the Greek population. To do so we studied genomic variation in 1,813 SNPs, genotyped by our group in 56 individuals of Greek origin, and compared them to the CEU and TSI genotypes (1,813 SNPs from the CEU HapMap dataset and 1,205 from the TSI dataset). The studied SNPs are spread over 13 autosomal chromosomes and 26 regions, ranging in size from 120Kb to more than 4Mb. Genotype, allele frequency, and pairwise LD measures were compared across all three populations. PCA was used in order to identify those markers that are responsible for the observed inter-sample variance. Tagging SNPs were selected in the CEU and TSI samples and their transferability to the Greek population was tested, using both the r2 metric as well as the efficiency of genotype imputation of the non-selected SNPs. Our results demonstrate that, although the CEU population can to some extent be used as reference for the Greek population, it is preferable to use as reference a European population of closer genetic ancestry, like the TSI. These results are applicable in medical genetics, in order to inform the design of genetic association studies, as well as in studies of evolutionary relationships of Southern European populations.
One of the great problems of Eurasian anthropology is whether the Uralic populations are simply variable admixtures of Caucasoids and Mongoloids or they contain a tertium quid in the form of a Proto-Uralic element. The latter need not be distinct from the other two, as it can also be an old or stabilized blend of the two major Eurasian races that later admixed with more recent groups on either side. The abstract does not seem promising in this respect, i.e., in identifying a common core of ancestry among Uralic speakers in addition to their variable east-west admixture, but it would be nice to see if anything like that exists in the paper.

K. Tambets et al. Haploid and autosomal variation within a linguistic continuum of the Uralic-speaking people of Eurasia.
For about last two decades the examination of uniparentally inherited genetic marker systems revealing the variation embedded in mtDNA and Y chromosome has been the main tool in the studies of human genetic origins. Within few recent years the analysis of the genome-wide SNP data of individuals from different populations has started to give promising new insights in the field of human population genetics. The uniparentally inherited markers have shown slightly different demographic scenarios for the maternal and paternal lineages of North Eurasian, particularly of European Uralic-speaking populations. The geographical location of a population has evidently been the most important component that dictates the proportion of western and eastern mtDNA types in the gene pool of Uralic-speakers. Thus, the palette of maternal lineages of the Uralic-speakers resembles that of their geographically close European or Western Siberian Indo-European and/or Altaic-speaking neighbours, respectively. At the same time, the most frequent North Eurasian Y chromosome type N1c, that is also a common link between almost all Uralic-speakers, is with few exceptions rare, if present at all, among Indo-European-speakers of Western and Southern Europe. Here we combine genome-wide high density SNP data (650 000 SNPs, Illumina) with uniparentally inherited mtDNA and Y-chromosome variation of 16 Uralic-speaking populations to assess their place on the genetic landscape of North Eurasia. By the use of principal component and structure-like analysis on the autosomal data we show that the proportions of western and eastern ancestry components among the Uralic-speakers are determined mostly by geographical factors. The westernmost populations from Europe, both Uralic- and Indo-European speakers, are similar in their pattern of ancestry components and show low levels (less than 10%) of the eastern component. Conversely, the eastern ancestry component is dominant (60-70%) in the gene pool of the Siberian Uralic-speakers. In general, the genome-wide analyses corroborate the results of mtDNA analysis and do not reflect the common genetic characteristics between western and eastern Uralic-speakers at the level seen in case of N1c. Interestingly, among Saami from North Europe, who are often considered as „outliers“ in genetic studies, the dominant western component is accompanied by 30% of eastern component making them more similar to Volga-Uralic populations than to their closest neighbours.



This seems to validate my thoughts on relics and their importance in age estimation.

U. A. Perego et al. The Initial Peopling Of The Americas: An Ever-Growing Number Of Founding Mitochondrial Genomes From Beringia
Genetic evidence based on mitochondrial DNA (mtDNA) has recently revealed the existence of additional founding lineages that have contributed to the first peopling of America’s double-continent in addition to the more popular five Native American haplogroups (A2, B2, C1, D1 and X2a), and has demonstrated as well the need for additional sampling and analysis to be performed for some of the already known but poorly characterized lineages. One paradigmatic example is represented by the pan-American haplogroup C1. Two of its sub-branches (C1b and C1c) harbor ages and geographical distributions that are indicative of an early arrival from Beringia about 15-17,000 years ago, concomitantly with the other currently accepted Paleo-Indian founders. However, the estimated age of C1d - the third Native American subset of C1 - is only 8-10,000 years, which is suggestive of a much later entry and spread in the Americas. In this study, we shed light on the origin of this enigmatic Native American branch of C1 by completely sequencing a large number of C1d mitochondrial genomes from a wide range of geographically diverse, mixed and indigenous American populations. The revised phylogeny shows that the age previously reported for C1d was heavily underestimated and indicate that C1d is ancient enough to be among the founding Paleo-Indian mtDNA lineages. Moreover, our results reveal that there were two C1d founder genomes for Paleo-Indians that most likely arose early (~16kya), either in the dynamic Beringian gene pool, or at a very initial stage of the Paleo-Indian southward migration. This brings the recognized maternal founding lineages of Native Americans to the unexpected number of 15, and indicates that the overall number of Beringian or Asian founder mitochondrial genomes will probably continue to increase as more Native American haplogroups reach the same level of phylogenetic resolution as we obtained here for C1d. Additionally, we have confirmed a nearly identical geographic distribution pattern for haplogroup C1d when comparing samples collected in the general mixed population with those from native tribal groups, as it was also reported previously for haplogroups X2a and D4h3. This substantiates the validity of searching large public mtDNA databases (such as the one available through the Sorenson Molecular Genealogy Foundation, www.SMGF.org) for novel founder candidates able to reveal unknown details concerning the ancient human history of the Americas.

Another interesting abstract. I've written before about the association of Y-chromosome haplogroups with the spread of Semitic speakers and the agreement with language phylogenetics.

N. Al-Zahery et al. The male gene pool of the contemporary Mesopotamia marsh population supports their Semitic origin.
The origin of the modern Mesopotamia marsh people, which are locally called “Ma’dan” or “Marsh’s Arabs”, is a question of great interest. Based on their life-style (living in reed houses, grazing of water buffalo and other aspects) and local archaeological sites, many historians and archaeologists believe they may have Sumerian ancestry. Although little is known about the origin of Sumerians themselves, two main hypotheses have been advanced in this regard. According to the first, Sumerians were a group of populations which migrated from the “South East” following a seashore route through the Arabian Gulf, and settled down in the southern marshes of Iraq. According to the second, the advancement of the Sumerian civilization is the result of migration from the mountainous area of Anatolia to the southern marshes of Iraq where they settled, adsorbing previous populations. In order to shed some light on the genetic origin of the Mesopotamia marsh population, we investigated the male gene pool of 145 DNA samples of modern Mesopotamia people, still living in marshes in the south of Iraq. The analyses of Single Nucleotide Polymorphisms (SNPs) and Short Tandem Repeats (STRs) of the paternally transmitted Male Specific region of the Y chromosome (MSY) revealed that more than 80% of marsh Y chromosomes belong to (Hg) J1-M267, the autochthonous haplogroup of Middle Eastern/Semitic speakers with possible recent expansion and/or founder effect reflected by the reduced STRs variability. In particular, 90% of them were assigned to the J1e-M267-PAGE08 sub-haplogroup, which is the predominant Y chromosome lineage among Middle Eastern Arab populations (Yemen, Qatar, UAE, and Levant). Thus, these findings testify, at least from the paternal side, a strong Semitic Arabian component in the contemporary Mesopotamia marshes population, whereas no clear Anatolian and/or South Asian genetic evidence has been detected.
The finding of haplogroup I in China is surprising, as I is not generally found that far away from Europe. It would be interesting to see what the actual haplotypes are.
Y. Lu et al. Western Eurasian Y chromosomes found in the Chinese Salar ethnic group
Salar is a small Western-Turkish-speaking population living mostly in Qinghai province of China. The most similar languages to Salar are all far in Turkmenistan. Historical records suggested that they may be descendants of the Turkic nomadic tribes in Central Asia. In this study, 141 Salar Y chromosomes were analyzed for 39 SNP and 14 STR markers to investigate the potential imprints of their western ancestors. The most frequent haplogroup (hg) in this population sample is Hg R, comprising 40% of all Y chromosomes. Most of these Hg R samples belong to R1a1 (M17), which distributes in a wide geographic region including South Asia, East Europe, Central Asia, and South Siberia. Other four Western Eurasian haplogroups (G-2%, H-5%, I-3%, J-3%) were also found in Salar Y chromosome gene pool. These paternal lineages of Salar are absent in their East Asian neighbors but frequent in Central Asia. Y-STR-based analyses also grouped Salar to Central Asians. On the other side, Salar also has low frequencies of the East Asian specific Hg D and Hg O, suggesting possible gene flow from their neighboring populations. This Y chromosome study demonstrated that Salar well keeps the Western Eurasian paternal lineages of their Central Asian ancestors although they may have migrated to Central China for about 800 years.

I wish that more "people pairs" would be studied this way, as it would give us some good insight of how migration affects gene pools (allele frequency changes, founder effects, possible social selection etc.)

M. Davis et al. Ancient and recent demographic events influence mitochondrial DNA diversity in an immigrant Basque population
The Basques are an ancient people, considered by many anthropologists to represent the oldest extant European population. Because of this, they have been the subject of numerous sociological and biological investigations. The Basque Diaspora, a relatively recent demographic expansion of the Basque population, has until now been overlooked in genetic studies. Samples were taken from 53 individuals with Basque ancestry in Boise, Idaho, and the mitochondrial DNA (mtDNA) sequence variation of the first and second hypervariable regions were determined. Thirty-six mtDNA haplotypes were detected in the sample. Comparing the genetic diversity in the Idaho sample with other Basque populations, signatures of founder effects were observed, consistent with both the recent and ancient history of Basque mitochondrial lineages. There has been a marked alteration of haplogroup frequency and diversity, and there is a slight reduction in other measures of diversity in the NW Basque population compared to the native Basque population. We have found a relatively high percentage of the Cambridge Reference Sequence (rCRS) haplotype for hypervariable regions I and II, which is absent in previous studies of Basque mtDNA, and rare in other Spanish populations. The amount of nucleotide diversity is consistent with a sample that is predominantly haplogroup H, which is especially common in the Basque regions of Europe, due to ancient migrations and expansions out of glacial refugia. This is the first report of mtDNA diversity in an immigrant Basque population, and we find that the diversity in NW Basques can be explained by the recent history of migration, as well as the phylogeography and diversity of the major European haplogroups.


W. S. Watkins et al. Admixture in New World populations: an analysis of Y-chromosome, mtDNA, and genome-wide microarray data
The first major interaction between Native Americans and Europeans is documented historically and occurred less than 550 years ago. This recent time frame provides an excellent opportunity to investigate the effects of admixture between two populations that were previously separated for hundreds of generations. To characterize European admixture in Native American populations, we sampled and analyzed a group of isolated Totonac agriculturists from tropical Mexico near Veracruz and a group of native Bolivians predominantly from the mountainous region near La Paz, Boliva. Mitochondrial sequencing of HVS1 showed that all samples had pre-Columbian mtDNA haplogroups (A, B, C, and D). Using a panel of 48 STRs or 12 Y-chromosome SNPs, Totonac Y-chromosomes lineages were all assigned to the pre-Columbian haplogroup Q1a3a, and Bolivian Y-chromosome lineages were assigned to haplogroups Q1a3a, R1, and J2. Haplogroups R1 and J2 are common in European populations. Principal components analysis (PCA) using >800K autosomal SNPs typed in 24 Totonacs and 23 Bolivians showed that all Totonacs and 14 Bolivians clustered distinctly from Eurasian individuals. Nine Bolivians, however, were positioned between the New World and European PCA clusters. Admixture analysis showed that these nine samples had 21 - 33% European admixture using a European reference population. All three observed Y-chromosome haplogroups, including the well-studied pre-Columbian haplogroup Q1a3a, occurred in the admixed individuals. Two of the nine admixed individuals had pre-Columbian mtDNA and Y-chromosome haplogroups but 21-23% European ancestry. This result demonstrates that Y-chromosome and mtDNA haplogroups are only partial indicators of an individual’s complete ancestry.

Readers of the blog know that I don't agree with the scenario presented in the followin abstract. The serial founder effect idea is used by geneticists to explain the overall reduced genetic diversity of our species (that we appear to be young, in evolutionary terms). Personally, I don't see how a smart, expanding species that all of the sudden had access to the resources of the landmass of Eurasia went through these extreme bottlenecks.
I think that the alternative of a larger human population, genetic diversity reduced across the species by ongoing climate- and culture-mediated selection, and admixture within Africa itself -where a particular expanding H. sapiens group must've co-existed with pre-existed hominids, anatomically modern or not- has merit.
J. Long et al. Evidence for archaic admixture in contemporary non-African human populations
Analyses of large-scale genetic data sets show evidence for a series of founder effects that occurred as modern humans left Africa and settled the rest of the world. Nonetheless, research on modern humans has not ruled out the possibility that other processes, such as local gene flow, or mixing between archaic and modern humans, have also contributed to modern human diversity. Recent analyses of the Neanderthal genome make archaic admixture a salient issue because they show evidence for mixing between Neanderthals and out-of-Africa migrants. The present study examines evidence for archaic admixture in genotypes for 619 microsatellite loci collected from over 2,000 individuals from 100 human populations. We obtained these data from the Marshfield Clinic collection. The populations analyzed represent all inhabited continents of the world. In our analysis, we formulate the serial founder effects (SFE) model as a special case of a phylogenetic model promoted by Cavalli-Sforza and his associates. In this light, the SFE process makes four predictions: 1) A tree of descent according to the pattern of fissions. 2) The root of the tree lies in Africa. 3) The length of each branch is proportional to ratio of evolutionary time to effective population size. 4) The gene identity between all pairs of populations that share the same most recent common ancestor is equal in expectation. Using hypothesis tests based on generalized hierarchical statistical models, we find good agreement between the SFE predictions and diversity within and between African populations, and we find good agreement between the SFE predictions and diversity between non-African populations. However, there is more diversity within the non-African populations than the SRE model can account for. This makes for greater genetic distance between Africans and non-Africans than otherwise expected. How and where did the non-Africans obtain this diversity? A simple explanation for the finding is that the earliest migrants out-of-Africa mixed with an archaic population such as Neanderthals prior to their expansion throughout Europe and Asia. Coalescent based computer simulations of the SFE model with mixing support our interpretation. The time and place that we detect mixing coincides perfectly with that detected in a recent examination of Neanderthal genome sequences. Our study shows that genomic diversity in modern humans still reflects ancient events and processes.

C. Flores et al. Using EuroAIMs to measure admixture proportions in atypical European populations: the case of Canary Islanders
Using ancestry informative markers (AIMs) allows reducing the number of makers needed for population stratification adjustments in association studies. As few as 100 AIMs are sufficient to adjust for the largest European axis of differentiation (i.e. EuroAIMs). However, their use for ancestry inference and adjustment in association studies in atypical European populations such as the Canary Islanders, a recently African-admixed population from Spain, needs to be addressed. We aimed to explore whether EuroAIMs were suitable both for the inference of Spanish and Northwest African admixture proportions and for ancestry adjustments in association studies including samples from Canary Islanders. We analyzed samples from Canary Islanders, mainland Spanish (IBE) and Northwest Africans (NWA) for 93 EuroAIMs and compared the data with CEU and YRI from HapMap, Basques and Mozabite from HGDP, as well as from previously analyzed European samples. The major genetic difference was observed between NWA and all European populations, preserving the northwest-to-southeast differentiation of European populations in the second axis. Analyses revealed that Canary Islanders were intermediate between IBE and NWA, and that direct sub-Saharan African influences were negligible. Assessment of individual admixtures without prior population information clearly identified two subpopulations corresponding to NWA and IBE, while Canary Islanders were admixed with an average of 17.4% Northwest African contribution varying largely among individuals (range 0-95.7%). As few as 23 EuroAIMs correctly estimated population membership to IBE and NWA, while 69 EuroAIMs were required to accurately estimate individual admixture proportions in Canary Islanders. Ancestry estimates based on a subset of 69 EuroAIMs also controlled significant allele frequency differences between IBE and Canary Islanders. These data suggest that a handful of EuroAIMs would be useful to control false-positives in association studies performed in Spanish populations. Supported by FUNCIS 23/07 and grants from the Spanish Ministry of Science and Innovation PI081383 and EMER07/001 to CF.
As I have I mentioned before, the Maasai (and many other east Africans in various degrees) are intermediate between Negroids and Caucasoids, and hence admixture estimates considering Yoruba Nigerians would tend to underestimate the African element. It's important to remember that extant Africans are not uniform, ranging from Caucasoids to Negroids, Pygmies, and Khoi-San, with multiple identifiable clusters within the major Negroid group itself, and all sorts of between-group gene flow in a regional basis. It is always useful (as is the case e.g., with African Americans) to both use historical knowledge about population sources, and also to validate historical narratives with the genetic evidence.
R. L. Raaum et al. Autosomal African admixture in Yemeni populations.
Approximately 30% of mtDNA lineages in South Arabian samples are African L haplotypes, whose origin has usually been attributed to migration and assimilation of African females into the Arabian population over approximately the last 2,500 years. Few In contrast, few Y chromosome lineages of clear recent sub-Saharan African origin have been found in Southern Arabian populations. This bias in maternal and paternal lineages is in accord with historical accounts of the female bias in the Middle Eastern slave trade. In order to evaluate autosomal African ancestry, we collected high-resolution SNP genotype data from a geographically representative set of 62 Yemenis selected from a collection of 552 samples acquired in the Spring of 2007. The ancestry of chromosomal segments in the Yemeni population was estimated using a haplotype-based local ancestry estimation method, HAPMIX. The HAPMIX method is based on a two way admixture model that requires two phased reference populations; we used the HapMap Yoruba in Ibadan, Nigeria (YRI), Luhya in Webuye, Kenya (LWK), Maasai in Kinyawa, Kenya (MKK), and CEPH US residents with ancestry from northern and western Europe (CEU) samples. The three African reference populations include two Bantu-speaking groups (YRI and LWK) and one Nilotic-speaking group (MKK). We estimated local ancestry in the Yemeni sample with all three European-African reference population combinations (CEU-YRI, CEU-LWK, CEU-MKK). The correlations among African ancestry calculated using all three reference population combinations are high (r > 0.98 in all pairwise correlations). Furthermore, there is no significant difference between the average proportion of African ancestry in Yemenis calculated using either of the two Bantu-speaking reference populations: CEU-YRI (mean 0.062, sd 0.044) and CEU-LWK (mean 0.076, sd 0.049) (p=0.13, two-tailed Welch two sample t-test). However, the average African ancestry calculated using the Maasai reference population (CEU-MKK, mean 0.148, sd 0.060) is significantly greater from that calculated using either the Yoruba or Luhya reference populations (p less than 0.0001 in both comparison, two-tailed Welch two sample t-test). These data suggest that the source population for the African ancestry of the Yemeni population is more similar to the contemporary Maasai population than either the Luhya or Yoruba.
The next abstract seems fun; it's always nice to see something that isn't like everything that came before it.
T. Rzeszutek et al. Music as a novel marker in the study of prehistoric human migrations.
The study of prehistoric human population history is often fraught with controversy owing to incongruent evidence among various markers of present-day genetic and cultural diversity. While archaeological evidence can be used to calibrate the conclusions drawn from present-day diversity, the fickle nature of the fossil record leaves some migration histories unresolved. Our work analyzes the potential of music - in particular, vocal music - to serve as novel migration marker, bolstering established migration work and shedding light on regions of the world whose settlement history is contested. One such migration is the recent expansion of Austronesian-speaking peoples across the Pacific within the last 6000 years. The dominant hypothesis posits a recent origin in Taiwan, with a rapid movement southwards and eastwards to populate Polynesia during the following 3500 years. While this model is strongly supported by both archaeological evidence and the present-day distribution of linguistic diversity, our goal was to analyze whether music could serve as a novel line of evidence in the study of Pacific prehistory. A critical concern regarding any migration marker is its time depth. In order to examine this for music, we analyzed correlations between musical diversity and mitochondrial-DNA diversity in 9 Taiwanese aboriginal tribes for which both types of data were available. A sample of 226 choral songs was analyzed using 39 binary characters representing significant structural features of music (e.g., rhythm, interval size, melodic contour, etc.). The musical samples were restricted to ritual musics, which constitute the most conservative (i.e., slowly changing) component of a culture’s repertoire. Mantel tests showed a significant correlation between musical distance and genetic distance among these 9 tribes, suggesting that music may have a time depth comparable to widely-used genetic markers like mitochondrial DNA. This work demonstrates that music has the potential to enrich the conclusions drawn from other markers, and establishes methods for employing it as a tool in the study of prehistoric human movements throughout the world. At the same time, we want to capitalize on music’s own unique dynamics of change over time and place, particularly its capacity for admixture. In other words, music might not only be able to support the narratives told by other migration markers but shed new light on the histories of population movement and cultural contact.


The bolded part in the following abstract makes sense, as it indicates (i) the distinctiveness of Ashkenazi Jews compared to CEU Europeans, and (ii) the fairly recent widespread formation of admixed individuals (in the last couple of generations) which generated individuals that are 1/4 1/2 and 3/4 AJ genomically.

V. Vacic et al., Admixture in Ashkenazi Jewish cohorts and implications for association studies.
Studies of complex genetic disorders may benefit from focusing on population isolates, such as Ashkenazi Jews (AJ). However, in order to truly exploit the advantages of reduced genetic diversity the self-declared AJ ancestry of study participants should be independently confirmed with available genetic data. We investigate whether the AJ cohorts display genetic heterogeneity, such as e.g. different rate of admixing in cases and controls, which could potentially confound disease association studies. We applied principal component analysis (PCA) to AJ cohorts ascertained in Israel and the US East Coast with the goal of characterizing population structure. As described previously, when compared to the HapMap samples with CEU, YRI and CHB/JPT ancestry, virtually all AJ samples cluster with the CEU. Similar analysis done on CEU and Jewish HapMap samples from Ashkenazi, Sephardic and Middle Eastern Jewish communities revealed that 97.8% of AJ samples cluster along the AJ-CEU axis, with modes at AJ and CEU cluster centers and at approximately quartile distances between them. We postulate that these groups correspond to 100-0, 75-25, 50-50, 25-75, and 0-100% AJ-CEU admixtures. Notably, only 91.7% of self-reported AJ individuals fall into the reference JHapMap panel AJ cluster, with 1.6, 3.3, 0.5 and 0.7% in the admixed modes ordered by decreasing fraction of AJ ancestry. We also observe admixing with the non-AJ Jewish communities: 0.7% of samples fall within the non-AJ clusters and 1.4% at a subgroup approximately halfway between the AJ and non-AJ cluster centers. In our dataset we found that when compared to the sample as a whole or only to controls, individuals with Crohn’s disease (CD) show significantly more admixing: 78.1, 3.1, 8.5, 2.0 and 0.9% in the 100, 75, 50, 25 and 0% AJ subgroups respectively. Also, CD samples show more admixing with non-AJ groups (2.8 and 1.0% in the 50-50 and 0-100 AJ-non-AJ subgroups). Isolates typically exhibit a greater amount of cryptic relatedness compared to outbred populations, which motivates an orthogonal method for verifying AJ ancestry based on identity-by-descent (IBD). The high background level of IBD within the Ashkenazi Jewish community can be used to estimate degree of AJ ancestry by averaging the IBD between a sample under study and the AJ individuals in the JHapMap panel. Our preliminary results show that this method recapitulates the high-level results from the PCA analysis and provides better resolution.

June 03, 2010

Two major groups of living Jews (Atzmon et al. 2010)

(Last Update Jun 10)

More on this paper as soon as I read it carefully. Nature News has an overview of the research. This study addresses my question about the extent of Southern vs. Central/European ancestry in Jews.

It is also entirely consistent with my theory that Diaspora Jews are to a certain extent descended from Italian-Balkan-Anatolian groups, among which they lived in Hellenistic-Roman-Late Antique times; my guess is that Middle-Eastern Jews form a distinct group in relation to European/Syrian ones because, unlike them, they had a smaller opportunity of absorbing Euranatolians, and their admixture -if any- came from linguistically (and probably genetically) related Semitic groups.

UPDATE I (Jun 4)

It is a bit frustrating how the authors did not limit themselves to HGDP but also included a wide variety of populations from POPRES, which they bundled together in a fairly arbitrary way:
Next, each of 2407 European subjects was assigned into one of 10 groups based on geographic region: South:Italy, Swiss-Italian; Southeast: Albania, Bosnia-Herzegovina, Bulgaria, Croatia, Greece, Kosovo, Macedonia, Romania, Serbia,Slovenia, Yugoslavia; Southwest: Portugal, Spain; East: CzechRepublic, Hungary; East-Southeast: Cyprus, Turkey; Central:Austria, Germany, Netherlands, Swiss-German; West: Belgium,France, Swiss-French, Switzerland; North: Denmark, Norway,Sweden; Northeast: Finland, Latvia, Poland, Russia, Ukraine;Northwest: Ireland, Scotland, UK.
I don't know exactly why Switzerland should be bundled with Belgium, while Austria with Netherlands, or that Finns would be bundled with Poles, or Albanians, Greeks, and Slavs would be bundled in a broad "Southeast" group. Anyway, the authors don't use this POPRES sample much in their actual paper, although they present some results in the supplement, so I won't dwell on it further.

UPDATE II (Jun 4)

On the left we have panel B from Figure 1 in the paper, which shows the first two principal components in a regional context. Capitalized labels represent Jewish groups. Note that Iranian and Iraqi Jews don't show a particular relationship to Arabs (Bedouins and Palestinians). It would be interesting to see if there is a relationship with Iraqis or Iranians, which might indicate whether admixture with these local groups is responsible for Iranian/Iraqi Jewish distinctiveness. As in previous studies, most other Jews, including Syrian Jews, are located between Europeans and Druze; the lack of non-Jewish Euranatolian populations is especially baffling.

I am particularly interested in the seemingly very close relationship between Greek, Turkish, and Italian Jews. The GRK-TUR relationship is not that puzzling, as these are mostly Ottoman Jews who found themselves on different sides of national borders, and we would not expect them to be any different. But, why would they be so similar to Italian Jews? Speaking of Greek Jews, how many of them are of Romaniote and how many of Sephardic extraction?

UPDATE III (Jun 4)

The STRUCTURE analysis is also quite interesting. Jews seem to lack appreciable levels of East Asian (orange) or Sub-Saharan African (yellow) admixture, or of Central/South Asian admixture (green). The lack of E/C/S Asian admixture is especially damning of the Khazar hypothesis.

We should probably not interpret the three main visible components ("European" blue, "Mozabite" purple, "Near Eastern" pink) as representing ancestral proportions of European, North African, and Near Eastern elements. For example, Mongoloids have some "purple" while it is unlikely that they have North African admixture; so, while purple has an obvious relationship to Mozabites, it is not a good fit for an ancestral population group. Its substantial presence in the Near East also precludes such an easy interpretation.

Nor can we easily infer the percentage of "European" and "Near Eastern" admixture in Jews. The "Pink" element seems to grade from prominence among Iranian Jews to insignificance among Basques, but what did the original European and Jewish groups look like? Depending on how close they were to the Basque and Iranian Jewish end of the gradient, quite different admixture proportions would arise.

In more mathematical terms, a gradient can be represented as a single variable x going from 0 to 1, e.g., pink/(pink+blue) in the STRUCTURE analysis or relative position between Basques and Druze in the PCA figure above. But x can be expressed in an infinite number of ways as a weighted summation of two other numbers between 0 and 1. If ancestral groups were exactly like Basques and Druze or they were exactly pure blue and pure pink, then we could arrive at exact ancestral proportions for living Jews, but unfortunately, unlike situations where clear-cut well-differentiated ancestral groups exist to act as yardsticks, this does not appear -as of yet- to be the case for intra-Caucasoid variation.

UPDATE IV (Jun 4)

From the paper:
Admixture with local populations, including Khazars and Slavs, may have occurred subsequently during the 1000 year (2nd millennium) history of the European Jews. Based on analysis of Y chromosomal polymorphisms, Hammer estimated that the rate might have been as high as 0.5% per generation or 12.5% cumulatively (a figure derived from Motulsky), although this calculation might have underestimated the influx of European Y chromosomes during the initial formation of European Jewry. Notably, up to 50% of Ashkenazi Jewish Y chromosomal haplogroups (E3b, G, J1, and Q) are of Middle Eastern origin,15 whereas the other prevalent haplogroups (J2, R1a1, R1b) may be representative of the early European admixture. The 7.5% prevalence of the R1a1 haplogroup among Ashkenazi Jews has been interpreted as a possible marker for Slavic or Khazar admixture because this haplogroup is very common among Ukrainians (where it was thought to have originated), Russians, and Sorbs, as well as among Central Asian populations, although the admixture may have occurred with Ukrainians, Poles, or Russians, rather than Khazars. In support of the ancestry observations reported in the current study, the major distinguishing feature between Ashkenazi and Middle Eastern Jewish Y chromosomes was the absence of European haplogroups in Middle Eastern Jewish populations.
I would not be so quick to assign haplogroups to European or Middle Eastern origin. For example, G seems to have originated in the Middle East, but it is quite plentiful in substantial parts of Europe. So, while its ultimate origins may be West Asian (it arose in a man who lived in West Asia thousands of years ago), its proximate origin may be European in some particular case.

As I have argued before, I doubt E3b (or E1b1b) was an original Jewish lineage, J2 probably represents Iranian/Euranatolian admixture in Jews, while J1 (or a subset thereof) has strong Semitic connotations.

UPDATE V (Jun 10):

Another paper by Behar et al. (2010) on the same topic.


AJHG doi:10.1016/j.ajhg.2010.04.015

Abraham's Children in the Genome Era: Major Jewish Diaspora Populations Comprise Distinct Genetic Clusters with Shared Middle Eastern Ancestry

Gil Atzmon et al.

Abstract

For more than a century, Jews and non-Jews alike have tried to define the relatedness of contemporary Jewish people. Previous genetic studies of blood group and serum markers suggested that Jewish groups had Middle Eastern origin with greater genetic similarity between paired Jewish populations. However, these and successor studies of monoallelic Y chromosomal and mitochondrial genetic markers did not resolve the issues of within and between-group Jewish genetic identity. Here, genome-wide analysis of seven Jewish groups (Iranian, Iraqi, Syrian, Italian, Turkish, Greek, and Ashkenazi) and comparison with non-Jewish groups demonstrated distinctive Jewish population clusters, each with shared Middle Eastern ancestry, proximity to contemporary Middle Eastern populations, and variable degrees of European and North African admixture. Two major groups were identified by principal component, phylogenetic, and identity by descent (IBD) analysis: Middle Eastern Jews and European/Syrian Jews. The IBD segment sharing and the proximity of European Jews to each other and to southern European populations suggested similar origins for European Jewry and refuted large-scale genetic contributions of Central and Eastern European and Slavic populations to the formation of Ashkenazi Jewry. Rapid decay of IBD in Ashkenazi Jewish genomes was consistent with a severe bottleneck followed by large expansion, such as occurred with the so-called demographic miracle of population expansion from 50,000 people at the beginning of the 15th century to 5,000,000 people at the beginning of the 19th century. Thus, this study demonstrates that European/Syrian and Middle Eastern Jews represent a series of geographical isolates or clusters woven together by shared IBD genetic threads.

Link

October 16, 2009

The emergence and dispersal of haplogroup J-P58 (aka J1e)

The paper uses the evolutionary mutation rate, which, as I have argued elsewhere overestimates time to most recent ancestor (TMRCA) by about a factor of 3. The evolutionary mutation rate is appropriate for haplogroups subject to strong genetic drift that have not grown to large numbers, but it is completely inappropriate under conditions of strong population growth.

To make things concrete, according to the model of drift-induced variance reduction proposed by Zhivotovsky, Underhill, and Feldman (2006), in 10,000 years (or 400 generations), J-P58 should have grown to the grand number of 200 men, or at least five orders of magnitude lower than the actual present-day haplogroup size. To account for the observed J-P58 size of millions of men, strong growth over time is needed, and with either the Z.U.F. (2006) analysis or my own, strong growth results in an accumulation of variance at close to the germline mutation rate.

With that said, all ages in this paper should be divided by a factor of 3. This is not only theoretically sound, but harmonizes better with other lines of evidence.

The paper studies Y-STR variance in several Middle Eastern populations. The lack of samples from the Caucasus does not allow us to infer the levels of Y-STR variance in that region. Arabian J-P58 from Saudi Arabia, Qatar, and UAE are pooled, resulting in low mean Y-STR variance of 0.16. This low value stems primarily from Qatar and UAE as the Saudi Arabian J-P58 makes a very small contribution (4 examples) in the pooled sample.

Unfortunately the authors just missed the very recent paper on Arabian DNA by Abu-Amero et al., which shows that J-M267 variance is 0.27-0.29 in Yemen and Saudi Arabia, and much lower (0.16-0.19) in UAE and Qatar. This severely weakens the case for an expansion of J1 from the northern to the southern Levant, as it reveals that not only Oman and Yemen (mentioned in the paper), but also the geographically dominant Saudi Arabia is a region of high Y-STR diversity. Thus it is not the case that:
The timing and geographical distribution of J1e is representative of a demic expansion of agriculturalists and herder–hunters from thePre-Pottery Neolithic B to the late Neolithic era.24,26 The higher variances observed in Oman, Yemen and Ethiopia suggest either sampling variability and/or demographic complexity associated with multiple founders and multiple migrations.
But rather Oman, Yemen and Ethiopia are not atypical for the southern J1e range, which also includes Saudi Arabia as a region of high Y-STR variance. It is rather only the small gulf states of UAE and Qatar that have lower variance.

An interesting find, however, is the fact of high Y-STR variance (0.37, 0.43) in Alawites from Syria and Assyrians from Syria and Iraq. These populations have an impeccable Semitic historical record, and, in the case of the Assyrians are one of the few non-Arabic populations included in the study. It is also interesting that Assyrians are said to be derived from both Assyrian- and Aramaic-speaking ancestors, and hence to potentially have a complex (both East- and Northwest- Semitic) origin. These facts probably explain their high Y-STR variance.

Translated into non-"evolutionary" years, the expansion time of 16.2ky for Assyrians, becomes ~5.4ky. This age is in uncanny agreement with the recently estimate age of Semitic languages 5.75ky ago.

The authors of the current paper cite the above-mentioned linguistic work, but have trouble bridging the gap between their own "evolutionary" dates and the date for the breakup of Proto-Semitic:
A recent Bayesian analysis of Semitic languages supports an originin the Levant 5750 years ago and subsequent arrival in the Horn of Africa from Arabia 2800 years ago,11 thus providing an indirect support of our phylogenetic clock estimates. It is important to note that the glottochronological dates yield estimates for the break-up and expansion of the Proto-Semitic language. Proto-Semitic, itself, may have been spoken in a localized linguistic community for millennia before its bifurcation into the East and West Semitic branches.
If one rejects the "evolutionary" rate, there is no need to postulate that Proto-Semitic was spoken (but did not disperse) for millennia; indeed, a "static" Proto-Semitic/J-P58 community would be difficult to explain in view of the fact that mobile herding was their main economic activity. In my view, The J-P58 bearing Proto-Semites emerge in the 4th millennium BC out of a general J1 Middle Eastern background, just as their TMRCA suggests. They begin to expand at that time, and emerge in the historical record 1-2 thousand years later in both their Eastern (Akkadian) and, later, Western (Aramaic and Canaanite) forms.

The authors also cite their own work with respect to the correlation of J1 distribution with semi-arid environments in the Middle East and cite evidence to the effect that:
archeological studies have shown an early presence (ca. 6000–7000 BCE) of domesticated herding in the arid steppe desert regions
The presence of a large frequency of undifferentiated J*(xJ1, J2) chromosomes in Soqotra suggests that the Arabian peninsula possessed such chromosomes, which now have a marginal status throughout the Middle East. I propose that a the early steppe desert herders of 6000-7000BC possessed J* chromosomes, that J1 arose in the Middle East, and its subclade J-P58 experienced rapid growth associated with the breakup and expansion of Semitic languages in the 4th millennium BC.

In conclusion: this paper gives us important new data on the origin and expansion of Y-chromosome J-P58, and strengthens the case that this haplogroup may be a diagnostic marker of the Proto-Semitic population of the Near East.

Related:

European Journal of Human Genetics doi: 10.1038/ejhg.2009.166

The emergence of Y-chromosome haplogroup J1e among Arabic-speaking populations

Jacques Chiaroni et al.

Abstract

Haplogroup J1 is a prevalent Y-chromosome lineage within the Near East. We report the frequency and YSTR diversity data for its major sub-clade (J1e). The overall expansion time estimated from 453 chromosomes is 10 000 years. Moreover, the previously described J1 (DYS388=13) chromosomes, frequently found in the Caucasus and eastern Anatolian populations, were ancestral to J1e and displayed an expansion time of 9000 years. For J1e, the Zagros/Taurus mountain region displays the highest haplotype diversity, although the J1e frequency increases toward the peripheral Arabian Peninsula. The southerly pattern of decreasing expansion time estimates is consistent with the serial drift and founder effect processes. The first such migration is predicted to have occurred at the onset of the Neolithic, and accordingly J1e parallels the establishment of rain-fed agriculture and semi-nomadic herders throughout the Fertile Crescent. Subsequently, J1e lineages might have been involved in episodes of the expansion of pastoralists into arid habitats coinciding with the spread of Arabic and other Semitic-speaking populations.

Link