September 05, 2009

More ASHG 2009 abstracts

For the first part, see here.

See Part I for another study on Ashkenazi Jews. We will have to look at the details of the study when it comes out, but the fact that Ashkenazi Jews are (a) between southern Europeans and Near Easterners, and (b) form a distinct cluster of their own at K=3 seems to support my theory that most of the European ancestry in Jews is of ancient origin in southern Europe rather than due to recent admixture with Central/Eastern Europeans: at K=2 the ancestral components are identified, but these components mixed a relatively long time ago, so that after a subsequent period of relative isolation, a distinctive pattern was formed out of the mixture which is identified at K=3.

Genome-wide SNP analysis of Ashkenazi Jews reveals unique population substructure
The Ashkenazi Jews (AJ) are a genetic isolate that has been widely utilized in genetic studies of both mendelian and complex disorders. However, the genetic variation and population structure of the AJ have been previously investigated with relatively few individuals and few genetic markers. We have now genotyped a large AJ cohort with the Affymetrix 6.0 genome-wide SNP array. After strict quality control filters, genotype data at 775K SNPs in 466 unrelated AJ individuals were available for analysis. To investigate the genetic structure of the AJ relative to other populations we used principle components analysis (PCA) as well as the frappe clustering algorithm. When merged with the worldwide Human Genome Diversity Project dataset, PCA shows the AJ are distinct from all other groups, including both European and Middle-Eastern populations. Further PCA using AJ genotypes combined with a large European dataset again validates the separation of AJ from European populations. Interestingly, principle component one seems to largely separate European and Middle-Eastern populations geographically according to latitude with the AJ fitting South of Europe and North of the Middle-East. Additional analysis using the frappe population clustering algorithm is consistent with a unique population signature for the AJ. Limiting the frappe clustering to only two population groups, specifying k=2, reveals that AJ cluster more closely to Europeans than Middle-Eastern populations but when allowing three populations, k=3, AJ form a group distinct from both the Middle-East and Europe. Compared to European populations, AJ also show an increase in genome-wide linkage disequilibrium, consistent with possible founder effects. These findings will aid in the design and use of AJ in case-control and association studies and clearly demonstrate the genetic separation of AJ from other populations.
Another paper on the topic, albeit one which uses HLA haplotypes to infer admixture and is limited to Jewish/Central European admixture.

Admixture between Ashkenazi Jews and Central Europeans
When distinct populations inhabit the same geographic space, culture often acts to restrict random mating in our species, while at the same preventing complete genetic privacy. The residency across Central Europe by the Ashkenazi Jews over the last thousand years is such a case. HLA typing from bone marrow donor registries in Israel, Poland and Germany were utilized to measure admixture between central European host populations and Ashkenazim. Inferred high resolution HLA A-B-DRB1 haplotype frequencies were generated from each population. A total of 1,676 Polishorigin- Ashkenazim and 13,556 Polish haplotypes were analyzed, along with a similar sample of ~5 million German haplotypes. The informativeness of HLA haplotypes is shown by the A-B-DRB1 haplotype 0101-0801-0301, the most common haplotype found in northern Europe. HLA B*0801 bearing haplotypes are present in the Near East, but those B*0801 haplotypes carry the HLA C allele Cw*0702 instead of the Cw*0701 found in 0101-0801- 0301. The 100 most common haplotypes constituted 53% of the total Ashkenazi, and 45% of the Polish, and 43% of the German samples, reflecting the sizeable total fraction of very rare haplotypes familiar in population samples of the diverse HLA system. The most common Ashkenazi haplotype had a frequency of 6.14% (n = 102.9) and the 100th haplotype was present at 0.29% (n = 4.86). Comparable values for the Polish sample were 5.83% (n = 790.3) and 0.13% (n = 17.6), respectively. Haplotypes from one population compared to those haplotypes in a second could be classified into three categories: less frequent, statistically identical or more frequent. In the graph of the ordered 100 Polish haplotypes, the less frequent Ashkenazi haplotypes supply a possible signature of admixture from the Poles into the Polish Ashkenazim, while the haplotypes more frequent in Ashkenazim than Poles are candidates for movement of genes from the Ashkenazim to the Poles. The averaged frequency differences between these categories give an indication of population admixture. The analysis showed that 1.8% of Polish haplotypes may be of Ashkenazi origin and 0.6% of Ashkenazi of Polish origin. The sample from Germany, in which the initial generations of Polish- Ashkenazi history was spent, was useful in demonstrating consistency of haplotype frequencies by rank order. The results show clear evidence of admixture occurring in both directions between two largely HLA-distinct populations.

The following study demonstrates a point I have argued several times before with Afrocentrists, namely the intermediate genetic position of Ethiopians between Caucasoids and Sub-Saharan Africans. It also underscores the difference between social and biological classifications: Ethiopians are undoubtedly "socially" black in most other societies, but intermediate between Negroids and Caucasoids anthropologically. This reality was recognized even by early anthropologists who coined the term of Ethiopids to describe them as a separate intermediate category between Caucasoids and Negroids.

The distribution of sex-specific human genetic variation in Ethiopia.
Ethiopia has been proposed as a candidate location for the emergence of anatomically modern humans, and the source region for the expansion out of Africa. It is also a region of substantial cultural diversity as expressed in languages (Nilo-Saharan, Cushitic, Semitic, and Omotic language families), religions (Christians, Jews, Moslems and Animists), ethnic identities (over 80 groups) as well as many marginalised groups socially excluded on grounds of caste-like occupation, supposed origin, or both. The demographic history of Ethiopia over the past several thousand years has involved both sustained migration of Semitic speakers from the Arabian Peninsula as well as internal conquests of lands in the south. To investigate the demographic histories of ethnic groups we analysed a battery of SNPs and microsatellites on the non-recombining portion of the Y chromosome (NRY) and sequence variation in the Hypervariable Segment 1 (HVS1) of mtDNA (5756 samples from 45 ethnic groups). Commonly used summary statistics (gene diversity h, genetic distance Fst) were analysed within the context of non Ethiopian data e.g. West Africa (Igbo, Nigeria) and Europeans. We present preliminary results reporting a wide range of genetic diversity values within ethnic groups (h: NRY = 0.743 - 0.972, HVS1 = 0.962 - 0.996) and pairwise genetic distance values between groups (Fst: NRY = 0.000 - 0.294, HVS1 = 0.000 - 0.035). A clustering of Ethiopian groups was observed when using principal coordinate analyses with genetic distances, appearing midway between a West African Niger-Congo speaking group (Igbo of Nigeria) and an Indo- European speaking group (Greek Cypriots). Some south-western groups (e.g. Anuak) showed greater similarity to West-Africans while the culturally influential Amhara were more similar to Europeans. Gene flow between dominant Dawuro agriculturalists and excluded members of the Manja was sex-biased, with many more NRY haplotypes common to the two groups than mtDNA haplotypes, relative to the distribution of the two systems across all the ethnic groups. The marginalised group had a particularly low level of mtDNA HVS1 diversity (h = 0.705). Of particular interest is the extensive sharing of discriminating NRY and mtDNA haplotypes across many ethnic groups, suggesting either a) the creation or preservation of cultural diversity despite substantial inter-group gene flow or b) recent ethnogenesis of the currently extant groups.
Yet another study of differences between ancient and modern mtDNA gene pools. I hope the 2012 crowd doesn't follow up on this for its own bizarre purposes...

Genetic Diversity of the Ancient People in Mesoamerica
DNAs were extracted from the human remains buried in the Moon Pyramidat archaeological Teotihuacan site in Mexico. Nucleotide sequences of theirmitochondrial D-loop and SNP sites were determined by the PCR-directsequencing. To reveal the genealogy of mitochondrial DNA sequences ofthe individuals buried in the Moon Pyramid and assess their positions amongNative Americans, we first constructed a network of the mitochondrial DNAfrom the contemporary Native Americans; the northern Native Americans(Haida, Bella Coola, and Nuu Chah Nulth), the central Native Americans(Huetar, Kuna, and Ngöbé), and the southern Native Americans (Yanomami,Zoro, Gavião, and Xavante), and compared them with those of the individualsfrom the Moon Pyramid. All of the mitochondrial DNA types from the MoonPyramid individuals were unique, and clear genetic affinities were notobserved between the Moon Pyramid individuals and any of the 10 NativeAmerican populations. To investigate genetic diversity among the contemporarycentral Native American populations, we constructed a phylogenetictree of their mitochondrial DNA sequences using the neighbor-joiningmethod. There was a major mitochondrial DNA sequence common to thesethree central Native American populations. However, there were a relativelysmall number of mitochondrial DNA types in each population, most of whichwere, moreover, unique to each Native American population. Next we comparedthe mitochondrial DNA sequences of the Moon Pyramid individualswith those of the ancient Mesoamerican people, ancient Maya people fromthe classic Copán site. We also used Huetar people as a reference for thecontemporary central Native Americans. The distribution of the mitochondrialDNA types found in the ancient Native Americans is greatly different fromthat found in the contemporary Native Americans. These results show thatgenetic diversity in the ancient Native Americans was not as low as that inthe contemporary Native Americans, suggesting an occurrence of bottleneckin the past.
This will be of great interest to Y chromosome enthusiasts.

Improved resolution of the human Y-chromosomal phylogeny using
targeted next-generation sequencing

The non-recombining part of the Y chromosome provides unique insights into male-specific aspects of human genetics and history. We are using next-generation Illumina sequencing to fully re-sequence targeted regions of the Y and resolve the Y-chromosomal phylogeny by characterization of additional single nucleotide polymorphisms (SNPs) on lineages of interest. Initially ~6 Mb of Y sequence (NCBI36:Y-chromosome: 12,308,579- 18,230,132) is being generated for an African haplogroup A male. The strategy involves sequence enrichment by long template PCR of genomic DNA (10-20 ng/reaction) using overlapping fragments of 5.5 - 6.5 kbp. Currently ~70% of primer pairs work using a standard touchdown PCR protocol. Fragments obtained from a single individual are pooled and used for library preparation and IIlumina sequencing. Re-sequencing generates accurate high coverage data; SNP calling and their subsequent validation will be presented. Most SNPs are expected to be rare but some are likely to resolve deep divisions within African populations. Subsequently, we aim to (1) determine the time depth of the human Y phylogeny, (2) resolve multifurcations in the major lineages by discovering additional SNPs on the relevant and (3) discover SNPs that mark any lineage of particular interest. In addition, we will be able to provide a subset of all primers that work well with this protocol to investigators who are interested in Y-chromosomal phylogenies so that comparable standard datasets can be generated for use by the community.
Female to male breeding ratio in the history of modern humans
Was the genetic contribution of men and women to successive generations the same? As a population, did we have fewer fathers than mothers? Was polygyny present among hominid lineages to influence relative divergence rates of autosomes and sex chromosomes? Students of genetic variation of the uniparentally inherited mitochondrial and Y-chromosome DNA confronted these questions, fewer addressed it by looking at the DNA diversity of autosomes and sex chromosomes (Hammer et al. 2009, Keinan et al. 2009) with equivocal results. Our approach is different: we analyzed the ratio of the population recombination rate, ρ, between autosomes and the X chromosome. The chromosome X recombines only in the female meiosis whereas autosomes undergo cross-overs in both male and female germ lines such that their relative ρ reflects changes in the breeding ratio, β. The estimate of β is calculated from the observed chromosomal ρ’s, obtained by InfRec (Lefebvre and Labuda 2008), after their calibration with the average chromosomal recombination rates known from pedigree data. We have tested our approach using coalescent simulations under different input parameters’ values and various demographic scenarios. For the HapMap populations we obtained β of 1.4 in Yoruba from West Africa, 1.2 in European and 1.0 in East Asian samples. This suggests that in the history of modern humans the reproductive variance between men and women did not drastically differ, thus consistent with the prevalence of monogamy or mild polygyny in the human lineage. Known incidences of polygyny may be of recent origin, related to raise of agriculture and shift from hunter-gathering to food producing economies, and therefore not sufficiently common to leave a strong genetic signature in the recombinational record. (Supported by GenomeQuebec/Genome Canada and Canadian Institutes of Health Research).


Accurate inference of individual ancestry geographic coordinates
within Europe using small panels of genetic markers

The study of genomewide datasets of thousands of individuals of European ancestry supports the close correspondence between genetic distances and geographic coordinates within Europe, especially when information from hundreds of thousands of genetic markers is used. In fact, Principal Components Analysis (PCA), summarizing genetic variation over the top two principal components (PCs), results in plots that are surprisingly reminiscent of geographic maps of Europe. We set out to discover those markers that are most closely correlated with geographic origin, seeking to predict individual ancestry at a fine level, and even for closely spaced populations. To this end we analyzed a previously described subset of the Population Reference Sample (POPRES). We focused on 12 populations and 1224 individuals for which geographic coordinates (longitude and latitude) of individual origin are given for at least 20 individuals per population. First, we performed a complete leave-one-out crossvalidation experiment using 447,212 SNPs, and a simple nearest neighbors approach to infer geographic coordinates. This resulted in extremely high accuracy, placing individuals within an average longitudinal error of 2.2 degrees, and an average latitudinal error of 0.88 degrees. Next, we applied an algorithm that we have previously described to select the top 5,000 SNPs that correlate well with population structure as captured by PCA. We then filtered highly correlated SNPs using standard linear algebraic algorithms for the column subset selection problem. We thus selected 500 maximally uncorrelated markers, which have a Pearson correlation coefficient of 0.92 with PC 1, and 0.83 with PC 2. We extensively validated the effectiveness of such SNP panels for genetic ancestry testing by once more performing a complete leave-one-out crossvalidation experiment on the 1224 studied individuals (approx. two weeks of CPU time in commodity hardware). Using 500 carefully selected SNPs we can place individuals within a few hundred kilometers of their reported origin (average longitudinal and latitudinal error of 4.7 and 1.9 degrees respectively). Finally, we crossvalidated our best panel of 500 SNPs on the HapMap CEPH European individuals, placing them accurately on the Northwestern corner of Europe. Not surprisingly, our SNP panel includes markers that are either within genes reported to be under selective pressure in Europeans, or in high LD with such genes.


Genetic relationships among the ancient Chinese populations viewed
from discrete cranial traits

The discrete cranial traits are informative in revealing the genetic relationship of human populations. Given little available knowledge on these traits, especially their underlying genetic determinants, the primary aim of this study is to select a small number of traits that are sufficiently informative to represent genetic differentiation among East Asian populations. We studied overall 51 traits for 1,578 skulls from 19 necropolises, and found that 5 traits could capture the largest variation in East Asian populations studied. They are accessory mandibular foramen, palatine torus, mandibular torus, mastoid foramen extra-sutural, and infraorbital suture. The analysis on these 5 traits resulted in similar population relationships to that using all 51 traits. The study on discrete cranial traits could not only facilitate exploration of the genetic relationship of populations, and could also allow identification of the genes underlying these anthropological traits.

Admixed ancestry and stratification of regional gene pools of Quebec
In Quebec, studies of different molecular polymorphisms have shown that the French Canadian gene pool is as diverse as its source European populations and, contrary to what was previously anticipated, does not display more homogeneity. To better understand the genetic structure of the contemporary population, we analyzed the origins and contribution of 7,798 immigrant founders identified in the genealogical ascendance of a sample of 2,221 subjects representative of the French Canadian population of Quebec. As expected, French founders are the most important in number (n=5,326) in all Quebec regions. They contribute for about 90% of the regional gene pools, except for regions located in the easternmost part of the province (76%), which are characterized by more diverse origins. Although this study supports the French founders’ importance, it also puts in the balance arguments in favor of the heterogeneity of the founding pool. The majority of immigrants landed as single member of their family, originating from all the regions of France. In addition, nearly all subjects have mixed origins, including French and non-French. Taken together, these results put into perspective the idea of the homogeneity of the origins of the French Canadians and of a pan-Quebec founder effect. The differential descent and genetic contribution of immigrant founders across regions points to the stratification of the French Canadian population of Quebec, showing a east-west gradient of diversity. These results will contribute to optimize study design in gene mapping studies relying on the founder effect in the French Canadian population of Quebec.

A nonsynonymous SNP in EDAR is associated with tooth shoveling
Teeth display variations among individuals in the size and the shape of cusps, ridges, grooves, and roots. In addition, there are certain dental characteristics which are predominant in certain human groups, such as tooth shoveling of upper incisors that is major in Asian populations but rare or absent in African and European populations. The common characteristics of dental morphology are thought to be determined mainly by genetic factors. However, genetic polymorphisms associated with dental morphology have not been elucidated yet. In humans, the ectodysplasin A receptor gene (EDAR) as well as the ectodysplasin A gene (EDA) is know to be responsible for hypohidrotic ectodermal dysplasia, a genetic disorder causing abnormal morphogenesis of teeth, hair, and eccrine sweat glands. Human genome diversity data have revealed that the derived allele of a nonsynonymous single nucleotide polymorphism (SNP), rs3827760 that is also called EDAR T1540C, is predominant in East Asian populations but absent in populations of African and European origins. It has recently been reported that the 1540C allele is associated with Asian-specific hair thickness. The aim of this study is to clarify whether the nonsynonymous polymorphism in EDAR is also associated with dental morphology in humans or not. For this purpose, we measured crown diameters and tooth shoveling grades, genotyped EDAR T1540C, and analyzed the correlations between them in Japanese populations. To comprehend individual patterns of dental morphology, we applied a principal component analysis (PCA) to individual-level metric data, the result of which implies that multiple types of factors affect the tooth size. This study clearly demonstrated that the number of the Asian-specific EDAR 1540C allele is strongly correlated with the tooth shoveling grade. The SNP significantly affected PC1 and PC2 in PCA, which denotes overall tooth size and the ratio of mesiodistal diameter to buccolingual diameter, respectively. Our study revealed a main genetic determinant of tooth shoveling that has classically received great attention from dental anthropologists. Further studies using powerful DNA technology will lead to clearer understanding about genetic factors for phenotypic variations in tooth morphology such as Carabelli’s tubercle, the numbers of cusps and roots, and the size balances shown in metric measurements.
Direct estimation of the microsatellite mutation rate
Characterizing the behavior of mutations is fundamental to our understanding of genetic variation. Attempts to directly observe DNA mutations arising from germline transmissions are confronted by two challenges: The large amount of DNA sequence that needs to be collected in order to observe a mutation (since the mutation rate in humans is estimated to be ~2x10-8 per generation), and a poor signal-to-noise ratio, due to the fact that any modern genotyping technology has an error rate far exceeding the mutation rate. Using deCODE Genetics’ database of over 95,000 Icelanders genotyped at over 3,000 microsatellite loci, we directly observed mutations in germline transmissions from pedigrees. Microsatellites are thought to have mutations rates as high as 10-3 per locus per generation. To overcome the genotyping error rate, which was estimated in this data set to be ≤10-2 per allele call after appropriate filtering, we carried out two independent analyses: (1) We restricted our analysis to mother-father-child trios, and required the mutated allele to be genotyped at least twice in both the child and in the transmitting parent to confirm mutant transmissions. This identified 2,124 mutant events from 5.62 million instances of parent-child transmissions, yielding a mutation rate estimate of 3.78x10-4 averaged across the markers that we analyzed. (2)Wetraced the haplotype affected by the mutation through local pedigrees, requiring that the mutant haplotype is observed in the affected proband’s children, and simultaneously, that the wildtype haplotype is observed in the affected proband’s siblings. This identified 788 mutant events from 1.59 million instances of parent-child transmissions, yielding a mutation rate of 4.96x10-4. Our collection of mutant events is significantly larger than previous studies. This allows for categorical analyses of microsatellite mutation rates partitioned based on the gender and age of the individual transmitting the allele, as well as the repeat type and cytogenetic position.

Eastern European ancestry in New Hampshire

There have been a bunch of studies on Hispanic Americans, Native Americans, African Americans, but very little work on European Americans (if we exclude the perennial fascination of the genetics community with Ashkenazi Jews and some studies which included European Americans of known European parentage).

This is one of the first studies I've seen where the objective was to look at a geographically definite population of European Americans and study its diverse origins in Europe itself. While there are many European Americans whose ancestry is no mystery at all (because their ancestors arrived within memory), there are also large numbers of them with much older ancestry, and these should sometime become the object of study, both for their own sake, but also because they may represent a separate evolutionary road of their ancestral European gene pool.

The STRUCTURE result, beautified by CLUMPP, is really fascinating. Unlike most studies where sub-population labels of clustered individuals are put on the chart, in this case individuals do not necessary report a single ancestry, so cannot be put on a single population label. Yet it is really evident that "European Americans from New Hampshire" can be broken down to several groups with a distinctive ancestry.

From the paper:
Bayesian clustering conducted using the structure software revealed distinct subpopulations, with the highest and most reliable probabilities between a K of 5 and 7. The bar plots are shown for K = 2 to K = 8 from the CLUMPP software (aligns multiple runs of structure) from 10 runs at each K (Figure 1a). As expected, individuals in the sample appear highly admixed; however distinct populations are discernible. The FST's increase consistently as K increases, with the average FST's for K = 4 to K = 7 around the level of “little genetic differentiation” as defined by Wright (approx. 0.05) (Figure 1c,d) [22]. The admixture values increase for lower K's, but begin to drop at K = 6 to values between 0.6–0.7 (Table S2). In selecting the most correct K, parsimony is an important consideration, i.e. that the simpler answer tends to be correct. Though there may be some validity to further subdividing the groups, the most statistically consistent and the most parsimonious K based on the structure output is K = 6. Further analysis using the ancestral data is used to describe the groupings and lends support to our selection of K = 6.
...
These results suggest that genetic population structure is detectable in a highly admixed US population and that this structure correlates with self-reported ancestry. To our knowledge, this is the first time such an investigation has uncovered a strong link between structure and ancestry in what would otherwise be assumed to be a homogeneous US state where most individuals are of European ancestry. Our data indicate that that admixture has not eliminated the genetic structure found within Europe, and descendants of the Russian, Polish and Lithuanian immigrants remain genetically distinct from the rest of the population and are closely related to one another.
...
Exploratory analysis revealed that among the ancestries, those reported by at least five individuals were: American Indian (n = 32), Austria (n = 5), Belgium (n = 5), Canadian Indian (n = 14), Canada (n = 113), Czech Republic (n = 5), England (n = 355), Finland (n = 7), French-Canadian (n = 54), France (n = 173), Germanic (countries where Germanic languages spoken) (n = 5), Germany (n = 110), Greece (n = 9), Ireland (n = 218), Italy (n = 41), Jewish (n = 6), Lithuania (n = 12), Canadian Maritime Provinces (n = 6), Netherlands (n = 25), Poland (n = 44), Russia (n = 13), Scotland (n = 157), Sweden (n = 24), Switzerland (n = 7), UK (n = 11), US (n = 42), Wales (n = 24).

PLoS ONE doi:10.1371/journal.pone.0006928

Genetic Population Structure Analysis in New Hampshire Reveals Eastern European Ancestry


Chantel D. Sloan et al.

Abstract

Genetic structure due to ancestry has been well documented among many divergent human populations. However, the ability to associate ancestry with genetic substructure without using supervised clustering has not been explored in more presumably homogeneous and admixed US populations. The goal of this study was to determine if genetic structure could be detected in a United States population from a single state where the individuals have mixed European ancestry. Using Bayesian clustering with a set of 960 single nucleotide polymorphisms (SNPs) we found evidence of population stratification in 864 individuals from New Hampshire that can be used to differentiate the population into six distinct genetic subgroups. We then correlated self-reported ancestry of the individuals with the Bayesian clustering results. Finnish and Russian/Polish/Lithuanian ancestries were most notably found to be associated with genetic substructure. The ancestral results were further explained and substantiated using New Hampshire census data from 1870 to 1930 when the largest waves of European immigrants came to the area. We also discerned distinct patterns of linkage disequilibrium (LD) between the genetic groups in the growth hormone receptor gene (GHR). To our knowledge, this is the first time such an investigation has uncovered a strong link between genetic structure and ancestry in what would otherwise be considered a homogenous US population.

Link

September 04, 2009

ASHG 2009 abstracts

It's that time of year again. Here is a list of abstracts from ASHG 2009 that caught my attention in three broad areas. It will be very interesting to see these when they become full papers, but if you are one of the lucky ones that goes to Hawaii this October and want to drop me a line about any of them, feel free to do so!

Population Genetics

Haplogroup H of mitochondrial DNA, a far echo of the West in the heart of Central Asia
Through the millennia, Inner Asia played a pivotal role in shaping the history that greatly added to the cultural, ethnic, and genetic diversity observed throughout present Eurasia. Perhaps the two most significant phenomena witnessed in this part of the world were the ambitious expansion strategy employed by Mongolia’s most prominent personality, Genghis Khan and the complex network known as the Silk Road that for nearly 3,000 years contributed to the exchange of goods and the transmission of philosophy, art, and science that laid the foundation for the great civilizations of China, India, Egypt, Persia, Arabia, and Rome, and in several respects to the modern world. Over the last few years, through an international collaborative effort, researchers at the Sorenson Molecular Genealogy Foundation were able to collect 2,727 DNA samples, informed consents, and genealogical data in Mongolia, Kyrgyzstan, and Kazakhstan. All the samples were sequenced for the three hypervariable segments of the mitochondrial DNA (mtDNA) control region to assess the genetic composition of the modern population of these countries. We identified ~600 different haplotypes that could be ascribed to more than 30 haplogroups and sub-haplogroups. As expected, most haplogroups are typical of modern East Asian populations, but intriguingly, many different Western Eurasian clades were also identified, with a particular high incidence of H (~8.0%), the most common haplogroup in Europe. This feature cannot be attributed to genetic drift since different H sub-lineages have also been identified, each of them represented by several different haplotypes. The mtDNA distribution profile in the heart of Central Asia suggests a direct link between this area and Western Eurasia that could be explained by ancient migrations or by more recent historical events, such as Genghis Khan’s conquering efforts and trade or cultural exchanges along the Silk Route. To discriminate between these two possible scenarios, we are now analyzing a subset of these samples at the highest possible level of resolution - that of complete mtDNA sequences - focusing particularly on those H mtDNAs that seem to be the most informative considering their control-region haplotypes. Our preliminary data seems to be in favor of rather ancient genetic inputs from the West in shaping the peculiar mtDNA gene pool of Inner Asia’s present-day populations.
The following study seems to do precisely what I recently asked for:
However, as the PCA analysis shows, Ashkenazi Jews are distinct from both Europeans and non-Jewish Middle Eastern populations and cannot be viewed as a simple mix of the two; their distinctiveness must be -in part- due to the specific features of the small founder population of that community after it became effectively reproductively semi-isolated from gentiles after Roman times. It would be interesting to see different Jewish communities studied in the context of a broad variety of European and Middle Eastern populations, to determine whether Ashkenazi distinctiveness is specifically Ashkenazi or more generally Jewish distinctiveness; I would bet on a combination of the two.

Abraham's children in the genome era: Major Jewish Diaspora populations comprise distinct genetic clusters with shared Middle Eastern ancestry
Despite residence all over the world, Jewish populations have maintained continuous genetic, cultural, and religious tradition over 4,000 years. The unique ethnic makeup and social practices provide an invaluable opportunity to understand their genetic origins and migrations and to elucidate the genetic basis of complex disorders. To generate a comprehensive HapMap of ethnically diverse, healthy Jewish populations, we used the Affymetrix array 6.0 to genotype 381 samples recruited from 7 Jewish communities with different geographic origins: Eastern European Ashkenazim; Italian, Greek and Turkish Sephardim; Iranian, Iraqi, and Syrian Mizrahim (Middle Easterners). Here, we present population structure results from compiled datasets after merging with the Human Genome Diversity Project and the Population Reference Sample studies, which consisted of 146 non-Jewish Middle Easterners (Druze, Bedouin and Palestinian), 30 northern Africans (Mozabite from Algeria), 1547 Europeans, and 653 individuals from other African, Asian, Latin American, and Oceanian populations. Both principal component analyses and multi-dimensional scaling analysis of pairwise Fst distance show that Jewish populations form a cluster clearly distinct from all major continental populations. The results also reveal a finer population substructure in which each of 7 Jewish populations studied here form distinctive clusters - in each instance within group Fst was smaller than between group, although some groups (Iranian, Iraqi) demonstrated greater within group diversity and even sub-clusters, based on village of origin. By pairwise Fst analysis, the Jewish groups are closest to Southern Europeans (i.e. Tuscan Italians) and to Druze, Bedouins, Palestinians. Interestingly, the distance to the closest Southern European population follows the order from proximal to distal: Ashkenazi, Sephardic, Syrian, Iraqi, and Iranian, which reflects historical admixture with local communities. STRUCTURE results show that the Jewish Diaspora groups all demonstrated Middle Eastern ancestry, but varied significantly in the extent of European admixture. There is almost no European ancestry in Iranian and Iraqi Jews, whereas Syrian, Sephardic, and Ashkenazi Jews have European admixture ranging from 30%~60%. Analysis of identity-by-descent provides further insight on recent and distinct history of such populations. These results demonstrate the shared and distinctive genetic heritage of Jewish Diaspora groups.
So, it seems that there will soon be real genomic data on the source and extent of admixture in Jews. The absence of Greek and Anatolian samples may be problematic in finding the sources of such admixture, but the presence of Tuscans, who are reasonably close to them in a pan-European context should do well to serve as a substitute. In a recent sutdy (in which Anatolians were not included), the closest populations to Ashkenazi Jews were Italians of mostly southern provenance (Fst=0.0040) and Greeks (Fst=0.0042) and fairly close to Tuscans (Fst=0.0066)


The following study seems to demonstrate my recent suggestion of archaic admixture in Africa itself:
It does not, however, tell us that this is because of archaic introgression in Europeans. The culprit could equally well be long-term population structure in Africa, i.e., the presence of "modern" and "archaic" populations in Africa itself.
Deep population structure in sub-Saharan African populations
We analyzed ~500 Kb of resequencing data from 91 different intergenic regions in samples from three sub-Saharan African populations: Mandenka from Senegal, Biaka pygmies from the Central African Republic and San from Namibia. We employed novel methodology to estimate the split times and migration rates between populations. We found strong evidence for split times that predate the exodus of modern humans out of Africa (e.g., > 100 Kya). In addition, we also found evidence of ancient admixture (with unknown ‘archaic’ human groups) in the recent history of both the Biaka and the San.
Analysis of Genomic Admixture in Costa Rica Population
Costa Rica (CR) population is a unique population representing a typical admixture of major continental ancestral populations. 1,301 samples collected from participants in a population-based study conducted in the Guanacaste region of CR were genotyped on a custom Illumina iSelect chip harboring 27,635 SNPs. The SNPs on the chip were selected based on multi-ethnic tagging strategy for three HapMap populations: CEU, YRI and JPT+CHB and cover 1,000 candidate genes/regions for a range of cancers. This data set was sufficiently large for the investigation of population substructure in our CR study and the examination of linkage disequilibrium (LD) patterns. Three HapMap major continental populations and a Native American population from the Illumina iControl DB were used as the reference populations for these analyses. Our preliminary results indicate that the Guanacaste CR population was formed mainly by a three-way admixture with 42.5%, 38.3% and 15.2% Native Indian, European, and African respectively. In addition, 4.0% residual genetic component derived from Asians was observed in our CR samples. Both model based STRUCTURE program and Principal Component Analysis (PCA) revealed consistent substructure pattern for the CR population. The magnitude of LD in the CR population seems to be smaller than all the reference populations except YRI. A more detailed knowledge of the underlying genetic structure of the CR population would be informative to assess its population genetic history and to assist in the interpretation of investigations of complex diseases in the CR or a comparably admixed population.
Analysis of Genetic Substructure of Han Chinese Using Genome-Wide SNP Arrays: Implication for Association Studies.
China will start this year a $30 million effort of genome-wide association studies (GWAS) of common diseases in Chinese populations which have been largely underrepresented in the similar effort worldwide. A general concern is population stratification (ancestry differences) among subpopulations which can cause false positive associations. Han Chinese is the largest ethnic group in the world, however, its population substructures are often expected and yet well characterized. In this study, we examined population substructures in a diverse set of >1,700 Han Chinese samples collected from 26 regions, each genotyped with at least 160K single nucleotide polymorphisms (SNPs). Our results showed that: (a) Han Chinese population is complicatedly substructured, with the main observed clusters roughly corresponding to northern Han, central Han and southern Han; (b) Han Chinese samples collected from large cities, such as Shanghai, Beijing and Guangzhou, show diverse source of ancestries including three aforementioned clusters; (c) HapMap samples (CHB & CHD) and HGDP samples (Han & Han-NChina) deliver a limited representation of Han Chinese people. Building on the above insights, we investigated false positive rates and statistical power in various study designs using both empirical and simulated data. We further explored sample collection strategies and public data usage for future association studies.
It will be interesting to see if the authors of the following study estimated gene flow in non-southern European populations as controls, to see what is the excess of Sub-Saharan admixture detected in the three southern European samples, and exactly what "methods that can infer admixture proportions in the absence of accurate ancestral populations" they used. Hopefully they will also extend their linkage disequilibrium analysis for the other populations besides Spaniards.

Characterizing the history of sub-Saharan African gene flow into southern Europe
Recent analyses of whole-genomeSNP data sets have suggested a history of sub-Saharan African ancestral contribution into southern Europe but not in northern Europe, consistent with previous analyses based on the Ychromosome and mitochondrial DNA. However, there has been no characterization of the proportion of African admixture in southern Europe, or of its date. Here we analyze data from ~450,000 autosomal SNPs in the Population Reference Sample, ~650,000 SNPs from the Human Genome Diversity Panel, and ~1.5 million SNPs from the HapMap Phase 3 Project, and studied patterns of correlation in allele frequencies across populations to confirm the evidence of African ancestry in many southern European populations but not in northern Europeans. Using methods that can infer admixture proportions in the absence of accurate ancestral populations, we estimated that the proportion of sub-Saharan African ancestry in Spain is 2.4 +/- 0.3%, in Tuscany 1.5 +/- 0.3%, and in Greece 1.9 +/- 0.7% (1 standard error). We also studied the decay of admixture linkage disequilibrium with genetic distance, which provided a preliminary estimate of the date of African gene flow into Spain of roughly 60 generations ago, or about 1,700 years ago assuming 28 years per generation. This date is consistent with the historically known movement of individuals of North African ancestry into Spain, although it is possible that this estimate also reflects a wider range of mixture times.
Genome-wide patterns of population structure and admixture among Hispanic/Latino populations
In order to document genome-wide patterns of variation in Hispanics/ Latinos (HL’s) we genotyped individuals from five distinct populations recruited in the US: Mexico, Colombia, Ecuador, Dominican Republic and Puerto Rico. We present population structure results from an extensive genome-wide SNP dataset compiled by merging Affymetrix 500K and Illumina 650K data from these populations together with the Human Genome Diversity Panel, HapMap, Mao et al (2005), and POPRES studies. We apply Principal Component Analysis (PCA) and a clustering method, frappe, to infer admixture and genetic relationships of 262 HL individuals with 467 Africans, 715 Europeans, and 210 Native Americans comprising a total of 88 populations. We observe substructure within Native Americans, and, as expected, find that the admixed HL populations show Native American ancestry derived from local Native American populations. We find striking differences in estimated population-wide mean African, European and Native American ancestry proportions which are consistent with historical admixture and proximity to slave trade routes. The Dominican Republic and Puerto Rico, located on islands along slave trade routes, show high levels of African Ancestry (means 41.7% and 23.6% respectively) with less Native American Ancestry (11.5% and 18.9%). Colombians show a wide range of both African and Native American ancestry, though they have an overall mean of slightly higher Native American ancestry (36.3%) and lower African ancestry (11.7%) than the highly-African Dominicans and Puerto Ricans. Ecuadorians show the highest Native American mean ancestry (54.0%) with low estimated mean African Ancestry (7.3%). Mexico shows the largest range of Native American ancestry (11.0% - 79.0%) with an overall mean of 50.1% Native American ancestry and the lowest African ancestry (5.6%). Our study shows a broad range in admixture proportions across different HL individuals as well as different admixture patterns across populations. We also compare this genotype data with mtDNA and Y chromosome genotypes and use simulations to estimate ancient male and female sex ratios in each HL population. Lastly, we discuss implications of population structure for genome-wide association studies in admixed populations such as HL’s, especially when recruited in the United States.
A new statistical method to infer population admixture events using genetic variation data
We present a novel statistical method that uses densely-spaced Single- Nucleotide-Polymorphism (SNP) data to identify the major admixture events occurring throughout a population’s history. The model has several advantages over leading available analytical approaches in this area, such as principal-components-analysis and STRUCTURE. In particular it can simultaneously (i) take advantage of the information inherent in patterns of linkage disequilibrium, i.e. non-random associations amongst neighbouring SNPs along a chromosome, (ii) efficiently analyse hundreds of individuals at hundreds of thousands of SNPs genome-wide, and (iii) allow for relatively straight-forward interpretation and direct inference of key historical parameters, such as the proportions and times of major admixture events. Using simulated data matched to currently available human datasets, we show that our model can identify and accurately date admixture events that have occurred between 7 and 150 generations ago. As our technique exploits the rich information in genetic data to infer details of a population’s admixture history, it marks a powerful complement to anthropological research and can help to resolve a number of existing controversies. We present results from applications of our model to two datasets: (1) SNP data from 22 distinct genetic regions for individuals from three chimpanzee populations in Africa; (2) genome-wide 650K SNP data for individuals from 53 world-wide populations of the Human Genome Diversity Panel (Science 319, 1100-1104). We highlight a number of intriguing new insights from these analyses. For example, the chimpanzee analysis showcases the model’s ability to infer the relative divergence among populations. The human analysis identifies several important admixture events, some of which are historically wellestablished (e.g. identification of recent European genetic influx into the Maya Native American population), others that can be placed into a clear historical context (e.g. an East Asian genetic influx into several Central and South Asian populations dated precisely to the era of the Mongol empire), and some that are to our knowledge novel (e.g. admixture in the Cambodian population between a Central/South Asian source and an East Asian source dated to around the period of the Cambodian Empire).
Bayesian methods of estimating ancestry using whole-genome SNP data
Estimation of the genetic ancestry of an individual is useful for association studies, disease risk prediction, population genetic analyses and is of inherent interest for the individual themselves. We have investigated methods of estimating ancestry using whole-genome SNP data on each individual. We focus on the scenario where the goal is to determine ancestry in relation to a set of genotype or haplotype data that is available from a set of distinct source populations, for example, the HapMap 2, HapMap 3 or 1000 Genomes datasets. Inference in this setting can focus either on the estimation of global ancestry, in which an overall estimate of the proportion of ancestry from the source populations is needed, or local ancestry, which aims to partition an individual genome into distinct segments of ancestry from the source populations. We have compared 2 models based on the estimated allele frequencies in the source populations at a set of unlinked SNPs. Model 1 only models global admixture, whereas Model 2 models both global and local admixture. Using simulated individuals with differing proportions of CEU and YRI admixture (based on HapMap3 data) we find that there is a relatively small difference in the mean square error of the estimates of global admixture from the 2 methods (1.16 10-4 and 8.88 10- 5 respectively). Since Model 1 is much faster to fit that Model 2 these results suggest that Model 1 can be used to estimate the level of global ancestry, or at the very least will be useful as an initial estimate for use in Model 2. Further investigation is required to see how these results hold for more genetically similar source populations. In contrast, the mean square error for the estimates of local admixture from the 2 methods is 0.298 and 0.0861 respectively, suggesting that an explicit model of local ancestry is needed to carry out this level of inference. We are also investigating the utility and practicality of using linked SNP data to estimate global and local admixture.
A detailed phylogeography of mtDNA haplogroup C1d: another piece in the Native American puzzle
Recent studies based on complete mitochondrial DNA (mtDNA) sequences revealed that two almost concomitant paths of migration from Beringia led to the dispersal of the first Americans (Paleo-Indians) approximately 15-17 thousand years ago (kya). This first expansion was followed by later more restricted diffusion events from the same dynamically changing Beringian source. Thus, five pan-American (A2, B2, C1, D1, and D4h3a) and four geographically confined (D2, D3, X2a, and C4c) mtDNA haplogroups represent the current female legacy of the ancient migratory events that gave rise to the native populations of the double continent. Regarding haplogroup C1, all its members appear to belong to one of three branches: C1b (characterized by the control-region transition at np 493), C1c, and C1d (with the control-region transition at np 16051). These three sub-haplogroups are found throughout the Americas, thus supporting the scenario that they most likely differentiated at the early stages of the Paleo-Indian southward migration. If considered as three separate founders, C1b, C1c, and C1d would bring the currently known number of native pan-American lineages to seven. As a whole, the C1 haplogroup has an estimated age of 17.0- 19.6 ky, while the three individual branches are dated 16.5-17.0 ky, 17.2- 17.6 ky, and 7.6-9.7 ky, respectively. The extremely young age estimate of C1d has been attributed, at least for the moment, to a major underrepresentation of C1d mtDNAs (only nine complete sequences published to date) in the current Native American mtDNA phylogeny. We have addressed this issue in the current study by completely sequencing more than 60 novel mtDNAs belonging to haplogroup C1d, which were carefully selected on the basis of both control-region variation and geographic/ethnic origin. Phylogeographic analyses have provided not only an accurate evaluation of the expansion time of C1d in the Americas, but also a detailed picture of its current distribution in both general mixed and indigenous populations.

Genetic diversity of European population isolates in the context of their geographic neighbors
Mapping traits in population isolates provides an opportunity to simplify the challenges of complex trait mapping because such populations likely have enhanced levels of linkage disequilibrium and reduced genetic heterogeneity for the underlying traits. Here we analyze high-throughput SNP genotyping data to compare genomic-scale patterns of variation in several European population isolates (Adygei, Basque, Orcadian, Roma from Slovakia, Sardinians, and Sorbs) and contrast their patterns of variation to geographical proximal populations. Our results reveal insights for the demographic history of each of these unique populations, suggest substantial variation among these population isolates in patterns of diversity, and highlight the importance of population selection in genome-wide association mapping.
Incompatibility of current Finnish mitochondrial diversity with simulations of assumed settlement history
Traditionally, geneticists studying Finnish population history have assumed a model where Northern and Eastern Finland were mostly uninhabited until the 16th Century A.D. and were then settled by small family groups from South-Western Finland. The reduced genetic diversity and the distinct Finnish disease heritage are seen as consequences of these founder effects. Y-chromosomal diversity is indeed reduced in the present population, especially in the eastern parts of the country. However, mitochondrial diversity is not heavily reduced compared to South-Western Finnish or other European populations. This discrepancy has been explained with the higher mitochondrial mutation rate having restored mitochondrial diversity in these populations since the founder effects.
In our view it seems unlikely that even with high mitochondrial mutation rates mtDNA diversity could be restored over a mere 17 generations after the alleged tight bottlenecks. Archaeological evidence also suggests a different settlement history, e.g. settlement beginning in South-Eastern instead of South-Western Finland.
In this study we use simuPOP, a state-of-the-art forward simulation tool, to simulate datasets corresponding to Finnish mitochondrial diversity under the traditional model and compare them with actual present-day Finnish data. We show that current mitochondrial variation is unlikely under this model, increasing the credibility of alternative hypotheses.
On the borderline between the east and the west: the maternal genetic background of Karelians
Introduction: The frontier between Finland and Russia represents one of the most conspicuous socioeconomic gaps in the world. Based on the mean gross national product, there is a ten-fold difference between Russian Karelian Republic and Finnish Karelia. Otherwise these populations share the same geophysical environment. For these reasons, Karelia has been a very interesting field of research for multifactorial disease studies. However, this area has undergone many demographic incidents, such as wars and famine, which may cause local differences in the gene pool. In this study, we wanted to elucidate the maternal genetic background of Karelians. Materials: Blood samples were collected from healthy unrelated individuals without known foreign background from four Karelian districts; Aunus(n=218), Viena(n= 87), Tver(n=61) and Finnish Karelia (n=70), The sample collection was performed according to the Basic Principles of the Declaration of Helsinki. Methods: The entire mitochondrial DNA was sequenced in 32 reactions per sample with the BigDye® Terminator v3.1 Cycle Sequencing Kit in the Applied Biosystem’s 3730 Genetic Analyzer sequencing machine. Sequence alignments were made by the SeqScape® Software, Version 2.5 (Applied Biosystem). Results: Haplogroup H was very common in all populations. However, H1a is almost absent in Finnish Karelia. Also U and its subhaplogroups were common. Specially U5b1b1 reached over 16% in Viena Karelians. U4 was most common among Tver Karelians. Conclusions: The maternal genetic background seem to be complex in this area. There is clear regional differences. Also there is solid evidence of gene flow from various sources. Representation of the clearly Asian haplogroups is strikingly low.
Genetic Landscape of Eurasia Viewed from Large Allele Frequency Differences.
The diversification leading to modern human populations in Eurasia is one of the most important topics in the study of human expansions after leaving Africa. Most studies of Eurasia populations have used either limited markers or involved insufficient population coverage. We chose 68 markers based on large allele frequency differences among a few Eurasian populations and then typed them on 1766 individuals from 34 populations representing all subdivisions of Eurasia. Analyses using the STRUCTURE program showed a clinal east-west division when K=2, with a median border dividing Central Asia along the Ob River, the Kazakh highland, the western side of Pamir Mountains, and the southwestern side of the Himalayas. We fit curves to the STRUCTURE loadings using distances of the population coordinates from the median border. The genetic structure changed dramatically only within 2000km on each side of the border. At higher values of K the western populations of East Asia are the first to be distinguished (at K=3): Mongols, Tibetans, Qiang, and Baima, are most distinct from the more eastern populations. At K=4 Southwest and South Asians are distinguished from the Europeans; At K=5 Southeast Asians and at K=6 Central Asians are successively distinguished from eastern East Asians. Several more isolated populations such as Samaritans, Atayals, or Micronesians were distinguished in different independent runs when K=7 providing no clear anthropological information. South Asians were always clustered with Southwest Asians with pronounced similarity to Central Asians. The failure to distinguish South Asians maybe due to the selection of the markers with large allele frequency differences specifically between Europeans and East Asians. We also tested for statistical differences in the allele frequencies for all pairs of clusters when K=6. The results showed significant borders (P less than 0.0001) including those between western East Asians and eastern East Asians or Central Asians; however, insignificant borders were observed between Southwest Asians and Southeast Asians or western East Asians, neither was between Central Asians and eastern East Asians. This indicates substantial gene flow in North Asia between eastern East Asians and Central Asians, and in South Asia between South Asians and Southeast Asians. Using increased population and marker coverage, this study helps to understand the details of genetic diversity and landscape of Eurasians.

Anthropometry

Dairy intake associates with the IGF2 rs680 polymorphism to height variation in Greek children. The GENDAI study
Objective: Height is a classic polygenic trait with a number of genes underlyingits variation. We evaluated the prospect of gene to diet interactions ina children cohort for the IGF2 rs680 polymorphism and height variation.Methods: We screened 795 peri-adolescent children (424 females) aged10-11 years old from the (Gene and Diet Attica Investigation; GENDAI)paediatric cohort for the IGF2 rs680 polymorphism. Results: Children homozygousfor common allele (GG) were taller (148.9 ± 7.9 cm) comparing tothose with the A allele (148.1 ± 7.9 cm), after adjusting for age, sex, anddairy intake (β±SE: 2.1± 0.95, p=0.026). A trend for interaction for theIgfrs680xdairy intake is also revealed (p=0.09). Stratification by IGF2 rs680genotype revealed a positive association between dairy products intakeand height only in A allele carriers, adjusted for the same confounders(standardized β=0.111, p=0.014). When dairy intake was classified, basedon the median value, into two equal groups of low (1.9 ± 0.7 servings/day)and high dairy products intake (4.4 ± 1.5 servings/day), it was found thatin A allele children high dairy eaters were significantly taller (p=0.05) comparedwith low dairy eaters (148.8 ± 7.9 cm vs 147.4 ± 7.7 cm respectively,adjusted for age and sex). Conclusion: A higher consumption of dairy productsassociated with increased height depending on the rs680 IGF2 genotype.Thus, exploring height variants and elucidating possible interactionswith environmental factors like diet could help us to design
A Non-synonymous HNF4A Variant is Associated with Glycemia During Pregnancy and Offspring Head Circumference in Populations of European Ancestry in the HAPO Study
The Hyperglycemia and Adverse Pregnancy Outcome (HAPO) study is a multicenter, international study, which examined the association of maternal glucose levels with fetal growth and outcome in 25,000 pregnant women from multiple ethnic groups to demonstrate a continuous relationship between maternal glucose measures and birth size throughout the range of glucose concentrations. We hypothesize genetic factors contribute to these phenotypes, and examined 1536 fetal and maternal SNPs in 79 candidate loci previously implicated in insulin secretion or sensitivity to determine associations with maternal glycemia and insulin secretion (fasting glucose and Cpeptide and 1-hr glucose from the OGTT) at ~28 weeks gestation and/or offspring size at birth (birth weight, length, head circumference, and sum of skinfolds) for HAPO mothers of European (Belfast and Manchester, UK, and Brisbane and Newcastle, Australia; N=3828) and Asian (Bangkok, Thailand; N=1813) ancestry and their offspring. Associations were assessed through linear regressions with the single trait/outcome under an additive genetic model adjusting for known confounders. Among our strongest signals was rs1800961G>A, which encodes a Thr>Ile amino acid change in exon 4 of HNF4A, recently identified in a GWAS meta-analysis as a variant associated with decreased HDL levels. In the HAPO study, this SNP was strongly associated with increased fetal head circumference (0.5cm [95%CI: 0.3-0.7] per maternal minor allele; P=1.2x10-7) in those of European descent. The maternal minor allele was also weakly associated with 1-hour glucose (4.3mg/dL [95%CI: 0.5-7.9]; P=0.03), birth length (0.7cm [95%CI: 0.2-1.1]; P=0.003), birth weight (52.6g [95%CI: -8.0-113.3]; P=0.09), and sum of skinfolds (0.3cm [95%CI: -0.1-0.6]; P=0.13). This same minor allele in the fetal genome was weakly associated with cord C-peptide (0.1ug/dL [95%CI: 0.01-0.22]; P=0.03), and head circumference (0.2cm [95%CI: -0.1-0.4]; P= 0.08). The same trends were observed among the Thai, although not significantly probably due to a reduction in power from the low risk allele frequency (<2%).>
Selection

In a recent study, Heyer used germline mutation rates to estimate time depth, so I am more inclined to take her dates at face value than in papers which used "evolutionary" rates. It will be interesting to see which Y-chromosome types the authors associates with the both the older and recent expansions.

Super Y-chromosomes in Eurasia and the impact of social selection and Neolithic transition
Some Y-chromosomal haplotypes have been found at unusually high frequenciesin Asian and European human populations. The massive spreadof these lineages has been explained by the impact of social selection i.e.the high reproductive success of some males and their relative/descendantsdue to their high social status. The most well-known examples are the “Khanhaplotype” and the “Manchou haplotype” in Asia, and the U’Neill haplotypein Ireland. But are these frequent haplotypes always associated with recentevents of social selection, or could they be linked to much older processes?To address this question, we have surveyed ~ 3500 males in 97 populationsfrom Turkey to Japan. We have focused on the 12 most frequently representedhaplotypes in Eurasia and tested whether their expansions are linkedto a specific factor such as language or subsistence methods. Our resultsshow that both recent and ancient processes are responsible for the expansionsof these lineages. The recent expansions (2000-3000 years) likely tobe linked to social selection are prevalent in Altaic-speaking and pastoralpopulations. This might indicate a recent cultural change in the social organizationof these populations. The ancient expansions (8000-10000 years)are over-represented in Indo-European speaking and sedentary farmer populations,and are likely to be the result of the Neolithic transition.

Lactase Persistence; Multiple causal mutations in sub-Saharan pastoralists
Background Milk is the primary source of nutrition for newborn mammals, including humans. The majority of human adults, estimated at approximately 65%, are unable to digest lactose (the main carbohydrate in milk) effectively since lactase expression is down-regulated after weaning, as it is in other mammals. In some humans however, lactase expression persists into adulthood (lactase persistence, LP) allowing adult consumption of milk from other species, and the frequencies of this trait vary throughout the world. A C-T SNP -13910 bases upstream from the lactase gene (LCT) is associated with LP in Europe. The -13910*T is rare in milk drinking groups in Africa although two other variants (-13915*G, -14010*C) have been shown previously to be significantly associated with LP and in an accompanying abstract (Ingram et al) we confirm a third locus (-13907*G) and present a fourth candidate SNP. However some LP individuals have also been identified who carry none of these alleles. Aims To examine the distribution across Africa of these and other allelic variants; to examine other regulatory regions in population groups in which enhancer alleles are lacking. Results The geographic and ethnic distribution of -13907*G, -13910*T, -13915*G, -14009*G, and -14010*C in 10 different countries and 15 distinct ethnic groups across Africa (n=1221 individuals) is presented here. Several other variants in this enhancer region are also described here for the first time. These tightly clustered enhancer variants are more frequent in pastoralist milk drinking groups than agriculturalist populations and are associated with several different LCT core haplotypes. Two further candidate regulatory regions have been sequenced in the same populations including a 1000bp region immediately upstream from LCT where novel variants have been found. Conclusions The data support the notion that many different mutations do have a functional role in LP, and that the trait has arisen independently several times, being subject to the positive selection conferred by the increased ability to digest milk lactose by people in pastoralist societies.
Extreme Evolutionary Disparities Seen in Positive Selection Across Seven Complex Diseases
Genome-wide association studies (GWASs) have successfully illuminated disease-associated variation. But whether human evolution is heading towards or away from disease susceptibility remains an open question. We analyzed the seven diseases studied by the Wellcome Trust Control Case Consortium (WTCCC), to calculate the relative selective pressure at every significant loci. Results reveal striking differences between the seven studied diseases. We find evidence of recent positive selection in favor of alleles increasing the risk of Type 1 Diabetes (T1D), Crohn’s Disease (CD), Hypertension (HT), Rheumatoid Arthritis (RA), and Bipolar Disorder (BD). Riskassociated alleles (defined as the allele most strongly associated with disease among associated SNPs) for Type 2 Diabetes (T2D) fall largely within the random neutral region, and Coronary Artery Disease (CAD) shows less positive selection than expected by random. When only protective alleles are considered (defined as the allele least strongly associated with disease among associated SNPs), we find that SNPs only associated with T1D, CD, and RA appear to exhibit significant signatures of positive selection. There is significant asymmetry in the 96 SNPs strongly associated with T1D (pvalue ≤0.005) showing strong signs of positive selection, with 79 SNPs selecting for the risky allele, and only 17 SNPs selecting for the protective allele. Furthermore, selection patterns of Coronary Artery Disease (CAD) fall far below the expected levels of random, implying stable allele frequencies. Results reveal the evolutionary trajectories of T1D and CD favor risk alleles, possibly due to their simultaneous role in protection from infectious diseases. These results inform on current understanding of disease etiology, thus aiding efforts to discover novel approaches to disease treatment and prevention.
Detecting Natural Selection in the Human Genome from Pilot1 Data in the 1000 Genomes Project
Identifying signatures of natural selection in the human genome is of fundamental implication for the study of population evolution and for the biomedical research. The distribution of selection in genome will provide important functional information. Natural selection modify the level of variability within and between populations and shapes the pattern of genetic variations in the genome. Genetic variation in genome is the raw data for detection of natural selection. The 1000 Genomes Project produces whole genome sequencing data and offers a unique and great opportunity to scan the genome for signature of natural selection. Five statistics: Tajima’D, Fu and Li’s F, Achaz’s Y, Fay and Wu’s H and Zeng et al.’s E (based on comparing the site frequency spectrum within population) and Fst statistic (based on the measure of population subdivision) were applied to Pilot 1 data in 1,000 genome project to scan the entire genome for detection of selection, where 344 chromosomes from ASI, CEU and YRI were sequenced. A total of more than 20 million of variant sites, 4.8 millions common in three populations were identified. We calculated seven statistics in 10 kb and 100 kb windows across the genome for each population and obtained their empirical distributions. Results show that two kinds of windows analyses lead to the similar distributions. The proportional rank of the test statistic in a particular window compared with the overall empirical genomic distribution was taken as empirical P-value for that window. We identified 3,046 candidate selection regions in ASI population, 2,015 selection regions in CEU, and 2,204 selection regions in YRI at 5% empirical significance level in 10 kb by five statistics based on differences in frequency spectrum. Among 457 candidate genes of selection reported from PubMed, we detected 102 selection genes in ASI, 53 selection genes in CEU, and 101 selection genes in YRI and 11 selection genes common in three populations by familiar Tajima D test. By comparison we obtained 3.9 million SNPs and the whole genome’s fixation index about 0.10~0.11. By compared with the empirical genome-wide distribution of FST, we identified 5, 278 candidate selection regions at an empirical significance level of 2.5% from each of the 22 autosomal chromosomes. Among 581 identified selection regions by FST which were reported from literatures, we found that 294 selection regions overlap our results.
Genomic Landscape of Positive Natural Selection in North European Populations
Analysing genetic variation of human populations to detect loci that have been affected by positive natural selection is important for understanding adaptive history and phenotypic variation in humans. In this study, we analysed recent positive selection in Northern Europe from genome-wide datasets of 250 000 and 500 000 single nucleotide polymorphisms in a total of over 1000 individuals from Great Britain, Northern Germany, Eastern and Western Finland, and Sweden. Coalescent simulations were used to demonstrate that the integrated haplotype score (iHS) and long-range haplotype (LRH) statistics have sufficient power in genome-wide datasets of different sample sizes and SNP densities. Furthermore, the behavior of the FST statistic in closely related populations was characterized by allele frequency simulations. In the analysis of the North European dataset, dozens of regions in the genome showed strong signs of recent positive selection. Most of these regions have not been discovered in previous scans, and many contain genes with interesting functions (e.g. RAB38, INFG, NOS1AP, and APOE). In the putatively selected regions, we observed a statistically significant overrepresentation of genetic association to complex disease, which emphasizes the importance of the analysis of positive selection in understanding the evolution of human disease. Altogether, this study demonstrates the potential of genome-wide datasets to discover loci that lie behind evolutionary adaptation in different human populations.
Evidence of Indigenous American specific selection in skin pigmentation genes
Recent studies of selection in human pigmentation genes have focused on Old World populations, neglecting the evolutionary changes that have occurred in Indigenous American populations since their migration into the Americas. Previous research shows correlations between Indigenous American ancestry and skin pigmentation variation, suggesting a genetic role in the determination of skin pigmentation among these populations. However, few genes contributing to these differences have been described. To identify genes that may have undergone Indigenous American specific changes, this work examines signatures of selection in 82 pigmentation candidate genes by genotyping 88 indigenous individuals from Central and South America using the Affymetrix Genomewide Human SNP Array 6.0. The resulting 906,600 single nucleotide polymorphisms (SNPs) were surveyed for signatures of selection in the Indigenous American populations compared to the HapMap Phase I populations. Evidence of selection was identified using four measures selected for the complementarity of their approaches, including the reduction in heterozygosity (lnRH), Locus-Specific Branch Length (LSBL), Tajima’s D, and by examination of the haplotype block structure. When computing lnRH and LSBL as well as when examining changes in haplotype frequency, the East Asian and European HapMap populations were included because they are the most closely related populations available. These analyses differentiate the selective changes that appear to be shared among East Asian and Indigenous American populations from those that are unique to the Indigenous American populations. For each test, the top5%of the empirical distribution of results was examined and pigmentation genes falling in this tail of the distribution were considered to show statistically significant evidence of selection. Based on these analyses, 12 genes - ADAM17, POMC, AP3B1,OPRM1, SILV, OCA2/HERC, PLDN, MYO5A, RAB27A, CYP1A2, ATRN, and ASIP - show evidence of selection unique to the Indigenous American populations. Many of these genes have known functional roles in melanogenesis and suggest potential pathways responsible for the observed differences in skin pigmentation between Indigenous American and Old World populations.
Patterns of correlation between genetic ancestry and facial features suggest selection on females is driving differentiation.
Human facial features show extensive variation within and among populations. By investigating the relationship between dimorphism in facial features and genetic ancestry in different populations, we can explore the roles of sexual and natural selection on the human face. We measured sexual dimorphism in facial traits while controlling for the effects of overall size differences and then tested for interactions between sex and genetic ancestry. The study sample consists of 254 subjects (n=170 females, n=84 males), ages 18-35, showing West African and European genetic ancestry sampled in the United States and Brazil. Maximum likelihood genetic ancestry estimates were determined from 176 ancestry informative markers (AIMs), which allowed for the proportional estimation of genetic ancestry from four parental populations (West African, European, East Asian, and Native American). Three-dimensional photographs of faces were acquired using the 3dMDface imaging system (Atlanta, GA). 22 standard anthropometric landmarks were placed on each image and XYZ coordinates were collected. All 231 possible pairwise inter-landmark distances were calculated and then log transformed. Using the pairwise distances, we tested whether some distances were larger in one sex than the other, having taken size into account, in a) African Americans sampled in the United States, b) Brazilians sampled in Brazil, and c) the combined African American and Brazilian sample. We found that several pairwise distances differed between the sexes. For example, the distance from the brow to nasal bridge was found to be more than 5% larger in females than males. We then tested for an interaction between sex and genetic ancestry by testing for differences in the slopes of the ancestry association between males and females. Although the pattern differed slightly between samples, after Bonferroni correction many correlations were the found to be same in both sexes. However, females in all three samples had many additional significant correlations that were not seen in males, while males had very few correlations that were not found in females. The results of these analyses suggest that selection on females is driving the differentiation in facial features among populations.
Effect of natural selection on North Asian mitochondrial haplogroup variation
The human mtDNA exhibits striking, region-specific sequence variation. The regional distribution of mtDNA haplogroups have attributed either to genetic drift assisted by purifying selection (Elson et al., 2004; Kivisild et al., 2006; Ingman, Gyllensten, 2007) or to an adaptation to different climates (Mishmar et al., 2003; Ruiz-Pesini et al., 2004). In an attempt to study the mode of selection in mtDNA variation in human populations we sequenced and analyzed 211 complete mtDNA sequences belonging to haplogroups A, C and D accounting in total for 49.3% of mtDNA lineages in North Asia. The North Asian haplogroups A, C and D showed a highly significant deviation from the standard neutral model as well as a bell-shaped distribution of pairwise differences consistent with rapid population expansion. To determine the overall importance of selection in shaping human mtDNA variation we calculated Ka/Ks ratio both for aggregated mtDNAs and for 13 proteinencoding genes within particular haplogroups (A, C and D). We have found a prevalence of Ks over Ka within haplogroups A, C and D indicating the influence of negative selection on mtDNA during evolution. Consistent with some previous reports we have found the Ka/Ks ratio for the ATP6 gene to be the highest among the North Asian sequences suggesting thereby that this gene has been subject to positive selection. We have also observed a set of genes with a somewhat higher Ka/Ks ratio relative to other mitochondrial genes - CO2 for haplogroup A, ND3 and ND4 for haplogroup C. Meanwhile the other approach taking into account the difference in NS/S ratios between the haplogroup-associated and private substitutions (Elson et al., 2004) shows the significant departures from neutrality only for haplogroup D and its subhaplogroup D4. Furthermore single gene analysis reveals the relatively strong influence of negative selection only in CYTb gene within haplogroupD(p=0.011, NI=14.1). In general, our results indicate that there is an evidence for both gene-specific and lineage-specific variation in selection acting on North Asian mtDNAs.
Selection for blue eyes in Europe and light skin pigmentation in East Asia at OCA2/HERC2
OCA2 and HERC2 are two genes on chromosome 15 separated by lessthan 10 kb. Mutations in this region have been shown to have an effect onpigmentation including causing oculocutaneous albinism type 2. In Europeans,a three SNP haplotype (rs4778138, rs4778241, rs7495174) and threeindividual SNPs (rs12913832, rs916977, rs1667394) have been associatedwith blue eyes. We have labeled the three SNP haplotype BEH1. We foundthat the first individual SNP, rs12913832, was in near complete LD withanother SNP (rs1129038). Wetreat these two SNPs together as a haplotype,BEH2. We also found that the other two individual SNPs were actually innear complete LD with each other and decided to label them BEH3. In EastAsians, a SNP (rs1800414) has been identified that is associated with alight skin pigmentation phenotype. We typed these eight SNPs in 64-70population samples. We then examined worldwide distribution of the fourpigmentation alleles. We saw that the light skin allele was at its highestfrequency in eastern East Asia, at midrange frequencies in Southeast Asia,and at lower frequencies in western East Asia. It is virtually absent from therest of the world. BEH1 and BEH3 show very similar global patterns, lowfrequencies to midrange frequencies in Africa and East Asia, midrangefrequencies in India and Eastern Siberia, and midrange to high frequenciesin Southwest Asia, Europe, Western Siberia, the Pacific Islands, and theAmericas. BEH2 shows a different pattern from the other two. It showslow frequencies in East Africa, India, Eastern Siberia, and the Americas,midrange frequencies in Southwest Asians and Southern Europeans, andhigh frequencies in Eastern and Northwestern Europe and Western Siberia.We then typed additional SNPs and test each pigmentation allele for selectionusing the Relative Extended Haplotype Homozygosity (REHH) test. Wefound that the light skin allele of rs1800414 is under selection in East Asiaand that the blue eye allele of BEH2 is under selection in Europe andSouthwest Asia. We show light skin pigmentation has been selected for inEast Asia. This is likely due to lower UV exposure at the higher latitudes(compared to equatorial Africa) and the need for lighter skin for vitamin Dproduction. We also show that blue eyes are selected for in Europe. Thisis most likely due to sexual selection, though another unknown effect of thisparticular allele could be selected for and the blues eyes are a side effect.
Ancestry variation along the genome in Latin American populations and implications for recent natural selection
Latin American populations stem from the admixture starting about 500 years ago of Europeans, Africans and Native Americans. Extreme deviation in ancestry estimates at certain genome locations (relative to the genomewide average) could reflect the action of recent natural selection. We evaluated the distribution of ancestry estimates along the genome using 678 microsatellite markers in 249 individuals sampled from 13 admixed populations across Latin America. We found a significant deviation in ancestry at two genomic locations with more than four times standard deviations from the genome-wide mean: an excess of European ancestry at 14q32 (Zscore = 4.14), and an excess of African ancestry at 6p22 (Z-score = 4.71). These deviations in ancestry were observed in the analysis of the combined dataset as well as in most of the individual populations examined. We showed that our findings are robust to the Native American ancestry populations used. We discussed the implications for recent natural selection in the context of the unique history of the New World, as well as the possibility of artifacts.

September 03, 2009

Central European farmers not descended from local hunter-gatherers (Bramanti et al. 2009)

This is the real power of DNA: the topic of whether central European farmers were the result of demic diffusion from the southeast or indigenous hunter-gatherers who adopted the agricultural economy has been endlessly debated in archaeological circles.

We are finally in a position to give an answer to the question, and the answer is in favor of the diffusionist camp and against the idea of acculturation by local hunter-gatherers. Surprisingly, modern Central Europeans do not appear to be a simple hunter-gatherer/farmer mix, suggesting that even later events (post-Neolithic) have shaped their genetic diversity.

This study is also a powerful argument against the idea of genetic continuity across long time spans. Most ancient DNA studies so far have reached a similar conclusion. Thus, it also destroys the supposed justification for continuity from Paleolithic Europe to modern times that early mtDNA work (of the Daughters of Eve variety) has proposed, hand in hand with the hunter acculturation hypothesis.

The paper is covered in National Geographic:
Central and western Europe's first farmers weren't crafty, native hunter-gatherers who gradually gave up their spears for seeds, a new study says.

Instead, they were experienced outsiders who arrived on the scene around 5500 B.C. with animals in tow—and the locals apparently didn't roll out the welcome wagon.

"Within a few generations, all the farmers—probably coming from southeast Europe—moved into central Europe bringing their culture, [livestock], and everything," Joachim Burger, a molecular archaeologist at the University of Mainz in Germany, said via email.

The finding is based on analysis of genetic material in the skeletal remains of ancient hunter-gatherers and early farmers found in Germany, Lithuania, Poland, and Russia—though farming is thought to have reached areas as far west as western France during the period of rapid expansion, about 7,500 years ago.

The study goes against a long-standing idea that Europe's first farmers were former hunter-gatherer populations that had settled the region after the last ice age, about 10,000 years ago.

Perhaps, the thinking went, the hunter-gatherers had observed farming practices during their travels or had learned from neighbors.

Instead, the researchers found, the hunter-gatherers and the early farmers remained segregated, according to the study, to be published tomorrow in the journal Science.
And the press release:
Analysis of ancient DNA from skeletons suggests that Europe's first farmers were not the descendants of the people who settled the area after the retreat of the ice sheets. Instead, the early farmers probably migrated into major areas of central and eastern Europe about 7,500 years ago, bringing domesticated plants and animals with them, says Barbara Bramanti from Mainz University in Germany and colleagues. The researchers analyzed DNA from hunter-gatherer and early farmer burials, and compared those to each other and to the DNA of modern Europeans. They conclude that there is little evidence of a direct genetic link between the hunter-gatherers and the early farmers, and 82 percent of the types of mtDNA found in the hunter-gatherers are relatively rare in central Europeans today.

For more than a century archaeologists, anthropologists, linguists, and more recently, geneticists, have argued about who the ancestors of Europeans living today were. We know that people lived in Europe before and after the last big ice age and managed to survive by hunting and gathering. We also know that farming spread into Europe from the Near East over the last 9,000 years, thereby increasing the amount of food that can be produced by as much as 100-fold. But the extent to which modern Europeans are descended from either of those two groups has eluded scientists despite many attempts to answer this question.

Now, a team from Mainz University in Germany, together with researchers from UCL (University College London) and Cambridge, have found that the first farmers in central and northern Europe could not have been the descendents of the hunter-gatherers that came before them. But what is even more surprising, they also found that modern Europeans couldn't solely be the descendents of either the hunter-gatherer alone, or the first farmers alone, and are unlikely to be a mixture of just those two groups. "This is really odd", said Professor Mark Thomas, a population geneticist at UCL and co-author of the study. "For more than a century the debate has centered around how much we are the descendents of European hunter-gatherers and how much we are the descendents of Europe's early farmers. For the first time we are now able to directly compare the genes of these Stone Age Europeans, and what we find is that some DNA types just aren't there - despite being common in Europeans today."

Humans arrived in Europe 45,000 years ago and replaced the Neandertals. From that period on, European hunter-gatherers experienced lots of climatic changes, including the last Ice Age. After the end of the Ice Age, some 11,000 years ago, the hunter-gatherer lifestyle survived for a couple of thousand years but was then gradually replaced by agriculture. The question was whether this change in lifestyle from hunter-gatherer to farmer was brought to Europe by new people, or whether only the idea of farming spread. The new results from the Mainz-led team seems to solve much of this long standing debate.

"Our analysis shows that there is no direct continuity between hunter-gatherers and farmers in Central Europe," says Prof Joachim Burger. "As the hunter-gatherers were there first, the farmers must have immigrated into the area."

The study identifies the Carpathian Basin as the origin for early Central European farmers. "It seems that farmers of the Linearbandkeramik culture immigrated from what is modern day Hungary around 7,500 years ago into Central Europe, initially without mixing with local hunter gatherers," says Barbara Bramanti, first author of the study. "This is surprising, because there were cultural contacts between the locals and the immigrants, but, it appears, no genetic exchange of women."

The new study confirms what Joachim Burger´s team showed in 2005; that the first farmers were not the direct ancestors of modern European. Burger says "We are still searching for those remaining components of modern European ancestry. European hunter-gatherers and early farmers alone are not enough. But new ancient DNA data from later periods in European prehistory may shed also light on this in the future."
And from archaeology.about.com:
A new study published by Barbara Bramanti and colleagues in Science Express on September 4, 2009, supports what some scholars have suspected all along—that the LBK likely were an in-migration of people from the Balkans, and that they did not, initially anyway, do much mixing at all with the earlier inhabitants of Europe.

Bramanti and her colleagues compared the mitochondrial DNA from 20 central European Upper Paleolithic, Mesolithic and Neolithic hunter-gatherers to that from 25 Neolithic farmers and 484 modern Europeans, spanning an age range from about 13,400 to 2,300 BC. The data shows that the early farmers and hunter-gatherers were from distinctively different populations.

This paper follows up on and to a degree contradicts with the hypothesis of an earlier paper that looked only at mtDA of the Neolithic farmers. That study (Haak et al. 2005) discovered that the farmers had a distinctive difference between the current residents of Europe, and hypothesized that that meant that the hunter-gatherers might have been more like the modern inhabitants, and thus, the LBK would have been only a minor component.
The earlier paper by Haak et al. they refer to.

(More technical details once I read the full paper)

UPDATE:

Pre-farming populations seem to have been dominated by mtDNA haplogroup U:
it is intriguing to note that 82% of our 22 hunter-gatherer individuals carried clade U (fourteen U5, two U4, and two unspecified U-types; table 1).
The hunter-gatherers had no N1a -which was a signature of early farmers in the Haak et al. paper- or of haplogroup H, the most common mtDNA haplogroup in Europeans today. The only non-U types in hunter-gatherers were all from the Ostorf site and included haplogroups T2e, J, and K.

The farmers:
In a previous study, we showed that the early farmers of Central Europe carried mainly N1a, but also H, HV, J, K, T, V, and U3 types (11, 12). We found no U5 or U4 types in that early farmer sample.
UPDATE I:

It is important to note the implications of this study: the most certain conclusion is that Neolithic farmers in Central Europe are very sharply differentiated from the Paleolithic-Mesolithic populations. This is clear evidence in favor of the diffusionist idea, since the acculturation hypothesis predicts that the mtDNA of the early farmers would be roughly that of the pre-farming population that picked up the new technology.

However, the evidence of this paper also contradicts the plain demic diffusion hypothesis. According to this hypothesis, farmer genes are gradually replaced by hunter genes as the farming economy spreads, because in each step there is a mix of farmer-indigenous populations which go on to colonize regions beyond the frontier. This is not what appears to have happened. Rather, it seems the farmers moved across Europe with very little interaction with pre-farmers. A long period of no contact between the LBK and foragers is actually supported by archaeology. I have termed this type of diffusion the "skipping stone":
In the Skipping Stone model, farmers move out in search of new territories before they have started to blend with the local foragers; the genetic impact of the initiators of the movement is preserved.
The great speed of the Linearbandkeramik farmers was also experienced by farmers who spread across the Mediterranean. The spread of agriculture in Europe does not appear to have been a slow process of interaction between farmer and forager, but rather a blitz by the first farmers, followed later, after the spread had already occurred by admixture with some of the foragers that remained.

We must also be certain not to jump into conclusions about the relative contributions of farmer and forager in the modern gene pool. Clearly both the idea of a predominantly "Paleolithic" and a predominantly "Neolithic" gene pool is problematic; such continuity is not really evident. However, the reasons for the discontinuity up to the present may be manifold: e.g., later population movements into Europe, or natural selection changing the gene pool without subsequent change of population.

What we do know is this: first farmers were not local foragers who abandoned the old ways for the new ones. Amalgamation between farmer and forager did not happen quickly as the farming economy spread. Finally it did happen, of course, and either because (i) there were few foragers in the mix, or (ii) their mtDNA was selected against, modern central Europeans have very little mitochondrial descent from the earliest European populations.

PS: Natural selection against forager mtDNA is not very outlandish. For example, a severe reduction of U5a1 and U5b haplogroup in Britain from ancient to modern times has been observed, which could potentially mark another data point in a process of selection against that haplogroup over time.

My personal guess is that both demography and selection may have played a role in the marginalization of hunter-gatherer mtDNA . LBK farmers were already 3 thousand years removed from the earliest agriculturalists of the Near East, so it is conceivable that they had evolved an mtDNA gene pool adapted to the new lifestyle that outcompeted the indigenous European one. But, the long period of isolation from foragers may mean that only farmer mtDNA benefited from the demographic boom associated with the new economy, and by the time relations between the two groups warmed up, the relatively few newcomers already dwarfed the older population demographically.

UPDATE II (Sep 4):

To understand the magnitude of the difference between farmers and hunter-gatherers, the authors calculate their Fst=0.163, which can be compared with a maximum value of 0.0327 among modern Europeans and 0.133 for modern Eurasians from Europe to Australia. Subsequently, the authors test the hypotheses of (a) continuity between hunter-gatherers and farmers, and (b) continuity between hunter-gatherers and modern Central Europeans, rejecting both.

This isn't very surprising in the light of the anthropological evidence in favor of diffusion of farmers from the Near East and against the acculturation hypothesis presented recently by Pinhasi et al. The very close relationship of the LBK skulls and their proximity to samples from Nea Nikomedeia in Greece and Catal Hoyuk in Anatolia contrasts with the Mesolithic populations.

UPDATE III (Sep 21):

Some possible anthropological evidence for post-LBK infusion into Central Europe:
Mesolithic Europeans display considerable variation in humero-clavicular and brachial indices yet none approach the extreme "hyper-polar" morphology of LBK humans from the MESV. In contrast, Late Neolithic and Early Bronze Age peoples display elongated brachial and crural indices reminiscent of terminal Pleistocene and "tropically adapted" recent humans. These marked morphological changes likely reflect exogenous immigration during the terminal Fourth millennium cal BC.

Science doi:10.1126/science.1176869

Genetic Discontinuity Between Local Hunter-Gatherers and Central Europe’s First Farmers

B. Bramanti et al.

Following the domestication of animals and crops in the Near East some 11,000 years ago, farming reached much of Central Europe by 7,500 years before present. The extent to which these early European farmers were immigrants, or descendants of resident hunter-gatherers who had adopted farming, has been widely debated. We compare new mitochondrial DNA (mtDNA) sequences from late European hunter-gatherer skeletons with those from early farmers, and from modern Europeans. We find large genetic differences between all three groups that cannot be explained by population continuity alone. Most (82%) of the ancient hunter-gatherers share mtDNA types that are relatively rare in Central Europeans today. Together, these analyses provide persuasive evidence that the first farmers were not the descendants of local hunter-gatherers but immigrated into Central Europe at the onset of the Neolithic.

Link

September 02, 2009

A Single Origin for Dogs South of Yangtze River, less than 16,300 Years Ago (Pang et al. 2009)

Another recent study by Boyko et al. raised some doubts about the strength of the evidence for dog domestication in Asia, by pointing out that including semi-feral village dogs may increase the observed Asian diversity. The advance access manuscript for this paper is free, so anyone interested in the sampling details. The authors do cite the other recent paper:
Notably, in a recent study of African village dogs (Boyko et al. 2009) it was claimed that the reported high diversity for mtDNA in East Asia compared to other parts of the world (Savolainen et al. 2002), was the result of sampling bias. However, in the present study (see “Results”) we show this assertion to be incorrect.

...

Thus, a direct comparison shows that the smaller South Chinese sample has 73% more haplotypes than the African one; the assertion by Boyko et al. (2009) is the result of not adequately compensating for differences in sample size between the relatively small East Asian samples in Savolainen et al. (2002) and the larger African samples. The African sample has also all the other characteristics of the “western” dog populations: The haplotypes fall in the same parts of the MS networks as for other western populations, leaving large parts unique to East Asia (data not shown); and values are high for UT (66.7%) and UTd (90.9%), and number of unique haplotypes low (12) (compare with e.g. South China: UT (42.0%), UTd (53.4%), and number of unique haplotypes (40; i.e. only one less than the total number of haplotypes in the African sample!)). To conclude, the sample of African village dogs in Boyko et al. (2009), like all “western” samples, has considerably lower genetic variation than the populations in ASY.
Molecular Biology and Evolution, doi:10.1093/molbev/msp195

mtDNA Data Indicates a Single Origin for Dogs South of Yangtze River, less than 16,300 Years Ago, from Numerous Wolves

Jun-Feng Pang et al.

Abstract

There is no generally accepted picture of where, when, and how the domestic dog originated. Previous studies of mitochondrial DNA (mtDNA) have failed to establish the time and precise place of origin because of lack of phylogenetic resolution in the so far studied control region (CR), and inadequate sampling. We therefore analysed entire mitochondrial genomes for 169 dogs to obtain maximal phylogenetic resolution, and the CR for 1,543 dogs across the Old World for a comprehensive picture of geographical diversity. Hereby, a detailed picture of the origins of the dog can for the first time be suggested. We obtained evidence that the dog has a single origin in time and space, and an estimation of the time of origin, number of founders and approximate region, which also gives potential clues about the human culture involved. The analyses showed that dogs universally share a common homogenous gene pool containing 10 major haplogroups. However, the full range of genetic diversity, all 10 haplogroups, was found only in south-eastern Asia south of Yangtze River, and diversity decreased following a gradient across Eurasia, through 7 haplogroups in Central China, and 5 in North China and Southwest Asia, down to only 4 haplogroups in Europe. The mean sequence distance to ancestral haplotypes indicates an origin 5,400-16,300 years ago from at least 51 female wolf founders. These results indicate that the domestic dog originated in southern China less than 16,300 years ago, from several hundred wolves. The place and time coincide approximately with the origin of rice agriculture, suggesting that the dogs may have originated among sedentary hunter-gatherers or early farmers, and the numerous founders indicate that wolf taming was an important culture trait.

Link

Y chromosomes and mtDNA of Central Asian Turkic and Iranian populations

Unfortunately this paper only studied 12 Y-STRs, reduced to 7 to compare them with previous studies. Moreover, as far as I can see, this data is not available in the journal website. ScienceDaily covers the paper with the totally unwarranted title of "No Such Thing As Ethnic Groups, Genetically Speaking, Researchers Say".

What this paper does show, as far as its limited marker set can, that some Turkic ethnic groups are aggregates of populations of unrelated origin, which is not particularly surprising. One has to look at the spread of Turks from Central Asia to Europe to see that the various "Turks" were usually opportunistic alliances of peoples of different stock. Perhaps these Central Asian ethnic groups will eventually be homogenized by continued in-group marriage.

This brings me to an important point of using genetic diversity to assess how long ago ethnic groups were formed. Some ethnic groups begin as homogeneous entities which become differentiated as they expand and undergo separate evolution/differential patterns of admixture in different localities. Other groups begin as heterogeneous groups of unrelated tribes that become united by some factor, e.g., the emergence of a strong king or dynasty around whom diverse peoples aggregate. In the first case, ethnic evolution is one of diversification over time, as the genetic legacy of the homogeneous founders is fragmented; in the latter, it is one of homogenization, as the genetic legacies of the heterogeneous constituents merge to form a single homogeneous group. Thus, one can't generally conclude, by looking at within-group differentiation whether the group is "old" or "recent" in origin; it could be an old group that has fragmented over a long period of time, or a new one that has not had enough time to become one.

UPDATE (Sep 3):

The study is also covered at the Spittoon under the title New Study on Genetics of Ethnic Groups Reveals We May Not Be So Different After All. Unfortunately, the author of the blog post gets the linguistic divisions wrong:
The Turks are largely nomadic herders. They speak Indo-Iranian languages like Azerbaijani, Turkish, and Altay. Their society is organized into clans, or “descent groups,” whose membership is passed down from father to children.

The Tajiks are, conversely, agriculturalists. They speak various dialects of the the Tajik, or Tajik Persian, language that may have arrived with Muslim invaders 1,000 years ago. Their society is largely patrilocal – meaning that when couples marry they put up residence near the husband’s family; and first cousin marriages are encouraged.
It is of course the Turks who speak Turkic languages, and the Tajiks who speak an Indo-Iranian (or more precisely Iranian) language.

As for the title, which, like the ScienceDaily title, seems to burst at the seams with delight that ethnic differences don't exist, a better angle on the topic would be to observe how prevalent ethnic differences are, if they can exist even among populations that are genetically non-differentiated. The pipe dream of some thinkers, that increased inter-ethnic and inter-racial mating will lead to an abolition of ethnic and racial genetic differences, and, thus, to a happy co-existence of people around the world, is refuted by the finding that humans happily self-segregate themselves along ethnic lines, even when there are no underlying genetic differences.

BMC Genetics 2009, 10:49doi:10.1186/1471-2156-10-49

Genetic diversity and the emergence of ethnic groups in Central Asia

Evelyne Heyer et al.

Abstract

Background

In this study, we used genetic data that we collected in Central Asia, in addition to data from the literature, to understand better the origins of Central Asian groups at a fine-grained scale, and to assess how ethnicity influences the shaping of genetic differences in the human species. We assess the levels of genetic differentiation between ethnic groups on one hand and between populations of the same ethnic group on the other hand with mitochondrial and Ychromosomal data from several populations per ethnic group from the two major linguistic groups in Central Asia.

Results

Our results show that there are more differences between populations of the same ethnic group than between ethnic groups for the Y chromosome, whereas the opposite is observed for mtDNA in the Turkic group. This is not the case for Tajik populations belonging to the Indo-Iranian group where the mtDNA like the Y-chomosomal differentiation is also significant between populations within this ethnic group. Further, the Y-chromosomal analysis of genetic differentiation between populations belonging to the same ethnic group gives some estimation of the minimal age of these ethnic groups. This value is significantly higher than what is known from historical records for two of the groups and lends support to Barth's hypothesis by indicating that ethnicity, at least for these two groups, should be seen as a constructed social system maintaining genetic boundaries with other ethnic groups, rather than the outcome of common genetic ancestry

Conclusions

Our analysis of uniparental markers highlights in Central Asia the differences between Turkic and Indo-Iranian populations in their sex-specific differentiation and shows good congruence with anthropological data.

Link

September 01, 2009

How humans differ from animals in height and mass variation (McKellar & Hendry 2009)

The authors found that humans within populations have more variation in mass than most animals do. In other words, there are many "thin" and "fat" people in human populations. This isn't very surprising to me, because in developing countries, socioeconomic differences may account for these differences (e.g., some people starve), while in developed countries, most people are employed in jobs and perform activities where having an optimal body mass is not that important for your survival and reproduction.

When it comes to height, humans show a very low within-population differentiation. In other words, most humans are around the "average" height, and really short and really tall ones are not that common. My guess is that this has something to do with the extreme socialization of humans; really short and really tall individuals (although the patterns are gender-specific) do have trouble in human society, if we judge from marriage ads where desired height is often specified, or from various pieces of technology (shields, spears, doors, steps, clothes, etc) which are designed for people of a particular height.

However, between-population differentiation in height is substantial. While we are in the 8th/4th percentiles in our within-population differentiation (very uniform), we are in the 47th/51st percentiles in our between-population differentiation. Thus, while in absolute terms between-population differences are average, these contrast greatly with our very low within-population differences: Human populations appear to be very different from each other in terms of their height.

From the paper:
One interesting result was that humans, in comparison to other animals, show a high level of within-population variation in mass considering their within-population variation in height (Figure 1). Specifically, when considering residuals from a regression of within-population CVs for mass on within-population CVs for length, human males and females fell into the 71st and 91st percentiles, respectively, for the entire distribution of animal species.

...

Another interesting result was that humans show low within-population variation in body height in comparison to body length in non-human animals (Figure 2), but the same was not true for human mass relative to animal mass (Figure S1). These differences can be quantified through several different comparisons. First, the mean within-population CVs for male and female human height correspond to the 8th and 4th percentiles, respectively, of the mean within-population CVs for animal length. In contrast, the mean within-population CVs for male and female human mass correspond to the 56th and 60th percentiles, respectively, of the within-population CVs for animal mass.
...
Specifically, the mean among-population CVs for male and female human height correspond to the 47th and 51st percentiles, respectively, of mean among-population CVs for animal length. Illustrated another way, humans show relatively low levels of within-population variation in height given their among-population variation in height (Figure 3).

PLoS ONE 4(9): e6876. doi:10.1371/journal.pone.0006876

How Humans Differ from Other Animals in Their Levels of Morphological Variation

Ann E. McKellar, Andrew P. Hendry

Abstract

Animal species come in many shapes and sizes, as do the individuals and populations that make up each species. To us, humans might seem to show particularly high levels of morphological variation, but perhaps this perception is simply based on enhanced recognition of individual conspecifics relative to individual heterospecifics. We here more objectively ask how humans compare to other animals in terms of body size variation. We quantitatively compare levels of variation in body length (height) and mass within and among 99 human populations and 848 animal populations (210 species). We find that humans show low levels of within-population body height variation in comparison to body length variation in other animals. Humans do not, however, show distinctive levels of within-population body mass variation, nor of among-population body height or mass variation. These results are consistent with the idea that natural and sexual selection have reduced human height variation within populations, while maintaining it among populations. We therefore hypothesize that humans have evolved on a rugged adaptive landscape with strong selection for body height optima that differ among locations.

Link

Coalescent-based serial founder model of migration outward from Africa (DeGiorgio et al. 2009)

Proc Natl Acad Sci U S A. 2009 Aug 17. [Epub ahead of print]

Out of Africa: Modern Human Origins Special Feature: Explaining worldwide patterns of human genetic variation using a coalescent-based serial founder model of migration outward from Africa.

Degiorgio M, Jakobsson M, Rosenberg NA.

Studies of worldwide human variation have discovered three trends in summary statistics as a function of increasing geographic distance from East Africa: a decrease in heterozygosity, an increase in linkage disequilibrium (LD), and a decrease in the slope of the ancestral allele frequency spectrum. Forward simulations of unlinked loci have shown that the decline in heterozygosity can be described by a serial founder model, in which populations migrate outward from Africa through a process where each of a series of populations is formed from a subset of the previous population in the outward expansion. Here, we extend this approach by developing a retrospective coalescent-based serial founder model that incorporates linked loci. Our model both recovers the observed decline in heterozygosity with increasing distance from Africa and produces the patterns observed in LD and the ancestral allele frequency spectrum. Surprisingly, although migration between neighboring populations and limited admixture between modern and archaic humans can be accommodated in the model while continuing to explain the three trends, a competing model in which a wave of outward modern human migration expands into a series of preexisting archaic populations produces nearly opposite patterns to those observed in the data. We conclude by developing a simpler model to illustrate that the feature that permits the serial founder model but not the archaic persistence model to explain the three trends observed with increasing distance from Africa is its incorporation of a cumulative effect of genetic drift as humans colonized the world.

Link