Showing posts with label Bulgaria. Show all posts
Showing posts with label Bulgaria. Show all posts

March 09, 2015

Sveti Ivan relics Middle Eastern, mid-1st century AD

Irrespective of one's religious beliefs, a genome from the mid-1st century AD would be interesting. I am personally in favor of the scientific study of saints' relics.

Bulgarian bones could be John the Baptist's, scientists say
So when Bulgarian archeologists announced in 2010 that they had found the bones of John the Baptist, Tom Higham was skeptical.

He got a surprise.

Higham, an Oxford University scientist and an atheist who doesn't believe in "any kind of religion or God or anything like that," was asked to test six small bone fragments found on an island named Sveti Ivan - St. John.

The bones turned out to be from a man who lived in the Middle East at the same time as Jesus, Higham said.

"We got a date that was exactly where it should be, right in the middle of the first century," said Higham, a radiocarbon dating expert.
...

DNA testing by colleagues at the University of Copenhagen suggested that the person was most likely to have been from the Middle East, he said.
"We have a complete genome. It's possible that we could step this a step further and see if there is any similarity," in the genetic material of all the relics.
St. John's father Zechariah belonged to to the Aaronic line of priests. While modern Jewish priests (Cohanim) belong to multiple Y-chromosomal lineages, I think it's a good bet that at least one or a few of these lineages could be traced to Jewish priests of 2,000 years ago (leaving the question of the ultimate Aaronic lineage open). If the remains from Sveti Ivan island match one of these lineages, then this would be a powerful piece of evidence in favor of their attribution to St. John the Baptist.

September 10, 2014

ASHG 2014 titles and abstracts

Some interesting titles from the ASHG 2014 conference.

UPDATE: I have added the abstracts.

The human X chromosome is the target of megabase wide selective sweeps associated with multi-copy genes expressed in male meiosis and involved in reproductive isolation. M. H. Schierup, K. Munch, K. Nam, T. Mailund, J. Y. Dutheil.
   The X chromosome differs from the autosomes in its hemizogosity in males and in its intimate relationship with the very different Y chromosome. It has a different gene content than autosomes and undergo specific processes such as meiotic sex chromosome inactivation (MSCI) and XY body formation. Previous studies have shown that natural selection is more efficient against deleterious mutations and, in chimpanzee, that positive selection is prevalent. We show that in all great apes species, megabase wide regions of the X chromosome has severely reduced diversity (by more than 80%). These regions are partly shared among species and indicate a large number of strong selective sweeps that have occurred independently on the same set of targets in different great apes species. We use simulations and deterministic calculations to show that background selection or soft selective sweeps are unlikely to be responsible. The regions also bear all the hallmarks of selective sweeps such as an increased proportion of singletons and higher divergence among closely related populations. Human populations are differently affected, suggesting that a large fraction of sweeps are private to specific human populations. The regions of reduced diversity correlates strongly with the position of X-ampliconic regions, which are 100-500 kb regions containing multiple copies of genes that are solely expressed during male meiosis. We propose that the genes in these regions escape MSCI and participate in an intragenomic conflict with regions of similar function on the Y chromosome for transmission of sex chromosomes to the next generation, i.e. sex chromosome meiotic drive. Recent results from Neanderthal introgression into humans point to the same regions as showing no introgression, consistent with the above process leading to reproductive isolation. Strikingly, the same regions of the X also shows much reduced divergence between human and chimpanzee, suggesting either that this speciation process was indeed complex or that the same regions were under strong selection in the human chimpanzee ancestor.
New insights on human de novo mutation rate and parental age. W. S. W. Wong, B. Solomon, D. Bodian, D. Thach, R. Iyer, J. Vockley, J. Niederhuber.
 Germline mutations have a major role to play in evolution. Much attention has been given to studying the pattern and rate of human mutations using biochemical or phylogenetic methods based on closely related species. Massively parallel sequencing technologies have given scientists the opportunity to study directly measured de novo mutations (DNMs) at an unprecedented scale. Here we report the largest study (to our knowledge) of de novo point mutations in humans, in which we used whole genome deep sequencing (~60x) data from 605 family trios (father, mother and newborn). These trios represent the first group of approximately 2,700 trios who have undergone whole-genome sequencing (WGS) through our pediatric-based WGS research studies. The fathers ages range from 17 to 63 years and the mothers ages range from 17 to 43 years. We identified over 23000 DNMs (~40 per newborn) in the autosomal chromosomes using a customized pipeline and infer that the mutation rate per basepair is around 1.2x10-8 per generation, well within the reported range in previous studies. We were also able to confirm that the total number of DNMs in the newborn was directly proportional to the paternal age (P  less than 2x10-16). Maternal age is shown to have a small but significant positive effect on the number of DNMs passed onto the offspring, (P =0.003) , even after accounting for the paternal age. This contradicts the prior dogma that maternal age only has an effect on chromosomal abnormalities related to nondisjunction events. Furthermore, 5% (22 total) of newborns in the analyzed group were conceived with assisted reproductive technologies (ARTs), and these infants have on average 5 more DNMs (Bias corrected and accelerated bootstrap 95% Confidence Interval, 1.24 to 8.00) than those conceived naturally, after controlling for both parents ages. Both parents ages remain significant as independently correlated with DNMs even after the families that used ARTs were removed from the analysis. Our study enhances current knowledge related to the human germline mutational rates.
Alignment to an ancestry specific reference genome discovers additional variants among 1000 Genomes ASW Cohort. R. A. Neff, J. Vargas, G. H. Gibbons, A. R. Davis.
   Whole genome sequencing studies across certain populations, such as those with African ancestry, are often underpowered due to a larger divergence between the common reference genome and the true genetic sequence of the population. However, a common reference genome is not designed to account for this divergence in population-specific studies. Strong signals from common (MAF>50%) single nucleotide polymorphisms (SNPs), insertion-deletions (indels), and structural variants (SVs) can make alignment and variant calling difficult by masking nearby variants with weaker genetic signals. We present the results generated from alignment to an African descent population-specific reference genome by applying variants present in a majority of individuals with African descent from all phases of the 1000 Genomes Project and the International HapMap Consortium. We identified 882,826 single nucleotide polymorphisms, short insertion-deletion events, and large structural variations present at MAF>50%; in the population, representing 2.39 MB of genetic variation changed from hg19. We demonstrate that utilization of a population-specific reference improves variant call quality, coverage level, and imputation accuracy. We compared alignment of 27 African-American SW population (ASW) samples from the 1000 Genomes Phase 1 project between the population-specific and the hg19 reference. We discovered an additional 443,036 SNPs by alignment to the population specific reference in union across all samples, including thousands of exonic variants that are non-synonymous and are clinically relevant to the study of disease.
Using compressed data structures to capture variation in thousands of human genomes. S. A. McCarthy, Z. Lui, J. T. Simpson, Z. Iqbal, T. M. Keane, R. Durbin.
   Currently the most widely used approach to catalogue variation amongst a set of samples is to align the sequencing reads to a single linear reference genome. This principle has been at the core of the 1000 Genomes data processing pipeline since the pilot phase of the project. However, there is now an increased awareness of the limitations of this approach, such as alignment artefacts, reference bias and unobserved variation on non-reference haplotypes. The Burrows-Wheeler transform and FM-index are compact data structures that have been successfully used in sequence alignment and assembly. One of the key features of these structures is that they are a searchable and reference-free representation of the raw sequencing reads. Our project aims to build a web server based on BWT data structures containing all the reads from many thousands of samples so as to efficiently retrieve matching reads and information about samples and populations. Enticingly, it is expected that data storage for this system would plateau as we collect more data since most new sequencing reads will have already been observed. We expect this to enable powerful new ways to query variation data from thousands of individuals. For the first phase of this project, we include all 87 Tbp of the low-coverage and exome data from the 2,535 samples in 1000 Genomes Phase 3. We envisage this would provide a means for researchers to easily check the prevalence of any human sequence in a control set of thousands of putatively healthy samples. We present our approaches and initial benchmarks on variant sensitivity and specificity against truth datasets and explore several applications for these structures such as validation of short insertion/deletion and structural variant calls, and rapid searching for traces of viral DNA.
Second-generation PLINK: Rising to the challenge of larger and richer datasets. C. C. Chang, C. C. Chow, L. C. A. M. Tellier, S. Vattikuti, S. M. Purcell, J. J. Lee.
   PLINK 1 is a widely used open-source C/C++ toolset for genome-wide association studies (GWAS) and research in population genetics. However, the steady accumulation of data from imputation and whole-genome sequencing studies has exposed a strong need for even faster and more scalable implementations of key functions. In addition, GWAS and population-genetic data now frequently contain probabilistic calls, phase information, and/or multiallelic variants, none of which can be represented by PLINK 1's primary data format.    To address these issues, we are developing a second-generation codebase for PLINK. The first major release from this codebase, PLINK 1.9, introduces extensive use of bit-level parallelism, O(sqrt(n))-time/constant-space Hardy-Weinberg equilibrium and Fisher's exact tests, and many other algorithmic improvements. In combination, these changes accelerate most operations by 1-4 orders of magnitude, and allow the program to handle datasets too large to fit in RAM. This will be followed by PLINK 2.0, which will introduce (a) a new data format capable of efficiently representing probabilities, phase, and multiallelic variants, and (b) extensions of many functions to account for the new types of information.    The second-generation versions of PLINK will offer dramatic improvements in performance and compatibility. For the first time, users without access to high-end computing resources can perform several essential analyses of the feature-rich and very large genetic datasets coming into use.
Exploring genetic variation and genotypes among millions of genomes. R. M. Layer, A. R. Quinlann.
Integrated analysis of protein-coding variation in over 90,000 individuals from exome sequencing data. D. G. MacArthur, M. Lek, E. Banks, R. Poplin, T. Fennell, K. Samocha, B. Thomas, K. Karczewski, S. Purcell, P. Sullivan, S. Kathiresan, M. I. McCarthy, M. Boehnke, S. Gabriel, D. M. Altshuler, G. Getz, M. J. Daly, Exome Aggregation Consortium.
   Rare, and thus largely unknown, variants are a major reason that, typically, less than 10% of the heritability of complex diseases currently can be explained by known genetic variation. While increasing the number of sequenced genomes may improve our ability to reveal this “hidden heritability,” the scale of the resulting dataset poses substantial storage and computational demands. Current efforts to sequence 100,000 genomes, and combined efforts that are likely to surpass 1 million genomes will identify hundreds of millions to billions of polymorphic loci. The minimum storage requirement for directly representing the variability found by these projects (1 bit per individual per variant, ignoring the necessary metadata) will range from terabytes to petabytes. Like most big-data problems, a balance must be found between optimizing storage and computational efficiency. For example, while compression can minimize storage by reducing file size, it can also cause inefficient computation since data must be decompressed before it can be analyzed. Conversely, highly structured data can reduce analysis times but typically require extra metadata that increase file size. Current variation storage schemes were not designed to quickly analyze massive datasets and fail to balance these competing goals. We present GENOTQ, an open source API and toolkit that reduces file size and data access time through use of a succinct data structure, a class of data structures that compress data such that operations can be performed without requiring the full decompression. Word aligned hybrid (WAH) bitmap compression is one such data structure that was developed to improve query times for relational databases. Binary values are encoded such that logical operations (AND, OR, NOT) can be performed on the compressed data. This encoding results in file sizes that are 20X smaller than uncompressed versions, and only 50% larger than the compressed version. Queries, such as finding shared variants among a subpopulation, are also 21X faster. Furthermore, representing the genotypes in this manner makes our method well suited to both distributed architectures like BigQuery and parallel processors like GPUs. We stress that this method is only part of a larger solution that would incorporate genomic annotations, medical histories, and pedigrees. Incorporating fast genotype queries with this web of metadata will provide a rich information source to both clinicians and researchers.
Capture of 390,000 SNPs in dozens of ancient central Europeans reveals a population turnover in Europe thousands of years after the advent of farming. I. Lazaridis, W. Haak, N. Patterson, N. Rohland, S. Mallick, B. Llamas, S. Nordenfelt, E. Harney, A. Cooper, K. W. Alt, D. Reich.
   To understand the population transformations that took place in Europe since the early Neolithic, we used a DNA capture technique to obtain reads covering ~390 thousand single nucleotide polymorphisms (SNPs) from a number of different archaeological cultures of central Europe (Germany and Hungary). The samples spanned the time period from 7,500 BP to 3,500 BP (Early Neolithic to Early Bronze Age periods) and most of them were previously studied using mtDNA (Brandt, Haak et al., Science, 2013). The captured SNPs include about 360,000 SNPs from the Affymetrix Human Origins Array that were discovered in African individuals, as well as about 30,000 SNPs chosen for other reasons (that are thought to have been affected by natural selection, or to have phenotypic effects, or are useful in determining Y-chromosome haplogroups). By analyzing this data together with a dataset of 2,345 present-day humans and other published ancient genomes, we show that late Neolithic inhabitants of central Europe belonging to the Corded Ware culture were not a continuation of the earlier occupants of the region. Our results highlight the importance of migration and major population turnover in Europe long after the arrival of farming. * Contributed equally to this work.
Insights into British and European population history from ancient DNA sequencing of Iron Age and Anglo-Saxon samples from Hinxton, England. S. Schiffels, W. Haak, B. Llamas, E. Popescu, L. Loe, R. Clarke, A. Lyons, P. Paajanen, D. Sayer, R. Mortimer, C. Tyler-Smith, A. Cooper, R. Durbin.
   British population history is shaped by a complex series of repeated immigration periods and associated changes in population structure. It is an open question however, to what extent each of these changes is reflected in the genetic ancestry of the current British population. Here we use ancient DNA sequencing to help address that question. We present whole genome sequences generated from five individuals that were found in archaeological excavations at the Wellcome Trust Genome Campus near Cambridge (UK), two of which are dated to around 2,000 years before present (Iron Age), and three to around 1,300 years before present (Anglo-Saxon period). Good preservation status allowed us to generate one high coverage sequence (12x) from an Iron Age individual, and four low coverage sequences (1x-4x) from the other samples.   By providing the first ancient whole genome sequences from Britain, we get a unique picture of the ancestral populations in Britain before and after the Anglo-Saxon immigrations. We use modern genetic reference panels such as the 1000 Genomes Project to examine the relationship of these ancient samples with present day population genetic data. Results from principal component analysis suggest that all samples fall consistently within the broader Northern European context, which is also consistent with mtDNA haplogroups. In addition, we obtain a finer structural genetic classification from rare genetic variants and haplotype based methods such as FineStructure. Reflecting more recent genetic ancestry, results from these methods suggest significant differences between the Iron Age and the Anglo-Saxon period samples when compared to other European samples. We find in particular that while the Anglo-Saxon samples resemble more closely the modern British population than the earlier samples, the Iron Age samples share more low frequency variation than the later ones with present day samples from southern Europe, in particular Spain (1000GP IBS). In addition the Anglo-Saxon period samples appear to share a stronger older component with Finnish (1000GP FIN) individuals. Our findings help characterize the ancestral European populations involved in major European migration movements into Britain in the last 2,000 years and thus provide more insights into the genetic history of people in northern Europe.
Fine-scale population structure in Europe. S. Leslie, G. Hellenthal, S. Myers, P. Donnelly, International Multiple Sclerosis Genetics Consortium.
   There is considerable interest in detecting and interpreting fine-scale population structure in Europe: as a signature of major events in the history of the populations of Europe, and because of the effect undetected population structure may have on disease association studies. Population structure appears to have been a minor concern for most of the recent generation of genome-wide association studies, but is likely to be important for the next generation of studies seeking associations to rare variants. Thus far, genetic studies across Europe have been limited to a small number of markers, or to methods that do not specifically account for the correlation structure in the genome due to linkage disequilibrium. Consequently, these studies were unable to group samples into clusters of similar ancestry on a fine (within country) scale with any confidence. We describe an analysis of fine-scale population structure using genome-wide SNP data on 6,209 individuals, sampled mostly from Western Europe. Using a recently published clustering algorithm (fineSTRUCTURE), adapted for specific aspects of our analysis, the samples were clustered purely as a function of genetic similarity, without reference to their known sampling locations. When plotted on a map of Europe one observes a striking association between the inferred clusters and geography. Interestingly, for the most part modern country boundaries are significant i.e. we see clear evidence of clusters that exclusively contain samples from a single country. At a high level we see: the Finns are the most differentiated from the rest of Europe (as might be expected); a clear divide between Sweden/Norway and the rest of Europe (including Denmark); and an obvious distinction between southern and northern Europe. We also observe considerable structure within countries on a hitherto unseen fine-scale - for example genetically distinct groups are detected along the coast of Norway. Using novel techniques we perform further analyses to examine the genetic relationships between the inferred clusters. We interpret our results with respect to geographic and linguistic divisions, as well as the historical and archaeological record. We believe this is the largest detailed analysis of very fine-scale human genetic structure and its origin within Europe. Crucial to these findings has been an approach to analysis that accounts for linkage disequilibrium.
The population structure and demographic history of Sardinia in relationship to neighboring populations. J. Novembre, C. Chiang, J. Marcus, C. Sidore, M. Zoledziewska, M. Steri, H. Al-asadi, G. Abecasis, D. Schlessinger, F. Cucca.
   Numerous studies have made clear that Sardinian populations are relatively isolated genetically from other populations of the Mediterranean, and more recently, intriguing connections between Sardinian ancestry and early Neolithic ancient DNA samples have been made. In this study, we analyze a whole-genome low-coverage sequencing dataset from 2120 Sardinians to more fully characterize patterns of genetic diversity in Sardinia. The study contains one subsample that contains individuals from across Sardinia and a second subsample that samples 4 villages from the more isolated Ogliastra region. We also merge the data with published reference data from Europe and North Africa. Overall Fst values of Sardinia to other European populations are low (less than 0.015); however using a novel method for visualizing genetic differentiation on a geographic map, we formally show how Sardinia is more differentiated than would be expected given its geographic distance from the mainland, consistent with periods of isolation. Applications of the software Admixture show how Sardinia populations differ in the levels of recent admixture with mainland European populations and that there are only minor contributions from North African populations to Sardinian ancestry. Notably the Sardinians from Ogliastra contain a distinct genetic cluster with minimal evidence of recent admixture with mainland Europe. We found frequency-based f3 tests and the tree-based algorithm Treemix both also show minimal evidence of recent admixture. Given the relative isolation, one might expect to see a unique demographic history from neighboring populations. Using coalescent-based approaches, we find Sardinian populations have had more constant effective sizes over the past several thousand years than mainland European populations, which typically show evidence for rapid growth trajectories in the recent past. This unique demographic history has consequences for the abundance of putatively damaging and deleterious variants, and we use our data to address the prediction that the genetic architecture of disease traits is expected to involve fewer loci with a greater proportion of variants at common frequencies in Sardinia.
Population structure in African-Americans. S. Gravel, M. Barakatt, B. Maples, M. Aldrich, E. E. Kenny, C. D. Bustamante, S. Baharian.
   We present a detailed population genetic study of 4 African-American cohorts comprising over 6000 genotyped individuals across US urban and rural communities: two nation-wide longitudinal cohorts, one biobank cohort, and the 1000 genomes ASW cohort. Ancestry analysis reveals a uniform breakdown of continental ancestry proportions across regions and urban/rural status, with 79% African, 19% European, and 1.5% Native American/Asian ancestries, with substantial between-individual variation. The Native Ancestry proportion is higher than previous estimates and is maintained after self-identified hispanics and individuals with substantial inferred Spanish ancestry are removed. This strongly supports direct admixture between Native Americans and African Americans on US territory, and linkage patterns suggest contact early after African-American arrival to the Americas. Local ancestry patterns and variation in ancestry proportions across individuals are broadly consistent with a single African-American population model with early Native American admixture and ongoing European gene flow in the South. The size and broad geographic sampling of our cohorts enables detailed analysis the geographic and cultural determinants of finer-scale population structure. Recent Identity-by-descent analysis reveals fine-scale structure consistent with the routes used during slavery and in the great African-American migrations of the twentieth century: east-to-west migrations in the south, and distinct south-to-north migrations into New England and the Midwest. These migrations follow transit routes available at the time, and are in stark contrast with European-American relatedness patterns.
Genetic testing of 400,000 individuals reveals the geography of ancestry in the United States. Y. Wang, J. M. Granka, J. K. Byrnes, M. J. Barber, K. Noto, R. E. Curtis, N. M. Natalie, C. A. Ball, K. G. Chahine.
   The population of the United States is formed by the interplay of immigration, migration and admixture. Recent research (R. Sebro et al., ASHG 2013) has shed light on the U.S. demography by studying the self-reported ethnicity from the 2010 U.S. Census. However, self-reported ethnicity may not accurately represent true genetic ancestry and may therefore introduce unknown biases. Since launching its DNA service in May 2012, AncestryDNA has genotyped over 400, 000 individuals from the United States. Leveraging this huge volume of DNA data, we conducted a large-scale survey of the ancestry of the United States. We predicted genetic ethnicity for each individual, relying on a rigorously curated reference panel of 3,000 single-origin individuals. Combining that with birth locations, we explored how various ethnicities are distributed across the United States Our results reveal a distinct spatial distribution for each ethnicity. For example, we found that individuals from Massachusetts have the highest proportion of Irish genetic ancestry and individuals from New York have the highest proportion of Southern European genetic ancestry, indicating their unique immigration and migration histories. We also performed pairwise IBD analysis on the entire sample set and identified over 300 million shared genomic segments among all 400,000 individuals. From this data, we calculated the average amount of sharing for pairs of individuals born within the same state or from two different states. In general, we found the genetic sharing decreases as the geographic distance between two states increases. However, the pattern also varies substantially among the 50 states. In summary, our analysis has provided significant insight on the biogeographic patterns of the ancestry in the United States.
Statistical inference of archaic introgression and natural selection in Central African Pygmies. P. Hsieh, J. D. Wall, J. Lachance, S. A. Tishkoff, R. N. Gutenkunst, M. F. Hammer.
   Recent evidence from ancient DNA studies suggests that genetic material introgressed from archaic forms of Homo, such as Neanderthals and Denisovans, into the ancestors of contemporary non-African populations. These findings also imply that hybridization may have given rise to some of adaptive novelties in anatomically modern humans (AMH) as they expanded from Africa into various ecological niches in Eurasia. Within Africa, fossil evidence suggests that AMH and a variety of archaic forms coexisted for much of the last 200,000 years. Here we present preliminary results leveraging high quality whole-genome data (>60X coverage) for three contemporary sub-Saharan African populations (Biaka, Baka, and Yoruba) from Central and West Africa to test for archaic admixture. With the current lack of African ancient DNA, especially in Central Africa due to its rainforest environment, our statistical inference approach provides an alternative means to understand the complex evolutionary dynamics among groups of the genus Homo. To identify candidate introgressive loci, we scan the genomes of 16 individuals and calculate S*, a summary statistic that was specifically designed by one of us (JDW) to detect archaic admixture. The significance of each candidate is assessed through extensive whole-genome level simulations using demographic parameters estimated by ∂a∂i to obtain a parametric distribution of S* values under the null hypothesis of no archaic introgression. As a complementary approach, top candidates are also examined by an approximate-likelihood computation method. The admixture time for each individual introgressive variant is inferred by estimating the decay of the genetic length of the diverged haplotype as a function of its underlying recombination rate. A neutrality test that controls for demography is performed for each candidate to test the hypothesis that introgressive variants rose to high frequency due to positive directional selection. Several genomic regions were identified by both selection and introgression scans, and we will discuss the possible genetic and functional properties of these “double-hits”. The present study represents one of the most comprehensive genomic surveys to date for evidence of archaic introgression to anatomically modern humans in Africa.
Inferences about human history and natural selection from 280 complete genome sequences from 135 diverse populations. S. Mallick, D. Reich, Simons Genome Diversity Project Consortium.
   The most powerful way to study population history and natural selection is to analyze whole genome sequences, which contain all the variation that exists in each individual. To date, genome-wide studies of history and selection have primarily analyzed data from single nucleotide polymorphism (SNP) arrays which are biased by the choice of which SNPs to include. Alternatively they have analyzed sequence data that have been generated as part of medical genetic studies from populations with large census sizes, and thus do not capture the full scope of human genetic variation. Here we report high quality genome sequences (~40x average) from 280 individuals from 135 worldwide populations, including 45 Africans, 26 Native Americans, 27 Central Asians or Siberians, 46 East Asians, 25 Oceanians, 46 South Asians, and 71 West Eurasians. All samples were sequenced using an identical protocol at the same facility (Illumina Ltd.). We modified standard pipelines to eliminate biases that might confound population genetic studies. We report novel inferences, as well as a high resolution map that shows where archaic ancestry (Neanderthal and Denisovan) is distributed throughout the world. We compare and contrast the genomic landscape of the Denisovan introgression into mainland Eurasians to that in island Southeast Asians. We are making this dataset fully available on Amazon Web Services as a resource to the community, coincident with the American Society of Human Genetics meeting.
Improved haplotype phasing using identity by descent. B. L. Browning, S. R. Browning.
   We present a new haplotype phasing method that achieves higher accuracy than existing methods. The method is based on the Beagle haplotype frequency model, but unlike the original Beagle phasing method, the new method incorporates genetic recombination, genotype error, and segments of identity by descent.     We compared the new haplotype phasing method to Beagle (r1230) and to SHAPEIT version 2 (r778) using Illumina Human 1M SNP data for chromosome 20. We phased 44 HapMap3 CEU trio offspring together with subsets of Wellcome Trust Case Control Consortium 2 controls (n=650, 1300, 2600, 5200). Phase error was measured at trio offspring genotypes on chromosome 20 that have phase determined by parental genotypes. The SHAPEIT “states” parameter was set at 6400 in order to increase its phasing accuracy.     The new haplotype phasing method produced haplotype switch error rates that were 20-25% lower than the error rates for the existing Beagle method and 1-7% lower than the error rates for SHAPEIT. The difference in switch error rates between the new method and SHAPEIT increased with increasing sample size.     The new haplotype phasing method will be incorporated into version 4 of the Beagle software package (http://faculty.washington.edu/browning/beagle/beagle.html).
Reducing pervasive false positive identical-by-descent segments detected by large-scale pedigree analysis. E. Y. Durand, N. Eriksson, C. Y. McLean.
   Analysis of genomic segments shared identical-by-descent (IBD) between individuals is fundamental to many genetic applications, from demographic inference to estimating the heritability of diseases. A large number of methods to detect IBD segments have been developed recently. However, IBD detection accuracy in non-simulated data is largely unknown. In principle, it can be evaluated using known pedigrees, as IBD segments are by definition inherited without recombination down a family tree. We extracted 25,432 genotyped European individuals containing 2,952 father-mother-child trios from the 23andMe, Inc. dataset. We then used GERMLINE, a widely used IBD detection method, to detect IBD segments within this cohort. Exploiting known familial relationships, we identified a false positive rate over 67% for 2-4 centiMorgan (cM) segments, in sharp contrast with accuracies reported in simulated data at these sizes. We show that nearly all false positives arise due to allowing switch errors between haplotypes when detecting IBD, a necessity for retrieving long (> 6 cM) segments in the presence of imperfect phasing. We introduce HaploScore, a novel, computationally efficient metric that enables detection and filtering of false positive IBD segments on population-scale datasets. HaploScore scores IBD segments proportional to the number of switch errors they contain. Thus, it enables filtering of spurious segments reported due to GERMLINE being overly permissive to imperfect phasing. We replicate the false IBD findings and demonstrate the generalizability of HaploScore to alternative genotyping arrays using an independent cohort of 555 European individuals from the 1000 Genomes project. HaploScore can be readily adapted to improve the accuracy of segments reported by any IBD detection method, provided that estimates of the genotyping error rate and switch error rate are available.
Parente2: A fast and accurate method for detecting identity by descent. S. Bercovici, J. M. Rodriguez, L. Huang, S. Batzoglou.
   Identity-by-descent (IBD) inference is the problem of establishing a direct and explicit genetic connection between two individuals through a genomic segment that is inherited by both individuals from a recent common ancestor. IBD inference is key to a variety of population genomic studies, ranging from demographic studies to linking genomic variation with phenotype and disease. The problem of both accurate and efficient IBD detection has become increasingly challenging with the availability of large collections of human genotypes and genomes: given a cohort’s size, as quadratic number of pairwise genome comparisons must be performed, in principle. Therefore, computation time and the false discovery rate can also scale quadratically. To enable practical large-scale IBD detection, we developed Parente2, a novel method for detecting IBD segments. Parente2 is based on an embedded log-likelihood ratio and uses an ensemble windowing approach to model complex linkage disequilibrium in the underlying studied population. Parente2 is applied directly on genotype data without the need to phase data prior to IBD inference. Through extensive simulations using real data, we evaluate Parente2’s performance. We show that Parente2 is superior to previous state-of-the-art methods, detecting pairs of related individuals sharing a 4 cM IBD segment with 99.9%; sensitivity at a 0.1%; false positive rate, and achieving 79.2%; sensitivity at a 1%; false positive rate for the more challenging case of pairs sharing a 2 cM IBD segment. Additionally, Parente2 is efficient, providing one to two orders of magnitude speedup compared to previous state of the art methods. Parente2 is freely available at http://parente.stanford.edu/.
Fast PCA of very large samples in linear time. K. J. Galinsky, P. Loh, G. Bhatia, S. Georgiev, S. Mukherjee, N. J. Patterson, A. L. Price.
   Principal components analysis (PCA) is an effective tool for inferring population structure and correcting for population stratification in genetic data. Traditionally, PCA runs in O(MN2+N3 ) time, where M is the number of variants and N is the number of samples. Here, we describe a new algorithm, fastpca, for approximating the top K PCs that runs in time O(MNK), making use of recent advances in random low-rank matrix approximation algorithms (Rokhlin et al. 2009). fastpca avoids computing the GRM and associated computational and memory storage costs, enabling PCA of very large datasets on standard hardware. We estimated the top 10 PCs of the WTCCC dataset (16k samples, 101k variants) in roughly 7 minutes while consuming 1GB of RAM, compared to 1 hour and 2.5GB for PLINK2. The fastpca approximation was extremely accurate (r2>99% between all fastpca and PLINK2 PCs). The improvement in running time becomes even larger at larger samples sizes; for example, fastpca estimated the top 10 PCs of a simulated data set with 100k samples and 300k variants in 135 minutes 8.5GB of RAM, vs. an estimated 350 hours and 85GB of RAM using PLINK2. A recently published O(MN2) time method, flashpca, did not complete on this data set due to exceeding 40GB memory requirement. All of these analyses were based on LD-pruning SNPs with r2>0.2, which leads to much more accurate PCs in simulations as compared to retaining all SNPs; more complex LD-adjustment strategies provide only a small further improvement.
Fast detection of IBD segments associated with quantitative traits in genome-wide association studies. Z. Wang, E. Kang, B. Han, S. Snir, E. Eskin.
   Recently, many methods have been developed to detect the identity-by-descent (IBD) segments between a pair of individuals. These methods are able to detect very small shared IBD segments between a pair of individuals up to 2 centimorgans in length. This IBD information can be used to identify recent rare mutations associated with phenotype of interest. Previous approaches for IBD association were applicable to case/control phenotypes. In this work, we propose a novel and natural statistic for the IBD association testing, which can be applied to quantitative traits. A drawback of the statistic is that it requires a large number of permutations to assess the significance of the association, which can be a great computational challenge. We make a connection between the proposed statistic and linear models so that it does not require permutations to assess the significance of an association. In addition, our method can control population structure by utilizing linear mixed models.
Long-range haplotype mapping in Hispanic/Latinos reveals loci for short stature. G. Belbin, D. Ruderfer, K. Slivinski, M.C. Yee, J. Jeff, O. Gottesman, E.A. Stahl, R.J.F. Loos, E.P. Bottinger, E.E. Kenny.
   The Hispanic/Latino (HL) population of Northern Manhattan represents a diverse recent diaspora population, with 95% of the individuals reporting having grandparents born outside of the United States. Of these 43% report grandparents born in Puerto Rico, 23% the Dominican Republic, 13% Central America, and 5%, 4%, and 2% from Mexico, South America, and Europe respectively. Despite complex patterns of migration, admixture, and diversity, strong signatures of cryptic relatedness persist amongst HLs. We have detected long-range genomic tract sharing (>3cM), or identity-by-descent (IBD), across 5,194 HL in the Mount Sinai BioMe Biobank. We observed an average population level IBD sharing of 0.0025 in HL, which is 2.5- and 5-fold higher than that observed in BioMe European- and African-American populations, respectively. We hypothesize that these patterns of recent migration and genetic drift may drive some otherwise rare functional alleles to detectable frequency. We clustered groups of homologous IBD tracts (n=112,250) segregating in this HL population. We observed that IBD clusters represent a class of low frequency alleles (median minor allele frequency =0.0077, s.d.=0.0015). We performed a genome-wide association of the IBD clusters, or ‘population-based linkage’, to detect loci implicated in height, a highly heritable polygenic trait. 15 independent loci surpassed our empirically derived genome-wide significance threshold of less than 4.4710-4, 11 of which replicated in an independent cohort of BioMe HLs. Strikingly, two regions confer strong recessive effects. In the case of the top hit on 9q32 (MAF less than 0.005; p less than8x10-6), homozygous non-referent individuals were shorter by 6” or 10”, for men or women, respectively, compared to the population mean (5’ 7” and 5’ 2” for men and women, respectively). In addition, IBD haplotypes in the 9q32 cluster harbored a significant enrichment of Native American ancestry (p less than 1x10-16). Finally, this interval contains a number of biologically compelling candidate genes, including COL27A1 and PALM2. This study demonstrates that rich population structure, rather than being a confounding factor in biomedical discovery efforts, may be leveraged to reveal novel genetic associations with complex human traits.
A haplotype reference panel of over 31,000 individuals and next-generation imputation methods. S. Das, on behalf of Haplotype Reference Consortium.
   Genotype imputation is now a key tool in the analysis of human genetic studies, enabling array-based genetic association studies to examine the millions of variants that are being discovered by advances in whole genome sequencing. Examining these variants increases power and resolution of genetic association studies and makes it easier to compare the results of studies conducted using different arrays. Genotype imputation improves in accuracy with increasing numbers of sequenced samples, particularly for low frequency variants. The goal of the Haplotype Reference Consortium is to combine haplotype information from ongoing whole genome sequencing studies to create a large imputation resource. To date, we have collected information on >31,500 sequenced whole genomes, aggregated over 20 studies of predominantly European ancestry, to create a very large reference panel of human haplotypes where ~50M genetic variants are observed 5 or more times. These haplotypes can be used to guide genotype imputation and haplotype estimation. In preliminary empirical evaluations, our panel provides substantial increases in accuracy relative to the 1000 Genomes Project Phase 1 reference panel and other smaller panels, particularly for variants with frequency less than 
5%. I will describe our evaluation of strategies for merging haplotypes and variant lists across studies and advances in methods for genotype likelihood-based haplotype estimation that can be applied to 10,000s of samples. I will also summarize new methods for next generation imputation that perform faster and require less memory than contemporary methods while attaining similar levels of imputation accuracy. Our full resource is available to the community through imputation servers that enable scientists to impute missing variants in any study and respect the privacy of subjects contributing to the studies that constitute the Haplotype Reference Consortium. The majority of haplotypes will also be deposited in the European Genotype Archive.
A rare variant local haplotype sharing method with application to admixed populations. S. Hooker, G. T. Wang, B. Li, Y. Guan, S. M. Leal.
   With the advent of next generation sequencing there is great interest in studying the involvement of rare variants in complex trait etiology. For many complex traits sequence data is being generated on DNA samples from African Americans and Hispanics to elucidate rare variant associations. Analyses of admixed populations present special challenges due to spurious associations which can occur because of confounding. However using information on admixture and local ancestry can also be highly beneficial and increase the power to detect associations in these populations. Here a local haplotype sharing (LHS) method (Xu and Guan 2014) was extended to test for rare variant (RV) associations in admixed populations. Previously the Weighted Haplotype and Imputation-based Test (WHAIT) (Li et al. 2010) was proposed to test for rare variant associations using haplotype data. The RV-LHS method unlike WHAIT, does not require reconstruction of haplotypes which can be both computationally intensive and error prone. Additionally the RV-LHS uses information on local ancestry which is particularly advantageous when analyzing admixed populations. Results will be shown from simulation studies performed for rare variant data from an admixed population. Both Type I and II errors are evaluated for the RV-LHS method. Additionally the power of the RV-LHS method is compared to WHAIT as well as several other non-haplotype-based rare variant association methods including the combined multivariate collapsing (CMC) (Li and Leal, 2008), Variable Threshold (VT) (Price et al. 2010) and Sequence Kernel Association Test (SKAT) (Wu et al. 2010). Several heart, lung and blood phenotypes were analyzed using sequence data on African-Americans from the NHLBI-Exome Sequencing Project to better evaluate the performance of the RV-LHS compared to other rare variant association methods.

May 09, 2014

Ancient DNA from the Balkans (Iron Age Thrace)

A new paper in PLoS Genetics presents new data from two Iron Age Thracian individuals and puts the Sardinian-ness of Oetzi in new context. The authors write:
The results of the analyses including additional ancient genomes provide mounting evidence that the Iceman's genetic affinity with Sardinians reflects an ancestry component that was widespread in Europe during the Neolithic. Despite their different geographic origins, both the Swedish farmer gok4 and the Thracian P192-1 closely resemble the Iceman in their relationship with Sardinians, making it unlikely that all three individuals were recent migrants from Sardinia. Furthermore, P192-1 is an Iron Age individual from well after the arrival of the first farmers in Southeastern Europe (more than 2,000 years after the Iceman and gok4), perhaps indicating genetic continuity with the early farmers in this region. The only non-HG individual not following this pattern is K8 from Bulgaria. Interestingly, this individual was excavated from an aristocratic inhumation burial containing rich grave goods, indicating a high social standing, as opposed to the other individual, who was found in a pit [15]. However, the DNA damage pattern of this individual does not appear to be typical of ancient samples (Table S4 in [15]), indicating a potentially higher level of modern DNA contamination. On the other hand, the Swedish and the Iberian hunter-gatherers show congruent patterns of relatedness to the modern populations of Northern Europe, which is consistent with the previous results using those samples.
Also of interest, given previous suggestions that the Iceman had more Neandertal ancestry than modern Europeans:
However, all D-tests involving another non-African population do not significantly deviate from zero, suggesting that the Iceman genome contains levels of archaic ancestry that are comparable to that of other non-African populations.
A model of European history is seen on the left. Some details are probably incorrect (e.g., Sardinian Neolithic probably followed the Cardial/Mediterranean route rather the one shown in C). There are no good ancient DNA from Cardial Neolithic farmers, so the fact that Sardinians are similar to the Iron Age Bulgarian, the Stuttgart LBK German, and the Swedish TRB farmers may mean that the Mediterranean/Cardial farmers were related to the ones that went into Europe following the inland route from the Balkans.

In any case, the fact that there are now data from Bulgaria is great, because it means that southern Europe is not hopeless for ancient DNA preservation and hopefully more is on its way.

UPDATE: Did anyone see a link to the new data? It appears that there are only ~1,000 SNPs in common with the HGDP. [A link to the data will become available at the Bustamante lab website]

PLoS Genet 10(5): e1004353. doi:10.1371/journal.pgen.1004353

Population Genomic Analysis of Ancient and Modern Genomes Yields New Insights into the Genetic Ancestry of the Tyrolean Iceman and the Genetic Structure of Europe

Martin Sikora et al.

Genome sequencing of the 5,300-year-old mummy of the Tyrolean Iceman, found in 1991 on a glacier near the border of Italy and Austria, has yielded new insights into his origin and relationship to modern European populations. A key finding of that study was an apparent recent common ancestry with individuals from Sardinia, based largely on the Y chromosome haplogroup and common autosomal SNP variation. Here, we compiled and analyzed genomic datasets from both modern and ancient Europeans, including genome sequence data from over 400 Sardinians and two ancient Thracians from Bulgaria, to investigate this result in greater detail and determine its implications for the genetic structure of Neolithic Europe. Using whole-genome sequencing data, we confirm that the Iceman is, indeed, most closely related to Sardinians. Furthermore, we show that this relationship extends to other individuals from cultural contexts associated with the spread of agriculture during the Neolithic transition, in contrast to individuals from a hunter-gatherer context. We hypothesize that this genetic affinity of ancient samples from different parts of Europe with Sardinians represents a common genetic component that was geographically widespread across Europe during the Neolithic, likely related to migrations and population expansions associated with the spread of agriculture.

Link

February 13, 2014

Human admixture common in human history (Hellenthal et al. 2014)

A string of recent papers argued for admixture in human populations at time scales from the Middle Pleistocene to recent centuries. A new paper in Science makes the point convincingly for extensive admixture in humans over the last few thousand years. The authors include the creators of Chromopainter/fineStructure software; the new "Globetrotter" method appears to be a natural extension of that method that seemed to work wonderfully well except for the limitation of producing only a tree of the studied populations.

The paper has a companion website in which you can look up the admixture history of individual populations.

While reading this study, it is important to remember its limitations. Two are immediately obvious: (i) admixture events can only be detected for the last few thousand years, as this method depends on pattern of linkage disequilibrium which decays exponentially with time due to recombination, and (ii) detection of admixture seems to depend on the presence of maximally differentiated populations from the edges of the human geographical range; for example, the Japanese appear unadmixed even though they are clearly of dual Jomon/Yayoi ancestry. On the other hand, the method does detect the admixture present in the San at a similar time scale.

The case of Northwestern Europe appears especially striking as none of the populations from the region show evidence of admixture. This may be because the mixtures taking place there (e.g., between "Celts" and "Anglo-Saxons" in Great Britain) involved populations that were not strongly differentiated. Alternatively, population admixture history may have preceded the last few thousand years and is thus beyond the temporal scope of this method.

An exception to the rule that populations at the edges of the human range appear to be unadmixed are the Armenians who appear to be the only * between the Atlantic and Pacific in Figure 2D (shown at the beginning of this post). The companion site lists their status as "uncertain".

Other results are more questionable; for example, the authors assert that Sardinians are an admixed population with one side being "Egyptian-like" and the other "French-like" whereas the ancient DNA evidence as it stands would rather indicate that Sardinians are the best approximation of Neolithic Europeans currently in existence and so are more likely to (mostly) possess a gene pool that traces back to ~8-9 thousand years in Europe. It will be quite the surprise if so many Europeans from 5kya or earlier look like modern Sardinians and ancient Sardinians don't!

The analysis of Eastern Europe is particularly interesting as it documents three way admixture (Northern/Southern/NE Asian) in most populations but two way admixture (Northern/Southern) in Greeks, estimated at ~37%. The authors claim that this is related to the Slavs, which seems reasonable given the 1,054AD age estimate. On the other hand, according to the companion website, the southern element in Greeks is inferred to be Cypriot-like and it's far from clear that the pre-Slavic population of Greece was Cypriot-like or indeed represented by any of the populations in the authors' dataset.

The three-way admixture in much of eastern Europe is not particularly surprising as history furnishes ample evidence for groups of steppe origin in the region during historical times. Some bequeathed their both language and name (e.g., Magyars), others only their name (e.g., Bulgarians) on the local Europeans, but records indicate a widespread presence of "eastern" groups in Europe from the time of the Huns to that of the Ottomans. A study of late Antique eastern Europeans from the Baltic to the Aegean may help better document how the twin phenomena of the eastern invasions and the spread of the Slavs shaped the present-day genetic diversity of the region.

I suspect that a few ancient samples will be far more informative for understanding the recent history of our species than the most sophisticated modeling of modern populations. Nonetheless, it's great to have a new method that maximizes what can be learned about the past from the messy palimpsest of the present.

Science 14 February 2014: Vol. 343 no. 6172 pp. 747-751 DOI: 10.1126/science.1243518

A Genetic Atlas of Human Admixture History

Garrett Hellenthal et al.

Modern genetic data combined with appropriate statistical methods have the potential to contribute substantially to our understanding of human history. We have developed an approach that exploits the genomic structure of admixed populations to date and characterize historical mixture events at fine scales. We used this to produce an atlas of worldwide human admixture history, constructed by using genetic data alone and encompassing over 100 events occurring over the past 4000 years. We identified events whose dates and participants suggest they describe genetic impacts of the Mongol empire, Arab slave trade, Bantu expansion, first millennium CE migrations in Eastern Europe, and European colonialism, as well as unrecorded events, revealing admixture to be an almost universal force shaping human populations.

Link

October 26, 2013

New aDNA capture method (plus some data on ancient individuals from Bulgaria, Denmark, and Peru)

This seems to present an alternative method for capture of ancient DNA libraries than the one used on the Tianyuan individual. It is mostly a methods paper, but also has some initial analysis of some ancient individuals. From the paper:
We were able to tentatively call mtDNA haplogroups for these samples (Table S1). The two Bulgarian Iron Age individuals (P192-1 and T2G5) fell into haplogroups U3b and HV(16311), respectively. Haplogroup U3 is especially common in the countries surrounding the Black Sea, including Bulgaria, and in the Near East, and HV is also found at low frequencies in Europe and peaks in the Near East.41 The three Peruvian mummies fell into haplogroups B2, M (an ancestor of D), and D1, all derived from founder Native American lineages and previously observed in both pre-Columbian and modern populations from Peru. 
P192-1 was an Iron Age Thracian; T2G5 was from an Iron Age Thracian tumulus burial.

Also:
For the Peruvian mummies, we also included 10 Native American individuals from Central and South America in the PCA (Figures 3E and 3F). Interestingly, all of the mummies fell between the Native American populations (KAR, MAY, AYM) and East Asian populations (JPT, CHS, CHB), as would be expected for a nonadmixed Native American individual (Figures 3E, 3F, and S2). These mummies belonged to the pre-Columbian Chachapoya culture, who, by some accounts, were unusually fair-skinned,39 suggesting a potential for pre- Columbian European admixture. However, based on our preliminary results, these individuals appear to have been ancestrally Native American. 
The Peruvian mummies were from 1000-1500AD, so it's not very surprising that they don't appear to have European admixture and to be "ancestrally Native American".

Hopefully a more complete analysis of this data and production of more data with this method will follow in the future.

The American Journal of Human Genetics (2013), http://dx.doi.org/10.1016/j.ajhg.2013.10.002

Pulling out the 1%: Whole-Genome Capture for the Targeted Enrichment of Ancient DNA Sequencing Libraries

Meredith L. Carpenter et al.

Most ancient specimens contain very low levels of endogenous DNA, precluding the shotgun sequencing of many interesting samples because of cost. Ancient DNA (aDNA) libraries often contain less than 1% endogenous DNA, with the majority of sequencing capacity taken up by environmental DNA. Here we present a capture-based method for enriching the endogenous component of aDNA sequencing libraries. By using biotinylated RNA baits transcribed from genomic DNA libraries, we are able to capture DNA fragments from across the human genome. We demonstrate this method on libraries created from four Iron Age and Bronze Age human teeth from Bulgaria, as well as bone samples from seven Peruvian mummies and a Bronze Age hair sample from Denmark. Prior to capture, shotgun sequencing of these libraries yielded an average of 1.2% of reads mapping to the human genome (including duplicates). After capture, this fraction increased substantially, with up to 59% of reads mapped to human and enrichment ranging from 6- to 159-fold. Furthermore, we maintained coverage of the majority of regions sequenced in the precapture library. Intersection with the 1000 Genomes Project reference panel yielded an average of 50,723 SNPs (range 3,062–147,243) for the postcapture libraries sequenced with 1 million reads, compared with 13,280 SNPs (range 217–73,266) for the precapture libraries, increasing resolution in population genetic analyses. Our whole-genome capture approach makes it less costly to sequence aDNA from specimens containing very low levels of endogenous DNA, enabling the analysis of larger numbers of samples.

Link (pdf)

March 07, 2013

Y chromosomes of Bulgarians (Karachanak et al. 2013)

Bulgaria had been something of a blank area in studies of uniparental markers, so it's nice to finally see a comprehensive Y-chromosome study of the country.

The dates in the paper are based on the "evolutionary mutation rate". I suspect that ancient DNA will be the final arbiter in this issue, because, for example, a Mesolithic TMRCA of E-V13 in Bulgaria implies that we'll find a lot of it in Neolithic contexts, whereas a Bronze Age one implies that we'll find a little if any of it, and a discontinuity across time.

Of interest is the occurrence of some E*(xM35, M2) in this sample in Burgas, Varna, and Plovdiv. It would be interesting to trace the ancestry of the bearers of these Y-chromosomes. I know that there still exists a minority-within-a-minority of Black Muslims in Greek Thrace, and it's not inconceivable that these Y-chromosomes may represent the legacy of a similar population; in any case, their haplotypes can be found in Table S5 for anyone wanting to investigate.

SNP Diversity within R seems substantial, and as always, it is difficult to say much, since this may be a consequence of either (i) a plausible role of the Balkans as a staging point of the likely invasion of Europe in late prehistory, or (ii) back-migration of derived R-bearers into the Balkans, be them Slavs or Goths or "eastern" folks of various stripes during history. Once again, I suspect that ancient DNA might solve this riddle, or, alternatively, routine high-coverage sequencing of the Y chromosome that might inform us, e.g., about the TMRCA of a Bulgarian and a German R-U152 or a Bulgarian and Polish R-M458.

PLoS ONE 8(3): e56779. doi:10.1371/journal.pone.0056779

Y-Chromosome Diversity in Modern Bulgarians: New Clues about Their Ancestry

Sena Karachanak et al

To better define the structure and origin of the Bulgarian paternal gene pool, we have examined the Y-chromosome variation in 808 Bulgarian males. The analysis was performed by high-resolution genotyping of biallelic markers and by analyzing the STR variation within the most informative haplogroups. We found that the Y-chromosome gene pool in modern Bulgarians is primarily represented by Western Eurasian haplogroups with ~ 40% belonging to haplogroups E-V13 and I-M423, and 20% to R-M17. Haplogroups common in the Middle East (J and G) and in South Western Asia (R-L23*) occur at frequencies of 19% and 5%, respectively. Haplogroups C, N and Q, distinctive for Altaic and Central Asian Turkic-speaking populations, occur at the negligible frequency of only 1.5%. Principal Component analyses group Bulgarians with European populations, apart from Central Asian Turkic-speaking groups and South Western Asia populations. Within the country, the genetic variation is structured in Western, Central and Eastern Bulgaria indicating that the Balkan Mountains have been permeable to human movements. The lineage analysis provided the following interesting results: (i) R-L23* is present in Eastern Bulgaria since the post glacial period; (ii) haplogroup E-V13 has a Mesolithic age in Bulgaria from where it expanded after the arrival of farming; (iii) haplogroup J-M241 probably reflects the Neolithic westward expansion of farmers from the earliest sites along the Black Sea. On the whole, in light of the most recent historical studies, which indicate a substantial proto-Bulgarian input to the contemporary Bulgarian people, our data suggest that a common paternal ancestry between the proto-Bulgarians and the Altaic and Central Asian Turkic-speaking populations either did not exist or was negligible.

Link

October 07, 2012

rolloff analysis of Bulgarians as Sardinian+Pathan

Continuing my rolloff experiments, I have taken the Yunusbayev et al. sample of Bulgarians. This is interesting because of the recent evidence of a Sardinian-like individual from Iron Age Bulgaria, and also as a complement to a similar analysis on the Greeks. Bulgarians are Slavic speaking, but their ethnogenesis owes a great deal to the Bulgars, adding another potential element of complication. However, the paucity of East Eurasian admixture in Bulgarians, together with their Slavic language, probably suggests that this element represented a small elite that did not have a substantial role in the genetic formation of the Bulgarian population.

The top f3 statistics can be seen below:

Kshatriya_M Sardinian Bulgarians_Y -0.003813 0.000295 -12.918 237507
Velamas_M Sardinian Bulgarians_Y -0.003783 0.000285 -13.287 238276
Piramalai_Kallars_M Sardinian Bulgarians_Y -0.003693 0.000306 -12.061 238106
Kanjars_M Sardinian Bulgarians_Y -0.003643 0.000298 -12.227 237838
GIH30 Sardinian Bulgarians_Y -0.003638 0.000259 -14.028 240548
North_Kannadi Sardinian Bulgarians_Y -0.00355 0.000317 -11.187 237882
Muslim_M Sardinian Bulgarians_Y -0.003542 0.000333 -10.632 236964
Chamar_M Sardinian Bulgarians_Y -0.003505 0.000303 -11.585 238882
INS30 Sardinian Bulgarians_Y -0.003467 0.000264 -13.153 240279
Dharkars_M Sardinian Bulgarians_Y -0.003452 0.000309 -11.155 238211
Brahmins_from_Uttar_Pradesh_M Sardinian Bulgarians_Y -0.003448 0.000278 -12.42 238041
Indian_D Sardinian Bulgarians_Y -0.003411 0.000256 -13.308 241225
Iyer_D Sardinian Bulgarians_Y -0.003364 0.000291 -11.568 237509
Jatt_D Sardinian Bulgarians_Y -0.003327 0.000289 -11.513 236735
Pathan Sardinian Bulgarians_Y -0.003212 0.000239 -13.444 240969
Iyengar_D Sardinian Bulgarians_Y -0.003209 0.000308 -10.416 236840
Dusadh_M Sardinian Bulgarians_Y -0.003181 0.000313 -10.172 237512
Sindhi Sardinian Bulgarians_Y -0.003094 0.000239 -12.919 241268
Balochi Sardinian Bulgarians_Y -0.002804 0.00024 -11.686 240924


To maximize the number of SNPs and number of individuals, I used the Sardinian+Pathan pair as reference populations. 509,395 SNPs were used for this experiment. The exponential fit can be seen below:
There was a technical issue with the jackknife which I am currently investigating, but the mean time of the admixture was estimated at 126.83004 generations, or 3,680 years. This is similar to the value of 3,850 years I obtained on the Greek sample.

If this date is accepted, then the interesting issue is why an individual from Bulgaria was Sardinian-like during the Iron Age. Possibly, either this individual was Sardinian-like in the broad sense, despite having  minority West Asian admixture, or a few centuries after the admixture event, there was still an uneven distribution of the constituent elements, with most individuals still predominantly Sardinian-like. Given that the indigenous element was probably most numerous, so only part of it would have the opportunity to admix with the intrusive West Asian-like population, and this influence would spread to the population-at-large over time.

In any case, this evidence, such as it is, appears consistent with my idea about a Bronze Age invasion of Europe from Asia.

Naturally, only a broad sampling of ancient DNA variation from the Balkans, perhaps targeting different sites, cultures, times, social status, and physical types will be sufficient to track the early appearance of an intrusive population.

September 06, 2012

ASHG 2012 abstracts are online!

There is so much good stuff there. This year I decided against posting the full abstracts, so I'll just link to a few, adding a few sentences on why they strike me as interesting. And, since there are so many interesting ones, I'll keep updating this entry.

On the Sardinian ancestry of the Tyrolean Iceman confirms that modern Sardinians are most similar to both the Tyrolean Iceman and the Swedish Neolithic TRB individual (presumably Gok4). You can find my analysis of both in the archives of the blog. But, look here:
Strikingly, an analysis including novel ancient DNA data from an early Iron Age individual from Bulgaria also shows the strongest affinity of this individual with modern-day Sardinians. Our results show that the Tyrolean Iceman was not a recent migrant from Sardinia, but rather that among contemporary Europeans, Sardinians represent the population most closely related to populations present in the Southern Alpine region around 5000 years ago. The genetic affinity of ancient DNA samples from distant parts of Europe with Sardinians also suggests that this genetic signature was much more widespread across Europe during the Bronze Age.
As you may have guessed, I can't wait to get my hands on that Iron Age Thracian. His similarity with Sardinians is striking, because by the Iron Age, I would have thought that something akin to the modern genetic landscape would have begun to crystallize in Europe.

Y Chromosome J Haplogroups trace post glacial period expansion from Turkey and Caucasus into the Middle East confirms what I have argued about, i.e., that the West Asian highlands are responsible for the spread of haplogroup J, including, it seems into the Middle East itself. The chronology presented probably assumes the evolutionary mutation rate; also, the lack of haplogroup J in Europe pre-5ka argues for a late expansion. I am fairly convinced that out of this West Asian highlander population came the two dominant groups of West Eurasian prehistory, the Indo-Europeans and the Semites, their spread associated with a "metallurgical edge" in technology and social complexity during the Late Neolithic and Bronze Age. The latter probably picked their language from a T- or E-bearing population of the southern Levant (Ghassulians?), as these two haplogroups might link the Proto-Semites with their African Afroasiatic brethren.

Analytical inference of human demographic history using multiple individual genome sequences:
We estimate that Eurasian populations split from ancient Africans at 58,000-120,800 years ago, and the divergence time of Europeans and Asians occurred at 35,750-70,500 years ago.
This sounds reasonable, and the wide confidence intervals probably reflect current uncertainties about the mutation rate. The European/Asian split time intersects the UP and postdates the ~70ka turning point (Toba + Drying up of Arabia/Sahara). The African/European split time intersects the ~106ka Nubian complex in Arabia.

A genomewide map of Neandertal ancestry in modern humans:
We identify around 35,000 Neandertal-derived alleles in Europeans and 21,000 in East Asians.
This might seem superficially at odds with the recent finding of greater Neandertal ancestry in East Asians than Europeans, but remember that levels of Neandertal admixture depend on allele frequencies of introgressed variants, and East Asians are generally less polymorphic than Europeans.

Analysis of contributions of archaic genome and their functions in modern non-Africans
Totally, we identified 410,683 archaic segments in 909 non-African individuals with averaged segment length 83,460bp. In the genealogy of each archaic segment with Neanderthal, Denisovan, African and chimpanzee segments, 77~81% archaic segment coalesced first with Neanderthal, 4~8% coalesced first with Denisovan, and 14% coalesced first with neither, validating the algorithm. Interestingly, a large proportion of all the archaic segments identified shared 88.9% similarity with Neanderthal, suggesting a single major admixture with Neanderthal at 82~121kya, right after the Africa exodus of the ancestors of modern humans.
It will be interesting to see what these authors get a different date than Sankararaman et al.  The mutation rate can't be at fault because these dates are mostly dependent on the recombination rate. My initial guess is that the lower rate of S. et al. may be due to limiting the analysis to alleles with MAF less than 0.1. As I said before, it is unclear whether admixture LD-based signals of admixture with Neandertals can account for the totality of the D-statistics of Non-Africans vs. Africans.

Sequencing of an extended pedigree in Western chimpanzees is interesting for a variety of reasons, but for me the primary one is the inevitable use of this pedigree to fix the chimpanzee autosomal mutation rate, which has so far been assumed to be similar to the human one. On a similar topic, Estimating human mutation rate using autozygosity in a founder population comes up with 1.21x10-8/bp/generation for humans, which is practically the same as that inferred for Iceland, and belongs to the class of slow mutation rates that have been inferred lately and which may reshape our understanding of events ranging from human-chimp speciation to the date of Out-of-Africa.

The genetic structure of Western Balkan populations based on autosomal and haploid markers
Comparison of the variation within autosomal and haploid data sets of studied Western Balkan populations revealed their genetic closeness regardless of a genetic system inspected, in particular among the Slavic speakers. Hence, culturally diverse Western Balkan populations are genetically very similar to each other. Only the Kosovars show slight differences both in the variance of autosomal and uniparentally inherited markers from the other populations of the region, possibly also due to their historically strict patrilineality. In a more general perspective, our results reveal clear genetic continuity between the Near Eastern and European populations, lending further credence to extensive, likely multiple and possibly bidirectional ancient gene flows between the Near East and Europe, cutting through the Balkans.
 Asian Expansion of Modern Human out of Africa is not very eloquent, and I think may be missing a zero in one of its numbers, but the point being made (that the major Y-haplogroup E found in Africans is descended from Asian back-migrants) is something which I also think very likely, for reasons explained here.

Paleolithic human migrations in East Eurasia by sequencing Y chromosomes:
Paleolithic human migrations in East Eurasia remains largely unknown due to the lack of sufficient markers derived from the mutations that occurred during that time frame. To tackle this problem, using the sequence capturing, barcoding technology and next-generation sequencing, we identified more than 4,000 new SNPs encompassing most single copy non-recombining region of human Y chromosome. New clades for haplogroups O, C, N, D, and Q could be geographically located. Especially, a few star-like expansions were unveiled, showing strong population growth. The phylogeny of Haplogroup N was radically rearranged, and all the N individuals could now be categorized into either a northern clade N1 or southern clade N2, revealing a Paleolithic migratory routes of the ancestors of Uralic speaking populations. Haplogroup C, especially the East Eurasia-dominant clade C3, could also be separated into at least two ancient clades, suggesting Paleolithic migrations in East Asia. Three major clades under O, M117+, M134xM117, and 002611+, each could be now further classified into several subclades. With these new findings, we proposed the modified the routes and dates for human populations’ migration, especially those in Paleolithic time. A few Y-chromosomal expansions could now be linked to certain prehistoric cultures or ancestors of language families.
Inferring and sequencing the founding bottleneck of Ashkenazim
Applying this methodology to data from self-identified AJ samples, we show 85-90% of them belong to a genetic isolate related to other Mid-Eastern populations. This group has experienced an extreme bottleneck 30-35 generations ago, with subsequent expansion greatly exceeding the growth rate across all humans. Data are consistent with bottleneck size of merely 400 founders.

July 26, 2012

A look at Y chromosomes of Romania via Count Dracula

In short: researchers tried to see whether they could identify a specific Y chromosome lineage associated with the House of Basarab in Romania, the most famous member of which is Vlad the Impaler, an inspiration for the mythical Count Dracula. To do this, they tested Basarab-surnamed individuals, as well as the general Romanian population.

The whole exercise was, in a sense, a failure, since it neither disclosed a Basarab-specific lineage, nor resolved the historical question about the origin of the House of Basarab (Vlach or Cuman). But, it gave us some wonderful new data on Romania that is, of course, quite welcome.

This seems like a good candidate for a future ancient DNA study, assuming of course, that Vlad and his family are still in their final resting place, and there are brave enough researchers to disturb them (j/k).

On a more serious note, the authors correctly state that even if the Basarab house was originally Turkic, they could still have carried West Eurasian chromosomes, since incoming Turkic groups in Europe were not purely Mongoloid like their more remote ancestors. On the other hand, I note that most of the Basarab-surnamed individuals belonged to E-V13, I-P37.2, J-M241 all of which are almost certainly native Romanian. If one of them carries the original chromosome, then the odds are in favor of a Romanian origin, although nothing short of ancient DNA work can resolve the issue, assuming that's possible.

Table S1 contains the new Romanian data, and Table S2 data from surrounding populations (Hungary, Bulgaria, Ukraine).

PLoS ONE 7(7): e41803. doi:10.1371/journal.pone.0041803

Y-Chromosome Analysis in Individuals Bearing the Basarab Name of the First Dynasty of Wallachian Kings

Begoña Martinez-Cruz et al.

Vlad III The Impaler, also known as Dracula, descended from the dynasty of Basarab, the first rulers of independent Wallachia, in present Romania. Whether this dynasty is of Cuman (an admixed Turkic people that reached Wallachia from the East in the 11th century) or of local Romanian (Vlach) origin is debated among historians. Earlier studies have demonstrated the value of investigating the Y chromosome of men bearing a historical name, in order to identify their genetic origin. We sampled 29 Romanian men carrying the surname Basarab, in addition to four Romanian populations (from counties Dolj, N = 38; Mehedinti, N = 11; Cluj, N = 50; and Brasov, N = 50), and compared the data with the surrounding populations. We typed 131 SNPs and 19 STRs in the non-recombinant part of the Y-chromosome in all the individuals. We computed a PCA to situate the Basarab individuals in the context of Romania and its neighboring populations. Different Y-chromosome haplogroups were found within the individuals bearing the Basarab name. All haplogroups are common in Romania and other Central and Eastern European populations. In a PCA, the Basarab group clusters within other Romanian populations. We found several clusters of Basarab individuals having a common ancestor within the period of the last 600 years. The diversity of haplogroups found shows that not all individuals carrying the surname Basarab can be direct biological descendants of the Basarab dynasty. The absence of Eastern Asian lineages in the Basarab men can be interpreted as a lack of evidence for a Cuman origin of the Basarab dynasty, although it cannot be positively ruled out. It can be therefore concluded that the Basarab dynasty was successful in spreading its name beyond the spread of its genes.