Showing posts with label fastIBD. Show all posts
Showing posts with label fastIBD. Show all posts

May 08, 2013

The Geography of Recent Genetic Ancestry across Europe (Ralph and Coop 2013)

This paper first came out last July on the arXiv and went through four versions there before its final form which has now appeared in PLoS Biology. It's great that its early release allowed other people to read it without having to wait for the completion of the peer review process.

I think that this is a good model: journals have the right and obligation to subject papers to close scrutiny according to their own procedures, but this process ought not interfere with the early availability of research results or the ability of anyone other than the chosen reviewers to comment on new results.

PLoS Biol 11(5): e1001555. doi:10.1371/journal.pbio.1001555

The Geography of Recent Genetic Ancestry across Europe

Peter Ralph, Graham Coop

The recent genealogical history of human populations is a complex mosaic formed by individual migration, large-scale population movements, and other demographic events. Population genomics datasets can provide a window into this recent history, as rare traces of recent shared genetic ancestry are detectable due to long segments of shared genomic material. We make use of genomic data for 2,257 Europeans (in the Population Reference Sample [POPRES] dataset) to conduct one of the first surveys of recent genealogical ancestry over the past 3,000 years at a continental scale. We detected 1.9 million shared long genomic segments, and used the lengths of these to infer the distribution of shared ancestors across time and geography. We find that a pair of modern Europeans living in neighboring populations share around 2–12 genetic common ancestors from the last 1,500 years, and upwards of 100 genetic ancestors from the previous 1,000 years. These numbers drop off exponentially with geographic distance, but since these genetic ancestors are a tiny fraction of common genealogical ancestors, individuals from opposite ends of Europe are still expected to share millions of common genealogical ancestors over the last 1,000 years. There is also substantial regional variation in the number of shared genetic ancestors. For example, there are especially high numbers of common ancestors shared between many eastern populations that date roughly to the migration period (which includes the Slavic and Hunnic expansions into that region). Some of the lowest levels of common ancestry are seen in the Italian and Iberian peninsulas, which may indicate different effects of historical population expansions in these areas and/or more stably structured populations. Population genomic datasets have considerable power to uncover recent demographic history, and will allow a much fuller picture of the close genealogical kinship of individuals across the world.

Link

March 28, 2013

Refined IBD in Beagle 4

The Beagle page doesn't show version 4 yet, but I'm sure it will eventually turn up there since this paper has just been published.

Genetics doi: 10.1534/genetics.113.150029

Improving the Accuracy and Efficiency of Identity by Descent Detection in Population Data

Brian L. Browning and Sharon R. Browning

Segments of identity by descent (IBD) detected from high-density genetic data are useful for many applications, including long-range phase determination, phasing family data, imputation, IBD mapping and heritability analysis in founder populations. We present Refined IBD, a new method for IBD segment detection. Refined IBD achieves both computational efficiency and highly accurate IBD segment reporting by searching for IBD in two steps. The first step (identification) uses the GERMLINE algorithm to find shared haplotypes exceeding a length threshold. The second step (refinement), evaluates candidate segments with a probabilistic approach to assess the evidence for IBD. Like GERMLINE, Refined IBD allows for IBD reporting on a haplotype level, which facilitates determination of multi-individual IBD and allows for haplotype-based downstream analyses. To investigate the properties of Refined IBD, we simulate SNP data from a model with recent super-exponential population growth that is designed to match UK data. The simulation results show that Refined IBD achieves a better power/accuracy profile than fastIBD or GERMLINE. We find that a single run of Refined IBD achieves greater power than 10 runs of fastIBD. We also apply Refined IBD to SNP data for samples from the UK and from Northern Finland, and describe the IBD sharing in these data sets. Refined IBD is powerful, highly accurate, easy to use, and is implemented in Beagle version 4.

Link

December 13, 2012

Daniel MacArthur's chromosome 10

In a previous post I confirmed Razib's observation on the South Asian ancestry of part of Dan MacArthur's chromosome 10. Razib has a new post up in which he argues that this type of ancestry is not of Romani origin. I figured it was my turn to continue the series of gratis ancestry analysis, with two goals in mind: (i) to let people know that if you put your genome online, you may find interesting things about it, and (ii) to indeed figure out what is going on with this interesting case of unexpected admixture.

I used Beagle/fastIBD with default parameters and with the HapMap recombination map to figure out the mean IBD sharing between Dan's chr10 and a number of different populations, including most of my available South Asian references. I also included YRI as an appropriate outgroup, as well as CEU30 and the 1000 Genomes British populations, given that Dr. MacArthur is Australian and has a Scottish surname, so it's a good bet that he has plenty of British-type ancestry.

The mean chr10 sharing is plotted below:



Now, it's cool that the top-2 matches are Argyll and Orkney, both of which are part of Scotland. But, what is interesting, is that North_Kannadi squeezes in ahead of CEU, with a very respectable mean of 1.4cM, and a number of other Indian populations are not far behind, while most of the ones from Pakistan are not. I'd say this looks consistent with an "Indian" origin of this type of ancestry.

A useful control is to repeat this experiment with a different chromosome, the similar-sized chr9 which lacks evidence for South Asian ancestry:

It is now visually clear that the difference between the British populations and the South Asian ones is greatly diminished. And, while the North Kannadi were near the top of the order for chr10, they are near the bottom for chr9, even lower than the YRI outgroup at the "noise left end" of the spectrum. Even the highest ranked South Asian population has about 1/3 estimated IBD sharing as the British ones. 

Moreover, whereas in chr10 the top-ranked South Asian populations were often from the south, in chr9 the situation is reversed, with most of the top-ranked ones being from Pakistan. Again, this suggests that there is no real South Asian admixture here, but just some low-level sharing with the (more West Eurasian) populations of the northern part of South Asia.

So, to make a long story short, it does look to me like an excellent suggestion that there is some type of peninsular or even south Indian ancestry in the chr10 in question.

August 08, 2012

fastIBD analysis of several Jewish and non-Jewish groups

This is more of a "just the data" kind of post, inspired by the two recent papers on Jewish origins. A few quick points:
  • fastIBD was run with default parameters over a dataset of 512 individuals/264,539 SNPs
  • fastIBD identifies segments of relatively recent origin that are shared by individuals. These results should not be construed as measures of overall genetic similarity or origins. Rather, they suggest which populations have exchanged genes in the relative recent past, say, the last two thousand years or so.
  • I included all Ashkenazi_D and North_African_Jews_D samples; of the other Dodecad and reference populations, I took random samples of 10 each; running time of fastIBD increases with the square of the number of individuals, so doing this allowed me to run this in less than a day as opposed to about a week.
With that said, you can get:
  • Spreadsheet of numeric results, showing sharing (in centi-Morgans, cM)
  • Population-level graphical results, showing an ordering of other populations based on mean IBD sharing.
The following heat map allows for a quick appraisal of populations sharing an excess of IBD sharing (read row-by-row)



And, here are a couple of the visualizations for a few Jewish populations:

Note that all sources of data are listed on the bottom left of the Dodecad blog.

August 07, 2012

Origins of North African and Central/East European Jews

A couple of new papers appeared yesterday. I'll just post the abstracts for the time being, and add any further comments as an update when I find the time for them.

Some related analyses of mine:

UPDATE (Aug 8):

Below is Fig. 3 from Campbell et al.:

One can see that Jewish groups have high degree of intra-population IBD sharing (A); many of the highest levels of IBD sharing is between Jewish groups (B and C).

This paper definitely shows that Jewish groups differ from non-Jewish North Africans. But, the lack of comparative samples from non-Jewish non-North Africans makes the interpretation of this result difficult. Both the PCA analysis, shown below, and the structure analysis indicates a significant Sub-Saharan component in North African non-Jewish populations.


So, it seems, based on these results, that Jewish groups are differentiated from North Africans due to their general lack of sub-Saharan admixture, and they also show a variable degree of affiliation to European groups; however, by "European" groups we go only as far as north Italy and Sardinia. What of the relationships of different Jewish groups to people from southern Italy, Greece, Anatolia, the Caucasus, or even Iranian speakers of the Near East?

Now, let's go to the Elhaik paper, which investigates a different problem altogether, trying to distinguish between the "Rhineland" and "Khazarian" hypotheses for the origins of Central-East European Jews. According to the paper:

Admixture calculations were carried out using a supervised learning approach in a structure-like analysis. This approach has many advantages over the unsupervised approach that not only traces ancestry to K abstract unmixed populations under the assumption that they evolved  independently (Chakravarti 2009; Weiss and Long 2009) but also problematic when applied to study Jewish ancestry, which can be dated as far back as 3,000 years (Figure 2). Admixture was calculated with a reference set of seven populations representing genetically distinct regions: Pygmies (Africa), French Basque (West Europe), Chuvash (East Europe), Han Chinese (Asia), Palestinians (Middle East), Turk-Iranians (Near East), and Armenians (Caucasus) (Figure 5).
But, Palestinians too have African admixture, so using them as a parental population conflates two separate issues: their old Near Eastern Semitic ancestors which could be reasonably inferred to be somewhat related to the Semitic ancestors of Jews, and their recent African admixture. Similarly, Turks have east Eurasian admixture, and Iranians have South Asian admixture.


The IBD sharing is probably the strongest piece of evidence in this paper for a Caucasian connection. Excess of IBD sharing with Caucasus and Palestinians relative to the other populations may indeed be a good indication of such admixture. On the other hand, the Khazarian Empire was primarily located in eastern Europe and the North Caucasus, not in Armenia and Georgia. Also, this analysis rejects the Greco-Roman hypothesis (whereby European Jews underwent admixture in Greco-Roman times when they were part of the Hellenistic and Roman Empires), but does not really include any Greco-Roman populations (for example, from Greece and Italy).

On the other hand, there may be something to the Khazar story (but in the sense of admixture, rather than replacement). High IBD sharing with Caucasians is one such piece of evidence. Another is the presence of Y-haplogroup Q and R-Z93+, both of which could in principle track a Central Asian Turkic influence (although Z93 could also track an Iranian influence). Then, there is the limited but persistent evidence for a little East Eurasian admixture present in Ashkenazi Jews and not in Sephardic Jews, which might also be consistent with a little Turkic influence.

Overall, I am convinced that most modern Jewish groups have some variable old Near Eastern Jewish ancestry, primarily on the basis of the elevated "Southwest Asian" that seems to correlate reasonably well with groups of Semitic speakers. But, it is difficult to say "how much" and to identify all the potential sources of admixture. Jews have been an international people for quite a long time, so I would guess that fragments of different peoples they encountered may remain in their genomes. Perhaps something akin to Ralph and Coop (2012) may give more information about the timing of these admixture events, as well as the date of the common ancestry of different Jewish groups.

PS: I started a small fastIBD analysis of different Jewish and non-Jewish groups with a fairly large assortment of populations, and will probably post it here in the next few days.

PNAS doi: 10.1073/pnas.1204840109

North African Jewish and non-Jewish populations form distinctive, orthogonal clusters

Christopher L. Campbell et al.

North African Jews constitute the second largest Jewish Diaspora group. However, their relatedness to each other; to European, Middle Eastern, and other Jewish Diaspora groups; and to their former North African non-Jewish neighbors has not been well defined. Here, genome-wide analysis of five North African Jewish groups (Moroccan, Algerian, Tunisian, Djerban, and Libyan) and comparison with other Jewish and non-Jewish groups demonstrated distinctive North African Jewish population clusters with proximity to other Jewish populations and variable degrees of Middle Eastern, European, and North African admixture. Two major subgroups were identified by principal component, neighbor joining tree, and identity-by-descent analysis—Moroccan/Algerian and Djerban/Libyan—that varied in their degree of European admixture. These populations showed a high degree of endogamy and were part of a larger Ashkenazi and Sephardic Jewish group. By principal component analysis, these North African groups were orthogonal to contemporary populations from North and South Morocco, Western Sahara, Tunisia, Libya, and Egypt. Thus, this study is compatible with the history of North African Jews—founding during Classical Antiquity with proselytism of local populations, followed by genetic isolation with the rise of Christianity and then Islam, and admixture following the emigration of Sephardic Jews during the Inquisition.

Link

arXiv:1208.1092v1 [q-bio.PE]

The Missing Link of Jewish European Ancestry: Contrasting the Rhineland and the Khazarian Hypotheses

Eran Elhaik

The question of Jewish ancestry has been the subject of controversy for over two centuries and has yet to be resolved. The "Rhineland Hypothesis" proposes that Eastern European Jews emerged from a small group of German Jews who migrated eastward and expanded rapidly. Alternatively, the "Khazarian Hypothesis" suggests that Eastern European descended from Judean tribes who joined the Khazars, an amalgam of Turkic clans that settled the Caucasus in the early centuries CE and converted to Judaism in the 8th century. The Judaized Empire was continuously reinforced with Mesopotamian and Greco-Roman Jews until the 13th century. Following the collapse of their empire, the Judeo-Khazars fled to Eastern Europe. The rise of European Jewry is therefore explained by the contribution of the Judeo-Khazars. Thus far, however, their contribution has been estimated only empirically; the absence of genome-wide data from Caucasus populations precluded testing the Khazarian Hypothesis. Recent sequencing of modern Caucasus populations prompted us to revisit the Khazarian Hypothesis and compare it with the Rhineland Hypothesis. We applied a wide range of population genetic analyses - including principal component, biogeographical origin, admixture, identity by descent, allele sharing distance, and uniparental analyses - to compare these two hypotheses. Our findings support the Khazarian Hypothesis and portray the European Jewish genome as a mosaic of Caucasus, European, and Semitic ancestries, thereby consolidating previous contradictory reports of Jewish ancestry.

Link

July 18, 2012

fastIBD over 2,257 Europeans

Razib points me towards a very interesting new paper that applies fastIBD over the large POPRES dataset of Europeans. The most interesting thing about this is that the authors develop techniques for estimating the time depth of the pattern of common ancestry across Europe, and hence are able to conclude that the Slavic expansion has played a bigger role in European history than the Germanic one.

A worthwhile improvement would be to apply a clustering algorithm like I did back in January over the fastIBD output; that way, one does not have to arbitrarily partition Europe into regions, but have the partitions jump out of the data.

A different idea to confirm the scenario presented in this paper would be to drill into different European populations. For example, in the case of the Italians, it would be worthwhile to identify whether there are particular sub-populations with likely Greek or Albanian ancestry who share an excess of IBD with modern Greeks and Albanians.

Population averages may mask such interesting patterns lurking in the data. For example, sub-clusters within populations can be identified with both fineSTRUCTURE and fastIBD, and the corresponding clusters can be assessed with supervised ADMIXTURE to detect how they differ from each other. For example, using this technique, I was able to infer 3 sub-clusters within the ethnic Greek population:

  • pop8 (mainland Greek) with ~23% North_European
  • pop11 (Greek Cypriot) with ~5% North_European
  • pop14 (Cretan, islander, mainland+Asia Minor) with ~12% North_European
  • I have also a strong hunch based on a few half Pontic Greek+half mainland Greek data points that unmixed Pontic Greeks would be related to pop22 (Northeastern Anatolia) with ~5% North_European
Based on these results and the fastIBD analysis of Ralph and Coop (the POPRES Greek sample is from northern Greece), it might appear that a hefty portion of the North_European component in Greeks may date to the medieval period, since it is relatively smaller in eastern Greeks and Cypriots and also in the South Italian/Sicilian cluster pop16 of a different analysis, with Italians as a whole lacking the eastern European affiliations of some Greek groups.

Interestingly, ~5% North_European levels would be similar to those of Armenians who are the closest linguistic cousins of the Greeks within the Indo-European family, as well as the the Anatolian Turkish cluster pop13 at ~9%.

Overall, it would appear that some mainland Greek groups received some input as the result of the medieval Slavic intrusions, since the mainland North_European excess appears as a "wedge" within the South Italy/Sicily/Crete/Anatolia/Armenia arc and the fastIBD pattern of sharing suggests that this is due to fairly recent connections.

As I have pointed out before, one limitation of the method of counting shared blocks of ancestry is that it does not disclose the directionality of gene flow. For example, gene flow between Germans and Slavs is detected in this study, which could be ascribed to Germans living in eastern Europe and/or to Slavs becoming acculturated Germans as a result of living within Germanic states or intermarrying with them prior to the age of the nation state.

Finally -and most interestingly- I hope that similar haplotype-based methods can be applied to a wider dataset, because, as it is becoming clear, Europe has not been isolated from Asia or Africa during its long history. The authors mention "Slavic or Hunnic" as an explanation for the pattern of shared ancestry in eastern Europe, but it is only by including Asian groups that we can detect the existence of real Hunnic (or Avar, or Mongol, or Pecheneg, or, ...) ancestry.

Moreover, I am confident that the Bronze Age is well within the power of haplotype-based methods to detect IBD. For example, South Asian populations clearly show differential patterns of affiliation with modern West Eurasian groups, most of which can date to no later than the Bronze Age. Together with the gradual incorporation of the new ancient DNA genomes that are bound to be coming our way soon, it seems that our picture of not only recent history, but also of late prehistory is bound to become much sharper.

arXiv:1207.3815v1 [q-bio.PE]


The geography of recent genetic ancestry across Europe

Peter Ralph, Graham Coop
(Submitted on 16 Jul 2012)

The recent genealogical history of human populations is a complex mosaic formed by individual migration, large-scale population movements, and other demographic events. Population genomics datasets can provide a window into this recent history, as rare traces of recent shared genetic ancestry are detectable due to long segments of shared genomic material. We make use of genomic data for 2,257 Europeans (the POPRES dataset) to conduct one of the first surveys of recent genealogical ancestry over the past three thousand years at a continental scale. We detected 1.9 million shared genomic segments, and used the lengths of these to infer the distribution of shared ancestors across time and geography. We find that a pair of modern Europeans living in neighboring populations share around 10-50 genetic common ancestors from the last 1500 years, and upwards of 500 genetic ancestors from the previous 1000 years. These numbers drop off exponentially with geographic distance, but since genetic ancestry is rare, individuals from opposite ends of Europe are still expected to share millions of common genealogical ancestors over the last 1000 years. There is substantial regional variation in the number of shared genetic ancestors: especially high numbers of common ancestors between many eastern populations likely date to the Slavic and/or Hunnic expansions, while much lower levels of common ancestry in the Italian and Iberian peninsulas may indicate weaker demographic effects of Germanic expansions into these areas and/or more stably structured populations. Recent shared ancestry in modern Europeans is ubiquitous, and clearly shows the impact of both small-scale migration and large historical events. Population genomic datasets have considerable power to uncover recent demographic history, and will allow a much fuller picture of the close genealogical kinship of individuals across the world.

Link

March 26, 2012

Similarity matrices and clustering (Lawson and Falush)

Lawson and Falush have a new review paper on different clustering methods using haplotype data such as their own ChromoPainter/fineSTRUCTURE methodology, as well as the MCLUST/fastIBD methods that I started playing with a while back.

I won't have much time for the next few days to comprehensively review this new work, but I will add one data point to the discussion, by pointing to my ChromoPainter and fastIBD analyses over the same dataset. I will also add any further comments on this blog post, once I get the opportunity to read the paper.

Another point that needs to be made is how commendable the ChromoPainter folks' attitude towards the topic has been. Not only did they post their ChromoPainter preprint and software online months before their original paper was published, but they quickly jumped on my comments and suggestions on their paper to write their new review paper, making at available as a preprint itself. I'm guessing this saved about a year or two over what would have been possible if all the formalities of "traditional" publishing had been observed. It's also a very nice example of synergy between professional and amateur science, that the Internet and social media has made possible.

Similarity matrices and clustering algorithms for population identification using genetic data


Daniel John Lawson and Daniel Falush

Abstract

A large number of algorithms have been developed to identify population
structure from genetic data. Recent results show that the information used
by both model-based clustering methods and Principal Components Analysis
can be summarised by a matrix of pairwise similarity measures between
individuals. Similarity matrices have been constructed in a number of ways,
usually treating markers as independent but differing in the weighting given
to polymorphisms of different frequencies. Additionally, methods are now being
developed that better exploit the power of genome data by taking linkage
into account. We review several such matrices and evaluate their ‘information
content’. A two-stage approach for population identification is to first construct
a similarity matrix, and then perform clustering. We review a range
of common clustering algorithms, and evaluate their performance through a
simulation study. The clustering step can be performed either directly, or
after using a dimension reduction technique such as Principal Components
Analysis, which we find substantially improves the performance of most algorithms.
Based on these results, we describe the population structure signal
contained in each similarity matrix, finding that accounting for linkage leads
to significant improvements for sequence data. We also perform a comparison
on real data, where we find that population genetics models outperform
generic clustering approaches, particularly in regards to robustness against
features such as relatedness between individuals.


Link

January 11, 2012

Clusters Galore (fastIBD edition)

(You can scroll down to the Results section, if you are not interested in the technical stuff)

When I proposed Clusters Galore  in November 2010, I was pleasantly surprised to see that very fine scale population structure could be uncovered using a combination of two algorithms:
  • A dimensionality reduction technique (such as PCA or MDS) applied to dense genotypic data
  • MCLUST, a state-of-the art model-based normal mixture clustering algorithm that had enough chops to uncover clusters of arbitrary size, shape, and orientation in multidimensional space
I explained how to carry out Clusters Galore analysis here. A most recent analysis of West Eurasians can be found here.

I have always thinking since that time of ways to improve the methodology. Since MCLUST is hard to best in my experience, I thought that improvement could be produced in the first step of the analysis; I have tried various ideas about choosing how many dimensions to retain, based on test of normality, Tracy-Widom statistics or some newer ideas. My conclusion has been that one could expect only delta improvements with any of these ideas.

After reading an abstract by Myers et al. in last years's ICHG, I realized that further improvement in resolution might be had by exploiting the linkage structure of dense genotype data, i.e., the pattern of co-inheritance of alleles along a stretch of chromosome.

Since, I didn't want to reinvent the wheel, I found the paintmychromosomes website, which I've also covered here, noting that it is unclear to what extent the ability of this methodology to infer fine-scale population structure is due to its exploitation of linkage.  Unfortunately, the processing pipeline for this technique is computationally daunting, and I would probably have to wait months to carry out any meaningful experiment on the types of datasets I'm used to working with.

An hour's worth of coding may save you a month of runtime, so I decided to look elsewhere.

I've been experimenting with fastIBD for a while now, so I invested some time to write some auxiliary code that would help me use it for my purposes. fastIBD finds identical-by-descent segments in a collection of individuals. It also has several attractive properties:

  • It is fast
  • It does its own phasing
  • It runs within BEAGLE, a very well-known genetic analysis software
In principle one could do a single fastIBD run over an entire dataset, but the memory footprint is prohibitive. So, rather than beg for money for a bigger computer, I took ten minutes to write some code that combines the results of 22 fastIBD runs (one per chromosome). I also wrote some code that calculates how much (in Morgans) IBD sharing exists in a pair of individuals.

fastIBD is tunable in various ways, but the default parameters seem to work fine for my purposes.

The end result of my labors is an NxN matrix (if there are N individuals) of IBD-based distance between individuals (in Morgans not IBD between individuals). This can then be fed into R's MDS routine, and then it's business as usual.

Results


I assembled a North European dataset for testing my ideas. A thing that always bugged me was the lack of ability to detect much population structure in the British Isles. So, I hoped that the added "punch" of using fastIBD would finally uncover this structure. All analyses were run with 256,932 SNPs.

Clusters Galore (fastIBD)


The 26 first MDS dimensions deviated from normality according to a Shapiro-Wilk test, and MCLUST found a total of 21 clusters using these dimensions.

Population N 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
Russian_D 22 2 20
Irish_D 22 19 2 1
Polish_D 23 23
German_D 21 1 18 2
Finnish_D 17 10 7
Swedish_D 13 1 12
English_D 12 12
British_D 13 1 12
Norwegian_D 11 11
Lithuanian_D 10 1 9
Dutch_D 9 5 3 1
British_Isles_D 8 8
Mixed_Scandinavian_D 4 4
Danish_D 3 3
Ukrainian_D 2 2
Latvian_D 1 1
Estonian_D 1 1
Russian 25 1 24
Orcadian 15 15
Lithuanians 10 1 9
Belorussian 9 8 1
Orkney_1KG 25 19 2 2 2
Kent_1KG 38 34 2 2
Cornwall_1KG 33 1 30 2
Argyll_1KG 4 1 3
FIN30 30 14 16
Ukranians_Y 20 20
Mordovians_Y 15 13 2

The clusters could be labeled as:
  1. Mordovian
  2. Slavic
  3. Irish
  4. English/British 
  5. German
  6. Mini-cluster of 2 related Germans?
  7. Scandinavian
  8. Finnish 1
  9. Finnish 2
  10. Lithuanian
  11. Orkney
  12. Vologda Russians (HGDP)
  13. Mini-cluster of 2 Vologda Russians
  14. Mini-cluster of 2 Vologda Russians
  15. Mini-cluster of 2 Vologda Russians
  16. Mini-cluster of 2 Kent English
  17. Mini-cluster of 2 Kent English
  18. Cornwall
  19. Mini-cluster of 2 Cornwall
  20. Argyll
  21. Mini-cluster of 2 Mordovians
So, it seems that my intuition was correct. There is a fairly clean division of Lithuanians and Slavs that was much more muddled whenever it came up before, a clean division of Mordvins and Russians, and a fairly comprehensive split of British Isles populations: a quite clean Irish cluster, a Cornwall cluster, an Argyll cluster, and a Kent/English cluster. Note that British_D and British_Isles_D populations consist mostly of English+some other British Isles, so I am not very surprised that they fall in the English main cluster.

The entire analysis (fastIBD + MCLUST) took a few hours to run.

Clusters Galore (PLINK/MDS)


For comparison, using the "classical" Clusters Galore with PLINK's MDS facility, there were 15 non-normal dimensions. A total of 16 clusters were inferred:

Population N 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
Russian_D 22 2 15 5
Irish_D 22 6 6 10
Polish_D 23 20 3
German_D 21 1 8 4 3 5
Finnish_D 17 17
Swedish_D 13 12 1
English_D 12 5 2 5
British_D 13 1 3 9
Norwegian_D 11 5 2 3 1
Lithuanian_D 10 10
Dutch_D 9 2 4 3
British_Isles_D 8 3 1 4
Mixed_Scandinavian_D 4 4
Danish_D 3 1 1 1
Ukrainian_D 2 1 1
Latvian_D 1 1
Estonian_D 1 1
Russian 25 8 1 16
Orcadian 15 6 9
Lithuanians 10 10
Belorussian 9 9
Orkney_1KG 25 2 14 5 2 2
Kent_1KG 38 15 8 11 2 2
Cornwall_1KG 33 17 1 13 2
Argyll_1KG 4 2 2
FIN30 30 1 29
Ukranians_Y 20 20
Mordovians_Y 15 11 1 3

Some of the distinctions lost: 2 Finnish clusters rolled into 1; Lithuanians and + Slavs rolled into 1; there are 3 British Isles clusters with substantial overlap between different populations as well as with Scandinavians; Mordovians and Russians overlap; no German cluster.

It seems pretty clear to me that Clusters Galore (fastIBD) is the way to go into the future for this type of analysis, and hopefully further refinements to the methodology and the addition of more project participants will add even more resolution.

Clustering relies on (i) the ability to detect "blobs" of individuals, and (ii) the existence of such "blobs" of individuals. Clusters Galore (fastIBD edition) seems to be pretty good at doing (i), but it's as good as the data it's fed. For example, currently the Dutch seem split between the English and the Germans, but I have little doubt that if their sample sizes were to grow, they would also form their own specific cluster.

If you are a Project participant from these groups, you can find the results of this run in this spreadsheet.