Showing posts with label Oroqen. Show all posts
Showing posts with label Oroqen. Show all posts

January 22, 2009

Best overall matching in Europeans

I was planning to do a post on the meaning of cluster separability, but this new paper actually demonstrates the main point I was going to investigate.

Over the last years, studies such as this have shown that is possible to distinguish between many European population groups.

However, clusters that are completely separable, may still harbor a substantial amount of internal variation. To see this, consider the following example of two groups, each consisting of 3 individuals and 5 markers:

G1:

a: ACGTA
b: AGACT
c: ACATT

G2:

d: CCGTA
e: CGACT
f: CCATT

It can be easily seen that group 1 is perfectly separable from group 2, in this case simply by looking at the first marker where group 1 has invariably an A, and group 2 has invariably a C.

But, if we try to see what the "best match" is for these individuals, we see that e.g., for individual a of group 1, the best match is d from group 2, for b it is e and for c it is f.

Now, for very distant populations, the scenario described above will almost never occur. Using 10K SNPs from the HGDP data, for the post I was preparing, I was able to conclude, for example, that any pair of Oroqen is always closer to each other than an Oroqen is to any Bantu sample.

This new study answers this question in the case of closely related European groups, showing that it is not the case that an individual will always have a member of his own group as his "best overall match" (BOM). Finns, for example, who appear as most distinct, have a Finnish BOM some 39 (out of 47) times, while some Finns have a Norwegian, German, or Polish BOM.

Moving into Central Europe, we see some counterintuitive results: no Austrian has an Austrian BOM, for example, but British, Danish, Dutch, German, Italian, and Polish ones.

Sample sizes play a role, however. For example, 25 out of 51 Greeks have Germans as their BOM, and only 7 out of 51 have Greek BOM's. But, since there is a sample of 983 Germans overall, Greeks are in fact 5.4 times more likely to match a Greek than a German.

Some markers and combinations of markers do differ between groups in a systematic way, like the A/C in the simple example above. Such markers allow us to separate groups, and distinguish between them. But, if we look at the overall genetic similarity between individuals, it turns out that members of one group may be more similar to some members of another than to their own.

So, if one were to be in a room with people from all over Europe, say during a meeting of the European Parliament, he might share some traits with people from his own country, but his best overall genetic match might be quite different.

Someone with the computing power and patience should carry out this investigation with the large HGDP dataset, to see which groups are strongly separable in the Oroqen-Bantu sense, and which ones are more weakly separable as in the European sense.

European Journal of Human Genetics doi: 10.1038/ejhg.2008.266

An evaluation of the genetic-matched pair study design using genome-wide SNP data from the European population

Timothy Tehva Lu et al.

Abstract

Genetic matching potentially provides a means to alleviate the effects of incomplete Mendelian randomization in population-based gene–disease association studies. We therefore evaluated the genetic-matched pair study design on the basis of genome-wide SNP data (309 790 markers; Affymetrix GeneChip Human Mapping 500K Array) from 2457 individuals, sampled at 23 different recruitment sites across Europe. Using pair-wise identity-by-state (IBS) as a matching criterion, we tried to derive a subset of markers that would allow identification of the best overall matching (BOM) partner for a given individual, based on the IBS status for the subset alone. However, our results suggest that, by following this approach, the prediction accuracy is only notably improved by the first 20 markers selected, and increases proportionally to the marker number thereafter. Furthermore, in a considerable proportion of cases (76.0%), the BOM of a given individual, based on the complete marker set, came from a different recruitment site than the individual itself. A second marker set, specifically selected for ancestry sensitivity using singular value decomposition, performed even more poorly and was no more capable of predicting the BOM than randomly chosen subsets. This leads us to conclude that, at least in Europe, the utility of the genetic-matched pair study design depends critically on the availability of comprehensive genotype information for both cases and controls.

Link

December 06, 2008

Genetic structure in East Asia using 200K SNPs

The table of paired Fst values for East Asian populations is here. The PCA plots are seen on the left.

PLoS ONE 3(12): e3862. doi:10.1371/journal.pone.0003862

Analysis of East Asia Genetic Substructure Using Genome-Wide SNP Arrays

Chao Tian et al.

Abstract

Accounting for population genetic substructure is important in reducing type 1 errors in genetic studies of complex disease. As efforts to understand complex genetic disease are expanded to different continental populations the understanding of genetic substructure within these continents will be useful in design and execution of association tests. In this study, population differentiation (Fst) and Principal Components Analyses (PCA) are examined using >200 K genotypes from multiple populations of East Asian ancestry. The population groups included those from the Human Genome Diversity Panel [Cambodian, Yi, Daur, Mongolian, Lahu, Dai, Hezhen, Miaozu, Naxi, Oroqen, She, Tu, Tujia, Naxi, Xibo, and Yakut], HapMap [ Han Chinese (CHB) and Japanese (JPT)], and East Asian or East Asian American subjects of Vietnamese, Korean, Filipino and Chinese ancestry. Paired Fst (Wei and Cockerham) showed close relationships between CHB and several large East Asian population groups (CHB/Korean, 0.0019; CHB/JPT, 00651; CHB/Vietnamese, 0.0065) with larger separation with Filipino (CHB/Filipino, 0.014). Low levels of differentiation were also observed between Dai and Vietnamese (0.0045) and between Vietnamese and Cambodian (0.0062). Similarly, small Fst's were observed among different presumed Han Chinese populations originating in different regions of mainland of China and Taiwan (Fst's less than 0.0025 with CHB). For PCA, the first two PC's showed a pattern of relationships that closely followed the geographic distribution of the different East Asian populations. PCA showed substructure both between different East Asian groups and within the Han Chinese population. These studies have also identified a subset of East Asian substructure ancestry informative markers (EASTASAIMS) that may be useful for future complex genetic disease association studies in reducing type 1 errors and in identifying homogeneous groups that may increase the power of such studies.

Link

September 27, 2008

More ASHG 2008 abstracts

The previous batch is here.

Analysis of East Asia Genetic Substructure: Population Differentiation and PCA Clusters Correlate with Geographic Distribution
Accounting for genetic substructure within European populations has been important in reducing type 1 errors in genetic studies of complex disease. As efforts to understand complex genetic disease are expanded to other continental populations an understanding of genetic substructure within these continents will be useful in design and execution of association tests. In this study, population differentiation(Fst) and Principal Components Analyses(PCA) are examined using >200K genotypes from multiple populations of East Asian ancestry(total 298 subjects). The population groups included those from the Human Genome Diversity Panel[Cambodian(CAMB), Yi, Daur, Mongolian(MGL), Lahu, Dai, Hezhen, Miaozu, Naxi, Oroqen, She, Tu, Tujia, Naxi, and Xibo], HapMap(CHB and JPT), and East Asian or East Asian American subjects of Vietnamese(VIET), Korean(KOR), Filipino(FIL) and Chinese ancestry. Paired Fst(Wei and Cockerham) showed close relationships between CHB and several large East Asian population groups(CHB/KOR, 0.0019; CHB/JPT, 00651; CHB/VIET, 0.0065) with larger separation with FIL(CHB/FIL, 0.014). Low levels of differentiation were also observed between DAI and VIET(0.0045) and between VIET and CAMB(0.0062). Similarly, small Fsts were observed among different presumed Han Chinese populations originating in different regions of mainland of China and Taiwan. For example, the four For PCA, the first two PCs showed a pattern of relationships that closely followed the geographic distribution of the different East Asian populations.corner groups were JPT, FIL, CAMB and MGL with the CHB forming the center group, and KOR was between CHB and JPT. Other small ethnic groups were also in rough geographic correlation with their putative origins. These studies have also enabled the selection of a subset of East Asian substructure ancestry informative markers(EASTASAIMS) that may be useful for future genetic association studies in reducing type 1 errors and in identifying homogeneous groups.

Worldwide Population Structure using SNP Microarray Genotyping
We genotyped 348 individuals sampled from 24 populations world-wide using the Affymetrix 250k NspI microarray chip. For context, we added matching genotypes from 210 HapMap individuals for a total of 250,823 loci genotyped in 543 individuals from 28 populations. We included populations from India and Daghestan to provide detail between the genetic poles of Western Europe, East Asia, and sub-Sahara Africa. With so many markers, principal components analyses reveal genetic differentiation between almost all identified populations in our sample. Northern and southern European populations (FST = 0.004, p <0.01) are statistically distinguishable, as are upper and lower caste groups in India (FST = 0.005, p <0.01). All individuals are accurately classified into continental groups, and even between closely-related populations, genetic- and self-classifications conflict for only a minority of individuals (e.g. ~2% between upper and lower Indian castes; k-means clustering.) As expected, the HapMap CHB+JPT, CEU, and YRI samples are most similar to our east Asian, west European, and African samples, respectively. The HapMap CEU samples and our northern European ancestry samples were both collected from Utah. Although individual samples cannot be reliably classified into their collection of origin, the groups are statistically distinguishable despite their high similarity (FST = 0.0005, n.s.). Our Japanese group is also statistically distinguishable from the HapMap JPT group (FST = 0.006, p <0.01), and in this comparison, most samples can be correctly classified. With such large numbers of genotypes, significant differences can be found even between very similar population samplings. Our results provide guidelines for researchers in selecting suitable control populations for case-control studies.


Frequency distribution and selection in 4 pigmentation genes in Europe
Pigmentation is one of the more obvious forms of variation in humans, particularly in Europeans where one sees more within group variation in hair and eye pigmentation than in the rest of the world. We studied 4 genes (SLC24A5, SLC45A2, OCA2 and MC1R) that are believed to contribute to the pigment phenotypes in Europeans. SLC24A5 has a single functional variant that leads to lighter skin pigmentation. Data on 83 populations worldwide (including 55 from our lab) show the variant (at rs1426654) has almost reached fixation in Europe, Southwest Asia, and North Africa, has moderate to high frequencies (.2-.9) throughout Central Asia, and has frequencies of .1-.3 in East and South Africa. The variant is essentially absent elsewhere. SLC45A2 also has a single functional variant (at rs16891982) associated with light skin pigmentation in Europe. Data on 84 populations worldwide show the light skin allele is nearly fixed in Northern Europe but has lower frequencies in Southern Europe, the Middle East and Northern Africa. In Central Asia the frequency of the SLC45A2 variant declines more quickly than the SLC24A5 variant. It is absent in both East and South Africa. In OCA2 we typed 4 SNPs (rs4778138, rs4778241, rs7495174, rs12913832) with a haplotype associated with blue eyes in Europeans. This haplotype shows a Southeastern to Northwestern pattern in Europe with frequencies of .25 (.05 homozygous) in the Adygei to .85 (.75 homozygous) in the Danes. In MC1R we typed 5 SNPs (rs3212345, rs3212357, rs3212363, C_25958294_10, rs7191944) that cover the entire MC1R gene and found a predominantly European haplotype that ranges in frequency from .35 to .65 in Europe, reaching its highest levels in Southwest Asia and Northwestern Europe. Extended Haplotype Heterozygosity (EHH) and normalized Haplosimilarity (nHS) show evidence of selection at SLC24A5 in not only our European and Southwest Asian populations but also our East African populations. Neither SLC45A2 or OCA2 showed evidence of selection in either test. MC1R did not show evidence of selection for our European specific haplotype but we did see some evidence both upstream and downstream in our nHS test in Europe.

Using principal components analysis to identify candidate genes for natural selection.
Genetic markers that differentiate populations are excellent candidates for natural selection due to local adaptation, and may shed light into physiological pathways that underlie disorders with varying frequencies around the world. Principal Components Analysis (PCA) has emerged as a powerful tool for the characterization and analysis of the structure of genomewide datasets. In prior work, we described an algorithm that can be used to select small subsets of genetic markers (SNPs) that correlate well with population structure, as captured by PCA. Our method can be used to detect SNPs that differentiate individuals from different geographic regions, or even neighboring subpopulations. We set out to explore the nature and properties of the genes where population-differentiating SNPs reside, by analyzing the publicly available Human Genome Diversity Panel dataset (650,000 SNPs for 1,043 individuals, 51 populations). Applying our SNP selection algorithms, we chose small subsets of SNPs that almost perfectly reproduce worldwide population structure as identified by PCA. We determined SNP panels both for population differentiation within seven geographic regions, as well as around the globe. We then explored the hypothesis that the selected SNPs attained their current worldwide allele frequency patterns as a response to the pressure of natural selection. Comparing our lists to recently published reports, we found a significant overlap with other genomewide scans for selection, thus validating our hypothesis. For example, EDAR (involved in the development of hair follicles) harbors the most differentiating SNPs in our world-wide panels. SNPs located in genes that are involved in skin and eye pigmentation (OCA2, MYO5C, HERC1, HERC2) are also among the top population differentiating markers. In East Asia, SNPs residing at the ADH cluster appear among the most important SNPs for population structure, while, in Europe, the same is true for genes that are involved in immune response to pathogens (CR1, DUOX2, TLR, and HLA). Finally, a comprehensive gene ontology analysis is presented.

August 09, 2006

August 1 update of YHRD

YHRD, the Y Chromosome Haplotype Reference Database has been updated on August 1:
The following populations were added today: Iceland, Elista (Russia, Kalmyks), Ecuador (Mestizo, Afroamerican, Quichua, Huaorani), Bama (China, Yao), Chengdu (China, Han), Zhenning (China, Buyi), Molidawa (China, Daur and Ewenki), Yuanjiang (China, Hani), Tongjiang (China, Hezhen), Tongxin (China, Hui), Yanji (China, Korean), Tongshi (China, Li), Xiuyan (China, Manchu), Alihe (China, Oroqen), Maowen (China, Qiang), Luoyuan (China, Fujian), Lhasa (China, Tibet), Yili (China, Xibe, Uigur and Han), Harbin (China, Han), Hailar (China, Mongolian), Lanzhou (China, Han), Liannan (China, Yao), Meixian (China, Han), Urumqi (China, Uigur), Mongolia, Japan, Korea, Gdansk (Poland), Nepal, Sao Paulo State (Brazil, European, African, Oriental and Pardo), Buenos Aires (Argentina), Santa Fe (Argentina), Mendoza (Argentina), Rio Negro (Argentina), Chubut (Argentina), Misiones (Argentina), Corrientes (Argentina), Formosa (Argentina), Chaco (Argentina), Salta (Argentina). We would like to thank the following colleagues for submitting these population samples: Daniel Corach and his group (Buenos Aires), Rune Andreassen and his group (Oslo, Norway), Ivan Nasidze and his group (Leipzig, Germany), Fabricio Gonzalez and his group (Quito, Ecuador), Chris Tyler-Smith, Yali Xue and their group (Cambridge, UK), Richard Pawlowski and his group (Gdansk, Poland), Rogerio Nogueira Oliveira and his group (Sao Paulo, Brazil), Gustavo Penacino and his group (Buenos Aires, Argentina), Emma Parkin, Mark Jobling and their group (Leicester, UK).

So, head on there to see if you get any new matches for your Y-chromosome samples.

October 21, 2005

The genetic legacy of the Manchu

A new paper in AJHG has identified a unique haplotype shared by many people from northeastern China and Mongolia. This haplotype belongs to haplogroup C3c, and is estimated to be about five centuries old. Its very recent spread corresponds with the rise to power of the Qing dynasty. As the authors write:
We reasoned that the events leading to the spread of this lineage might have been recorded in the historical record, as well as in the genetic record. The spread must have occurred after the cluster's TMRCA (∼500 years ago, corresponding to about A.D. 1500) and, most likely, before the Xibe migration in 1764. Notable features are the occurrence of the lineage in seven different populations but its apparent absence from the most populous Chinese ethnic group, the Han. A major historical event took place in this part of the world during this period—namely, the Manchu conquest of China and the establishment of the Qing dynasty, which ruled China from 1644 to 1912. This dynasty was founded by Nurhaci (1559–1626) and was dominated by the Qing imperial nobility, a hereditary class consisting of male-line descendants of Nurhaci's paternal grandfather, Giocangga (died 1582), with >80,000 official members by the end of the dynasty (Elliott 2001). The nobility were highly privileged; for example, a ninth-rank noble annually received ∼11 kg of silver and 22,000 liters of rice and maintained many concubines. A central part of the Qing social system was the army, the Eight Banners, which was made up of separate Manchu, Mongolian, and Chinese (Han) Eight Banners. The nobility occupied high ranks in the Manchu Eight Banners but not in the Mongolian or Chinese Eight Banners; the Manchu Eight Banners were recruited from the Manchu, Mongolian, Daur, Oroqen, Ewenki, Xibe, and a few other populations. A social mechanism was thus established that would have led to the increase of the specific Y lineage carried by Giocangga and Nurhaci and to its spread into a limited number of populations. We suggest that this lineage was the Manchu lineage.
Due to the very recent spread of this "Manchu" haplotype, it may be possible to find descendants of the Qing imperial nobility and test them. Presumably, they should have a high frequency of this haplotype. As the authors point out, the tumultuous events of recent times, have obscured the geneaological records. It may still be worthwhile to track such descendants though, or even remains of Qing noblemen. This would be a test for the validity of this theory.

An alternative explanation may be that the haplotype has been positively selected. There is, however, no direct evidence for this, and the limited geographical distribution of the haplotype may argue against this explanation:

Image Hosted by ImageShack.us

American Journal of Human Genetics (early view)

Recent Spread of a Y-Chromosomal Lineage in Northern China and Mongolia

Yali Xue et al.

We have identified a Y-chromosomal lineage that is unusually frequent in northeastern China and Mongolia, in which a haplotype cluster defined by 15 Y short tandem repeats was carried by ∼3.3% of the males sampled from East Asia. The most recent common ancestor of this lineage lived 590 ± 340 years ago (mean ± SD), and it was detected in Mongolians and six Chinese minority populations. We suggest that the lineage was spread by Qing Dynasty (1644–1912) nobility, who were a privileged elite sharing patrilineal descent from Giocangga (died 1582), the grandfather of Manchu leader Nurhaci, and whose documented members formed ∼0.4% of the minority population by the end of the dynasty.

Link