Showing posts with label Lego. Show all posts
Showing posts with label Lego. Show all posts

April 03, 2011

Rare and common genetic variants in humans

The findings of this paper are not very surprising, but it is nice to see them formally expressed and tested.

Common variants are usually not functional; this makes sense as alleles that differ functionally from their competitors at a given locus are likely to win/lose in the evolutionary game and either become fixed, disappear, or be maintained at an extremely low frequency (and hence not be common).

The finding that major derived alleles are more functional also makes sense: a new allele at a given locus may either disappear, attain a non-trivial frequency by random drift, or even achieve a high frequency, pushing the ancestral allele to a lower one. In the latter case, the derived allele becomes the major (most frequent) allele in the population. A high frequency can be attained by either drift or selection: drift is slow and effective in small populations, whereas selection is faster, depending on the relative advantage of the new allele.

Polymorphism in the human genome can be maintained by either mutation-selection, whereby new variants appear constantly due to mutations in individuals, and selection culls many of them, or by balancing selection, whereby different variants have important functions but involve tradeoffs so that a tug-of-war between them results in an equilibrium.

What this paper suggests is that:
  1. Most common variants are not functional
  2. Common derived variants that are functional attain a high frequency and become the major alleles
  3. Rare variants are more likely to be functional than common ones
  4. Selection operates mostly by the constant culling of new alleles, sometimes by favoring new derived alleles, but, not so much by a tug-of-war between competing alleles
I'd say that this is quite consistent with my lego-block paradigm:
This Lego-block paradigm is based on the notion that most of our alleles are commodity"building blocks"; if they are brought together harmoneously, they produce positive results. The occasional allele may have a large effect, and some alleles fit better together than others. Yet, most of the success or failure of a construction depends on how the components fit together, and not what they are.
That, in my opinion, is where the "hidden heritability" mostly hides: part of it is due to the degradation of function by rare mutations that run in families or small populations, and part of it is due to the fortuitous combination of commodity alleles that have no long-term evolutionary advantage/disadvantage (function), but are co-inherited in the short-term from parents to offspring.

The authors have this to say:
Our analyses suggest that most of the functional variation carried by humans is likely to be rare genetic variation that is at least moderately deleterious and held to low frequency by selection.43,44 These analyses therefore provide a possible explanation for the relatively limited role of common genetic variation in most human diseases identified by genome-wide association studies.1,3
From the press release:
"The more common a variant is, the less likely it is to be found in a functional region of the genome," said senior author David Goldstein, Ph.D., director of the Duke Center for Human Genome Variation. "Scientists have reported this observation before, but this study is the most comprehensive effort to date using annotations of the functional regions of the human genome and fully sequenced genomes."

Goldstein said that "the magnitude of the effect is dramatic and is consistent across all frequencies of variants we looked at." He also said he was surprised by the notable consistency of the finding. "It's not just that the most rare variants are different from the most common, it's that at every increase in frequency, a variant is less and less likely to be found in a functional region of the DNA," Goldstein said. "This analysis is consistent with what appears to be a growing consensus that common variants are less important in common diseases than many had originally thought."



The American Journal of Human Genetics, 31 March 2011
doi:10.1016/j.ajhg.2011.03.008

A Genome-wide Comparison of the Functional Properties of Rare and Common Genetic Variants in Humans

Qianqian Zhu et al.

Abstract
One of the longest running debates in evolutionary biology concerns the kind of genetic variation that is primarily responsible for phenotypic variation in species. Here, we address this question for humans specifically from the perspective of population allele frequency of variants across the complete genome, including both coding and noncoding regions. We establish simple criteria to assess the likelihood that variants are functional based on their genomic locations and then use whole-genome sequence data from 29 subjects of European origin to assess the relationship between the functional properties of variants and their population allele frequencies. We find that for all criteria used to assess the likelihood that a variant is functional, the rarer variants are significantly more likely to be functional than the more common variants. Strikingly, these patterns disappear when we focus on only those variants in which the major alleles are derived. These analyses indicate that the majority of the genetic variation in terms of phenotypic consequence may result from a mutation-selection balance, as opposed to balancing selection, and have direct relevance to the study of human disease.

Link

November 18, 2010

Joint effects explain some hidden variance (Culverhouse et al. 2010)

It seems that my lego-block paradigm has found support after all. The authors found genetic effects for their trait of interest (nicotine dependence) for pairs of SNPs, where the SNPs themselves showed no individual effect. Thus: unremarkable "commodity" building blocks combined to produce a particular effect.

Note that this was done on only pairs of SNPs. But, there is no reason to think that there won't be triads, or tetrads, or n-nads of SNPs having such effect, and the important thing is: if n-1 SNPs have no effect, n might. To give an everyday analogy: press Ctrl: nothing happens; press Alt: nothing happens, press Del: nothing happens, or perhaps a character is deleted, but press them all together, and all of the sudden something big happens.

There is a catch, however: for independent SNPs, the number of individuals that possess a particular n-long combination decreases exponentially with n. In short, you'd need to sample the whole population of the Earth, and you'd still not be able to find any individuals having some particular effective n-long combination, let alone a large enough sample to establish a statistical dependence with the trait of interest.

To reiterate: genome-wide association studies treat humans like black boxes: flip a SNP and see if the person is nicotine dependent or not. Or, as in this study, flip two SNPs that looked like "dead switches" when you tried to flip them individually. That approach is a dead end for most complex traits of interest, and the way forward is to get into the box, and see what genes actually do.

HUMAN GENETICS DOI: 10.1007/s00439-010-0911-7

Uncovering hidden variance: pair-wise SNP analysis accounts for additional variance in nicotine dependence

Robert C. Culverhouse et al.

Abstract

Results from genome-wide association studies of complex traits account for only a modest proportion of the trait variance predicted to be due to genetics. We hypothesize that joint analysis of polymorphisms may account for more variance. We evaluated this hypothesis on a case–control smoking phenotype by examining pairs of nicotinic receptor single-nucleotide polymorphisms (SNPs) using the Restricted Partition Method (RPM) on data from the Collaborative Genetic Study of Nicotine Dependence (COGEND). We found evidence of joint effects that increase explained variance. Four signals identified in COGEND were testable in independent American Cancer Society (ACS) data, and three of the four signals replicated. Our results highlight two important lessons: joint effects that increase the explained variance are not limited to loci displaying substantial main effects, and joint effects need not display a significant interaction term in a logistic regression model. These results suggest that the joint analyses of variants may indeed account for part of the genetic variance left unexplained by single SNP analyses. Methodologies that limit analyses of joint effects to variants that demonstrate association in single SNP analyses, or require a significant interaction term, will likely miss important joint effects.

Link

October 04, 2009

DRD2 and Uralic admixture in Eastern European plain

From the paper:
The East European (Russian) Plain is a region in which peoples of the Indo-European and Uralic language families have come into contact over an extended period. Uralic-speaking peoples have the longest validated archaeological record in this region [17]. The most recent large-scale migration to this region involved the movement of Slavs (the Indo-European language family) to the east and northeast of their presumed homeland in Central Europe about 500 AD [18,19]. Slavs were not the first Indo-European-speaking people who arrived in the Russian Plain: in the firstmillennium BC, Baltic-speaking tribes occupied a large part of the East European Plain [17]. They were later displaced by Slavic tribes. According to the widely accepted hybridization theory of the origin of Eastern Slavs [20], Slavic populations arriving in the East European Plain were mixed with indigenous Uralic- and, probably, Baltic-speaking people.
...
Populations in the northwestern (Byelorussians 2 from Mjadel’), northern (Russians from Mezen’ and 6 from Oshevensk; Komi 3), and eastern parts (Russians 4 from Puchezh and Chuvash) of the East European Plain have relatively high frequencies of haplotype B2-D2-A2, which may reflect admixture with Uralic-speaking populations.Uralic genetic substratum in these regions, which were inhabited by Uralic-speaking tribes as late as the Early Middle Ages, was also shown by studies in which other genetic markers were used (mtDNA, Y-chromosome, and autosomal). Thus, the analysis of DRD2 haplotypes supports results on Slavic-Uralic admixture obtained using other markers, mainly neutral and sex-specific markers.
BMC Genet. 2009 Sep 30;10(1):62. [Epub ahead of print]

Haplotype frequencies at the DRD2 locus in populations of the East European Plain.

Flegontova OV, Khrunin AV, Lylova OI, Tarskaia LA, Spitsyn VA, Mikulich AI, Limborska SA.

ABSTRACT: BACKGROUND: It was demonstrated previously that the three-locus RFLP haplotype, TaqI B-TaqI D-TaqI A (B-D-A), at the DRD2 locus constitutes a powerful genetic marker and probably reflects the most ancient dispersal of anatomically modern humans. RESULTS: We investigated TaqI B, BclI, MboI, TaqI D, and TaqI A RFLPs in 17 contemporary populations of the East European Plain and Siberia. Most of these populations belong to the Indo-European or Uralic language families. We identified three common haplotypes, which occurred in more than 90% of chromosomes investigated. The frequencies of the haplotypes differed according to linguistic and geographical affiliation. CONCLUSIONS: Populations in the northwestern (Byelorussians from Mjadel'), northern (Russians from Mezen' and Oshevensk), and eastern (Russians from Puchezh) parts of the East European Plain had relatively high frequencies of haplotype B2-D2-A2, which may reflect admixture with Uralic-speaking populations that inhabited all of these regions in the Early Middle Ages.

September 19, 2008

Carl Zimmer article on Intelligence (and some thoughts on nature/nurture and IQ)

Carl Zimmer blogs about his Scientific American article on Intelligence. From the article:
It was with great delight that Plomin got his hands on microarrays that could detect 500,000 genetic markers--hundreds of times more than he had previously used. He and his colleagues got cheek swabs from 7,000 children, isolated their DNA, and ran it through the microarrays. And once more the results were disappointing.

“I’m not willing to say that we have found genes for intelligence,” Plomin declares, “because there have been so many false positives. They’re such small effects that you’re going to have to replicate them in many studies to feel very confident about them.”
I had blogged about this study when it came out. I repeat my comments from 2006 which are still valid today:
It appears that the hunt for genes affecting intelligence is not going well. I can't say that I'm surprised, because I have always maintained that intelligence is an emergent property of a set of co-operating genes during development in a particular environment and I don't anticipate that the geno-centric approach will take us closer to understanding it.

Intelligence, and -I believe- other complex traits are like complex dishes with many ingredients. The ingredients themselves (e.g., salt, lettuce, or chicken) are themselves unremarkable, but it is the way that they are put together and turned on and off by internal and external stimuli (the pot, the temperature, time, etc.) that makes a good dish.
I have expressed the same view in the recent entry on genome-wide association studies:
This Lego-block paradigm is based on the notion that most of our alleles are commodity "building blocks"; if they are brought together harmoneously, they produce positive results. The occasional allele may have a large effect, and some alleles fit better together than others. Yet, most of the success or failure of a construction depends on how the components fit together, and not what they are.
From the Carl Zimmer article:
Researchers have made images of their developing brains once a year, and Shaw has focused much of his attention on what the pictures reveal about the growth of the cortex, the outer rind of the brain where the most sophisticated information processing takes place.

...

In all children the cortex gets thicker as new neurons grow and produce new branches. Then the cortex thins out as branches are pruned. But in some parts of the cortex, Shaw found, development took a different course in children with different levels of intelligence. “The superclever kids started off very thin,” Shaw says. “They got really relatively thicker, but in adolescence they got thinner again very quickly.”

I had blogged about this study in 2006; check out that blog entry to see the thickness curves of cortex in development.

At the dawn of the genetics era, physical anthropologists' ideas that intelligence was correlated with the brain's observable properties were often ridiculed. And, yet neuronatomical correlates are pretty much the only game in town when it comes to giving a prediction (admittedly a very coarse one) of a person's IQ

That doesn't mean that genes don't play a role in intelligence; they do, and it's a sizeable one. But that role is hidden in a gene-gene and gene-environment interaction web of thousands of factors, where the individual components aren't really important, but the way they are put together are.

This realization also leads one to question genetic fetishists' conclusions about environmental influences on IQ.

It is true that scientists have looked at a lot of possible environmental influences on IQ and have come up short on significant environmental factors that can boost a person's IQ. There is simply very limited evidence that any particular environment can achieve this --sort of really bad influences such as malnutrition or some infectious diseases in childhood. And, yet we know that part of the variation of IQ is due to environmental influences. What gives?

What scientists have looked at are recognizable, "obvious", environmental influences (parenting style, schooling, etc.), which are analogous to the "common variants" in genetics.

Just as a microarray-based genome-wide association study has no clue about the rare family-level gene complexes and disease factors, so studies of environmental influences have no clue about the rare family/school/peer group micro-environments affecting a person's development.

Thus, the failure to find strong environmental influences on IQ doesn't strengthen the nature side of the nature-nurture divide, just as the failure to find strong genetic influences on IQ doesn't strengthen the nurture side.

The truth is, that Intelligence is an emergent property of a complex web of genetic and non-genetic interactions.

A human being is like a black box with zillions of inputs, some of them genetic, others environmental. We know that the box's output, e.g. its IQ score on a test is related to its inputs; but the relationship isn't linear and tidy: you can try different inputs from here to eternity, but you won't be able to figure out what the output is.

As I wrote in my post on height and body mass index, real progress will come about only when we finally look into the box:
Real progress will only come about with more developmental and functional studies, i.e. studies that actually look at what genes do in the body.

Figuring out how humans "work" is easier said than done. But, I believe, there is no shortcut.

September 16, 2008

Why genome-wide association studies don't really work (and how human evolution really happens)

In case you are skeptical about my gloomy assessment of the power of genome-wide association studies, here is Nicholas Wade in today's New York Times, profiling David Goldstein (via john hawks):
This idea, called the common disease/common variant hypothesis, drove major developments in biology over the last five years. Washington financed the HapMap, a catalog of common genetic variation in the human population. Companies like Affymetrix and Illumina developed powerful gene chips for scanning the human genome. Medical statisticians designed the genomewide association study, a robust methodology for discovering true disease genes and sidestepping the many false positives that have plagued the field.

But David B. Goldstein of Duke University, a leading young population geneticist known partly for his research into the genetic roots of Jewish ancestry, says the effort to nail down the genetics of most common diseases is not working. “There is absolutely no question,” he said, “that for the whole hope of personalized medicine, the news has been just about as bleak as it could be.”

...

The reason for this disappointing outcome, in his view, is that natural selection has been far more efficient than many researchers expected at screening out disease-causing variants. The common disease/common variant idea is largely wrong. What has happened is that a multitude of rare variants lie at the root of most common diseases, being rigorously pruned away as soon as any starts to become widespread.

I would only add that the common variant idea is probably wrong for neutral or positive traits as well. While negative traits are culled from the gene pool by purifying selection, advantageous traits (muscularity, beauty, intelligence, etc.) are positively selected.

But, what is selected? Surely, "a multitude of rare variants" can't exist for positive traits: most mutation is deleterious; the good stuff doesn't appear de novo very often. In my opinion, by and large, it is not common variants behind these traits. Rather, it is fortuitous combinations of unexceptional alleles.

This Lego-block paradigm is based on the notion that most of our alleles are commodity"building blocks"; if they are brought together harmoneously, they produce positive results. The occasional allele may have a large effect, and some alleles fit better together than others. Yet, most of the success or failure of a construction depends on how the components fit together, and not what they are.

I had previously made the point that evolution doesn't require mutation, selection, or drift but can be effected by the self-segregation of individuals into geographical or social niches for which they are better adapted. Differential reproduction of the semi-segregated geographical or social groups (i.e. group selection) then ensues.

There has doubtlessly been recent selection in humans, particularly because of feedback from the changed environments humans created for ourselves.

But, individuals don't differ from each other primarily because of genes that have undergone population-wide selection. Rather, we differ from each other first because we belong to a specific hierarchy of groups (race, subrace, ethnic group, etc.) and foremost because we have inherited a particular combination of alleles, our "family lego shape" from our parents.