This looks like a must-read for anyone starting out in computational biology without extensive experience at the command line. The 135-page document linked at the bottom of the Google Group page looks like an excellent primer with lots of examples that could probably be completed in a day or two, and provides a great start for working in a Linux/Unix environment and programming with Perl. This started out as a graduate student course at UC Davis, and is now freely available for anyone who wants to learn Unix and Perl. Also, don't forget about the printable linux command line cheat sheet I posted here long ago.
Google Groups: Unix and Perl for Biologists
Wednesday, April 21, 2010
Friday, April 16, 2010
My thoughts on King's "Genetic Heterogeneity" essay in Cell
Update Thursday, April 29, 2010: See further commentary at a newer post here.
Just finished reading Jon McClellan and Mary-Claire King's Genetic Heterogeneity in Human Disease essay in Cell. It's definitely one of the most forthright and compelling essays I've read on the subject of the inadequacy of GWAS for identifying genes that cause complex human disease. The essay starts with an evolutionary perspective. Most human variation is relatively ancient - originating in ancient human populations long before the migration out of Africa. Yet new alleles arise constantly, and because of the relatively recent human population growth, we can be certain that most alleles are actually recent and rare. For a common allele to remain in the population it must withstand evolutionary pressure. If the variation is pathogenic, it must either (1) lead to disease later in life so as not to affect fitness (e.g. Alzheimer's Disease, AMD), or (2) it must be balanced by positive selection (e.g. hemoglobin genes which cause sickle cell anemia are balanced by positive selection from malaria resistance).
The authors then dive into heterogeneity, citing many examples of human diseases which display both locus heterogeneity (mutations in many different genes lead to the same disease), and allelic heterogeneity (many mutations in the same gene cause the same disease). The authors discuss early-onset breast and ovarian cancer, inherited hearing loss, genetics of lipid metabolism, and severe mental illnesses such as autism or schizophrenia.
Next comes a very nice discussion of the common-disease-common-variant (CDCV) hypothesis and GWAS. Thousands of "risk variants" have been identified from GWAS, yet most of these have no apparent biological function. Since most genotyping platforms select for common variants, and because evolution has ensured that most common variants are neutral, then it follows that most GWAS findings are neutral, stemming from factors other than a true association with disease risk.
For one, the authors cite a problem we're all well aware of: population stratification. Yet we tend to think that if we eliminate ethnic outliers or control for stratification with PCA or the like, then we've eliminated the problem. Yet the authors point to a recently Nature-published GWAS in autism that provides a striking example of the problem hypervariable alleles can cause. The authors found an association with a SNP which had a frequency in cases of 0.65, and a frequency in controls of 0.61. All cases and controls were of European descent. Yet the frequency of the risk variant varies from 0.21 to 0.77 across European populations! (N.b. - see the discussion of this point in a newer post). This difference in frequency across European populations is 14 times higher than the frequency difference between cases and controls! Even very minimal differences in ancestry between cases and controls could have explained this association rather than true association with autism.
The authors do give a few examples of where common variants truly affect a common disease (hemoglobin genes and sickle-cell anemia, autoimmune disorders and the MHC region, Alzheimer's disease and APOE, lactose intolerance and alleles in the lactase gene enhancer region). Yet these examples prove two points. (1) All the variants in these genes have a demonstrable effect on the protein or its expression, as opposed to most GWAS findings, and (2) back to the evolutionary perspective, all of these genes have reason to remain common because of their evasion of evolutionary pressure, because they either do not affect reproductive fitness, or are balanced by positive selection.
The authors conclude by offering potential paths going forward, utilizing high-throughput sequencing technologies. One of the problems with sequencing data is not just finding potentially deleterious mutations, but determining which of the many potentially deleterious mutations actually play a role in human disease. One of the most promising strategies is to use next-gen sequencing to trace coinheritance of potential disease causing alleles with disease within affected families - essentially linkage analysis. Finally the authors assert that replication in genetics studies should focus on the identification and confirmation of multiple biologically relevant mutations in the same gene. This would provide both biological and epidemiological support for the causality of the gene or pathway in the pathogenesis of the disease.
This essay is definitely worth a read.
Cell: Genetic Heterogeneity in Human Disease
Update Tuesday, April 27, 2010: Keep an eye out over at Genetic Future for an upcoming post pointing out some of the problems with this paper I didn't consider here.
Just finished reading Jon McClellan and Mary-Claire King's Genetic Heterogeneity in Human Disease essay in Cell. It's definitely one of the most forthright and compelling essays I've read on the subject of the inadequacy of GWAS for identifying genes that cause complex human disease. The essay starts with an evolutionary perspective. Most human variation is relatively ancient - originating in ancient human populations long before the migration out of Africa. Yet new alleles arise constantly, and because of the relatively recent human population growth, we can be certain that most alleles are actually recent and rare. For a common allele to remain in the population it must withstand evolutionary pressure. If the variation is pathogenic, it must either (1) lead to disease later in life so as not to affect fitness (e.g. Alzheimer's Disease, AMD), or (2) it must be balanced by positive selection (e.g. hemoglobin genes which cause sickle cell anemia are balanced by positive selection from malaria resistance).
The authors then dive into heterogeneity, citing many examples of human diseases which display both locus heterogeneity (mutations in many different genes lead to the same disease), and allelic heterogeneity (many mutations in the same gene cause the same disease). The authors discuss early-onset breast and ovarian cancer, inherited hearing loss, genetics of lipid metabolism, and severe mental illnesses such as autism or schizophrenia.
Next comes a very nice discussion of the common-disease-common-variant (CDCV) hypothesis and GWAS. Thousands of "risk variants" have been identified from GWAS, yet most of these have no apparent biological function. Since most genotyping platforms select for common variants, and because evolution has ensured that most common variants are neutral, then it follows that most GWAS findings are neutral, stemming from factors other than a true association with disease risk.
For one, the authors cite a problem we're all well aware of: population stratification. Yet we tend to think that if we eliminate ethnic outliers or control for stratification with PCA or the like, then we've eliminated the problem. Yet the authors point to a recently Nature-published GWAS in autism that provides a striking example of the problem hypervariable alleles can cause. The authors found an association with a SNP which had a frequency in cases of 0.65, and a frequency in controls of 0.61. All cases and controls were of European descent. Yet the frequency of the risk variant varies from 0.21 to 0.77 across European populations! (N.b. - see the discussion of this point in a newer post). This difference in frequency across European populations is 14 times higher than the frequency difference between cases and controls! Even very minimal differences in ancestry between cases and controls could have explained this association rather than true association with autism.
The authors do give a few examples of where common variants truly affect a common disease (hemoglobin genes and sickle-cell anemia, autoimmune disorders and the MHC region, Alzheimer's disease and APOE, lactose intolerance and alleles in the lactase gene enhancer region). Yet these examples prove two points. (1) All the variants in these genes have a demonstrable effect on the protein or its expression, as opposed to most GWAS findings, and (2) back to the evolutionary perspective, all of these genes have reason to remain common because of their evasion of evolutionary pressure, because they either do not affect reproductive fitness, or are balanced by positive selection.
The authors conclude by offering potential paths going forward, utilizing high-throughput sequencing technologies. One of the problems with sequencing data is not just finding potentially deleterious mutations, but determining which of the many potentially deleterious mutations actually play a role in human disease. One of the most promising strategies is to use next-gen sequencing to trace coinheritance of potential disease causing alleles with disease within affected families - essentially linkage analysis. Finally the authors assert that replication in genetics studies should focus on the identification and confirmation of multiple biologically relevant mutations in the same gene. This would provide both biological and epidemiological support for the causality of the gene or pathway in the pathogenesis of the disease.
This essay is definitely worth a read.
Cell: Genetic Heterogeneity in Human Disease
Update Tuesday, April 27, 2010: Keep an eye out over at Genetic Future for an upcoming post pointing out some of the problems with this paper I didn't consider here.
Tags:
GWAS,
Recommended Reading
Cell: Genetic Heterogeneity in Human Disease Review
Check out this review essay in Cell: Genetic Heterogeneity in Human Disease, by Jon McClellan and Mary-Claire King. (King's lab, incidentally, was the group who discovered via linkage analysis that the gene for early-onset breast and ovarian cancer on chromosome 17q21, nearly 5 years before Myriad Genetics filed for patent protection on the BRCA1/2 genes). Anyhow, looks like a great review on genetic heterogeneity and GWAS. Thanks @JVJAI.
Cell: Genetic Heterogeneity in Human Disease
UPDATE 4/16/2010: See my synopsis and thoughts on this essay here.
UPDATE 4/29/2010: See further thoughts on this essay here.
Cell: Genetic Heterogeneity in Human Disease
UPDATE 4/16/2010: See my synopsis and thoughts on this essay here.
UPDATE 4/29/2010: See further thoughts on this essay here.
Tags:
GWAS,
Recommended Reading
Wednesday, April 14, 2010
Cancer Biostatistics Workshop: Overfitting
This month's cancer biostatistics workshop on overfitting will be given by Fei Ye and Zhiguo (Alex) Zhao, both in the Department of Biostatistics and the Cancer Biostatistics Center. This looks like a good one, especially after attending Frank Harrell's regression modeling strategies course a few weeks ago. See the link below for the full 2010 series.
2010 Cancer Biostatistics Works Series
2010 Cancer Biostatistics Works Series
Tags:
Announcements,
Statistics
Journal Club 4/16/2010
Our Program in Computation Genomics Journal Club is starting again, now the 3rd Friday of each month. The next meeting is this Friday, April 16, at 3pm in the CHGR conference room. As usual, please bring in any articles you've found recently and give a brief overview of why you thought it was interesting. Also, take a look at these papers related to BioVU by investigators at Vanderbilt:
The BioVU demonstration project:
Ritchie MD, Denny JC, Crawford DC, Ramirez AH, Weiner JB, Pulley JM, Basford MA, Brown-Gentry K, Balser JR, Masys DR, Haines JL, Roden DM. Robust Replication of Genotype-Phenotype Associations across Multiple Diseases in an Electronic Medical Record. Am J Hum Genet. 2010 Mar 31.
The PheWAS:
Denny JC, Ritchie MD, Basford M, Pulley J, Bastarache L, Brown-Gentry K, Wang D, Masys DR, Roden DM, Crawford DC. PheWAS: Demonstrating the feasibility of a phenome-wide scan to discover gene-disease associations. Bioinformatics. 2010 Mar 24.
And finally, the PNAS paper from Brad Malin's group, which has generated quite a lot of press (Nature News, Technology Review, Genomics Law Report):
Loukides G, Gkoulalas-Divanis A, Malin B. Anonymization of electronic medical records for validating genome-wide association studies. PNAS.
Tuesday, April 13, 2010
Efficient Mixed-Model Association in GWAS using R
I recently did an analysis for the eMERGE network where I had lots of individuals from a small town in central Wisconsin where many of the subjects were related to one another. The subjects could not be treated as independent, but I could not use a family-based design either. I ended up using a mixed model approach using previously mentioned GenABEL. You can read about the method here (PubMed).
While researching which methods to use, I ran into what could be a potential problem. All of the methods that examine relatedness (including the method mentioned above), assume you have an ethnically homogeneous population. Yet all of the methods which look for population stratification (Eigenstrat, Structure, etc) assume samples are unrelated. So what do you do if you have both population stratification AND a high level of relatedness among your samples?
A few weeks ago our graduate student association invited and hosted Dr. Elaine Ostrander here from the NIH to talk about her work with gene mapping in dogs. She mentioned a method she used called Efficient Mixed-Model Association (EMMA) for performing association mapping while simultaneously correcting for relatedness and population structure. Using multiple highly inbred dog breeds represents the extreme case of simultaneously having to deal with substructure, inbreeding, and relatedness. If this method works for association mapping combining several purebred dog breeds, it should work for a less problematic human dataset as well.
EMMA is also implemented in R. You can download the necessary R package from the project's website below.
Efficient Mixed-Model Association (EMMA) website
PubMed: Efficient control of population structure in model organism association mapping.
Abstract: Genomewide association mapping in model organisms such as inbred mouse strains is a promising approach for the identification of risk factors related to human diseases. However, genetic association studies in inbred model organisms are confronted by the problem of complex population structure among strains. This induces inflated false positive rates, which cannot be corrected using standard approaches applied in human association studies such as genomic control or structured association. Recent studies demonstrated that mixed models successfully correct for the genetic relatedness in association mapping in maize and Arabidopsis panel data sets. However, the currently available mixed-model methods suffer from computational inefficiency. In this article, we propose a new method, efficient mixed-model association (EMMA), which corrects for population structure and genetic relatedness in model organism association mapping. Our method takes advantage of the specific nature of the optimization problem in applying mixed models for association mapping, which allows us to substantially increase the computational speed and reliability of the results. We applied EMMA to in silico whole-genome association mapping of inbred mouse strains involving hundreds of thousands of SNPs, in addition to Arabidopsis and maize data sets. We also performed extensive simulation studies to estimate the statistical power of EMMA under various SNP effects, varying degrees of population structure, and differing numbers of multiple measurements per strain. Despite the limited power of inbred mouse association mapping due to the limited number of available inbred strains, we are able to identify significantly associated SNPs, which fall into known QTL or genes identified through previous studies while avoiding an inflation of false positives. An R package implementation and webserver of our EMMA method are publicly available.
While researching which methods to use, I ran into what could be a potential problem. All of the methods that examine relatedness (including the method mentioned above), assume you have an ethnically homogeneous population. Yet all of the methods which look for population stratification (Eigenstrat, Structure, etc) assume samples are unrelated. So what do you do if you have both population stratification AND a high level of relatedness among your samples?
A few weeks ago our graduate student association invited and hosted Dr. Elaine Ostrander here from the NIH to talk about her work with gene mapping in dogs. She mentioned a method she used called Efficient Mixed-Model Association (EMMA) for performing association mapping while simultaneously correcting for relatedness and population structure. Using multiple highly inbred dog breeds represents the extreme case of simultaneously having to deal with substructure, inbreeding, and relatedness. If this method works for association mapping combining several purebred dog breeds, it should work for a less problematic human dataset as well.
EMMA is also implemented in R. You can download the necessary R package from the project's website below.
Efficient Mixed-Model Association (EMMA) website
PubMed: Efficient control of population structure in model organism association mapping.
Abstract: Genomewide association mapping in model organisms such as inbred mouse strains is a promising approach for the identification of risk factors related to human diseases. However, genetic association studies in inbred model organisms are confronted by the problem of complex population structure among strains. This induces inflated false positive rates, which cannot be corrected using standard approaches applied in human association studies such as genomic control or structured association. Recent studies demonstrated that mixed models successfully correct for the genetic relatedness in association mapping in maize and Arabidopsis panel data sets. However, the currently available mixed-model methods suffer from computational inefficiency. In this article, we propose a new method, efficient mixed-model association (EMMA), which corrects for population structure and genetic relatedness in model organism association mapping. Our method takes advantage of the specific nature of the optimization problem in applying mixed models for association mapping, which allows us to substantially increase the computational speed and reliability of the results. We applied EMMA to in silico whole-genome association mapping of inbred mouse strains involving hundreds of thousands of SNPs, in addition to Arabidopsis and maize data sets. We also performed extensive simulation studies to estimate the statistical power of EMMA under various SNP effects, varying degrees of population structure, and differing numbers of multiple measurements per strain. Despite the limited power of inbred mouse association mapping due to the limited number of available inbred strains, we are able to identify significantly associated SNPs, which fall into known QTL or genes identified through previous studies while avoiding an inflation of false positives. An R package implementation and webserver of our EMMA method are publicly available.
Tags:
GWAS,
R,
Software,
Statistics
Tuesday, April 6, 2010
ProbABEL - R package for GWAS data imputation
I've been using GenABEL for some time now for GWAS analysis using related individuals. It has an excellent set of functions for estimating a kinship matrix from a dense marker panel and then using this in a linear mixed effects model to allow for related individuals in the analysis of a quantitative trait. GenABEL also has many other nice features for analysis and visualization of GWAS data that you can't find in PLINK, it's free, cross-platform, and implemented in R. I'll write another post about GenABEL later, but here I wanted to note that GenABEL's creator, Yurii Aulchenko, released another package called ProbABEL for genome-wide association of imputed data. ProbABEL can perform imputation analyzing quantitative, binary, and survival outcomes while taking imputation uncertainty into account.
BMC Bioinformatics - ProbABEL package for genome-wide association analysis of imputed data
GenABEL homepage
GenABEL tutorial and reference manual
ProbABEL manual
BMC Bioinformatics - ProbABEL package for genome-wide association analysis of imputed data
GenABEL homepage
GenABEL tutorial and reference manual
ProbABEL manual
Tags:
GWAS,
R,
Software,
Statistics
Friday, April 2, 2010
Professor Bush
A short announcement - my friend, colleague, running partner, and GGD contributor Will Bush is now an assistant professor in the Department of Biomedical Informatics, and investigator in the Center for Human Genetics Research here at Vanderbilt.
Tags:
Announcements
Subscribe to:
Posts (Atom)