Iterative Usage of Fixed and Random Effect Models for Powerful and Efficient Genome-Wide Association Studies
PLoS Genetics · 2016 · 1.6K citations · 38 references
GeneticsGenomicsGenomic PredictionGenome-wide Association StudiesGenome-wide Association StudyGenetic AnalysisGenotype-phenotype AssociationComputational GenomicsBiostatisticsPublic HealthRandom Effect ModelsFalse PositivesPersonal GenomicsMedicineStatistical GeneticsFixed Effect ModelBioinformaticsFunctional GenomicsEpidemiologyIterative UsageFixed EffectLinkage Analysis
GWAS false positives can be controlled by mixed linear models that account for population structure and kinship, but this adjustment can reduce true positives; a modified MLMM approach includes multiple markers as covariates to partially mitigate this confounding. The study aims to eliminate confounding in GWAS by iteratively applying a fixed‑effect model and a random‑effect model. The method iteratively applies a fixed‑effect model that tests one marker while controlling for associated markers, and a random‑effect model that estimates those markers to define kinship, unifying p‑values at each step; this iterative procedure is called FarmCPU. FarmCPU outperforms existing methods in statistical power and achieves linear‑time computation, enabling analysis of datasets with half a million individuals and markers in under three days.
False positives in a Genome-Wide Association Study (GWAS) can be effectively controlled by a fixed effect and random effect Mixed Linear Model (MLM) that incorporates population structure and kinship among individuals to adjust association tests on markers; however, the adjustment also compromises true positives. The modified MLM method, Multiple Loci Linear Mixed Model (MLMM), incorporates multiple markers simultaneously as covariates in a stepwise MLM to partially remove the confounding between testing markers and kinship. To completely eliminate the confounding, we divided MLMM into two parts: Fixed Effect Model (FEM) and a Random Effect Model (REM) and use them iteratively. FEM contains testing markers, one at a time, and multiple associated markers as covariates to control false positives. To avoid model over-fitting problem in FEM, the associated markers are estimated in REM by using them to define kinship. The P values of testing markers and the associated markers are unified at each iteration. We named the new method as Fixed and random model Circulating Probability Unification (FarmCPU). Both real and simulated data analyses demonstrated that FarmCPU improves statistical power compared to current methods. Additional benefits include an efficient computing time that is linear to both number of individuals and number of markers. Now, a dataset with half million individuals and half million markers can be analyzed within three days.
38
PLINK: A Tool Set for Whole-Genome Association and Population-Based Linkage Analyses
Shaun Purcell, Benjamin M. Neale, Katherine EO Todd-Brown et al. · The American Journal of Human Genetics · 2007
Genome-wide Association StudyWhole-genome AssociationLinkage Disequilibrium+12
34.9K citations
Principal components analysis corrects for stratification in genome-wide association studies
Alkes L. Price, Nick J. Patterson, Robert M. Plenge et al. · Nature Genetics · 2006
Genome-wide Association StudyHaplotype DeterminationGenotype-phenotype Association+8
10.5K citations
8K citations
LD Score regression distinguishes confounding from polygenicity in genome-wide association studies
Brendan Bulik‐Sullivan, Po‐Ru Loh, Hilary K. Finucane et al. · Nature Genetics · 2015
Genome-wide Association StudyGenetic DeterminantGenotype-phenotype Association+9
6K citations
Efficient Methods to Compute Genomic Predictions
P.M. VanRaden · Journal of Dairy Science · 2008
6K citations