mixOmics: An R package for ‘omics feature selection and multiple data integration
PLoS Computational Biology · 2017 · 3.6K citations · 37 references
EngineeringGenomicsMultiple Data IntegrationR PackageData ScienceOmics TechnologyBiostatisticsDimension ReductionMulti-omics StudySmall SubsetsOmicsPathway AnalysisMulti-omicsFunctional GenomicsBioinformaticsOmics DatasetsComputational BiologySystems BiologyMedicineOmics Integration
High‑throughput technologies generate abundant transcriptomic, proteomic, and metabolomic data, and integrating these large‑scale datasets can uncover biological insights, yet existing methods largely identify small, univariate molecular signatures from single omics types. mixOmics is an R package designed for multivariate exploration, dimensionality reduction, and visualization of biological data sets. It implements a systems‑biology approach that statistically integrates multiple heterogeneous omics data sets simultaneously to probe inter‑dataset relationships. The package extends PLS models for discriminant analysis, multi‑omics or multi‑study integration, and molecular signature discovery, and demonstrates these integrative frameworks on available omics data.
The advent of high throughput technologies has led to a wealth of publicly available 'omics data coming from different sources, such as transcriptomics, proteomics, metabolomics. Combining such large-scale biological data sets can lead to the discovery of important biological insights, provided that relevant information can be extracted in a holistic manner. Current statistical approaches have been focusing on identifying small subsets of molecules (a 'molecular signature') to explain or predict biological conditions, but mainly for a single type of 'omics. In addition, commonly used methods are univariate and consider each biological feature independently. We introduce mixOmics, an R package dedicated to the multivariate analysis of biological data sets with a specific focus on data exploration, dimension reduction and visualisation. By adopting a systems biology approach, the toolkit provides a wide range of methods that statistically integrate several data sets at once to probe relationships between heterogeneous 'omics data sets. Our recent methods extend Projection to Latent Structure (PLS) models for discriminant analysis, for data integration across multiple 'omics data or across independent studies, and for the identification of molecular signatures. We illustrate our latest mixOmics integrative frameworks for the multivariate analyses of 'omics data available from the package.
37
Regression Shrinkage and Selection Via the Lasso
Robert Tibshirani · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1996
50.3K citations
Regularization Paths for Generalized Linear Models via Coordinate Descent
Jerome H. Friedman, Trevor Hastie, Robert Tibshirani · Journal of Statistical Software · 2010
16.3K citations
Regularization Paths for Generalized Linear Models via Coordinate Descent.
Jerome H. Friedman, Trevor Hastie, Rob Tibshirani · PubMed · 2010
14K citations
Enterotypes of the human gut microbiome
Manimozhiyan Arumugam, Jeroen Raes, Éric Pelletier et al. · Nature · 2011
7.4K citations
Classification and diagnostic prediction of cancers using gene expression profiling and artificial neural networks
Javed Khan, Jun S. Wei, Markus Ringnér et al. · Nature Medicine · 2001
2.6K citations