IEEE/ACM Transactions on Computational Biology and Bioinformatics · 2021 · 18 citations · 44 references
EngineeringData ScienceClass ImbalanceRandom Forest ClassificationUnder-sampling StrategyBiostatisticsRandom Forest AlgorithmProteomicsProtein Interaction SitesInteractomicsKnowledge DiscoveryProtein ModelingOmicsProtein Structure PredictionFunctional GenomicsBioinformaticsProtein BioinformaticsTarget PredictionOmics DatasetsComputational BiologySystems BiologyMedicine
The computational methods of protein-protein interaction sites prediction can effectively avoid the shortcomings of high cost and time in traditional experimental approaches. However, the serious class imbalance between interface and non-interface residues on the protein sequences limits the prediction performance of these methods. This work therefore proposed a new strategy, NearMiss-based under-sampling for unbalancing datasets and Random Forest classification (NM-RF), to predict protein interaction sites. Herein, the residues on protein sequences were represented by the PSSM-derived features, hydropathy index (HI) and relative solvent accessibility (RSA). In order to resolve the class imbalance problem, an under-sampling method based on NearMiss algorithm is adopted to remove some non-interface residues, and then the random forest algorithm is used to perform binary classification on the balanced feature datasets. Experiments show that the accuracy of NM-RF model reaches 87.6% and 84.3% on Dtestset72 and PDBtestset164 respectively, which demonstrate the effectiveness of the proposed NM-RF method in differentiating the interface or non-interface residues.
44
Leo Breiman · Machine Learning · 2001 · 119.3K citations · Full text
Helen M. Berman · Nucleic Acids Research · 2000 · 38.9K citations
Biological Database, Biochemistry, Structural Bioinformatics +12
Addressing the Curse of Imbalanced Training Sets: One-Sided Selection.
Miroslav Kubát, Stan Matwin · 1997 · 2.2K citations