Journal of Proteome Research · 2015 · 49 citations · 18 references
GeneticsMolecular BiologyBiological SignificanceBiostatisticsProtein-level Fdr AnalysisBiomarker DiscoveryPublic HealthProteomicsProtein-level FdrProtein FdrStatistical GeneticsStatistical ChallengesProtein ModelingOmicsFunctional GenomicsBioinformaticsProtein BioinformaticsComputational BiologySystems BiologyMedicine
In any high-throughput scientific study, it is often essential to estimate the percent of findings that are actually incorrect. This percentage is called the false discovery rate (abbreviated "FDR"), and it is an invariant (albeit, often unknown) quantity for any well-formed study. In proteomics, it has become common practice to incorrectly conflate the protein FDR (the percent of identified proteins that are actually absent) with protein-level target-decoy, a particular method for estimating the protein-level FDR. In this manner, the challenges of one approach have been used as the basis for an argument that the field should abstain from protein-level FDR analysis altogether or even the suggestion that the very notion of a protein FDR is flawed. As we demonstrate in simple but accurate simulations, not only is the protein-level FDR an invariant concept, when analyzing large data sets, the failure to properly acknowledge it or to correct for multiple testing can result in large, unrecognized errors, whereby thousands of absent proteins (and, potentially every protein in the FASTA database being considered) can be incorrectly identified.
18
A Statistical Model for Identifying Proteins by Tandem Mass Spectrometry
Alexey I. Nesvizhskii, Andrew Keller, Eugene Kolker et al. · Analytical Chemistry · 2003 · 4.9K citations
Semi-supervised learning for peptide identification from shotgun proteomics datasets
Lukas Käll, Jesse D. Canterbury, Jason Weston et al. · Nature Methods · 2007 · 2.5K citations