Bioinformatics · 2014 · 38 citations · 45 references
We were able to qualitatively replicate the raw experimental pyrosequencing data after rigorously adjusting existing simulation software. This allowed us to simulate datasets of real-life complexity, which we used to assess the influence and performance of two widely used pre-processing methods along with 11 clustering algorithms. We show that the choice, order and mode of the pre-processing methods have a larger impact on the accuracy of the clustering pipeline than the clustering methods themselves. Without pre-processing, the difference between the performances of clustering methods is large. Depending on the clustering algorithm, the most optimal analysis pipeline resulted in significant underestimations of the expected number of clusters (minimum: 3.4%; maximum: 13.6%), allowing us to make quantitative estimations of the bacterial complexity of real microbiome samples.
45
Basic local alignment search tool
Stephen F. Altschul, Warren Gish, Webb Miller et al. · Journal of Molecular Biology · 1990 · 92.8K citations
The SILVA ribosomal RNA gene database project: improved data processing and web-based tools
Christian Quast, Elmar Pruesse, Pelin Yilmaz et al. · Nucleic Acids Research · 2012 · 32.2K citations · Full text
Ribosomal Rna, Genetics, Genomics +20
Search and clustering orders of magnitude faster than BLAST
R. C. Edgar · Bioinformatics · 2010 · 21.1K citations · Full text
UCHIME improves sensitivity and speed of chimera detection
R. C. Edgar, Brian J. Haas, José C. Clemente et al. · Bioinformatics · 2011 · 15.2K citations · Full text