Briefings in Bioinformatics · 2014 · 39 citations · 47 references
EngineeringGeneticsGenomicsSequence AlignmentString-searching AlgorithmData ScienceData MiningString ProcessingStemmingSuffix ArraySequence AnalysisKnowledge DiscoveryMorphologyOmicsComputer ScienceBioinformaticsFunctional GenomicsBiologyModified Suffix ArraysComputational BiologyCombinatorial Pattern MatchingSuffix Array ConstructionSystems BiologyMedicine
The suffix array and its variants are text-indexing data structures that have become indispensable in the field of bioinformatics. With the uninitiated in mind, we provide an accessible exposition of the SA-IS algorithm, which is the state of the art in suffix array construction. We also describe DisLex, a technique that allows standard suffix array construction algorithms to create modified suffix arrays designed to enable a simple form of inexact matching needed to support 'spaced seeds' and 'subset seeds' used in many biological applications.
47
Fast and accurate short read alignment with Burrows–Wheeler transform
Heng Li, Richard Durbin · Bioinformatics · 2009 · 60.7K citations · Full text
Fast gapped-read alignment with Bowtie 2
Ben Langmead, Steven L. Salzberg · Nature Methods · 2012 · 58.3K citations · Full text
Long-read Sequencing, Sequence Assembly, Natural Sciences +7
<tt>BLAT</tt>—The <tt>BLAST</tt>-Like Alignment Tool
W. James Kent · Genome Research · 2002 · 8.3K citations · Full text
dbSNP: the NCBI database of genetic variation
Stephen T. Sherry · Nucleic Acids Research · 2001 · 7.7K citations · Full text
SOAP: short oligonucleotide alignment program
Ruiqiang Li, Yingrui Li, Karsten Kristiansen et al. · Bioinformatics · 2008 · 3.6K citations · Full text