Nucleic Acids Research · 2006 · 4.6K citations · 17 references
BiologyReference SequenceRefseq RecordsBiological DatabaseBioinformatics DatabaseGene Sequence AnnotationGeneticsComputational GenomicsSequence AnalysisNcbi Reference SequencesMicrobiologyGenomicsNcbi StaffMedicineBioinformatics
NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.
17
Basic local alignment search tool
Stephen F. Altschul, Warren Gish, Webb Miller et al. · Journal of Molecular Biology · 1990 · 92.8K citations
Basic Local Alignment Search Tool
Stephen F. Altschul · Journal of Molecular Biology · 1990 · 13.8K citations
dbSNP: the NCBI database of genetic variation
Stephen T. Sherry · Nucleic Acids Research · 2001 · 7.7K citations · Full text
Entrez Gene: gene-centered information at NCBI
Donna Maglott, James Ostell, Kim D. Pruitt et al. · Nucleic Acids Research · 2010 · 2.2K citations · Full text