Ecology and Evolution · 2021 · 27 citations · 59 references
DNA barcoding has become one of the most important techniques in plant species identification. Successful application of this technology is dependent on the availability of reference database of high species coverage. Unfortunately, there are experimental and data processing challenges to construct such a library within a short time. Here, we present our solutions to these challenges. We sequenced six conventional DNA barcode fragments (ITS1, ITS2, <i>matK</i>1, <i>matK</i>2, <i>rbcL</i>1, and <i>rbcL</i>2) of 380 flowering plants on next-generation sequencing (NGS) platforms (Illumina Hiseq 2500 and Ion Torrent S5) and the Sanger sequencing platform. After comparing the sequencing depths, read lengths, base qualities, and base accuracies, we conclude that Illumina Hiseq2500 PE250 run is suitable for conventional DNA barcoding. We developed a new "Cotu" method to create consensus sequences from NGS reads for longer output sequences and more reliable bases than the other three methods. Step-by-step instructions to our method are provided. By using high-throughput machines (PCR and NGS), labeling PCR, and the Cotu method, it is possible to significantly reduce the cost and labor investments for DNA barcoding. A regional or even global DNA barcoding reference library with high species coverage is likely to be constructed in a few years.
59
MEGA X: Molecular Evolutionary Genetics Analysis across Computing Platforms
Sudhir Kumar, Glen Stecher, Michael Li et al. · Molecular Biology and Evolution · 2018 · 37.2K citations · Full text
DADA2: High-resolution sample inference from Illumina amplicon data
Benjamin J. Callahan, Paul J. McMurdie, Michael Rosen et al. · Nature Methods · 2016 · 33.5K citations · Full text
Illumina Amplicon Data, Omics Datasets, Spectral Searching +5
FLASH: fast length adjustment of short reads to improve genome assemblies
Tanja Magoč, Steven L. Salzberg · Bioinformatics · 2011 · 15.1K citations · Full text