Concepedia

Publication | Open Access

Estimating Phred scores of Illumina base calls by logistic regression and sparse modeling

32

Citations

21

References

2017

Year

Abstract

The L <sub>1</sub>-regularized logistic regression improves the empirical discrimination power by as large as 14 and 25% respectively for two kinds of preprocessed sequencing signals, compared to the Illumina scoring method. Namely, the L <sub>1</sub> method identifies more base calls of high fidelity. Computationally, the L <sub>1</sub> method can handle large dataset and is efficient enough for daily sequencing. Meanwhile, the logistic model resulted from BIC is more interpretable. The modeling suggested that the most prominent quenching pattern in the current chemistry of Illumina occurred at the dinucleotide "GT". Besides, nucleotides were more likely to be miscalled as the previous bases if the preceding ones were not "G". It suggested that the phasing effect of bases after "G" was somewhat different from those after other nucleotide types.

References

YearCitations

Page 1