2017 · 41 citations · 16 references
We present an unsupervised contextsensitive spelling correction method for clinical free-text that uses word and character n-gram embeddings. Our method generates misspelling replacement candidates and ranks them according to their semantic fit, by calculating a weighted cosine similarity between the vectorized representation of a candidate and the misspelling context. We greatly outperform two baseline off-the-shelf spelling correction tools on a manually annotated MIMIC-III test set, and counter the frequency bias of an optimized noisy channel model, showing that neural embeddings can be successfully exploited to include context-awareness in a spelling correction model. Our source code, including a script to extract the annotated test data, can be found at https://github.com/ pieterfivez/bionlp2017.
16
Scikit-learn: Machine Learning in Python
Fabián Pedregosa, Gaël Varoquaux, Alexandre Gramfort et al. · arXiv (Cornell University) · 2012 · 63.3K citations · Full text
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado · arXiv (Cornell University) · 2013 · 18.1K citations · Full text
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado et al. · arXiv (Cornell University) · 2013 · 11.7K citations · Full text
MIMIC-III, a freely accessible critical care database
Alistair E. W. Johnson, Tom Pollard, Lu Shen et al. · Scientific Data · 2016 · 7.7K citations · Full text