arXiv (Cornell University) · 2018 · 18 citations · 33 references
EngineeringMachine LearningCorpus LinguisticsText MiningWord EmbeddingsNatural Language ProcessingSpeech RecognitionNeural Network ApproachesInformation RetrievalData ScienceComputational LinguisticsLanguage EngineeringLanguage StudiesNamed-entity RecognitionMachine TranslationEntity DisambiguationNlp TaskKnowledge DiscoveryTerminology ExtractionF1 ScoreRobust Lexical FeaturesLinguisticsPo Tagging
Neural network approaches to Named-Entity Recognition reduce the need for carefully hand-crafted features. While some features do remain in state-of-the-art systems, lexical features have been mostly discarded, with the exception of gazetteers. In this work, we show that this is unfair: lexical features are actually quite useful. We propose to embed words and entity types into a low-dimensional vector space we train from annotated data produced by distant supervision thanks to Wikipedia. From this, we compute - offline - a feature vector representing each word. When used with a vanilla recurrent neural network model, this representation yields substantial improvements. We establish a new state-of-the-art F1 score of 87.95 on ONTONOTES 5.0, while matching state-of-the-art performance with a F1 score of 91.73 on the over-studied CONLL-2003 dataset.
33
Glove: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, Christopher D. Manning · 2014 · 33.2K citations
ADADELTA: An Adaptive Learning Rate Method
Matthew D. Zeiler · arXiv (Cornell University) · 2012 · 5.5K citations · Full text
Natural Language Processing (almost) from Scratch
Ronan Collobert, Jason Weston, Léon Bottou et al. · arXiv (Cornell University) · 2011 · 5.2K citations · Full text