2017 · 25 citations · 22 references
Topic models jointly learn topics and document-level topic distribution. Extrinsic evaluation of topic models tends to focus exclusively on topic-level evaluation, e.g. by assessing the coherence of topics. We demonstrate that there can be large discrepancies between topic-and documentlevel model quality, and that basing model evaluation on topic-level analysis can be highly misleading. We propose a method for automatically predicting topic model quality based on analysis of documentlevel topic allocations, and provide empirical evidence for its robustness.
22
Distributed Representations of Words and Phrases and their Compositionality
Tomáš Mikolov, Ilya Sutskever, Kai Chen et al. · arXiv (Cornell University) · 2013 · 18.1K citations · Full text
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado · arXiv (Cornell University) · 2013 · 18.1K citations · Full text
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado et al. · arXiv (Cornell University) · 2013 · 11.7K citations · Full text
The Stanford CoreNLP Natural Language Processing Toolkit
Christopher D. Manning, Mihai Surdeanu, John Bauer et al. · 2014 · 7.2K citations · Full text