2014 · 184 citations · 20 references
Topic models such as the latent Dirichlet allo-cation (LDA) have become a standard staple in the modeling toolbox of machine learning. They have been applied to a vast variety of data sets, contexts, and tasks to varying degrees of success. However, to date there is almost no formal theory explicating the LDA’s behavior, and despite its familiarity there is very little systematic analysis of and guidance on the properties of the data that affect the inferential performance of the model. This paper seeks to address this gap, by providing a systematic analysis of factors which character-ize the LDA’s performance. We present theorems elucidating the posterior contraction rates of the topics as the amount of data increases, and a thor-ough supporting empirical study using synthetic and real data sets, including news and web-based articles and tweet messages. Based on these re-sults we provide practical guidance on how to identify suitable data sets for topic models, and how to specify particular model parameters. 1.
20
Inference of Population Structure Using Multilocus Genotype Data
Jonathan K. Pritchard, Matthew Stephens, Peter Donnelly · Genetics · 2000 · 33.7K citations · Full text
Thomas L. Griffiths, Mark Steyvers · Proceedings of the National Academy of Sciences · 2004 · 5.9K citations · Full text
Collaborative topic modeling for recommending scientific articles
Chong Wang, David M. Blei · 2011 · 1.6K citations
Empirical study of topic modeling in Twitter
Liangjie Hong, Brian D. Davison · 2010 · 1.1K citations