2013 · 472 citations · 13 references
EngineeringSocial Medium MonitoringTwitter ContentCommunicationCorpus LinguisticsJournalismText MiningNatural Language ProcessingTopic CoherenceComputational Social ScienceSocial MediaInformation RetrievalData ScienceComputational LinguisticsContent AnalysisLda Topic ModelsSocial Medium MiningKnowledge DiscoveryMessy TextTweet PoolingAutomatic LabelingTopic ModelSocial Medium DataArts
Twitter, or the world of 140 characters poses serious challenges to the efficacy of topic models on short, messy text. While topic models such as Latent Dirichlet Allocation (LDA) have a long history of successful application to news articles and academic abstracts, they are often less coherent when applied to microblog content like Twitter. In this paper, we investigate methods to improve topics learned from Twitter content without modifying the basic machinery of LDA; we achieve this through various pooling schemes that aggregate tweets in a data preprocessing step for LDA. We empirically establish that a novel method of tweet pooling by hashtags leads to a vast improvement in a variety of measures for topic coherence across three diverse Twitter datasets in comparison to an unmodified LDA baseline and a variety of pooling schemes. An additional contribution of automatic hashtag labeling further improves on the hashtag pooling results for a subset of metrics. Overall, these two novel schemes lead to significantly improved LDA topic models on Twitter content.
13
Introduction to information retrieval
Choice Reviews Online · 2009 · 12.5K citations
Jianshu Weng, Ee‐Peng Lim, Jing Jiang et al. · 2010 · 1.7K citations · Full text
Computational Social Science, Social Media, Social Networks +15
Empirical study of topic modeling in Twitter
Liangjie Hong, Brian D. Davison · 2010 · 1.1K citations