2013 · 198 citations · 29 references
While various claims have been made about text in social media text being noisy, there has never been a systematic study to investigate just how linguistically noisy or otherwise it is over a range of social media sources. We explore this question empirically over popular social media text types, in the form of YouTube comments, Twitter posts, web user forum posts, blog posts and Wikipedia, which we compare to a reference corpus of edited English text. We first extract out various descrip-tive statistics from each data type (includ-ing the distribution of languages, average sentence length and proportion of out-of-vocabulary words), and then investigate the proportion of grammatical sentences in each, based on a linguistically-motivated parser. We also investigate the relative similarity between different data types. 1
29
Named Entity Recognition in Tweets: An Experimental Study
Alan Ritter, Sam Clark, Oren Etzioni · 2011 · 1.2K citations
Improved Part-of-Speech Tagging for Online Conversational Text with Word Clusters
Olutobi Owoputi, Brendan O’Connor, Chris Dyer et al. · Figshare · 2013 · 679 citations · Full text
Using Social Media to Enhance Emergency Situation Awareness
Jie Yin, Andrew Lampert, Mark Cameron et al. · IEEE Intelligent Systems · 2012 · 645 citations
A Latent Variable Model for Geographic Lexical Variation
Jacob Eisenstein, Brendan O’Connor, Noah A. Smith et al. · Figshare · 2010 · 607 citations · Full text