Publication | Closed Access
Geolocation Prediction in Social Media Data by Finding Location Indicative Words
168
Citations
37
References
2012
Year
Location InformationEngineeringGeographic Information RetrievalLocal Event DetectionLocation-aware Social MediumPrediction ConfidenceLocalizationCorpus LinguisticsText MiningNatural Language ProcessingComputational Social ScienceSocial MediaInformation RetrievalData ScienceLanguage StudiesGeographyGeosocial NetworkSocial Media DataLocation Indicative WordsGeospatial SemanticsGeolocation PredictionLinguistics
Geolocation prediction is vital to geospatial applications like localised search and local event detection. Predominately, social media geolocation models are based on full text data, including common words with no geospatial dimension (e.g. today) and noisy strings (tmrw), potentially hampering prediction and leading to slower/more memory-intensive models. In this paper, we focus on finding location indicative words (LIWs) via feature selection, and establishing whether the reduced feature set boosts geolocation accuracy. Our results show that an information gain ratiobased approach surpasses other methods at LIW selection, outperforming state-of-the-art geolocation prediction methods by 10.6% in accuracy and reducing the mean and median of prediction error distance by 45km and 209km, respectively, on a public dataset. We further formulate notions of prediction confidence, and demonstrate that performance is even higher in cases where our model is more confident, striking a trade-off between accuracy and coverage. Finally, the identified LIWs reveal regional language differences, which could be potentially useful for lexicographers.
| Year | Citations | |
|---|---|---|
Page 1
Page 1