2019 · 25 citations · 29 references
EmojisEngineeringMachine LearningCross-lingual RepresentationMultilingual PretrainingMultimodal Sentiment AnalysisCorpus LinguisticsText MiningWord EmbeddingsNatural Language ProcessingSocial MediaData ScienceComputational LinguisticsLanguage StudiesContent AnalysisMachine TranslationNlp TaskEmoji PredictionDimension Reduction MethodsLinguistics
In this study, we aim to predict the most likely emoji given only a short text as an input. We extract a Hebrew political dataset of user comments for emoji prediction. Then, we investigate highly sparse n-grams representations as well as denser character n-grams representations for emoji classification. Since the comments in social media are usually short, we also investigate four dimension reduction methods, which associates similar words to similar vectorial representation. We demonstrate that the common Word Embedding dimension reduction method is not optimal. We also show that the character n-grams representations outperform all the other representation for the task of emoji prediction for Hebrew political domain.
29
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
Scikit-learn: Machine Learning in Python
Fabián Pedregosa, Gaël Varoquaux, Alexandre Gramfort et al. · arXiv (Cornell University) · 2012 · 63.3K citations · Full text
Indexing by latent semantic analysis
Scott Deerwester, Susan Dumais, George W. Furnas et al. · Journal of the American Society for Information Science · 1990 · 12.7K citations
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado et al. · arXiv (Cornell University) · 2013 · 11.7K citations · Full text
Thomas L. Griffiths, Mark Steyvers · Proceedings of the National Academy of Sciences · 2004 · 5.9K citations · Full text