2018 · 148 citations · 4 references
EngineeringWord VectorCorpus LinguisticsText MiningWord EmbeddingsNatural Language ProcessingInformation RetrievalData ScienceData MiningPattern RecognitionComputational LinguisticsDocument ClassificationText ClassificationLanguage StudiesContent AnalysisAutomatic ClassificationKnowledge DiscoveryIntelligent ClassificationVector Space ModelKeyword ExtractionVector Representation
In recent years, with the rapid development of Internet Technology, text data is growing rapidly every day. Users need to filter out the information they need from a large amount of text. Therefore, automatic text classification technology can help users find information. In order to address problems, such as ignoring contextual semantic links and different vocabulary importance in traditional text classification techniques, this paper presents a vector representation of feature words based on the deep learning tool Word2vec, and the weight of the feature words is calculated by the improved TF-IDF algorithm. By multiplying the weight of the word and the word vector, the vector representation of the word is realized. Finally, each text is represented by accumulating all the word vectors. Thus, text classification is carried out.
4
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado et al. · arXiv (Cornell University) · 2013 · 11.7K citations · Full text