EngineeringSimilarity MeasureSemantic WebTf-idf AlgorithmsCorpus LinguisticsText MiningNatural Language ProcessingInformation RetrievalData ScienceData MiningPattern RecognitionSimilarity ComputingDocument ClassificationDocument ClusteringDocuments SimilarityKnowledge DiscoveryComputer ScienceContent Similarity DetectionVector Space ModelSimilarity Search
The precision and efficiency of the similarity computing of documents is the foundation and key of other documents processing. In this paper, the DF and TF-IDF algorithms are improved. First, DF's time complexity is linear which suits mass documents processing, but it has the fault that exceptional useful features may be deleted, so we make up that by adding the count of the words at the important places. Second, we rectify the weight of feature by the result of feature selection phase. In this way, we improve the precision of documents similarity without adding much time and space complexity.
5
A Comparative Study on Feature Selection in Text Categorization
Yiming Yang, Jan Pedersen · 1997 · 4.8K citations
Feature selection for classification
Manoranjan Dash, Hua Liu · Intelligent Data Analysis · 1997 · 2.6K citations
An evaluation on feature selection for text clustering
Tao Liu, Shengping Liu, Zheng Chen et al. · 2003 · 246 citations