2009 · 11 citations · 10 references
Natural Language ProcessingEngineeringInformation RetrievalData ScienceData MiningVector Space ModelComputational LinguisticsKnowledge DiscoveryQuery ModelRelevance FeedbackText Document SearchKeyword SearchQuery VectorQuery ExpansionRelevance WeightingCorpus LinguisticsText Mining
The vector space model is one of the most common information retrieval (IR) methods for text document search. The cosine of the angle or the Euclidean distance between the query vector and each document vector is commonly used to measure similarity for query matching. Even though the vector space model starts with a term-by-document matrix, it inevitably loses the information of relations between query terms in the document in the first place. This paper presents a modified vector space model for measuring similarity between the query and the document when responding to a multi-term query. More weight is assigned to the keywords based on the adjacency between the terms in the documents. Thus, when a document contains the adjacency terms, its vector will typically move closer to the query vector to show stronger relevancy between query and the document.
10
Indexing by latent semantic analysis
Scott Deerwester, Susan Dumais, George W. Furnas et al. · Journal of the American Society for Information Science · 1990 · 12.7K citations
Introduction to information retrieval
Choice Reviews Online · 2009 · 12.5K citations