2016 · 12 citations · 4 references
EngineeringIntelligent Information RetrievalEntity SummarizationCorpus LinguisticsText MiningAutomatic SummarizationNatural Language ProcessingLanguage DocumentationInformation RetrievalData ScienceText SummarizationComputational LinguisticsTopical AnalysisData RetrievalLanguage StudiesContent AnalysisText AnalysisKnowledge DiscoveryText IndexingInformation ExtractionPos TaggingKeyword ExtractionTerm FrequencyText ProcessingLinguistics
Text summarization combines the process of POS tagging, term frequency and topical analysis. All these together are used to produce insightful summary of the document/documents. The concise version of the text document can be made using the concept of frequency of the terms and inverse frequency of documents. Text summarization is useful for bring the short story of all the newspaper articles, email correspondence or to extract key elements for the search engine. To compact the size, the sentences which are not near to the centroid is not to be considered in the output. To do that, the data which does not relate to the centroid topic has to be pruned. The output consists of only important data useful to the user. Large unstructured data can be converted in such form that can be used for report making, compacting of web pages and review of the book. In this, the summary from documents contains significant information, and is less than half of the original size. The output should be such that it fully satisfies the user's query and understands the answer given to it.
4
Indexing by latent semantic analysis
Scott Deerwester, Susan Dumais, George W. Furnas et al. · Journal of the American Society for Information Science · 1990 · 12.7K citations
A Corpus Factory for Many Languages
Adam Kilgarriff, Siva Reddy, Jan Pomikálek et al. · 2010 · 80 citations