International Journal of Artificial Intelligence Tools · 2004 · 10 citations · 1 references
EngineeringMining MethodsCorpus LinguisticsLanguage ProcessingText MiningNatural Language ProcessingKnowledge Discovery In DatabasesInformation RetrievalData ScienceData MiningComputational LinguisticsData ResourcesDocument ClassificationKnowledge Discovery ProcessKnowledge DiscoveryText Mining AlgorithmsWeb Text MiningInformation ExtractionText Mining ResearchKeyword ExtractionStructure MiningTextual Data Mining
Few tools exist that address the challenges facing researchers in the Textual Data Mining (TDM) field. Some are too specific to their application, or are prototypes not suitable for general use. More general tools often are not capable of processing large volumes of data. We have created a Textual Data Mining Infrastructure (TMI) that incorporates both existing and new capabilities in a reusable framework conducive to developing new tools and components. TMI adheres to strict guidelines that allow it to run in a wide range of processing environments – as a result, it accommodates the volume of computing and diversity of research occurring in TDM. A unique capability of TMI is support for optimization. This facilitates text mining research by automating the search for optimal parameters in text mining algorithms. In this article we describe a number of applications that use the TMI. A brief tutorial is provided on the use of TMI. We present several novel results that have not been published elsewhere. We also discuss how the TMI utilizes existing machine-learning libraries, thereby enabling researchers to continue and extend their endeavors with minimal effort. Towards that end, TMI is available on the web at .
1
Indexing by latent semantic analysis
Scott Deerwester, Susan Dumais, George W. Furnas et al. · Journal of the American Society for Information Science · 1990 · 12.7K citations