Journal of the American Society for Information Science and Technology · 2003 · 67 citations · 19 references
EngineeringCongress ClassificationCorpus LinguisticsLcc HierarchyText MiningNatural Language ProcessingClassification MethodInformation RetrievalData ScienceData MiningComputational LinguisticsDocument ClassificationLcc TreeLanguage StudiesHierarchical ClassificationAutomatic ClassificationPredictive AnalyticsKnowledge DiscoveryIntelligent ClassificationComputer ScienceCongress Subject HeadingsCongress ClassificationsLinguistics
Abstract This paper addresses the problem of automatically assigning a Library of Congress Classification (LCC) to a work given its set of Library of Congress Subject Headings (LCSH). LCCs are organized in a tree: The root node of this hierarchy comprises all possible topics, and leaf nodes correspond to the most specialized topic areas defined. We describe a procedure that, given a resource identified by its LCSH, automatically places that resource in the LCC hierarchy. The procedure uses machine learning techniques and training data from a large library catalog to learn a model that maps from sets of LCSH to classifications from the LCC tree. We present empirical results for our technique showing its accuracy on an independent collection of 50,000 LCSH/LCC pairs.
19
Ian H. Witten, Eibe Frank · ACM SIGMOD Record · 2002 · 5.2K citations
Inductive learning algorithms and representations for text categorization
Susan Dumais, John Platt, David Heckerman et al. · 1998 · 1.5K citations
Hierarchically Classifying Documents Using Very Few Words
Daphne Koller, Mehran Sahami · 1997 · 840 citations
Hierarchical classification of Web content
Susan Dumais, Hao Chen · 2000 · 803 citations · Full text