2001 · 58 citations · 7 references
EngineeringCommunicationCorpus LinguisticsJournalismText MiningNatural Language ProcessingInformation RetrievalData ScienceTopic IdentificationSeveral Statistical MethodsComputational LinguisticsDocument ClassificationNews AnalyticsLanguage StudiesContent AnalysisDocument ClusteringKnowledge DiscoveryTerminology ExtractionInformation ExtractionTopic ModelKeyword ExtractionE-mail CorpusLinguistics
This work presents several statistical methods for topic identification on two kinds of textual data: newspaper articles and e-mails. Five methods are tested on these two corpora: topic unigrams, cache model, TFIDF classijier, topic peqdexity, and weighted model. Our work aims to study these methods by confronting them to very diferent data. This study is very fruitful for our research. Statistical topic identiJication methods depend not only on a corpus, but also on its type. One of the methods achieves a topic identiJcation of 80% on a general newspaper corpus but does not exceed 30% on e-mail corpus. Another method gives the best result on e-mails, but has not the same behavior on a newspaper corpus. We also show in this paper that almost all our methods achieve good results in retrieving the first two manually annotated labels.
7
A re-examination of text categorization methods
Yiming Yang, Xin Liu · 1999 · 2.7K citations · Full text
Developments in Automatic Text Retrieval
Gerard Salton · Science · 1991 · 621 citations
Large Text Files, Engineering, Intelligent Information Retrieval +20
Approaches to topic identification on the switchboard corpus
J. McDonough, Kenney Ng, P. Jeanrenaud et al. · 2002 · 61 citations
A maximum likelihood model for topic classification of broadcast news
Richard Schwartz, Toru Imai, Francis Kubala et al. · 1997 · 57 citations