2011 · 15 citations · 12 references
EngineeringCross-lingual RepresentationMultilingualismTextual ComponentsCorpus LinguisticsText MiningApplied LinguisticsNatural Language ProcessingLanguage DocumentationData SciencePattern RecognitionBilingual DocumentsComputational LinguisticsLanguage EngineeringTextual EntitiesDocument ClassificationLanguage StudiesMachine TranslationAutomatic ClassificationCross-language RetrievalLanguage RecognitionLinguisticsChinese Components
In this paper, we propose a method for classifying textual entities of bilingual documents written in Chinese and English. In contrast to earlier works that performed classification on the level of text lines or documents, we apply our method to the level of textual components, as we must first identify Chinese components before merging them into intact characters and sending the latter characters to a Chinese recognizer. To cope with a large training data set containing 365,672 samples, we employ a decision-tree support vector machine (DTSVM) method, which decomposes a given data space into small regions and trains local SVMs on those regions. By applying this method to train classifiers on various combinations of feature types, we were able to complete each training process within 3,500 seconds and achieve higher than 99.6% test accuracy in classifying a textual component into Chinese, alphanumeric, and punctuation. Moreover, the classification had no strong bias towards any of the three categories.
12
Classification and Regression Trees.
John Van Ryzin, Leo Breiman, Jerome H. Friedman et al. · Journal of the American Statistical Association · 1986 · 21K citations
J. R. Quinlan · Machine Learning · 1986 · 12.3K citations · Full text
Classification and regression trees
J. Praagman · European Journal of Operational Research · 1985 · 10.2K citations