ACM Transactions on Asian Language Information Processing · 2002 · 45 citations · 4 references
EngineeringScripting LanguageDetection MechanismCorpus LinguisticsText MiningSpeech RecognitionApplied LinguisticsNatural Language ProcessingSyntaxLanguage DocumentationInformation RetrievalString-searching AlgorithmText SegmentationComputational LinguisticsLanguage TestingLanguage EngineeringGrammarLanguage StudiesTarget Text DocumentCharacter RecognitionMachine TranslationComputer ScienceN-gram-based LanguageDetermination MethodLanguage RecognitionLanguage CorpusN-gram StatisticsText ProcessingLinguistics
An N-gram-based language, script, and encoding scheme-detection method is introduced in this article. The method detects language, script, and encoding schemes using a target text document encoded by computer by checking how many byte sequences of the target match the byte sequences that can appear in the texts belonging to a language, script, and encoding scheme. This detection mechanism is different from conventional N-gram-based methods in that its threshold for any category is uniquely predetermined. The method was originally created for a survey of web pages conducted to find how many web pages are written in a particular language, script, and encoding scheme. The requirement is that the method must be able to respond to either "correct answer" or "unable to detect" where "unable to detect" includes "other than registered." There are some minor problems with this method, but its effectiveness as a language, script, and encoding scheme-detection method has been confirmed by experiments.
4
N-gram-based text categorization
William B. Cavnar, John M. Trenkle · 1994 · 1.5K citations
Text categorization using compression models
Eibe Frank, Chang Chui, Ian H. Witten · 2002 · 84 citations · Full text
Proceedings IEEE Data Compression Conference
James A. Elliott, Peter Grant, Garrett Sexton · 1991 · 81 citations