Europhysics Letters (EPL) · 2002 · 155 citations · 17 references
EngineeringPart-of-speech TaggingDna SequencesSemanticsCorpus LinguisticsText MiningNatural Language ProcessingKeyword DetectionSyntaxInformation RetrievalData ScienceData MiningString ProcessingComputational LinguisticsAutomatic Keyword ExtractionLanguage StudiesComputational LexicologyStrong Self-attractionTerminology ExtractionDistributional SemanticsKeyword ExtractionText ProcessingLinguistics
We show that words in a text present long-range frequency fluctuations due to a strong self-attraction, that is directly related to the relevance of the term to the text considered. The standard deviation of the distance between successive occurrences of a word is an excellent parameter to quantify this self-attraction, and provides us with an effective tool for automatic keyword extraction. DNA sequences also present the same features: "words", for example codons in the coding part of the sequences, attract between themselves.
17
The Complete Genome Sequence of <i>Escherichia coli</i> K-12
Frederick R. Blattner, Guy Plunkett, Craig A. Bloch et al. · Science · 1997 · 7.4K citations
The Automatic Creation of Literature Abstracts
H. P. Luhn · IBM Journal of Research and Development · 1958 · 3.2K citations
Random-matrix physics: spectrum and strength fluctuations
Thomas A. Brody, J. Flores, J.B. French et al. · Reviews of Modern Physics · 1981 · 2.1K citations
Long-range correlations in nucleotide sequences
Chung‐Kang Peng, Sergey V. Buldyrev, Ary L. Goldberger et al. · Nature · 1992 · 1.4K citations