Knowledge discovery in Textual Databases (KDT)

TLDR

The information age has produced a rapid growth of electronic data that traditional data‑handling methods cannot manage, prompting the need for Knowledge Discovery in Databases (KDD) techniques to uncover patterns in large structured datasets and, increasingly, in unstructured textual information. This work aims to extend KDD to textual data by imposing a structured, hierarchically organized annotation of text articles, thereby enabling data summarization, pattern exploration, and trend analysis. The authors employ a text categorization approach that automatically assigns hierarchical concepts to articles, providing a lightweight, extractable structure suitable for KDD operations.

Abstract

The information age is characterized by a rapid growth in the amount of information available in electronic media. Traditional data handling methods are not adequate to cope with this information flood. Knowledge Discovery in Databases (KDD) is a new paradigm that focuses on computerized exploration of large amounts of data and on discovery of relevant and interesting patterns within them. While most work on KDD is concerned with structured databases, it is clear that this paradigm is required for handling the huge amount of information that is available only in unstructured textual form. To apply traditional KDD on texts it is necessary to impose some structure on the data that would be rich enough to allow for interesting KDD operations. On the other hand, we have to consider the severe limitations of current text processing technology and define rather simple structures that can be extracted from texts fairly automatically and in a reasonable cost. We propose using a text categorization paradigm to annotate text articles with meaningful concepts that are organized in hierarchical structure. We suggest that this relatively simple annotation is rich enough to provide the basis for a KDD framework, enabling data summarization, exploration of interesting patterns, and trend analysis. This research combines the KDD and text categorization paradigms and suggests advances to the state of the art in both areas.

References

Page 1

	Year	Citations

Page 1