Mining Complex Data Streams: Discretization, Attribute Selection and Classification

Dewan Md. Farid, Chowdhury Mofizur Rahman

Journal of Advances in Information Technology · 2013 · 42 citations · 19 references

DOIFull text

Open access

Abstract

<span style="font-family: &quot;Times New Roman&quot;,&quot;serif&quot;; font-size: 10pt; mso-bidi-font-size: 9.0pt; mso-fareast-font-family: &quot;Times New Roman&quot;; mso-ansi-language: EN-US; mso-fareast-language: EN-US; mso-bidi-language: AR-SA;" lang="EN-US">Due to the large volume of data set as well as complex and dynamic properties of data instances, several data mining algorithms have been applied for mining complex data streams in the last decades. Now a day, knowledge extraction from data streams is getting more complex because the structure of the data instance does not match the attribute values when considering the tabulated data, texts, web, images or videos etc. In this paper, we address some difficulties of mining complex data streams such as dealing with continuous attributes, input attribute selection, and classifier construction. The proposed discretization algorithm finds the possible cut points in continuous attributes using information gain heuristic and na&iuml;ve Bayesian classifier that can separate the class distributions. We evaluate the proposed algorithms on several benchmark data sets from UCI machine learning repository. The experimental results demonstrate that the proposed method improves the quality of discretization of continuous attributes and scales up the classification accuracy for different types of classification problem.</span>

References

19