Next Article in Journal
Action and Entropy in Heat Engines: An Action Revision of the Carnot Cycle
Previous Article in Journal
Area Entropy and Quantized Mass of Black Holes from Information Theory
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Measuring the Effectiveness of Adaptive Random Forest for Handling Concept Drift in Big Data Streams

by
Abdulaziz O. AlQabbany
1,2 and
Aqil M. Azmi
1,*
1
Department of Computer Science, College of Computer & Information Sciences, King Saud University, Riyadh 11543, Saudi Arabia
2
King Abdulaziz City for Science and Technology, Riyadh 12371, Saudi Arabia
*
Author to whom correspondence should be addressed.
Entropy 2021, 23(7), 859; https://doi.org/10.3390/e23070859
Submission received: 6 April 2021 / Revised: 23 June 2021 / Accepted: 30 June 2021 / Published: 4 July 2021
(This article belongs to the Section Information Theory, Probability and Statistics)

Abstract

We are living in the age of big data, a majority of which is stream data. The real-time processing of this data requires careful consideration from different perspectives. Concept drift is a change in the data’s underlying distribution, a significant issue, especially when learning from data streams. It requires learners to be adaptive to dynamic changes. Random forest is an ensemble approach that is widely used in classical non-streaming settings of machine learning applications. At the same time, the Adaptive Random Forest (ARF) is a stream learning algorithm that showed promising results in terms of its accuracy and ability to deal with various types of drift. The incoming instances’ continuity allows for their binomial distribution to be approximated to a Poisson(1) distribution. In this study, we propose a mechanism to increase such streaming algorithms’ efficiency by focusing on resampling. Our measure, resampling effectiveness (ρ), fuses the two most essential aspects in online learning; accuracy and execution time. We use six different synthetic data sets, each having a different type of drift, to empirically select the parameter λ of the Poisson distribution that yields the best value for ρ. By comparing the standard ARF with its tuned variations, we show that ARF performance can be enhanced by tackling this important aspect. Finally, we present three case studies from different contexts to test our proposed enhancement method and demonstrate its effectiveness in processing large data sets: (a) Amazon customer reviews (written in English), (b) hotel reviews (in Arabic), and (c) real-time aspect-based sentiment analysis of COVID-19-related tweets in the United States during April 2020. Results indicate that our proposed method of enhancement exhibited considerable improvement in most of the situations.
Keywords: adaptive random forest; data stream; concept drift; online learning; resampling; Poisson distribution adaptive random forest; data stream; concept drift; online learning; resampling; Poisson distribution

Share and Cite

MDPI and ACS Style

AlQabbany, A.O.; Azmi, A.M. Measuring the Effectiveness of Adaptive Random Forest for Handling Concept Drift in Big Data Streams. Entropy 2021, 23, 859. https://doi.org/10.3390/e23070859

AMA Style

AlQabbany AO, Azmi AM. Measuring the Effectiveness of Adaptive Random Forest for Handling Concept Drift in Big Data Streams. Entropy. 2021; 23(7):859. https://doi.org/10.3390/e23070859

Chicago/Turabian Style

AlQabbany, Abdulaziz O., and Aqil M. Azmi. 2021. "Measuring the Effectiveness of Adaptive Random Forest for Handling Concept Drift in Big Data Streams" Entropy 23, no. 7: 859. https://doi.org/10.3390/e23070859

APA Style

AlQabbany, A. O., & Azmi, A. M. (2021). Measuring the Effectiveness of Adaptive Random Forest for Handling Concept Drift in Big Data Streams. Entropy, 23(7), 859. https://doi.org/10.3390/e23070859

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop