Skip to Content
AlgorithmsAlgorithms
  • Article
  • Open Access

10 January 2022

Graph Based Feature Selection for Reduction of Dimensionality in Next-Generation RNA Sequencing Datasets

,
and
1
Department of Computing and Information Technology, University of Embu, P.O. Box 6-60100, Embu 60100, Kenya
2
School of Computing and Information Technology, Jomo Kenyatta University of Agriculture and Technology, P.O. Box 62000-00200, Nairobi 00200, Kenya
3
Biotechnology Research Institute, Kenya Agricultural and Livestock Research Organization, P.O. Box 362-00902, Kikuyu 00902, Kenya
*
Author to whom correspondence should be addressed.

Abstract

Analysis of high-dimensional data, with more features ( p ) than observations ( N ) ( p > N ), places significant demand in cost and memory computational usage attributes. Feature selection can be used to reduce the dimensionality of the data. We used a graph-based approach, principal component analysis (PCA) and recursive feature elimination to select features for classification from RNAseq datasets from two lung cancer datasets. The selected features were discretized for association rule mining where support and lift were used to generate informative rules. Our results show that the graph-based feature selection improved the performance of sequential minimal optimization (SMO) and multilayer perceptron classifiers (MLP) in both datasets. In association rule mining, features selected using the graph-based approach outperformed the other two feature-selection techniques at a support of 0.5 and lift of 2. The non-redundant rules reflect the inherent relationships between features. Biological features are usually related to functions in living systems, a relationship that cannot be deduced by feature selection and classification alone. Therefore, the graph-based feature-selection approach combined with rule mining is a suitable way of selecting and finding associations between features in high-dimensional RNAseq data.

1. Introduction

Analysis of high-dimensional data, with more features ( p ) than observations ( N ) ( p > > N ), places significant demand in cost and memory computational usage attributes. Analysis of high-dimensional data requires both high computational costs and computer memory usage [1]. Dimensionality reduction techniques that minimize the features in the original data without losing any important information have been used to address this challenge [2]. This reduces storage space and computational time and removes redundant, noisy and irrelevant data while improving algorithm efficiency as well as accuracy [3]. Feature selection and extraction constitute common approaches used to reduce dimensionality with the former identifying subsets of sufficient informative features that define the data. Major feature-selection techniques include filters, wrappers and embedded/hybrid techniques [4]. Wrappers follow a greedy search approach that wraps the feature selection around a learning algorithm and subsequently uses the accuracy of performance or error rate of the classification process as a criterion of feature evaluation. However, this approach is slower for a large feature space because every feature must be evaluated using a trained classifier. The filter method checks the features relying on the intrinsic characteristics (information, dependency, consistency and distance) prior to the learning tasks [5], which requires lower computational resources. Filters are thus faster but inefficient in classification relative to wrapper methods. Hybrid/embedded techniques are a combination of filters and wrappers [6]. We refer readers to the following articles and references therein for detailed reviews of feature-selection methods [7,8,9,10,11,12,13].
Feature-extraction techniques obtain the most important features from a dataset [14]. Principal component analysis (PCA), for example, uses covariance to extract relevant features from high-dimensional data [15]. A major drawback of feature-extraction techniques is their inaccuracy when mapping from a high-dimensional space to a low-dimensional space, which leads to the loss of data interpretability [16]. Another approach used in the analysis of big data is machine learning which has been defined as the ability of machines to learn without being programmed [17]. Classification is a supervised machine learning approach where labels are provided with the input data. Trained classifiers have been used in areas such as diabetes studies [18], cancer studies [19] and medical diagnosis [14]. Metrics such as accuracy, sensitivity, specificity, recall, robustness, computational scalability and computational cost are then used to evaluate the performance of the classifiers [20].
Microarray and next-generation sequencing (NGS) are two high-throughput technologies that generate large volumes of high-dimensional biological data [21]. Discovering meaningful associations in these kinds of data is time-consuming and computationally demanding. Association rule mining [22] is a data mining approach that has been widely used to discover high-frequency co-occurrence of items in databases. The Apriori algorithm iteratively finds associations among or between features in a dataset [23]. A major advantage of this method is that it does not require prior knowledge (training) on the dataset. However, a discretization step is required to transform continuous variables into one-hot encoding, a data structure recognized by Apriori. When biological data are analyzed using this approach, the output of association rule mining reflects expected biological associations between different features. In this study, we highlighted the effects of various feature-selection methods on classification and association rule mining. Based on the results, we recommend a graph-based feature-selection method as a more suitable dimensionality reduction strategy when selecting features that can be used for classification and association rule mining from high-dimensional RNA-Seq data. The proposed graph-based approach is important because (1) only informative features are selected from the high-dimensional data based on their associations in the graph (nodes and edges), (2) the association between features in the graph (nodes and edges) reflects potential biological associations in vivo, and (3) association rule mining confirms the associations observed in the graphs, and this can be used to predict phenotypes of unknown features.

3. Materials and Methods

3.1. Data Source and Data Type

Two RNA-Seq datasets were used in this study (Table 1). The first dataset referred to as small-cell lung cancer (SCLC) and had 86 samples with two classes: 79 cancer cells and 7 normal cells. The second dataset denoted as non-small-cell lung cancer (NSCLC) had a total of 218 samples whereby 199 samples were non-small-cell lung cancer and 19 were normal cells. Both datasets used in this study were highly imbalanced whereby the cancer samples were the majority class while normal samples were the minority class. We used synthetic minority oversampling technique (SMOTE) algorithm to balance the datasets. The minority classes were increased based on the 5 k-nearest neighbors to nearly equal classes. After addressing class imbalance, we then performed 70:30 subsets on data and performed 10-fold cross validation of the accuracy and then recorded the accuracy and F-measure. This was repeated 20 times across each feature-selection method. We then used Kruskal–Wallis H-statistic to test if there was significant difference in the mean ranks of the groups.
Table 1. Summary of the datasets used in this study.

3.2. Data Preprocessing

The raw count data were preprocessed to filter out any features with zero counts. This was achieved via normalization using the upper quartile method implemented in edgeR package [55]. A scaling factor of 75th percentile of every count was calculated after removing features with zero counts using Equation (4):
d j U Q = U Q k g j g = 1 G K g j
where U Q X is the upper quartile of sample X of j th sample of normalized counts and K g j > 0 .

3.3. Feature Selection

After normalization, we performed feature selection on both datasets as a means of dimensionality reduction. We used principal component analysis (PCA), recursive feature elimination (RFE) and a graph-based approach.

3.3.1. Principal Component Analysis (PCA)

Assuming dataset x 1 x 2 , . x m has inputs of n dimensions, these n dimension data must be reduced to k d i m e n s i o n a l   k n using PCA.
The first step in PCA is raw data standardization whereby the raw data should have unit variance and zero mean defined in Equation (5).
x j i = x j i x ¯ j σ j j
In the second step, a covariance matrix of the raw data is calculated as shown in Equation (6).
Σ = 1 m i m x i x i   T Σ R n n
The third stage is calculation of eigenvector and eigenvalue of the covariance using Equation (7).
u T Σ = λ , μ
U = u 1 ,   u 2 ,   . u n , ,   u i ,     R n
In the fourth step, raw data are projected into a k-dimensional subspace, and this is followed by choosing the top k eigenvector of covariance matrix. The corresponding vector is calculated as shown in Equation (8).
x i n e w = u 1 T x i u 2 T x i u k T x i R k
The raw data with n dimensionality is reduced to a new k dimensional representation of data.

3.3.2. Recursive Feature Elimination

The second feature-selection approach used was recursive feature elimination (RFE). This is a recursive process where features are ranked based on their importance [56]. RFE employs machine learning models in computing feature-relevant scores. RFE first trains the model using all features and then computes the relevance score of every feature in the dataset. All the features with lowest relevance score are ignored, and this is followed by model retraining for computation of new relevant feature scores. This process is repeated until the final desired features are obtained [57].

3.3.3. Proposed Graph-Based Approach

The following steps were used in graph-based feature selection.
i. 
Normalization
Calculate scaling factor (Equation (9)).
d j U Q = U Q K g j g = 1 G K g j  
ii. 
Network construction
Calculate PCC of the normalized features (Equation (10)):
r x i , x j = k = 1 m x k i , x ¯ i y k y ¯ k = 1 m x k i , x i ¯ 2 Σ k = 1 m y k y ¯ 2
iii. 
Determine threshold (Equation (11)):
a i j = p o w e r ( S i j , β ) = s i j β
iv. 
Construct a topological overlap matrix TOM based on the adjacency of a i j (Equation (12)):
T o m i j = Σ u 1 , j a u i a u j + a i j m i n k i , k j + 1 a i j
v. 
Filter the resulting network using maximal cliques:
  • Apply Bron–Kerbosch algorithm to find all possible cliques within the graph that has been filtered where a clique is a complete subgraph C G .
  • Determine a set of C of the maximal cliques C i m a x where a maximal clique is a complete subgraph C i G which is not a subset of another complete subgraph C i G .
  • Whenever there is C i m a x > 1 , where C i m a x   C a rating function r :   C i m a x   , C i M a x     C is applied and maximal clique with the highest score selected.

3.4. Classification

Experiments were performed in WEKA framework and RStudio on a PC intel i3 CPU, 2 cores and 8 GB of RAM. Three classifiers namely Naïve Bayes, sequential minimal optimization (SMO) and multilayer perceptron were used to compare the classification accuracy of the different feature-selection methods. Features were the dependent variables while class was the independent variable. In this case, the expression levels (counts) of the features (genes) were used to predict the tissue type (diseased/normal). Classification accuracy, root mean squared error (RMSE), mean absolute error (MAE), kappa statistic (KS) and the time taken to build the model for every classifier were recorded before and after feature selection using the various approaches.

3.5. Discretization

To discretize the data, we used equal-width interval discretization. This algorithm divides the range of values for a feature into equally sized bins that are represented by k parameter provided by the user. The Algorithm 1 finds minimum and maximum observed values through sorting of continuous features A = a 0 , a 1 , .. a n where a m i n , = a 0 , and a m a x = a n , . To compute the interval, the range of observed values for the variable is divided into equally sized bins as described in Equations (13) and (14) below:
I n t e r v a l = a m a x a m i n k
B o u n d a r i e s = a m i n + i i n t e r v a l
Algorithms 1: Equal-Width Interval Discretization Steps
Input: continuous values of A = a 0 , a 1 , .. a n , with k being number of parts where k > 0
Output: discretized values
Step 1. Sort values of A in ascending order
Step 2: Calculate interval using Equation (19)
Step 3: Divide the data into 3 bins
Step 4: Place the values of the array in the same boundary

3.6. Association Rule Mining (ARM)

ARM is a process of determining possible association rules amongst items in a large database or dataset. Let I = l 1 , l 2 l k be a set of k features and T be a transaction that contains items such that T I and D is a database of transaction records. An association rule is of type X a n t e c e d e n t     Y C o n s e q u e n t where X   a n d   Y are features such that X     Y = ø .

4. Results and Discussion

4.1. Data Preprocessing

The datasets used in this study had 28,089 initial features as summarized in Table 2. Preprocessing involved elimination of non-differentially expressed genes and normalization which resulted in 12.2% features for small-cell lung cancer and 43.2% features for non-small-cell lung cancer after preprocessing.
Table 2. Output normalization and feature selection using PCA, RFE and a graph-based approach.

4.2. Feature Selection

After preprocessing, we used three feature-selection approaches to filter the features further to retain only informative features. The RFE retained the highest number of features in both datasets followed by PCA as shown in Table 2 and Figure 1.
Figure 1. Features selected by each of the methods from the two datasets.
The graph-based feature-selection approach retained 80 for the SCLC dataset and 134 features for the NSCLC dataset. This is because the graph considers only the connecting features and the filtering step using maximal cliques retaining only the features that had the highest maximal clique score. Figure 2 presents the networks for the two datasets before and after filtering using maximal cliques.
Figure 2. Network diagrams for the 2 datasets: figures on the left side represent the network before filtering while those on the right show the networks after filtering with maximal clique. On the filtered networks, different colors denote expression levels with red color showing features that were highly expressed.

4.3. Classification

In the next step, we compared the performance of three classifiers with the features selected in the previous step as input while the raw features were our baseline. Table 3 summarizes the performance of the various classifiers before and after feature selection. The accuracy value, root mean squared error (RMSE), mean absolute error (MAE), kappa statistic (KS) and the time taken to build the model for every classifier, arising from the 10-fold cross validation are given. Overall, accuracy levels after selection ranged between 94.186 and 100% depending on the classification method used and the dataset (Table 3).
Table 3. Classification results after feature selection.
NB performed better on features selected using PCA and a graph-based approach whereby accuracy, MAE, kappa and time taken improved as compared to unfiltered features and RFE-selected features where there was no difference. This can be attributed to the working principle of RFE where the optimal number of features is not known apriori (in advance) [58]. A Kruskal–Wallis test showed that there is no significant difference between the mean ranks of the groups ( p < 0.05 ), i.e., 20 iterations for each of the feature-selection methods.
A study by Furat and Ibrikci [59] used five tumor types of gene expression cancer RNA-Seq data and, using Naïve Bayes with 10-fold cross validation, achieved an accuracy of 98.7516%. This shows that NB accuracy levels will vary with the dataset being analyzed. In dataset GSE81089, which had a larger sample size of selected features, SMO and MLP achieved 100% accuracy when feature selection was performed prior to classification, and in fact, the PCA-selected features could be classified at 100% accuracy by all three classifiers. In the smaller dataset, SCLC dataset, accuracy levels were also lower. Notable is that a graph-based feature-selection approach gave the best classification results in the two datasets and also took the least time to execute. The time required to build the model improved after feature selection across the three classifiers though MLP required the longest duration and a graph-based approach the shortest (Table 3).

4.4. Association Rule Mining

Features from PCA, RFE and graph-based selection methods were discretized and analyzed to find possible associations using Apriori. The resulting number of rules, maximum confidence, support and lift values are summarized in Table 4.
Table 4. Rules generated using Apriori from features selected using different approaches.
The graph-based feature-selection approach gave 15 and 36 non-redundant rules, respectively, from the two datasets at a support of 0.5 confidence value of 0.9 and a lift of 2. The other feature-selection methods did not generate any rules at a support of 0.5. Features selected by RFE had the lowest maximum support and lift, and this led to the generation of too many redundant rules. For the PCA-based feature selection, support ranged between 0.405 and 0.425 with a total of 38 rules for the first dataset and 36 rules for the second dataset (Table 4). The top 10 rules are shown in Table 5.
Table 5. A summary of top ten rules generated from the two datasets after graph-based.
Association rules are represented as X => Y, where X and Y are items contained within a dataset/database, and XY = ø. X is the antecedent, and Y is the consequent (Table 5). It means that whenever X, which is the antecedent, is present, even Y, which is the consequent, will be present. Support indicates the frequency of the itemset appearance in the dataset, and the confidence indicates how often a rule has been found to be true. A support value of 0.5 means 50% of the items (genes) are found in the transaction and 90% of the rules are true (Confidence). The lower support means that most of the items are not frequently found together. The lift value is used to measure the rule importance. A lift of greater than 2 achieved by the graph-based feature-selection approach indicates the degree to which any two occurrences depend on each other, and this is an indication those rules are useful in consequent prediction.
External factors that play a critical role in the success of this type of study include the choice of the technology used in generating the data since it determines the volume and quality of the data across the replicates. The cost of generating the data also limits the number of samples as well as the volume of data. In the biological space, what can be defined as control/normal samples is a gray area, and therefore this may have a bearing on downstream analysis while at the same time bringing about class imbalance. Internal threats to this kind of study include experimental noise in the data as well as the assumption that gene expression level is uniform across cells. Co-regulation and co-expression between the features can also lead to redundancy in the datasets. Features without an assigned biological function may also not be informative unless in vitro experiments are designed to validate the function.

5. Conclusions

In this study, we used three different feature-selection methods to select informative features from two different cancer datasets. Most existing feature-selection techniques assume that features are independent of one another. However, this assumption ignores the fact that biological features are usually related because of their function in living systems. We evaluated the performance of three classifiers on the selected features. Features selected using a graph-based approach with maximal clique could be classified with high accuracy when compared to PCA and RFE. These features also gave informative rules with a higher support and lift as compared to those selected using PCA and RFE. Therefore, the proposed graph-based feature-selection approach combined with rule mining is a suitable way of selecting and finding associations between features in high-dimensional RNA-Seq data.

Author Contributions

Conceptualization, C.G. and R.R.; methodology, C.G., P.O.M. and R.R.; formal analysis, C.G.; writing—original draft preparation, C.G.; writing—review and editing, C.G., R.R. and P.O.M.; supervision, R.R and P.O.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

Datasets used in this study are publicly available at https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE81089 (accessed on 2 August 2021); https://www.ncbi.nlm.nih.gov/bioproject/?term=GSE60052 (accessed on 2 August 2021).

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Jindal, P.; Kumar, D. A review on dimensionality reduction techniques. Int. J. Comput. Appl. 2017, 173, 42–46. [Google Scholar] [CrossRef] [Scilit]
  2. Nguyen, L.H.; Holmes, S. Ten quick tips for effective dimensionality reduction. PLoS Comput. Biol. 2019, 15, e1006907. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Zebari, R.; Abdulazeez, A.; Zeebaree, D.; Zebari, D.; Saeed, J. A comprehensive review of dimensionality reduction techniques for feature selection and feature extraction. J. Appl. Sci. Technol. Trends 2020, 1, 56–70. [Google Scholar] [CrossRef] [Scilit]
  4. Abdulrazzaq, M.B.; Saeed, J.N. A Comparison of Three Classification Algorithms for Handwritten Digit Recognition. In Proceedings of the 2019 International Conference on Advanced Science and Engineering (ICOASE), Zakho-Duhok, Iraq, 2–4 April 2019; pp. 58–63. [Google Scholar]
  5. Mafarja, M.; Mirjalili, S. Whale optimization approaches for wrapper feature selection. Appl. Soft Comput. 2018, 62, 441–453. [Google Scholar] [CrossRef] [Scilit]
  6. Yu, L.; Liu, H. Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution. In Proceedings of the 20th international conference on machine learning (ICML-03), Washington, DC, USA, 21–24 August 2003; pp. 856–863. [Google Scholar]
  7. Chandrashekar, G.; Sahin, F. A survey on feature selection methods. Comput. Electr. Eng. 2014, 40, 16–28. [Google Scholar] [CrossRef] [Scilit]
  8. Jović, A.; Brkić, K.; Bogunović, N. A review of feature selection methods with applications. In Proceedings of the 2015 38th International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO), Opatija, Croatia, 25–29 May 2015; pp. 1200–1205. [Google Scholar]
  9. Mlambo, N.; Cheruiyot, W.K.; Kimwele, M.W. A survey and comparative study of filter and wrapper feature selection techniques. Int. J. Eng. Sci. 2016, 5, 57–67. [Google Scholar]
  10. Urbanowicz, R.J.; Meeker, M.; La Cava, W.; Olson, R.S.; Moore, J.H. Relief-based feature selection: Introduction and review. J. Biomed. Inform. 2018, 85, 189–203. [Google Scholar] [CrossRef] [Scilit]
  11. Abiodun, E.O.; Alabdulatif, A.; Abiodun, O.I.; Alawida, M.; Alabdulatif, A.; Alkhawaldeh, R.S. A systematic review of emerging feature selection optimization methods for optimal text classification: The present state and prospective opportunities. Neural Comput. Appl. 2021, 33, 15091–15118. [Google Scholar] [CrossRef] [Scilit]
  12. Piles, M.; Bergsma, R.; Gianola, D.; Gilbert, H.; Tusell, L. Feature Selection Stability and Accuracy of Prediction Models for Genomic Prediction of Residual Feed Intake in Pigs Using Machine Learning. Front. Genet. 2021, 12, 137. [Google Scholar] [CrossRef] [Scilit]
  13. Yang, P.; Huang, H.; Liu, C. Feature selection revisited in the single-cell era. Genome Biol. 2021, 22, 321. [Google Scholar] [CrossRef] [Scilit]
  14. Arowolo, M.O.; Adebiyi, M.O.; Adebiyi, A.A.; Olugbara, O. Optimized hybrid investigative based dimensionality reduction methods for malaria vector using KNN classifier. J. Big Data 2021, 8, 1–14. [Google Scholar] [CrossRef] [Scilit]
  15. Cateni, S.; Vannucci, M.; Vannocci, M.; Colla, V. Variable Selection and Feature Extraction through Artificial Intelligence Techniques. Available online: https://www.intechopen.com/chapters/41752 (accessed on 7 December 2021).
  16. Kim, K. An improved semi-supervised dimensionality reduction using feature weighting: Application to sentiment analysis. Expert Syst. Appl. 2018, 109, 49–65. [Google Scholar] [CrossRef] [Scilit]
  17. Samuel, A.L. Some studies in machine learning using the game of checkers. IBM J. Res. Dev. 1959, 3, 210–229. [Google Scholar] [CrossRef] [Scilit]
  18. Das, H.; Naik, B.; Behera, H. Classification of diabetes mellitus disease (DMD): A data mining (DM) approach. In Progress in Computing, Analytics and Networking; Springer: Singapore, 2018; pp. 539–549. [Google Scholar]
  19. Mazumder, D.H.; Veilumuthu, R. An enhanced feature selection filter for classification of microarray cancer data. ETRI J. 2019, 41, 358–370. [Google Scholar] [CrossRef] [Scilit]
  20. Sun, S.; Zhu, J.; Ma, Y.; Zhou, X. Accuracy, robustness and scalability of dimensionality reduction methods for single-cell RNA-seq analysis. Genome Biol. 2019, 20, 1–21. [Google Scholar] [CrossRef] [Scilit]
  21. Ai, D.; Pan, H.; Li, X.; Gao, Y.; He, D. Association rule mining algorithms on high-dimensional datasets. Artif. Life Robot. 2018, 23, 420–427. [Google Scholar] [CrossRef] [Scilit]
  22. Agrawal, R.; Imieliński, T.; Swami, A. Mining Association Rules between Sets of Items in Large Databases. In Proceedings of the 1993 ACM SIGMOD international conference on Management of Data, Washington, DC, USA, 25–28 May 1993; pp. 207–216. [Google Scholar]
  23. Liu, X.; Sang, X.; Chang, J.; Zheng, Y.; Han, Y. The water supply association analysis method in Shenzhen based on kmeans clustering discretization and apriori algorithm. PLoS ONE 2021, 16, e0255684. [Google Scholar] [CrossRef] [Scilit]
  24. Saeys, Y.; Inza, I.; Larrañaga, P. A review of feature selection techniques in bioinformatics. Bioinformatics 2007, 23, 2507–2517. [Google Scholar] [CrossRef] [Scilit]
  25. Ang, J.C.; Mirzal, A.; Haron, H.; Hamed, H.N.A. Supervised, unsupervised, and semi-supervised feature selection: A review on gene selection. IEEE/ACM Trans. Comput. Biol. Bioinform. 2016, 13, 971–989. [Google Scholar] [CrossRef] [Scilit]
  26. Ray, R.B.; Kumar, M.; Rath, S.K. Fast In-Memory Cluster Computing of Sizeable Microarray Using Spark. In Proceedings of the 2016 International Conference on Recent Trends in Information Technology (ICRTIT), Chennai, India, 8–9 April 2016; pp. 1–6. [Google Scholar]
  27. Lokeswari, Y.; Jacob, S.G. Prediction of child tumours from microarray gene expression data through parallel gene selection and classification on spark. In Computational Intelligence in Data Mining; Springer: Singapore, 2017; pp. 651–661. [Google Scholar]
  28. Peralta, D.; Del Río, S.; Ramírez-Gallego, S.; Triguero, I.; Benitez, J.M.; Herrera, F. Evolutionary feature selection for big data classification: A mapreduce approach. Math. Probl. Eng. 2015, 2015. [Google Scholar] [CrossRef] [Scilit]
  29. Sonnenburg, S.; Franc, V.; Yom-Tov, E.; Sebag, M. Pascal Large Scale Learning Challenge. In Proceedings of the 25th International Conference on Machine Learning (ICML2008) Workshop, Helsinki, Finland, 5–9 July 2008. [Google Scholar]
  30. Alghunaim, S.; Al-Baity, H.H. On the scalability of machine-learning algorithms for breast cancer prediction in big data context. IEEE Access 2019, 7, 91535–91546. [Google Scholar] [CrossRef] [Scilit]
  31. Turgut, S.; Dağtekin, M.; Ensari, T. Microarray Breast Cancer Data Classification Using Machine Learning Methods. In Proceedings of the 2018 Electric Electronics, Computer Science, Biomedical Engineerings’ Meeting (EBBT), Istanbul, Turkey, 18–19 April 2018. [Google Scholar]
  32. Matamala, N.; Vargas, M.T.; Gonzalez-Campora, R.; Minambres, R.; Arias, J.I.; Menendez, P.; Andres-Leon, E.; Gomez-Lopez, G.; Yanowsky, K.; Calvete-Candenas, J. Tumor microRNA expression profiling identifies circulating microRNAs for early breast cancer detection. Clin. Chem. 2015, 61, 1098–1106. [Google Scholar] [CrossRef] [Scilit]
  33. Morovvat, M.; Osareh, A. An ensemble of filters and wrappers for microarray data classification. Mach. Learn. Appl. An. Int. J. 2016, 3, 1–17. [Google Scholar] [CrossRef] [Scilit]
  34. Goswami, S.; Das, A.K.; Guha, P.; Tarafdar, A.; Chakraborty, S.; Chakrabarti, A.; Chakraborty, B. An approach of feature selection using graph-theoretic heuristic and hill climbing. Pattern Anal. Appl. 2019, 22, 615–631. [Google Scholar] [CrossRef] [Scilit]
  35. Zhang, Z.; Hancock, E.R. A Graph-Based Approach to Feature Selection. In International Workshop on Graph-Based Representations in Pattern Recognition; Springer: Berlin/Heidelberg, Germany, 2011; pp. 205–214. [Google Scholar]
  36. Schroeder, D.T.; Styp-Rekowski, K.; Schmidt, F.; Acker, A.; Kao, O. Graph-Based Feature Selection Filter Utilizing Maximal Cliques. In Proceedings of the 2019 Sixth International Conference on Social Networks Analysis, Management and Security (SNAMS), Granada, Spain, 22–25 October 2019; pp. 297–302. [Google Scholar]
  37. Roffo, G.; Castellani, U.; Vinciarelli, A.; Cristani, M. Infinite feature selection: A graph-based feature filtering approach. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 43, 12. [Google Scholar] [CrossRef] [Scilit]
  38. Rana, P.; Thai, P.; Dinh, T.; Ghosh, P. Relevant and Non-Redundant Feature Selection for Cancer Classification and Subtype Detection. Cancers 2021, 13, 4297. [Google Scholar] [CrossRef] [Scilit]
  39. Nguyen, H.; Thai, P.; Thai, M.; Vu, T.; Dinh, T. Approximate k-Cover in Hypergraphs: Efficient Algorithms, and Applications. arXiv 2019, arXiv:1901.07928. [Google Scholar]
  40. Lu, S.J.; Xie, J.; Li, Y.; Yu, B.; Ma, Q.; Liu, B.Q. Identification of lncRNAs-gene interactions in transcription regulation based on co-expression analysis of RNA-seq data. Math. Biosci. Eng. 2019, 16, 7112–7125. [Google Scholar] [CrossRef] [Scilit]
  41. Chiclana, F.; Kumar, R.; Mittal, M.; Khari, M.; Chatterjee, J.M.; Baik, S.W. ARM–AMO: An efficient association rule mining algorithm based on animal migration optimization. Knowl. Based Syst. 2018, 154, 68–80. [Google Scholar]
  42. Wen, F.; Zhang, G.; Sun, L.; Wang, X.; Xu, X. A hybrid temporal association rules mining method for traffic congestion prediction. Comput. Ind. Eng. 2019, 130, 779–787. [Google Scholar] [CrossRef] [Scilit]
  43. Shui, Y.; Cho, Y.-R. Filtering Association Rules in GENE Ontology Based on Term Specificity. In Proceedings of the 2016 IEEE international conference on bioinformatics and biomedicine (bibm), Shenzhen, China, 15–18 December 2016; pp. 1314–1321. [Google Scholar]
  44. Agapito, G.; Cannataro, M.; Guzzi, P.H.; Milano, M. Using GO-WAR for mining cross-ontology weighted association rules. Comput. Methods Programs Biomed. 2015, 120, 113–122. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Bhavsar, H.; Ganatra, A. A comparative study of training algorithms for supervised machine learning. Int. J. Soft Comput. Eng. (IJSCE) 2012, 2, 2231–2307. [Google Scholar]
  46. Han, J.; Pei, J.; Kamber, M. Data Mining: Concepts and Techniques; The Morgan Kaufmann Series in Data Management Systems 5.4; Morgan Kaufmann Publishers: Waltham, MA, USA, 2011; pp. 83–124. [Google Scholar]
  47. Hall, M.; Frank, E.; Holmes, G.; Pfahringer, B.; Reutemann, P.; Witten, I.H. The WEKA data mining software: An update. ACM SIGKDD Explor. Newsl. 2009, 11, 10–18. [Google Scholar] [CrossRef] [Scilit]
  48. Vapnik, V.N. An overview of statistical learning theory. IEEE Trans. Neural Netw. 1999, 10, 988–999. [Google Scholar] [CrossRef] [Scilit]
  49. Tanwani, A.K.; Afridi, J.; Shafiq, M.Z.; Farooq, M. Guidelines to Select Machine Learning Scheme for Classification of Biomedical Datasets. In Proceedings of the European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics; Springer: Berlin/Heidelberg, Germany, 2009; pp. 128–139. [Google Scholar]
  50. Carletta, J. Assessing agreement on classification tasks: The kappa statistic. arXiv 1996, arXiv:9602004. [Google Scholar]
  51. Chai, T.; Draxler, R.R. Root mean square error (RMSE) or mean absolute error (MAE)?—Arguments against avoiding RMSE in the literature. Geosci. Model Dev. 2014, 7, 1247–1250. [Google Scholar] [CrossRef] [Scilit]
  52. Dunham, M.H.; Sridhar, S. Data Mining: Introductory and Advanced Topics, Dorling Kindersley; Pearson Education: New Delhi, India, 2006. [Google Scholar]
  53. Jiang, L.; Huang, J.; Higgs, B.W.; Hu, Z.; Xiao, Z.; Yao, X.; Conley, S.; Zhong, H.; Liu, Z.; Brohawn, P. Genomic landscape survey identifies SRSF1 as a key oncodriver in small cell lung cancer. PLoS Genet. 2016, 12, e1005895. [Google Scholar] [CrossRef] [Scilit]
  54. Djureinovic, D.; Hallström, B.M.; Horie, M.; Mattsson, J.S.M.; La Fleur, L.; Fagerberg, L.; Brunnström, H.; Lindskog, C.; Madjar, K.; Rahnenführer, J. Profiling cancer testis antigens in non–small-cell lung cancer. JCI Insight 2016, 1, e86837. [Google Scholar] [CrossRef] [Scilit]
  55. Bullard, J.; Purdom, E.; Hansen, K.D.; Dudoit, S. Evaluation of statistical methods for normalization and differential expression in mrna-seq experiments. BMC Bioinform. 2010, 11, 94. [Google Scholar] [CrossRef] [Scilit]
  56. Ustebay, S.; Turgut, Z.; Aydin, M.A. Intrusion Detection System with Recursive Feature Elimination by Using Random Forest and Deep Learning Classifier. In Proceedings of the International Congress on Big Data, Deep Learning and Fighting Cyber Terrorism (IBIGDELFT), Ankara, Turkey, 3–4 December 2018; pp. 71–76. [Google Scholar]
  57. Gunduz, H. An efficient stock market prediction model using hybrid feature reduction method based on variational autoencoders and recursive feature elimination. Financ. Innov. 2021, 7, 1–24. [Google Scholar] [CrossRef] [Scilit]
  58. Artur, M. Review the performance of the Bernoulli Naïve Bayes Classifier in Intrusion Detection Systems using Recursive Feature Elimination with Cross-validated selection of the best number of features. Procedia Comput. Sci. 2021, 190, 564–570. [Google Scholar] [CrossRef] [Scilit]
  59. Furat, F.G.; İbrikçi, T. Tumor Type Detection Using Naïve Bayes Algorithm on Gene Expression Cancer RNA-Seq Data Set. Lung Cancer 2019, 10, 13. [Google Scholar]
Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.