Next Article in Journal
The Combined Effect of Silver Precursor and Sodium Salt on the Structure and Crystallization Behavior of Photo-Thermo-Refractive Glass
Previous Article in Journal
Deep Eutectic Solvents for Sustainable Extraction of Bioactive Compounds from Biomass: Mechanistic Insights and Scale-Up Challenges
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Novel Strategy for Rapid Quantification of Multiple Quality Indicators and Grade Discrimination of Atractylodis macrocephalae Rhizoma Based on Electronic Nose, Electronic Tongue and Machine-Learning Algorithms

1
School of Chinese Materia Medica, Beijing University of Chinese Medicine, Beijing 102401, China
2
School of Traditional Chinese Materia Medica, Shanxi University of Chinese Medicine, Taiyuan 030619, China
*
Authors to whom correspondence should be addressed.
These authors contributed equally to this work.
Molecules 2026, 31(5), 881; https://doi.org/10.3390/molecules31050881
Submission received: 27 December 2025 / Revised: 3 March 2026 / Accepted: 4 March 2026 / Published: 6 March 2026

Abstract

Atractylodes macrocephala Rhizoma (AMR) is a frequently used medicinal herb for treating gastrointestinal disorders, with its quality influenced by factors such as origin and cultivation duration. Traditional quality control methods for AMR are time-consuming and invasive, making the development of faster and more efficient alternatives urgently needed. This study aims to utilize electronic nose (E-nose) and electronic tongue (E-tongue) to achieve the acquisition of odor–taste two-dimensional information of AMR. Integrating this approach with machine learning (ML) enables intelligent transformation from “experience-driven” to “data-driven” quality assessment, thereby developing a rapid and cost-effective quality control strategy for AMR. Feature-extraction and feature-selection techniques were employed to optimize back-propagation neural network (BPNN) classification and regression models for eight key quality markers, selecting the optimal feature subset. Additionally, nine machine-learning algorithms were applied with the optimal feature subset to establish classification models for different AMR grades and quantitative regression models for eight components based on E-nose and E-tongue data. The results demonstrated that the E-tongue combined with the k-nearest neighbors (KNN) algorithm could achieve a rapid classification of AMR grades with an accuracy of 95.56%. It also successfully predicted the contents of the extract, volatile oil, polysaccharides, atractylenolide I, atractylenolide II, atractylenolide III, bis-atractylenolide, and atractylone, with the test set’s coefficient of determination (R2) values of 0.8874, 0.8313, 0.9628, 0.8406, 0.8736, 0.8532, 0.7758, and 0.8101, respectively. In conclusion, this study provides a comprehensive and rapid solution for AMR grade classification and quality evaluation, significantly improving efficiency compared with traditional methods. This strategy holds substantial promise for real-world applications, as it enables a high-throughput, non-destructive screening of AMR in settings such as post-harvest processing and market quality surveillance, thereby supporting the sustainable and intelligent development of the herbal medicine industry.

Graphical Abstract

1. Introduction

Atractylodes macrocephala Rhizoma (AMR), commonly known as Baizhu in Chinese, is the dried rhizome of the Asteraceae plant Atractylodes macrocephala Koidz and is a traditional Chinese medicine (TCM) used for treating gastrointestinal disorders. Its main chemical constituents include volatile oils, lactones, and polysaccharides, which have been proven to possess broad pharmacological effects [1,2,3,4,5]. It is mainly produced in the Zhejiang, Anhui, Henan, and Hebei provinces in China [6]. Variations in geographical growing conditions and agricultural practices across different regions may lead to inconsistencies in the quality of the medicinal material, potentially affecting its safety and efficacy [7]. Currently, the Chinese Pharmacopoeia (2020 edition) only records one quality control indicator for AMR, alcohol-soluble extractive, with a lower limit requirement. This standard is disconnected from clinical value and demand, as both high- and low-quality medicinal materials could meet the pharmacopoeial criteria, failing to reflect the advantages of high-quality materials and seriously impacting the healthy development of the AMR industry.
Volatile oil, polysaccharides, atractylenolide I, atractylenolide II, atractylenolide III and atractylone were the quality markers [8,9] in AMR, and various analytical methods are employed for the quality monitoring and evaluation of these components. Although several reports have utilized high-performance liquid chromatography (HPLC), gas chromatography (GC), and liquid chromatography–mass spectrometry (LC-MS) techniques to determine the multi-component content of AMR [10,11,12], these methods face challenges when handling large-scale testing tasks, being operationally complex, time-consuming, costly, and reliant on large volumes of organic solvents, which are environmentally unfriendly. Traditional methods for distinguishing superior and inferior medicinal materials often rely on macroscopic identification based on appearance characteristics, requiring specialized technicians in TCM authentication. However, these methods are inherently subjective and lack objective, digitized expressions. Therefore, there is an urgent need to explore a rapid and simple approach for the quantitative analysis of multiple quality markers and the classification of varying AMR quality levels (i.e., grades) to ensure the consistency of AMR quality and the sustainable development of its industry.
In recent years, the integration of sensor technologies such as electronic noses (E-nose) and electronic tongue (E-tongue) with chemometrics and machine-learning algorithms has become widespread in the field of food [13] and pharmaceutical quality control [14]. They are favored for their speed and simplicity, as they require minimal sample preparation, such as crushing, and do not involve chemical reagents. Li et al. [15] employed metabolomics, DNA barcoding, and E-nose technology to differentiate between various cultivars of Citri Reticulatae Pericarpium. The results revealed that E-nose technology was capable of rapidly and accurately distinguishing among four types of Citri Reticulatae Pericarpium. Zhou et al. [16] utilized a metal-oxide-based E-nose sensor system to analyze the odor profiles of Amomi Fructus. By employing the maximum sensor response values in conjunction with a convolutional neural network, they successfully predicted the grades of Amomi Fructus as well as the contents of bornyl acetate and camphor within the herb. Lei et al. [17] employed a potentiometric E-tongue to detect and assess the flavor profiles of genuine bear bile powder and its counterfeit products. After converting the membrane potential values obtained from the E-tongue into taste values based on the Weber–Fechner law, they effectively differentiated between authentic bear bile powder and its counterfeits by integrating multiple machine-learning algorithms. For the quality control of AMR, current studies primarily rely on traditional physicochemical analysis and sensory experience. Additionally, our preliminary research has demonstrated the ability to differentiate the geographical origins of AMR using an E-nose [18], while Wei et al. employed an E-tongue to explore the material basis of its bitter taste [19]. However, the systematic application of both E-nose and E-tongue technologies for the comprehensive quality evaluation of AMR remains unexplored.
Despite these advantages, the literature research reveals that there is no unified consensus on data-analysis methods for applying biomimetic sensor technology to the quality evaluation of TCM. Taking feature extraction from E-nose data as an example, typically, a single sensor can generate 120 data points (i.e., 120 values) over a 120 s sample information acquisition period. When considering a scenario with 10 samples and 18 sensors, the resulting data matrix dimensions are 120 × 18 × 10. Extracting meaningful information from such a data matrix for subsequent analysis is of paramount importance. In our previous research applying E-nose and E-tongue technologies, we frequently extracted the maximum value from the E-nose response curves and the steady-state value from the E-tongue response curves for analysis [20]. However, beyond these commonly used features, extracting dynamic characteristics (such as integral values, slopes, curve fitting parameters, etc.) from the sensor response curves is also an important area of research [21]. Nevertheless, some of the literature does not specify the feature extraction method employed. This raises questions about whether the maximum values derived from the E-nose truly encapsulate the holistic odor and flavor information of TCM. Are there more suitable feature-extraction techniques tailored for sensor technology data analysis? Moreover, given the vast and multidimensional nature of the data obtained through feature extraction, how can dimensionality reduction be effectively performed, and how can sensor feature values that are highly correlated with TCM quality be screened out? These issues remain areas that warrant further exploration and refinement. However, according to our investigation, there are few reports on the quality control of AMR raw herbs using E-nose and E-tongue technology.
The objectives of this study are as follows: ① to clarify the differences in chemical composition of AMR across different grades; ② to establish a sensory information database for the odor and taste profiles of AMR based on E-nose and E-tongue; ③ to develop a complete data-analysis pipeline encompassing raw signal processing, feature selection, and model construction; and ④ to build grade-classification and content-prediction models integrating E-nose, E-tongue, and machine learning, thereby achieving rapid and non-destructive evaluation of AMR quality—from “sensory grade” to “key component content.” A schematic diagram of this study is shown in Figure 1.

2. Results and Discussion

2.1. Sensory Evaluation Results of AMR

An evaluation team consisting of 10 evaluators assessed AMR samples based on sectional color, chrysanthemum pattern, number of oil spots on the section, and number per kilogram. The results showed that according to the fuzzy sensory evaluation method and evaluation criteria, AMR samples could be divided into four grades, as shown in Table 1. As a genuine regional medicinal material, Zhejiang generally scored higher and was classified as first-grade. Anhui and Henan, with similar processing methods and appearance of the medicinal material, were classified into one specification and further divided into two grades based on their scores, namely AH-HN-1 and AH-HN-2. Hebei scored lower and was classified as fourth-grade.
There is a correlation between grade and geographical origin, which aligns with established knowledge in traditional Chinese medicine. As we know, authentic medicinal materials (such as “Zhe Baizhu” from Zhejiang Province) have historically been recognized as superior in quality due to their favorable genetic characteristics and growing environment. Our research findings (Table 1), as well as previous studies, both corroborate this traditional understanding. However, this correlation is not absolute or one-to-one. In the sample set, not all samples from the same origin were classified into the same grade. For example, samples from Anhui Province were assigned to the second and third grades based on their specific sensory scores, and the same applied to samples from Henan Province. Therefore, the grading of the samples is not merely labeling based on origin but is determined by the individual quality performance of each sample.

2.2. Results of Quality Marker Content of AMR

The results of extract content are shown in Figure 2A. There were significant differences in extract content between HB and AH-HN1, but no significant differences were found in extract content among other grades of AMR. The average extract content in samples of different grades was ranked from high to low as follows: HB (48.48%) > ZJ (45.10%) > AH-HN-2 (44.95%) > AH-HN-1 (44.22%).
The results of volatile oil content are shown in Figure 2B. ZJ had the highest essential oil content and showed significant differences with the other three grades. AH-HN-1 and AH-HN-2 ranked second and showed significant differences with HB in essential oil content. The average essential oil content in samples of different grades was ranked from high to low as follows: ZJ (1.471%) > AH-HN-1 (1.225%) > AH-HN-2 (1.172%) > HB (0.9846%).
The relative standard deviation (RSD) values for the precision, repeatability, and stability of the polysaccharide testing method were all below 3.00%, indicating good instrument precision, reproducibility of analytical conditions, and stability of the sample test solution within 240 min. The glucose standard curve was y = 4.8715x + 0.0752, with R2 = 0.9997 and an average recovery rate of 103.77%, as well as an RSD of 1.655%, indicating good method accuracy. The results of polysaccharide content are shown in Figure 2C. AH-HN-1 had the highest polysaccharide content and showed significant differences with ZJ and HB samples. The average total polysaccharide content in samples of different grades was ranked from high to low as follows: AH-HN-1 (456.6 mg/g) > AH-HN-2 (397.7 mg/g) > ZJ (395.5 mg/g) > HB (370.3 mg/g).
The RSD values for the precision, repeatability, and stability of the HPLC analysis procedure were all below 3.00%, indicating good instrument precision, reproducibility of analytical conditions, and stability of the sample solution within 48 h. The R2 value of the standard calibration curve for compound quantification based on peak area concentration is no less than 0.9995. The sample recovery ranged from 102.7% to 105.9%, indicating satisfactory quantitative accuracy of the analysis process. The results of the lactone components and atractylenolide content are shown in Figure 2D–H. The ZJ sample had the highest content of atractylenolide I, II, and III, showing significant differences with other grade samples. There were no significant differences in the content of bis-atractylenolide among all grade samples. The content of atractylone in ZJ, AH-HN-1, and AH-HN-2 was higher, showing significant differences with HB samples. The content ranges of atractylenolide I, II, III, bis-atractylenolide, and atractylone in different grades of AMR are shown in Table S1.
The above results indicate that the contents of extracts, polysaccharides, bis-atractylenolide, and atractylone are not strongly correlated with grade and could be used as a qualified indicator for the inspection of AMR but not as a basis for grade classification. The reason may be that extract and polysaccharides are comprehensive indicators. The extract reflects the total amount of alcohol-soluble substances, while polysaccharides represent the sum of a broad category of compounds. Their responses to growth environments (e.g., fertilization, water availability) may be more direct and universal, allowing them to accumulate to relatively high levels under different origins and cultivation conditions. As a result, their ability to differentiate between finer grades is limited. In contrast, The ZJ samples had the highest content of volatile oil and atractylenolide I, II, and III (which have been confirmed as the primary active components in AMR [1,3,4,5]), and showed significant differences with other grade samples, which is consistent with the characteristics of genuine high-quality medicinal materials, and could be used as an indicator to distinguish ZJ from other grade samples. This study clarified the distribution pattern of quality markers of AMR in different grade samples, providing data support for the subsequent establishment of a rapid quality prediction model based on E-nose and E-tongue.

2.3. Analysis of Sensor Response Values of E-Nose and E-Tongue

As shown in Figure S1A, the response values of the E-nose sensors displayed a clear trend of first increasing and then decreasing. The values increased rapidly from 0 s to 20 s, peaked within 20–30 s, and then decreased gradually afterward. Differently, the E-tongue sensors (Figure S1B) showed relatively high initial response values, which tended to stabilize after 30 s. Box plots of typical characteristics for the E-nose S13 sensor and E-tongue PKS sensor in this study visually revealed significant feature discrepancies among different samples (Figure 3 and Figure S2). Results demonstrated that a series of parameters presented significant differences among samples: For the E-nose S13 sensor, these parameters included maximum response value, integral area at maximum response, maximum and minimum values of the first derivative, response time corresponding to the maximum value, and time corresponding to the maximum first derivative. For the E-tongue PKS sensor, the distinguishable parameters consisted of response value at 120 s, steady-state response, average response within 120 s, integral area of the whole response period, and average value of the first derivative. These results provide a solid foundation for establishing accurate quality evaluation models for different grades of AMR using E-nose and E-tongue technologies.

2.4. Feature Extraction of E-Nose and E-Tongue

The classification model for AMR grades and the regression quantitative model for quality markers based on the E-nose showed that the prediction performance of the three feature values corresponding to the time of specific response values was the worst (Table 2 and Table 3). A total of 72 feature values from four groups exhibited favorable predictive performances, including the maximum value, the steady-state response value, the average value of the response curve exceeding 120 s, and the average first-order differential value of the response curve. The classification and regression prediction models (Table 2 and Table 3) based on the optimal feature set have significantly improved performance. The classification model has the highest accuracy in the test set, and the regression model has the highest R2 value in the test set, which verifies the rationality and effectiveness of the selected feature set.
The classification model for AMR grades and the regression quantitative model for quality markers based on the E-tongue showed that the average value of the first derivative exhibited the poorest predictive performance (Table 4 and Table 5). By combining 28 feature values that demonstrated relatively good predictive performances, including the average value over the last 10 s, the value at the 120th second, the average value over 120 s, and integral values into an optimal set, further classification and regression prediction models were established (Table 4 and Table 5). The results indicated that, despite the superior performance of the optimal feature set (with 28 features), the 120 s average value, comprising only 7 features by comparison, also demonstrated a relatively good predictive performance. To reduce computational time and data dimensionality, the 120 s average value was selected as the feature value for the subsequent feature selection process.
This study conducted a comprehensive extraction of sensor feature values when using E-noses and E-tongues for the quality evaluation of traditional Chinese medicine (TCM) for the first time. The results showed that, for both E-tongue and E-nose sensors, the commonly used steady-state values from the last 10 s and the maximum response values did not exhibit optimal performances in classification and regression models. The feature values corresponding to specific sensor response times and the differential values of sensor curves demonstrated poor predictive performances, making them unsuitable for feature-extraction methods for TCM, which has complex odor and taste profiles. Additionally, it was found that more feature value information does not necessarily lead to better results; the inclusion of a large amount of redundant information and noise in the feature values can actually degrade model performance, which is consistent with the “dimensional catastrophe” theory [22].

2.5. Feature Selection of E-Nose and E-Tongue

2.5.1. Feature Selection of E-Nose

The results of feature selection for the grade-classification model based on the E-nose (Figure S3) indicate that when the number of features is 13, the accuracy of the test set is close to that achieved using the optimal set. Among the three feature-selection methods, MI and RF exhibit better classification performances, with test set accuracies of 85.11% and 85.78%, respectively. In contrast, SVM-RFE demonstrates a poorer classification performance, with a test set accuracy of 69.33%. The results of feature selection for the compound regression model of the E-nose (Figure S3) indicate that when the number of features ranges from 12 to 15, the R2 value on the test set is close to that achieved using the optimal set. The feature selection results for both the classification model and the regression model are presented in Table S2. By counting the frequency of occurrence of each selected feature in the classification and regression models across the three feature-selection methods, features with a frequency of occurrence greater than or equal to 3 were screened. MI selected 15 features, namely S8SS, S8D-av, S10D-av, S10SS, S10AV, S13Max, S15Max, S17Max, S17AV, S17SS, S17D-av, S18Max, S18AV, S18SS, and S18D-av, involving six sensors. SVM-RFE selected 16 features, namely S6Max, S7AV, S7SS, S7D-av, S9D-av, S10Max, S12SS, S13D-av, S14Max, S14D-av, S17SS, S17AV, S17D-av, S18AV, S18SS, and S18D-av, involving nine sensors. RF selected 14 features, namely S1Max, S8SS, S8D-av, S10Max, S10SS, S10D-av, S17Max, S17SS, S17AV, S17D-av, S18Max, S18AV, S18SS, and S18D-av, involving five sensors.
Based on the features selected by the aforementioned feature-selection methods, classification and regression models were once again established using BPNN (Figure S4). The results indicated that the features selected by MI generally performed better in both classification and regression models. A total of 15 features were selected using the MI method, involving six sensors (S8, S10, S13, S15, S17, S18) that corresponded to sensitive materials detecting hydrocarbons, methane, fluorine, polar compounds, alcohol and chlorinated compounds, respectively. While reducing the data dimensionality by 79.17% (from 72 to 15), the accuracy of the classification model decreased slightly from 84.67% to 83.78%, retaining 98.95% of the original information. The R2 of the test set for the extract content regression model increased from 0.5461 to 0.5927, a growth of 8.53%. For the volatile oil content regression model, the R2 of the test set increased from 0.4730 to 0.4862, a growth of 2.79%. For the polysaccharide content regression model, the R2 of the test set increased from 0.5144 to 0.5198, a growth of 1.05%. For the atractylenolide I content regression model, the R2 of the test set increased from 0.6352 to 0.6554, a growth of 3.18%. For the atractylenolide II content regression model, the R2 of the test set decreased from 0.5346 to 0.5136, retaining 96.07% of the original information. For the atractylenolide III content regression model, the R2 of the test set decreased from 0.5260 to 0.5125, retaining 97.43% of the original information. For the bis-atractylenolide content regression model, the R2 of the test set decreased from 0.4353 to 0.3874, retaining 89.00% of the original information. For the atractylon content regression model, the R2 of the test set decreased from 0.5067 to 0.4783, retaining 94.40% of the original information.

2.5.2. Feature Selection of E-Tongue

The results of feature selection for the grade-classification model based on E-tongue (Figure S5) revealed that when the number of features was 2, the accuracy of the test set for MI nearly matched that of the initial feature set, reaching 90.44%. For SVM-RFE and RF, the accuracy of the test set approached that of the initial feature set when the number of features was 4, achieving 91.56% and 92.44%, respectively. The results of feature selection for the compounds regression model based on the E-tongue (Figure S5) indicated that when the number of features ranged from 2 to 5, the R2 of the test set for the three feature-selection methods was close to that of the initial feature set. The results of feature selection for both the classification and regression models are presented in Table S3. By statistically analyzing the frequency of each selected feature in the classification and regression models across the three feature-selection methods, the top four features with the highest selection frequencies were identified. Specifically, the features selected by MI were ANS, NMS, PKS, and SCS; those selected by SVM-RFE were PKS, CTS, SCS, and ANS; and those selected by RF were PKS, ANS, CTS, and NMS.
Based on the features selected by the aforementioned feature-selection methods, classification and regression models were re-established using BPNN (Figure S6). The results indicated that the features selected by SVM-RFE generally performed better in classification and regression models. The E-tongue feature-selection results showed that a total of four features were selected using the SVM-RFE method, including PKS, CTS, SCS, and ANS, corresponding to four sensors that respectively detected general, salty, bitter, and sweet tastes. While reducing the data dimension by 42.86% (from 7 to 4), the accuracy of the classification model decreased from 94.67% to 93.56%, retaining 98.83% of the original information. For the regression models: the R2 of the test set for the extract content regression model increased from 0.6878 to 0.6979, a growth of 1.47%; the R2 of the test set for the volatile oil content regression model decreased from 0.7151 to 0.6765, retaining 94.60% of the original information; the R2 of the test set for the polysaccharide content regression model increased from 0.6304 to 0.7099, a growth of 12.61%; the R2 of the test set for the atractylenolide I content regression model decreased from 0.7533 to 0.6227, retaining 96.81% of the original information; the R2 of the test set for the atractylenolide II content regression model decreased from 0.7488 to 0.5136, retaining 82.66% of the original information; the R2 of the test set for the atractylenolide III content regression model decreased from 0.6266 to 0.5798, retaining 92.53% of the original information; the R2 of the test set for the bis-atractylolide content regression model increased from 0.5386 to 0.5887, a growth of 9.30%; and the R2 of the test set for the atractylone content regression model decreased from 0.6086 to 0.6065, retaining 99.65% of the original information.
The feature-selection results of the E-nose and E-tongue showed that while feature dimension was significantly reduced, the selected low-dimensional features could still retain most of the information from the original feature set. In some content regression models, the prediction performance is improved compared with using the original feature set, as redundant data and noise interference were eliminated.
Among the three feature-selection methods, the MI based on the filter approach and the SVM-RFE based on the wrapper approach yield better feature selection results. The RF feature-selection method based on the embedded approach may lead the embedded method to split nodes along non-critical features due to the model’s lack of bias, thereby affecting the model’s decision-making process [23]. By comparison, the MI and SVM-RFE methods are more suitable for feature selection in this study.

2.6. Establishment of Classification and Regression Models Based on Machine-Learning

Based on the features selected in the previous step (Table 6), classification models for differentiating AMR of various grades and regression quantitative models for eight components were established using nine machine-learning methods (Tables S4 and S5). Among the classification and regression models based on the E-nose data, the KNN, LSSVM, and PSO-SVM algorithms demonstrated the best predictive performances, while the RBF algorithm exhibited a relatively poor predictive performance. The reason may be that the E-nose data encompass 15 features, while the RBF algorithm is highly sensitive to the distribution of input data. The presence of noise in the dataset could alter the distribution of the input data, thereby affecting the performance of the model [24]. Among the classification and regression models based on the E-tongue, the KNN algorithm and the PSO-SVM algorithm demonstrate the best predictive performances, followed by the RF algorithm, while the PLS algorithm exhibits the worst predictive performance. The possible reason is that the relationship between the input data and the output data is not a simple linear one, and there exist nonlinear relationships in the data that the PLS algorithm fails to effectively capture. The BPNN algorithm optimized by the PSO algorithm does not improve the model’s predictive performance to a certain extent. The possible reason is that the current task may not be suitable for the BPNN model. Each algorithm possesses distinct characteristics in terms of generalization ability, computational complexity, sensitivity to outliers, and other aspects, as well as exhibits varying performances according to different evaluation metrics. Therefore, it is necessary to select an appropriate algorithm based on specific conditions such as the target detection object, sensor array, and application site.
Whether in E-nose models or E-tongue models, the KNN algorithm demonstrates good performances. It is a straightforward and effective non-parametric method for classification and regression. The core idea is to calculate the distances between samples, identify the k-nearest neighbors to the target sample, and then predict the class or value of the target sample based on the labels or values of these neighbors. Its advantages lie in its simple algorithm principle, high accuracy, and insensitivity to outliers, while its disadvantages include high computational complexity and the need to select an appropriate k value [25,26]. In this study, an optimal k value of 2 was determined through trial-and-error before model establishment.
Table 7 summarizes the accuracy, weighted F1 score, precision and recall for each category of the KNN classification models. Precision refers to the proportion of samples predicted to belong to a certain category that actually belong to that category, while recall indicates the proportion of samples that actually belong to a certain category and are correctly predicted as such. The F1 score is the harmonic mean of these two metrics. When dealing with imbalanced classes, the weighted F1 score is often used to evaluate classification models, as it can more accurately reflect the true performance of the model [27]. The data indicate that E-tongue technology provides good classification rates, with both the weighted F1 score and accuracy exceeding 95.0% for both the training and test sets. To further elucidate the model’s classification performance, we conducted additional multivariate analyses. A confusion matrix for the KNN classification on the test set is provided (Figure S7). This matrix details the model’s performance for each specific category and quantitatively illustrates the relationship between predicted and actual class distributions. In particular, no misclassification examples were observed for ZJ and HB samples, demonstrating the superior reliability of the E-tongue-based classification model. In addition, to visually demonstrate the performances of the E-nose and E-tongue in classifying different AMR grades, canonical discriminant analysis was applied for visualization. The resulting CDA score plot (Figure S8) clearly shows the clustering of samples from different grades, with the least overlap observed between clusters from the E-tongue.
Figure 4 and Figure 5 demonstrate the performance of the KNN regression models for the E-nose and E-tongue. The R2tr (training set R2) values are close to the R2te (testing set R2) values, indicating that the established models do not suffer from overfitting [28]. The correlation plots for the quantitative models of the eight component contents in AMR reveal that the data points are primarily clustered around the diagonal, suggesting a high degree of correlation between the predicted values and the actual values. In the component regression models based on the E-nose, except for bis-atractylenolide, the R2te values all exceed 0.5. For the component regression models based on the E-tongue, except for bis-atractylenolide, the R2te values all exceed 0.8.
Volatile oils and atractylenolide I, II, and III, the key active substances in AMR, are significantly enriched in authentic medicinal materials (ZJ samples) and show significant differences compared with other grades. The accumulation of these components constitutes the core chemical markers for distinguishing grades, especially for identifying authentic medicinal materials. Based on the regression models combining E-nose and chemical composition, the R2 values for atractylenolide I, II, and III all exceed 0.7 (Table S5), suggesting that they are the primary chemical substances eliciting the characteristic response profiles of the E-nose. Overall, the E-tongue demonstrates a superior predictive performance, with most R2 values exceeding 0.8. The four sensors selected by the E-tongue—PKS, CTS, SCS, and ANS—represent general-purpose, salty, bitter, and sweet taste sensors, respectively, indicating their sensitivity to active components in the solution. The response of the bitter taste sensor is likely correlated with the content of bitter terpenoid lactones (such as atractylenolide I, II and III), while the response of the sweet taste sensor may be associated with the content of polysaccharides and other water-soluble flavor compounds. However, these inferences are based solely on observational data. Rigorous causal validation requires controlled experiments using standard substances such as pure atractylenolide I or atractylone to directly observe sensor responses. This study provides key markers and correlative evidence for designing such validation experiments.

2.7. Results of Data Fusion

Data-level fusion: The results of data-level fusion based on grade classification (Figure 6A) showed that the test set accuracies of the fused data and the E-tongue data were 97.78% and 95.34%, respectively, with no significant difference. The results of data-level fusion based on content prediction (Figure 6B–I) indicated that the test set R2 of the E-tongue data alone showed no significant difference compared with the test set R2 of the data-level fusion.
Feature-level fusion: Principal component analysis was used to extract three principal components (cumulative contribution rate: 98.04%) from the E-nose feature set and four principal components (cumulative contribution rate: 100%) from the E-tongue feature set. Therefore, for feature-level fusion, the feature layers of the E-nose (three principal components) and the E-tongue (four principal components) were selected for analysis. The results of feature-level fusion based on grade classification (Figure 6A) showed that the test set accuracies for the fused features and the E-tongue features were 98.89% and 95.11%, respectively, with no significant difference. The results of feature-level fusion based on content prediction (Figure 6B–I) indicated that, except for atractylenolide II, the test set R2 of the electronic tongue feature data alone showed no significant difference compared with the test set R2 of the fused features.
It can be seen that both data-level fusion and feature-level fusion results indicate that electronic tongue data contributes more significantly to the performance of classification and content prediction models. When comparing the same evaluation metrics, feature-level fusion outperforms data-level fusion, which is consistent with findings in the literature [29]. Comparatively, feature-level fusion retains the characteristic information of the original data while reducing the complexity of information processing, thereby enhancing model performance [30].
This study is the first to combine electronic senses with machine learning, systematically achieving both grade classification of AMR and quantitative prediction of eight key active components, filling the gap in rapid, non-destructive detection methods in this field. Currently, there is limited research on extracting and screening feature values for traditional Chinese medicine quality evaluation. In the future, portable and miniaturized devices could be developed based on the screened sensors and feature values to promote on-site testing.
Although the sample size in the current study is representative, constructing a more universally applicable model in the future will require collecting samples from a wider range of origins, more cultivation years, and those influenced by other factors (such as processing methods). While we have established correlations between sensor signals and chemical components through association analysis, this relationship remains largely statistical. The precise mechanism of interaction between sensors and specific molecules requires further in-depth fundamental research to be elucidated.
It should be noted that due to the limited dataset size, k-fold cross-validation was not employed in this study, although this method is helpful for obtaining more robust performance estimates. In future research, we will further expand the dataset and supplement the analysis with k-fold cross-validation, while fully reporting uncertainty measures (e.g., standard deviation), to further validate the advantages of the proposed method.

3. Materials and Methods

3.1. Chemicals, Reagents and Samples

Chemical reference standards of D-glucose anhydrous, atractylenolide I, atractylenolide II, and atractylenolide III were obtained from the National Institutes for Food and Drug Control (Beijing, China). Chemical reference standards of atractylone and bis-atractylenolide were obtained from Shanghai yuanye Bio-Technology Co., Ltd. (Shanghai, China). HPLC grade acetonitrile was purchased from Thermo Fisher Scientific Co., Ltd. (Shanghai, China). Analytical pure phenol and ethanol absolute was purchased from Fuchen Chemical Reagent Co., Ltd. (Tianjin, China). Analytical pure sulfuric acid was from Beijing Tongguang Chemical Co. (Beijing, China). Purified water was purchased from Hangzhou Wahaha Group Co., Ltd. (Hangzhou, China).
All 180 batches of AMR samples were collected from the Zhejiang, Anhui, Henan and Hebei provinces in China, and the information is listed in the Table S6. Their authenticity was authenticated as correct by Professor Yonghong Yan from Department of Chinese Materia Medica of Beijing University of Chinese Medicine according to their morphological features. All batches of AMR samples were pulverized and sieved through a 50 mesh sieve and then stored properly in a refrigerator at −20 °C for further analysis.

3.2. Sensory Evaluation of AMR

A fuzzy mathematics method was used to establish sensory evaluation standards for the sectional color, chrysanthemum pattern, number of oil spots and number per kilogram of AMR. The different evaluation criteria were divided into four grades, and the binary comparison determination method was used to determine the proportion of each criterion in the comprehensive evaluation [31]. In each comparison, the more important factor receives 1 point, while the less important factor receives 0 points. The weight of each of the four indicators is calculated as the ratio of its total score to the sum of all scores. A total of 10 previously trained professionals were invited to comprehensively evaluate the AMR grades based on the defined descriptive terms for trait characteristics and the scoring criteria. The order of sample presentation followed a randomized complete block design, with each sample assigned a 3-digit random code. After reviewing all AMR samples, the sensory evaluation panel members evaluated each sample individually according to the trait identification scoring sheet. The trait scoring criteria are detailed in Table S7. The number of evaluators for each criterion was counted, and the membership degree matrix as well as the sensory evaluation scores were calculated. The grades were then classified based on the score values.

3.3. Determination of Quality Markers Content of AMR

3.3.1. Determination of Extract Content

A total of 2.0 g of AMR powder was accurately weighed into a 100 mL Erlenmeyer flask. Then, 100 mL of 60% ethanol was added, and the total weight was recorded. After standing for 1 h, a reflux condenser was connected, and the mixture was heated to boiling while maintaining a gentle simmer for 1 h. After cooling, the Erlenmeyer flask was removed and reweighed. The weight loss was compensated by adding 60% ethanol. The mixture was then filtered through a dried filter. Subsequently, 25 mL of the filtrate was measured and transferred to a pre-dried and constant-weight evaporating dish. The filtrate was evaporated to dryness on a water bath, followed by drying at 105 °C for 3 h. Finally, the dish was placed in a desiccator to cool for 30 min before being quickly and accurately weighed. Each sample was measured three times, and the average value was taken.

3.3.2. Determining of Volatile Oil Content

In total, 100 g of AMR powder was accurately weighed into a round-bottom flask. Then, 500 mL of water was added, and the volatile oil determinator and reflux condenser were connected. Water was added from the upper end of the condenser until the graduated portion of the volatile oil determinator was filled and overflowed into the flask. The flask was placed in an electric heating jacket and heated slowly until boiling, maintaining a gentle simmer for about 5 h. After heating was stopped and the flask was allowed to stand for a moment, the piston at the lower end of the determinator was opened to slowly release the water until the upper end of the oil layer reached 5 mm above the 0 mark on the scale. After standing for more than 1 h, the piston was opened again to lower the oil layer until its upper end was exactly level with the 0 mark on the scale. The volume of the volatile oil was then read, and the content of volatile oil in the test sample was calculated as a percentage. Each sample was measured three times, and the average value was taken.

3.3.3. Determination of Polysaccharides Content

The 5% phenol solution: A total of 5.0 g of phenol was accurately weighed and an appropriate amount of water was added to dissolve it. The solution was then transferred to a 100 mL brown volumetric flask and diluted to the mark with water.
Reference solution: A total of 10.65 mg of anhydrous glucose reference standard was accurately weighed into a 100 mL volumetric flask and water was added to dissolve it.
Preparation of sample solution: A total of 0.5 g of AMR powder was accurately weighed into a 250 mL round-bottom flask. Then, 100 mL of 80% ethanol was added, and the mixture was heated to reflux at 85 °C for 1.5 h before being filtered while still hot. The filter residue along with the filter paper was placed back into the flask, and 150 mL of water was added. The mixture was heated to reflux at 90 °C for 3 h before being filtered again while still hot. The filter was washed with a small amount of hot water, and the filtrate and washings were combined and allowed to cool. The combined solution was transferred to a 250 mL volumetric flask, diluted to the mark with water, and mixed well. Finally, 1 mL of the solution was accurately transferred to a 25 mL volumetric flask, diluted to the mark with water, and mixed well to obtain the sample solution.
Preparation of standard curve: First, 20, 50, 80, 110, 140, 170, and 200 μL of the reference solution were accurately pipetted into 2 mL centrifuge tubes, and water was added to bring the volume to 200 μL. Then, 200 μL of 5% phenol solution and 700 μL of concentrated sulfuric acid were accurately added to each tube, and the mixtures were rapidly mixed. The tubes were placed in a water bath at 80 °C for 10 min, removed, and rapidly cooled in an ice bath for 5 min to stop the reaction. Then, 150 μL of each reaction mixture was accurately pipetted into a 96-well microplate, and the absorbance was measured at a wavelength of 490 nm. Finally, a standard curve was plotted with absorbance on the y-axis and glucose concentration on the x-axis.
Determination of samples: A total of 200 μL of the sample solution was accurately pipetted into a 2 mL centrifuge tube. Then, following the “preparation of standard curve “method, starting from the step of “adding 200 μL of 5% phenol solution to each tube”, the absorbance was measured as per the method. The glucose concentration (mg/mL) in the test solution was calculated from the standard curve. Each sample was measured three times, and the average value was taken.
The precision, repeatability, stability, and recovery rate of the polysaccharide determination method were tested by analyzing six samples. To check the precision, the same sample solution was measured six times consecutively. Additionally, the repeatability was evaluated by independently preparing six identical samples. The stability of the test solution was assessed by storing it at room temperature for six time points (0, 30, 60, 120, 180, and 240 min). Finally, the recovery rate was determined by using a sample with a known content and adding an equal amount of the standard.

3.3.4. Determination of Lactones and Atractylone

Reference solutions: Atractylenolide I, atractylenolide II, atractylenolide III, bis-atractylenolide, and atractylone standards were accurately weighed and dissolved in methanol to obtain stock solutions with concentrations of 0.522, 0.528, 0.574, 0.882, and 1.322 mg/mL, respectively.
Sample solution: A total of 1.0 g of AMR powder was accurately weighed into a 100 mL conical flask. Then, 20 mL of methanol was added and the flask was sealed. The mixture was subjected to ultrasonic treatment (250 W, 40 kHz) for 30 min. After cooling to room temperature, the solution was mixed well and filtered through a microporous membrane (0.45 μm). The subsequent filtrate was collected as the sample solution.
The HPLC analyses in this study were conducted using the DGU403-HPLC system equipped with an ultraviolet detector (Shimadzu, Tokyo, Japan). The HPLC test conditions were as follows: a Kromasil 100-5C18 column (250 × 4.6 mm, 5 μm) was used; mobile phase A was water; mobile phase B was acetonitrile; an elution gradient was used (0–10 min, 32% A;10–15 min, 32% A–20% A; 15–20 min, 20% A–15% A; 20–25 min, 15% A–10% A; and 25–40 min, 10% A); the detection wavelength was set at 220 nm for atractylenolide II, atractylenolide III, bis-atractylenolide, and atractylone, and at 276 nm for atractylenolide I. The volume flow, column temperature, and injection volume were set to 1.0 mL/min, 30 °C, and 10 μL, respectively. Each sample was measured three times, and the average value was taken.
The precision, repeatability, stability, and recovery rate of the HPLC analysis method were tested by analyzing six samples. The linearity of the chromatographic procedure was determined using the standard solution. To check the precision, the same sample solution was measured six times consecutively. Additionally, the repeatability was evaluated by independently preparing six identical samples. The stability of the test solution was assessed by storing it at room temperature for six time points (0, 30, 60, 120, 180, and 240 min). Finally, the recovery rate was determined by using a sample with a known content and adding an equal amount of the standard.

3.4. Odor Information Collection

An α-Fox4000 E-nose (Alpha MOS, Co., Ltd., Toulouse, France) with 18 metal-oxide gas sensors was employed to obtain the odor information of AMR, the names and response characteristics of each sensor are presented in Table S8. The E-nose was self-checked, and the sensor array was preheated for 2–3 h before each sampling experiment. Processed pure air was used as carrier gas to clean the sensor array, returning the signal response back to baseline. In the experiment, the environmental temperature for E-nose analysis was controlled at 25 °C ± 2 °C. To ensure the reliability of the signals under these conditions, we conducted a repeatability validation of the analysis process. The RSD value for six repeated measurements was below 3%, indicating good repeatability. AMR powder (0.4 g) was sealed into a 20 mL vial and incubated for 10 min at 45 °C (250 rpm). The temperature and volume of injection were set at 45 °C and 1500 μL. The flow rate of carrier gas was 150 mL/min. The data acquisition interval and data acquisition cycle were 1 s and 120 s, respectively. The fixed sampling rate for data acquisition was 1.0 Hz (i.e., one data point recorded per second). This sampling rate is sufficient to capture the relatively stable response signals of the olfactory sensors and aligns with the standard operating protocol for this type of electronic nose. Every AMR sample was continuously sampled 3 times, and the average response value (feature values for each feature-extraction method) was calculated and used for further preprocessing.

3.5. Taste Information Collection

ASTREE II E-tongue (Alpha MOS, Co., Ltd., Toulouse, France) with 7 sensors was employed to obtain the taste information of AMR; the names and response characteristics of each sensor are presented in Table S9. A total of 1.0 g of AMR powder was accurately weighed and placed in a 100 mL conical flask. Then, 50 mL of water was added, and the mixture was subjected to ultrasonic extraction for 30 min (250 W, 40 kHz). It was then centrifuged at 5000 r/min for 10 min, and the supernatant was collected. The supernatant was passed through a 0.45 μm microporous filter membrane, and the subsequent filtrate was collected as the test solution. The sensors were pre-equilibrated and calibrated with 0.01 mol/L hydrochloric acid. After successful pre-equilibration and calibration, testing can be performed. For each sample tested, a cleaning step was set, with a cleaning time of 10 s, and the sample acquisition time was 120 s. The fixed sampling rate for data acquisition was 1.0 Hz (i.e., one data point recorded per second). This sampling rate is sufficient to capture the relatively stable response signals of the taste sensors and aligns with the standard operating protocol for this model of electronic tongue. In the experiment, the environmental temperature for E-tongue analysis was controlled at 25 °C ± 2 °C. To ensure the reliability of the signals under these conditions, we validated the precision, repeatability, and stability of the electronic tongue analysis process. The precision test results showed that after excluding the first four measurements, the RSD of the sensor response values at each time point was less than 3%. Therefore, in subsequent experiments, each sample was measured 9 times, with the first 4 measurements discarded, and the average value of the last 5 measurements was used for data analysis. The RSD value for six repeated measurements was below 3%, indicating good repeatability. The stability test results demonstrated that the RSD of the sensor response values at 120 s over a 12 h period was less than 3%, confirming that the samples remained stable within 12 h.

3.6. Feature Extraction and Feature Selection for Odor and Taste Information

3.6.1. Feature Extraction

In the data analysis of E-nose and E-tongue, raw sensor response signals often contain noise, baseline drifts, and nonlinear characteristics. Directly modeling these signals may lead to degraded model performance. However, through appropriate preprocessing and feature extraction, the accuracy and robustness of gas and taste classification/identification can be significantly improved. To comprehensively investigate the impact of different feature-extraction methods on model performance, we searched the literature [21,32,33] and employed various techniques to preprocess the raw sensor response curves. The feature-extraction methods for the E-nose and E-tongue are summarized in Table 8. These features reflect different aspects of reaction kinetics from various perspectives. The steady-state value represents the final steady-state characteristic of the entire dynamic response process, while the maximum value reflects the maximum degree of change in the sensor’s response to odor. The first derivative can indicate the sensor’s reaction rate. The purpose of these feature-extraction methods is to simplify the data while preserving as much useful information as possible.
BPNN is widely used to handle the nonlinear and high-dimensional data in electronic nose and tongue systems. As a baseline model, it effectively evaluates the performances of different feature-extraction methods. Through feature extraction of the response values from the E-nose and E-tongue, back-propagation neural network (BPNN) classification models for different grades of AMR and regression quantitative models for eight components were established. The performances of various feature-extraction methods was compared, and an optimal feature set was constructed by combining the feature values with superior performances.

3.6.2. Feature Selection

Undoubtedly, the complex and redundant information could cause interferences in recognition. Therefore, it is necessary to perform dimensionality reduction on the data to reduce the number of features while retaining as much valid information as possible. Currently, there are mainly two categories of feature dimensionality reduction methods: feature selection and feature extraction. Feature selection involves selecting a subset of the most informative features from the original feature set, a process that involves evaluating the importance of each feature for the prediction target and choosing the most significant ones. Common feature selection methods include filter methods, wrapper methods, and embedded methods, among others. Feature extraction, on the other hand, maps high-dimensional feature vectors into a lower-dimensional space through some mathematical transformation, thereby constructing a new and typically smaller set of low-dimensional features, which may be linear or nonlinear combinations of the original features. Common feature-extraction methods include principal component analysis, linear discriminant analysis, etc. [34,35].
The purpose of this study is to provide references for the application of bionic technology in the precise quality evaluation of TCM and to lay the foundation for the development of portable detection devices for precise TCM quality evaluation, with particular focus on reducing the sensor array. Therefore, this section employs feature-selection methods to reduce feature dimensionality. One method was selected from each of the three categories of feature selection methods: filter methods, wrapper methods, and embedded methods [36]. Specifically, the mutual information (MI), support vector machine-based recursive feature elimination (SVM-RFE), and random forest (RF) [37] were chosen. These methods were combined with the BPNN algorithm to compare the performances of different feature-selection methods, identify the most suitable one for AMR quality evaluation, and select more effective features to form the optimal feature set.

3.7. Machine-Learning Classification and Regression Models

Commonly used machine-learning methods can be roughly divided into three categories based on their principles: statistical theory-based pattern recognition, neural network-based pattern recognition, and ensemble algorithms. In this section, we selected widely used and representative algorithms from each category to perform classification and regression modeling after feature selection. From the statistical theory-based pattern recognition methods, the partial least squares (PLS) algorithm, the k-nearest neighbors (KNN) algorithm, and the support vector machine (libSVM) were selected. To address the issue of SVM being significantly affected by its parameter selection, a method based on particle swarm optimization (PSO) was proposed to optimize the feature classification and regression methods of SVM, and a PSO-SVM model was constructed to improve prediction performance. Additionally, the least squares support vector machine (LSSVM) algorithm based on different kernel functions was applied for classification and regression modeling [38]. From the neural-network-based pattern recognition methods, the BPNN and the radical basis function (RBF) were selected and from the ensemble algorithms, the random forest (RF) algorithm was chosen to establish classification and regression models, respectively. BPNNs have strong approximation and generalization capabilities for nonlinear and large-scale systems [39]; the RBF network is a special artificial neural network that excels in local approximation [40]; RF is an ensemble classifier based on decision trees and has the advantages of simple parameter settings and the ability to effectively capture nonlinear relationships between variables [41]; and all of these algorithms could be used to solve problems such as classification, regression, and pattern recognition.
To avoid bias from relying on a single method, we employed three feature-selection approaches for screening. MI was implemented by core computational modules such as MI select. SVM-RFE utilized a linear kernel SVM as the base classifier to recursively eliminate the least contributive features. RF quantified feature importance based on the average impurity reduction brought by each feature when splitting nodes in decision trees. By calculating the frequency of each selected feature across the three feature-selection methods in both classification and regression models, features with a frequency greater than or equal to three were screened out for the comparative analysis of the three feature-selection techniques.

3.8. Data Fusion

Multi-source information fusion could be categorized into three types based on the level of fusion: data-level fusion, feature-level fusion, and decision-level fusion [42]. According to the characteristics of data fusion, this study adopts two fusion approaches—data-level fusion and feature-level fusion—to perform an integrated analysis of the features extracted from the E-nose and E-tongue. The objective is to explore whether data fusion could enhance the performance of classification models and content prediction models.
Principal component analysis (PCA) was applied separately to the feature sets selected from the E-nose and E-tongue to extract features for feature-level fusion. To retain the maximum amount of information from the original data, principal components with a cumulative variance contribution rate of 95.0% or higher were extracted. The pattern-recognition method employed was the optimal algorithm selected in the previous section. Classification models for grade and prediction models for content were established separately for the electronic nose, the electronic tongue, and their fusion analysis. Each model was predicted ten times.

3.9. Data Analysis

The signal acquisition duration for both the electronic nose and electronic tongue was 120 s. All features were extracted based on the sensor response curves within this time period (the feature-extraction method is described in Section 3.6.1). To eliminate the influence of differences in sensor magnitude and scale on the model, Z-score normalization was applied to all extracted raw feature values.
During feature extraction, feature selection, and machine-learning classification and regression modeling, the sample dataset was consistently randomly divided into two fixed subsets: a training set comprising 75% of the total samples and an independent test set comprising 25% of the total samples.
The optimal sensor feature-extraction and feature-selection methods were determined by testing with a BPNN 10 times. The classification model used accuracy as the evaluation criterion, while the regression model used the coefficient of determination (R2) as the evaluation criterion.
The development of classification and regression models was implemented using the MATLAB 2022b software (Natick, MA, USA), while data analysis and visualization were performed with the GraphPad Prism 9.4.0 software (San Diego, CA, USA).

4. Conclusions

Rapid determination of the content of multiple active components in AMR and grade classification are critical issues to be addressed in the quality control of AMR. On this basis, a rapid classification model for AMR grades based on E-tongue and the KNN algorithm, as well as a KNN regression rapid quantitative model for the content of active components in AMR, were established. These models enable rapid prediction of the contents of extract, volatile oil, polysaccharides, atractylenolide I, atractylenolide II, atractylenolide III, bis-atractylenolide and atractylone in AMR. Furthermore, a systematic analysis of feature-extraction and feature-selection methods for E-nose and E-tongue was conducted, providing a reference for research on data-analysis methods in the field of TCM quality evaluation using sensor technologies. By comparing the performances of E-nose and E-tongue in AMR grade classification and compound regression prediction, it was found that the E-tongue was a powerful tool suitable for detecting AMR quality. The rapid detection capability of the E-tongue, combined with the data-mining potential of machine-learning algorithms, represents a feasible solution for rapid qualitative and quantitative quality control of AMR.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/molecules31050881/s1, Figure S1: Response values of E-nose and E-tongue sensors; Figure S2: Response values analysis of E-nose S13 sensor and E-tongue PKS sensor; Figure S3: Feature selection for E-nose; Figure S4: Comparison of three feature-selection methods based on E-nose; Figure S5. Feature selection for E-tongue; Figure S6. Comparison of three feature-selection methods based on E-tongue. Figure S7. The confusion matrices of KNN models; Figure S8. The plot based on canonical discriminant analysis. Table S1: Content range of quality markers of different grades; Table S2: Feature composition of different feature selection methods based on E-nose; Table S3: Feature composition of different feature selection methods based on E-tongue; Table S4: Accuracy of the classification model for machine learning; Table S5: The coefficient of determination (R2) of the regression model for machine learning; Table S6: The AMR samples’ information; Table S7: Trait characteristic scoring criteria; Table S8: Detailed information of 18 metal-oxide sensors; Table S9: Detailed information of seven E-tongue sensors.

Author Contributions

R.Y.: Methodology, validation, investigation, and writing—original draft; J.W.: methodology, investigation, and writing—review and editing; Y.W., X.G., Y.S. and Z.S.: validation, investigation, and writing—review and editing; K.Z. and Y.Z.: investigation and writing—review and editing; Y.Y.: conceptualization, supervision, and funding acquisition. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Shanxi Province 2022–2023 Traditional Chinese Medicine Technology Innovation Project, “2100601 Traditional Chinese Medicine (Ethnic Medicine) Special Project”; the 2023 Beijing University of Chinese Medicine “unveiling and leading” project, Research on Intelligent Engineering for Odor Diagnosis of Traditional Chinese Medicine Advantageous Diseases (grant number: 2023-JYB-JBQN-058); and research on the Quality Evaluation of “Imported Calculus Bovis” and Development of Expert AI Identification Equipment for Authenticable and Detectable Bovis-Related Medicinal Materials (90020271720383).

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article/Supplementary Materials, and further inquiries can be directed to the corresponding authors.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AMRAtractylodes macrocephala Rhizoma
TCMtraditional Chinese medicine
HPLChigh-performance liquid chromatography
E-noseelectronic noses
E-tongueelectronic tongue
MImutual information
SVM-RFEsupport vector machine-based recursive feature elimination
RFrandom forest
BPNNback-propagation neural network

References

  1. Gu, S.; Li, L.; Huang, H.; Wang, B.; Zhang, T. Antitumor, Antiviral, and Anti-Inflammatory Efficacy of Essential Oils from Atractylodes Macrocephala Koidz. Produced with Different Processing Methods. Molecules 2019, 24, 2956. [Google Scholar] [CrossRef]
  2. Liu, C.; Wang, S.; Xiang, Z.; Xu, T.; He, M.; Xue, Q.; Song, H.; Gao, P.; Cong, Z. The Chemistry and Efficacy Benefits of Polysaccharides from Atractylodes Macrocephala Koidz. Front. Pharmacol. 2022, 13, 952061. [Google Scholar] [CrossRef]
  3. Ji, G.; Chen, R.; Zheng, J. Atractylenolide I Inhibits Lipopolysaccharide-Induced Inflammatory Responses via Mitogen-Activated Protein Kinase Pathways in RAW264.7 Cells. Immunopharmacol. Immunotoxicol. 2014, 36, 420–425. [Google Scholar] [CrossRef] [PubMed]
  4. Ye, Y.; Wang, H.; Chu, J.-H.; Chou, G.; Chen, S.-B.; Mo, H.; Fong, W.; Yu, Z.-L. Atractylenolide II Induces G1 Cell-Cycle Arrest and Apoptosis in B16 Melanoma Cells. J. Ethnopharmacol. 2011, 136, 279–282. [Google Scholar] [CrossRef]
  5. Huang, M.; Jiang, W.; Luo, C.; Yang, M.; Ren, Y. Atractylenolide III Inhibits Epithelial-mesenchymal Transition in Small Intestine Epithelial Cells by Activating the AMPK Signaling Pathway. Mol. Med. Rep. 2022, 25, 98. [Google Scholar] [CrossRef] [PubMed]
  6. Zhao, L.; Wu, W.; Guo, L.; Yang, S.; Zhi, H.; Yang, X.; Zhou, S.; Liu, Z.; Zhang, M. Changes of Producing Regions of Atractylodes macrocephala Koidz. in China. Mod. Chin. Med. 2023, 25, 2434–2444. [Google Scholar] [CrossRef]
  7. Gong, F.; Sun, J.; Wu, M.; Wang, P.; Dong, Y.; Jiang, J.; Wang, Z. Analysis of Quality Variation of Baizhu from Different Producing Areas. J. Zhejiang Agric. Sci. 2023, 64, 1481–1487. [Google Scholar] [CrossRef]
  8. Yao, Z.; Chen, W.; Yang, Z.; Jiang, C.; Li, N.; Guo, Y.; Wang, D.; Liu, C. Research progress in Atractylodes macrocephala and predictive analysis on Q-marker. Chin. Tradit. Herb. Drug 2019, 50, 4796–4807. [Google Scholar]
  9. Liu, D.; He, L.; Cui, N.; Wang, Q. Research Progress of Chemical Constituents, Pharmacological Action and Quality Marker Prediction of Largehead Atractylodes Rhizome. Inf. Tradit. Chin. Med. 2024, 41, 65–78. [Google Scholar] [CrossRef]
  10. Huang, S. Determination of Atractylenolide I, II and III in Rhizoma Atractylodis Macrocephalae by High Performance Liquid Chromatography. Chin. J. Vet. Drug 2024, 58, 15–20. [Google Scholar]
  11. Shi, Y.-Y.; Guan, S.-H.; Tang, R.-N.; Tao, S.-J.; Guo, D.-A. Simultaneous Determination of Four Sesquiterpenoids in Atractylodes Macrocephala Rhizoma by GC-FID: Optimisation of an Ultrasound-Assisted Extraction by Central Composite Design. Phytochem. Anal. 2012, 23, 408–414. [Google Scholar] [CrossRef] [PubMed]
  12. Chen, Q.; He, H.; Li, P.; Zhu, J.; Xiong, M. Identification and Quantification of Atractylenolide I and Atractylenolide III in Rhizoma Atractylodes Macrocephala by Liquid Chromatography–Ion Trap Mass Spectrometry. Biomed. Chromatogr. 2013, 27, 699–707. [Google Scholar] [CrossRef]
  13. Gil, M.; Rudy, M.; Duma-Kocan, P.; Stanisławczyk, R. Electronic Sensing Technologies in Food Quality Assessment: A Comprehensive Literature Review. Appl. Sci. 2025, 15, 1530. [Google Scholar] [CrossRef]
  14. Zhao, X.; Xu, L.; Yang, X.; Wang, Y.; Luo, X.; Li, J.; Zhang, L.; Kang, S.; Ma, S. Research Progress on the Application of Intelligent Sensory Technology in Traditional Chinese Medicine. Chin. Pharm. Aff. 2025, 39, 96–104. [Google Scholar] [CrossRef]
  15. Li, S.-Z.; Zeng, S.-L.; Wu, Y.; Zheng, G.-D.; Chu, C.; Yin, Q.; Chen, B.-Z.; Li, P.; Lu, X.; Liu, E.-H. Cultivar Differentiation of Citri Reticulatae Pericarpium by a Combination of Hierarchical Three-Step Filtering Metabolomics Analysis, DNA Barcoding and Electronic Nose. Anal. Chim. Acta 2019, 1056, 62–69. [Google Scholar] [CrossRef]
  16. Zhou, H.; Li, Z.; Luo, D.; Gholamhosseini, H.; Han, B.; Wang, H. Rapid Qualitative and Quantitative Analysis of Volatile Components and Quality Identification of Amomi Fructus Based on Bionic Olfactory System. Rapid Qual. Quant. Anal. Volatile Compon. Qual. Identif. Amomi Fruct. Based Bionic Olfactory Syst. 2021, 17, 223–230. [Google Scholar] [CrossRef]
  17. Lei, K.; Yuan, M.; Li, S.; Zhou, Q.; Li, M.; Zeng, D.; Guo, Y.; Guo, L. Performance Evaluation of E-Nose and E-Tongue Combined with Machine Learning for Qualitative and Quantitative Assessment of Bear Bile Powder. Anal. Bioanal. Chem. 2023, 415, 3503–3513. [Google Scholar] [CrossRef]
  18. Yang, R.; Wang, Y.; Wang, J.; Guo, X.; Zhao, Y.; Zhu, K.; Zhu, X.; Zou, H.; Yan, Y. Geographical Origin Traceability of Atractylodis Macrocephalae Rhizoma Based on Chemical Composition, Chromaticity, and Electronic Nose. Molecules 2024, 29, 4991. [Google Scholar] [CrossRef]
  19. Wei, Z.-L.; Riya, A.; Sun, X.-Y.; Zhao, S.; Xie, J.-B.; Lai, C.-J.-S.; Zhang, Y.-Q. Exploration on Bitter Substance Basis of Atractylodes Macrocephala Rhizoma Based on the“Spectral Taste”Relationship between Taste Information and Chemical Components. J. Instrum. Anal. 2023, 42, 952–959. [Google Scholar] [CrossRef]
  20. Yang, R.; Zhu, K.; Zhao, Y.; Guo, X.; Wang, Y.; Wang, J.; Zou, H.; Yan, Y. Rapid Classification and Quantitative Prediction of Aflatoxin B1 Content and Colony Counts in Nutmeg Based on Electronic Nose. Molecules 2025, 30, 2538. [Google Scholar] [CrossRef] [PubMed]
  21. Yan, J.; Guo, X.; Duan, S.; Jia, P.; Wang, L.; Peng, C.; Zhang, S. Electronic Nose Feature Extraction Methods: A Review. Sensors 2015, 15, 27804–27831. [Google Scholar] [CrossRef] [PubMed]
  22. Hu, Y.; Zhao, L.; Li, Z.; Dong, X.; Xu, T.; Zhao, Y. Classifying the Multi-Omics Data of Gastric Cancer Using a Deep Feature Selection Method. Expert Syst. Appl. 2022, 200, 116813. [Google Scholar] [CrossRef]
  23. Liu, S.; Qi, X.; Xing, C.; Ming, X.; Lv, X. Research on Feature Selection for AC Contactor Vibration Signals Based on Regularized Random Forest with Recursive Selection. PLoS ONE 2024, 19, e0310110. [Google Scholar] [CrossRef] [PubMed]
  24. Gu, T.; Wang, J.; Tang, D.; Wang, J.; Guo, T. Reconstruction of Measurement Data with Multiple Outliers Using Novel Domain-Based RBF. Mech. Syst. Signal Process. 2024, 214, 111385. [Google Scholar] [CrossRef]
  25. Kaushal, S.; Nayi, P.; Rahadian, D.; Chen, H.-H. Applications of Electronic Nose Coupled with Statistical and Intelligent Pattern Recognition Techniques for Monitoring Tea Quality: A Review. Agriculture 2022, 12, 1359. [Google Scholar] [CrossRef]
  26. Feng, W.; Zhou, L.; Han, Y.; Zhang, T.; Wen, J.; Chen, C.; Wang, Y.; He, Y. Combing Chemical Composition Profiling with Machine Learning for Geographical Origins Identification of Nardostachys Jatamansi DC. Microchem. J. 2024, 207, 112087. [Google Scholar] [CrossRef]
  27. Du, H.; Xu, J.; Du, Z.; Chen, L.; Ma, S.; Wei, D.; Wang, X. MF-MNER: Multi-Models Fusion for MNER in Chinese Clinical Electronic Medical Records. Interdiscip. Sci. Comput. Life Sci. 2024, 16, 489–502. [Google Scholar] [CrossRef]
  28. Cui, T.; Chen, H.; Li, J.; Zhou, J.; Han, L.; Tian, X.; He, F.; Chen, X.; Wang, H. A Novel Strategy for Rapid Quantification of Multiple Quality Markers and Authenticity Identification Based on Near-Infrared Spectroscopy and Machine Learning Algorithms, Fructus Gardeniae as a Case Study. Microchem. J. 2025, 209, 112697. [Google Scholar] [CrossRef]
  29. Jing, W.; Zhao, X.; Li, M.; Hu, X.; Cheng, X.; Ma, S.; Wei, F. Application of Multiple-Source Data Fusion for the Discrimination of Two Botanical Origins of Magnolia Officinalis Cortex Based on E-Nose Measurements, E-Tongue Measurements, and Chemical Analysis. Molecules 2022, 27, 3892. [Google Scholar] [CrossRef]
  30. Xia, H.; Chen, W.; Hu, D.; Miao, A.; Qiao, X.; Qiu, G.; Liang, J.; Guo, W.; Ma, C. Rapid Discrimination of Quality Grade of Black Tea Based on Near-Infrared Spectroscopy (NIRS), Electronic Nose (E-Nose) and Data Fusion. Food Chem. 2024, 440, 138242. [Google Scholar] [CrossRef] [PubMed]
  31. Cai, H.; Liu, Y.; Jin, W.; Li, F.; Chen, X.; Yang, G.; Shen, W. Construction of Sensory Evaluation System of Purple Sweet Potato Rice Steamed Sponge Cake Based on Fuzzy Mathematics. Foods 2024, 13, 3527. [Google Scholar] [CrossRef]
  32. Zhang, J.; Xue, Y.; Sun, Q.; Zhang, T.; Chen, Y.; Yu, W.; Xiong, Y.; Wei, X.; Yu, G.; Wan, H.; et al. A Miniaturized Electronic Nose with Artificial Neural Network for Anti-Interference Detection of Mixed Indoor Hazardous Gases. Sens. Actuators B Chem. 2021, 326, 128822. [Google Scholar] [CrossRef]
  33. Rusinek, R.; Gancarz, M.; Krekora, M.; Nawrocka, A. A Novel Method for Generation of a Fingerprint Using Electronic Nose on the Example of Rapeseed Spoilage. J. Food Sci. 2019, 84, 51–58. [Google Scholar] [CrossRef] [PubMed]
  34. Li, J.; Othman, M.S.; Chen, H.; Yusuf, L.M. Optimizing IoT Intrusion Detection System: Feature Selection versus Feature Extraction in Machine Learning. J. Big Data 2024, 11, 36. [Google Scholar] [CrossRef]
  35. Li, Z.; Du, J.; Nie, B.; Xiong, W.; Huang, C.; Li, H. Summary of Feature Selection Methods. Comput. Eng. Appl. 2019, 55, 10–19. [Google Scholar] [CrossRef]
  36. Liu, X.; Tang, H.; Ding, Y.; Yan, D. Investigating the Performance of Machine Learning Models Combined with Different Feature Selection Methods to Estimate the Energy Consumption of Buildings. Energy Build. 2022, 273, 112408. [Google Scholar] [CrossRef]
  37. Li, Y.; Li, T.; Liu, H. Recent Advances in Feature Selection and Its Applications. Knowl. Inf. Syst. 2017, 53, 551–577. [Google Scholar] [CrossRef]
  38. Ye, S.; Weng, H.; Xiang, L.; Jia, L.; Xu, J. Synchronously Predicting Tea Polyphenol and Epigallocatechin Gallate in Tea Leaves Using Fourier Transform–Near-Infrared Spectroscopy and Machine Learning. Molecules 2023, 28, 5379. [Google Scholar] [CrossRef]
  39. Huo, Z.; Liu, Y.; Yang, R.; Dong, G.; Lin, X.; Yang, Y.; Yang, F. Qualitative and Quantitative Analysis of Microplastics in Chicken Meat Using Near-Infrared Spectroscopy. Microchem. J. 2025, 210, 112979. [Google Scholar] [CrossRef]
  40. Chen, H.; Tan, C.; Lin, Z. Application of Subspace Ensemble Radical Basis Function Networks to Quantitative Analysis of Near-Infrared and Mid-Infrared Spectroscopy. Microchem. J. 2025, 212, 113354. [Google Scholar] [CrossRef]
  41. Feng, Y.; Wang, J.; Tang, Y. Estimation and Inversion of Soil Heavy Metal Arsenic (As) Based on UAV Hyperspectral Platform. Microchem. J. 2024, 207, 112027. [Google Scholar] [CrossRef]
  42. Zou, H.; Li, R.; Xuan, X.; Jiang, Y.; Yuan, H.; An, T. Rapid and Quantitative Prediction of Tea Pigments Content During the Rolling of Black Tea by Multi-Source Information Fusion and System Analysis Methods. Foods 2025, 14, 2829. [Google Scholar] [CrossRef] [PubMed]
Figure 1. The illustration of the study on quality control strategies for rapid multi-component quantification and rapid identification of AMR grades based on E-nose, E-tongue and machine-learning algorithms.
Figure 1. The illustration of the study on quality control strategies for rapid multi-component quantification and rapid identification of AMR grades based on E-nose, E-tongue and machine-learning algorithms.
Molecules 31 00881 g001
Figure 2. Distribution of eight components in AMR at different grades. ((A): Extract content; (B): volatile oil content; (C): polysaccharide content; (D): atractylenolide I content; (E): atractylenolide II; (F): atractylenolide III; (G): bis-atractylenolide; (H): atractylone.) (Different letters indicate significant differences between the two groups, p < 0.05. The box plot displays the median (center line within the box), interquartile range (box boundaries), and whiskers (minimum and maximum values). Statistical significance was determined using the Kruskal–Wallis test).
Figure 2. Distribution of eight components in AMR at different grades. ((A): Extract content; (B): volatile oil content; (C): polysaccharide content; (D): atractylenolide I content; (E): atractylenolide II; (F): atractylenolide III; (G): bis-atractylenolide; (H): atractylone.) (Different letters indicate significant differences between the two groups, p < 0.05. The box plot displays the median (center line within the box), interquartile range (box boundaries), and whiskers (minimum and maximum values). Statistical significance was determined using the Kruskal–Wallis test).
Molecules 31 00881 g002
Figure 3. Response values analysis of E-nose S13 sensor and E-tongue PKS sensor ((AF): E-nose; (GK): E-tongue; (A): maximum value; (B): integral area corresponding to the maximum response time; (C): maximum value of the first derivative; (D): minimum value of the first derivative; (E): time corresponding to the maximum response value; (F): time corresponding to the maximum value of the first derivative; (G): response value at the 120th second; (H): steady-state response value; (I): average value of the response curve over 120 s; (J): integral area of the total response period; (K): average value of the first derivative.). (Different letters indicate significant differences between the two groups, p < 0.05).
Figure 3. Response values analysis of E-nose S13 sensor and E-tongue PKS sensor ((AF): E-nose; (GK): E-tongue; (A): maximum value; (B): integral area corresponding to the maximum response time; (C): maximum value of the first derivative; (D): minimum value of the first derivative; (E): time corresponding to the maximum response value; (F): time corresponding to the maximum value of the first derivative; (G): response value at the 120th second; (H): steady-state response value; (I): average value of the response curve over 120 s; (J): integral area of the total response period; (K): average value of the first derivative.). (Different letters indicate significant differences between the two groups, p < 0.05).
Molecules 31 00881 g003
Figure 4. Correlation plots between actual values and predicted values of the KNN regression quantitative model for eight components based on E-nose. ((A): Extract; (B): volatile oil; (C): polysaccharides; (D): atractylenolide I; (E): atractylenolide II; (F): atractylenolide III; (G): bis-atractylenolide; (H): atractylone).
Figure 4. Correlation plots between actual values and predicted values of the KNN regression quantitative model for eight components based on E-nose. ((A): Extract; (B): volatile oil; (C): polysaccharides; (D): atractylenolide I; (E): atractylenolide II; (F): atractylenolide III; (G): bis-atractylenolide; (H): atractylone).
Molecules 31 00881 g004
Figure 5. Correlation plots between actual values and predicted values of the KNN regression quantitative model for eight components based on E-tongue. ((A): Extract; (B): volatile oil; (C): polysaccharides; (D): atractylenolide I; (E): atractylenolide II; (F): atractylenolide III; (G): bis-atractylenolide; (H): atractylone).
Figure 5. Correlation plots between actual values and predicted values of the KNN regression quantitative model for eight components based on E-tongue. ((A): Extract; (B): volatile oil; (C): polysaccharides; (D): atractylenolide I; (E): atractylenolide II; (F): atractylenolide III; (G): bis-atractylenolide; (H): atractylone).
Molecules 31 00881 g005
Figure 6. Data fusion analysis. ((A): Classification model; (BI): content-prediction model ((B): extract; (C): volatile oil; (D): polysaccharides; (E): atractylenolide I; (F): atractylenolide II; (G): atractylenolide III; (H): bis-atractylenolide; (I): atractylone) (Data are presented as mean ± SD from ten independent runs. Statistical significance was assessed by Kruskal–Wallis test with Dunn’s post hoc test. Different letters indicate significant differences between the two groups, p < 0.05).
Figure 6. Data fusion analysis. ((A): Classification model; (BI): content-prediction model ((B): extract; (C): volatile oil; (D): polysaccharides; (E): atractylenolide I; (F): atractylenolide II; (G): atractylenolide III; (H): bis-atractylenolide; (I): atractylone) (Data are presented as mean ± SD from ten independent runs. Statistical significance was assessed by Kruskal–Wallis test with Dunn’s post hoc test. Different letters indicate significant differences between the two groups, p < 0.05).
Molecules 31 00881 g006
Table 1. AMR grade-classification results.
Table 1. AMR grade-classification results.
Producing AreaGradeAbbreviationThe Number of Samples
ZhejiangFirst classZJ60
Anhui, HenanSecond classAH-HN-155
Third classAH-HN-230
HebeiFourth classHB35
Table 2. Accuracy of the classification model for feature extraction based on E-nose.
Table 2. Accuracy of the classification model for feature extraction based on E-nose.
MaxSSAVIAIA-T-MaxD-MaxD-MinD-AvT-MaxDt-MaxDt-MinAll 11 TypesThe Optimal Set
The accuracy of training set87.41%89.11%89.93%90.59%70.30%82.00%74.07%91.33%69.85%57.26%51.04%86.59%93.85%
The accuracy of test set80.22%83.78%82.89%81.56%57.11%70.22%64.44%78.67%49.33%47.11%34.44%69.33%84.67%
Note: Maximum value of the response curve (Max); steady-state response value (average of the last 10 s, SS); average value of the response curve over 120 s (AV); integral area of the total response period (IA); integral area corresponding to the maximum response time (IA-T-max); maximum value of the first derivative (D-max); minimum value of the first derivative (D-min); average value of the first derivative (D-av); time corresponding to the maximum response value (T-max); time corresponding to the maximum value of the first derivative (Dt-max); time corresponding to the minimum value of the first derivative (Dt-min). (All the maximum values mentioned above refer to absolute values.)
Table 3. The coefficient of determination (R2) of the regression model for feature extraction based on E-nose.
Table 3. The coefficient of determination (R2) of the regression model for feature extraction based on E-nose.
MaxSSAVIAIA-T-MaxD-MaxD-MinD-AvT-MaxDt-MaxDt-MinAll 11 TypesThe Optimal Set
Extract0.40840.36880.40670.48890.22330.18710.23050.42810.25620.16070.15160.28790.5462
Volatile oil0.39010.39670.43260.39970.24070.16860.27600.41800.23630.12440.13010.27680.4730
Polysaccharides0.42050.34600.33880.33140.20240.11970.21800.37010.17740.13420.13590.18760.5144
Atractylenolide I0.45430.48690.54820.51810.23530.20470.40490.45960.24490.15910.19390.26030.6352
Atractylenolide II0.36150.40540.41570.41070.18870.30470.32210.40020.20500.13000.11270.16060.5346
Atractylenolide III0.39500.42630.36470.45930.18350.21870.33520.39280.25150.10380.14790.21510.5260
Bis-atractylenolide0.29170.37120.33180.37680.24430.19830.27320.37490.18730.11790.14840.23690.4353
Atractylone0.31610.42900.41630.43230.25970.25070.21000.39250.20800.15050.13250.29350.5067
Note: Maximum value of the response curve (Max); steady-state response value (average of the last 10 s, SS); average value of the response curve over 120 s (AV); integral area of the total response period (IA); integral area corresponding to the maximum response time (IA-T-max); maximum value of the first derivative (D-max); minimum value of the first derivative (D-min); average value of the first derivative (D-av); time corresponding to the maximum response value (T-max); time corresponding to the maximum value of the first derivative (Dt-max); time corresponding to the minimum value of the first derivative (Dt-min). (All the maximum values mentioned above refer to absolute values.).
Table 4. Accuracy of the classification model for feature extraction based on E-tongue.
Table 4. Accuracy of the classification model for feature extraction based on E-tongue.
Response Value at the 120th SecondSteady-State Response ValueAverage Value of the Response Curve Over 120 sIntegral Area of the Total Response PeriodAverage Value of the First DerivativeThe Optimal Set
The accuracy of training set94.37%95.11%93.85%95.63%82.15%94.67%
The accuracy of test set93.78%94.89%94.67%93.11%80.44%93.56%
Table 5. The R2 of the regression model for feature extraction based on E-tongue.
Table 5. The R2 of the regression model for feature extraction based on E-tongue.
Response Value at the 120th SecondSteady-State Response ValueAverage Value of the Response Curve Over 120 sIntegral Area of the Total Response PeriodAverage Value of the First DerivativeThe Optimal Set
Extract0.69840.64160.68780.68400.42220.6724
Volatile oil0.69440.68730.71510.68870.50500.7069
Polysaccharides0.70720.65330.63050.63070.32990.6702
Atractylenolide I0.69290.73250.75330.72300.54780.7519
Atractylenolide II0.65880.70280.74880.65960.31460.7165
Atractylenolide III0.68880.65340.62660.67120.36670.7030
Bis-atractylenolide0.54780.55800.53860.55770.33850.6187
Atractylone0.60630.61710.60860.65300.34620.6826
Table 6. Selected features used for building machine-learning models.
Table 6. Selected features used for building machine-learning models.
Feature IDFeature NameFeature DescriptionSource
F1S8SSaverage of the last 10 s of S8E-nose
F2S8D-avaverage value of the first derivative of S8
F3S10D-avaverage value of the first derivative of S10
F4S10SSaverage of the last 10 s of S10
F5S10AVaverage value of the response curve over 120 s of S10
F6S13Maxmaximum value of the absolute response curve of S13
F7S15Maxmaximum value of the absolute response curve of S15
F8S17Maxmaximum value of the absolute response curve of S17
F9S17AVaverage value of the response curve over 120 s of S17
F10S17SSaverage of the last 10 s of S817
F11S17D-avaverage value of the first derivative of S17
F12S18Maxmaximum value of the absolute response curve of S18
F13S18AVaverage value of the response curve over 120 s of S18
F14S18SSaverage of the last 10 s of S18
F15S18D-avaverage value of the first derivative of S18
T1PKS-AVaverage value of the response curve over 120 s of PKSE-tongue
T2CTS-AVaverage value of the response curve over 120 s of CTS
T3SCS-AVaverage value of the response curve over 120 s of SCS
T4ANS-AVaverage value of the response curve over 120 s of ANS
Table 7. Comparison of KNN classification models based on E-nose and E-tongue.
Table 7. Comparison of KNN classification models based on E-nose and E-tongue.
Training SetsTest Sets
GradesPrecisionRecallF1W-F1 1AccuracyPrecisionRecallF1W-F1 1Accuracy
E-noseZJ100.0%97.83%98.90%91.44%91.85%93.33%100.0%96.55%89.04%88.89%
AH-HN-197.56%80.00%87.91%85.71%80.00%82.76%
AH-HN-260.87%100.0%75.68%71.43%71.43%71.43%
HB96.15%100.0%98.04%100.0%100.0%100.0%
E-tongueZJ100.0%100.0%100.0%99.26%99.26%100.0%100.0%100.0%95.56%95.56%
AH-HN-1100.0%97.62%98.80%92.86%92.86%92.86%
AH-HN-295.65%100.0%97.78%85.71%85.71%85.71%
HB100.0%100.0%100.0%100.0%100.0%100.0%
W-F1 1, Weighted F1.
Table 8. The feature-extraction methods for the E-nose and E-tongue.
Table 8. The feature-extraction methods for the E-nose and E-tongue.
SourceFeature NameDescriptionFeature Category
E-noseMaxmaximum value of the absolute response curveSteady-state value
SSsteady-state response value (average of the last 10 s)
AVaverage value of the response curve over 120 s
IAintegral area of the total response periodTransient value
IA-T-maxintegral area corresponding to the maximum response time
D-maxmaximum value of the first derivative
D-minminimum value of the first derivative
D-avaverage value of the first derivative
T-maxtime corresponding to the maximum response valueThe time point corresponding to the specific response value
Dt-maxtime corresponding to the maximum value of the first derivative
Dt-mintime corresponding to the minimum value of the first derivative
E-tongue120thresponse value at the 120th secondSteady-state value
SSsteady-state response value (average of the last 10 s)
AVaverage value of the response curve over 120 s
IAintegral area of the total response periodTransient value
D-avaverage value of the first derivative
Note: All the maximum values mentioned above refer to absolute values.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Yang, R.; Wang, J.; Wang, Y.; Guo, X.; Sun, Y.; Song, Z.; Zhu, K.; Zhao, Y.; Yan, Y. A Novel Strategy for Rapid Quantification of Multiple Quality Indicators and Grade Discrimination of Atractylodis macrocephalae Rhizoma Based on Electronic Nose, Electronic Tongue and Machine-Learning Algorithms. Molecules 2026, 31, 881. https://doi.org/10.3390/molecules31050881

AMA Style

Yang R, Wang J, Wang Y, Guo X, Sun Y, Song Z, Zhu K, Zhao Y, Yan Y. A Novel Strategy for Rapid Quantification of Multiple Quality Indicators and Grade Discrimination of Atractylodis macrocephalae Rhizoma Based on Electronic Nose, Electronic Tongue and Machine-Learning Algorithms. Molecules. 2026; 31(5):881. https://doi.org/10.3390/molecules31050881

Chicago/Turabian Style

Yang, Ruiqi, Jiayu Wang, Yushi Wang, Xingyu Guo, Yunqi Sun, Ziyue Song, Keyao Zhu, Yuanyu Zhao, and Yonghong Yan. 2026. "A Novel Strategy for Rapid Quantification of Multiple Quality Indicators and Grade Discrimination of Atractylodis macrocephalae Rhizoma Based on Electronic Nose, Electronic Tongue and Machine-Learning Algorithms" Molecules 31, no. 5: 881. https://doi.org/10.3390/molecules31050881

APA Style

Yang, R., Wang, J., Wang, Y., Guo, X., Sun, Y., Song, Z., Zhu, K., Zhao, Y., & Yan, Y. (2026). A Novel Strategy for Rapid Quantification of Multiple Quality Indicators and Grade Discrimination of Atractylodis macrocephalae Rhizoma Based on Electronic Nose, Electronic Tongue and Machine-Learning Algorithms. Molecules, 31(5), 881. https://doi.org/10.3390/molecules31050881

Article Metrics

Back to TopTop