Skip to Content
FoodsFoods
  • Article
  • Open Access

15 April 2026

23 Pages

Multi-Granularity Domain Adversarial Learning for Cross-Domain Tea Classification Using Electronic Nose Signals

and
1
College of Information Science and Technology, Beijing University of Chemical Technology, Beijing 100029, China
2
School of Biomedical Engineering, Capital Medical University, Beijing 100069, China
*
Author to whom correspondence should be addressed.

Abstract

Rapid and reliable tea classification is valuable for routine product screening, yet conventional sensory or physicochemical methods are subjective or time-consuming. Electronic nose (E-nose) sensing provides a fast alternative, but performance often degrades under domain shifts caused by different tea types, commercial categories, or acquisition conditions. This study proposes MGDA-Net, a multi-granularity domain adversarial network for cross-domain tea classification using E-nose time-series signals. MGDA-Net learns local temporal dynamics via a CNN branch and global contextual dependencies via a self-attention branch, and fuses them through an adaptive gating module. A branch-level adversarial alignment strategy is introduced to reduce source–target discrepancy at both local and global feature levels. A three-stage training procedure, consisting of source pretraining, adversarial alignment, and target fine-tuning, enables knowledge transfer from a labeled green tea source-domain to two target tasks. Experiments on oolong tea commercial-category classification (6 classes) and jasmine tea retail price-level classification (8 classes) show that MGDA-Net achieves mean accuracies of 99.31 ± 0.69% and 99.38 ± 0.51% over 10 independent runs, substantially outperforming all compared baseline methods. Ablation studies, feature-space analyses, and label-efficiency experiments further confirm the contribution of each component and show that MGDA-Net maintains mean accuracies above 87% when only 40% of the target-domain labels are used for fine-tuning. These findings suggest that MGDA-Net is a promising approach for cross-domain tea classification using E-nose data.

1. Introduction

Tea is one of the most widely consumed beverages worldwide and also a high-value agricultural commodity with substantial economic, cultural, and commercial significance [1,2]. According to official statistics from the National Bureau of Statistics of China, national tea production reached 3.74 million tons in 2024, representing a 5.5% year-on-year increase [3]. At the global level, tea production generates more than USD 18 billion annually and supports the livelihoods of approximately 13 million people [4]. With the continuous expansion and diversification of the tea market, reliable tea classification and product differentiation have become increasingly important for product grading, commercial management, market supervision, and consumer protection. In practice, premium teas may be misclassified, mislabeled, or mixed with lower-value products, while adjacent commercial categories often exhibit only subtle sensory differences [5,6,7]. These issues create a strong demand for rapid, objective, and scalable methods for tea classification under realistic industrial conditions.
Conventional tea evaluation still relies mainly on sensory assessment and physicochemical analysis [8,9]. Sensory evaluation remains an important reference approach in routine grading, but it depends heavily on trained panelists and is therefore inherently subjective, labor-intensive, and vulnerable to inter-observer variability [10,11]. Physicochemical methods, such as chromatographic and spectrometric analyses of catechins, amino acids, and volatile compounds, can provide more objective evidence, but they usually require expensive instrumentation, complex sample preparation, destructive testing, and time-consuming analytical procedures [12,13]. Several rapid and non-destructive alternatives, including near-infrared spectroscopy and hyperspectral imaging, have also been explored for tea analysis, yet their performance may still be sensitive to measurement geometry, instrument calibration, and environmental fluctuations [14,15,16]. Against this background, the electronic nose (E-nose), which captures aroma-response patterns through sensor arrays, has attracted increasing attention because of its rapid response, non-destructive measurement, operational simplicity, and potential for low-cost on-site screening [17,18]. Representative studies have already demonstrated the applicability of E-nose methods in tea-related tasks. For example, Kang et al. [19] combined an E-nose with an adaptive pooling attention mechanism to discriminate teas from different picking periods, Yu and Gu [20] developed a CNN–SVM framework for fine-grained green tea classification with geographical indication, and Chang and Lu [21] proposed RLCSA-Net for E-nose-based tea identification. These studies collectively indicate that E-nose signals contain informative patterns for tea category discrimination, commercial-label-related classification, and origin identification.
Despite these advances, most reported E-nose models are still developed and evaluated under single-domain or closed-set conditions, where the training and test samples are collected under highly similar acquisition settings [22,23]. In practical deployment, however, E-nose measurements may vary across acquisition days, batches, tea types, grading criteria, and ambient conditions, leading to domain shifts that can substantially degrade model performance when a trained classifier is transferred to a new target task [24,25,26]. This challenge becomes more pronounced in cross-category tea classification, where the source and target tasks may differ not only in data distribution but also in class definition. Domain adaptation has been widely studied to mitigate cross-domain distribution discrepancy in machine learning [27,28]. In E-nose research, existing adaptation methods have mainly focused on sensor drift correction or single-condition transfer [29,30,31,32,33]. However, E-nose signals are multichannel time-series in which domain discrepancy may appear simultaneously in local transient dynamics (e.g., rise and recovery patterns) and global contextual structure (e.g., long-range temporal trends and inter-sensor dependencies) [34,35]. Methods that align only a single representation space may therefore overlook complementary cross-domain cues and be insufficient for fine-grained classification scenarios [36,37,38,39]. Consequently, an important remaining research gap is the lack of domain adaptation frameworks that explicitly model and align E-nose features at multiple granularities for robust cross-domain tea classification.
To address this gap, this study proposes a multi-granularity domain adversarial network (MGDA-Net) for cross-domain tea classification using E-nose time-series signals. Green tea is used as the labeled source-domain, and knowledge is transferred to two heterogeneous target tasks, namely oolong tea commercial-category classification and jasmine tea retail price-level classification. Accordingly, the present study focuses on cross-domain classification of heterogeneous commercial label sets rather than on independently validated sensory or chemical quality grading. MGDA-Net combines a convolutional neural network (CNN) branch to learn local temporal dynamics, a self-attention branch to model global contextual dependencies, an adaptive fusion module to integrate the two branches, and a branch-level adversarial alignment strategy within a progressive three-stage transfer framework. The main contributions of this work are summarized as follows:
1.
A multi-granularity feature learning framework is proposed to jointly capture local temporal patterns and global contextual dependencies from multichannel E-nose signals. The framework integrates a CNN branch for local dynamic feature extraction, a self-attention branch for global contextual modeling, and an adaptive fusion module with learnable gating for dynamic integration of heterogeneous branch features.
2.
A branch-level adversarial domain adaptation strategy is developed to reduce source–target discrepancy simultaneously at local and global feature levels. Combined with a three-stage transfer procedure, this design improves the transferability and robustness of the learned representations in cross-domain tea classification.
3.
Comprehensive evaluations are conducted on two heterogeneous target tasks with different classification criteria and difficulty levels. Comparisons with five target-only baselines and three representative domain adaptation baselines, together with ablation studies, multi-seed robustness analysis, feature-level analyses, and label-efficiency experiments, validate the effectiveness and stability of MGDA-Net.

2. Materials and Methods

2.1. Tea Samples

Three tea datasets were used in this study to evaluate the cross-domain classification performance of the proposed model, as summarized in Table 1.
Table 1. Overview of the three tea datasets used in this study.
The source-domain dataset was adopted from our group’s previous study [20]. It consisted of 12 varieties of Chinese green tea (Camellia sinensis) collected in Beijing, China, in 2020. The samples were sourced from major tea-producing provinces, including Anhui, Zhejiang, Sichuan, Guizhou, Henan, Guangxi, Hunan, and Shaanxi. The 12 varieties comprised six Maofeng-type teas (Huangshan Maofeng, Yandang Maofeng, Emei Maofeng, Lanxi Maofeng, Qishan Maofeng, and Meitan Maofeng) and six Maojian-type teas (Xinyang Maojian, Guilin Maojian, Duyun Maojian, Guzhang Maojian, Queshe Maojian, and Ziyang Maojian). In total, 1440 green tea samples were used as the source-domain dataset.
Target-domain 1 consisted of six classes of oolong tea collected in Xiamen, China, in 2022. All samples were provided by the Xiamen Institute of Quality Inspection and Testing. The six classes included four non-branded yancha (rock tea) samples, namely Kuidou Shuixian (Premium, hereafter O1), Kuidou Shuixian (Grade I, hereafter O2), Foguoyan Shuixian (A++, hereafter O3), and Foguoyan Shuixian (A, hereafter O4), as well as two branded products, Xupinhao Shuixian (hereafter O5) and Xupinhao Rougui (hereafter O6). The first two Shuixian classes originated from Kuidou Village, Yongchun County, Fujian Province, while the latter two Shuixian classes originated from the Foguoyan area of Wuyi Mountain, Fujian Province. Each class contained 60 samples, yielding 360 samples in total.
Target-domain 2 consisted of eight classes of jasmine tea collected in Beijing, China, in 2025. All samples were produced by Wu Yutai, a well-known Chinese tea brand based in Beijing. The eight classes were defined by their retail price levels, ranging from CNY 50 to CNY 400 per 500 g at increments of CNY 50, representing different commercial price tiers. Each class contained 60 samples, resulting in 480 samples in total.
Green tea was selected as the source-domain because it shares a common botanical origin and partially overlapping volatile organic compound profiles with both oolong tea and jasmine tea, while exhibiting processing-induced differences [40,41,42]. This design enables the evaluation of cross-category domain adaptation, where knowledge learned from one tea category is transferred to classify chemically related categories under distribution shifts.

2.2. E-Nose System and Data Acquisition

All E-nose measurements were performed using a PEN3 portable E-nose system (Airsense Analytics GmbH, Schwerin, Germany), which is equipped with an array of 10 metal oxide semiconductor (MOS) gas sensors. Each sensor exhibits selective sensitivity to different groups of volatile compounds. The source-domain data adopted from our previous study were acquired using the same PEN3 instrument type and a consistent acquisition protocol as those used for the target-domain datasets.
For each measurement, approximately 4.0 g of tea sample was placed in a sealed 50-mL glass vial and equilibrated at room temperature for 3 min to allow headspace volatiles to accumulate. All measurements were conducted under controlled laboratory conditions at 25 ± 1 °C and 40 ± 2% relative humidity. Filtered ambient air was used as the carrier gas at a constant flow rate of 600 mL/min. Each acquisition cycle consisted of two stages. In the cleaning stage, filtered air was passed through the sensor chamber for 100 s to restore the sensor baseline. In the subsequent sampling stage, the headspace gas from the sample vial was introduced into the sensor chamber for 120 s. The sampling interval was 1 s, resulting in 120 time points for each of the 10 sensors. Therefore, each measurement was represented as a 10 × 120 response matrix describing the temporal evolution of the sensor array.
Based on the above acquisition procedure, three datasets were constructed for the cross-domain experiments. The source-domain dataset consisted of 1440 measurements from 12 green tea varieties, collected over 12 days, with 10 samples per variety measured each day (one measurement per sample). The oolong tea dataset contained 360 samples from 6 classes (60 samples per class), and the jasmine tea dataset contained 480 samples from 8 classes (60 samples per class).
For each target-domain dataset, measurements were collected over 10 consecutive days, with 6 measurements acquired per class per day. To simulate a realistic cross-session deployment scenario, a day-wise split was adopted. Specifically, the data collected on days 1–8 were used as the target-domain training portion, whereas the data collected on days 9–10 were reserved as the target-domain test portion. Accordingly, the oolong tea dataset was divided into 288 training measurements and 72 test measurements, and the jasmine tea dataset was divided into 384 training measurements and 96 test measurements. This temporal split ensured that the final evaluation was performed on measurements obtained in independent acquisition sessions rather than on randomly mixed data from the same session.
For the oolong tea dataset, each of the 60 samples per class was treated as a physically independent unit. The four non-branded yancha classes were provided as individually packaged quality inspection samples by the Xiamen Institute of Quality Inspection and Testing, with each measured sample originating from a distinct production lot. The two branded products (Xupinhao Shuixian and Xupinhao Rougui) were sourced from separate retail packages. No sample was measured more than once. The day-wise train–test split therefore ensures that training and test partitions reflect separate acquisition sessions with no overlapping measured samples.
For the jasmine tea dataset, each of the 60 samples per class was prepared from a separately opened retail package of Wu Yutai jasmine tea at the corresponding price level. This package-level sampling introduced within-class variability beyond repeated measurement of a single package. Each sample was measured exactly once. As with the oolong tea dataset, the day-wise temporal split ensures that training and test samples were collected in independent acquisition sessions.
In the cross-domain experiments, the green tea dataset served as the labeled source-domain, while the oolong tea and jasmine tea datasets were treated as two independent target domains. During Stage 2 domain adversarial training, the target-domain training samples participated in feature alignment without using their class labels. During Stage 3 fine-tuning, the same target-domain training samples were used with their class labels for supervised classifier adaptation. The held-out test set was used exclusively for final evaluation and was not involved in any stage of model training or model selection. Accordingly, the proposed setting should be interpreted as a supervised three-stage transfer framework rather than an unsupervised target-domain adaptation setting.

2.3. Evaluation Metrics

To comprehensively assess the classification performance of the proposed model, five evaluation metrics were adopted in this study. Overall accuracy (OA) measured the proportion of correctly classified samples among all test samples, providing the most intuitive indication of model performance. It is defined as:
O A = ∑ i = 1 N 1 ( y ^ i = y i ) N
where N denotes the total number of test samples, y i and y ^ i are the true and predicted labels of the i-th sample, and 1 ( · ) is the indicator function.
For each class c, precision P c and recall R c were defined as:
P c = T P c T P c + F P c , R c = T P c T P c + F N c
where T P c , F P c , and F N c denote the numbers of true positives, false positives, and false negatives for class c, respectively. Macro-averaged precision ( P m a c r o ) and macro-averaged recall ( R m a c r o ) were computed by averaging per-class values equally across all C classes:
P m a c r o = 1 C ∑ c = 1 C P c , R m a c r o = 1 C ∑ c = 1 C R c
Macro-averaged precision reflected the model’s ability to avoid false positive errors across all classes, while macro-averaged recall captured its ability to correctly identify samples from each class.
The macro-averaged F1-score (F1-macro) was obtained by first computing the harmonic mean of P c and R c for each class and then averaging across all classes
F 1 m a c r o = 1 C ∑ c = 1 C 2 P c R c P c + R c
By treating all classes equally regardless of sample size, this metric provided a balanced evaluation of the trade-off between precision and recall, making it particularly suitable for multi-class classification tasks.
Cohen’s kappa coefficient ( κ ) quantified the degree of agreement between predicted and true labels beyond what would be expected by chance, offering a more rigorous evaluation than accuracy alone. It is defined as:
κ = p 0 − p e 1 − p e
where p 0 is the observed agreement (equivalent to OA) and p e is the expected agreement under random classification.

3. Proposed Method

3.1. Overview of MGDA-Net

The proposed multi-granularity domain adversarial network (MGDA-Net) is designed to transfer discriminative knowledge learned from a labeled source-domain to a target domain for E-nose-based tea classification. The overall architecture, illustrated in Figure 1, consists of four main components: (1) a multi-granularity feature extractor comprising a CNN branch and a self-attention branch, (2) an adaptive fusion module that dynamically combines local and global features, (3) a multi-granularity domain adversarial module with two independent discriminators operating at different feature granularities, and (4) a multilayer perceptron (MLP) classifier with center loss regularization.
Figure 1. Overall architecture of the proposed MGDA-Net.
Given an input E-nose measurement X ∈ R C × T , where C = 10 denotes the number of sensor channels and T = 120 denotes the number of time steps, the network first applies instance normalization across each channel to reduce sensor-level distributional variability. The normalized input is then processed in parallel by the CNN branch and the self-attention branch to extract local and global feature representations, respectively. These two representations are combined through the adaptive fusion module to produce a fused feature vector, which is subsequently fed into the classifier for category prediction. During domain adversarial training, the local and global features are independently passed through gradient reversal layers and domain discriminators to align the source- and target-domain distributions at multiple granularities.

3.2. Multi-Granularity Feature Extraction

3.2.1. CNN Branch for Local Feature Extraction

The CNN branch is designed to capture local temporal patterns and short-range dependencies in the E-nose response signals. It consists of three successive convolutional blocks followed by a fully connected layer. Each convolutional block contains a one-dimensional convolution, batch normalization, ReLU activation, and max pooling. The first two blocks employ a kernel size of 5 with 32 and 64 filters, respectively, while the third block uses a kernel size of 3 with 128 filters. After the third convolutional block, adaptive average pooling is applied to compress the temporal dimension to a single value, resulting in a 128-dimensional vector. This vector is then projected through a fully connected layer with batch normalization and ReLU activation to produce the local feature representation f c n n ∈ R d , where d = 128 . The local feature extraction process can be summarized as:
f c n n = FC AdaptiveAvgPool ConvBlock 3 ( ConvBlock 2 ( ConvBlock 1 ( X ) ) )
The relatively small kernel sizes and hierarchical pooling operations enable the CNN branch to focus on localized sensor response patterns and fine-grained temporal variations, which are essential for distinguishing tea samples with subtle differences in volatile compound profiles.

3.2.2. Self-Attention Branch for Global Feature Extraction

The self-attention branch is designed to model long-range temporal dependencies and capture global contextual information across the entire measurement sequence. The input X ∈ R C × T is first transposed to R T × C and projected to a higher-dimensional space through a linear layer with ReLU activation, yielding H 0 ∈ R T × d m , where d m = 64 is the model dimension.
Sinusoidal positional encodings are then added to preserve temporal order information
P E ( p o s , 2 i ) = sin p o s 10,000 2 i / d m , P E ( p o s , 2 i + 1 ) = cos p o s 10,000 2 i / d m
where p o s denotes the time step index and i denotes the dimension index. The position-aware representations are processed by a Transformer encoder consisting of L = 2 layers [43], each containing a multi-head self-attention mechanism with h = 4 heads and a feed-forward network with a hidden dimension of 4 d m = 256 . Pre-layer normalization is adopted to stabilize training. The multi-head self-attention is computed as:
MultiHead ( Q , K , V ) = Concat ( head 1 , … , head h ) W O
head j = Attention ( Q W j Q , K W j K , V W j V ) = softmax Q W j Q ( K W j K ) ⊤ d k V W j V
where d k = d m / h = 16 is the dimension per head. After the Transformer encoder, temporal mean pooling is applied across all time steps to aggregate the sequence-level information into a fixed-length vector. This vector is then projected through a fully connected layer with batch normalization and ReLU activation to produce the global feature representation f a t t n ∈ R d .
By attending to all time steps simultaneously, the self-attention branch captures holistic sensor response trends and inter-channel correlations that span the entire measurement duration, providing complementary information to the local features extracted by the CNN branch.

3.3. Adaptive Fusion Module

To effectively integrate the local and global feature representations, an adaptive fusion module is introduced that dynamically adjusts the relative contribution of each branch based on the input. Rather than using a fixed weighting scheme, the module employs a learnable gating mechanism that computes a sample-dependent fusion coefficient.
The local and global features are first concatenated and passed through a gating network consisting of two fully connected layers with ReLU activation and a sigmoid output:
α = σ W 2 · ReLU ( W 1 [ f c n n ; f a t t n ] + b 1 ) + b 2
where [ · ; · ] denotes concatenation, α ∈ ( 0 , 1 ) is a scalar gating coefficient shared across all feature dimensions, and σ ( · ) is the sigmoid function. The weighted features are then concatenated and projected to form the fused representation:
f f u s e d = ReLU BN W p [ α f c n n ; ( 1 − α ) f a t t n ] + b p
where f f u s e d ∈ R d f with d f = 256 . This adaptive mechanism allows the network to emphasize local features for samples where fine-grained patterns are more discriminative, while relying more on global features for samples where long-range dependencies are more informative.

3.4. Multi-Granularity Domain Adversarial Learning

To reduce the distribution discrepancy between the source and target domains, MGDA-Net employs a multi-granularity domain adversarial strategy that aligns feature distributions at both the local and global levels simultaneously. Unlike conventional domain adversarial methods that apply a single discriminator to the fused or final-layer features, the proposed approach operates two independent domain discriminators on the CNN branch features and the self-attention branch features, respectively. This design enables finer-grained domain alignment by encouraging cross-domain feature consistency at each feature granularity before fusion.
Each domain discriminator is implemented as an MLP with two hidden layers of 128 and 64 units, ReLU activations, and dropout regularization, followed by a single output unit. The discriminators receive features through a gradient reversal layer (GRL), which acts as an identity function during forward propagation but reverses the gradient sign during backpropagation [44]:
GRL ( f ) = f , ∂ GRL ( f ) ∂ f = − λ I
where λ is the reversal intensity that controls the strength of domain adversarial learning. The domain adversarial loss for each granularity is formulated as a binary cross-entropy loss:
L domain ( g ) = − 1 n s + n t ∑ i = 1 n s log 1 − D g GRL ( f i s ) + ∑ j = 1 n t log D g GRL ( f j t )
where g ∈ { local , global } denotes the granularity, D g is the corresponding domain discriminator, and n s and n t are the numbers of source and target samples in each mini-batch. The source and target domains are assigned binary domain labels 0 and 1, respectively.
To prevent the domain adversarial signal from destabilizing training in the early stages, a progressive scheduling strategy is adopted for the reversal intensity. The values of λ local and λ global are gradually increased from zero to their respective maximum values following a sigmoid-shaped schedule:
λ ( e ) = λ max 2 1 + exp ( − 10 · e / E ) − 1
where e is the current epoch and E is the total number of training epochs. The maximum reversal intensities are set to λ local max = 0.2 and λ global max = 0.1 , with a higher value assigned to the local branch to encourage stronger alignment of fine-grained local patterns that are more susceptible to domain shift in E-nose data.

3.5. Classification with Center Loss Regularization

The fused feature f fused is fed into an MLP classifier consisting of a fully connected layer with 128 hidden units, batch normalization, ReLU activation, dropout ( p = 0.5 ), and a final linear layer that outputs class logits. The classification loss is the cross-entropy loss with label smoothing ( ϵ = 0.05 ):
L c l s = − 1 N ∑ i = 1 N ∑ c = 1 C y ˜ i , c log ( p ^ i , c )
where y ˜ i , c = ( 1 − ϵ ) y i , c + ϵ / C is the smoothed label and p ^ i , c is the predicted probability for class c.
To further enhance intra-class compactness in the feature space, a center loss is incorporated as an auxiliary regularization term [45]:
L c e n t e r = 1 N ∑ i = 1 N h i − c y i 2 2
where h i is the hidden feature of sample i extracted from the penultimate layer of the classifier, and c y i is the learnable center vector for class y i . The center vectors are updated using a dedicated SGD optimizer with a learning rate of 0.5.

3.6. Three-Stage Training Strategy

The training of MGDA-Net follows a three-stage progressive strategy that sequentially addresses source-domain learning, cross-domain alignment, and target-domain adaptation.
Stage 1 (Source-domain pretraining). In the first stage, the entire network is trained on the labeled source-domain data to learn general-purpose feature representations for tea classification. The training objective combines the classification loss and the center loss:
L s t a g e 1 = L c l s + λ c L c e n t e r
where λ c = 0.01 is the center loss weight. The model is optimized using the Adam optimizer with a learning rate of 1 × 10 − 3 and cosine annealing scheduling for 200 epochs. The model checkpoint with the highest validation accuracy on the source-domain is retained for the next stage.
Stage 2 (Multi-granularity domain adversarial training). In the second stage, the feature extractors and domain discriminators are jointly trained to align the source- and target-domain distributions. The source-domain classification loss L c l s s is maintained to preserve discriminative capacity, while the domain adversarial losses are introduced at both the local and global feature levels. The total training objective is:
L s t a g e 2 = L c l s s + λ d L d o m a i n l o c a l + L d o m a i n g l o b a l
where λ d = 0.5 is the domain adversarial weight. Separate optimizers are used for the feature extractors (learning rate 2 × 10 − 4 ) and the domain discriminators (learning rate 1 × 10 − 3 ), with the discriminators trained at a higher learning rate to ensure they provide a strong adversarial signal. This stage is trained for 200 epochs with progressively increasing reversal intensity.
Stage 3 (Target-domain fine-tuning). In the third stage, the classifier head is replaced with a new MLP adapted to the target-domain class count, and the center loss module is reinitialized with the corresponding number of target classes. The entire network is then fine-tuned on the labeled target-domain training data. Differential learning rates are applied to different components to preserve the domain-invariant features learned during Stage 2 while allowing the new classifier to adapt quickly. Specifically, the CNN and self-attention branches use a learning rate of 3 × 10 − 5 , the fusion module uses 5 × 10 − 5 , and the new classifier uses 1 × 10 − 4 . The training objective follows the same form as Stage 1 with the center loss regularization. The model is trained for 300 epochs with cosine annealing. To reduce overfitting during target-domain fine-tuning, 20% of the labeled target-domain training samples were randomly split off as a stratified validation subset. The checkpoint with the best validation accuracy was then selected for final evaluation on the held-out target-domain test set.

4. Results and Discussion

4.1. PCA Visualization of E-Nose Data

To obtain an initial understanding of the data distribution and class separability in the two target-domain datasets, principal component analysis (PCA) was performed on the steady-state E-nose responses of the oolong tea and jasmine tea samples. Specifically, the responses recorded during the last 20 s of each 120 s measurement cycle were used to represent the stable sensor characteristics. The responses of each sensor over the last 20 s were averaged to obtain a single steady-state value per channel, yielding a 10-dimensional feature vector for each sample. These vectors were standardized to zero mean and unit variance before PCA projection. The two-dimensional PCA results are shown in Figure 2.
Figure 2. PCA visualization of the steady-state E-nose responses for (a) the oolong tea dataset and (b) the jasmine tea dataset.
For the oolong tea dataset, the first two principal components explained 52.9% and 29.5% of the total variance, respectively, corresponding to a cumulative explained variance of 82.4%. The PCA plot revealed moderate class separability among the six oolong tea categories. In particular, Xupinhao Shuixian and Xupinhao Rougui formed relatively distinct clusters in the lower-right region of the feature space, indicating that their volatile profiles differed markedly from those of the other oolong tea classes. By contrast, substantial overlap was observed among several yancha-related classes, especially Kuidou Shuixian (Premium), Foguoyan Shuixian (A++), and Foguoyan Shuixian (A), which were distributed in closely adjacent regions. These observations suggest that although certain oolong tea categories exhibit distinguishable aroma-response characteristics, fine-grained discrimination among high-end rock tea categories remains challenging when using linear projection alone.
For the jasmine tea dataset, the first two principal components accounted for 64.2% and 24.7% of the total variance, respectively, yielding a cumulative explained variance of 88.9%. Despite the relatively high variance captured by the first two components, the PCA scatter plot showed substantial inter-class overlap among the eight retail price-level categories. Most samples were concentrated in the central region of the projected space, while only a few classes, such as CNY 50 and CNY 100, exhibited partially distinguishable clusters at the periphery. This result indicates that the volatile profiles of jasmine tea samples from different retail price levels are highly similar, and that the differences among adjacent price levels are too subtle to be effectively separated by linear dimensionality reduction.
The PCA results of both datasets demonstrate that conventional linear projection is insufficient for satisfactory separation of the target-domain tea samples, particularly in fine-grained retail price-level classification. These findings highlight the necessity of nonlinear feature learning for robust cross-domain tea classification based on E-nose data.

4.2. Comparison with Baseline Methods

To evaluate the effectiveness of the proposed MGDA-Net, eight representative baseline methods were considered, including both classical machine learning classifiers and deep learning-based approaches. The target-only methods comprised three traditional classifiers—support vector machine (SVM), random forest (RF), and k-nearest neighbor (KNN)—and two deep learning models—long short-term memory network (LSTM) and one-dimensional residual network (1D-ResNet). These five methods were trained using only the target-domain training set without access to source-domain data. The classical classifiers were included to quantify how much discrimination can be achieved from target-domain data alone using conventional chemometric features, whereas the transfer-learning baselines were included to assess whether the gains of MGDA-Net arise merely from using source-domain knowledge or specifically from multi-granularity adversarial alignment. LSTM was selected as a widely adopted recurrent architecture for temporal E-nose signal modeling, while 1D-ResNet was included to provide a representative convolutional baseline with residual learning capability. Together, they cover two mainstream deep learning paradigms (recurrent and convolutional) for one-dimensional time-series classification.
For transfer-learning comparison, three widely used domain adaptation methods—deep adaptation network (DAN), domain adversarial neural network (DANN), and conditional domain adversarial network (CDAN)—were implemented under the same source-to-target transfer setting as MGDA-Net, using the three-stage protocol (source pretraining, domain adversarial alignment, and target fine-tuning with classifier replacement). To ensure fairness, all methods used the same day-wise train/test split. Target-only methods were standardized using statistics fitted on the available target-domain training set, whereas transfer-learning methods were standardized using source-domain training statistics because their Stages 1–2 were source-driven. In all cases, preprocessing parameters were estimated without using any test data. For target-only deep models and transfer-learning methods, the same 20% stratified validation subset from the target-domain training portion was used for checkpoint selection during fine-tuning. For SVM, RF, and KNN, hyperparameters were selected on the target-domain training set only. The detailed data usage and preprocessing protocols for each method are summarized in Table 2.
Table 2. Summary of data usage and preprocessing for each compared method.
For SVM, a radial basis function (RBF) kernel was used. The regularization parameter C and kernel coefficient γ were optimized via 5-fold cross-validation on the target-domain training set, with C ∈ { 0.1 , 1 , 10 , 100 } and γ ∈ { 0.001 , 0.01 , 0.1 , 1 } . RF was configured with 200 estimators, fully grown trees (no maximum depth constraint), a minimum of two samples per leaf, and the number of features considered at each split set to p , where p is the feature dimension. KNN used k = 5 neighbors with distance-based weighting and Euclidean distance. LSTM consisted of two bidirectional layers with 128 hidden units, dropout of 0.5 between layers, and a fully connected output layer; it was trained for 500 epochs using the Adam optimizer with a learning rate of 1 × 10 − 3 and cosine annealing scheduling. 1D-ResNet comprised an initial convolutional layer followed by three residual stages with 64, 128, and 256 filters, global average pooling, and a fully connected classifier; it was trained under the same optimization setting as LSTM. For DAN, DANN, and CDAN, the backbone feature extractor was a three-layer one-dimensional CNN consistent with the CNN branch of MGDA-Net, and the same three-stage training protocol was applied for fair comparison under heterogeneous source–target label spaces.
To assess robustness to training stochasticity, all methods involving random initialization or stochastic training (RF, LSTM, 1D-ResNet, DAN, DANN, CDAN, and MGDA-Net) were evaluated over 10 independent runs with different random seeds, and the results are reported as mean ± standard deviation. SVM and KNN are deterministic and are therefore reported from a single run.
On the oolong tea dataset (Table 3), MGDA-Net achieved the best performance across all five evaluation metrics, with a mean OA of 99.31 ± 0.69%, a mean F1-macro of 99.30 ± 0.70%, and a mean Kappa coefficient of 0.9917 ± 0.0083 over 10 independent runs. Among the target-only methods, 1D-ResNet achieved the highest mean accuracy (91.60 ± 1.06%), followed by SVM (90.29%) and LSTM (86.18 ± 1.05%). RF and KNN yielded lower accuracies of 85.60 ± 0.46% and 83.27%, respectively. Among the transfer-learning baselines, DAN achieved the highest mean accuracy (88.96 ± 1.05%), slightly outperforming DANN (87.92 ± 1.17%) and CDAN (87.78 ± 1.08%). Compared with the best overall baseline (1D-ResNet), MGDA-Net improved the mean accuracy by 7.71 percentage points. Compared with the best domain adaptation baseline (DAN), the improvement was 10.35 percentage points. The fact that 1D-ResNet outperformed all three standard transfer-learning baselines suggests that single-backbone domain adaptation with only CNN features may not fully exploit the cross-domain structure of the E-nose signals. In contrast, MGDA-Net benefits from multi-granularity feature learning and branch-level adversarial alignment, which together enable more effective knowledge transfer.
Table 3. Comparison of classification performance on the oolong tea dataset.
On the jasmine tea dataset (Table 4), MGDA-Net again achieved the best results across all metrics, reaching a mean OA of 99.38 ± 0.51%, a mean F1-macro of 99.37 ± 0.51%, and a mean Kappa coefficient of 0.9929 ± 0.0058. This task proved substantially more challenging for all baseline methods, consistent with the PCA in Section 4.1. SVM and KNN achieved accuracies of only 82.35% and 75.27%, respectively, while RF yielded 75.88 ± 0.48%. LSTM achieved 75.56 ± 1.72%, the lowest among the deep learning methods. 1D-ResNet again attained the highest target-only accuracy (87.50 ± 1.52%), confirming that residual learning provides a strong inductive bias for E-nose time-series data even without cross-domain knowledge. Among the domain adaptation baselines, DAN reached the highest mean accuracy (86.94 ± 1.64%), followed by CDAN (86.28 ± 2.31%) and DANN (85.73 ± 1.08%). Compared with the best overall baseline (1D-ResNet), MGDA-Net improved the mean accuracy by 11.88 percentage points. Compared with the best domain adaptation baseline (DAN), the improvement was 12.44 percentage points, representing the largest gain observed across both tasks.
Table 4. Comparison of classification performance on the jasmine tea dataset.
To further examine the class-wise prediction behavior of MGDA-Net, the confusion matrices of a representative run on the oolong tea and jasmine tea test sets are shown in Figure 3.
Figure 3. Confusion matrices of MGDA-Net on the target-domain test sets: (a) oolong tea task and (b) jasmine tea task.
In the representative run shown in Figure 3, near-perfect classification performance was observed on both target-domain tasks. On the oolong tea test set, only one misclassification occurred among all samples: one Kuidou Shuixian (Grade I) sample was incorrectly predicted as Xupinhao Rougui, while all remaining samples were correctly classified. On the jasmine tea test set, only one error was also observed, where one CNY 100 sample was misclassified as CNY 400. Apart from these isolated cases, no obvious systematic confusion was found among the remaining categories. These representative confusion matrices illustrate that MGDA-Net not only achieves strong overall performance in terms of accuracy, F1-score, and Kappa, but also maintains reliable class-level discrimination on both datasets.
Across both datasets, the results consistently demonstrate the superiority of MGDA-Net over all baseline methods. The improvement was particularly pronounced on the jasmine tea dataset, where fine-grained retail price-level classification posed a greater challenge for conventional classifiers and standard domain adaptation methods. The strong performance and low standard deviation across runs suggest that jointly modeling local and global feature representations, together with branch-level adversarial alignment and adaptive fusion, is both effective and stable for cross-domain tea classification.

4.3. Ablation Study

To quantify the contribution of each key component in MGDA-Net, a series of ablation experiments was conducted on both target-domain datasets. Seven ablation variants were designed by removing or modifying one specific component of the full model at a time. Because each ablation variant requires retraining a modified architecture, the results in Table 5 are reported from a representative run under the default experimental setting, whereas the robustness of the full model is quantified separately in Section 4.2.
Table 5. Ablation study results of MGDA-Net on the oolong tea and jasmine tea datasets. Results are reported from a representative single run; multi-seed robustness of the full model is presented in Table 3 and Table 4.
Effect of dual-branch feature extraction. To evaluate the necessity of jointly modeling local and global representations, two single-branch variants were considered. CNN-only retained only the CNN branch, while Attn-only retained only the self-attention branch. On the oolong tea dataset, CNN-only and Attn-only achieved accuracies of 87.15% and 82.64%, respectively, both substantially lower than the full MGDA-Net (98.61%). A similar trend was observed on the jasmine tea dataset, where CNN-only and Attn-only achieved 85.42% and 79.17%, respectively, compared with 98.96% for the full model. Across both datasets, CNN-only consistently outperformed Attn-only, indicating that the CNN branch provides stronger discriminative cues when used alone. However, the large performance gap between either single-branch variant and the full model demonstrates that the two branches are complementary, and that jointly exploiting local and global features is essential for achieving optimal performance.
Effect of adaptive fusion. To examine the contribution of the adaptive fusion mechanism, a Concat Fusion variant was implemented by directly concatenating the CNN and self-attention features without learnable reweighting. On the oolong tea dataset, Concat Fusion achieved 84.72% accuracy, which was 13.89 percentage points lower than the full model and even lower than CNN-only (87.15%). On the jasmine tea dataset, Concat Fusion achieved 86.81%, corresponding to a 12.15 percentage point decrease relative to the full model. These results indicate that naive feature concatenation is insufficient for effectively combining heterogeneous branch representations. In contrast, the adaptive fusion module enables dynamic reweighting of local and global features, thereby better exploiting their complementary information.
Effect of multi-granularity domain adversarial (DA) alignment. Three variants were designed to investigate the role of domain adaptation and the benefit of feature alignment at different granularities. The w/o DA variant completely removed the domain adversarial training stage, whereas Local-only DA and Global-only DA retained only the local or global domain discriminator, respectively. On the oolong tea dataset, w/o DA achieved 83.68% accuracy, while Local-only DA and Global-only DA achieved 84.38% and 86.11%, respectively. On the jasmine tea dataset, w/o DA yielded 84.03%, Local-only DA yielded 81.94%, and Global-only DA yielded 84.72%. Compared with the full model, removing domain adversarial learning caused substantial performance degradation on both datasets, confirming the importance of cross-domain feature alignment. Moreover, Global-only DA consistently outperformed Local-only DA, suggesting that aligning global contextual representations plays a more important role than aligning local features alone in the present tasks. Nevertheless, neither single-granularity variant matched the full model, demonstrating that local and global alignment provide complementary benefits. Notably, on the jasmine tea dataset, Local-only DA performed even worse than w/o DA, indicating that local-only alignment is insufficient for this more challenging fine-grained classification task.
Effect of target-domain fine-tuning. The w/o Fine-tuning variant omitted Stage 3 and directly evaluated the model after Stage 2, without target-domain supervised adaptation of the classifier. This variant produced the lowest performance among all ablation settings, with accuracies of 71.18% on the oolong tea dataset and 73.61% on the jasmine tea dataset. This result is expected because the source-domain and target domains differ not only in data distribution but also in class definition and label space. Without target-domain fine-tuning, the classifier cannot be adequately adapted to the target classification task, even if the learned feature representations have been partially aligned across domains. The marked performance drop further highlights the necessity of Stage 3 for achieving accurate target-domain classification.
Overall, the ablation study demonstrates that each component of MGDA-Net makes a meaningful and non-redundant contribution to the final performance. The dual-branch architecture provides complementary local and global representations, the adaptive fusion module improves feature integration, the multi-granularity adversarial alignment effectively reduces domain discrepancy, and the target-domain fine-tuning stage adapts the classifier to the target task. Removing any of these components leads to a clear reduction in performance, thereby confirming the effectiveness of the proposed design.

4.4. Label-Efficiency Analysis

To evaluate the dependence of MGDA-Net on the amount of labeled target-domain data, a series of label-efficiency experiments was conducted. The Stage 1 (source pretraining) and Stage 2 (domain adversarial alignment) models were kept fixed, and only Stage 3 (target fine-tuning) was repeated with progressively reduced proportions of labeled target-domain training data. Specifically, the labeled target-domain training samples were subsampled at ratios of 100%, 80%, 60%, 40%, and 20% using stratified sampling to preserve class balance. For each ratio, 10 independent runs with different random seeds were performed and the results are reported as mean ± standard deviation. The test set (days 9–10) remained unchanged throughout. The results are summarized in Table 6.
Table 6. Label efficiency results of MGDA-Net on the oolong tea and jasmine tea datasets.
Several observations can be drawn from the label-efficiency analysis. First, MGDA-Net exhibits graceful degradation in performance as the proportion of labeled target data decreases. Even when only 40% of the target-domain labels are available (115 samples for oolong tea and 154 samples for jasmine tea), the model still achieves mean accuracies of 88.19% and 87.92%, respectively. These values remain competitive with or superior to several fully supervised baseline methods reported in Table 3 and Table 4. For example, on the oolong tea task, MGDA-Net with 40% labels (88.19%) outperformed LSTM (86.18%), KNN (83.27%), and RF (85.60%), while remaining close to DAN (88.96%). On the jasmine tea task, MGDA-Net with 40% labels (87.92%) slightly outperformed 1D-ResNet (87.50%) and DAN (86.94%).
Second, the performance gap between 100% and 80% labels is relatively small on both tasks, suggesting that the source-domain pretraining and domain-aligned representations already provide a strong initialization, and that the full target-domain training set contains some redundancy for the fine-tuning stage. Third, a more pronounced decline is observed below 40%, particularly for the jasmine tea task, where the standard deviation also increases notably at the 20% level. This indicates that when the number of labeled samples per class becomes very small, the fine-tuning stage becomes more sensitive to the specific samples selected.
These results indicate meaningful label efficiency within the supervised fine-tuning stage of MGDA-Net. At the same time, they should not be interpreted as evidence of unsupervised or few-shot target adaptation, because Stage 3 still relies on labeled target-domain data.

4.5. Feature Alignment and Visualization Analysis

To further investigate the effectiveness of MGDA-Net in reducing cross-domain discrepancy and improving feature organization, a series of quantitative and visual analyses was conducted, including multi-kernel maximum mean discrepancy (MMD) and t-distributed stochastic neighbor embedding (t-SNE). Results on both the oolong tea and jasmine tea tasks are discussed below.

4.5.1. Multi-Kernel MMD Analysis

The domain discrepancy between the source-domain (green tea) and the target-domain training set was quantified using multi-kernel MMD with five Gaussian bandwidths at three feature levels, namely the CNN branch output (local feature), the self-attention branch output (global feature), and the fused representation. The MMD values were measured before domain adversarial training (after Stage 1) and after domain adversarial training (after Stage 2).
For the oolong tea task (Figure 4a), domain adversarial training resulted in a modest reduction in the local feature MMD from 2.261 to 2.157, while the global feature MMD remained nearly unchanged (1.655 vs. 1.652). This suggests that the global branch already exhibited a relatively small cross-domain discrepancy before adaptation in the oolong tea task, leaving limited room for further reduction. In contrast, for the jasmine tea task (Figure 4b), domain adversarial training led to clearer reductions in branch-level MMD values: the local feature MMD decreased from 2.300 to 2.004 (−12.9%), and the global feature MMD decreased from 1.830 to 1.631 (−10.9%). These quantitative results are consistent with reduced source–target discrepancy at the branch level after adversarial training. Across both datasets, the global branch exhibited a smaller domain gap than the local branch both before and after adaptation, which is consistent with the ablation results in Section 4.3 showing that Global-only DA outperformed Local-only DA.
Figure 4. Multi-kernel MMD comparison before and after domain adaptation: (a) oolong tea task and (b) jasmine tea task.
It should also be noted that the fused feature MMD did not decrease after Stage 2 on either dataset. For the oolong tea task, the fused MMD increased from 1.382 to 2.033, whereas for the jasmine tea task, it increased slightly from 1.734 to 1.781. This behavior is reasonable because the adversarial alignment in Stage 2 is imposed directly on the branch-level features rather than on the fused representation itself. Meanwhile, the adaptive fusion module dynamically adjusts branch contributions to improve classification-oriented feature integration, which may not always correspond to a monotonic reduction in the fused-space MMD. The final target-domain classification performance is therefore jointly determined by branch-level alignment and subsequent Stage 3 fine-tuning.

4.5.2. t-SNE Visualization

To provide an intuitive view of the learned feature organization, t-SNE was applied to the fused features of the final model after Stage 3 fine-tuning. Two types of visualizations were considered, namely domain-level embeddings showing source and target samples in a shared space, and class-level embeddings showing only the target-domain test samples colored by class label.
As shown in Figure 5, the source-domain and target-domain samples in both tasks were not completely separated into isolated regions at the domain level. Instead, the target samples were distributed in regions adjacent to portions of the source-domain distribution, suggesting partial cross-domain overlap in the learned feature space rather than complete alignment. At the class level, Figure 6 indicates that the target-domain samples formed more structured class-dependent groupings in both tasks. For the oolong tea dataset, six visually distinguishable regions can be observed, whereas the jasmine tea classes remain more closely distributed, which is consistent with the greater difficulty of this fine-grained task. These plots provide descriptive visual evidence of the learned feature organization, but they should not be interpreted as formal proof of discriminative validity or alignment quality.
Figure 5. t-SNE visualization of source- and target-domain distributions: (a) oolong tea task and (b) jasmine tea task.
Figure 6. t-SNE visualization of target-domain class distributions: (a) oolong tea task and (b) jasmine tea task.
The quantitative and visual analyses presented above provide complementary evidence consistent with the effectiveness of MGDA-Net. The MMD results suggest that domain adversarial learning reduces source–target discrepancy at the branch level, while the t-SNE visualizations indicate improved feature organization in the projected two-dimensional space. It should be noted that PCA and t-SNE are descriptive dimensionality reduction tools rather than formal tests of discriminative validity or alignment quality; their visualizations should therefore be interpreted as supportive illustrations rather than definitive proof.

5. Conclusions

This study proposed MGDA-Net, a multi-granularity domain adversarial framework for cross-domain tea classification using electronic nose signals. By jointly modeling local temporal dynamics and global contextual dependencies and performing branch-level adversarial alignment, MGDA-Net enhances feature transferability from a labeled green tea source-domain to target domains, followed by target-domain fine-tuning to adapt the decision boundary to the target label space.
Comprehensive experiments on two target-domain tasks demonstrated the effectiveness and robustness of MGDA-Net. On the oolong tea dataset (six commercial categories), MGDA-Net achieved a mean accuracy of 99.31 ± 0.69% and a mean Kappa coefficient of 0.9917 ± 0.0083 over 10 independent runs. On the jasmine tea dataset (eight retail price levels), MGDA-Net achieved a mean accuracy of 99.38 ± 0.51% and a mean Kappa coefficient of 0.9929 ± 0.0058. Compared with the respective best overall baseline method (1D-ResNet), MGDA-Net improved the mean accuracy by 7.71 and 11.88 percentage points on the two tasks, respectively. Compared with the best domain adaptation baseline (DAN), the corresponding gains were 10.35 and 12.44 percentage points.
Ablation experiments verified that the dual-branch architecture, adaptive fusion, multi-granularity adversarial alignment, and target-domain fine-tuning each contribute non-redundantly to the overall performance. Label-efficiency experiments further showed that MGDA-Net maintains mean accuracies above 87% even when only 40% of the target-domain labels are used for fine-tuning, indicating that the source-domain knowledge and domain-aligned representations reduce—but do not eliminate—the dependence on labeled target-domain data. Additional distribution and visualization analyses (multi-kernel MMD and t-SNE) provided complementary evidence that the proposed adversarial training improves cross-domain feature consistency.
Several limitations should be acknowledged. First, the target-domain labels used in this study are commercial categories and retail price levels rather than independently validated quality grades based on sensory panel assessment or chemical quality markers. The results therefore demonstrate classification of these specific commercial label sets, and further validation with established quality standards would be needed to extend the conclusions to objective tea quality grading. Second, the method was validated on two target-domain scenarios using a single E-nose instrument type under controlled laboratory conditions; its generalizability to broader tea categories, different instruments, and real-world environmental variations (e.g., temperature and humidity fluctuations) remains to be evaluated. Third, although the label-efficiency experiments show reasonable performance at reduced annotation levels, the current fine-tuning stage still relies on a moderate amount of labeled target data; investigating semi-supervised or few-shot strategies to further reduce labeling requirements is a meaningful direction for future research.
Future work will extend MGDA-Net to additional food and beverage classification tasks, incorporate label-efficient learning strategies, explore model interpretability techniques such as SHapley Additive exPlanations (SHAP) to relate learned features to specific sensor channels and volatile compound profiles, and assess performance under real-world deployment settings with portable electronic nose devices.

Author Contributions

Conceptualization, X.W. and Y.G.; methodology, X.W.; software, X.W.; investigation, X.W.; visualization, X.W.; writing—original draft preparation, X.W.; writing—review and editing, X.W. and Y.G.; supervision, Y.G.; project administration, Y.G.; funding acquisition, Y.G. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China [Grant No. 62276174], the National Key Research and Development Program of China [Grant No. 2023YFC3604200], the Key Research and Development Program of Xinjiang Uygur Autonomous Region (2024B03033), and the Beijing Municipal Natural Science Foundation [Grant No. Z22015].

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The code, implementation details, and processed data that support the findings of this study are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
E-noseElectronic nose
CNNConvolutional neural network
MLPMultilayer perceptron
GRLGradient reversal layer
DADomain adaptation
MMDMaximum mean discrepancy
t-SNEt-distributed stochastic neighbor embedding
PCAPrincipal component analysis
SVMSupport vector machine
RFRandom forest
KNNK-nearest neighbor
LSTMLong short-term memory
1D-ResNetOne-dimensional residual network
DANDeep adaptation network
DANNDomain adversarial neural network
CDANConditional domain adversarial network
MOSMetal oxide semiconductor
OAOverall accuracy
BNBatch normalization
ReLURectified linear unit
SHAPSHapley Additive exPlanations

References

  1. Moreira, J.; Aryal, J.; Guidry, L.; Adhikari, A.; Chen, Y.; Sriwattana, S.; Prinyawiwatkul, W. Tea quality: An overview of the analytical methods and sensory analyses used in the most recent studies. Foods 2024, 13, 3580. [Google Scholar] [CrossRef] [Scilit]
  2. Chen, Z.; Sui, Y.; Wisniewski, M. Current and future perspectives on tea production. Ind. Crops Prod. 2025, 235, 121663. [Google Scholar] [CrossRef] [Scilit]
  3. National Bureau of Statistics of China. Statistical Bulletin of the People’s Republic of China on the 2024 National Economic and Social Development. 2025. Available online: https://www.stats.gov.cn/sj/zxfb/202502/t20250228_1958817.html (accessed on 10 March 2026).
  4. Food and Agriculture Organization of the United Nations. International Tea Day 2023: Supporting Smallholder Tea Producers Is an Integral Part in the Transformation of Agrifood Systems. 2023. Available online: https://www.fao.org/newsroom/detail/international-tea-day-2023--supporting-smallholder-tea-producers-is-an-integral-part-in-the-transformation-of-agrifood-systems/en (accessed on 10 March 2026).
  5. Li, Y.; Elliott, C.T.; Petchkongkaew, A.; Wu, D. The classification, detection and ‘SMART’control of the nine sins of tea fraud. Trends Food Sci. Technol. 2024, 149, 104565. [Google Scholar] [CrossRef] [Scilit]
  6. Tan, H.R.; Zhou, W. Metabolomics for tea authentication and fraud detection: Recent applications and future directions. Trends Food Sci. Technol. 2024, 149, 104558. [Google Scholar] [CrossRef] [Scilit]
  7. Fernandes, J.S.; de Sousa Fernandes, D.D.; Pistonesi, M.F.; Diniz, P.H.G.D. Tea authentication and determination of chemical constituents using digital image-based fingerprint signatures and chemometrics. Food Chem. 2023, 421, 136164. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Chen, Y.; Guo, M.; Chen, K.; Jiang, X.; Ding, Z.; Zhang, H.; Lu, M.; Qi, D.; Dong, C. Predictive models for sensory score and physicochemical composition of Yuezhou Longjing tea using near-infrared spectroscopy and data fusion. Talanta 2024, 273, 125892. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. He, Y.; Zhang, Q.; Inostroza, A.C.; Kierszniowska, S.; Liu, L.; Li, Y.; Ruan, J. Application of metabolic fingerprinting in tea quality evaluation. Food Control 2024, 160, 110361. [Google Scholar] [CrossRef] [Scilit]
  10. Liu, Y.; Luo, L.; Zeng, L. Guidelines for Sensory Evaluation of Tea: Traditional Chinese Method and Quantitative Descriptive Analysis. AgriFood 2025, 1, 15–20. [Google Scholar] [CrossRef] [Scilit]
  11. Sipos, L.; Nyitrai, Á.; Hitka, G.; Friedrich, L.F.; Kókai, Z. Sensory panel performance evaluation—Comprehensive review of practical approaches. Appl. Sci. 2021, 11, 11977. [Google Scholar] [CrossRef] [Scilit]
  12. Zan, J.; Chen, W.; Yuan, H.; Jiang, Y.; Zhu, H. Evaluation of key taste components in Huangjin green tea based on electronic tongue technology. Food Res. Int. 2025, 201, 115569. [Google Scholar] [CrossRef] [Scilit]
  13. Liu, N.; Shen, S.; Huang, L.; Deng, G.; Wei, Y.; Ning, J.; Wang, Y. Revelation of volatile contributions in green teas with different aroma types by GC–MS and GC–IMS. Food Res. Int. 2023, 169, 112845. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Chen, Y.; Gao, R.; Xiang, J.; Lin, Z.; Chen, C.; Yu, W. Progress of near infrared spectroscopy in tea research. Food Chem. 2025, 487, 144691. [Google Scholar] [CrossRef] [Scilit]
  15. Hu, Y.; Yu, H.; Song, X.; Chen, W.; Ding, L.; Chen, J.; Liu, Z.; Guo, Y.; Xu, D.; Zhu, X.; et al. Comprehensive assessment of matcha qualities and visualization of constituents using hyperspectral imaging technology. Food Res. Int. 2024, 196, 115110. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Medina-García, M.; Amigo, J.M.; Martínez-Domingo, M.A.; Valero, E.M.; Jiménez-Carvelo, A.M. Strategies for analysing hyperspectral imaging data for food quality and safety issues–A critical review of the last 5 years. Microchem. J. 2025, 214, 113994. [Google Scholar] [CrossRef] [Scilit]
  17. Li, Z.; Gao, Z.; Yu, J.; Shi, H.; Ling, J.; Zhang, G. Applications of E-nose, GC-MS, and GC-IMS in tea volatile components analysis. J. Food Compos. Anal. 2025, 149, 108764. [Google Scholar] [CrossRef] [Scilit]
  18. Sanislav, T.; Mois, G.D.; Zeadally, S.; Folea, S.; Radoni, T.C.; Al-Suhaimi, E.A. A comprehensive review on sensor-based electronic nose for food quality and safety. Sensors 2025, 25, 4437. [Google Scholar] [CrossRef] [Scilit]
  19. Kang, S.; Zhang, Q.; Li, Z.; Yin, C.; Feng, N.; Shi, Y. Determination of the quality of tea from different picking periods: An adaptive pooling attention mechanism coupled with an electronic nose. Postharvest Biol. Technol. 2023, 197, 112214. [Google Scholar] [CrossRef] [Scilit]
  20. Yu, D.; Gu, Y. A machine learning method for the fine-grained classification of green tea with geographical indication using a MOS-based electronic nose. Foods 2021, 10, 795. [Google Scholar] [CrossRef] [Scilit]
  21. Chang, J.; Lu, A. RLCSA-Net: A new deep learning method combined with an electronic nose to identify the quality of tea. J. Food Meas. Charact. 2024, 18, 5222–5231. [Google Scholar] [CrossRef] [Scilit]
  22. Kombo, K.O.; Ihsan, N.; Syahputra, T.S.; Hidayat, S.N.; Puspita, M.; Roto, R.; Triyana, K. Enhancing classification rate of electronic nose system and piecewise feature extraction method to classify black tea with superior quality. Sci. Afr. 2024, 24, e02153. [Google Scholar] [CrossRef] [Scilit]
  23. Kaushal, S.; Rana, P.; Chung, C.C.; Chen, H.H. Geographical Origin Classification of Oolong Tea Using an Electronic Nose: Application of Machine Learning and Gray Relational Analysis. Chemosensors 2025, 13, 295. [Google Scholar] [CrossRef] [Scilit]
  24. Vanaraj, R.; IP, B.; Mayakrishnan, G.; Kim, I.S.; Kim, S.C. A systematic review of the applications of electronic nose and electronic tongue in food quality assessment and safety. Chemosensors 2025, 13, 161. [Google Scholar] [CrossRef] [Scilit]
  25. Cai, M.; Xu, S.; Zhou, X.; Lu, H. Electronic nose humidity compensation system based on rapid detection. Sensors 2024, 24, 5881. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Pang, X.; Yan, Z.; Zou, J.; Miao, P.; Cheng, W.; Zhou, Z.; Ye, J.; Wang, H.; Jia, X.; Li, Y.; et al. Combined analysis of grade differences in lapsang souchong black tea using sensory evaluation, electronic nose, and HS-SPME-GC-MS, based on Chinese national standards. Foods 2024, 13, 3433. [Google Scholar] [CrossRef] [Scilit]
  27. Wang, C.; Wang, Z.; Liu, Q.; Dong, H.; Liu, W.; Liu, X. A comprehensive survey on domain adaptation for intelligent fault diagnosis. Knowl.-Based Syst. 2025, 327, 114109. [Google Scholar] [CrossRef] [Scilit]
  28. Wu, T.; Chen, Q.; Zhao, D.; Wang, J.; Jiang, L. Domain adaptation of time series via contrastive learning with task-specific consistency. Appl. Intell. 2024, 54, 12576–12588. [Google Scholar] [CrossRef] [Scilit]
  29. Yan, J.; Chen, F.; Liu, T.; Zhang, Y.; Peng, X.; Yi, D.; Duan, S. Subspace alignment based on an extreme learning machine for electronic nose drift compensation. Knowl.-Based Syst. 2022, 235, 107664. [Google Scholar] [CrossRef] [Scilit]
  30. Wu, Z.; Tian, F.; Covington, J.A.; Li, H.; Deng, S.; Liu, Y.; Li, J. Triplet calibration for general-purpose electronic noses. Microchem. J. 2025, 209, 112652. [Google Scholar] [CrossRef] [Scilit]
  31. Zhu, H.; Wu, Y.; Yang, G.; Song, R.; Yu, J.; Zhang, J. Electronic nose drift suppression based on smooth conditional domain adversarial networks. Sensors 2024, 24, 1319. [Google Scholar] [CrossRef] [Scilit]
  32. Jiang, Q.; Zhang, Y.; Zhang, Y.; Liu, J.; Xu, M.; Ma, C.; Jia, P. An adversarial network used for drift correction in electronic nose. Sens. Actuators A Phys. 2024, 377, 115720. [Google Scholar] [CrossRef] [Scilit]
  33. Sun, J.; Zheng, H.; Diao, W.; Sun, Z.; Qi, Z.; Wang, X. Prototype-Optimized unsupervised domain adaptation via dynamic Transformer encoder for sensor drift compensation in electronic nose systems. Expert Syst. Appl. 2025, 260, 125444. [Google Scholar] [CrossRef] [Scilit]
  34. Jiang, M.; Li, N.; Li, M.; Wang, Z.; Tian, Y.; Peng, K.; Sheng, H.; Li, H.; Li, Q. E-nose: Time–frequency attention convolutional neural network for gas classification and concentration prediction. Sensors 2024, 24, 4126. [Google Scholar]
  35. Sun, J.; Dong, M.; Lu, F.; Wang, Z.; Zhu, J.; Wang, Y.; Zheng, L.; Zhao, F. Transformer-driven multi-scale contextual alignment for robust E-nose drift compensation. Knowl.-Based Syst. 2025, 327, 114136. [Google Scholar] [CrossRef] [Scilit]
  36. Ren, L.; Cheng, G.; Chen, W.; Li, P.; Wang, Z. Advances in drift compensation algorithms for electronic nose technology. Sens. Rev. 2024, 44, 733–745. [Google Scholar] [CrossRef] [Scilit]
  37. Sun, F.; Sun, R.; Yan, J. Cross-domain active learning for electronic nose drift compensation. Micromachines 2022, 13, 1260. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Wen, J.; Yuan, J.; Zheng, Q.; Liu, R.; Gong, Z.; Zheng, N. Hierarchical domain adaptation with local feature patterns. Pattern Recognit. 2022, 124, 108445. [Google Scholar] [CrossRef] [Scilit]
  39. Zhang, J.Y.; Liu, X.Q.; Xue, Z.Y.; Luo, X.; Xu, X.S. MAGIC: Multi-granularity domain adaptation for text recognition. Pattern Recognit. 2025, 161, 111229. [Google Scholar] [CrossRef] [Scilit]
  40. He, C.; Zhou, J.; Li, Y.; Ntezimana, B.; Zhu, J.; Wang, X.; Xu, W.; Wen, X.; Chen, Y.; Yu, Z.; et al. The aroma characteristics of oolong tea are jointly determined by processing mode and tea cultivars. Food Chem. X 2023, 18, 100730. [Google Scholar] [CrossRef] [Scilit]
  41. Zhang, S.; Xia, H.; Sun, S.; Ma, C.; Li, F.; Pan, S.; Zhang, Z.; Lai, X.; Li, Q.; Hao, M.; et al. Analysis of volatile metabolites and differences in aroma quality of green tea, black tea and oolong tea prepared from highly aromatic tea variety (Dancong). Food Res. Int. 2025, 221, 117402. [Google Scholar] [CrossRef] [Scilit]
  42. Zhang, Y.; Xiong, Y.; An, H.; Li, J.; Li, Q.; Huang, J.; Liu, Z. Analysis of volatile components of jasmine and jasmine tea during scenting process. Molecules 2022, 27, 479. [Google Scholar] [CrossRef] [Scilit]
  43. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
  44. Ganin, Y.; Ustinova, E.; Ajakan, H.; Germain, P.; Larochelle, H.; Laviolette, F.; March, M.; Lempitsky, V. Domain-adversarial training of neural networks. J. Mach. Learn. Res. 2016, 17, 1–35. [Google Scholar]
  45. Wen, Y.; Zhang, K.; Li, Z.; Qiao, Y. A discriminative feature learning approach for deep face recognition. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2016; pp. 499–515. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.