1. Introduction
Grain crops are fundamental to global food security. In China, the world’s most populous country, reducing paddy rice harvest losses by just 1% could save approximately 10 billion kilograms of grain annually—enough to feed 20 million people [
1]. This imperative for minimizing harvest loss extends beyond cereals: potato, a staple food for over 70 countries including China, India, and the United States, is equally vital for both food supply and industrial processing [
2]. Combine harvesters, by integrating cutting, threshing, separating, and cleaning into a single pass, are the core equipment for ensuring timely harvest completion and minimizing field losses [
3]. The automation and intelligent control of these machines further reduce operator labor intensity and help alleviate agricultural labor shortages [
4]. However, the harsh and variable field conditions under which combine harvesters operate pose significant challenges to their structural reliability, making fault diagnosis a critical enabler of harvest efficiency and food supply resilience.
Combine harvesters typically operate for no more than two months per year [
5]; however, during the harvest season, harvester operators often operate the combine for more than 12 h per working day [
6]. Harvesting is a highly time-sensitive operation; failing to complete the harvest within the narrow window of the optimal harvest timing will result in significant “timeliness loss” [
7]. Within this harvest window, any unplanned downtime will lead to economic losses. When insufficient technical service support causes delays in the supply of spare parts, the downtime of the harvester can account for 10–50% of its total usage time [
5]. Moreover, the skill level of maintenance service providers directly affects the efficiency of fault recovery; targeted training enhances the responsiveness of agricultural socialized services [
8]. The economic consequence of harvest loss is substantial: in China, the average wheat harvest loss rate reaches 2.43% [
9], underscoring the urgent need for real-time monitoring of machine performance during harvesting operations. When the harvest time exceeds the optimal window, each hour of downtime increases the grain loss rate by 0.004% to 0.006% [
10]. Beyond harvest delays, structural issues during threshing can also cause internal grain damage that degrades crop quality [
11]. In addition to crop losses, the materials and energy consumed for repair and component replacement further exacerbate the economic losses caused by failures [
12].
During rice harvesting, excessive stalk breakage during threshing not only increases energy consumption but also elevates the volume of small straw fragments entering the cleaning system, thereby increasing the cleaning burden and reducing overall efficiency [
13]. Crop physical characteristics further complicate machine design. For thick-stalked crops such as industrial hemp, the tall, rigid stems tend to break and twist during clamping and conveying, causing inconsistent laying angles that hinder subsequent field operations [
14]. Crops with unconventional physical characteristics require dedicated machine designs to achieve viable mechanized harvesting. For example, tiger nut (Cyperus esculentus) has irregular seed morphology, an uneven surface, and an entangled root–soil–tuber matrix that render conventional root crop harvesters ineffective [
15].
However, the complex kinematics and structure of combine harvesters, their high-intensity use, and the harsh and variable operating environments [
16,
17] pose technical challenges for Fault Detection and Diagnosis (FDD). In this review, we use fault diagnosis interchangeably with FDD to refer specifically to the identification and classification of structural and mechanical faults in combine harvesters through sensor-based monitoring and machine learning (ML) methods. We use structural health monitoring (SHM), a broader term encompassing continuous condition assessment of civil and mechanical structures, only when discussing general methodological frameworks that extend beyond agricultural machinery. Structural failures are one of the major factors causing unplanned downtime and impairing harvesting performance.
Artificial intelligence has increasingly permeated agricultural production, with multi-source sensor fusion, deep learning (DL), and adaptive control enabling advances from remote sensing to real-time field perception [
18]. In crop monitoring, Unmanned Aerial Vehicle-based hyperspectral imaging combined with ML now permits non-destructive estimation of water status, while Light Detection and Ranging-based 3D point clouds map height distributions in rice [
19], and unmanned aerial vehicle-based structure-from-motion techniques have enabled maize plant height estimation from RGB imagery [
20]. In real-time field perception, vision-based DL measures the lateral deviation from the grain divider to the harvesting boundary via inverse perspective mapping [
21] and detects all-day tea shoots under both natural and artificial lighting using a lightweight YOLOv4-based model [
22]. As a single-stage real-time object detector, the YOLO architecture predicts bounding boxes and class probabilities in a single forward pass, making it well-suited for time-sensitive agricultural perception tasks. Across crops and production stages, fused deep networks further classify nursery tree species and segment crowns and trunks from point clouds [
23], recognize rice false smut offline under variable illumination [
24], and detect branch-infected mulberries in aeroponic cultivation using hybrid convolutional neural network-gated recurrent unit (CNN-GRU) architectures [
25].
While this review draws on international literature to provide comprehensive methodological coverage, this review anchors its analytical focus in the technological and operational realities most relevant to the Chinese agricultural context—the world’s largest combine harvester market, where medium-to-large wheeled and tracked machines operate across diverse cropping systems spanning from the North China Plain to the hilly southwest. Regional differences in farm scale, crop varieties, field conditions, and maintenance infrastructure—within China and globally—can significantly influence the applicability and prioritization of diagnostic approaches. For example, harvesters operating in the small, fragmented, high-moisture paddy fields of southern China—characterized by heavy clay soils and limited bearing capacity [
26]—face fundamentally different fault mode distributions and sensor installation constraints than large machines working extensive wheat or corn fields in the northeast. High-moisture rice, in particular, is prone to forming clumps that clog conveying, threshing, and cleaning devices [
27], a challenge far less pronounced in dryland crop harvesting. Similarly, crop-specific harvesting challenges impose distinct mechanical loads and failure patterns: the thin, fragile seed coat of sunflower seeds makes them highly susceptible to mechanical damage during conveying [
28], while the hilly and mountainous terrain prevalent in parts of China’s maize-growing regions demands specialized low-center-of-gravity harvester designs to maintain stability and reduce grain loss [
29]. Where relevant, the review notes these regional and crop-specific factors throughout, but a comprehensive treatment of geographic, climatic, and socio-economic determinants of harvester reliability lies beyond its scope and warrants dedicated comparative investigation.
This trend places unprecedented demands on the reliability of combine harvesters, as any unplanned downtime directly interrupts the automated precision operation chain. Developing advanced fault diagnosis technologies that match these demands is therefore essential for both harvest efficiency and the broader goal of alleviating agricultural labor shortages.
During field operations, agricultural machinery endures harsh environmental stressors—high dust concentrations, rapid humidity swings, broadband mechanical vibrations, and transient impact loads—which collectively induce continuous, non-periodic fatigue in structural components, accelerating cumulative damage and performance degradation [
30,
31] while also compromising electronic reliability. Common structural failure modes span five categories: (i) wear and fatigue fracture of transmission components, with the power train being the most frequently failing system in combine harvesters [
32]; (ii) loosening of critical fasteners under vibratory loads, a dominant failure mode at connection points of working parts such as the vibrating screen [
33]; (iii) abrasive wear and surface degradation of grain-contacting parts, which the silica-rich rice–steel friction pair drives [
34]; (iv) fatigue failure of reciprocating components under prolonged alternating loads, for example rubber bearing wear, bearing housing fracture, and weld failure in the cleaning sieve [
35]; and (v) blockage resulting from abnormal material accumulation, including straw winding on rotary tillage components in high-stubble clay soils [
36] and vine clogging in peanut clamping–conveying systems [
37].
In 2023, the Department of Agricultural Mechanization Management of China’s Ministry of Agriculture and Rural Affairs conducted a quality survey on corn harvesters (grain harvesting type) in nine major corn-producing regions. The results are shown in
Figure 1. During operation, 22.98% of the surveyed machines experienced a total of 298 failures. Among these, failures of the threshing unit, header system, and transmission system accounted for the dominant share (
Figure 2), indicating a high degree of concentration at these three failure sites.
Traditional methods rely on manual inspection and empirical judgment, which offer limited efficiency and accuracy. The structural complexity of combine harvesters, variable working conditions, scarcity of fault samples, non-stationary signals, and strong background noise all pose significant challenges to FDD. In manned operations, fault identification depends primarily on operators’ auditory, tactile, and visual experience—detecting abnormal noise, vibration, or sudden increases in resistance. Recent advances in sensor technology, data science, and computational power have driven fault diagnosis toward higher levels of intelligence.
2. Review Methodology
2.1. Search Strategy
The literature search for this review was conducted across six major academic databases: Web of Science Core Collection, Scopus, IEEE Xplore, ScienceDirect, SpringerLink, and Google Scholar. The search aimed to identify peer-reviewed journal articles and conference papers addressing fault diagnosis, condition monitoring (CM), and structural health management of combine harvesters. The search strategy combined terms related to the target machinery with terms describing diagnostic methodologies and enabling technologies. To reduce the risk of omission, the reference lists of all retrieved papers and relevant review articles were manually examined for additional studies that may not have been captured by the initial database queries. All retrieved records were subsequently deduplicated and consolidated.
The search period was set from January 1980 to May 2026, with the final update conducted in April 2026. This time frame captures the major developments in intelligent fault diagnosis and sensor-based CM of agricultural machinery, from early statistical reliability analysis and expert systems in the 1980s–1990s to the DL and multi-sensor fusion approaches of the past decade.
2.2. Inclusion and Exclusion Criteria
Studies were selected according to predefined inclusion criteria: (i) the work focused on fault diagnosis, failure detection, or health monitoring of structural components in combine harvesters or comparable agricultural machinery, including headers, threshing drums, cleaning sieves, transmission systems, chassis frames, and bearings; (ii) sensor-based data acquisition was employed, encompassing vibration, acoustic, speed, temperature, strain, or visual signals; (iii) the methodology involved ML, DL, multi-sensor fusion, or advanced signal processing; and (iv) the work was published as a peer-reviewed journal article or conference paper in English or Chinese.
Studies were excluded if they: (i) addressed purely agronomic perception tasks without diagnostic intent; (ii) focused exclusively on structural optimization or durability design without a CM component; (iii) were non-peer-reviewed preprints, dissertations, or patents; (iv) represented duplicate publications of the same work; or (v) lacked full-text accessibility.
2.3. Screening Process
The study selection followed a four-stage procedure based on the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 framework. First, we merged records retrieved from different databases and removed duplicate entries. Second, we screened titles and abstracts to exclude studies clearly irrelevant to the structural fault diagnosis of combine harvesters. Third, we assessed the full texts of the remaining studies and retained only those that explicitly addressed the target diagnostic tasks and provided sufficient methodological detail or experimental validation. Finally, we conducted backward reference tracking on the included studies to supplement the dataset and further reduce the risk of missing relevant publications.
2.4. Analytical Framework and Limitations
For the synthesis and structured comparison of the selected literature, we established a three-level taxonomy. Studies were grouped by the physical subsystem under diagnosis (structural module level;
Figure 3), by methodological approach spanning traditional ML, DL, and multi-sensor fusion (methodological level), and by practical challenges constraining field deployment (deployment-readiness level).
Notably, we did not perform a formal quality assessment of individual studies. This review prioritizes methodological coverage and technological evolution over quantitative evidence synthesis; accordingly, we treated all peer-reviewed publications as equally credible sources of methodological insight, which may introduce bias. Furthermore, the literature surveyed predominantly reports successful diagnostic outcomes with high accuracy or F1 scores, likely reflecting publication bias—studies with negative or inconclusive results are less likely to be submitted—and an editorial preference for novel algorithmic contributions over replication studies. Both mechanisms skew the published literature toward optimistic performance estimates, and the diagnostic performance figures compiled in this review may therefore overestimate the real-world effectiveness of certain methods. Future systematic reviews in this field would benefit from including grey literature and preregistered studies to mitigate these biases.
First, the scope of this review is constrained to structural and mechanical fault diagnosis of combine harvesters; related topics—hydraulic and pneumatic system diagnostics, risk analysis, onboard computer architectures, communication networks, satellite positioning, and anthropogenic factors—are beyond its methodological focus and warrant dedicated reviews. Second, the literature search was limited to publications in English and Chinese, which may introduce language bias by omitting relevant work in Russian, German, and Japanese. Third, non-peer-reviewed sources such as preprints, dissertations, patents, and manufacturer technical reports were excluded by design, meaning that practical diagnostic solutions from industry may be underrepresented. Fourth, the literature search was conducted with a cutoff date of May 2026; any relevant studies published after this date are not included. The coverage of this review is therefore systematic but not exhaustive.
5. Deep Learning and Multi-Sensor Fusion-Based Intelligent Diagnosis Frontiers
The previous chapter detailed the inherent limitations of traditional ML diagnostic methods: strong dependence on handcrafted feature engineering, limited ability to model complex nonlinear relationships, and poor adaptability to time-varying operating conditions. Consequently, they cannot independently perform whole-machine-level intelligent diagnosis of combine harvesters in real field environments. Recent advances in DL have opened new pathways to break through these bottlenecks, and DL has gradually become a mainstream research method in agriculture [
157,
158,
159,
160]. The joint deployment of multimodal sensors and the deep fusion of heterogeneous information further enable comprehensive sensing of the multi-component coupled operating state of combine harvesters. Industry-specific DL methods are becoming standard for fault diagnosis [
161,
162,
163]. This chapter sequentially introduces the general paradigm of DL fault diagnosis, reviews the specific applications of deep models in vibration, acoustic, and visual signals, discusses the evolution of multi-source signal fusion strategies, and systematically compares them with traditional methods.
5.1. General Paradigm of Deep Learning Fault Diagnosis
Unlike the traditional pipeline of “signal preprocessing → handcrafted feature extraction → feature dimensionality reduction → shallow classification,” deep diagnostic models form an end-to-end trainable system: raw time-domain signals serve as the network input, and through serialized nonlinear transformations across multiple hidden layers, fault features are progressively abstracted from low-level textures to high-level semantics, with the output layer directly yielding fault category or health indicator predictions. This paradigm jointly optimizes feature extraction and classification, eliminating the need to predefine sensitive frequency bands or statistical quantities based on specific fault mechanisms [
164].
Among typical network structural components, convolutional layers, with their local receptive fields and weight sharing, excel at capturing local impulse patterns in vibration signals and spatial textures in time–frequency images. Pooling layers reduce dimensionality and enhance translation invariance while also suppressing noise. Recurrent layers (e.g., LSTM, GRU) and their variants, with memory and forgetting mechanisms for sequential states, naturally handle time-dependent degradation processes caused by speed fluctuations and gradual load changes. Autoencoders learn low-dimensional representations of data manifolds via reconstruction error and are particularly valuable in unsupervised pre-training and anomaly detection. Attention mechanisms can adaptively weight different channels, time steps, or sensor nodes, enabling the model to focus on the signal components with the strongest fault discriminability. These components can be combined according to task requirements into convolutional neural network (CNN), residual network (ResNet), vision transformer (ViT), or hybrid architectures to handle different forms of input.
Training strategies for deep diagnostic models also vary significantly. Supervised learning relies on complete fault category labels, and most current combine harvester fault diagnosis studies adopt this approach. When labels are scarce, unsupervised pre-training can leverage large amounts of unlabeled operating data to initialize network parameters, followed by supervised fine-tuning with few labeled samples. Transfer learning can alleviate the scarcity of fault samples in the target domain (e.g., field combine harvesters) by reusing model weights pre-trained on source–domain data (e.g., laboratory bench data or similar rotating machinery). In addition, advanced training paradigms such as meta-learning and domain adaptation [
165,
166] are being introduced to enhance the model’s small-sample adaptability under variable operating conditions.
The choice of input data form directly influences network architecture design. In current combine harvester research, the most common forms include: (1) one-dimensional raw vibration/acoustic signals fed directly into 1D-CNN or LSTM networks; (2) two-dimensional time–frequency images, where STFT or CWT maps the signal to the time–frequency plane for feature extraction by 2D-CNN or ViT; and (3) multi-channel fusion input, where data from different sensors or modalities are concatenated along the channel dimension and fed into a multi-branch network. The following sections detail the application of these architectures to specific signal modalities in combine harvester fault diagnosis.
5.2. Application of Deep Neural Networks in Vibration, Acoustic, and Visual Signal-Based Diagnosis
5.2.1. Deep Diagnostic Models Based on Vibration Signals
As discussed in
Section 4.2.1, vibration signals are the primary monitoring modality for combine harvester fault diagnosis, and DL models have been extensively applied across a range of vibration-based architectures. CNNs were among the earliest: the 1D-CNN approach demonstrated real-time structural damage detection directly from raw time-domain acceleration signals without handcrafted feature extraction [
167]. In parallel, 2D-CNNs emerged by integrating CWT scalograms as inputs to simultaneously capture temporal and spectral fault features directly from the full wavelet coefficient matrix [
168].
Autoencoders and their variants have shown unique advantages in whole-machine multi-component CM. Stacked denoising autoencoders (SDAE) inject random noise into the input and train the network to reconstruct noise-free signals, forcing learning of disturbance-insensitive features. An SDAE-SVM hybrid architecture using deep features from multi-component speed signals achieved 95.31% accuracy, significantly outperforming standard SVM (77.03%) and BP network (74.61%) under strong noise [
169]. Similarly, the SDAE-BP model took the rotational speeds of six key components—feeding auger, tailings auger, grain auger, fan, threshing drum, and conveyor chain—as input, using a layer-wise DAE strategy with differentiated Gaussian noise centers to integrate local and global information for whole-machine state identification, achieving 99.00% accuracy [
170].
The combination of ResNet and ViT represents a more recent advance. A diagnostic method integrating SVD-EDS denoising, generalized S-transform time–frequency imaging, and a ResViT model—using ResNet-34 for feature pre-extraction and an improved ViT to integrate global and local information—achieved 99.08% average accuracy for rolling bearing diagnosis under dust and varying loads, significantly outperforming pure CNN, LSTM, and the original ViT.
Gated recurrent architectures, particularly LSTM and Bidirectional LSTM, capture long-term temporal dependencies in fault evolution under complex excitations such as feed rate fluctuations, ground bumps, and strong background noise. For example, a deep feature extraction module integrating a multi-scale 1D-CNN with multi-head attention and an adaptive soft-thresholding function was combined with a Bidirectional LSTM classifier to achieve over 98% accuracy for transmission system defect identification [
171].
5.2.2. Deep Diagnostic Methods Based on Acoustic Signals
Acoustic signals offer non-contact acquisition, providing a supplementary sensing channel for combine harvester fault diagnosis. Unlike traditional methods relying on handcrafted features such as MFCC, DL can directly learn high-level representations from raw acoustic waveforms or spectrograms, though strong field background noise demands higher model robustness. Cross-modal joint learning of acoustic and vibration signals has become the mainstream approach. For example, a multi-channel multi-scale spatiotemporal convolutional cross-attention fusion network converts synchronously collected vibration and acoustic signals into time–frequency images, and a dual cross-attention mechanism adaptively fuses inter-modal information, outperforming single-modality approaches in bearing fault diagnosis [
172]. Nevertheless, DL studies targeting whole combine harvesters using only acoustic signals remain rare.
5.2.3. Diagnostic Exploration Based on Visual Signals
Visual signals—visible light images and infrared thermal images—are still at an early exploratory stage for combine harvester fault diagnosis. Visible light images can intuitively reflect faults with obvious appearance characteristics such as header blockage, belt slippage, and material accumulation, while infrared thermal imaging is sensitive to thermophysical features such as abnormal temperature rise and local overheating before bearing failure. Some studies have employed USB cameras for surveillance video of working areas and CNNs or object detection models for abnormal state identification [
105]. Infrared thermal images have also been used for bearing fault classification, with CNNs learning temperature distribution patterns corresponding to fault modes in an end-to-end manner [
106,
107,
108].
However, field conditions pose significant challenges: high dust reduces visibility, mud occludes targets, lighting varies drastically, and early internal faults such as bearing micro-pitting cannot be captured externally. These factors constrain the reliability of visual diagnosis as an independent modality.
Beyond environmental challenges, visual diagnosis also faces fine-grained discrimination difficulty—fault types often exhibit high inter-class similarity and intra-class variance under different conditions, a problem well-recognized in crop disease diagnosis where multi-granularity feature aggregation and self-attention mechanisms have been applied. These approaches provide useful references for combine harvester visual fault diagnosis.
Although visible light images can intuitively reveal faults with obvious visual manifestations such as header blockage and material accumulation, mature vision-based object detection in agriculture still focuses mainly on agronomic processes—crop monitoring, weed identification, and pest and disease diagnosis [
173]. These efforts have accumulated practical experience with field visual challenges, providing a methodological basis for transferring object detection frameworks to abnormal state identification on combine harvesters. However, early internal faults such as bearing micro-pitting and gear micro-cracks are difficult to capture externally, restricting independent pure visual diagnosis. Currently, visual signals serve primarily as auxiliary verification for vibration- or acoustic-based diagnosis.
5.3. Evolution of Multi-Source Heterogeneous Signal Fusion Strategies
The operational state information of a combine harvester’s core components—header, threshing drum, cleaning sieve, transmission system—is distributed across heterogeneous signals such as vibration, rotational speed, acoustics, temperature, and images. Multi-sensor information fusion has also benefited field environmental sensing [
174]—for example, estimating soil surface roughness from ultrasonic echoes [
175]—yet most diagnostic studies still rely on single-sensor data, limiting information diversity and accuracy. Existing multi-source fusion methods generally apply basic feature concatenation or statistical weighting, failing to fully exploit the complementary features embedded in heterogeneous data sources [
176].
A single sensor modality captures only limited information dimensions and cannot comprehensively reflect the system-level, multi-component coupled fault state. Multi-sensor fault diagnosis is generally more reliable than single-sensor methods [
177]—an understanding already established during the traditional method stage (
Section 4.4)—and DL has further enhanced the automation and adaptability of multi-source fusion. Current fusion strategies are divided into three levels: data-level, feature-level, and decision-level.
Data-level fusion concatenates or stacks raw signals from multiple homogeneous or heterogeneous sensors into a multi-dimensional input matrix after synchronization, which a deep network then processes. For example, raw waveforms from AE and vibration sensors fed into a one-dimensional CNN (DE-1D-CNN) achieved strong results in extremely low-speed bearing fault diagnosis [
178]. The principal advantage of data-level fusion is maximal preservation of original signal information, avoiding information loss from preprocessing; its drawbacks include stringent cross-modal synchronization requirements and, when sensor heterogeneity is pronounced—as with vibration acceleration versus temperature—differences in physical dimensions and dynamic ranges that complicate network training.
Feature-level fusion dominates current DL applications and offers a key advantage over traditional methods. Where traditional approaches manually extract time-domain, frequency-domain, and time–frequency-domain features from each modality for concatenation and linear dimensionality reduction, deep networks achieve end-to-end adaptive fusion through multi-branch architectures: each modality passes through independent convolutional or recurrent sub-networks, and the resulting deep features undergo concatenation, weighted summation, or cross-attention interaction in a common latent space before final classification.
CWT offers advantages for analyzing the non-stationary vibration signals encountered in combine harvester field operations. Unlike the fixed-resolution spectrograms from short-time Fourier transform (STFT), CWT provides variable time–frequency resolution through scalable wavelet basis functions, capturing transient high-frequency fault impulses—such as those from bearing pitting or gear tooth breakage—alongside the slowly varying low-frequency components that arise from engine speed fluctuations and feed rate variations. This multiresolution capability exceeds what STFT can achieve in a single representation. In the study by Li et al. [
172], vibration and acoustic signals undergo CWT to produce time–frequency images; parallel ResNet branches then extract visual features, followed by a multi-head spatiotemporal attention module and multi-scale temporal convolution for cross-modal enhancement and fusion. Separately, a parallel GRU architecture with multi-head attention models the spatial dimension (multi-sensor channels) and temporal dimension (sequence evolution) independently, enabling flexible information scheduling among sensors with strong stability under small-sample and imbalanced data conditions [
179].
Decision-level fusion integrates soft probability outputs or hard labels from independent diagnostic channels—for example, separate deep models for vibration, acoustics, and temperature—into a comprehensive decision via weighted voting, Bayesian inference, or D-S evidence theory. An existing combine harvester remote monitoring system fused speed indicators, component slip rates, and adaptive threshold judgments to identify operating conditions, achieving 97.46% accuracy [
180]. The key advantage of decision-level fusion is full channel decoupling: failure of a single sensor does not paralyze the entire system. However, information interaction among channels occurs only at the final decision stage, precluding the fine-grained representation-level collaboration that feature-level fusion achieves.
Table 6 and
Table 7 systematically compare the three fusion strategies across fusion stage, end-to-end property, synchronization requirements, information retention, and applicable scenarios in combine harvester fault diagnosis, providing a reference framework for fusion architecture selection.
5.4. Systematic Comparison with Traditional Machine Learning Methods
This section provides a systematic comparison between the traditional ML paradigm (
Section 4.5) and the DL-based approaches surveyed in this chapter.
Table 8 summarizes the differences across five dimensions—feature learning capability, generalization under variable conditions, diagnostic coverage, computational cost, and interpretability—each discussed in the following subsections.
5.4.1. Feature Learning Capability
As analyzed in
Section 4.5, the quality of handcrafted features is fundamentally constrained by the completeness of prior domain knowledge, and traditional feature libraries cannot automatically discover discriminative information for novel or compound faults. DL methods address this limitation by learning hierarchical fault representations directly from raw data. For example, Liu et al. [
181] developed an autoencoder network that achieves unknown anomaly detection without fault samples, using only frequency-domain features and reconstruction error; integrating convolutional and LSTM structures further enhanced autonomous spatiotemporal feature extraction. Sun et al. [
182] proposed an open-set classification method based on time–frequency fusion and latent representation prompts, adaptively constructing complementary time–frequency joint representations and guiding the clustering of similar latent representations to automatically identify unknown faults without requiring domain experts to predefine fault modes. This end-to-end representation learning capability reduces reliance on domain expert knowledge and enables automatic discovery of unknown fault modes.
However, the advantages of DL are context-dependent. When fault modes are well-defined and their physical signatures are well understood—bearing characteristic frequencies, gear mesh harmonics, bolt loosening patterns—handcrafted features with optimized SVM or RF classifiers can match or exceed DL accuracy. A one-versus-one SVM using multi-point vibration features achieved 96.9–99.7% accuracy for bolt state identification on a combine harvester conveyor trough [
138], and an improved RF algorithm reached 97.9% for silage harvester blockage diagnosis [
139]—on par with the 98–99% reported for DL-based approaches [
170,
171]. Under small-sample conditions, the norm in field fault diagnosis, shallow models with appropriate regularization often generalize more robustly than overparameterized deep networks prone to overfitting. The principal contribution of DL therefore lies not in marginal accuracy gains on well-characterized component faults, but in extending diagnostic capability to scenarios where manual feature engineering is infeasible—compound faults under multi-source coupled excitation, end-to-end learning from heterogeneous multi-sensor streams, and cross-condition generalization under continuously time-varying operations. Traditional ML and DL are thus complementary rather than competing paradigms, a distinction essential for guiding practical method selection.
Beyond the traditional-versus-DL debate lies a further methodological concern: published diagnostic accuracy figures almost universally lack uncertainty quantification. With few exceptions, the studies reviewed in
Section 4 and
Section 5 report point-estimate accuracies—often to two or three decimal places—without confidence intervals, standard deviations from repeated trials, or per-class metrics such as precision, recall, and F1-score. This has two consequences. First, it renders performance comparisons across studies unreliable: a reported 97.5% versus 96.8% may suggest a genuine gap, but without error bars one cannot determine whether the difference is statistically significant or merely reflects variation in train–test splits, sensor placement, or specific fault instances. Second, omitting per-class metrics is particularly problematic in the combine harvester context, where severe class imbalance allows a model to achieve high overall accuracy by correctly classifying abundant healthy instances while systematically misclassifying rare but critical fault modes. Reporting per-class precision, recall, and F1-score would expose such behavior; aggregate accuracy alone conceals it. Future studies should, as a minimum standard, report results with ± one standard deviation over at least five independent trials with different random seeds, and provide confusion matrices or per-class metrics for all fault categories. Bootstrap confidence intervals should also be reported where computational resources permit, following established practices in the broader ML reproducibility literature.
5.4.2. Generalization Ability Under Variable Operating Conditions
As discussed in
Section 4.5, continuously varying field conditions cause significant drift in the statistical distribution of handcrafted features, limiting cross-condition generalization of traditional models. DL offers training strategies to mitigate this challenge. Domain adaptation and transfer learning, for instance, reduce the distribution discrepancy of deep features between source and target domains, encouraging the model to learn fault-discriminative features less sensitive to specific operating conditions. These techniques do not yet guarantee a solution for the full spectrum of time-varying field conditions combine harvesters encounter; robust generalization under continuous multi-factor coupling remains an active research frontier. A multi-stage transfer learning framework progressively adapted a CNN pre-trained on laboratory or public datasets through three stages—freezing general feature layers, adjusting for specific vibration representations, and task-focused fine-tuning—enabling effective reuse under different rotational speeds while preserving general low-level representations and optimizing high-level speed-related fault features [
183]. To address extremely scarce fault samples in combine harvester gearboxes under variable conditions, a meta-transfer learning-driven few-shot diagnosis method combines meta-learning for rapid cross-task adaptation, a conditional domain adversarial network for cross-domain discriminative features, and multi-step loss optimization to resolve gradient instability, preliminarily verifying effective fault diagnosis under variable-condition, few-shot scenarios [
184].
Deep models can also alleviate bias from imbalanced fault categories through data augmentation (random noise injection, time warping, masking), resampling, and cost-sensitive loss functions, offering advantages for scenarios with scarce, unevenly distributed fault samples—though their practical deployment still faces the challenges discussed in the following chapter.
5.4.3. Diagnostic Level and Coverage
The literature clearly shows that traditional methods succeed most on component-level faults—particularly standard parts such as rolling bearings and gearboxes with well-defined characteristic frequencies and mature diagnostic knowledge bases (
Section 4.5). DL, in contrast, extends toward whole-machine-level, multi-component synchronous diagnosis. Architectures such as SDAE-SVM and SDAE-BP simultaneously take the rotational speed signals of six core working components—feeding auger, tailings auger, grain auger, fan, threshing drum, and conveyor chain—as input, enabling collaborative monitoring and fault prediction across multiple components at the whole-machine level [
170]. Whole-machine-oriented multi-sensor fusion and multi-task learning architectures can potentially achieve synchronous identification of fault-prone parts such as the threshing unit, header system, and transmission system [
180,
185]. This broader diagnostic coverage extends the potential of deep models for full-lifecycle engineering applications. For well-characterized component-level faults, however, traditional methods and DL achieve comparable accuracy, yet traditional methods do so with fewer training samples and lower computational cost—making them the more practical choice under the data-scarce conditions typical of field diagnosis. The principal contribution of DL lies instead in extending diagnostic capability to complex, multi-fault, whole-machine scenarios where manual feature engineering is infeasible.
5.4.4. Computational Cost and Deployment Feasibility
Computational cost remains a primary bottleneck for deploying deep models in the field. Traditional shallow models (SVM, RF, KNN) offer fast training, few parameters, and efficient execution on low-power embedded microcontrollers. Deep networks, especially large Transformer-based architectures, rely on GPU acceleration, and their memory footprint and power consumption challenge edge computing platforms in mobile scenarios. Recent years have seen rapid development of compression techniques: knowledge distillation trains a lightweight student network to inherit the teacher’s diagnostic capability; network pruning and quantization accelerate inference by removing redundant connections or reducing numerical precision. Cloud–edge collaborative architectures provide a practical system-level compromise. In aircraft fuel pump monitoring, a three-tier architecture performs multi-type signal acquisition at the sensor layer, real-time anomaly detection at the edge layer, and computationally intensive fault classification in the cloud [
186]. Complementarily, lightweight ML models on embedded hardware rank data by novelty at the source, using prediction error and confidence quantification to transmit only information-rich data and reduce communication bandwidth by orders of magnitude in gas turbine monitoring [
187]. Nevertheless, running highly complex real-time diagnostic models stably under the harsh power supply, heat dissipation, and network conditions of agricultural machinery remains an open practical challenge.
5.4.5. Interpretability
A key advantage of traditional methods is the explicit physical interpretability of handcrafted features: an increase in kurtosis directly indicates impulse-type faults, an amplitude rise at specific bearing characteristic frequency pinpoints the defect location, and abnormal temperature signals poor lubrication or overload friction. These indicators maintain clear causal chains with failure mechanisms, making them intuitive for operators to understand and trust.
In contrast, the black-box nature of deep neural networks obscures the decision-making process. Which input patterns lead the model to decide “drum blockage is imminent” or “transmission chain wear is intensifying”? Without additional tools, the answers remain hidden among millions of weights. For safety-critical agricultural machinery, this opacity can directly undermine operator trust and adoption of automated warnings.
Post hoc interpretability techniques have seen increasing adoption in rotating machinery fault diagnosis. Gradient-weighted Class Activation Mapping (Grad-CAM) has been integrated into multi-scale CNN frameworks with vibration image encoding to visualize the time–frequency features driving diagnostic decisions [
188], applied to explain 1D-CNN classifications on synthetic rotor fault data [
189], and comparatively evaluated alongside Layer-wise Relevance Propagation (LRP) and a modified Local Interpretable Model-agnostic Explanations (LIME) algorithm under variable-speed conditions [
190]. Integrated Gradients (IG) with SmoothGrad (SG) have been employed to guide data preprocessing by selecting informative frequency ranges for CWT input to CNN [
191]. SHapley Additive exPlanations (SHAP)-based recursive feature elimination (RFE) has enabled transparent feature selection in SVM classifiers for bearing fault diagnosis, achieving high accuracy and model-level interpretability without deep network opacity [
192].
A further step toward intrinsic interpretability integrates physical knowledge directly into DL architectures. Physics-informed neural networks for fault diagnosis embed governing equations—bearing characteristic frequency relationships, gear mesh dynamics, rotor vibration models—into the loss function or network architecture, constraining learned representations to remain consistent with known physical laws. For rotating machinery, two complementary paradigms have demonstrated success. The first embeds physical knowledge into the network structure: a physics-informed feature weighting method for bearing diagnostics assigns higher weights to features near bearing fault characteristic frequencies in the order spectrum, guiding the network toward physically meaningful signatures and producing more interpretable outputs than purely data-driven CNN [
193]. The second embeds physical knowledge into the loss function and encoding layers: a physics-informed attention LSTM framework incorporating the single-degree-of-freedom dynamic equation of rolling bearings into an attention-based LSTM encoder constructed a dual-knowledge fusion model, achieving over 99% diagnostic accuracy under small-sample conditions while endowing the learned representations with physical significance [
194]. These implementations demonstrate that physics-informed architectures can provide intrinsic interpretability—model representations anchored to physically meaningful quantities—without sacrificing DL’s representational power. However, physics-informed diagnostic models for the multi-component coupled dynamics of combine harvesters remain an open research frontier; the successful demonstrations above come primarily from single-component rotating machinery and have yet to be systematically extended to the whole-machine, multi-fault scenarios characteristic of field harvesting operations.
5.5. Chapter Summary
This chapter has reviewed the progress of DL in combine harvester fault diagnosis, progressing from general paradigms to specific applications, from single-modality perception to multi-source fusion, and concluding with a systematic comparison with traditional methods.
The general technical framework of DL fault diagnosis replaces the traditional separated process with end-to-end joint optimization. Through the flexible combination of convolution, recurrence, autoencoders, and attention mechanisms, hierarchical representations are automatically learned from low-level signal textures to high-level fault semantics. The main training strategies span supervised learning, unsupervised pre-training, transfer learning, and meta-learning, with input forms including one-dimensional waveforms, two-dimensional time–frequency images, and multi-channel fusion.
Among the three signal modalities reviewed, vibration diagnosis has the richest signal information and most mature technical foundation, with architectures such as CNN, SDAE, ResViT, and LSTM validated on key combine harvester components including bearings, gearboxes, and threshing drums. Acoustic diagnosis, enabling non-contact acquisition, is advancing through vibration–acoustic cross-modal fusion. Visual diagnosis remains at an exploratory stage, serving mainly as an auxiliary means for detecting faults with obvious appearance characteristics; its field reliability requires further improvement.
At the multi-source sensor fusion level, DL has enabled a shift from data-level concatenation to adaptive feature-level fusion, where multi-branch networks with attention mechanisms learn optimal fusion weights and modality interactions end-to-end. Decision-level fusion meanwhile provides the basis for highly reliable distributed diagnostic architectures. The three strategies each have distinct applicable boundaries, jointly constituting a multi-level state perception system.
Finally, a systematic five-dimension comparison with the traditional ML methods of
Section 4—feature learning capability, generalization under variable conditions, diagnostic coverage, computational cost, and interpretability—defines the relative advantages and complementary nature of the two paradigms. Traditional methods retain practical value under stable conditions with well-defined fault modes, scarce data, and high interpretability requirements. In complex scenarios characterized by multiple coexisting faults, severe time-varying conditions, and multi-source heterogeneous data, DL offers irreplaceable advantages through representation learning, adaptive fusion, and cross-domain generalization. Physics-informed deep diagnosis, whole-machine-level multi-task collaborative perception, lightweight edge deployment, and trustworthy interpretable diagnosis will be key directions for advancing this field from theory to reliable field application.
6. Challenges for Real-Time Fault Diagnosis
Although DL and multi-sensor fusion have significantly enhanced diagnostic intelligence, bridging the gap from laboratory validation to real-time field deployment remains a substantial engineering challenge. The particularity of combine harvester operating conditions, the evolution of power architectures, and the shift toward unmanned operation confront real-time diagnostic systems with increasingly complex bottlenecks.
6.1. Separation of Weak Fault Features Under Strong Background Noise
Combine harvester fault diagnosis shares fundamental challenges with CM in adjacent domains—wind turbine drivetrains, construction and mining equipment, and off-road heavy vehicles. These domains all involve rotating machinery under non-stationary speed and load, where broadband noise masks fault signatures, fault samples are scarce and imbalanced, and operating environments are harsh and uncontrolled [
104,
195,
196]. Wind turbine CM, particularly well-studied, has identified challenges—data scarcity, low cross-condition generalization, insufficient explainability—that closely parallel those discussed below for combine harvesters [
197]. The non-stationary nature of field operation is a shared cross-domain concern: vibration signals under continuously variable speed and load violate the stationarity assumptions of conventional spectral analysis, and traditional diagnostic techniques designed for steady-state operation do not transfer well to nonlinear, non-stationary processes [
198], motivating the development of advanced time–frequency methods and transfer learning strategies directly relevant to combine harvester diagnostics [
199]. Combine harvesters, however, face additional specificities—extreme non-stationarity from continuous feed rate and terrain fluctuations, multi-crop variability, seasonal duty cycles with long idle periods, and the paramount economic sensitivity of harvest timeliness—that collectively define a more complex diagnostic problem than those in manufacturing or power generation, and motivate the detailed examination of individual challenges in the following subsections.
The dust-laden airflow, mechanical vibrations, and EMI generated during field operations constitute broadband, high-energy background noise that readily submerges the vibration or acoustic signatures of early-stage weak faults such as bearing pitting and gear micro-cracks, severely degrading the SNR at their characteristic frequencies. Existing denoising algorithms mostly rely on the stationarity assumption, whereas the non-stationary noise from drastic fluctuations in harvester operating conditions challenges adaptive filtering. Furthermore, the transient response of the threshing drum under feed rate shocks and fault impulse features often occupy the same analysis scale, with no effective criteria to separate “damaging impulses” from “process impulses” within mode mixing.
Beyond signal-level interference, field conditions impose direct physical challenges on mechanical components. In high-moisture clay-loam soils, severe soil adhesion on tillage and harvesting components reduces working efficiency and increases energy consumption—a problem that has motivated biomimetic surface designs modeled on soil-burrowing animals such as the badger [
200].
6.2. Data Distribution Shift and Model Generalization Under Multi-Condition Coupling
As discussed in
Section 3.3.1 and
Section 4.5, the time-varying field conditions of combine harvesters cause significant distributional shifts in monitoring data, degrading the cross-condition performance of both traditional and DL-based diagnostic models. A more fundamental challenge, however, is that field operating conditions exhibit continuous multi-factor coupling rather than discrete source–target domain switching. Achieving adaptive model generalization to continuously time-varying conditions—without large volumes of fully condition-labeled data—therefore remains a largely unsolved challenge.
6.3. Data Hunger and Extreme Imbalance of Abnormal Samples
Reliable DL models typically require massive fault samples, yet long periods of normal operation are the norm on modern high-reliability combine harvesters. Real physical fault data—especially full-lifecycle data capturing natural degradation to failure—are extremely scarce and costly to acquire, leading to typical “zero-shot” or “few-shot” learning dilemmas. Moreover, fault simulation experiments often obtain samples by artificially implanting damage, and the fault characteristics generated by such “accelerated failure” differ in distribution from those of natural gradual wear in the field. Models trained on artificial data may therefore exhibit high false alarm rates when deployed. Addressing deficiencies in both the quantity and quality of fault samples is key to transitioning from data-driven approaches to actual field deployment.
Promising steps toward overcoming this data bottleneck have been taken for combine harvester gearboxes. A recent meta-transfer learning method combining multi-step loss optimization and a conditional domain adversarial network demonstrated effective cross-condition generalization for gearbox fault diagnosis with few-shot data, simultaneously addressing domain shift and sample scarcity [
184].
Several methodological strategies address class imbalance at the algorithmic level. Data-level techniques include oversampling the minority class through synthetic sample generation, successfully applied to bearing fault diagnosis under imbalanced data [
201], and undersampling of abundant normal-condition data, either randomly or through informed selection, to balance the training distribution. At the algorithm level, cost-sensitive learning [
202] assigns higher misclassification penalties to the minority class, forcing the classifier to attend to fault samples even when vastly outnumbered. In the DL context, specialized loss functions such as focal loss [
203] down-weight well-classified majority-class examples, focusing optimization on challenging minority-class instances. These strategies are not mutually exclusive—combining synthetic oversampling with cost-sensitive training, for instance, can improve diagnostic performance under extreme imbalance. Despite demonstrated effectiveness in related domains, the systematic application and comparative evaluation of these strategies specifically for combine harvester fault diagnosis remains an open research gap and a necessary direction for advancing field-deployable diagnostic systems.
More fundamentally, the field lacks publicly available benchmark datasets and standardized evaluation protocols for combine harvester fault diagnosis. The studies reviewed in
Section 4 and
Section 5 used proprietary datasets from different harvester models, operating conditions, sensor configurations, and fault severity levels, rendering reported accuracies incomparable—97% on one dataset does not imply superiority over 95% on another when the underlying classification task difficulty is unknown. The contrast with neighboring fields is instructive. In rolling-element bearing diagnostics, open benchmark datasets have decisively accelerated methodological progress: the Case Western Reserve University (CWRU) bearing dataset became a de facto standard, used in at least 41 papers published in Mechanical Systems and Signal Processing alone between 2004 and early 2015 [
204]; the Paderborn University (PU) bearing dataset extended this paradigm to motor current signal-based diagnosis, showing that controlled fault generation and systematic labeling enable reproducible evaluation across sensing modalities [
205]. A task-oriented characterization of widely used bearing benchmarks further revealed that dataset properties—fault generation mechanisms, temporal structure, labeling granularity—directly determine which diagnostic tasks they can validly support, underscoring the need for task-oriented benchmark design [
206]. Standardized evaluation frameworks for domain adaptation in fault diagnosis have since been proposed, with parametric partitioning schemes and controlled protocols that isolate individual domain shift factors such as operating condition and fault severity [
207], and benchmark studies for domain generalization have demonstrated the value of open-source datasets and reproducible code frameworks for rigorous cross-domain evaluation [
208].
No comparable infrastructure exists for combine harvesters. Establishing such benchmarks would require multi-institutional collaboration to acquire instrumented field data across representative harvester models, fault types, crop varieties, and harvesting environments—a non-trivial undertaking whose payoff in methodological rigor and accelerated research progress would be substantial. Until such benchmarks emerge, adopting minimum reporting standards—confusion matrices, per-class metrics, and cross-condition evaluation results, as advocated in
Section 5.4.1—would improve the comparability of results across existing proprietary-dataset studies. As a minimum reporting standard for the field, we propose that future diagnostic studies on combine harvesters report: (1) confusion matrices for all fault categories, including the healthy/normal class; (2) per-class precision, recall, and F1-score; (3) mean and standard deviation of accuracy over at least five-fold cross-validation with different random seeds; and (4) metadata on operating conditions—crop type and moisture content, terrain slope, harvester model and engine speed, sensor types and installation positions—under which data were acquired. These requirements mirror reporting standards adopted in related fields such as rotating machinery diagnostics, where benchmark datasets with controlled fault generation and systematic labeling have enabled reproducible evaluation across different sensing modalities. Their adoption in combine harvester fault diagnosis would substantially improve the comparability and reproducibility of published results, even in the absence of shared benchmark datasets.
6.4. Conflict Between Limited Onboard Computing Power and Real-Time Requirements
Computationally intensive DL models—deep Transformers and multimodal fusion architectures—conflict with the reality that combine harvesters cannot accommodate expensive, high-power industrial computers. The hot, high-vibration field environment strictly limits the computing power, memory, and thermal stability of edge computing hardware such as GPU embedded modules.
Real-time diagnosis demands not only fast single inference but also sustained computation without thermal throttling under continuous high load. A structural tension therefore exists between high-precision large models and low-latency real-time inference. Bridging this gap requires structurally lightweight, hardware-friendly efficient inference networks and cloud–edge collaborative elastic computing frameworks.
Beyond general computational efficiency, safety-critical real-time fault diagnosis imposes specific constraints on inference latency, jitter, and determinism. Inference latency—from sensor acquisition to diagnostic output—determines how quickly a developing fault can be detected and acted upon. For rapidly propagating faults such as bearing seizure or belt slippage under high load, latency on the order of seconds may be acceptable; for impending threshing unit blockage, where feed rate can surge within milliseconds, sub-second latency is essential to enable preventive action before material accumulation reaches a critical level. Lightweight DL models on edge platforms have demonstrated sub-second inference for rotating machinery fault diagnosis: an efficient CNN achieved 0.31 s per sample on an edge device while maintaining high accuracy [
209], and a lightweight FT–CNN–Transformer on a Raspberry Pi achieved significantly faster inference than previous 2D image-based approaches by directly processing one-dimensional time-domain signals, showing that careful architectural design can reconcile accuracy with stringent latency demands [
210]. Jitter—inference latency variability across diagnostic cycles—is equally important: high jitter introduces timing uncertainty, potentially causing missed intervention windows or false alarms from timing artifacts. Determinism—the guarantee that inference completes within a bounded time under all conditions—is a prerequisite for integrating diagnostic outputs into automated control loops, since non-deterministic delays could cause safety-critical actuation such as emergency power reduction or header disengagement to occur too late. These timing requirements constitute functional safety constraints: standards such as ISO 25119 [
211] specify acceptable failure rates and response times for safety-related control systems, and diagnostic modules feeding into such systems must meet commensurate temporal specifications. Addressing these constraints demands co-design of diagnostic algorithms and execution hardware—the algorithm’s worst-case execution time must be bounded, and the hardware must provide sufficient computational margin to absorb transient load spikes without violating latency bounds. This co-design paradigm departs fundamentally from the accuracy-centric optimization prevalent in current research and is a necessary step toward field-deployable, safety-qualified diagnostic systems.
6.5. Trust and False Alarm Problems Caused by Lack of Interpretability
The black-box nature of DL models means that even high-confidence fault warnings come with opaque decision logic, breeding operator distrust. Random transient interference can produce isolated false alarms that, if they frequently trigger shutdown inspections, exacerbate timeliness loss. Currently, post hoc interpretability methods often generate saliency maps (heatmaps) that appear scattered and blurred due to physically meaningless noise interference. Developing physics-constrained neural network architectures jointly trained with data and domain knowledge is essential to achieve transparent, reliable decision logic.
Beyond the technical dimensions of interpretability, the human side of fault diagnosis has received little attention in the combine harvester literature. Operator trust is not a simple function of accuracy—it reflects prior experience with false alarms, the perceived transparency of decision logic, and the alignment between machine recommendations and the operator’s own sensory observations. In field environments where experienced operators have long relied on auditory, tactile, and visual cues, a diagnostic system that cannot explain its reasoning in terms that resonate with that experiential knowledge risks being ignored regardless of its statistical performance. Alarm fatigue—progressive desensitization to frequent warnings—poses a related challenge: even a modest false alarm rate over hundreds of operating hours can condition operators to disregard all alerts, including genuine early warnings. The design of the human–machine interface—how diagnostic information is prioritized, formatted, and presented to operators with varying technical training—directly mediates these risks, yet human-factors evaluation of diagnostic interfaces for agricultural machinery remains largely unexplored. As combine harvesters transition toward higher levels of autonomy, the human role shifts from direct operator to supervisory controller, and the interaction design problem becomes one of maintaining appropriate trust and situation awareness rather than simply conveying fault codes. Design recommendations for diagnostic HMIs in agricultural machinery include tiered alert displays that distinguish critical safety warnings from advisory notifications, confidence-level indicators that convey diagnostic uncertainty to operators, and multimodal alert delivery—visual, auditory, and haptic—to accommodate the high-noise cabin environments typical of harvester operation. While human-factors standards for diagnostic interfaces exist in adjacent domains—such as ISO 9241 for ergonomics of human–system interaction [
212] and ISO 11064 for control center design [
213]—agricultural machinery currently lacks equivalent guidelines tailored specifically to fault diagnostic displays, and human-factors evaluation of diagnostic interfaces for harvesters remains an open research area. Systematic investigation of these human-factors dimensions is an essential complement to the algorithmic advances surveyed in this review.
6.6. New Diagnostic Complexity Introduced by Electrified Power Architectures
The power architecture of combine harvesters is gradually evolving toward hybrid electric and fully electric drive for constant engine speed operation, braking energy recovery, and reduced fuel consumption [
214]. Battery electric drive systems have also been validated in other agricultural machinery such as small orchard tractors [
215]. However, high-voltage battery packs, drive motors, inverters, and DC-DC converters significantly expand the fault mode space. Electrical faults—inter-turn short circuits and demagnetization of permanent magnet synchronous motors, thermal cycling fatigue of power modules, aging of DC-link capacitors—produce characteristic signals (current ripple, partial discharge, hot spots) that differ from traditional mechanical vibrations, compelling diagnostic systems to integrate electrical monitoring channels. More critically, the high-intensity EMI generated by high-frequency switching of high-power inverters couples into vibration and acoustic sensor circuits through conducted and radiated paths, further degrading the already weak SNR of mechanical fault signals. Existing shielding and filtering measures offer limited effectiveness against such broadband electromagnetic aggression. Integrated electro-mechanical diagnosis and electromagnetic compatibility (EMC) co-design therefore constitutes a new and necessary challenge for hybrid harvester diagnosis.
Mitigating these effects requires a layered EMC strategy spanning sensor selection, signal conditioning, and algorithmic robustness. At the sensor and cabling level, differential-output accelerometers and microphones with balanced line drivers offer substantially greater common-mode noise rejection than single-ended alternatives; in extreme cases, fiber-optic acoustic or strain sensors provide galvanic isolation that eliminates conducted interference paths entirely. At the signal conditioning level, anti-aliasing filters with cutoff frequencies tuned to the diagnostic bandwidth of interest—rather than the full Nyquist bandwidth—suppress high-frequency EMI before digitization. At the algorithmic level, diagnostic models can be explicitly trained or fine-tuned on data acquired under representative electromagnetic conditions, including during maximum-power inverter operation, to learn interference-robust feature representations. The automotive industry has developed standardized EMC test protocols—ISO 11452 [
216] for component-level radiated immunity and CISPR 25 for emission limits in vehicles—that could be adapted to agricultural machinery to establish minimum electromagnetic immunity requirements for diagnostic sensor systems. Systematic investigation of EMI effects on diagnostic signal integrity and the development of agricultural-specific EMC design guidelines for onboard sensing systems are necessary steps toward robust diagnosis on electrified harvesters, yet remain absent from the current literature.
The distinction between retrofit and new-design scenarios is particularly relevant for EMI management. New harvester designs can integrate shielding, filtering, and harness layout optimization at the manufacturing stage, whereas retrofitting diagnostic systems onto existing machines must contend with pre-existing cable routing and limited access for EMI countermeasures. Beyond EMI, electrified powertrains introduce diagnostic challenges from novel fault modes—including power electronic device degradation, motor winding insulation breakdown, and battery state-of-health deterioration [
217,
218]—that produce electrical signature patterns fundamentally different from the vibration-based fault features discussed in
Section 4 and
Section 5. These fault modes demand diagnostic approaches distinct from the vibration-based methods that dominate the existing combine harvester literature, and their gradual degradation trajectories suggest that future diagnostic frameworks must incorporate life-cycle monitoring capabilities beyond traditional fault classification. The harsh agricultural environment—extreme temperature variations, high humidity, dust, and persistent broadband vibration—further compounds these challenges through multi-stress coupling effects on power converters and motor drive systems [
217], representing a necessary direction for future research as electrification progresses.
6.7. Special Requirements for Diagnostic System Autonomy in Unmanned Operations
Deploying unmanned combine harvesters eliminates the driver’s role as a multimodal perception and decision-making center. Perceptual dimensions once provided by the driver—detecting abnormal noise, burnt smell, and abnormal vibration—must now be fully inherited by microphone arrays, electronic noses, thermal imagers, and multi-axis vibration sensors. The diagnostic system thus shifts from assisting the operator to replacing the operator: its output becomes not a suggestion correctable by a human but a safety command directly triggering power reduction, emergency shutdown, or remote intervention.
This shift imposes near-absolute demands on diagnostic reliability. A false alarm causing full machine shutdown inflicts irreparable timeliness loss, while a missed detection allows the machine to continue operating with faults until catastrophic failure occurs without human intervention.
Furthermore, in remote fields with weak or no network connectivity, the edge diagnostic unit must possess fully local autonomous decision-making capability and comply with the strict failure rate constraints that agricultural machinery functional safety standards—ISO 25119—impose on safety-related control systems. Compressing high-precision DL models into embedded systems that meet functional safety integrity levels while maintaining extremely low false alarm and missed detection rates presents a multi-objective optimization challenge.
6.8. Model Degradation and Adaptation over Operational Lifetimes
An equally critical but less explored challenge is ensuring that a deployed diagnostic model maintains its performance over the multi-year operational lifetime of a combine harvester. Two interrelated phenomena drive model degradation: concept drift and the emergence of novel fault modes.
Concept drift occurs when the statistical properties of sensor data gradually shift over time. In wind turbine CM, concept drift has been systematically defined as machine performance deviating significantly from an established baseline, often signaling incipient faults [
219]. For combine harvesters, similar drift arises from multiple sources: component wear shifts the vibration signature of a healthy transmission away from its factory-calibrated baseline; seasonal variations in crop properties alter dynamic loads on structural components; changes in field terrain introduce new excitation patterns; and sensor degradation over time introduces measurement bias. A model trained on wheat-harvesting data may therefore exhibit systematically degraded performance when deployed for soybean or corn, owing to differences in crop physical characteristics and resulting machine dynamics.
The emergence of novel fault modes—failure types absent from the original training data—poses an even steeper challenge. As the open-set fault diagnosis literature has comprehensively documented, conventional diagnostic models operate under a closed-set assumption, presuming all possible fault classes are known and represented during training [
220]. Over a harvester’s service life, previously undocumented failure patterns may surface due to design modifications, new materials in replacement parts, or evolving operational practices. A classifier trained to distinguish healthy, bearing fault, and gear fault states cannot recognize a previously unseen seal leak or solenoid valve malfunction; it will either misclassify the novel fault as a known category or, if equipped with open-set recognition, correctly flag it as “unknown” without providing actionable diagnostic information.
Promising directions for addressing model degradation have emerged in incremental learning-enabled fault diagnosis [
221]. These include periodic retraining with field data accumulated during normal operation, incremental architectures that update model parameters without full retraining from scratch, and open-set recognition frameworks that flag distributional shifts for expert review. Each approach, however, introduces its own difficulties: incremental learning risks catastrophic forgetting, where previously acquired diagnostic knowledge erodes as the model adapts to new data [
220]; retraining demands ground-truth labels that are scarce in field environments, since harvesting operations are rarely interrupted for detailed failure analysis; and open-set recognition generates “unknown” alarms that, without subsequent expert annotation and model updating, contribute little to improving diagnostic coverage. For combine harvesters operating in remote fields with intermittent connectivity, any model updating strategy must also contend with bandwidth constraints for transmitting training data and the limited computational budget of onboard edge hardware.
The systematic investigation of model updating and lifelong learning strategies for agricultural machinery fault diagnosis remains a nascent research area [
221], yet advancing it is essential for transitioning diagnostic systems from one-time deployment to sustained, trustworthy operation over a combine harvester’s entire service life.
6.9. Functional Safety Compliance for Diagnostic Systems
Beyond achieving high diagnostic accuracy, any fault diagnosis system feeding into safety-related control functions must satisfy agricultural machinery functional safety standards. ISO 25119, “Tractors and machinery for agriculture and forestry—Safety-related parts of control systems,” is the principal normative framework. The standard comprises four parts—general principles (Part 1), concept phase (Part 2), hardware development (Part 3), and software development (Part 4)—and defines Safety Integrity Levels from SIL 1 (lowest) to SIL 4 (most stringent).
For combine harvester fault diagnosis, ISO 25119 is relevant in two respects. First, when a diagnostic module directly triggers safety actions—emergency engine power reduction upon detecting imminent threshing drum seizure, or automatic header disengagement upon critical structural failure—that module becomes part of a safety-related control system and must meet the appropriate SIL. The standard requires that software-based diagnostic functions undergo rigorous verification and validation, that their failure modes and rates be systematically analyzed, and that their response times be deterministically bounded. The real-time latency and determinism requirements discussed in
Section 6.4 are thus functional safety prerequisites, not merely performance preferences. Second, even a purely advisory diagnostic system—issuing alerts without automated intervention—remains subject to risk assessment and hazard analysis. A false alarm distracting the operator during a safety-critical maneuver, or a missed detection allowing a degrading component to fail catastrophically, can contribute to hazardous situations even though the diagnostic system itself is not a safety actuator.
Despite its centrality to deploying intelligent diagnostic systems on agricultural machinery, the literature surveyed in this review is almost entirely silent on functional safety. None of the studies reviewed explicitly addresses ISO 25119 compliance or conducts a formal hazard and risk analysis for their proposed diagnostic methods. This gap between the data-driven diagnostic research community and the functional safety engineering discipline represents a significant barrier to industrial adoption. Future work, particularly that targeting deployment on unmanned combine harvesters (
Section 6.7), must bridge this gap by incorporating functional safety considerations from the earliest design stages: defining safety goals and corresponding SIL requirements; designing diagnostic architectures with hardware and software redundancy commensurate with the target SIL; and providing evidence that diagnostic performance metrics—false alarm rate, missed detection rate, and inference latency—satisfy the quantitative targets derived from the hazard analysis. Translating ISO 25119 into specific diagnostic system design requirements remains an unexplored research area. Key questions that future work must address include determining appropriate SIL targets for different failure severities in harvesting operations, defining diagnostic coverage requirements for each target SIL, and establishing proof-test intervals for periodic validation of diagnostic functions. These requirements collectively inform sensor redundancy, voting architectures, fail-safe design, and the rigor of verification and validation activities—an integration that, despite established parallels in automotive and industrial machinery functional safety, has yet to be addressed in the combine harvester diagnostic literature.
7. Future Outlook
Before addressing individual research directions,
Figure 10 provides a panoramic overview of the challenges analyzed in
Section 6 and their relationship to the future directions discussed in this section. The figure organizes challenges and directions into four thematic groups—signal and data, computation and deployment, trust and safety, and lifecycle and extension. The grouping reflects thematic affinity rather than strict one-to-one mapping. Notably, physics-informed approaches (
Section 7.1) serve a dual role—enhancing signal robustness against field noise and model interpretability for operator trust—and this dual contribution is reflected in their appearance across two groups. Subsection numbers annotated in each block allow readers to locate the corresponding detailed discussion in
Section 6 and
Section 7.
To overcome these challenges, advancing combine harvester fault diagnosis to practical application requires systematic exploration of physics-data fusion paradigms, autonomous diagnostic architectures, edge intelligent hardware, and support for new power systems and unmanned operations.
7.1. Hybrid-Driven Diagnosis with Embedded Physical Information
Purely data-driven approaches struggle with generalization and interpretability, while purely physical models cannot capture the full complexity of field conditions. A promising direction embeds dynamic mechanism equations and wear degradation models as physical constraints or inductive biases into neural network architectures or loss functions, constructing physics-informed ML models that reduce data dependence, enhance extrapolation across operating conditions, and improve decision interpretability. For hybrid systems, an electromechanical coupling physical model can serve as a regularization term, guiding reasoning logic under combined electromagnetic and mechanical faults.
7.2. Self-Supervised Pre-Training and Few-Shot Diagnostic Paradigm
Given the scarcity of fault samples in high-end combine harvesters, the field must move beyond the fully supervised annotation paradigm. The vast quantities of unlabeled normal operating data generated daily can be exploited for self-supervised pre-training based on contrastive learning or masked autoencoders, enabling the model to learn generalized signal representations. Subsequent fine-tuning with very few fault samples can then achieve few-shot or even zero-shot anomaly detection. This “general education first, specialization later” paradigm promises to replace the current closed setting of “one fault, one training session” with a general whole-machine state representation foundation model, and is particularly suited to hybrid and unmanned agricultural machinery scenarios where data are even scarcer.
7.3. Model Lightweighting and Hardware Acceleration for the Edge
Bridging the gap between high-precision models and limited onboard computing power requires collaborative software and hardware optimization. On the algorithm side, techniques such as neural architecture search and dynamic width/depth inference can seek a Pareto optimum between diagnostic accuracy and inference latency. For real-time diagnosis in unmanned contexts, model inference must meet functional safety constraints on response time and determinism. On the deployment side, low-power edge AI hardware—processing-in-memory and event-driven brain-inspired chips—requires exploration for agricultural machinery, along with EMC reinforcement to cope with the interference environment of hybrid electric powertrains.
7.4. Staged Verification Strategy for Pre-Deployment Validation
A structured verification framework is essential to ensure diagnostic models perform reliably before entrusting them with real-time decisions on operating harvesters. Drawing on established practices in aerospace and industrial CM [
222], we propose a three-stage verification pipeline for agricultural machinery fault diagnosis.
The first stage is offline validation on curated datasets. Diagnostic models should be evaluated not merely on aggregate accuracy but on per-class metrics—precision, recall, and F1-score—computed for each fault mode under multiple operating conditions. Cross-condition testing protocols, training on one set of operating parameters and testing on another, can expose brittleness to distributional shift before field deployment [
223]. Publicly available benchmark datasets, such as those from the Case Western Reserve University bearing data center and the Paderborn University bearing dataset [
204], can serve as standardized testbeds for initial algorithm comparison, though their transferability to agricultural machinery must be critically assessed.
The second stage is hardware-in-the-loop (HIL) simulation [
224]. After compression into their deployable lightweight form, diagnostic algorithms execute on the target edge computing hardware, fed with real sensor signals replayed from pre-recorded field data or generated by real-time simulators. This stage verifies that inference latency, memory consumption, and thermal behavior remain within acceptable bounds under sustained operation, and provides a controlled environment for injecting simulated fault signatures into normal-condition recordings, enabling quantitative assessment of detection sensitivity without requiring actual component damage.
The third stage is controlled field testing under instrumented conditions. Prior to full autonomous deployment, the diagnostic system operates in a passive monitoring mode on a manned harvester, logging its predictions alongside conventional inspection records. Comparing diagnostic outputs with ground-truth maintenance findings over an entire harvest season yields estimates of false alarm rate, missed detection rate, and mean time to first detection. Only after these metrics satisfy predefined acceptance criteria—established in consultation with domain engineers and aligned with relevant functional safety standards—does the diagnostic system transition to active decision-making roles.
This staged approach acknowledges that no single validation metric or laboratory benchmark can capture the full complexity of field conditions. It provides a pragmatic roadmap from algorithm development to trustworthy deployment, while generating the real-world case studies whose absence the existing literature currently notes as a limitation.
7.5. Fault Data Sample Augmentation Enhanced by Digital Twins
A comprehensive review of digital twin (DT) in agriculture has outlined the generic framework and potential applications, which can be adapted for the SHM of combine harvesters [
225]. High-fidelity DT of key components—including the electric motor, inverter, and transmission mechanism in hybrid electric systems—can be constructed for virtual-real interactive mapping. Injecting multiple fault modes and random operating condition combinations into the virtual space then generates massive quantities of high-fidelity, labeled fault simulation data. Combined with domain randomization, models extensively trained on virtual data require only fine-tuning to transfer to physical entities, potentially overcoming the data hunger caused by extreme operating conditions and scarce fault samples, and providing a full-lifecycle offline training and verification environment for the diagnostic systems of unmanned agricultural machinery.
These opportunities, however, require careful contextualization. DT technology in general—and its application to agricultural machinery in particular—remains at a relatively early stage. A cross-industry umbrella review found DT research fragmented across domains, with inconsistencies in definitions, methodologies, and recognized challenges [
226]. Even in manufacturing, the most mature DT domain, core challenges persist in modeling fidelity, real-time synchronization, and verification and validation, with no structured validation framework yet established [
227]. This contrasts with domains such as wind energy, where DTs have progressed to operational deployment for performance optimization and predictive maintenance [
228]. In agricultural machinery, a systematic review concluded that practical, IoT-powered DT applications remain nascent, limited by data integration, optimization, and communication constraints, and the literature lacks structured classification of physical twin components and standardized definitions of physical–virtual integration levels for off-road machinery [
225]. For combine harvesters specifically, DT systems remain scarce and limited to specific subsystems or offline simulation; no peer-reviewed study has demonstrated a fully operational DT for combine harvester SHM that generates validated fault simulation data for diagnostic model training [
229]. The DT direction outlined here should therefore be viewed as a strategic research vision rather than a near-term deployable solution. Incremental progress is more realistic, beginning with hybrid models that embed partial physical constraints into data-driven architectures, validated on individual subsystems under controlled conditions, before scaling toward whole-machine, multi-physics, full-season DT implementations. In the near term, physics-informed hybrid models (
Section 7.1) and staged verification on individual subsystems (
Section 7.4) offer more immediately achievable pathways toward simulation-driven diagnostics. Full DT integration represents a longer-term vision, contingent on advances in modeling fidelity, real-time synchronization, and validation frameworks that are only beginning to emerge in agricultural machinery.
7.6. Resilient Diagnostic Interaction Architecture for Diverse Human–Machine Relationships
The operation mode of combine harvesters is gradually transitioning from manual driving and remote control toward full autonomy, with multiple modes coexisting for the foreseeable future. The diagnostic system must accordingly shift from an “assistant tool” to a “resilient agent.”
In manned or semi-autonomous scenarios, decision-level human–machine collaboration is central. The diagnostic output should convey not merely a fault probability but the fault location, severity, and recommended actions in an interpretable manner—for instance, through feature attribution grounded in physical constraints or natural language descriptions. A human-in-the-loop mechanism allows operators to provide annotation feedback for ambiguous alarms, while active learning enables the model to evolve online, mitigating the trust crisis that false alarms create.
In fully autonomous operation, the diagnostic system must possess self-validation capability—assessing the uncertainty of its own judgments in real time before issuing safety-critical decisions. This requires a confidence modeling layer based on Bayesian inference, evidence theory, or conformal prediction at the output of the diagnostic network. When uncertainty exceeds a safety threshold, the system should automatically degrade operation or request remote intervention rather than blindly executing high-risk commands.
This self-validation engine also serves human–machine collaboration: when the diagnostic system lacks confidence in its judgment, it actively requests operator intervention. The future diagnostic interaction architecture should thus synchronously increase diagnostic autonomy and output confidence—autonomous execution at high confidence, intervention request at low confidence. Human–machine collaboration and self-confirming diagnosis are not mutually exclusive but rather adjacent segments on a single spectrum of resilient autonomous diagnosis.
7.7. Cross-Energy-Domain Diagnosis and Fault-Tolerant Control Adapted to New Power Architectures
As combine harvester electrification deepens, future diagnostic frameworks must fuse mechanical, hydraulic, and electrical physical quantities. Mechanic–electronic–hydraulic powertrain systems in agricultural tractors have demonstrated measurable energy savings, offering a reference for next-generation hybrid combine harvesters [
230]. Energy optimization control based on quasi-cycle power demand estimation has also been explored for extended-range hybrid harvesters to balance efficiency and operational demands [
231]. Future work should model fault propagation based on power flow topology, locating root causes through energy flow anomalies rather than analyzing single-component signals in isolation. Diagnostic results should directly drive fault-tolerant control strategies: upon detecting abnormal temperature in a motor winding, for example, torque redistribution can derate the faulty motor while the remaining motors compensate, allowing harvesting to continue uninterrupted while ensuring safety. This integrated diagnosis-to-fault-tolerance design will become a key technology for improving the mission reliability of hybrid and electric combine harvesters.
7.8. Harvest Quality-Aware Diagnostics
Beyond preventing mechanical downtime, future diagnostic systems could monitor parameters directly affecting harvested grain quality. Post-harvest storage conditions—particularly elevated temperature and moisture—accelerate nutrient degradation and the production of acidic compounds in rice, paddy, and soybean, compromising freshness and market value [
232]. The nutritional value of harvested grain—including bioactive compounds such as bound polyphenols with documented antioxidant and antitumor properties [
233] and the health benefits of grain-derived dietary fiber [
234]—suggests that future harvest quality metrics should encompass not only physical integrity but also the preservation of nutritional and functional components. Integrating quality-centric sensing into harvester health management frameworks would thus extend onboard diagnostics from machine reliability to crop quality assurance, creating a unified platform for both operational efficiency and food quality protection.
7.9. From Research to Practice
Translating the diagnostic advances reviewed in this paper into operational systems on working harvesters requires bridging gaps that extend well beyond algorithmic performance. The deployment pathway spans sensor selection and placement that balance diagnostic coverage against cost and installation complexity; edge computing platforms that meet the latency, thermal, and vibration constraints of field environments; and validation protocols that establish trustworthiness before diagnostic outputs guide maintenance or control decisions.
Practical deployment also demands economic viability—the upfront hardware and integration cost must be justified by demonstrable reductions in downtime-related losses—yet cost–benefit analyses specific to agricultural machinery fault diagnosis remain conspicuously absent from the literature. A structured cost–benefit analysis for agricultural machinery fault diagnosis would account for cost components—sensor hardware, installation labor, edge computing units, data transmission infrastructure, and ongoing maintenance—and benefit categories—reduced downtime, repair cost savings, avoided timeliness loss, extended component service life, and lower insurance premiums. While such analyses remain absent from the agricultural machinery literature, cost–benefit frameworks from adjacent domains provide methodological references. A systematic review of 42 CMS evaluation studies across industrial applications identified established approaches—including cost–benefit analysis, cost-effectiveness analysis, and net present value methods—though it found that only one-third of studies comprehensively incorporate equipment, maintenance, and CMS-related costs [
235]. A review of SHM cost-effectiveness in aviation similarly noted that maintenance manpower, fuel consumption, and sensor costs are the most commonly modeled cost variables, while downtime costs and sensor reliability are frequently neglected despite their significant economic impact, and called for standardized frameworks to ensure consistency across future studies [
236]. A concrete example from wind turbine CM demonstrated a net present value model for journal bearing CMS that identified failure detection rate, sensor hardware cost, and bearing failure rate as the most influential parameters, and explicitly distinguished between retrofit and new-installation scenarios—a distinction equally relevant to agricultural machinery [
237]. These precedents collectively provide both a reference framework and cautionary guidance for structuring cost–benefit analyses for agricultural machinery fault diagnosis.
While comprehensive multi-season field deployment studies remain scarce, individual studies have demonstrated promising field validation results that offer proof-of-concept evidence for broader deployment. For instance, a one-versus-one SVM model using multi-point vibration features achieved 96.9–99.7% accuracy for bolt state identification on a combine harvester conveyor trough under field conditions [
138], and a dual-speed robust balancing method achieved residual unbalances of 37 g and 45 g on two combine harvester models in field tests [
59]. These examples, though limited in scope, demonstrate that data-driven diagnostic and physics-based correction methods can perform effectively outside the laboratory when properly validated.
Similarly, documented case studies of multi-season field deployments with transparent reporting of both successes and failures are urgently needed to ground the research literature in operational reality. A further rarely addressed dimension is the distinction between retrofitting diagnostic systems onto existing harvester fleets—where sensor installation is constrained by pre-existing mechanical and electrical architectures—and designing diagnostic capability into new harvester models from the outset, where sensors, wiring harnesses, and computing modules can be optimally integrated during manufacturing. These two deployment scenarios impose fundamentally different constraints on sensor selection, signal conditioning, and system cost, yet the current literature offers no systematic comparison between them. The staged verification framework outlined in
Section 7.4 and the human–machine interaction principles discussed in
Section 7.6 provide starting points for translational work, but systematic, cross-disciplinary efforts engaging equipment manufacturers, maintenance service providers, and agricultural economists alongside diagnostic algorithm developers will be essential to move the field from laboratory demonstrations to commercially viable, trusted products.
8. Conclusions
This review has systematically examined the progress of structural fault diagnosis for combine harvesters, tracing a trajectory from passive structural redundancy to active state perception. Structural optimization strategies remain an irreplaceable first line of defense, yet time-varying operating conditions, system degradation, coupling transmission, and random overloads constrain their effective boundaries. These boundaries have driven the shift toward data-driven fault diagnosis. Traditional ML methods offer advantages in physical interpretability and small-sample performance, but exhibit fundamental limitations: high dependence on expert knowledge, inability to capture deep-level features, and poor cross-condition generalization.
DL with multi-sensor fusion has recently overcome many bottlenecks of traditional methods, providing a new diagnostic paradigm. Moving toward field deployment, however, real-time diagnostic systems face both persistent obstacles—strong noise interference, cross-condition generalization, scarce fault samples, limited onboard computing power, and insufficient interpretability—and new challenges from powertrain electrification: electromechanically coupled faults, EMI, and the stringent requirements for diagnostic autonomy and functional safety in unmanned operations.
These research directions fall along a temporal horizon that can guide future efforts. Short-term priorities—immediately actionable, building on existing methods—include physics-informed hybrid models, self-supervised pre-training and few-shot learning, and staged verification frameworks. Medium-term priorities, requiring further methodological development and field validation, include lightweight edge inference, incremental learning for model updating, and open-set recognition for novel fault modes. Long-term visions, contingent on advances in enabling technologies, include full DT integration and cross-energy-domain diagnosis with fault-tolerant control. This temporal framing highlights that near-term progress on physics-informed and self-supervised approaches can yield practical benefits while laying the groundwork for more ambitious long-term goals.
Overall, the future reliability assurance system for combine harvesters must integrate a robust physical structure, an efficient electrified drive, a multimodal sensing system, safety-critical autonomous diagnosis, and cloud-based intelligence. Whether for manned or unmanned operation, purely mechanical transmission or hybrid drive, this system must take mission reliability as its core and shift from single-fault detection to holistic production risk management. Such a multi-layered virtual–physical fusion architecture can approach the ideal of “zero unplanned downtime,” providing key technical support for global grain supply chain resilience and sustainable agricultural development.