Next Article in Journal
Tropical Agriculture, Scientific and Technological Cooperation, and Knowledge Transfer Between China and Latin America: The China–Ecuador Case
Previous Article in Journal
Hyperspectral Prediction of Variety, SPAD Value, and Water Content of Oilseed Rape Leaves Using an Improved WGAN-GP
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

A Comprehensive Review of Deep Learning in Agricultural Visual Perception: Progress, Bottlenecks, and Emerging Trends

1
School of Low-Altitude Technology, Jiangsu University, Xuefu Road 301, Zhenjiang 212013, China
2
School of Electrical and Information Engineering, Jiangsu University, Xuefu Road 301, Zhenjiang 212013, China
*
Author to whom correspondence should be addressed.
Agriculture 2026, 16(17), 1826; https://doi.org/10.3390/agriculture16171826
Submission received: 12 July 2026 / Revised: 7 August 2026 / Accepted: 24 August 2026 / Published: 26 August 2026
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)

Abstract

This paper provides a comprehensive review of deep learning in agricultural visual perception. Climate change and land degradation demand a shift from experience-driven to data-driven intelligent agriculture. Traditional manual inspections remain subjective and unscalable. Computer vision offers a non-invasive solution for precision crop and livestock management, while challenges also exist. Its application in actual agricultural scenarios faces specific challenges such as severe occlusion, drastic changes in lighting, and non-rigid deformation of biological targets. To systematically summarize how these perception bottlenecks are being resolved, this review explores deep learning architectures featuring spatial extraction and spatio-temporal modeling. This comprehensive review first constructs a progressive analytical framework from low-level data augmentation and static spatial cognition to high-level dynamic spatio-temporal reasoning, and then systematically classifies existing literature. In-depth analysis reveals that these three core challenges can be computationally resolved using three unified technologies. Domain drift and label scarcity, rather than baseline accuracy, are the main obstacles to practical deployment. We further identified four priority research directions, i.e., architectural efficiency, data-level annotation efficiency, deployment-level privacy and simulation, and trust-level security, to guide future research.

1. Introduction

Driven by escalating climate change, land degradation, and resource scarcity, precision agriculture has emerged as a strategic imperative for ensuring global food security [1]. Within the realms of broad-acre farming and advanced animal husbandry, achieving precise, instantaneous observation of plant developmental traits and specific animal actions remains essential for refining operational strategies. Nevertheless, automated monitoring in such environments encounters a fundamental obstacle in computer vision: the intensely shifting spatial arrangements and continuous physical transformations of the observed subjects. Accurate and real-time perception of crop phenotypes and livestock behavior is fundamental to optimizing production decisions and improving yield in modern agriculture [2]. However, intelligent perception in these scenarios faces a core bottleneck: the highly dynamic spatial and morphological evolution of targets [3]. For crop phenotyping, plant canopies undergo drastic spatiotemporal changes throughout their growth cycle, such as wheat transitioning from creeping to upright posture and from green to yellow [4]. For livestock analysis, individual animals exhibit severe non-rigid deformations during posture transitions, e.g., pigs standing, lying down, or walking [5]. These behavioral state variations produce substantially different visual representations, posing significant challenges for robust visual recognition [6]. The non-rigid deformation of pigs when standing, lying down, walking, and interacting with environmental enrichment objects leads to huge differences in their visual representations under different behavioral states [7]. Traditional operators with fixed geometric topology fail under such conditions [8]. Deep learning has emerged as the dominant paradigm to overcome this limitation, with architectures such as Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), and Graph Neural Networks (GNNs) learning adaptive hierarchical representations from massive data.
Besides dynamic morphological evolution, high density and severe occlusion from spatial overlap pose another major challenge for visual perception. Intertwined targets severely degrade object boundaries in agricultural images. For example, dense cotton bolls create substantial visual clutter that obscures individual plant structures [9,10]. Similarly, highly aggregated fish during feeding cause severe inter-individual overlap, making it difficult for standard tracking pipelines to maintain consistent target identities and ultimately leading to irreversible loss of structural information [11,12]. To overcome this limitation, modern visual perception systems in precision agriculture are transforming discrete individual identification toward continuous spatial density estimation [13]. This density map-based modeling method mitigates occlusion-induced interference by bypassing the need for explicit segmentation of fragmented boundaries, enabling high-precision macroscopic perception even in extremely crowded environments [14]. Finally, image enhancement and domain adaptation technologies have been increasingly adopted to address visual perception failures caused by degradation in harsh agricultural environments [15].
For example, Yu et al. (2024) proposed a lightweight image enhancement network and a nighttime dairy cow detection method, eliminating light degradation interference [16]. Zhao et al. (2025) introduced an enhancement architecture with dynamic residual integration and ultra-lightweight upsampling modules to address rain and fog degradation. This approach overcomes harsh weather interference for strawberry maturity assessment [17]. Furthermore, Kaplan et al. (2026) developed a color correction-based image enhancement scheme for underwater aquaculture. This scheme corrects underwater imaging distortion and greatly improves target recognition systems [18]. In summary, traditional computer vision methods show significant limitations in complex and dynamic agricultural production environments. Current visual perception is undergoing a paradigm transition from single-task precision toward integrated optimization of physics-driven low-level visual restoration and high-level semantic perception [19]. This architecture integrates task-specific deep networks (e.g., low-light enhancement, image dehazing, and underwater restoration) with downstream perception tasks, featuring a clear division of labor and high targeted efficacy. Breaking these environmental bottlenecks is essential for achieving robust intelligent perception.
This review systematically covers four major categories of deep learning architectures that have driven recent developments in agricultural vision. Convolutional Neural Networks, including classic backbone networks and detection-specific variants, remain the workhorse for local pattern encoding and real-time edge deployment. Recurrent Neural Networks and Sequence Models are used to address temporal dependencies in growth monitoring and behavior analysis. Transformer-based architectures and their efficient successor, Vision Mamba, have emerged in recent years, providing global context modeling and linear time-scale scaling capabilities for high-resolution images. Generative models, including Generative Adversarial Networks and diffusion models, address data scarcity through synthetic augmentation and achieve zero-shot transfer. Furthermore, graph-based methods can dynamically model non-Euclidean relationships between discrete entities (e.g., livestock skeletons), while deep reinforcement learning shifts the paradigm from passive perception to active decision-making under uncertainty. This review selects these architectures based on two principles: (1) their prevalence in agricultural vision literature over the past five years; and (2) their effectiveness in addressing the three core challenges defining the field: non-rigid target deformation, severe occlusion, and environmental degradation.
Unlike previous reviews that categorized methods by crop type or application area, this review makes three contributions: (1) it constructs a progressive four-layer framework covering data foundation, static spatial cognition, dynamic spatiotemporal reasoning, and embodied deployment; (2) it critically compares different architectures within each layer, clarifying their advantages, limitations, and application scenarios; and (3) it systematically explores emerging visual foundation models and embodied intelligence. This review aims to reveal the mathematical unity of agricultural perception problems and provide practical guidance for future research.

2. Review Framework and Article Structure

This review aims to systematically survey and critically evaluate deep learning techniques for agricultural visual perception, focusing on how modern architectures address three persistent bottlenecks: non-rigid target deformation, severe occlusion, and environmental image degradation. Unlike previous reviews that categorize methods by crop species or application scenarios, this review offers three distinctive contributions: a progressive four-layer analytical framework spanning data foundation, static spatial cognition, dynamic spatiotemporal reasoning, and future outlook; critical comparisons of competing architectures within each layer to clarify their respective advantages, limitations, and application scenarios; and a systematic discussion of emerging vision foundation models and embodied intelligence, which have been largely overlooked in prior agricultural surveys. The remainder of this review is organized as follows. Section 3 covers the foundational data pipeline. It includes multi-modal imaging acquisition (Section 3.1), low-level enhancement under extreme conditions (Section 3.2), and cross-domain transfer learning for data-scarce scenarios (Section 3.3). Together, these three components form a coherent “acquisition-restoration-generalization” workflow, as depicted in Figure 1. Section 4 moves from data foundation to static spatial cognition. It examines three core perceptual tasks: non-rigid deformation detection, dense crowd counting, and fine-grained classification. Section 5 extends static perception to high-level spatiotemporal reasoning. It addresses short-term behavioral dynamics, long-term lifecycle analysis, heterogeneous fusion, and embodied decision-making. Section 6 synthesizes persistent challenges and future directions. Section 7 concludes the review.
Figure 1 presents the schematic framework for deep learning-driven closed-loop visual perception in precision agriculture, highlighting the integration of deep learning models within the perception cycle. This system demonstrates a four-layer progressive technological evolution, from bottom-level multimodal data acquisition, static spatial analysis, and dynamic spatio-temporal inference to top-level embodied execution and data feedback. (1) Multidimensional data acquisition sources: macroscopic satellite remote sensing (a) [20], microscopic multispectral imaging (b) [21], heterogeneous data of laboratory microscopic images (c) [22]. (2) Image enhancement under extreme environments: image dehazing (d) [23], low-light enhancement (e) [16], underwater reconstruction algorithm (f) [24]. (3) Solving non-rigid deformation problems: using key point skeleton estimation technology to reconstruct the topological structure of pigs (g) [25], deep feature extraction and coordinate regression architecture (h) [26]. (4) Solving extreme dense occlusion problems: late stage of wheat planting (i) [27], density problem in intensive fish farming (j) [28]. (5) Solving the problem of low interclass variance: introducing fine-grained visual classification (FGVC), focusing on microscopic local features to accurately distinguish crop varieties or disease categories with extremely similar appearances (k) [29]. (6) Cross-scale spatio-temporal dynamic cognition: covering high-frequency dynamic behavior localization at the millisecond/second level (e.g., tracking the high dynamic operation of greenhouse personnel/agricultural machinery) (l) [30], as well as long-term evolution and three-dimensional morphological simulation prediction of crops throughout the growing season (m) [31]. (7) Embodied Intelligent Closed-Loop Deployment: A large-scale vision-language foundation model (e.g., Vision-Language Model, Large Language Model) integrating prior agricultural knowledge is responsible for global semantic reasoning and command issuance (n) [32], as well as edge-side driving of agricultural automated machinery/robots to perform precise physical interactions (o) [33]. (8) Right-side Feedback Loop: The physical trial-and-error and interactive feedback of the robot in a real unstructured environment are transformed into high-value self-supervised data and continuously fed back to the underlying visual model, forming a closed-loop evolutionary ecosystem. As the foundational layer of the proposed framework (Figure 1), the following sections will first systematically discuss multi-modal data acquisition and low-level image enhancement techniques.

3. Multi-Modal Data and Low-Level Enhancement

This section establishes the foundational visual workflow for agricultural perception, comprising three stages that transform raw sensory input into actionable semantic information. First, we outline multi-scale, multimodal imaging methods (Section 3.1), ranging from satellite remote sensing to microscopic imaging. Second, we explore enhancement techniques (Section 3.2) for recovering degraded images under extreme conditions (rain, fog, low light, and underwater scattering). Third, we address data scarcity through cross-domain transfer learning (Section 3.3). These three parts constitute a coherent “acquisition–restoration–generalization” workflow, laying the foundation for subsequent perception tasks.
Agricultural visual perception relies heavily on specific imaging modalities designed for different spatial scales and specific perception tasks—from satellite-based regional crop monitoring to Unmanned Aerial Vehicle (UAV)-based plot-level perception and ground-based proximal inspection of crops and livestock. Satellite remote sensing and Unmanned Aerial Vehicle (UAV) remote sensing respectively construct complementary observation systems for macro-strategic monitoring and local fine-grained observation. Awais et al. (2022) confirmed through regression analysis that satellite imagery excels in regional farmland assessment, while UAVs capture local heterogeneous details lost by satellites [34]. Building a robust visual foundation for precision agriculture requires a systematic approach that combines high-fidelity multimodal data acquisition with targeted environmental degradation mitigation measures. Figure 2 illustrates the overall architecture and workflow of this process, depicting the step-by-step construction from raw multi-source input data to a high-quality visual foundation.
As shown in Figure 2, this framework integrates multi-source raw data input with task-specific enhancement modules. (1) Multimodal data acquisition sources: macroscopic remote sensing data for regional monitoring (a) [35], mesoscopic multispectral and thermal imaging data for crop phenotypic analysis (b) [36], and microscopic digital optical microscopy data (c) [37]. (2) Image enhancement under extreme environments: atmospheric dehazing of images (d) [38], low-light enhancement processing for enclosed aquaculture farms (e) [39], and underwater image restoration processing for aquatic environment perception (f) [25]. These modules work together to mitigate environmental degradation and ensure the robust visual input required for subsequent semantic analysis.

3.1. Multi-Modal Imaging Classification

Agricultural multi-modal imaging encompasses sensor technologies across multiple spatial scales—from satellite and UAV remote sensing to ground-based RGB-D and hyperspectral cameras, and down to microscopic imaging. Each modality offers distinct advantages: satellite and UAV platforms enable large-area monitoring, RGB-D provides depth information for 3D structure analysis, hyperspectral imaging captures biochemical signatures, and microscopy reveals cellular-level indicators. These modalities are often combined to compensate for individual limitations. Section 3.1 surveys representative studies, progressing from macro remote sensing to meso UAV-based sensing, ground-level RGB-D and hyperspectral imaging, and finally to ultra-microscopic imaging.
For macro-level monitoring, Cui et al. (2023) demonstrated that fusing phenological and multi-temporal Sentinel-2 spectral data significantly improves large-scale soil salinization estimation [40]. At the meso/micro level, unmanned aerial vehicles (UAVs) have high revisit rates and high resolution. Khodaei et al. (2026) used thermal infrared sensors carried by UAVs to accurately estimate dynamic parameters such as water temperature and dissolved oxygen [41]. Wei et al. (2023) fused UAV images with a convolutional neural network-long short-term memory network (CNN-LSTM) to achieve high-precision prediction of wheat yield before harvest [42]. To overcome the data discontinuity caused by weather and fixed satellite revisit cycles, Yu et al. (2026) proposed a multimodal generative adversarial network (GAN) that synthesizes high-fidelity and continuous multispectral images from a limited input to achieve weather-independent all-weather monitoring [43]. Ground-based sensing is rapidly evolving towards 3D and multi-source fusion. To overcome the limitations of 2D RGB (Red-Green-Blue) cameras in dense environments, Tang et al. (2024) combined RGB-D (Red-Green-Blue Depth) data with an improved Segmenting Objects by Locations version 2 (SOLOv2) method, successfully resolving overlap-induced segmentation and 3D positioning errors [44]. Similarly, Kang et al. (2025) utilized RGB-D features to accurately estimate the complex geometry of lettuce leaves [45]. This spatial geometry approach is also applicable to open fields. Niu et al. (2024) constructed a high-precision digital terrain model to improve the adaptability of UAV RGB images to corn height estimation [46]. Babalola et al. successfully fused RGB processing with Convolutional Neural Networks (CNNs) for soil texture classification [47]. Despite their low cost, broadband RGB cameras present perceptual blind spots for microscopic physiological tasks [48]. Hyperspectral imaging (HSI) overcomes this by capturing hundreds of narrow bands, providing “fingerprint” spectral data to distinguish crops with similar appearances but different internal compositions [49]. In terms of physiological monitoring, Zhao et al. (2022) used HSI to predict the canopy water content of lettuce [21]. Guzmán Q. et al. (2021) used a wavelet spectrum-based partial least squares regression (PLSR) model to predict leaf water content and achieved better prediction performance [50]. Lam et al. (2026) extended HSI to an airborne platform, enabling targeted application of pesticides for early detection of potato alternaria solani, thereby reducing the use of chemicals [51].
As precision agriculture shifts towards complex biochemical monitoring, single modalities become insufficient. Chen et al. (2020) combined a Visual Geometry Group 16-layer (VGG16) network with a long short-term memory (LSTM) model to achieve accurate recognition of aggressive episodes in pigs [52]. Tu et al. (2021) pointed out that pure visual information is easily affected by environmental interference in complex biochemical processes (e.g., potato browning), so it is necessary to combine machine vision with electronic nose olfactory data [53]. Consequently, multimodal fusion has become an inevitable trend. Ouyang et al. (2026) proposed a multimodal evaluation approach based on computer vision and near-infrared spectroscopy, integrating both qualitative classification and quantitative prediction during the black tea withering process [54]. Similarly, Zhao et al. (2026) developed a deep learning-based multimodal data fusion framework to address the challenge of real-time monitoring in solid-state fermentation, using lactic acid bacteria-fermented barley bran as the substrate [55].
In modern aquaculture, underwater perception is severely hindered by light scattering and absorption, rendering traditional RGB cameras unsuitable for non-invasive health assessment of farmed fish species such as salmon and tilapia [38]. To eliminate noise from turbid water, Picardi et al. (2025) designed an autonomous underwater vision system that uses a moderate optical correction algorithm for robust sorting [56]. Lu et al. (2026) developed Foreground-Panorama Prompt Learning for Zero-Shot Anomaly Detection in freshwater fish disease monitoring [57].
At the ultra-micro scale, macroscopic observations often miss the optimal window for disease control [58]. Li et al. (2023) utilized digital optical microscopy for spore morphology capture and ultra-early diagnosis [59]. For post-harvest storage, Luo et al. (2026) combined structured illumination reflectance imaging (SIRI) with an ultra-lightweight You Only Look Once version 11 with Spatial and Channel Reconstruction Convolution (YOLOv11-SSConv) model to detect weak decay features in citrus with high precision [60]. In field environments, Han et al. (2026) constructed a Multi-Scale Edge-Enhanced Lightweight Network (MSS-YOLO) framework with multi-scale feature fusion for Fusarium spore separation [22]. Finally, Wang et al. (2021) integrated a microfluidic chip with lensless diffraction imaging for portable, low-cost microscopic disease perception [61]. A comprehensive summary of these technologies applied in precision agriculture is presented in Table 1.
In summary, multimodal imaging follows a “scale adaptation” logic in actual production. At the macro level, satellite-UAV collaboration has been demonstrated for large-scale crop monitoring and disaster assessment, with Sentinel-2 time-series data enabling soil salinity mapping and UAV-based multispectral imagery supporting precision re-irrigation decisions [40]. At the meso level, RGB-D imaging has been embedded in automated platforms for crop-weed discrimination and 3D structure analysis in densely planted crops such as corn and greenhouse lettuce [46]. In the post-harvest stage, hyperspectral imaging has been applied to quality grading and disease detection, identifying subsurface lesions in potato and early blight in airborne platforms [51]. At the micro level, portable lensless diffraction imaging has enabled rapid spore detection and ultra-early disease diagnosis, supporting field-deployable plant protection [61]. Despite these advances, key gaps remain in multimodal data registration. Most studies process various modalities independently, with limited real-time spatiotemporal registration capabilities, especially between UAV images and ground sensors. Future research should prioritize hardware-synchronized multi-sensor platforms and lightweight cross-attention fusion architectures capable of dynamically weighting various modalities based on environmental reliability.

3.2. Low-Level Enhancement for Extreme Techniques

Image enhancement in precision agriculture addresses visual degradation caused by three environmental conditions: rain and fog in open fields, low light at dawn/dusk, and light scattering in underwater aquaculture. Section 3.2 reviews representative enhancement methods organized by degradation type.
Low-order image enhancement must be deeply integrated with downstream high-order perception tasks. Otherwise, directly feeding severely degraded images into subsequent networks introduces significant feature noise and impairs recognition accuracy [62].
Atmospheric scattering caused by outdoor rain and fog can lead to severe loss of visual details. Traditional dark channel prior (DCP) algorithms rely on fixed empirical formulas and often produce edge artifacts and color distortion under varying illumination conditions [63]. To overcome this problem, Gao et al. (2025) proposed an improved DCP algorithm that optimizes the computational architecture and successfully solves the problems of fog concentration variation and overexposure in the sky region, thereby significantly improving the performance of downstream detection [23]. To address the scarcity of paired training data and to coordinate the augmentation-detection task, Li et al. (2025) proposed the progressive augmentation network Lightweight Framework for PConv, EfficientNetV2, and ADown based YOLO (PED-YOLO) [64]. This method generates simulated paired data by transferring the model, achieving co-optimization and significantly accelerating the dehazing process. In addition, deep learning is shifting toward explicit physical guidance frameworks. Lyu et al. (2026) proposed embedding a physical degradation model of atmospheric scattering directly into the network topology to accurately decouple the fog region under non-uniform dense fog to achieve excellent color fidelity [65].
In low-light scenes, target features are masked by high-frequency noise. Traditional Retinex methods improve visibility but cause color distortion [66]. To address this issue, the Low-Light Neural Network (LLNet) first utilizes a deep autoencoder to jointly optimize brightness enhancement and adaptive denoising [67]. Yu et al. (2023) improved the U-shaped Net model by using the Convolutional Block Attention Module (CBAM) attention mechanism and adjusted the EnlightenGAN parameters, transitioning from supervised training to unsupervised training through transfer learning [68]. However, the large computational overhead of GANs limits real-time deployment. Consequently, the Zero-Reference Deep Curve Estimation (Zero-DCE) paradigm achieved extremely high computational efficiency through iterative nonlinear curve adjustments [69]. Extending highly efficient enhancement paradigms to crop phenotyping, Wang et al. (2026) introduced an image restoration framework to tackle low-light constraints (e.g., dawn/dusk field measurements), significantly improving the accuracy of maize Leaf Area Index (LAI) estimation under poor illumination [70].
In underwater imaging, suspended particles cause extreme color casts and image blur. To address these complex image degradation issues, Lin et al. (2021) proposed a multi-scale deformable convolutional network with an attention mechanism. This architecture utilizes a skip-fusion scheme based on attention, leveraging deformable local receptivity at different scales to reconstruct high-quality underwater images, thereby improving the performance of downstream visual tasks in fields such as aquaculture [71]. Early methods relied on simplified physical models [72]. To overcome this problem, Guo et al. (2026) introduced a texture map-guided multi-scale enhancement architecture that uses CNNs to jointly estimate background light and transmission maps to outperform traditional priors in heterogeneous environments [73].To reduce dependence on specific data distributions, Shao et al. (2026) proposed the physically guided Robust-SeaThru framework, which utilizes Graphics Processing Unit(GPU)-accelerated vectorized mesh search to significantly improve generalization robustness [74]. Since classical mathematical models fail under non-uniform illumination, Pérez-Zarate (2025) et al. proposed the Underwater Non-uniform Illumination Restoration Network (UNIR-Net) for non-uniform illumination restoration, improving downstream semantic segmentation [24]. For underwater depth estimation (e.g., key capabilities for autonomous aquaculture net inspection and underwater robot navigation), Lv et al. (2026) proposed the RGB-Sonar Fusion Network (RSFNet) for heterogeneous fusion of RGB and sonar data using Transformer-based global context modeling [75]. Ultimately, these underlying restoration technologies provide the visual foundation for autonomous underwater operations [76]. A summary of the representative deep learning architectures for environmental enhancement is presented in Table 2.
In summary, image enhancement in extreme agricultural environments follows a logic targeting specific degradation. For open fields shrouded in fog, the PED-YOLO framework supports real-time defogging by drones, enabling precise spraying [64]. Physically embedded defogging methods can improve color fidelity under non-uniform dense fog conditions [65]. Under low light conditions, Zero-DCE iterative enhancement technology supports nighttime greenhouse monitoring through leaf area index inversion [69]. A dedicated image restoration framework further enhances the effectiveness of maize phenotypic analysis at dawn/dusk [70]. For underwater aquaculture, texture-guided restoration technology can restore turbid images for salmon lesion detection [73]. However, autonomous operation requires other functions. Specifically, RSFNet fuses RGB images and sonar images for depth estimation in fishing net inspection [75]. Robust-SeaThru ensures the algorithm’s generalization ability across different water types [74]. UNIR-Net handles non-uniform lighting to improve image segmentation [24]. Despite these advances, the models remain task-specific, standard metrics often fail to match downstream semantic accuracy, and the coupling between augmentation and perception remains weak. Therefore, future research should focus on task-driven joint optimization and adaptive architectures to address various degradation types without requiring retraining.

3.3. Cross-Domain Transfer Learning

Cross-domain transfer learning addresses data scarcity and domain shift in agricultural vision. Section 3.3 reviews strategies ranging from fine-tuning to unsupervised adaptation and vision-language integration.
Data scarcity and domain shift caused by variable weather, long growth cycles and high labeling costs remain the core bottlenecks of agricultural artificial intelligence [77]. Although transfer learning usually relies on fine-tuning ImageNet pre-trained models [78], improving its generalization ability largely depends on the visual similarity between the source model and the target model [79]. Hussin et al. (2026) identified a “domain impedance” phenomenon in microalgae and plant pathogen image classification, demonstrating that general ImageNet models retain too many macroscopic features and yield limited generalization gains for microscopic agricultural tasks such as spore recognition and cellular-level disease diagnosis [80].
To break this bottleneck, modern research focuses on domain-specific architectures. Qiao et al. (2026) constructed the Multi-scale Dilated Convolutional Visual Geometry Group Network 16-layer (MDCVggNet16) for wheat pest and disease recognition, introducing multi-scale dilated convolution and CBAM attention into ImageNet pre-training to focus on weak lesion features and suppress complex field backgrounds [81]. Hukkeri et al. (2024) systematically evaluated pre-trained CNNs, confirming that EfficientNet exhibits superior parameter efficiency when underlying general features are frozen and high-level specific features are fine-tuned [82].
Furthermore, unsupervised domain adaptation techniques based on GANs offer new solutions for scarce annotations and edge deployment [83]. Wen et al. (2025) used unpaired style transfer from the Cycle-Consistent Generative Adversarial Network (CycleGAN) to increase training sample diversity, constructing a lightweight YOLO version 5 with MobileNet and Segmentation (YOLOv5-Mobile-Seg network) that reduces computational load while maintaining high edge accuracy [84]. Ranario et al. (2026) extended style transfer to cross-modal alignment, enabling RGB-trained models to generalize to zero-annotation thermal scenarios, decoupling the scarcity of multimodal annotations [36].
Because explosively growing agricultural video streams cannot be manually labeled in time, self-supervised and semi-supervised learning are becoming cutting-edge paradigms [85]. Jin et al. (2026) proposed the Momentum Contrast Prototypical Network (MoCoProto) framework, performing large-scale contrastive learning on over 350,000 unlabeled images followed by meta-learning for few-shot classification [86].
Besides classification, Yang et al. (2026) proposed a collaborative framework integrating Convolutional Neural Networks (CNNs), Vision-Language Models (VLMs), and Large Language Models (LLMs) [87]. By using CycleGAN for unsupervised style transfer, this framework bridges the domain gap between laboratories and farmlands. It systematically links low-order detection (e.g., lesion spotting), mid-order semantic description (e.g., disease type reporting), and high-order embodied decision-making (e.g., robotic spraying), enabling end-to-end autonomous disease control in field conditions. Building on these advanced architectures, Musazade et al. (2026) further integrated VLMs and LLMs into plant phenotyping [32].
The scarcity of labeled data is a major obstacle to model deployment. For emerging pathogens, self-supervised comparative frameworks (such as MoCoProto) can build small-sample classifiers within 24 h using only 50–100 field images, supporting emergency control. Domain-specific pre-trained models (such as EfficientNet) are packaged into mobile applications, allowing for fine-tuning of local models during flight and achieving over 90% accuracy in files smaller than 50 MB. Furthermore, drone service providers are leveraging CycleGAN’s unpaired style transfer to synthesize thermal infrared images from labeled RGB images, eliminating the need for expensive thermal imaging annotations for nighttime livestock screening. The CNN-VLM-LLM collaborative framework is being piloted in orchards in Europe and the US, enabling farmers to generate geotagged application maps via voice queries, thus promoting the widespread adoption of AI-based plant protection technologies.
In summary, Section 3 has established the foundational data pipeline for agricultural visual sensing. Multi-modal imaging provides the raw sensory inputs across spatial scales, low-level enhancement ensures their quality under extreme conditions, and cross-domain transfer learning extends their usability under label scarcity. Collectively, these three components constitute the “acquisition-restoration-generalization” workflow previewed in Figure 1 and the Introduction. This data foundation lays the groundwork for the spatial cognition tasks to be discussed in Section 4.

4. Static Spatial Cognition

In agricultural visual perception, the physical characteristics of biological targets and the complexity of unstructured environments present unique spatial challenges. This section analyzes the inherent limitations of traditional rigid bounding box detection methods when applied to non-rigid morphologies and cluttered scenes of agricultural organisms, and explores how modern deep learning architectures have evolved to overcome these bottlenecks through geometrically adaptive algorithms that handle severe occlusion, extreme crowding, and low inter-class variance.
Traditional bounding box detection paradigms rely on strict geometric assumptions, which fail to adapt to the highly non-rigid morphology of agricultural organisms (e.g., curved leaves, irregular lesions, and intertwined livestock) [88]. These rigid localization mechanisms remain sensitive to environmental clutter. For instance, in densely housed pig groups, overlapping and occlusion during feeding can cause detection accuracy to plummet [89]. Although deep learning models perform well in controlled laboratory environments, they lack robustness under intense light fluctuations and complex background clutter inherent in open-field scenes [90]. A comprehensive summary of the evolving static cognitive paradigms is presented in Table 3.

4.1. Non-Rigid Deformation Detection

To overcome these perception bottlenecks, deep learning architectures featuring spatial extraction and spatio-temporal modeling have been introduced. These architectures have profoundly accelerated research progress in agricultural pest and disease detection [91].
In field and greenhouse planting environments, crop and weed morphologies exhibit pronounced non-rigid deformations due to varying growth cycles, environmental stress, and mechanical disturbances, posing direct challenges for precision spraying and mechanical weeding operations that require accurate plant localization [92]. Traditional rectangular positioning methods frequently fail under these conditions, particularly when leaf curling, stem bending, or plant entanglement occurs [93].
To enhance feature discriminative power in competitive growth environments, Wang et al. (2021) combined multi-feature fusion technology with Back Propagation (BP) neural networks for weed-crop discrimination in asparagus fields, overcoming the limitations of single-modality visual features under varying light and soil background conditions [94]. Peng et al. (2021) extended this strategy to the deep feature level, leveraging multi-stream deep features to drive a support vector machine (SVM) for accurate grape leaf disease identification, enabling early fungicide intervention and reducing reliance on expert visual scouting [95].
To accurately locate disease regions in complex backgrounds, Zuo et al. (2022) proposed a multi-granularity feature aggregation method for crop disease recognition in complex field backgrounds. By employing block-level channel self-attention to enhance species-specific features, they simultaneously introduced a spatial reasoning module to model the spatial geometric relationships between image patches, thereby improving lesion localization accuracy for early disease warning [37]. Furthermore, Lu et al. (2024) introduced a geometry-based balanced feature amplification method, which improved the detection stability of YOLOv7 under shading and complex light variations, ensuring reliable fruit localization for robotic picking under fluctuating natural illumination [26].
To address intertwined branches in dense peach orchards, Tao et al. (2022) proposed an improved CNN incorporating attention mechanisms and multi-scale feature fusion to achieve accurate flower counting, enabling early-season yield forecasting and informed thinning decisions [96]. Abbas et al. (2021) systematically evaluated four mainstream CNN models for strawberry leaf scorch in uncontrolled fields, confirming that EfficientNet-B3 exhibited the best performance across all disease stages, supporting field-deployable disease surveillance without requiring laboratory-based diagnosis [97]. Regarding algorithm architectures, Ji et al. (2025) proposed a detection method combining multi-dimensional feature extraction with a Detection Transformer (DETR) framework for green apple detection under orchard canopy occlusion, achieving robust fruit localization without relying on hand-crafted anchor designs [98].
Livestock husbandry faces similar challenges, where non-rigid deformations and high-density crowding make it difficult for traditional rectangular boxes to distinguish individual identities [99]. Consequently, the community is shifting toward keypoint pose estimation and spatio-temporal modeling. Chen et al. (2020) proposed a spatio-temporal feature extraction method that combines CNN and long short-term memory networks (LSTM), which solves the misjudgment problem caused by overlapping pig bodies and playing with waterers in group rearing environments, and finally achieves accurate classification of the actual drinking and playing behaviors of pigs [100]. For lameness detection in dairy cows, Russello et al. (2024) developed the Temporal LEAP(T-LEAP) framework, which can accurately extract the trajectories of nine skeletal key points. This framework calculates six core kinematic indices (e.g., back posture and stride), effectively eliminating the subjective bias of human scoring [101]. Chen et al. (2026) pushed pig behavior recognition to a deeper semantic level by combining YOLOv8 with Depth Channel Temporal Module-UniformerV2(DCTM-UniformerV2), enabling accurate individual detection that decouples genuine aggressive acts from other confounding interactions [102].
For robotic harvesting, He et al. (2025) combined 2D instance segmentation with 3D point cloud pose estimation for broccoli harvesting robots, using spherical fitting and Basis Spline (B-spline) surface reconstruction to reverse the stem orientation and cutting points, enabling automated selective harvesting with minimal crop damage [103]. Finally, Bumbalek et al. (2025) systematically benchmarked multiple generations of YOLO models for individual dairy cow identification in commercial barns, finding that YOLOv12 Medium (YOLOv12m) offers the highest accuracy under severe occlusion—a critical requirement for automated feeding monitoring and individual health tracking in group-housed herds [104]. Henrich et al. constructed the PigDetect and PigTrack datasets to standardize the systematic evaluation of advanced Multiple Object Tracking (MOT) algorithms in dynamic breeding environments [105].

4.2. Dense Crowd Counting

When targets are highly stacked or highly compressed, the individual boundaries become visually indistinguishable, rendering the traditional “count after detection” paradigm logically invalid [29]. Density map regression addresses this problem by generating a crowded spatial heatmap in which the probability of each pixel belonging to a target is evaluated, and the total count is derived by integration [106].
For large-scale counting, especially in high-density scenarios where individual boundaries are indistinguishable due to severe overlapping, shifting from direct target detection to count prediction via multi-feature fusion and regression models has become a critical development direction [107]. Early multi-column convolutional networks (MCNNs) suffered from problems such as bloated architecture and excessive computational overhead [108]. However, the Multi-Scale Feature Enhancement Network (MFNet) developed by Qian et al. (2024) effectively addressed these limitations for in-field wheat ear counting by introducing a deformable spatial attention mechanism (DSAM) to suppress complex soil and straw backgrounds, improving feature extraction under varying head densities and growth stages [27]. Sun et al. (2022) further demonstrated that combining a single-branch CNN with dilated convolution and multi-scale feature fusion can effectively solve the problems of scale variation and severe occlusion, and significantly reduce computational cost [109]. In addition to counting, Xu et al. (2026) proposed a lightweight oriented bounding box (OBB) detection model for underwater fish behavior analysis in recirculating aquaculture systems, addressing centroid shift and background mixing caused by fish body bending during swimming, which is essential for accurate individual tracking and feeding behavior quantification [110]. Modern networks can physically segment continuous density maps by integrating spatial topology segmentation algorithms (e.g., watershed algorithm), thereby inferring each center point and overcoming the traditional limitation that density maps can only be counted [111].
Despite the strong performance of density maps, they often rely on manual annotation. To reduce labeling costs, simple point annotations—where only the center point of each target is clicked—have been widely adopted. Yang et al. (2023) demonstrated that such point-based labels can effectively train cross-platform wheat ear counting models across UAV and ground-based imaging systems, facilitating field data collection for breeding trials and yield mapping at extremely low cost [112]. Ji et al. (2026) demonstrated that their weakly supervised method effectively reduces labeling costs for high-density fish counting in aquaculture by combining a small number of point labels with image-level density labels, thus facilitating routine population estimation without intensive manual annotation [28]. In addition, Zhang et al. (2025) compared 2D and 3D vision for quality grading of granular agricultural products such as grains and seeds, highlighting how 3D structured light reconstruction can mitigate the depth ambiguity inherent in 2D vision by directly measuring volume and thickness—key parameters for post-harvest sorting and premium pricing [113].

4.3. Fine-Grained Classification and Anomaly Screening

Fine-grained visual classification and anomaly screening are complementary core technologies [114]. Unlike coarse-grained recognition, agricultural FGVC often deals with high intra-class variance and extremely low inter-class variance, necessitating a focus on highly localized, weak visual details.
To identify Fuji apple varieties for post-harvest sorting and authenticity verification, Guo et al. fused MobileNetV2 with Gray-Level Co-occurrence Matrix (GLCM) texture features and a multi-head attention mechanism, systematically verifying that local micro-texture is the core decision-making factor in variety identification [115]. Similarly, Ray et al. (2025) combined Local Binary Pattern (LBP) texture features, GLCM statistics and Hue, Saturation, Value (HSV) color histograms for guava leaf disease detection, revealing how SVM kernel selection affects recognition performance under varying field illumination [116].
As the focus shifts from species classification to individual tracing, large-model paradigms have become essential. Quan et al. (2026) proposed a two-stage fine-tuning paradigm using the Contrastive Language-Image Pre-training (CLIP) large model for individual fish re-identification in aquaculture tanks, enabling tracking of specific individuals for health monitoring and breeding management [117]. Finally, Wei et al. proposed an ultra-lightweight Multi-scale Adaptive Screening and Matching YOLO(MASM-YOLO) model for bovine behavior screening on edge devices in pasture environments, which is designed with a multi-scale focusing network and an adaptive decomposition alignment head. This model achieves an optimal Pareto balance between computational power and perception performance, providing a feasible hardware solution for real-time disease anomaly warning in unstructured environments [118].
Figure 3 illustrates the paradigm shift from traditional detection methods, which are constrained by rigid boundaries and discrete segmentation and struggle with non-rigid deformation (a) [119], dense overlapping (b) [109], and visual ambiguity (c) [120], to advanced adaptive cognitive models. These include keypoint pose estimation for morphological adaptation (d) [119], density map regression for robust counting in high-density occluded scenes (e) [109], and attention-based fine-grained feature aggregation for precise diagnostic screening (f) [120].
In summary, these three static spatial cognition paradigms address different agricultural challenges, but each paradigm has inherent trade-offs that determine its application scenarios. Keypoint-based pose estimation offers the highest morphological accuracy, making it an indispensable tool for robotic manipulation (e.g., harvesting, weeding) and behavioral understanding (e.g., livestock lameness) [103]. However, it relies on expensive joint-level annotation and high computational overhead, thus limiting its application to high-value crops and high-end livestock farming where unit profits can offset its costs [98]. Density map regression excels in high-throughput population estimation, especially in situations where individual detection is impossible due to extreme crowding. Its low annotation cost and computational efficiency make it ideal for large-scale yield forecasting, breeding trials, and aquaculture resource management [112]. Its fundamental drawback is the loss of individual identity information, making it unsuitable for tasks requiring instance-by-instance health recording or robot interaction [111]. Fine-Grained Visual Classification (FGVC) offers the highest semantic resolution in distinguishing visually similar varieties, detecting subtle lesions, and re-identifying individual animals [117]. However, it is extremely sensitive to image quality; motion blur, insufficient lighting, and occlusion can easily destroy the local microtexture it relies on [116]. Therefore, its deployment is limited to controlled environments, such as post-harvest grading lines and semi-enclosed barns [104]. In practice, the choice of these paradigms should depend on the primary mission objective: physical interactions tend towards attitude estimation, population censuses towards density regression, and quality certification towards FGVC. Hybrid approaches combining two or more paradigms are emerging, but their computational cost remains prohibitive for most edge deployments [118].

5. High-Level Spatio-Temporal Cognition

Relying solely on static spatial features is insufficient to capture complex and continuous changes in biological targets and the environment. This section transitions from static spatial extraction to advanced spatiotemporal cognition. Section 5.1 traces the evolution of spatiotemporal architectures, exploring dynamic operators and active perception loops. Section 5.2 discusses lifecycle refinement and global analysis, comparing short-term kinematic anomalies to long-term seasonal modeling. Section 5.3 examines cross-domain coupling and adaptive decision-making, focusing on DRL and vision foundation models. Table 4 provides a comprehensive summary of these paradigms.

5.1. Evolution of Spatio-Temporal Architectures

Traditional convolution operators with fixed sampling grids struggle with agricultural morphological diversity, including leaf curling in crops and dynamic occlusion in livestock [121]. Consequently, dynamic spatial perception and spatio-temporal embodied decision-making have emerged as the core drivers for advancing precision agricultural AI. To overcome these rigid geometric constraints, the prioritization of dynamic operators with inherent morphological adaptation is being adopted by advanced methodologies [122]. For tea quality monitoring during processing, Li et al. (2025) integrated deformable convolution into a ResNet architecture to adaptively capture non-rigid leaf deformations, embedding an attention module for dynamic spatio-temporal perception [123]. In addition, Li et al. (2024) proposed the Compact Dynamic Faster Network (DyFasterNet) architecture, which dynamically aggregates multi-scale kernel functions and utilizes a deformable attention mechanism to capture irregular variations in fruit size. The results show that, when combined with Inner Intersection over Union (Inner-IoU) Loss, Frequency Modulation-based Wise Intersection over Union (Wise-IoU), and semantic frequency cue distillation, the method enables lightweight edge deployment of high-precision models [124].
When extreme occlusion exceeds algorithmic restoration limits, proactive intervention becomes necessary. Huang et al. (2026) proposed YOLOv8 with fuzzy attention enhancement for cucumber harvesting under dense foliage, quantifying occlusion uncertainty and driving an intelligent end effector that performs physical leaf-pushing to overcome blind spots—constructing a physical-level “perception-cognition-action” closed loop [125]. Similarly, Zhang et al. (2025) raised the robustness ceiling for complex background tasks through the Efficient Multi-scale Attention with Grouping and Encoder–Decoder Swin (ED-Swin) Transformer framework [126].
Biological behavior is essentially a continuous temporal evolutionary process, rather than a series of isolated frames. Liu et al. (2026) used a combination of One-Dimensional Convolutional Neural Network (1D-CNN) and LSTM architectures to predict the long-term crop coefficient of spring maize [127]. Subsequently, Bahrami et al. (2026) extended this paradigm to macro-coastal evolution modeling [128]. For fine behavior detection in precision spraying operations, Pu et al. (2025) utilized YOLOv8-n to localize targets, introducing a spatio-temporal graph convolutional network (ST-GCN) to model temporal topological features for accurate spraying behavior recognition, enabling real-time monitoring of sprayer operator performance [30]. Against this backdrop, a spatio-temporal integrated analysis framework has emerged. It effectively coordinates spatial coordinate perception, motion offset tracking, and attention filtering mechanisms, ultimately achieving highly sensitive recognition [129]. Zhao et al. (2025) proposed a ResNet50-LSTM fusion model, demonstrating that multi-source heterogeneous feature fusion significantly outperforms single visual data sources in monitoring fermentation physicochemical dynamics [17]. Finally, Zheng et al. (2026) constructed a three-level “detect-track-trigger” pipeline for tree-trunk detection in orchard automation, converting continuous perception streams into discrete hardware execution events for selective pruning or targeted spraying [130].
Sumana et al. (2026) demonstrated that Partially Observable Markov Decision Processes (POMDPs)-based sampling platforms enable autonomous underwater vehicles to navigate unknown environments, achieving high-precision spatial flow field reconstruction for precision feeding and cage management with minimal physical and energy costs [131]. This trajectory marks a disruptive shift in agricultural visual intelligence, moving from unstructured environmental perception to adaptive cognition and embodied AI decision-making through physical intervention [132].
In summary, current agricultural perception has not yet formed a universal temporal network architecture, but rather focuses more on adapting to specific local environments at different time scales. Therefore, this field is undergoing a fundamental technological leap, gradually transitioning from isolated single-frame detection to a new paradigm of continuous spatio-temporal understanding and unified proactive perception-decision-action closed loop.

5.2. Lifecycle Refinement and Global Analysis

In video-based agricultural perception, short-term emergencies—such as livestock birthing warnings or equipment malfunctions—require analyzing instantaneous kinematic deformations rather than static appearances [133]. While optical flow and ST-GCNs pioneered this field, the paradigm has been fundamentally reshaped by Transformer-based architectures. Their global receptive fields excel at capturing non-local spatiotemporal dependencies in noisy farm environments.
For fine-grained motion perception, the evolution from variational frameworks to end-to-end architectures, such as FlowNet, RAFT, and Transformer-based FlowFormer, has established a solid methodological foundation [134]. Building upon this, modern video Transformers with self-attention mechanisms now directly reconstruct occluded parturition postures, bridging pixel-level motion and high-level pathological semantics. Wu et al. (2026) proposed DA-MOT for dairy cow tracking, integrating an enhanced YOLO detector with an angle-aware Kalman tracker, while systematically reviewing Transformer-based tracking paradigms [135]. Wang et al. (2026) developed a three-stage framework employing Video Masked Autoencoder V2 (VideoMAE V2) for feature extraction and ActionMamba for temporal localization; their benchmarking against Transformer-based methods like ActionFormer demonstrates the complementary strengths of Transformers and SSM architectures for agricultural video analysis [136].
Beyond second-level emergencies, season-scale modeling supports long-term system decision-making. Lakhirar et al. (2024) highlighted the role of Internet of Things-based (IOT-based) closed-loop irrigation systems in managing water resources throughout the growing season [137]. Wang et al. (2025) demonstrated a paradigm shift in fertilization strategies from experience-driven to data-driven [138]. Critically, this data-driven transition is now accelerated by Transformer-based time-series forecasting models, which dynamically reweight historical sensor data across phenological stages to optimize water-fertilizer coupling. For behavioral recognition, Yuan et al. (2026) reviewed the spatial dimensionality upgrade of cattle behavior recognition from 2D CNNs to 3D Transformers, identifying poor cross-scenario generalization and insufficient fine-grained quantification as the primary bottlenecks for large-scale production [139]. These bottlenecks are being actively tackled by foundation vision Transformers and hybrid CNN-Transformer architectures, which leverage large-scale pre-training and demonstrate remarkable few-shot adaptability to novel farm layouts. As a spatial complement, Bah et al. (2023) proposed an unsupervised hierarchical graph representation for UAV crop row detection, inferring global topology from individual plant arrangements [140]. Recent extensions incorporate Transformer-based spatial topology inference to dynamically track structural deformations caused by wind or mechanical operations, ensuring robust long-term field structure monitoring.

5.3. Cross-Domain Coupling and Adaptive Decisions

In precision agriculture, crop phenological evolution is slow and subtle, making it difficult for traditional visual algorithms to detect such gradual state changes. Therefore, generating temporally continuous and interpretable simulations of crop phenology remains a core challenge, as purely data-driven artificial intelligence models often lack consideration for underlying growth processes [141].
To address this highly dynamic and unstructured environment, deep reinforcement learning (DRL) has become a key solution. Zhao et al. (2025) demonstrated that DRL can effectively overcome the performance bottlenecks of agricultural machinery, especially in dynamic path planning and multi-machine collaboration [142]. Similarly, in facility agriculture, modern DRL-based methods have achieved a historic shift from traditional single-setpoint tracking (e.g., proportional-integral-derivative, model predictive control) to multi-objective global collaborative optimization in greenhouse management [143].
With its trial-and-error and policy adaptation capabilities, DRL breaks down barriers to cross-scenario macroscopic control [144]. A study by Rajagukguk et al. (2025) reviewed the application of deep learning in livestock monitoring, emphasizing that high-level behavioral classification in complex farm environments is highly dependent on the stability of previous low-level tasks (e.g., detection and tracking) [145]. Once the macro-technical landscape is established, building a scalable and efficient spatiotemporal modeling foundation becomes the core driver of technology implementation.
At the remote sensing level, a synchronous spatiotemporal downsampling strategy was designed by Li et al. (2025) to achieve the comprehensive unified modeling of cross-platform, multi-source satellite data [146]. Thus, a highly scalable spatiotemporal base for large-scale agricultural mapping was provided. The efficiency of spatiotemporal feature extraction for large-scale crop identification was optimized by Feng et al. (2026), greatly satisfying the computing power requirements for all-weather real-time monitoring [147]. Furthermore, the generalization performance of remote sensing large models under diverse scenarios and complex tasks was systematically evaluated by Zhang et al. (2024). The underlying shortcomings of general models regarding spatial geometric reasoning and fine-grained agricultural identification were profoundly revealed [148]. In response to this bottleneck, the largest multidimensional wheat image dataset globally was constructed by Han et al. (2025). A wheat-specific vision foundation model was successfully trained using self-supervised learning, providing a powerful domain-specific engine for organ-level micro-dynamic change analysis [149].
As perception engines are developed toward specialization, the final form of temporal cognition has evolved from “post-event passive diagnosis” to “high-dynamic real-time positioning and pre-event prediction of future forms.” Regarding high-dynamic video behavior localization, an adaptive deep spatiotemporal network integrating RGB and high-frequency optical flow was proposed by Yan et al. (2024) [150]. By designing a dual attention mechanism of modality and temporality, the absolute start and end times of pig aggression behavior in untrimmed complex videos were accurately located. Consequently, livestock behavior analysis was successfully promoted from simple “fragment classification” to a new stage of “complete temporal semantic understanding”.
From post-event response to pre-event proactive prediction, dynamic modeling has driven breakthroughs at different spatial scales. Delving into the micro-pixel level, Wang et al. (2022) constructed a Spatiotemporal Long Short-Term Memory (ST-LSTM) and Memory-In-Memory (MIM) network architecture. This method successfully realized dynamic simulation and prediction of the future morphological evolution of plant organs based on time-series images. Through these multi-scale advances, plant phenotypic analysis has been formally elevated from “static geometric measurement” to a new scientific height of “dynamic growth visualization prediction” [151]. Meanwhile, Lwaho et al. (2026) used a dynamic regression model with AutoRegressive Integrated Moving Average (ARIMA) error to achieve highly interpretable dynamic prediction of maize yield by deeply integrating external multivariate variables including climate fluctuations and fertilization intensity [152]. A comprehensive summary of the evolving dynamic spatiotemporal cognitive paradigms, which systematically address challenges ranging from kinematic feature extraction to autonomous embodied decision-making, is presented in Table 4.
Table 4. Evolution of dynamic spatiotemporal cognitive paradigms and their algorithmic architectures.
Table 4. Evolution of dynamic spatiotemporal cognitive paradigms and their algorithmic architectures.
Cognitive ParadigmCore Technical ObjectivesKey Algorithmic ArchitecturesRef.
Spatiotemporal AdaptationResolving geometric constraints
Handle non-rigid morphological changes
Deformable convolution
Attention mechanisms
DyFasterNet
[121,122,146]
Active Perception LoopOvercoming occlusion via “perception-cognition-action”Fuzzy attention
Hybrid visual servoing
[125,130]
Continuous ModelingModeling behavior as continuous temporal evolution1D-CNN + LSTM
ST-GCN
High-frequency optical flow
[127,128,147,149,151]
Heterogeneous FusionCoupling multimodal data for robust inferenceResNet50-LSTM
Wasserstein Generative Adversarial Network
Kalman filters
[17,129]
Kinematic Feature ExtractionCapturing instantaneous velocity & trajectoryOptical Flow
Pyramid, Warping, and Cost volume Network
RAFT
FlowFormer
[133,134,135,150]
Lifecycle AnalysisLong-term series modeling for managementIoT-AI Integration
Convolutional Long Short-Term Memory
ARIMA
[136,137,138,139,140,145,152]
Embodied AI & OptimizationAutonomous exploration & cross-modal decision makingMulti-modal Large Language Model
DRL
Transformer-based spatiotemporal reasoning
[131,132,141,142,143,144,148]
These four spatiotemporal architectures exhibit complementary advantages and disadvantages in practice. CNN-LSTM, which combines convolutional spatial extraction and cyclic temporal modeling, excels at handling regular grid data of moderate sequence length (e.g., RGB video, multispectral time series), making it suitable for growth stage classification and fermentation monitoring [17]. However, its sequential looping limits parallelization in the temporal dimension, and it cannot model non-Euclidean relationships (e.g., skeletal joints).
Graph-based methods (e.g., ST-GCN) compensate for this deficiency by handling topological representations. They achieve superior parameter efficiency in the dynamics of relationships between discrete entities (e.g., livestock skeletons, plant rows, or multi-robot teams), making them ideal for behavior recognition and pose analysis [140]. Their key bottleneck lies in their dependence on detectors; that is, the graph collapses if the detector fails due to occlusion [125].
Transformer models overcome the sequential bottleneck of LSTM and the local receptive field limitations of CNNs through global self-attention to spatiotemporal markers. They perform well in large-scale farmland mapping and pre-training of base models [149]. However, quadratic complexity limits their deployment in edge environments, and they have high data requirements and inference latency. Currently, Transformer models remain an offline, cloud-based solution [147].
Deep reinforcement learning (DRL) shifts focus from perception to sequential decision-making under uncertainty. It is particularly suitable for autonomous navigation, multi-robot collaboration, and dynamic greenhouse control—tasks where optimal action depends on long-term rewards [142]. However, low sample efficiency, the gap between simulation and reality, and barriers to reward engineering limit its current deployment, confining it to simulation-validated path planning and controlled environments [131].
In practice, the following hierarchy is chosen: CNN-LSTM for fixed-camera mesh structure monitoring; graph-based models for skeleton/relationship dynamics; Transformer for cloud-based global modeling when data allow; and DRL only when sequential policy optimization and simulators are available.

6. Challenges and Future Directions

Agricultural visual perception has achieved remarkable progress, yet several persistent challenges remain unresolved before these technologies can be widely deployed in real-world farming. This section synthesizes these challenges and outlines future directions across three interconnected dimensions: domain generalization, lightweight deployment, and embodied intelligence. These three dimensions form a logical progression from ensuring models work across diverse field conditions, through enabling cost-effective edge deployment, to closing the loop from perception to autonomous action.

6.1. Universal Domain Generalization

Agricultural visual perception models are gradually moving from the laboratory to the complex and ever-changing open fields. The core challenge is no longer just improving baseline accuracy, but also maintaining robustness in constantly changing environments while achieving low-cost engineering under strict boundary constraints.
Open farmland is a highly dynamic system where light, temperature, humidity, and crop phenotypes undergo irreversible, continuous evolution over time. Although recent continuous testing time-domain adaptation methods have shown promise in addressing the dynamic drift problem, these methods are largely based on theoretical assumptions and lack validation in long-term field data spanning the entire crop growth cycle. An even more serious challenge lies in long-tailed open-set identification. Unknown disease mutations can appear at any time in farmland. Dong et al. (2025) also warned that the standard classification model practice of forcibly classifying unknown categories into known categories poses a significant risk in actual decision-making [153]. However, most existing methods remain at the post-processing or threshold setting level and fail to establish inherent robustness for rare agricultural samples.
Despite progress in domain adaptation, several weaknesses persist. Most studies emphasize benchmark accuracy while overlooking robustness, calibration, and interpretability. Models trained under controlled conditions often degrade in real farms where lighting, weather, and noise violate standard assumptions. Test-time adaptation methods frequently suffer from instability and catastrophic forgetting of domain-invariant knowledge. The domain gap in agriculture is particularly severe—models perform well on source domains but drop significantly when applied to unmanned aerial vehicle imagery or different seasons and varieties.
Looking ahead, research should focus on developing sufficiently robust test-time adaptive protocols to handle multi-season field data, striking a balance between adapting to changing environmental conditions and maintaining the stability of core knowledge. Foundational models tailored for agricultural vision, such as the emerging Agri-FM+, consistently outperform ImageNet pre-trained models even with limited labeled data. In the long term, directly embedding biological growth patterns and the physics of optical imaging into the model architecture holds promise for achieving true zero-shot generalization capabilities for previously unseen agricultural scenarios.

6.2. Multi-Objective Co-Optimization

The practical application of agricultural visual perception technology often faces challenges in balancing engineering feasibility and economic feasibility. Due to limitations in computing power, power consumption, and network latency, large cloud models cannot be deployed directly. Currently, research on lightweight deployment has made some progress. For example, Nugroho et al. (2024) even quantized the model to 8 bits and ran it on a microcontroller costing hundreds of dollars [154]. Gao et al. (2025) deployed an improved YOLO algorithm on a Jetson Nano to detect wheat scab [155]. Similarly, highly optimized lightweight YOLOv5 models have been developed. These models enable rapid real-time detection of tea shoots under complex field conditions [156]. However, existing literature often separates “model compression” from “economic feasibility.” As Shao et al. (2022) pointed out in their empirical analysis of cotton fields in Xinjiang, fixed equipment costs, labor costs, and hydropower costs are key factors restricting the promotion of this technology [157].
Despite these advances, deployment on edge devices remains challenging due to the fundamental conflict between detection accuracy and computational overhead. Existing models have an excessive number of parameters, making them difficult to deploy on resource-constrained devices. Compression strategies often prioritize parameter reduction at the expense of accuracy and lack end-to-end hardware-software co-design. Most research focuses on accuracy under controlled conditions rather than field-level comparative evaluations under complex backgrounds, varying lighting, and occlusion.
Next-generation lightweight deployments are moving towards more complex strategies. In the short term, system evaluations of YOLOv8n to YOLOv12n in embedded frameworks show that YOLOv12n achieves the best balance between accuracy and efficiency on platforms such as NVIDIA Jetson Nano devices. Novel architectures such as CEG-YOLO achieve an average accuracy of 93.9% on edge devices at 20 frames per second with only 7.8 million parameters. In the long term, research should shift towards heterogeneous model deployment and dynamic routing techniques, coordinating model branches of varying complexity based on the urgency of real-time tasks and hardware status to maximize energy efficiency with limited power and cost.

6.3. Synergy Between Embodied AI and Foundation Models

The introduction of embodied intelligence marks a leap from “passive observation” to “active execution” in agricultural vision. This transition perfectly aligns with the developmental trend of modern agricultural robots [158]. Recent perception advances have demonstrated reliability under real farm conditions, with Wang et al. (2026) addressing long-term re-identification in dairy cow tracking [159] and Shi et al. (2024) tackling subtle behavioral recognition in group-housed ewes [160]. These successes confirm that perception modules are approaching practical robustness. First, the parameters of existing basic models are too large, which contradicts the requirements of edge deployment. Second, low-cost sorting systems still heavily rely on specialized models and lack scene generalization capabilities. More importantly, as Tian et al. (2026) pointed out, agricultural scenes have extreme dynamic characteristics (e.g., unstructured, highly occluded, and muddy conditions), leading to catastrophic failures of execution strategies learned in simulated environments in real farmland. The industry currently lacks agricultural embodied datasets with high-quality physical interaction annotations [161].
Future research focuses on building a high-fidelity, physically interactive agricultural simulation platform that deeply integrates the realism of visual rendering with the physical realism of plant biomechanics. Simultaneously, it is essential to move away from the one-way data acquisition model and transform every physical interaction of the embodied robot in the real environment into high-value self-supervised data. The combination of the basic model and the embodied robot should not be a simple patchwork of algorithms and hardware, but rather a continuous learning loop of “perception-decision-execution-trial and error feedback re-evolution.” Only through closed-loop adversarial learning in the real world and lightweight feedback can a general-purpose agricultural robot capable of flexibly handling different crops and complex agricultural operations be ultimately developed. To provide a clear and holistic overview of the aforementioned discussions, Table 5 summarizes the major challenges, current solutions, remaining limitations, and future research directions across the key dimensions of agricultural visual perception.

7. Conclusions

This paper systematically reviews the latest advancements in deep learning for agricultural visual perception, breaking through traditional biological species-based classification systems and constructing a novel four-layer framework encompassing low-level augmentation, static spatial cognition, high-level spatiotemporal reasoning, and cross-domain deployment. The challenges of non-rigid deformation, dense occlusion, and severe environmental degradation can be mathematically unified at the computational level. The unification is achieved through three techniques: keypoint localization, density regression, and sequential time modeling.
Currently, bottlenecks in agricultural visual perception remain concentrated in domain transformation, dataset scarcity, and the economic constraints of edge deployment. To overcome these obstacles, future research must explore multiple domains. At the architectural level, Vision Transformer and Vision Mamba provide long-range dependency modeling capabilities and achieve a good balance between accuracy and efficiency on edge and drone platforms. Self-supervised vision foundation models such as DINOv2/DINOv3 have demonstrated strong transferability to agricultural tasks including disease classification, root phenotyping, and weed identification with limited labels. The Segment Anything Model (SAM) and SAM 2 have further enabled zero-shot segmentation across diverse scenarios, from grape cluster detection to cropland parcel extraction. However, their billion-parameter scale limits edge deployment, necessitating future work on model compression and domain-specific adaptation. At the data level, low-label learning paradigms, namely semi-supervised, weakly supervised, and self-supervised methods, offer solutions to alleviate label scarcity, while diffusion models, as powerful generative tools, are beginning to play a role in agricultural data augmentation and image inpainting. At the deployment level, federated learning enables privacy-preserving distributed model training across farms without sharing raw data; digital twin technology, by generating high-fidelity synthetic environments, validates algorithms before field deployment, thus bridging the gap between simulation and reality. At the trust and execution level, interpretable AI must evolve from post hoc visualization to quantitative reliability assessment to enhance farmer trust; while open-set recognition improves safety by enabling models to reject unknown categories rather than misclassifying emerging diseases. The close integration of vision-based autonomous navigation and motion planning remains crucial for field robots to achieve a closed loop of perception and action.

Author Contributions

Conceptualization, C.C.; methodology, C.C. and R.L.; investigation, C.C. and R.L.; writing—original draft preparation, R.L.; writing—review and editing, C.C. and L.X.; visualization, R.L.; supervision, C.C. and L.X. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China (Grant No. 32102598).

Institutional Review Board Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
CNNsConvolutional Neural Networks
ViTsVision Transformers
GNNsGraph Neural Networks
ViMVision Mamba
GANsGenerative Adversarial Networks
ST-GCNSpatiotemporal Graph Convolutional Network
FGVCFine-Grained Visual Classification
UAVUnmanned Aerial Vehicle
RGB-DRed-Green-Blue Depth
CNN-LSTMConvolutional Neural Network-Long Short-Term Memory Network
SOLOv2Segmenting Objects by Locations version 2
HSIHyperspectral Imaging
PLSRPartial Least Squares Regression
VGG16Visual Geometry Group 16-layer
CNN-GRUConvolutional Neural Network-Gated Recurrent Unit
ResNet50Residual Network 50
SIRIStructured-Illumination Reflectance Imaging
YOLOv11-SSConvYou Only Look Once version 11 with Spatial and Channel Reconstruction Convolution
MSS-YOLOMulti-Scale Edge-Enhanced Lightweight Network
DCPDark Channel Prior
PED-YOLOPConv, EfficientNetV2, and ADown-based YOLO
LLNetLow-Light Neural Network
CBAMConvolutional Block Attention Module
Zero-DCEZero-Reference Deep Curve Estimation
LAILeaf Area Index
GPUGraphics Processing Unit
UNIR-NetUnderwater Non-uniform Illumination Restoration Network
RSFNetRGB-Sonar Fusion Network
PSNRPeak Signal-to-Noise Ratio
MDCVggNet16Multi-scale Dilated Convolutional Visual Geometry Group Network 16-layer
CycleGANCycle-Consistent Generative Adversarial Network
MoCoProtoMomentum Contrast Prototypical Network
VLMsVision-Language Models
LLMsLarge Language Models
SVMSupport Vector Machines
DETRDEtection TRansformer
T-LEAPTemporal LEAP
DCTM-UniformerV2Depth Channel Temporal Module-UniformerV2
B-splineBasis Spline
YOLOv12mYOLOv12 Medium
MOTMultiple Object Tracking
MCNNsMulti-Column Convolutional Networks
MFNetMulti-Scale Feature Enhancement Network
DSAMDeformable Spatial Attention Mechanism
OBBOriented Bounding Box
GLCMGray-Level Co-occurrence Matrix
LBPLocal Binary Pattern
HSVHue, Saturation, Value
CLIPContrastive Language-Image Pre-training
MASM-YOLOMulti-scale Adaptive Screening and Matching YOLO
DRLDeep Reinforcement Learning
DyFasterNetDynamic Faster Network
Inner-IoUInner Intersection over Union
Wise-IoUWise Intersection over Union
ED-SwinEncoder–Decoder Swin
1D-CNNOne-Dimensional Convolutional Neural Network
VideoMAE V2Video Masked Autoencoder V2
POMDPsPartially Observable Markov Decision Processes
RAFTRecurrent All-Pairs Field Transforms
HF2-VADHybrid Framework integrating Flow reconstruction and Frame prediction for Video Anomaly Detection
PigVADNetPig Video Anomaly Detection Network
IOT-basedInternet of Things-based
ST-LSTMSpatiotemporal Long Short-Term Memory
MIMMemory-In-Memory
ARIMAAutoRegressive Integrated Moving Average
DINOv2/DINOv3Distillation with No Labels v2/Distillation with No Labels v3
SAMSegment Anything Model

References

  1. Taha, M.F.; Mao, H.; Zhang, Z.; Elmasry, G.; Awad, M.A.; Abdalla, A.; Mousa, S.; Elwakeel, A.E.; Elsherbiny, O. Emerging Technologies for Precision Crop Management Towards Agriculture 5.0: A Comprehensive Overview. Agriculture 2025, 15, 582. [Google Scholar] [CrossRef] [Scilit]
  2. Pan, Y.; Zhang, Y.; Wang, X.; Gao, X.X.; Hou, Z. Low-Cost Livestock Sorting Information Management System Based on Deep Learning. Artif. Intell. Agric. 2023, 9, 110–126. [Google Scholar] [CrossRef] [Scilit]
  3. Chen, C.; Zhu, W.; Liu, D.; Steibel, J.; Siegford, J.; Wurtz, K.; Han, J.; Norton, T. Detection of Aggressive Behaviours in Pigs Using a RealSence Depth Sensor. Comput. Electron. Agric. 2019, 166, 105003. [Google Scholar] [CrossRef] [Scilit]
  4. Ahmed, F.; Li, D.; Zhao, B.; Wang, Z.; Huang, J.; Li, T.; Huang, J.; Hou, J.; Jobaer, S.; Yan, H. Pepper-4D: Spatiotemporal 3D Pepper Crop Dataset for Phenotyping. Plants 2026, 15, 599. [Google Scholar] [CrossRef] [Scilit]
  5. Cai, D.; Zhu, C.; Wang, X.; Manceau, L.; Jezequel, S.; Marguerie, M.; De Solan, B.; Baret, F.; Buis, S.; Liu, S.; et al. Predictions of Wheat Phenotypic Variability by Integrating High-Throughput Phenotyping Observations into a Crop Growth Model. Plant Phenomics 2026, 8, 100149. [Google Scholar] [CrossRef] [Scilit]
  6. Chen, C.; Zhu, W.; Norton, T. Behaviour Recognition of Pigs and Cattle: Journey from Computer Vision to Deep Learning. Comput. Electron. Agric. 2021, 187, 106255. [Google Scholar] [CrossRef] [Scilit]
  7. Chen, C.; Zhu, W.; Oczak, M.; Maschat, K.; Baumgartner, J.; Larsen, M.L.V.; Norton, T. A Computer Vision Approach for Recognition of the Engagement of Pigs with Different Enrichment Objects. Comput. Electron. Agric. 2020, 175, 105580. [Google Scholar] [CrossRef] [Scilit]
  8. Xu, L.; Shi, X.; Wang, Y.; Wu, Z.; Zhao, Y.; He, Y. Edge-RegNet: An Adaptive Edge-Aware Fusion and Task Alignment Framework for Detecting Papilionidae Larvae on Citrus. Comput. Electron. Agric. 2026, 246, 111649. [Google Scholar] [CrossRef] [Scilit]
  9. Wu, N.; Wu, J.; Zhang, C.; Zhao, Y.; Xu, X.; Yang, J.; Gao, P.; Xiao, Q.; He, Y. Feature Extraction Combined with Data Generation Powers the Grading of Cotton Verticillium Wilt under Small-Sample Condition. Smart Agric. Technol. 2026, 14, 102122. [Google Scholar] [CrossRef] [Scilit]
  10. Luqiang, Z.; Jianming, K.; Yulong, C.; Bin, H.; Yingkai, C.; Jie, H.; Qiangji, P. Real-Time Recognition and Dynamic Positioning Method for Cotton Terminal Buds Based on CottonBud-YOLOv5s Algorithm and RGBD Camera. Smart Agric. Technol. 2025, 11, 100975. [Google Scholar] [CrossRef] [Scilit]
  11. Mei, Y.; Chen, Y.; Liu, Y.; Yu, H.; Yang, L.; Li, D. MFSD-YOLO: A Multi-Scenario Fish Small Target Detection Method in Aquaculture. Aquac. Eng. 2026, 113, 102677. [Google Scholar] [CrossRef] [Scilit]
  12. Xiong, Y.; McCarthy, C.; Humpal, J.; Percy, C. Pre-Visual Soilborne Common Root Rot Disease Detection in Wheat Using UAV Multispectral Imagery and Deep Neural Networks. Biosyst. Eng. 2026, 264, 104403. [Google Scholar] [CrossRef] [Scilit]
  13. Khaki, S.; Safaei, N.; Pham, H.; Wang, L. WheatNet: A Lightweight Convolutional Neural Network for High-Throughput Image-Based Wheat Head Detection and Counting. Neurocomputing 2022, 489, 78–89. [Google Scholar] [CrossRef] [Scilit]
  14. Chauhdary, J.N.; Li, H.; Jiang, Y.; Pan, X.; Hussain, Z.; Javaid, M.; Rizwan, M. Advances in Sprinkler Irrigation: A Review in the Context of Precision Irrigation for Crop Production. Agronomy 2024, 14, 47. [Google Scholar] [CrossRef] [Scilit]
  15. Liu, Y.; Yan, H.; Zhang, C.; Zhang, J.; Wang, G.; Zhang, D.; Bao, R.; Han, Y. Application of Irrigation Decision-Making Methods for Smart Irrigation Systems: A Review. Irrig. Drain. 2026, 75, 1250–1264. [Google Scholar] [CrossRef] [Scilit]
  16. Yu, Z.; Guo, Y.; Zhang, L.; Ding, Y.; Zhang, G.; Zhang, D. Improved Lightweight Zero-Reference Deep Curve Estimation Low-Light Enhancement Algorithm for Night-Time Cow Detection. Agriculture 2024, 14, 1003. [Google Scholar] [CrossRef] [Scilit]
  17. Zhao, C.; Lei, Y.; Wang, J.; Li, Z.; Li, J.; Wang, Y.; Bai, H. DynaFogBerryNet: A Unified Model for Image Dehazing and Object Detection toward Accurate Strawberry Maturity Monitoring. Smart Agric. Technol. 2025, 12, 101335. [Google Scholar] [CrossRef] [Scilit]
  18. Kaplan, N.H.; Kucuk, S.; Severoglu, N.; Demir, Y. VIRTUE: Color Correction Guided Virtual Exposure Based Underwater Image Enhancement. J. Vis. Commun. Image Represent. 2026, 119, 104870. [Google Scholar] [CrossRef] [Scilit]
  19. Liang, X.; Liang, Z.; Li, L.; Chen, J. AODs-CLYOLO: An Object Detection Method Integrating Fog Removal and Detection in Haze Environments. Appl. Sci. 2024, 14, 7357. [Google Scholar] [CrossRef] [Scilit]
  20. Gadiraju, K.K.; Vatsavai, R.R. Remote Sensing Based Crop Type Classification Via Deep Transfer Learning. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2023, 16, 4699–4712. [Google Scholar] [CrossRef] [Scilit]
  21. Zhao, J.; Li, H.; Chen, C.; Pang, Y.; Zhu, X. Detection of Water Content in Lettuce Canopies Based on Hyperspectral Imaging Technology under Outdoor Conditions. Agriculture 2022, 12, 1796. [Google Scholar] [CrossRef] [Scilit]
  22. Han, Y.; Mao, K.; Wang, X.; Li, X.; Gao, G.; Zhou, F.; Shi, Y. YOLO-MSS: A Lightweight Multi-Scale Detection Model for Dense Multi-Species Fusarium Spores. Smart Agric. Technol. 2026, 14, 102294. [Google Scholar] [CrossRef] [Scilit]
  23. Gao, Q.; Qian, B.; Yu, F.; Chen, L.; Gao, P.; Wu, J.; Li, Z.; Wang, W.; Xie, C.V.J. Improved Adaptive FPGA Dark Channel Prior Dehazing Algorithm for Edge Applications in Agricultural Scenarios. Smart Agric. Technol. 2025, 12, 101285. [Google Scholar] [CrossRef] [Scilit]
  24. Pérez-Zarate, E.; Liu, C.; Ramos-Soto, O.; Oliva, D.; Pérez-Cisneros, M. UNIR-Net: A Novel Approach for Restoring Underwater Images with Non-Uniform Illumination Using Synthetic Data. Image Vis. Comput. 2025, 163, 105734. [Google Scholar] [CrossRef] [Scilit]
  25. Hu, Y.; Dai, X.; Dai, B.; Li, R.; Fang, J.; Yin, Y.; Liu, H.; Shen, W. Feeding Behavior Recognition of Group-Housed Pigs Based on Pose Estimation and Keypoint Features Discrimination. Comput. Electron. Agric. 2025, 239, 111039. [Google Scholar] [CrossRef] [Scilit]
  26. Lu, P.; Zheng, W.; Lv, X.; Xu, J.; Zhang, S.; Li, Y.; Zhangzhong, L. An Extended Method Based on the Geometric Position of Salient Image Features: Solving the Dataset Imbalance Problem in Greenhouse Tomato Growing Scenarios. Agriculture 2024, 14, 1893. [Google Scholar] [CrossRef] [Scilit]
  27. Qian, Y.; Qin, Y.; Wei, H.; Lu, Y.; Huang, Y.; Liu, P.; Fan, Y. MFNet: Multi-Scale Feature Enhancement Networks for Wheat Head Detection and Counting in Complex Scene. Comput. Electron. Agric. 2024, 225, 109342. [Google Scholar] [CrossRef] [Scilit]
  28. Ji, S.; Li, D.; Wang, C.; Lyu, T.; Wang, B.; Xiao, M. Toward Fish Counting in Large-Scale Aquaculture: A Novel Weakly Supervised Density Estimation Framework. Comput. Electron. Agric. 2026, 252, 111872. [Google Scholar] [CrossRef] [Scilit]
  29. Xu, X.; Geng, Q.; Gao, F.; Xiong, D.; Qiao, H.; Ma, X. Segmentation and Counting of Wheat Spike Grains Based on Deep Learning and Textural Feature. Plant Methods 2023, 19, 77. [Google Scholar] [CrossRef] [Scilit]
  30. Pu, L.; Zhao, Y.; Hua, Z.; Han, M.; Song, H. Multi-Target Spraying Behavior Detection Based on an Improved YOLOv8n and ST-GCN Model with Interactive of Video Scenes. Expert Syst. Appl. 2025, 262, 125668. [Google Scholar] [CrossRef] [Scilit]
  31. Taha, M.F.; Mao, H.; Mousa, S.; Zhou, L.; Wang, Y.; Elmasry, G.; Al-Rejaie, S.; Elwakeel, A.E.; Wei, Y.; Qiu, Z. Deep Learning-Enabled Dynamic Model for Nutrient Status Detection of Aquaponically Grown Plants. Agronomy 2024, 14, 2290. [Google Scholar] [CrossRef] [Scilit]
  32. Musazade, E.; Mrisho, I.I.; Gao, J.; Feng, X. Integration of LLMs and VLMs in Plant Stress Phenotyping: From Trait Recognition to Decision Support. Plant Phenomics 2026, 8, 100161. [Google Scholar] [CrossRef] [Scilit]
  33. Visentin, F.; Cremasco, S.; Sozzi, M.; Signorini, L.; Signorini, M.; Marinello, F.; Muradore, R. A Mixed-Autonomous Robotic Platform for Intra-Row and Inter-Row Weed Removal for Precision Agriculture. Comput. Electron. Agric. 2023, 214, 108270. [Google Scholar] [CrossRef] [Scilit]
  34. Awais, M.; Li, W.; Hussain, S.; Cheema, M.J.M.; Li, W.; Song, R.; Liu, C. Comparative Evaluation of Land Surface Temperature Images from Unmanned Aerial Vehicle and Satellite Observation for Agricultural Areas Using in Situ Data. Agriculture 2022, 12, 184. [Google Scholar] [CrossRef] [Scilit]
  35. Dai, T.; Chen, J.; Liu, H.; Liu, L.; Fan, S.; Bai, X.; Qian, L.; Liu, H.; Ba, Y. Qiongda Research on the Potential of Reconstructing Spectral Indices and Dividing Crop Growth Stages Based on Satellite Remote Sensing for Monitoring Soil Moisture in Farmland. Comput. Electron. Agric. 2025, 239, 110943. [Google Scholar] [CrossRef] [Scilit]
  36. Ranario, E.; Mayanja, I.; Yun, H.; Bailey, B.N.; Mason Earles, J. Thermal Image Segmentation in Weedy Fields via Synthetic RGB-Trained Models and GAN-Based Cross-Modality Alignment. Plant Phenomics 2026, 8, 100214. [Google Scholar] [CrossRef] [Scilit]
  37. Zuo, X.; Chu, J.; Shen, J.; Sun, J. Multi-Granularity Feature Aggregation with Self-Attention and Spatial Reasoning for Fine-Grained Crop Disease Classification. Agriculture 2022, 12, 1499. [Google Scholar] [CrossRef] [Scilit]
  38. Chen, J.; Liu, H.; Zhao, H.; Yang, G.; Ren, W. Follow Your Prompts: Controllable Image Dehazing via Latent Space Manipulation. Pattern Recognit. 2026, 180, 114303. [Google Scholar] [CrossRef] [Scilit]
  39. Xing, C.; Wang, Z.; Teng, D.; Meng, Y.; Shi, F. Low-Light Image Enhancement with Shallow Diffusion Model Based on Collaborative Sampling. Comput. Electron. Agric. 2026, 250, 111903. [Google Scholar] [CrossRef] [Scilit]
  40. Cui, X.; Han, W.; Zhang, H.; Dong, Y.; Ma, W.; Zhai, X.; Zhang, L.; Li, G. Estimating and Mapping the Dynamics of Soil Salinity under Different Crop Types Using Sentinel-2 Satellite Imagery. Geoderma 2023, 440, 116738. [Google Scholar] [CrossRef] [Scilit]
  41. Khodaei, M.S.; Nikoo, M.R.; Boudaghpour, S.; Al-Wardy, M. Assessing Reservoir Water Quality Using UAV-Based Multispectral and Thermal Imagery Integrated with Machine and Deep Learning Approaches. J. Hydrol. 2026, 677, 135892. [Google Scholar] [CrossRef] [Scilit]
  42. Wei, L.; Yang, H.; Niu, Y.; Zhang, Y.; Xu, L.; Chai, X. Wheat Biomass, Yield, and Straw-Grain Ratio Estimation from Multi-Temporal UAV-Based RGB and Multispectral Images. Biosyst. Eng. 2023, 234, 187–205. [Google Scholar] [CrossRef] [Scilit]
  43. Yu, T.; Chen, J.; Cherney, J.H.; Zhang, Z. AMGAN: A Multimodal Generative Adversarial Network for near-Daily Alfalfa Multispectral Image Reconstruction. Comput. Electron. Agric. 2026, 244, 111468. [Google Scholar] [CrossRef] [Scilit]
  44. Tang, S.; Xia, Z.; Gu, J.; Wang, W.; Huang, Z.; Zhang, W. High-Precision Apple Recognition and Localization Method Based on RGB-D and Improved SOLOv2 Instance Segmentation. Front. Sustain. Food Syst. 2024, 8, 1403872. [Google Scholar] [CrossRef] [Scilit]
  45. Kang, Z.; Zhou, B.; Fei, S.; Wang, N. Predicting the greenhouse crop morphological parameters based on RGB-D Computer Vision. Smart Agric. Technol. 2025, 11, 100968. [Google Scholar] [CrossRef] [Scilit]
  46. Niu, Y.; Han, W.; Zhang, H.; Zhang, L.; Chen, H. Estimating Maize Plant Height Using a Crop Surface Model Constructed from UAV RGB Images. Biosyst. Eng. 2024, 241, 56–67. [Google Scholar] [CrossRef] [Scilit]
  47. Babalola, E.-O.; Asad, M.H.; Bais, A. Soil Surface Texture Classification Using RGB Images Acquired Under Uncontrolled Field Conditions. IEEE Access 2023, 11, 67140–67155. [Google Scholar] [CrossRef] [Scilit]
  48. Ji, K.; Chen, Z.; Niu, Z.; Mo, A.; Qin, Z.; DeCamp, K.; Wang, C.; Xu, W.; Young, J.M.; Young, B.G.; et al. Variational Autoencoder Enables Unsupervised Leaf Diagnosis via Hyperspectral Imaging. Comput. Electron. Agric. 2026, 251, 111971. [Google Scholar] [CrossRef] [Scilit]
  49. Ji, X.; Tan, Z.; Wang, J.; Zhang, W.; Gouda, M.; He, Y.; Ye, G.; Li, X. Exploring Ensemble Learning Methods with Labor-Free Line-by-Line Radiometric Calibration for Classifying the Damage Types of Rice Planthopper from Hyperspectral Images. Smart Agric. Technol. 2026, 14, 102242. [Google Scholar] [CrossRef] [Scilit]
  50. Guzmán Q., J.A.; Sanchez-Azofeifa, G.A. Prediction of Leaf Traits of Lianas and Trees via the Integration of Wavelet Spectra in the Visible-near Infrared and Thermal Infrared Domains. Remote Sens. Environ. 2021, 259, 112406. [Google Scholar] [CrossRef] [Scilit]
  51. Lam, O.H.Y.; Gevens, A.J.; Jordan, S.; Heberlein, B.; Hills, W.B.; Özdoğan, M.; Townsend, P.A. Early Season Detection of Early Blight (Alternaria solani) in Potato Crops Using Airborne Hyperspectral Imagery. ISPRS J. Photogramm. Remote Sens. 2026, 238, 542–557. [Google Scholar] [CrossRef] [Scilit]
  52. Chen, C.; Zhu, W.; Steibel, J.; Siegford, J.; Wurtz, K.; Han, J.; Norton, T. Recognition of Aggressive Episodes of Pigs Based on Convolutional Neural Network and Long Short-Term Memory. Comput. Electron. Agric. 2020, 169, 105166. [Google Scholar] [CrossRef] [Scilit]
  53. Tu, H.; Huang, D.; Huang, X.; Aheto, J.H.; Ren, Y.; Wang, Y.; Liu, J.; Niu, S.; Xu, M. Detection of browning of fresh-cut potato chips based on machine vision and electronic nose. J. Food Process Eng. 2021, 44, e13631. [Google Scholar] [CrossRef] [Scilit]
  54. Ouyang, Q.; Chang, H.; Li, D.; Xu, Z.; She, Y.; Liu, Z. Intelligent Evaluation of Black Tea Withering Quality via Image-Spectrum Fusion Using a Multimodal Attention Network with Contrastive Learning. Food Chem. 2026, 505, 147771. [Google Scholar] [CrossRef] [Scilit]
  55. Zhao, Y.; Jin, P.; Xiong, G.; Cai, J.; Bai, J.; Ding, C.; Xiao, X. Multimodal Data Fusion and Deep Learning for Predicting Phenolics Dynamics in Barley Bran Solid-State Fermentation. Food Meas. 2026, 20, 8093–8107. [Google Scholar] [CrossRef] [Scilit]
  56. Picardi, G.; Astolfi, A.; Seemakurthy, K.; Hurst, B.; Piana, E.; Bosilj, P.; Calisti, M. Visual servoing of an underwater robotics arm for automatic sorting of crustaceans. Smart Agric. Technol. 2025, 12, 101168. [Google Scholar] [CrossRef] [Scilit]
  57. Lu, A.; Liu, J.; Wei, Y.; Meng, Y.; An, D. FP-CLIP: Foreground-Panorama Prompt Learning for Zero-Shot Anomaly Detection. Digit. Signal Process. 2026, 170, 105798. [Google Scholar] [CrossRef] [Scilit]
  58. De Wit, J.; Tonn, S.; Shao, M.-R.; Van Den Ackerveken, G.; Kalkman, J. Revealing Real-Time 3D in Vivo Pathogen Dynamics in Plants by Label-Free Optical Coherence Tomography. Nat. Commun. 2024, 15, 8353. [Google Scholar] [CrossRef] [Scilit]
  59. Li, K.; Zhu, X.; Qiao, C.; Zhang, L.; Gao, W.; Wang, Y. The Gray Mold Spore Detection of Cucumber Based on Microscopic Image and Deep Learning. Plant Phenomics 2023, 5, 0011. [Google Scholar] [CrossRef] [Scilit]
  60. Luo, W.; Li, Q.; Zhang, H.; Diao, Z.; Guo, Z.; Cai, Z.; Zhang, Y.; Li, J. Detection of Early Decay in Dekopon Fruit Based on Structured-Illumination Reflectance Imaging Combined with a Comparison between Traditional Machine Learning and Deep Learning Models. Postharvest Biol. Technol. 2026, 233, 114044. [Google Scholar] [CrossRef] [Scilit]
  61. Wang, Y.; Mao, H.; Zhang, X.; Liu, Y.; Du, X. A Rapid Detection Method for Tomato Gray Mold Spores in Greenhouse Based on Microfluidic Chip Enrichment and Lens-Less Diffraction Image Processing. Foods 2021, 10, 3011. [Google Scholar] [CrossRef] [Scilit]
  62. Yin, M.; Ling, M.; Chang, K.; Yuan, Z.; Qin, Q.; Chen, B. Joint Image and Feature Enhancement for Object Detection under Adverse Weather Conditions. In Proceedings of the 2024 International Joint Conference on Neural Networks (IJCNN), Yokohama, Japan, 30 June–5 July 2024; IEEE: New York, NY, USA, 2024; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
  63. Li, Z.; Ma, L.; Hou, Y.; Li, J.; Tao, T. An Improved Dehazing Algorithm Combining Dark Channel Prior and Artificial Bee Colony Algorithm. In Proceedings of the 2024 International Applied Computational Electromagnetics Society Symposium (ACES-China), Xi’an, China, 16–19 August 2024; IEEE: New York, NY, USA, 2024; pp. 1–3. [Google Scholar] [CrossRef] [Scilit]
  64. Li, Z.; Wu, J.; Lin, S.; Wang, Z.; Jin, X.; Geng, G.; Huang, F.; Weng, J. Progressive Enhancement Dehazing for Object Detection in Extreme Weather. Eng. Appl. Artif. Intell. 2025, 155, 110903. [Google Scholar] [CrossRef] [Scilit]
  65. Lyu, Z.; Gao, X.; An, Q. From Structure to Semantics: A Fusion Network with Chrominance Guidance and Physics-Inspired for Single Image Dehazing. Signal Process. 2026, 250, 110783. [Google Scholar] [CrossRef] [Scilit]
  66. Yang, Y.; Zhao, J.; Han, L. PhyReNet: Physics-Guided Retinex Network with Illumination-Aware Contrastive Learning for Low-Light Image Enhancement. Pattern Recognit. Lett. 2026, 207, 130–136. [Google Scholar] [CrossRef] [Scilit]
  67. Xu, Y.; Cui, S.; Feng, F.; Wang, H. Dark Light Image Recognition Technology Based on Improved Ssa and Object Detection. Signal Process. Image Commun. 2026, 140, 117427. [Google Scholar] [CrossRef] [Scilit]
  68. Yu, Q.; Wang, J.; Tang, H.; Zhang, J.; Zhang, W.; Liu, L.; Wang, N. Application of Improved UNet and EnglightenGAN for Segmentation and Reconstruction of In Situ Roots. Plant Phenomics 2023, 5, 0066. [Google Scholar] [CrossRef] [Scilit]
  69. Xu, X.; Wu, T.; Ge, W.; Dai, L.; Gong, J.; Chen, D.; Li, S. Enhanced Pothole Detection in Low-Light Conditions Using DCE-Pothole and YOLOv10 Fusion. Results Eng. 2025, 28, 107322. [Google Scholar] [CrossRef] [Scilit]
  70. Wang, J.; Liu, S.; Yu, X.; Tang, Z.; Nan, F.; Yan, P.; Zheng, J.; Tan, G.; Jin, X.; Wang, C. Image Enhancement and Restoration for Maize LAI Estimation under Low-Light Conditions. Smart Agric. Technol. 2026, 14, 102333. [Google Scholar] [CrossRef] [Scilit]
  71. Lin, Y.; Zhou, J.; Ren, W.; Zhang, W. Autonomous Underwater Robot for Underwater Image Enhancement via Multi-Scale Deformable Convolution Network with Attention Mechanism. Comput. Electron. Agric. 2021, 191, 106497. [Google Scholar] [CrossRef] [Scilit]
  72. Wang, Y.; Shang, W. DiffPUIR: A Plug-and-Play Underwater Image Restoration Model with Dual Diffusion Model Prior Constraints. Expert Syst. Appl. 2026, 319, 132172. [Google Scholar] [CrossRef] [Scilit]
  73. Guo, J.; Zhu, Y.; Wu, Y.; Wang, J.; Wang, H.; Lu, T. UPMC-Net: Underwater Physical Model Constraint Network for Underwater Image Enhancement. Digit. Signal Process. 2026, 168, 105673. [Google Scholar] [CrossRef] [Scilit]
  74. Shao, J.; Zhang, H.; Miao, J. Robust-SeaThru: A Physics-Guided Underwater Image Restoration Model with Percentile-Based Backscatter Estimation. Appl. Soft Comput. 2026, 192, 114763. [Google Scholar] [CrossRef] [Scilit]
  75. Lv, P.; Liu, Y.; Jiang, Z.; Chai, S.; Zha, F.; Wang, P.; Guo, W. An RGB-Sonar Fusion Framework for Underwater Depth Estimation Enabled by Cross-Modal Alignment. Opt. Lasers Eng. 2026, 205, 109879. [Google Scholar] [CrossRef] [Scilit]
  76. Gao, X.; Zhao, W.; Li, D.; Liang, Z.; Zhang, W. URDNet: Unsupervised Retinex Decomposition Network for Low-Light Image Enhancement. Inf. Sci. 2026, 755, 123758. [Google Scholar] [CrossRef] [Scilit]
  77. Xing, S.; Lee, H.J. Crop Pests and Diseases Recognition Using DANet with TLDP. Comput. Electron. Agric. 2022, 199, 107144. [Google Scholar] [CrossRef] [Scilit]
  78. Vinay, K.; Surya, V.; Thushar, S.; Singh, T.; Sahay, A. A Deep Learning Framework for Early Detection and Diagnosis of Plant Diseases. Procedia Comput. Sci. 2025, 258, 1435–1445. [Google Scholar] [CrossRef] [Scilit]
  79. Hnida, Y.; Mahraz, M.A.; Achebour, A.; Yahyaouy, A.; Riffi, J.; Tairi, H. OliveTreeCrownsDb: A High-Resolution UAV Dataset for Detection and Segmentation in Agricultural Computer Vision. Data Brief 2025, 60, 111515. [Google Scholar] [CrossRef] [Scilit]
  80. Hussin, A.A.B.; Shapiai, M.I.; Mohamad, S.E.; Iwamoto, K.; Kamaroddin, M.F.; Takemoto, K. A Morphologically Diverse Freshwater Microalgae Dataset for Deep Learning-Based Classification with Transfer Learning Analysis. Ecol. Inform. 2026, 94, 103655. [Google Scholar] [CrossRef] [Scilit]
  81. Qiao, X.; Chen, Y.; Lou, P.; He, B.; Tie, B.; Zhao, Y.; Feng, M.; Xiao, L.; Yang, W.; Song, X.; et al. Research on Precise Identification of Wheat Pests and Diseases Based on an Enhanced VggNet16 Model and Transfer Learning. Smart Agric. Technol. 2026, 14, 102185. [Google Scholar] [CrossRef] [Scilit]
  82. Hukkeri, G.S.; Soundarya, B.C.; Gururaj, H.L.; Ravi, V. Classification of Various Plant Leaf Disease Using Pretrained Convolutional Neural Network on Imagenet. Open Agric. J. 2024, 18, e18743315305194. [Google Scholar] [CrossRef] [Scilit]
  83. Min, X.; Ye, Y.; Xiong, S.; Chen, X. Computer Vision Meets Generative Models in Agriculture: Technological Advances, Challenges and Opportunities. Appl. Sci. 2025, 15, 7663. [Google Scholar] [CrossRef] [Scilit]
  84. Wen, F.; Wu, H.; Zhang, X.; Shuai, Y.; Huang, J.; Li, X.; Huang, J. Accurate Recognition and Segmentation of Northern Corn Leaf Blight in Drone RGB Images: A CycleGAN-Augmented YOLOv5-Mobile-Seg Lightweight Network Approach. Comput. Electron. Agric. 2025, 236, 110433. [Google Scholar] [CrossRef] [Scilit]
  85. Chai, A.Y.H.; Lee, S.H.; Tay, F.S.; Bonnet, P.; Joly, A. Beyond Supervision: Harnessing Self-Supervised Learning in Unseen Plant Disease Recognition. Neurocomputing 2024, 610, 128608. [Google Scholar] [CrossRef] [Scilit]
  86. Jin, D.; Yin, H.; Piao, X.; Gu, Y.H. MoCoProto: Enhancing Few-Shot Pest Image Classification with Self-Supervised Representation Learning. Plant Phenomics 2026, 8, 100242. [Google Scholar] [CrossRef] [Scilit]
  87. Yang, S.; Guo, J.; Yu, J. A Generative AI-Driven Framework Integrating CNN-VLM-LLM for Intelligent Crop Disease Diagnosis and Control Strategy Generation. Comput. Electron. Agric. 2026, 244, 111475. [Google Scholar] [CrossRef] [Scilit]
  88. Khan, Z.; Shen, Y.; Liu, H. ObjectDetection in Agriculture: A Comprehensive Review of Methods, Applications, Challenges, and Future Directions. Agriculture 2025, 15, 1351. [Google Scholar] [CrossRef] [Scilit]
  89. Chen, C.; Zhu, W.; Steibel, J.; Siegford, J.; Han, J.; Norton, T. Recognition of Feeding Behaviour of Pigs and Determination of Feeding Time of Each Pig by a Video-Based Deep Learning Method. Comput. Electron. Agric. 2020, 176, 105642. [Google Scholar] [CrossRef] [Scilit]
  90. Wei, L.; Jianping, H.; Jiaxin, L.; Rencai, Y.; Tengfei, Z.; Mengjiao, Y.; Jing, L. Method for the Navigation Line Recognition of the Ridge without Crops via Machine Vision. Int. J. Agric. Biol. Eng. 2024, 17, 230–239. [Google Scholar] [CrossRef] [Scilit]
  91. Wu, Y.; Chen, L.; Yang, N.; Sun, Z. Research Progress of Deep Learning-Based Artificial Intelligence Technology in Pest and Disease Detection and Control. Agriculture 2025, 15, 2077. [Google Scholar] [CrossRef] [Scilit]
  92. Guo, Y.; Wang, Y.; Wang, X.; Peng, B.; Ran, X.; Mao, R. MS-HA-DETR: A Multi-Scale and Hybrid-Attention Model for Enhancing Pig Abnormal Behavior Detection in Dense Farming. In Proceedings of the 2025 IEEE 23rd International Conference on Industrial Informatics (INDIN), Kunming, China, 12–15 July 2025; IEEE: New York, NY, USA, 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  93. Savvakis, E.; Kapetas, D.; Martínez-Ballesta, M.d.C.; Katsoulas, N.; Pechlivani, E.M. AI-Based Potato Crop Abiotic Stress Detection via Instance Segmentation. AI 2026, 7, 111. [Google Scholar] [CrossRef] [Scilit]
  94. Wang, Y.; Zhang, X.; Ma, G.; Du, X.; Shaheen, N.; Mao, H. Recognition of Weeds at Asparagus Fields Using Multi-Feature Fusion and Backpropagation Neural Network. Int. J. Agric. Biol. Eng. 2021, 14, 190–198. [Google Scholar] [CrossRef] [Scilit]
  95. Peng, Y.; Zhao, S.; Liu, J. Fused-Deep-Features Based Grape Leaf Disease Diagnosis. Agronomy 2021, 11, 2234. [Google Scholar] [CrossRef] [Scilit]
  96. Tao, K.; Wang, A.; Shen, Y.; Lu, Z.; Peng, F.; Wei, X. Peach Flower Density Detection Based on an Improved CNN Incorporating Attention Mechanism and Multi-Scale Feature Fusion. Horticulturae 2022, 8, 904. [Google Scholar] [CrossRef] [Scilit]
  97. Abbas, I.; Liu, J.; Amin, M.; Tariq, A.; Tunio, M.H. Strawberry Fungal Leaf Scorch Disease Identification in Real-Time Strawberry Field Using Deep Learning Architectures. Plants 2021, 10, 2643. [Google Scholar] [CrossRef] [Scilit]
  98. Ji, W.; Zhai, K.; Xu, B.; Wu, J. Green Apple Detection Method Based on Multidimensional Feature Extraction Network Model and Transformer Module. J. Food Prot. 2025, 88, 100397. [Google Scholar] [CrossRef] [Scilit]
  99. Lu, J.; Chen, Z.; Li, X.; Fu, Y.; Xiong, X.; Liu, X.; Wang, H. ORP-Byte: A Multi-Object Tracking Method of Pigs That Combines Oriented RepPoints and Improved Byte. Comput. Electron. Agric. 2024, 219, 108782. [Google Scholar] [CrossRef] [Scilit]
  100. Chen, C.; Zhu, W.; Steibel, J.; Siegford, J.; Han, J.; Norton, T. Classification of Drinking and Drinker-Playing in Pigs by a Video-Based Deep Learning Method. Biosyst. Eng. 2020, 196, 1–14. [Google Scholar] [CrossRef] [Scilit]
  101. Russello, H.; Van Der Tol, R.; Holzhauer, M.; Van Henten, E.J.; Kootstra, G. Video-Based Automatic Lameness Detection of Dairy Cows Using Pose Estimation and Multiple Locomotion Traits. Comput. Electron. Agric. 2024, 223, 109040. [Google Scholar] [CrossRef] [Scilit]
  102. Chen, C.; Liu, R.; Zhu, W.; Norton, T. Aggression Recognition for Individual Pigs Based on YOLOv8 and DCTM-UniformerV2. Biosyst. Eng. 2026, 261, 104345. [Google Scholar]
  103. He, Z.; Yuan, F.; Zhou, Y.; Cui, B.; He, Y.; Liu, Y. Stereo Vision Based Broccoli Recognition and Attitude Estimation Method for Field Harvesting. Artif. Intell. Agric. 2025, 15, 526–536. [Google Scholar] [CrossRef] [Scilit]
  104. Bumbálek, R.; Ufitikirezi, J.D.D.M.; Umurungi, S.N.; Zoubek, T.; Kuneš, R.; Stehlík, R.; Bartoš, P. Computer Vision in Precision Livestock Farming: Benchmarking YOLOv9, YOLOv10, YOLOv11, and YOLOv12 for Individual Cattle Identification. Smart Agric. Technol. 2025, 12, 101208. [Google Scholar] [CrossRef] [Scilit]
  105. Henrich, J.; Post, C.; Zilke, M.; Shiroya, P.; Chanut, E.; Yamchi, A.M.; Yahyapour, R.; Kneib, T.; Traulsen, I. Benchmarking Pig Detection and Tracking under Diverse and Challenging Conditions. Comput. Electron. Agric. 2026, 241, 111264. [Google Scholar] [CrossRef] [Scilit]
  106. Bai, X.; Liu, P.; Cao, Z.; Lu, H.; Xiong, H.; Yang, A.; Cai, Z.; Wang, J.; Yao, J. Rice Plant Counting, Locating, and Sizing Method Based on High-Throughput UAV RGB Images. Plant Phenomics 2023, 5, 0020. [Google Scholar] [CrossRef] [Scilit]
  107. Zhu, H.; Dong, Z.; Wei, L.; Qin, S.; Qin, X.; He, Y. From Canopy Segmentation to Accurate Prediction: An UAV-Based Multi-Feature Fusion Framework for Plot-Scale Ratoon Sugarcane Seedling Counting. Int. J. Appl. Earth Obs. Geoinf. 2026, 147, 105183. [Google Scholar] [CrossRef] [Scilit]
  108. Gonzalez-Trejo, J.A.; Mercado-Ravell, D.A. Dense Crowds Detection and Counting with a Lightweight Architecture. J. Intell. Robot. Syst. 2021, 102, 7. [Google Scholar] [CrossRef] [Scilit]
  109. Sun, J.; Yang, K.; Chen, C.; Shen, J.; Yang, Y.; Wu, X.; Norton, T. Wheat head counting in the wild by an augmented feature pyramid networks-based convolutional neural network. Comput. Electron. Agric. 2022, 193, 106705. [Google Scholar] [CrossRef] [Scilit]
  110. Xu, W.; Lang, P.; Li, J.; Li, D. Multi-Object Fish Behavior Analysis in Recirculating Aquaculture Systems Using Oriented Bounding Boxes and Model Pruning. Comput. Electron. Agric. 2026, 251, 112000. [Google Scholar] [CrossRef] [Scilit]
  111. Wu, W.; Zhong, X.; Lei, C.; Zhao, Y.; Liu, T.; Sun, C.; Guo, W.; Sun, T.; Liu, S. Sampling Survey Method of Wheat Ear Number Based on UAV Images and Density Map Regression Algorithm. Remote Sens. 2023, 15, 1280. [Google Scholar] [CrossRef] [Scilit]
  112. Yang, B.; Pan, M.; Gao, Z.; Zhi, H.; Zhang, X. Cross-Platform Wheat Ear Counting Model Using Deep Learning for UAV and Ground Systems. Agronomy 2023, 13, 1792. [Google Scholar] [CrossRef] [Scilit]
  113. Zhang, D.; Huang, S.; Sun, X.; Zou, X.; Battino, M.; Katona, J.; Shen, L. Advances in 2D and 3D Machine Vision Technologies for Morphological Characterization of Granular Food Products: From Laboratory to Application. Food Meas. 2025, 19, 9292–9318. [Google Scholar] [CrossRef] [Scilit]
  114. Guo, J.; Zhang, K.; Adade, S.Y.S.; Lin, J.; Lin, H.; Chen, Q. Tea Grading, Blending, and Matching Based on Computer Vision and Deep Learning. J. Sci. Food Agric. 2025, 105, 3239–3251. [Google Scholar] [CrossRef] [Scilit]
  115. Guo, Z.; Xiao, H.; Dai, Z.; Wang, C.; Sun, C.; Watson, N.; Povey, M.; Zou, X. Identification of Apple Variety Using Machine Vision and Deep Learning with Multi-Head Attention Mechanism and GLCM. Food Meas. 2025, 19, 6540–6558. [Google Scholar] [CrossRef] [Scilit]
  116. Ray, K.K.; Kumari, A.; Kumar, S.; Machavaram, R.; Shekh, I.; Deshmukh, S.M.; Tadge, P. Guava Leaf Disease Detection Using Support Vector Machine (SVM). Smart Agric. Technol. 2025, 12, 101190. [Google Scholar] [CrossRef] [Scilit]
  117. Quan, J.; Wang, C.; Tian, Y. CLIP-AFIR: A Contrastive Language-Image Pretraining Model for Accurate Fish Individual Fine-Grained Re-Identification. Aquaculture 2026, 610, 742885. [Google Scholar] [CrossRef] [Scilit]
  118. Wei, P.; Sun, W.; Cao, S.; Kong, F. Lightweight Model for Beef Cattle Behavior Recognition from Quadruped Robot Video in Grassland Pastures. Comput. Electron. Agric. 2026, 242, 111329. [Google Scholar] [CrossRef] [Scilit]
  119. Fu, C.; Wang, A.; Tian, A. DCS-YOLO: A Blueberry Maturity Detection Model in Complex Environments. Comput. Electron. Agric. 2026, 248, 111830. [Google Scholar] [CrossRef] [Scilit]
  120. Yuan, M.; Wang, Z.; Liu, Z.; Fang, C.; Yang, S.; Hu, X.; Ning, J. DSLNDD-Net: A Multi-Scale Edge-Aware Attention Network for Nutrient Deficiency Detection in Dense Strawberry Leaves. Smart Agric. Technol. 2026, 14, 102102. [Google Scholar] [CrossRef] [Scilit]
  121. Wu, H.; Wang, X.; Chen, X.; Zhang, Y.; Zhang, Y. Review on Key Technologies for Autonomous Navigation in Field Agricultural Machinery. Agriculture 2025, 15, 1297. [Google Scholar] [CrossRef] [Scilit]
  122. Liu, B.; Huang, X.; Sun, L.; Wei, X.; Ji, Z.; Zhang, H. MCDCNet: Multi-Scale Constrained Deformable Convolution Network for Apple Leaf Disease Detection. Comput. Electron. Agric. 2024, 222, 109028. [Google Scholar] [CrossRef] [Scilit]
  123. Li, D.; Yang, Y.; Chang, H.; Xu, Z.; Ouyang, Q. Improved ResNet Deep Learning Model-Assisted Computer Vision for Intelligent Assessment of Tencha Chlorophyll Content during Drying. J. Food Compos. Anal. 2025, 148, 108545. [Google Scholar] [CrossRef] [Scilit]
  124. Li, A.; Wang, C.; Ji, T.; Wang, Q.; Zhang, T. D3-YOLOv10: Improved YOLOv10-Based Lightweight Tomato Detection Algorithm Under Facility Scenario. Agriculture 2024, 14, 2268. [Google Scholar] [CrossRef] [Scilit]
  125. Huang, Z.-H.; Chen, C.-T.; Ikegaya, N.; Chang, T.; Ke, K.; Chen, Y.-C. A Robotic Harvesting System for Occluded Cucumbers Using F2SA-YOLOv8 and HVSC. Comput. Electron. Agric. 2026, 246, 111616. [Google Scholar] [CrossRef] [Scilit]
  126. Zhang, J.; Zhou, H.; Liu, K.; Xu, Y. ED-Swin Transformer: A Cassava Disease Classification Model Integrated with UAV Images. Sensors 2025, 25, 2432. [Google Scholar] [CrossRef] [Scilit]
  127. Liu, J.; Han, X.; Zhou, Q.; Zhang, B.; Yang, J.; Gao, Q. Integrating Attention Mechanism into CNN-LSTM for Spring Maize Crop Coefficient Prediction. Smart Agric. Technol. 2026, 14, 102311. [Google Scholar] [CrossRef] [Scilit]
  128. Bahrami, N. Deep Learning Enhanced CNN-LSTM Framework for Quantitative Assessment of Coastal Changes Induced by Sea Level Variations Using Remote Sensing. Estuar. Coast. Shelf Sci. 2026, 338, 109912. [Google Scholar] [CrossRef] [Scilit]
  129. Vysyaraju, U.S.R.; Moon, S.; Mendes, E.D.M.; Adekunle, A.J.; Kaniyamattam, K.; Pi, Y. 47. Activity Recognition in Beef Cattle Calan Gate Using CLIP and YOLO V11. Anim. Sci. Proc. 2025, 16, 578–579. [Google Scholar] [CrossRef] [Scilit]
  130. Zheng, K.; Yang, S.; Wang, Z.; Fu, H.; Wang, X.; Zou, W.; Zhai, C.; Chen, L. Real-Time Detection and Validation of a Target-Oriented Model for Spindle-Shaped Tree Trunks Leveraging Deep Learning. Agronomy 2026, 16, 210. [Google Scholar] [CrossRef] [Scilit]
  131. Sumana, S.L.; Jing, X.; Hu, H.; Zhang, C.; Abdullateef, M.M.; Mani, E.K.F.; Wu, X.; Shuaibu, A.; Su, S. Embodied AI for Aquaculture Monitoring: From Stationary Sensing to Mobile Robotic Observation. Aquac. Eng. 2026, 115, 102766. [Google Scholar] [CrossRef] [Scilit]
  132. Yang, G.; He, Y.; Ye, L.; Luo, Y.; Kim, H.; Feng, X. Improving Wheat Yield Prediction in Breeding under Climate Change via a New Method of Correcting and Reconstructing Meteorological Data and Weighing Multiple Crop Model Outputs. Comput. Electron. Agric. 2026, 248, 111734. [Google Scholar] [CrossRef] [Scilit]
  133. Fiqar, T.P.; Wang, F.; Shimasaki, K.; Ishii, I.; Sugino, T. Cattle Trembling Detection Using HFR-Video-Based DIC Analysis. IEEE Sens. Lett. 2025, 9, 14. [Google Scholar] [CrossRef] [Scilit]
  134. Alfarano, A.; Maiano, L.; Papa, L.; Amerini, I. Estimating Optical Flow: A Comprehensive Review of the State of the Art. Comput. Vis. Image Underst. 2024, 249, 104160. [Google Scholar] [CrossRef] [Scilit]
  135. Wu, T.; Wang, R.; Song, H.; Zhang, Z.; Li, Q.; Wang, Z.; Shi, Y.; Gao, R. DA-MOT: A Directed-Angle-Aware Multi-Object Tracking Framework for Dairy Cows. Comput. Electron. Agric. 2026, 254, 112180. [Google Scholar] [CrossRef] [Scilit]
  136. Wang, H.; Sun, B.; Zeng, Y.; Yang, X.; Liang, C.; Ma, J.; Qi, R.; Wang, C. Spatial Prior-Enhanced Temporal Action Localization: A Novel Three-Stage Framework for Pig Behavior Recognition. Smart Agric. Technol. 2026, 14, 102010. [Google Scholar] [CrossRef] [Scilit]
  137. Lakhiar, I.A.; Yan, H.; Zhang, C.; Wang, G.; He, B.; Hao, B.; Han, Y.; Wang, B.; Bao, R.; Syed, T.N.; et al. A Review of Precision Irrigation Water-Saving Technology under Changing Climate for Enhancing Water Use Efficiency, Crop Yield, and Environmental Footprints. Agriculture 2024, 14, 1141. [Google Scholar] [CrossRef] [Scilit]
  138. Wang, L.; Gao, J.; Qureshi, W.A. Evolution and Application of Precision Fertilizer: A Review. Agronomy 2025, 15, 1939. [Google Scholar] [CrossRef] [Scilit]
  139. Yuan, L.; Li, W.; Yue, J.; Li, Z.; Zhao, Y.; Li, Q.; Feng, W.; Kou, G. A Review of Cattle Behavior Recognition Methods Based on Computer Vision and Deep Learning. Smart Agric. Technol. 2026, 14, 101867. [Google Scholar] [CrossRef] [Scilit]
  140. Bah, M.D.; Hafiane, A.; Canals, R. Hierarchical Graph Representation for Unsupervised Crop Row Detection in Images. Expert Syst. Appl. 2023, 216, 119478. [Google Scholar] [CrossRef] [Scilit]
  141. Ou, J.; Chen, F.; Zhang, M.; Batchelor, W.D.; Wang, B.; Wu, D.; Ma, X.; Zhang, Z.; Hu, K.; Feng, P. A Continuous Learning Framework Based on Physics-Guided Deep Learning for Crop Phenology Simulation. Agric. For. Meteorol. 2025, 368, 110562. [Google Scholar] [CrossRef] [Scilit]
  142. Zhao, J.; Fan, S.; Zhang, B.; Wang, A.; Zhang, L.; Zhu, Q. Research Status and Development Trends of Deep Reinforcement Learning in the Intelligent Transformation of Agricultural Machinery. Agriculture 2025, 15, 1223. [Google Scholar] [CrossRef] [Scilit]
  143. Li, K.; Shi, J.; Hu, C.; Xue, W. The Intelligentization Process of Agricultural Greenhouse: A Review of Control Strategies and Modeling Techniques. Agriculture 2025, 15, 2135. [Google Scholar] [CrossRef] [Scilit]
  144. Xu, S.; Zhao, H.; Wang, L. CNN-DQN Fusion for Precision Agriculture: A Reinforcement Learning Approach to Yield Prediction and Resource Optimization. In Proceedings of the 2025 2nd International Conference on Intelligent Computing and Robotics (ICICR), Dalian, China, 16–18 May 2025; IEEE: New York, NY, USA, 2025; pp. 774–781. [Google Scholar] [CrossRef] [Scilit]
  145. Rajagukguk, R.A.; Lee, S.; Park, J.; Daniel, K.F.; Lee, C.; Chen, Z.; Liu, D.; Norton, T.; Park, J.; Hong, S. Deep Learning for Visual Animal Monitoring (Detection, Tracking, Pose Estimation, and Behavior Classification): A Comprehensive Review. Smart Agric. Technol. 2025, 12, 101539. [Google Scholar] [CrossRef] [Scilit]
  146. Al-Najadi, R.; Al-Mulla, Y.; Goher, K. Advances in Intelligent and Autonomous Greenhouse Systems: A Comprehensive Review of Internet of Things, Artificial Intelligence, and Robotics Integration. Smart Agric. Technol. 2026, 13, 101670. [Google Scholar] [CrossRef] [Scilit]
  147. Feng, L.; Zong, S. Large-Scale Crop Image Recognition Based on ConvLSTM Spatiotemporal Feature Extraction and LCMST-Net Model. IEEE Access 2026, 14, 10719–10734. [Google Scholar] [CrossRef] [Scilit]
  148. Zhang, M.; Yang, B.; Hu, X.; Gong, J.; Zhang, Z. Foundation model for generalist remote sensing intelligence: Potentials and prospects. Sci. Bull. 2024, 69, 3652–3656. [Google Scholar] [CrossRef] [Scilit]
  149. Han, B.; Zhu, C.; Han, D.; Yu, R.; Cao, S.; Wu, J.; Chapman, S.; Zheng, B.; Guo, W.; Weiss, M.; et al. FoMo4Wheat: Toward Reliable Crop Vision Foundation Models with Globally Curated Data. arXiv 2025, arXiv:2509.06907. [Google Scholar]
  150. Yan, K.; Dai, B.; Liu, H.; Yin, Y.; Li, X.; Wu, R.; Shen, W. Deep Neural Network with Adaptive Dual-Modality Fusion for Temporal Aggressive Behavior Detection of Group-Housed Pigs. Comput. Electron. Agric. 2024, 224, 109243. [Google Scholar] [CrossRef] [Scilit]
  151. Wang, C.; Pan, W.; Song, X.; Yu, H.; Zhu, J.; Liu, P.; Li, X. Predicting Plant Growth and Development Using Time-Series Images. Agronomy 2022, 12, 2213. [Google Scholar] [CrossRef] [Scilit]
  152. Lwaho, J.; Ilembo, B. The Inclusion of External Variables in Forecasting Maize Production: A Dynamic Regression Model with ARIMA Errors. J. Saudi Soc. Agric. Sci. 2026, 25, 54. [Google Scholar] [CrossRef] [Scilit]
  153. Dong, J.; Jin, W.; Fuentes, A.; Lee, J.; Jeong, Y.; Yoon, S.; Park, D.S. Harnessing Prototype Networks for Novel Plant Species and Disease Classification in Open-World Scenarios. Eng. Appl. Artif. Intell. 2025, 156, 111016. [Google Scholar] [CrossRef] [Scilit]
  154. Nugroho, H.; Xan, C.J.; Siang, T.F.; Eswaran, S. Optimizing Convolutional Neural Networks for Plant Disease Detection on ARM-M Microcontrollers. In Proceedings of the 2024 IEEE International Conference on Agrosystem Engineering, Technology & Applications (AGRETA), Kuala Lumpur, Malaysia, 7 September 2024; IEEE: New York, NY, USA, 2024; pp. 107–110. [Google Scholar] [CrossRef] [Scilit]
  155. Gao, X.; Gao, J.; Qureshi, W.A. Applications, Trends, and Challenges of Precision Weed Control Technologies Based on Deep Learning and Machine Vision. Agronomy 2025, 15, 1954. [Google Scholar] [CrossRef] [Scilit]
  156. Zhang, Z.; Lu, Y.; Peng, Y.; Yang, M.; Hu, Y. A Lightweight and High-Performance YOLOv5-Based Model for Tea Shoot Detection in Field Conditions. Agronomy 2025, 15, 1122. [Google Scholar] [CrossRef] [Scilit]
  157. Shao, L.; Gong, J.; Fan, W.; Zhang, Z.; Zhang, M. Cost Comparison between Digital Management and Traditional Management of Cotton Fields—Evidence from Cotton Fields in Xinjiang, China. Agriculture 2022, 12, 1105. [Google Scholar] [CrossRef] [Scilit]
  158. Deng, J.; Ni, L.; Bai, X.; Jiang, H.; Xu, L. Simultaneous Analysis of Mildew Degree and Aflatoxin B1 of Wheat by a Multi-Task Deep Learning Strategy Based on Microwave Detection Technology. LWT 2023, 184, 115047. [Google Scholar] [CrossRef] [Scilit]
  159. Wang, M.; Zandona, B.; Räisänen, S.E.; Serviento, A.M.; Niu, M. A Tracking-by-Classification Approach for Continuous Individual Monitoring of Holstein Dairy Cows in a Free-Stall Barn. Biosyst. Eng. 2026, 265, 104420. [Google Scholar] [CrossRef] [Scilit]
  160. Shi, J.; Chen, X.; Zhang, Y.; Gong, P.; Xiong, Y.; Shen, M.; Norton, T.; Gu, X.; Lu, M. Detection of Estrous Ewes’ Tail-Wagging Behavior in Group-Housed Environments Using Temporal-Boost 3D Convolution. Comput. Electron. Agric. 2025, 234, 110283. [Google Scholar] [CrossRef] [Scilit]
  161. Tian, M.; Guo, H.; Long, C.; Zhu, J.; Zhao, W.; Li, H.; Ruchay, A.; Pezzuolo, A.; Li, B. How Large-Scale Foundation Models Benefit Precision Livestock Farming: A Survey. Artif. Intell. Agric. 2026, 16, 1109–1132. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Deep learning-driven closed-loop visual perception in precision agriculture. (ac) Multi-modal imaging classification; (d) Dehazing; (e) Low-light enhancement; (f) Underwater restoration; (g,h) Deformation/Non-rigid; (i,j) Crowding/Severe Occlusion; (k) Low inter-class variance; (l) Short-term behaviors; (m) Long-term evolution; (n) Vision foundation; (o) Agricultural robots/Automated machinery.
Figure 1. Deep learning-driven closed-loop visual perception in precision agriculture. (ac) Multi-modal imaging classification; (d) Dehazing; (e) Low-light enhancement; (f) Underwater restoration; (g,h) Deformation/Non-rigid; (i,j) Crowding/Severe Occlusion; (k) Low inter-class variance; (l) Short-term behaviors; (m) Long-term evolution; (n) Vision foundation; (o) Agricultural robots/Automated machinery.
Agriculture 16 01826 g001
Figure 2. Workflow of multimodal data acquisition and low-level physical enhancement processes for precision agriculture. (a) Satellite remote sensing; (b) UAV multispectral imaging; (c) Microscopic imaging; (d) Dehazing; (e) Low-light Enhancement; (f) Underwater image restoration.
Figure 2. Workflow of multimodal data acquisition and low-level physical enhancement processes for precision agriculture. (a) Satellite remote sensing; (b) UAV multispectral imaging; (c) Microscopic imaging; (d) Dehazing; (e) Low-light Enhancement; (f) Underwater image restoration.
Agriculture 16 01826 g002
Figure 3. Static Spatial Cognitive Paradigms for Agricultural Visual Perception. (a) Rigid constraints; (b) Dense overlap; (c) Visual Ambiguity; (d) Adaptive pose; (e) Dense map; (f) Fine-grained features.
Figure 3. Static Spatial Cognitive Paradigms for Agricultural Visual Perception. (a) Rigid constraints; (b) Dense overlap; (c) Visual Ambiguity; (d) Adaptive pose; (e) Dense map; (f) Fine-grained features.
Agriculture 16 01826 g003
Table 1. Summary of multi-scale and multimodal visual sensing technologies in precision agriculture.
Table 1. Summary of multi-scale and multimodal visual sensing technologies in precision agriculture.
Spatial Scale & ModalityTarget ApplicationsTraditional BottlenecksDeep Learning & Fusion SolutionsRef.
Macro (Satellite)Soil salinization
Yield estimation
Temporal gaps
Cloud occlusion
Multi-temporal spectral fusion[40,43]
Meso (UAV)Fine-grained yield
Water quality
Insufficient resolution from single viewsCNN-LSTM[41,42]
Ground (RGB-D)Canopy height
Livestock poses
Lack of 3D depth 3D point clouds & digital terrain models[44,45,46]
Ground (HSI)Early disease
Grading
Physiological blind spots in RGBSpatial-spectral joint extraction[48,49,50,51]
Micro (Microscopy)Spore typing
Weak decay
Macroscopic invisibilityLightweight YOLO
SIRI integration
[58,59,60,61]
Table 2. Summary of image enhancement literature in extreme agricultural environments.
Table 2. Summary of image enhancement literature in extreme agricultural environments.
EnvironmentVisual DegradationTraditional LimitsDeep Learning ArchitecturesRef.
Rain & FogDetail loss
Contrast attenuation
Edge artifacts
Color distortion
Progressive enhancement
Physics-guided fusion
[63,64,65]
Low Light/NightPhoton starvation
High-frequency noise
Color distortion
Detail loss
Autoencoder joint optimization, Zero-DCE nonlinear iterations[66,67,68,69,70]
UnderwaterSevere backscattering Red-light attenuationHigh physical model dependenceTexture-guided multi-scale enhancement
Sonar–optical fusion
[24,71,72,73,74,75,76]
Table 3. Evolution of static spatial cognitive paradigms and their technical solutions.
Table 3. Evolution of static spatial cognitive paradigms and their technical solutions.
Spatial ChallengeBiological TargetsTraditional BBox LimitsModern Cognitive Paradigms & MechanismsRef.
Non-rigid DeformationCurved crop stems
Interaction with livestock
Background noise
Localization failure
Keypoint-based Pose Estimation[91,92,93,94,95,96,97,98,99,100,101,102,103,104]
Extreme Crowding & Severe OcclusionOverlapping wheat
Congested fish schools
Dense canopies
Missed detections
Structural data loss
Density Map Regression[105,106,107,108,109,110,111,112,113]
Low Inter-class VarianceSimilar crop varieties
Tiny localized lesions
Global features fail to capture subtle discrepanciesFine-Grained Visual Categorization (FGVC)[114,115,116,117]
Table 5. Summary of major challenges, current solutions, limitations, and future directions.
Table 5. Summary of major challenges, current solutions, limitations, and future directions.
DimensionMajor ChallengesCurrent SolutionsRemaining LimitationsFuture Research Directions
Domain GeneralizationDynamic field shifts
Open-set mutations
Domain gaps (UAV, seasons)
Test-time adaptation
Agricultural foundation models
Lab-biased assumptions
Catastrophic forgetting
Poor calibration in real farms
Multi-season adaptation
Physics-embedded zero-shot learning
Inherent open-set recognition
Lightweight DeploymentEdge computing limits
High economic/energy costs
8-bit quantization
Lightweight architectures for edge devices
Accuracy loss
Lack of hardware–software co-design
Heterogeneous model deployment
Dynamic routing for energy–accuracy optimization
Embodied IntelligenceMoving from passive vision to active execution
Unstructured, highly occluded scenes
Multimodal large models
Diffusion models for zero-shot transfer
Excessive parameters
Sim-to-real failures
No interactive datasets.
High-fidelity physical interactive simulations
Closed-loop “perception-decision-execution” learning
Data Scarcity & TrustScarcity of labeled data
Privacy in data sharing
Low farmer trust in AI
ViT/Mamba for spatial reasoning
Semi/self-supervised learning
Privacy risks in centralization
Post-hoc interpretability
Federated learning Quantitative reliability assessment for trust
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Chen, C.; Liu, R.; Xu, L. A Comprehensive Review of Deep Learning in Agricultural Visual Perception: Progress, Bottlenecks, and Emerging Trends. Agriculture 2026, 16, 1826. https://doi.org/10.3390/agriculture16171826

AMA Style

Chen C, Liu R, Xu L. A Comprehensive Review of Deep Learning in Agricultural Visual Perception: Progress, Bottlenecks, and Emerging Trends. Agriculture. 2026; 16(17):1826. https://doi.org/10.3390/agriculture16171826

Chicago/Turabian Style

Chen, Chen, Runlin Liu, and Leijun Xu. 2026. "A Comprehensive Review of Deep Learning in Agricultural Visual Perception: Progress, Bottlenecks, and Emerging Trends" Agriculture 16, no. 17: 1826. https://doi.org/10.3390/agriculture16171826

APA Style

Chen, C., Liu, R., & Xu, L. (2026). A Comprehensive Review of Deep Learning in Agricultural Visual Perception: Progress, Bottlenecks, and Emerging Trends. Agriculture, 16(17), 1826. https://doi.org/10.3390/agriculture16171826

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop